What is this?

This project makes it easy to analyze the Python ecosystem by providing of all the code ever published to PyPI via git, parquet datasets with file metadata, and a set of tools to help analyze the data.

Thanks to the power of git the contents of PyPI takes up only 380.0 GB on disk, and thanks to tools like libcst every Python file can be analysed on a consumer-grade laptop in a few hours.

Download all the codeExplore the datasets


Stats for nerds 🤓

Total files
1.13 Billion
83,587,205 unique
Total lines of text
355.6 Billion
355,597,944,412 to be precise
Total uncompressed size
62.1 TiB
That is ~46,530,159.893 floppy disks
Lines of code added per second
3,472
In the month 2023-09-01
Click here for lots more stats!