The Stack Source Code Dataset
The Stack is a large, permissively licensed source code dataset from the BigCode project, covering hundreds of programming languages. It is widely used to pre-train and evaluate code generation models, and includes opt-out mechanisms and deduplication tooling to support responsible data practices.
Overview
"The Stack Source Code Dataset" is a "Dataset" resource curated by AI Resource Hub, filed under the Datasets category and suited to Advanced-level learners. It is provided by BigCode, was last updated on 2026-06-22, and holds an editorial score of 4.5/5 from our team. Click "Visit Resource" on the right to open the original page.
Our Verdict
The Stack is the default starting corpus for anyone training or evaluating code models: a huge, permissively licensed collection spanning mainstream and niche languages alike, already foundational to the field. Treat it as industrial material, not a download-and-go asset. You must filter licenses before commercial training, audit for residual issues, and budget serious storage and compute for its full volume. For research and pretraining it is close to indispensable; for casual experimentation, smaller curated sets are kinder.
Tags
Key Features
- ▹Large corpus of source code
- ▹Many programming languages
- ▹Used to train code models
Pros
- +Foundational for code LLMs
- +Broad language coverage
- +Foundational training data for code models
Cons
- −License filtering is important
- −License filtering is essential before use
- −Huge volume needs serious storage and compute
FAQ
Details
- Pricing
- Free / open dataset
- Author
- BigCode
- Editorial score
- ★ 4.5 / 5
- Last updated
- Jun 22, 2026