Datasets
High-quality open-source AI training datasets
Hugging Face Datasets Collection
The largest open-source dataset hub on the web, hosted by Hugging Face. It spans NLP, computer vision, audio, and multimodal tasks, offering hundreds of thousands of community-contributed datasets with efficient streaming and loading utilities. Whether you need text classification benchmarks, image-caption pairs, or speech-to-text corpora, this is the go-to infrastructure for sourcing quality training data.
Kaggle Datasets Platform
Kaggle hosts tens of thousands of public datasets spanning machine learning competitions, real-world business problems, and academic research. All datasets are available for online analysis through built-in notebooks and can be downloaded directly for local experimentation and model training.
Papers with Code Datasets
Papers with Code links benchmark datasets from academic papers to their corresponding state-of-the-art models and open-source implementations. This makes it easy to compare methods, reproduce results, and find the right dataset and code baseline for your research.
Google Dataset Search
Google Dataset Search is a dedicated search engine that indexes public datasets across thousands of sources, letting researchers and developers quickly locate the specific data they need for training, analysis, or validation regardless of where it is hosted.
Common Crawl Web Corpus
Common Crawl is an open, petabyte-scale web crawl corpus that serves as a foundational pre-training data source for many large language models. The corpus is freely available for download and widely used in NLP research, data mining, and large-scale machine learning experiments.
LAION Open Multimodal Datasets
LAION (Large-Scale Artificial Intelligence Open Network) provides massive open-source image-text pair datasets that have been widely used to train multimodal models such as CLIP and Stable Diffusion. The datasets are freely available and serve as key infrastructure for open-source generative AI research.
UCI Machine Learning Repository
The UCI Machine Learning Repository is a classic collection of hundreds of structured datasets maintained since the late 1990s, widely used for hands-on practice, algorithm benchmarking, and teaching fundamental machine learning concepts in academic courses worldwide.
The Stack Source Code Dataset
The Stack is a large, permissively licensed source code dataset from the BigCode project, covering hundreds of programming languages. It is widely used to pre-train and evaluate code generation models, and includes opt-out mechanisms and deduplication tooling to support responsible data practices.
ImageNet
ImageNet is a landmark large-scale image dataset with millions of labeled images spanning thousands of categories. It powered the deep learning revolution in computer vision through the ILSVRC benchmark competition and remains foundational for pre-training, transfer learning, and evaluating image classification models.
Alpaca-Cleaned
Alpaca-Cleaned is a language modeling dataset by yahma on Hugging Face.
Alpaca CoT
Alpaca CoT is a Chinese open dataset by QingyiSi on Hugging Face.
Alpaca
Alpaca is a language modeling dataset by tatsu-lab on Hugging Face.
Bad Prompt
Bad Prompt is an open dataset by Nerfgun3 on Hugging Face.
c4
c4 is a multilingual language modeling dataset by allenai on Hugging Face.
Chinese DeepSeek R1 Distill data 110k
Chinese DeepSeek R1 Distill data 110k is a Chinese language modeling dataset by Congliu on Hugging Face.
COIG CQIA
COIG CQIA is a Chinese question answering dataset by m-a-p on Hugging Face.
Cosmopedia
Cosmopedia is an open dataset by HuggingFaceTB on Hugging Face.
CulturaX
CulturaX is a multilingual language modeling dataset by uonlp on Hugging Face.
Databricks Dolly 15k
Databricks Dolly 15k is a question answering dataset by databricks on Hugging Face.
DiffusionDB
DiffusionDB is an image captioning dataset by poloclub on Hugging Face.
Dolma
Dolma is a language modeling dataset by allenai on Hugging Face.
EasyNegative
EasyNegative is an open dataset by gsdf on Hugging Face.
Falcon RefinedWeb
Falcon RefinedWeb is a language modeling dataset by tiiuae on Hugging Face.
📄 FinePDFs
📄 FinePDFs is a language modeling dataset by HuggingFaceFW on Hugging Face.
FineWeb
FineWeb is a language modeling dataset by HuggingFaceFW on Hugging Face.
🥂 FineWeb 2
🥂 FineWeb 2 is a language modeling dataset by HuggingFaceFW on Hugging Face.
FineWeb-Edu
FineWeb-Edu is a language modeling dataset by HuggingFaceFW on Hugging Face.
GAIA
GAIA is an open dataset by gaia-benchmark on Hugging Face.
GLUE (General Language Understanding Evaluation benchmark)
GLUE (General Language Understanding Evaluation benchmark) is a text classification dataset by nyu-mll on Hugging Face.
Grade School Math 8K
Grade School Math 8K is a language modeling dataset by openai on Hugging Face.
Hh Rlhf
Hh Rlhf is an open dataset by Anthropic on Hugging Face.
hle
hle is an open dataset by cais on Hugging Face.
ImageNet
ImageNet is an image classification dataset by ILSVRC on Hugging Face.
IMDB
IMDB is a text classification dataset by stanfordnlp on Hugging Face.
Infinity Instruct
Infinity Instruct is a Chinese language modeling dataset by BAAI on Hugging Face.
Internship Warehouse
Internship Warehouse is an open dataset by FlyRank on Hugging Face.
Llama Nemotron Post Training Dataset
Llama Nemotron Post Training Dataset is an open dataset by nvidia on Hugging Face.
LLaVA Instruct 150K
LLaVA Instruct 150K is a question answering dataset by liuhaotian on Hugging Face.
Lmsys Chat 1m
Lmsys Chat 1m is an open dataset by lmsys on Hugging Face.
medical o1 reasoning SFT
medical o1 reasoning SFT is a Chinese question answering dataset by FreedomIntelligence on Hugging Face.
Medical Prescription English Audio Dataset
Medical Prescription English Audio Dataset is an open dataset by HumynLabs on Hugging Face.
Measuring Massive Multitask Language Understanding
Measuring Massive Multitask Language Understanding is a question answering dataset by cais on Hugging Face.
MNBVC
MNBVC is a Chinese language modeling dataset by liwu on Hugging Face.
NuminaMath CoT
NuminaMath CoT is a language modeling dataset by AI-MO on Hugging Face.
OpenAssistant Conversations
OpenAssistant Conversations is a multilingual open dataset by OpenAssistant on Hugging Face.
OpenHermes 2.5
OpenHermes 2.5 is an open dataset by teknium on Hugging Face.
OpenOrca
OpenOrca is a text classification dataset by Open-Orca on Hugging Face.
OpenR1 Math 220k
OpenR1 Math 220k is an open dataset by open-r1 on Hugging Face.
OpenThoughts 114k
OpenThoughts 114k is an open dataset by open-thoughts on Hugging Face.
Opus 4.6 Reasoning 3000x filtered
Opus 4.6 Reasoning 3000x filtered is an open dataset by nohurry on Hugging Face.
PersonaHub
PersonaHub is a Chinese language modeling dataset by proj-persona on Hugging Face.
PhysicalAI Autonomous Vehicles
PhysicalAI Autonomous Vehicles is an open dataset by nvidia on Hugging Face.
Prompts.chat
Prompts.chat is a question answering dataset by fka on Hugging Face.
RedPajama Data 1T
RedPajama Data 1T is a language modeling dataset by togethercomputer on Hugging Face.
ShareGPT Vicuna unfiltered
ShareGPT Vicuna unfiltered is an open dataset by anon8231489123 on Hugging Face.
SQuAD
SQuAD is a question answering dataset by rajpurkar on Hugging Face.
Synthetic Text To Sql
Synthetic Text To Sql is a question answering dataset by gretelai on Hugging Face.
The-Stack-v2
The-Stack-v2 is a language modeling dataset by bigcode on Hugging Face.
TinyStories
TinyStories is a language modeling dataset by roneneldan on Hugging Face.
UltraChat 200k
UltraChat 200k is a language modeling dataset by HuggingFaceH4 on Hugging Face.
+ 1,766 more — VIEW ALL →