Skip to content
📊

Datasets

High-quality open-source AI training datasets

4.7

Hugging Face Datasets Collection

The largest open-source dataset hub on the web, hosted by Hugging Face. It spans NLP, computer vision, audio, and multimodal tasks, offering hundreds of thousands of community-contributed datasets with efficient streaming and loading utilities. Whether you need text classification benchmarks, image-caption pairs, or speech-to-text corpora, this is the go-to infrastructure for sourcing quality training data.

FreeDatasetIntermediateDatasetsHugging Face
Hugging FaceUpdated Jun 30, 2026
4.8

Kaggle Datasets Platform

Kaggle hosts tens of thousands of public datasets spanning machine learning competitions, real-world business problems, and academic research. All datasets are available for online analysis through built-in notebooks and can be downloaded directly for local experimentation and model training.

FreeDatasetBeginnerKaggleDatasets
KaggleUpdated Jun 18, 2026
4.7

Papers with Code Datasets

Papers with Code links benchmark datasets from academic papers to their corresponding state-of-the-art models and open-source implementations. This makes it easy to compare methods, reproduce results, and find the right dataset and code baseline for your research.

FreeDatasetIntermediateDatasetsPapers
Papers with CodeUpdated Jun 20, 2026
4.5

Google Dataset Search

Google Dataset Search is a dedicated search engine that indexes public datasets across thousands of sources, letting researchers and developers quickly locate the specific data they need for training, analysis, or validation regardless of where it is hosted.

FreeDatasetBeginnerGoogleDatasets
GoogleUpdated Jun 22, 2026
4.6

Common Crawl Web Corpus

Common Crawl is an open, petabyte-scale web crawl corpus that serves as a foundational pre-training data source for many large language models. The corpus is freely available for download and widely used in NLP research, data mining, and large-scale machine learning experiments.

FreeDatasetAdvancedCommon CrawlPre-training
Common CrawlUpdated Jun 12, 2026
4.5

LAION Open Multimodal Datasets

LAION (Large-Scale Artificial Intelligence Open Network) provides massive open-source image-text pair datasets that have been widely used to train multimodal models such as CLIP and Stable Diffusion. The datasets are freely available and serve as key infrastructure for open-source generative AI research.

FreeDatasetAdvancedLAIONMultimodal
LAIONUpdated Jun 16, 2026
4.4

UCI Machine Learning Repository

The UCI Machine Learning Repository is a classic collection of hundreds of structured datasets maintained since the late 1990s, widely used for hands-on practice, algorithm benchmarking, and teaching fundamental machine learning concepts in academic courses worldwide.

FreeDatasetBeginnerUCIDatasets
UC IrvineUpdated Jun 10, 2026
4.5

The Stack Source Code Dataset

The Stack is a large, permissively licensed source code dataset from the BigCode project, covering hundreds of programming languages. It is widely used to pre-train and evaluate code generation models, and includes opt-out mechanisms and deduplication tooling to support responsible data practices.

FreeDatasetAdvancedThe StackCode
BigCodeUpdated Jun 22, 2026
4.7

ImageNet

ImageNet is a landmark large-scale image dataset with millions of labeled images spanning thousands of categories. It powered the deep learning revolution in computer vision through the ILSVRC benchmark competition and remains foundational for pre-training, transfer learning, and evaluating image classification models.

FreeDatasetIntermediateImageNetComputer Vision
Stanford Vision LabUpdated Jun 14, 2026
NEW

Alpaca-Cleaned

Alpaca-Cleaned is a language modeling dataset by yahma on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

Alpaca CoT

Alpaca CoT is a Chinese open dataset by QingyiSi on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

Alpaca

Alpaca is a language modeling dataset by tatsu-lab on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

Bad Prompt

Bad Prompt is an open dataset by Nerfgun3 on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

c4

c4 is a multilingual language modeling dataset by allenai on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

Chinese DeepSeek R1 Distill data 110k

Chinese DeepSeek R1 Distill data 110k is a Chinese language modeling dataset by Congliu on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

COIG CQIA

COIG CQIA is a Chinese question answering dataset by m-a-p on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

Cosmopedia

Cosmopedia is an open dataset by HuggingFaceTB on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

CulturaX

CulturaX is a multilingual language modeling dataset by uonlp on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

Databricks Dolly 15k

Databricks Dolly 15k is a question answering dataset by databricks on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

DiffusionDB

DiffusionDB is an image captioning dataset by poloclub on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

Dolma

Dolma is a language modeling dataset by allenai on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

EasyNegative

EasyNegative is an open dataset by gsdf on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

Falcon RefinedWeb

Falcon RefinedWeb is a language modeling dataset by tiiuae on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

📄 FinePDFs

📄 FinePDFs is a language modeling dataset by HuggingFaceFW on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

FineWeb

FineWeb is a language modeling dataset by HuggingFaceFW on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

🥂 FineWeb 2

🥂 FineWeb 2 is a language modeling dataset by HuggingFaceFW on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

FineWeb-Edu

FineWeb-Edu is a language modeling dataset by HuggingFaceFW on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

GAIA

GAIA is an open dataset by gaia-benchmark on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

GLUE (General Language Understanding Evaluation benchmark)

GLUE (General Language Understanding Evaluation benchmark) is a text classification dataset by nyu-mll on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

Grade School Math 8K

Grade School Math 8K is a language modeling dataset by openai on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

Hh Rlhf

Hh Rlhf is an open dataset by Anthropic on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

hle

hle is an open dataset by cais on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

ImageNet

ImageNet is an image classification dataset by ILSVRC on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

IMDB

IMDB is a text classification dataset by stanfordnlp on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

Infinity Instruct

Infinity Instruct is a Chinese language modeling dataset by BAAI on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

Internship Warehouse

Internship Warehouse is an open dataset by FlyRank on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

Llama Nemotron Post Training Dataset

Llama Nemotron Post Training Dataset is an open dataset by nvidia on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

LLaVA Instruct 150K

LLaVA Instruct 150K is a question answering dataset by liuhaotian on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

Lmsys Chat 1m

Lmsys Chat 1m is an open dataset by lmsys on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

medical o1 reasoning SFT

medical o1 reasoning SFT is a Chinese question answering dataset by FreedomIntelligence on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

Medical Prescription English Audio Dataset

Medical Prescription English Audio Dataset is an open dataset by HumynLabs on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

Measuring Massive Multitask Language Understanding

Measuring Massive Multitask Language Understanding is a question answering dataset by cais on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

MNBVC

MNBVC is a Chinese language modeling dataset by liwu on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

NuminaMath CoT

NuminaMath CoT is a language modeling dataset by AI-MO on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

OpenAssistant Conversations

OpenAssistant Conversations is a multilingual open dataset by OpenAssistant on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

OpenHermes 2.5

OpenHermes 2.5 is an open dataset by teknium on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

OpenOrca

OpenOrca is a text classification dataset by Open-Orca on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

OpenR1 Math 220k

OpenR1 Math 220k is an open dataset by open-r1 on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

OpenThoughts 114k

OpenThoughts 114k is an open dataset by open-thoughts on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

Opus 4.6 Reasoning 3000x filtered

Opus 4.6 Reasoning 3000x filtered is an open dataset by nohurry on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

PersonaHub

PersonaHub is a Chinese language modeling dataset by proj-persona on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

PhysicalAI Autonomous Vehicles

PhysicalAI Autonomous Vehicles is an open dataset by nvidia on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

Prompts.chat

Prompts.chat is a question answering dataset by fka on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

RedPajama Data 1T

RedPajama Data 1T is a language modeling dataset by togethercomputer on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

ShareGPT Vicuna unfiltered

ShareGPT Vicuna unfiltered is an open dataset by anon8231489123 on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

SQuAD

SQuAD is a question answering dataset by rajpurkar on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

Synthetic Text To Sql

Synthetic Text To Sql is a question answering dataset by gretelai on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

The-Stack-v2

The-Stack-v2 is a language modeling dataset by bigcode on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

TinyStories

TinyStories is a language modeling dataset by roneneldan on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026
NEW

UltraChat 200k

UltraChat 200k is a language modeling dataset by HuggingFaceH4 on Hugging Face.

DatasetIntermediateDatasetHugging Face
AI Resource HubUpdated Sep 10, 2026

+ 1,766 more — VIEW ALL →