What this chart shows. The fifteen dataset repositories with the most likes on Hugging Face, read on 21 September 2026 at 08:19 UTC. The leader is not a training corpus at all: prompts.chat, whose entire repository is one CSV file of prompt texts released under a CC0 licence, holds 9,839 likes, 2.93 times the 3,357 of the web-scale corpus fineweb. Anthropic's hh-rlhf is third on 2,091. The fifteen hold 30,995 likes between them.
How the data is collected. Likes work the same way for datasets as for models: one signed-in account, one repository, a public count exposed as the likes field of the API, and the list here is the Hub's own sort by that field. Dataset downloads, quoted below for comparison, are counted differently from model downloads. Since September 2024 the Hub treats every file a user pulls from one dataset repository within a five-minute window, identified by IP address, as a single download, counted server-side as the files are served. Before that date it counted calls to the datasets library's load_dataset function instead, and manual downloads through the web interface or tools such as wget were not counted at all.
Background. The list is a picture of what the open model-training stack is built from. Alongside prompts.chat sit web-scale pretraining corpora such as fineweb, fineweb-edu, dolma, RedPajama-Data-1T and Wikipedia; human feedback and instruction sets including hh-rlhf, oasst1, alpaca, databricks-dolly-15k and OpenOrca; the maths benchmark gsm8k, which the Hub tags as an official benchmark; the small synthetic corpus TinyStories; and EasyNegative, whose files are an embedding and two sample images rather than text. Twelve of the fifteen repositories were created before 2024, and none were created in 2025 or later.
Limits. A like is a bookmark by one account, not a measure of use, and the two diverge here as widely as they do for models: the most liked dataset drew 22,609 downloads in the last thirty days while OpenAI's gsm8k, fourth by likes, drew 1,221,924. Because likes never expire, age helps, which is part of why nothing published in the last year appears. Dataset download figures are also counted per IP address within a five-minute window rather than per person, so shared networks and cloud jobs are not separated.
Analysis
prompts.chat holds 31.7 per cent of the likes in this group and is the only entry above 5,000. The drop from first to second place is 6,482 likes, larger than the entire total of the fifteenth entry, dolma, on 1,087. Only one publisher appears twice, HuggingFaceFW, with fineweb and fineweb-edu. The oldest repository in the group is wikimedia's Wikipedia dump, carried over in the Hub's March 2022 migration, and the newest is medical-o1-reasoning-SFT from December 2024. Downloads tell a different story again: OpenAI's gsm8k leads the group on 1,221,924 downloads in thirty days, 54.05 times the leader by likes.
Values are shown as reported by the source. Nothing is rescaled, rebased or converted — the only changes are thousand separators and decimals shown to two places.
Entities are ranked highest-first on their value in 2026.
all-MiniLM-L6-v2 was downloaded 251,048,129 times from Hugging Face in thirty days. Five public rankings from the Hub show that the models the industry runs and...
Rankings & Full Data
15 records
· Likes on Hugging Face · ranked by 2026