NewHow to Remain Valuable When Intelligence Becomes Cheap — 224 pages on the work AI can’t commoditize. Get the ebook →

Reference facts for your agents, in fewer tokens.

185 public-domain datasets — HTTP codes, SQL joins, ports, cognitive biases — as readable tables and as JSON, YAML, or TOON from one URL. TOON uses 32% fewer tokens than the JSON we serve, and 60% fewer on the best dataset. No key, no signup.

Looking for something specific? Press / to search all 185.

Datasets in 27 categories
185
Tokens, TOON vs JSON
−32%
API keys needed
0
tcp-ip-protocol-suite
Loading live response…
  • JSON—
  • YAML—
  • TOON—

— Open full response →

Browse

185 datasets on 27 shelves.

Each spine is a category, as tall as the number of datasets it holds, and keeps its color wherever it appears on this site. Point at one to read it; select it to open it. Type / to search instead.

Code & models

Programming Languages

Syntax, features, and comparisons of major programming languages.

Open all 10 datasets →

Code & models

Data Structures

Fundamental data structures with properties and complexity analysis.

Open all 10 datasets →

Code & models

Algorithms

Classic algorithms, strategies, and complexity.

Open all 7 datasets →

Code & models

Design Patterns

Software design patterns with descriptions and use cases.

Open all 7 datasets →

Code & models

AI & ML Terms

Key terminology in artificial intelligence and machine learning.

Open all 10 datasets →

Systems & the web

Unix Commands

Common Unix/Linux shell commands and their usage.

Open all 9 datasets →

Systems & the web

DevOps

DevOps practices, CI/CD, and infrastructure as code.

Open all 6 datasets →

Systems & the web

Cloud Computing

Cloud service models, providers, and deployment patterns.

Open all 7 datasets →

Systems & the web

Networking

Network protocols, models, and architecture.

Open all 7 datasets →

Systems & the web

HTTP Status Codes

Complete reference of HTTP response status codes.

Open all 10 datasets →

Systems & the web

Web Technologies

Standards, protocols, and APIs for web development.

Open all 9 datasets →

Systems & the web

Database Types

Different database systems and their characteristics.

Open all 9 datasets →

Systems & the web

Cybersecurity

Security principles, threats, and defensive practices.

Open all 6 datasets →

Natural science

Mathematics

Core mathematical concepts, theorems, and notation.

Open all 7 datasets →

Natural science

Physics

Core physics concepts and laws.

Open all 6 datasets →

Natural science

Biology

Fundamental concepts in biology: cell biology, genetics, evolution, and ecology.

Open all 6 datasets →

People & society

History

Major historical events, civilizations, eras, and turning points that shaped the modern world.

Open all 6 datasets →

People & society

Philosophy

Major philosophical traditions, schools of thought, key thinkers, and fundamental questions.

Open all 6 datasets →

People & society

Psychology

Psychological concepts, disorders, and theories.

Open all 6 datasets →

People & society

Literature

Literary movements, genres, narrative techniques, and major works across history.

Open all 6 datasets →

What TOON removes

Same facts. Less syntax.

Scroll to watch eight real rows of the NATO phonetic alphabet lose everything a reader does not need — the keys repeated on every row, then the braces and quotes — until only the header and the values remain.

  • The JSONEight rows, one per line, exactly as the API serves them.
  • Keys, spelled on every row"letter" and "code_word" appear eight times each. They are the same eight times.
  • Delimiters that delimit nothingBraces, quotes and trailing commas go. A comma and a newline are enough.
  • TOONThe keys, once, in a header. Every row is just values.
phonetic-alphabet · first 8 rows 108bytes
{"data":[data[8]{letter,code_word}:{"letter":"A","code_word":"Alpha"},{"letter":"B","code_word":"Bravo"},{"letter":"C","code_word":"Charlie"},{"letter":"D","code_word":"Delta"},{"letter":"E","code_word":"Echo"},{"letter":"F","code_word":"Foxtrot"},{"letter":"G","code_word":"Golf"},{"letter":"H","code_word":"Hotel"}]}

292 bytes as minified JSON, 108 as TOON: 63% smaller for these rows. Whole-catalog figures use a real tokenizer and are in the table below.

Visual browser

185 datasets, one cube each.

Every cube is one dataset. Every stack is one category, and its height is how many datasets that category holds — so the silhouette below is the catalog itself. Drag to turn it, hover to identify a cube, select it to open that dataset's page.

TOON Studio

See the saving on your own data.

Paste JSON and this page encodes it in your browser with the same encoder the API uses. Byte counts are exact; token counts are labelled estimates. Table-shaped data saves the most — try a flat list of records, then a deeply nested config, and watch the number move.

Samples
JSON in llm-glossary
  • —
  • —
  • —
  • —

Token counts here are estimates from a character heuristic; the site-wide figures above use a real tokenizer. The browser encoder (js/toon.js) is held byte-identical to the server's by scripts/tests/toon-parity.php and checked against the reference TOON implementation by scripts/tests/toon-conformance.mjs.

Choose a format

One dataset, three encodings.

The same data lives on the server once. What changes is what it costs in a prompt and how it reads in an editor.

Tokens to send all 185 datasets (o200k_base, as served)

  • JSONas served 246,704
  • YAML 197,814−20% vs JSON
  • JSON, minifiedno whitespace 179,087−27% vs JSON
  • TOON 166,870−32% vs JSON
Measured on all 185 datasets as served (1,199,518 bytes of JSON) and on the combined catalog index. Tokens: o200k_base.
PropertyJSONYAMLTOON
Readable by a human Yes, with punctuation noise The easiest to read Yes — uniform rows become a table
Parser required Built into every language A YAML library A TOON library, or a line parser — this site publishes a small Python one
Nested data Native Native Native, plus a tabular form for uniform arrays
Size of the catalog index 66 KB 44 KB 33 KB (−49%)
Tokens, all 185 datasets 246,704 197,814 (−20%) 166,870 (−32%)
Against minified JSON Baseline: 179,087 tokens Larger −7% overall; 161 of 185 datasets smaller
Reach for it when Services exchange data and every language must parse it A human edits the file Table-shaped data goes into a prompt on every request

Best case: phonetic-alphabet uses 60% fewer tokens in TOON. Every dataset page shows its own sizes in all three formats. Pick one and compare →

About

Reference data should be boring in the best way: stable, cited, and cheap to put into a prompt. YJTOON keeps 185 public-domain datasets that way — stored once as JSON, served as JSON, YAML, or TOON, no authentication, no keys.

Also from the author

How to Remain Valuable When Intelligence Becomes Cheap

Intelligence is becoming cheap. Judgment, ownership, and trust are not — this book maps the advantages that stay scarce.

Format
PDF + EPUB
Length
224 pages
Price
$3.84
  • The human edge: accountability, taste, and trust in an age of fluent machines
  • The economics: what gets commoditized, what stays scarce, and where to stand
  • Practical moves for people who build with AI every day

Start free: Decision biases checklist · How this catalog pays for itself · Grounding LLM answers

How to Remain Valuable When Intelligence Becomes Cheap, book cover

Open Data for AGI — Why Public, Free, and Open Datasets Matter, book cover

The book behind this catalog

Open Data for AGI

A practical brief on why public, free, and open datasets decide what AI systems become.

eBook · PDF · $4.99 · All site data stays free; the book is the argument, not a paywall.

The five machines0:00 / 8:50

Film · 8:50 · English captions

Five products gave the chatbot its own computer. Rented or owned?

Grok Bot, Meta Muse, ChatGPT Dots, OpenClaw and Hermes compared in an 8-minute film: what each one does, who it is for, what it costs, and rented versus owned.

FAQ

Short answers.

Is YJTOON free to use?

Yes. All 185 datasets are released under CC0, which places them in the public domain. There is no signup, no API key, and no paid tier. The only limit is a rate limit of 120 requests per minute per IP.

What is TOON?

TOON (Token-Oriented Object Notation) is a compact text encoding for structured data, aimed at prompts rather than at storage. A uniform array of objects becomes a table: the field names are written once in a header and every row carries only values. The encoder here follows the public TOON format and is checked against its reference implementation on every build.

Read the grammar overview → · Try it on your own JSON →

How much does TOON actually save?

Measured with the o200k_base tokenizer across all 185 datasets: 166,870 tokens in TOON versus 246,704 in the JSON files this site serves — 32% fewer, and up to 60% fewer on the best dataset (prime-numbers-list). Against minified JSON the overall saving is only 7%, because prose-heavy datasets gain little; 161 of 185 datasets are still smaller in TOON than in minified JSON. TOON pays off most on table-shaped data.

See the measurement →

Do I need an API key?

No. Every endpoint is anonymous. Requests are rate limited by a hash of the IP address, and raw IP addresses are never stored.

Can I use the data in a commercial product?

Yes. CC0 waives copyright entirely, so there is no attribution requirement and no restriction on commercial use. Attribution is welcome, but it is never required.

What CC0 actually covers →

How is this different from a vector database?

A vector index returns passages that look similar to a query. Reference data answers questions that have one correct answer. For those, an exact lookup against a known dataset is cheaper and more reliable than a nearest-neighbour search.

When not to use a vector database →

How should I cite a dataset?

Every dataset page includes a suggested citation and a permanent URL. Attribution is not required under CC0, but it helps other people find the source.

See a dataset page and its citation →