Need more data?

Exabite (formerly Calaveras) provides pre-training and post-training data for frontier AI labs. Trillion-dollar corporations are training the most powerful AI models on earth with our data.

Schedule a chat
01X0+ Exabytes

Total OTS data

02X0+ Trillion

Corpus-ready tokens (fully incremental to CommonCrawl, GitHub, and other common sources)

03Endless

Supply of new novel datasets spanning pretraining, SFT, RL, and subject-specific data.

01 / WHAT WE PROVIDE

Every stage of training.
One data partner.

01

Large OTS Catalog

Move faster with a deep off-the-shelf catalog spanning pre-training, post-training, and RL—including video, documents, and global web text beyond Common Crawl. We can deliver exabytes of data and 60T+ corpus-ready tokens.

02

Web Scraping

Our ethical scraping infrastructure is built for petabyte- and exabyte-scale procurements. We deliver faster, cheaper, and at higher volume than any competitor.

03

Novel Data

We continually develop new ways to source hard-to-find data for pre-training, post-training, RL, and evals. Ask for our latest catalog!

hi@exabite.ai

02 / EXABITE IS FOR YOU

Built for customers.
Not VCs.

Our team is entirely technical, from Stanford, MIT, Magic.dev, and Pika. We are laser focused on helping you build the best model.

We have invested substantially in building some of the best technical scraping infrastructure in our industry. Ask for a sample!

TECHNICAL TEAMCUSTOMER FUNDEDPROVEN AT SCALE

03 / PRICE MATCH

10–20%

Save on equivalent data.

For equivalent internet-sourced data on equivalent timelines, we aim to beat any competitor quote by 10–20%.

Request a price match
READY WHEN YOU ARE

Let's talk data.

Book time with us