Pay once per dataset and own it. Every price includes the perpetual commercial license, and there is no subscription.
Read the real thing before you talk to anyone.
Take the set you need and own it for good.
Your dialect or domain, first specimen in a week.
Payment and delivery both happen over WhatsApp. No account, no invoice portal.
No. Every row in every dataset is synthetically generated. There is no scraped text, no real user data and no PII — safe under NDMO, GDPR and any comparable framework.
Because it is honest about what exists. Entries marked “in production” are on the roadmap and carry no price. Only “available now” means we can send you the file today — and a commission moves a roadmap entry to the front of the queue.
You describe the dataset — domain, dialects, label space, volume. We write a recipe: a taxonomy, a prompt set and validators for your rules. You get a small paid pilot to check the quality before committing to the full run.
JSONL by default, plus JSON, CSV and Parquet, with per-field selection at export. Ready for Hugging Face, Axolotl, LLaMA-Factory or any pipeline that reads a line-delimited file.
20+ validators run on every row: dialect contamination, brand-capability rules, phrase-frequency caps, structural ordering and domain vocabulary. Failures go back to the model with the specific complaint, up to three attempts; what still fails is tagged and excluded.
Yes. Every purchased dataset comes with a commercial use licence — train models, ship products, publish research. No restrictions and no revenue share.
Yes. The flagship dataset has a free 500-row download on this page and live samples on its detail page. For anything else, ask and we will cut you a sample.
Message us on WhatsApp. We confirm the dataset and the price, you pay, and the file comes back in the same thread. No account, no sales call.