The data your research needs.One query away.

Describe what you need in plain English. Datta searches a consented corpus of conversational voice and video, and shows you the exact timestamped moments, with transcript, emotion, and prosody, before you license anything.

m. okafor · video3f81aa · 12:04
a. reyes · voicec09e44 · 31:55
j. lindqvist · video5d217b · 18:21
maya r. · voice84c2f1 · 09:12
s. tanaka · voiceb7d90a · 06:30
p. duval · voicee63f08 · 22:47

session_84c2f1 · maya r. · 02:14

“I’ve already explained this twice, but sure, let’s go through it again.”

consent · recordedlicensed · cleared for research use

you ask in plain English

illustration · queries and sessions from the live demo corpus

the workbench

Research is your job. Providing data is ours.

tool.datta.ml · live demo · canned datainteractive · click anything to take over
01 / 06Search in plain English.

Describe the data you need. Datta reads back consented, well-described sessions you can preview and license.

voice · consented corpus
try
Sessionsrun a query to read sessions from the corpus
early access

results appear here

live demo · canned data · the real console works like this · “early access” beats are our hands-on service

search in plain English · see the data before you license it · annotations built in · save it to a project · projects are versioned · request a license · the loop stays open

01/the failure modes

Where today's training data fails you.

fault_01

The sample lied.

The eval set was clean. The delivery wasn't. You found out three weeks into training.

fault_02

The vendor vanished.

Invoice paid, contact gone. No follow-up on provenance questions, no fixes, no recourse.

fault_03

The labels were wrong.

"Annotated" meant a spreadsheet from a different vendor, on a different schema, at a different frame rate.

fault_04

You couldn't find what you already had.

Somewhere in last year's purchase is the exact clip you need today. Nobody can find it.

fault_05

Your GPU money went to fixing data you already paid for.

Compute and grant budgets are finite. Every hour of cleaning and re-labeling someone else's data is money that should have bought training runs.

02/where the data comes from

Natural data. Nobody is performing.

Most conversational emotion data is acted: people paid to perform feelings on cue. Models trained on performance learn performance. Sessions in this corpus are real people in real conversations, recorded the way your research needs them.

Contribute data instead? For contributors →

source

Real people, real conversations.

Contributors record actual exchanges, natural speech on both sides, or upload conversations they already have. Nobody is reading emotion prompts to a webcam.

signal

Natural distributions.

Emotion, prosody, hesitation, and disfluency are captured as they occur in conversation, not as they are performed on cue, and annotated on one clock.

search

Findable by what actually happened.

Transcript, emotion, and prosody annotations make natural moments searchable. You get data that matches your query, not a sample deck.

Every session is fully consented at the source and cleared for research use.

03/get started

Early access is open.

We’re opening onboarding to research teams now. Licensing runs through us while the marketplace is in early access: you get the console, the corpus, and a human on the other end.

Book a walkthrough

get early access

we onboard research teams and contributors in small batches · a real person replies