Kivotos

Multilingual natural speech

Commissionable

Spontaneous, natural speech from native speakers outside studio conditions — real acoustic environments with engine noise, traffic and overlapping speakers — across languages underrepresented in commercial speech corpora.

Specifications

Dataset IDKV-SPC-001
AvailabilityCommissionable
Intended usesVoice AI and ASR · Conversational models · Multimodal alignment
ModalityNatural speech audio; optional transcription and annotation
Source categoryConsumer platform networks; native speakers in real acoustic conditions
GeographyPending verificationDefined per program across the operating footprint
LanguagesPending verificationLanguage and market subset defined per program, including low-resource languages
Volume / capacityPending verificationDefined per program: fixed monthly hours
TimelinePending verificationStated at specification; confirmed before activation
Formats & schemaPending verificationSampling rate, channels and schema agreed in the capture specification
ProcessingRaw, or transcribed and annotated (diarization, labels) to buyer specification

Why this data is novel

Most available speech data is scripted, narrow in language coverage, or studio-recorded. Scripted corpora do not close the low-resource gap; this program class produces natural speech in deployment-like conditions under a defined capture specification.

Source methodology and collection context

Native speakers within established networks opt into defined capture tasks. Acoustic conditions, prompts or task structure, and session parameters follow the capture specification.

Composition (organic-data statement)

Human speech only; no synthetic voices. AI-assisted transcription, where used, is disclosed and quality-checked against human review samples.

Public availability and prior licensing

Pending verificationCommissioned output is new capture; prior-licensing status stated per program

Duplication, overlap and contamination

Pending verificationOverlap with public speech corpora assessed and reported per program

Rights and permitted uses

Consent at task acceptance for the declared purpose. All-party consent requirements are resolved in the jurisdiction review before any recording.

Privacy, PII and de-identification

Voice is personal data: jurisdiction-specific review covers biometric and voice-data handling, retention and de-identification options before activation.

Quality, acceptance criteria and known limitations

Two-stage verification: automated audio integrity checks plus independent adjudication against written acceptance criteria, with transcription quality sampling where applicable.

Delivery

Secure transfer with versioned manifests, checksums, consent attestation and rejection analysis.

Commercial structure

Commissioned program (buyer-specific). Exclusivity where required, priced as a rights grade.

Evaluation pack

Buyer diligence is productized. A qualified request receives an evaluation pack containing:

  • Representative stratified sample
  • Dataset card
  • Data dictionary and schema
  • Provenance summary
  • Rights and consent summary
  • Privacy and PII assessment
  • Quality and novelty report
  • Known limitations
  • Security and delivery sheet
  • Version and checksum manifest

Samples are never cherry-picked: each pack states whether its sample is representative, illustrative, anonymized, or structurally simulated.