Open Speech. A Brighter AI Future.

Human Voices Build Better AI

Collect high-quality, multilingual voice data for TTS and STT. Simple, fast, and reliable.

PDF & Text

Turn into recording clips

Guided Studio

4–15s auto-save takes

Export Ready

Train TTS & STT models

10K+

Contributors

∞Open

Extensible Dialects

100%

Open Data Options

Built for

Researchers & Builders

QAI Voice
v1.4.2 · online
Dashboard
Projects
Recordings
Review Queue
Speakers
Datasets & Export
Analytics
Settings
H

Good morning, Hawar

Let's make voices matter.

3

Active Projects

142

Completed Clips

28

Remaining

2.4 h

Total Recorded

Continue RecordingView Project
ckb

Kurdish Common Texts

14%
Your Progress342 / 2,500
Continue Recording
Your ProgressLast 7 days
150750
M
T
W
T
F
S
S
Recent ProjectsView All
::

Kurdish Common Texts

2,500 clips · ckb

68%Continue
ar

Arabic News Dataset

1,800 clips · ar

24%Continue
Recent ActivityView All
•

New recording submitted

2 minutes ago

✓

Clip approved

12 minutes ago

Real People

Real Voices

A Better Tomorrow.

See QAI Voice in action

From source text to a useful dataset

Five short guides show how a project moves through collection, recording, review, and export.

Sound effects included

Workspace Preview

Explore the Dashboard

Test drive the self-hosted command center. Click between navigation tabs inside the dashboard preview to inspect real-time telemetry, dataset projects, audio verification, and exporter tools.

QAI Voice
Live Preview
A

Good morning, Admin

Here's what's happening with your voice datasets.

Sep 1, 2026 – Sep 30, 2026
Total Projects

12

Recorded Clips

8,421

Approved Clips

6,973

Hours of Audio

48.2

Recording Activity
2K1.5K1K5000
Sep 1
.
.
Sep 7
.
.
Sep 14
.
.
Sep 21
.
Sep 28
Language Distribution
Kurdish
42%
Arabic
33%
English
25%
Total 48.2 Hours Lossless Audio
Recent Activity
H

New recording submitted

2 minutes ago

S

Clip approved

12 minutes ago

Approved
A

New speaker joined

1 hour ago

Frontier Speech Engineering

Engineered for Every Voice Pipeline

From low-resource dialect preservation to high-volume commercial TTS dataset generation, QAI Voice delivers verifiable audio.

Text-to-Speech

TTS Voice Synthesis

Produce clean 24kHz studio datasets with phoneme-aligned transcripts for training XTTS, VITS, and FastSpeech acoustic models.

Speech-to-Text

Automatic Speech Recognition

Collect thousands of varied acoustic takes across Sorani & Kurmanji Kurdish, Arabic dialects, and accented English for Whisper fine-tuning.

Preservation

Language & Dialect Preservation

Safeguard endangered oral heritage and regional languages with structured community recording pipelines and full speaker consent.

Open Data

Academic & Corporate Research

Export ethical, license-compliant audio corpora with rich metadata, verifiable SNR acoustic stats, and standardized train/dev/test splits.

Interactive Acoustic Studio

Listen to Broadcast-Quality Takes

Test drive our studio recording engine. Toggle between raw microphone takes and normalized audio inspected by real-time acoustic quality gates.

H

Hawar Ahmad

SPK-0012

Sorani (Central) · Classical Literature & Lexicon

TELEPROMPTER PROMPT #428CKB · Sorani

ئەمڕۆ کەشوهەوا زۆر جوان و سازگارە، و دەمەوێت زمانی شیرینی کوردی بە جوانی فێرببم.

“Today the weather is very pleasant and refreshing, and I want to learn the sweet Kurdish language gracefully.”

00:02.45/00:07.41
Processed with 3-Band Parametric Filter
Real-Time Acoustic TelemetryGATE PASSED
Integrated Loudness
-16.0LUFS

Target: -16.0 ± 0.5 LUFS

Signal-to-Noise Ratio
+34.6dB

Studio Reference

True Peak (Clipping)
-1.0dBFS

0 Clipped Samples

Audio Format Spec

24.0 kHz

16-bit PCM · Mono (1.0)

WAV PCM Uncompressed

Automated Gate Verifications
Trailing Silence Normalization
0.32s Pass
Spectral DC Bias Removal
0.00% Offset
Phoneme Alignment Check
99.4% Match
Self-Hosted & Scalable

Transparent Deployment Models

100% self-hosted on your own hardware or deployed on managed distributed GPU clusters.

Open SourceCommunity Edition
$0/ forever (Apache 2.0)

Deploy QAI Voice in minutes on your local workstations or on-premise servers with Docker Compose.

Unlimited speech recording clips & projects
Local MinIO S3 & PostgreSQL storage
IndexedDB offline-first recording studio
Export to LJSpeech, Common Voice & Parquet
Deploy Free Community
Enterprise
Custom Cluster
Tailored/ multi-team SLA

Designed for large research institutes, language authorities, and commercial voice AI laboratories.

Everything in Community, plus:
Distributed S3 bucket synchronization
Custom automated acoustic quality gate webhooks
Multi-tenant RBAC & SSO integration
Dedicated technical support & migration engineer
Contact Enterprise Sales→