Skip to content
Shahryar Sabir Muhammed
Shahryar Sabir
Software & AI Engineer
← Back to Systems
RegaLabs R&D Hugging Face Model
PythonPyTorchCosyVoice 3TTSSorani KurdishHugging FaceAI/ML

RegaLabs-TTS (CosyVoice 3 Central Kurdish)

State-of-the-art Central Kurdish (Sorani / سۆرانی) text-to-speech and one-shot voice cloning adaptation on Hugging Face, trained on 53 hours of curated speech data.

$ pip install git+https://github.com/RegaLabs/RegaLabs-TTS.git

RegaLabs-TTS: CosyVoice 3 Central Kurdish (Sorani) Adaptation

RegaLabs-TTS is a high-quality Central Kurdish (Sorani / سۆرانی) text-to-speech and voice cloning adaptation developed at RegaLabs based on CosyVoice 3 (FunAudioLLM/Fun-CosyVoice3-0.5B-2512).

  • Hugging Face Model: RegaLabs/RegaLabs-TTS
  • GitHub Repository: RegaLabs/RegaLabs-TTS
  • Base Architecture: CosyVoice 3 (0.5B Flow Matching + Multi-speaker Conditioning)
  • License: Apache 2.0 (Checkpoints & Inference Code)

Dataset & Training Milestones

  • 53 Total Hours of cleaned, phonetically aligned Sorani Kurdish speech data.
    • Male Speakers: ~35–40 hours.
    • Female Speakers: ~13–18 hours.
  • Flow Adaptation Model: Step 2300 checkpoint (cosyvoice3_sorani_flow_best_step2300.pt) paired with matching YAML architecture definition.
  • One-Shot Voice Cloning: High similarity, prosody accuracy, and natural vocal timbre from single reference audio samples.

Interactive Audio Sample

  • Prompt: "سڵاو، بەخێربێن بۆ پڕۆژەی RegaLabs-TTS" (Hello, welcome to the RegaLabs-TTS project)
  • Sample File: aran_en021.wav

Quick Installation & CLI Inference

# 1. Install via pip directly from GitHub
pip install git+https://github.com/RegaLabs/RegaLabs-TTS.git

# 2. Clone base engine and model weights
git clone --recursive https://github.com/FunAudioLLM/CosyVoice.git
cd CosyVoice
pip install -r requirements.txt

git clone https://huggingface.co/RegaLabs/RegaLabs-TTS regalabs-tts-weights

# 3. Synthesize speech in Kurdish
python regalabs-tts-weights/infer.py \
  --text "سڵاو، بەخێربێن بۆ پڕۆژەی RegaLabs-TTS" \
  --prompt-wav regalabs-tts-weights/samples/aran_en021.wav \
  --prompt-text "ئەمە دەنگی نموونەیە" \
  --out output_sorani.wav

Model Artifacts & Architecture

  • cosyvoice3_sorani_flow_best_step2300.pt — Sorani acoustic/flow adaptation model weights.
  • cosyvoice3_sorani_flow_best_step2300.yaml — Matching flow architecture configuration.
  • sorani/frontend.py — Central Kurdish text normalizer and grapheme-to-phoneme engine.
  • infer.py — Ready-to-run Sorani inference CLI.
  • app.py — Gradio Web UI Live Demo script.

System Specifications

0
E
1
n
2
g
3
i
4
n
5
e
6
e
7
r
8
e
9
d
10
11
a
12
n
13
d
14
15
p
16
u
17
b
18
l
19
i
20
s
21
h
22
e
23
d
24
25
R
26
e
27
g
28
a
29
L
30
a
31
b
32
s
33
-
34
T
35
T
36
S
37
,
38
39
a
40
41
p
42
r
43
o
44
d
45
u
46
c
47
t
48
i
49
o
50
n
51
-
52
g
53
r
54
a
55
d
56
e
57
58
C
59
e
60
n
61
t
62
r
63
a
64
l
65
66
K
67
u
68
r
69
d
70
i
71
s
72
h
73
74
(
75
S
76
o
77
r
78
a
79
n
80
i
81
)
82
83
a
84
c
85
o
86
u
87
s
88
t
89
i
90
c
91
92
a
93
d
94
a
95
p
96
t
97
a
98
t
99
i
100
o
101
n
102
103
b
104
a
105
s
106
e
107
d
108
109
o
110
n
111
112
F
113
u
114
n
115
A
116
u
117
d
118
i
119
o
120
L
121
L
122
M
123
124
C
125
o
126
s
127
y
128
V
129
o
130
i
131
c
132
e
133
134
3
135
.
136
137
T
138
r
139
a
140
i
141
n
142
e
143
d
144
145
o
146
n
147
148
5
149
3
150
151
h
152
o
153
u
154
r
155
s
156
157
o
158
f
159
160
p
161
h
162
o
163
n
164
e
165
t
166
i
167
c
168
169
K
170
u
171
r
172
d
173
i
174
s
175
h
176
177
s
178
p
179
e
180
e
181
c
182
h
183
184
d
185
a
186
t
187
a
188
189
(
190
3
191
5
192
-
193
4
194
0
195
h
196
197
m
198
a
199
l
200
e
201
,
202
203
1
204
3
205
-
206
1
207
8
208
h
209
210
f
211
e
212
m
213
a
214
l
215
e
216
)
217
,
218
219
a
220
c
221
h
222
i
223
e
224
v
225
i
226
n
227
g
228
229
f
230
l
231
a
232
w
233
l
234
e
235
s
236
s
237
238
m
239
a
240
l
241
e
242
243
v
244
o
245
i
246
c
247
e
248
249
c
250
l
251
o
252
n
253
i
254
n
255
g
256
,
257
258
p
259
r
260
e
261
c
262
i
263
s
264
e
265
266
p
267
r
268
o
269
s
270
o
271
d
272
i
273
c
274
275
t
276
i
277
m
278
i
279
n
280
g
281
,
282
283
a
284
n
285
d
286
287
n
288
a
289
t
290
u
291
r
292
a
293
l
294
295
i
296
n
297
t
298
o
299
n
300
a
301
t
302
i
303
o
304
n
305
306
f
307
o
308
r
309
310
l
311
o
312
w
313
-
314
r
315
e
316
s
317
o
318
u
319
r
320
c
321
e
322
323
K
324
u
325
r
326
d
327
i
328
s
329
h
330
331
s
332
p
333
e
334
e
335
c
336
h
337
338
s
339
y
340
n
341
t
342
h
343
e
344
s
345
i
346
s
347
.
348
349
R
350
e
351
l
352
e
353
a
354
s
355
e
356
d
357
358
o
359
p
360
e
361
n
362
l
363
y
364
365
o
366
n
367
368
H
369
u
370
g
371
g
372
i
373
n
374
g
375
376
F
377
a
378
c
379
e
380
381
a
382
n
383
d
384
385
G
386
i
387
t
388
H
389
u
390
b
391
392
u
393
n
394
d
395
e
396
r
397
398
A
399
p
400
a
401
c
402
h
403
e
404
-
405
2
406
.
407
0
408
409
w
410
i
411
t
412
h
413
414
c
415
o
416
m
417
p
418
l
419
e
420
t
421
e
422
423
i
424
n
425
f
426
e
427
r
428
e
429
n
430
c
431
e
432
433
s
434
c
435
r
436
i
437
p
438
t
439
s
440
441
a
442
n
443
d
444
445
G
446
r
447
a
448
d
449
i
450
o
451
452
d
453
e
454
m
455
o
456
.

Need this in production?

I build and deploy custom speech AI models and persistent agent systems for businesses and teams.

Schedule Architecture Call →