#speech (3)

Whisper OSS

OpenAIの音声認識モデル。多言語の文字起こし・翻訳に対応し、日本語の精度も高い定番OSS。

stars: 108.6k price: open_source license: MIT ja: full level: beginner
faster-whisper OSS

Whisperを最大4倍高速化した再実装。同じ精度でメモリ消費も少なく、文字起こしの実運用での標準になっている。

stars: 25.3k price: open_source license: MIT ja: full level: intermediate
Kokoro TTS OSS

わずか82Mパラメータで高品質な音声を生成する軽量TTSモデル。速度・コスト・品質のバランスで注目を集める。

stars: 8.7k price: open_source license: Apache-2.0 ja: partial level: intermediate