offers speech quality and intelligibility that were previously achievable only at nearly double the bitrate. Below are the results of comparative tests with the MELPe 2400 bps vocoder, which operates at twice the bitrate.
Deployed in Digital HF Radio, tactical communications, secure voice, satellite communications, and other bandwidth-constrained communication applications.
- Advanced Speech Coding Approach: Pitch-synchronous processing and proprietary Tri-Wave excitation, providing an alternative to MELPe, AMBE+2 and ACELP-based approaches.
- Mature & Proven: Nearly two decades of continuous development and real-world deployment in professional communication systems.
- Speech & Beyond: High-quality speech plus improved reproduction of tones, sirens, music and other non-speech signals.
Note:
The original ITU-T P.50 speech files were resampled to 8 kHz. Files for each language were combined and inter-speech pauses were shortened to minimize the effect of silence on the evaluation results.
| Language | MELPe 2400 bps | TWELP 1200 bps | MELPe 1200 bps |
|---|---|---|---|
| American | 2.967 | 2.899 | 2.711 |
| Arabic | 3.024 | 2.868 | 2.722 |
| British | 2.885 | 2.791 | 2.689 |
| Chinese | 2.871 | 2.939 | 2.607 |
| Danish | 2.946 | 2.963 | 2.720 |
| Dutch | 2.856 | 2.772 | 2.644 |
| Finnish | 2.820 | 2.791 | 2.630 |
| French | 3.018 | 2.943 | 2.818 |
| German | 2.986 | 2.923 | 2.793 |
| Greek | 2.950 | 2.921 | 2.698 |
| Hindi | 3.070 | 2.968 | 2.847 |
| Hungarian | 3.060 | 2.944 | 2.793 |
| Italian | 3.213 | 3.092 | 2.983 |
| Japanese | 3.133 | 3.021 | 2.898 |
| Norwegian | 3.007 | 2.948 | 2.745 |
| Polish | 3.027 | 2.943 | 2.758 |
| Portuguese | 3.113 | 3.009 | 2.894 |
| Russian | 2.863 | 2.849 | 2.689 |
| Spanish | 3.046 | 2.928 | 2.816 |
| Swedish | 3.119 | 3.019 | 2.907 |
| Average | 2.999 | 2.927 | 2.768 |
|
The results confirm that TWELP 1200 bps provides speech quality close to that of MELPe 2400 bps while operating at half the bitrate. | |||
STOI and ESTOI were calculated using the updated 8 kHz test set described above. ASR-based Speech Intelligibility was evaluated separately using the original ITU-T P.50 speech files at their original 16 kHz sampling rate.
ASR reference transcripts were generated by whisper.cpp with the Whisper large-v3 model directly from the original, unprocessed 16 kHz P.50 speech files. For vocoder evaluation, the same speech files were converted to 8 kHz, processed by the vocoder, converted back to 16 kHz and transcribed using the same ASR system.
ASR-based Speech Intelligibility is calculated as 100 - CER (%) by comparing each processed-speech transcript with its corresponding ASR reference transcript.
CER is used instead of WER because it measures character-level recognition errors without treating an entire word as incorrect due to a single-character error.
CER scoring ignores letter case, whitespace and punctuation.
Control measurement: the same 16 kHz → 8 kHz → 16 kHz conversion path, but without any vocoder processing, resulted in an overall ASR-based intelligibility of 99.41%. This shows that narrowband conversion alone produces a small measurable reduction in the ASR-based intelligibility measure, even for clean, unencoded speech.
All vocoder results below are scored against the ASR reference transcripts generated from the original, unprocessed 16 kHz speech, not against transcripts from the converted signal.
| Language | MELPe 2400 bps | TWELP 1200 bps | MELPe 1200 bps |
|---|---|---|---|
| American | 99.43 | 98.92 | 98.51 |
| Arabic | 96.31 | 97.65 | 94.80 |
| British | 100.0 | 99.41 | 99.41 |
| Chinese | 90.89 | 97.04 | 86.70 |
| Danish | 93.94 | 90.73 | 87.35 |
| Dutch | 99.01 | 98.68 | 96.75 |
| Finnish | 98.20 | 95.17 | 92.71 |
| French | 98.97 | 98.93 | 98.22 |
| German | 99.22 | 98.91 | 98.83 |
| Greek | 98.34 | 97.01 | 94.69 |
| Hindi | 96.00 | 93.10 | 90.03 |
| Hungarian | 98.93 | 98.50 | 96.86 |
| Italian | 99.06 | 99.18 | 98.43 |
| Japanese | 96.58 | 96.72 | 96.58 |
| Norwegian | 99.52 | 98.61 | 98.55 |
| Polish | 98.56 | 97.48 | 96.16 |
| Portuguese | 99.06 | 99.38 | 98.05 |
| Russian | 99.78 | 99.49 | 99.13 |
| Spanish | 95.07 | 98.95 | 94.02 |
| Swedish | 98.15 | 96.42 | 94.25 |
| Overall | 97.97 | 97.29 | 95.53 |
|
Overall ASR-based speech intelligibility: 97.29% for TWELP 1200 bps, versus 97.97% for MELPe 2400 bps and 95.53% for MELPe 1200 bps. | |||
| Language | MELPe 2400 bps | TWELP 1200 bps | MELPe 1200 bps |
|---|---|---|---|
| American | 87.43 | 88.54 | 85.20 |
| Arabic | 87.70 | 87.61 | 84.76 |
| British | 85.10 | 85.87 | 80.97 |
| Chinese | 86.78 | 87.85 | 83.74 |
| Danish | 87.30 | 88.87 | 84.86 |
| Dutch | 86.19 | 86.84 | 82.66 |
| Finnish | 82.94 | 84.17 | 80.12 |
| French | 87.35 | 88.40 | 84.28 |
| German | 87.10 | 88.23 | 85.01 |
| Greek | 86.96 | 87.84 | 84.03 |
| Hindi | 86.89 | 88.18 | 84.38 |
| Hungarian | 87.85 | 87.90 | 84.72 |
| Italian | 87.57 | 88.19 | 85.00 |
| Japanese | 88.88 | 87.96 | 86.01 |
| Norwegian | 87.97 | 88.89 | 85.12 |
| Polish | 87.66 | 88.18 | 83.93 |
| Portuguese | 87.68 | 87.90 | 85.05 |
| Russian | 86.89 | 86.55 | 82.55 |
| Spanish | 86.57 | 87.25 | 83.57 |
| Swedish | 85.92 | 86.48 | 83.47 |
| Average | 86.94 | 87.59 | 83.97 |
|
Average STOI score: 87.59 for TWELP 1200 bps, versus 86.94 for MELPe 2400 bps and 83.97 for MELPe 1200 bps. | |||
| Language | MELPe 2400 bps | TWELP 1200 bps | MELPe 1200 bps |
|---|---|---|---|
| American | 81.43 | 81.12 | 78.01 |
| Arabic | 82.61 | 80.87 | 78.27 |
| British | 79.35 | 78.18 | 75.02 |
| Chinese | 82.03 | 82.25 | 78.08 |
| Danish | 81.22 | 82.08 | 77.70 |
| Dutch | 80.88 | 79.76 | 76.97 |
| Finnish | 77.26 | 77.01 | 73.56 |
| French | 82.03 | 81.71 | 77.91 |
| German | 80.14 | 80.54 | 76.58 |
| Greek | 82.55 | 81.68 | 78.51 |
| Hindi | 80.48 | 79.56 | 76.53 |
| Hungarian | 80.57 | 80.09 | 76.09 |
| Italian | 81.95 | 80.81 | 77.88 |
| Japanese | 83.96 | 82.20 | 80.21 |
| Norwegian | 82.53 | 82.23 | 78.86 |
| Polish | 82.26 | 81.39 | 77.75 |
| Portuguese | 81.94 | 80.90 | 78.04 |
| Russian | 80.86 | 79.11 | 76.02 |
| Spanish | 81.18 | 80.76 | 77.57 |
| Swedish | 79.60 | 78.10 | 75.91 |
| Average | 81.24 | 80.52 | 77.27 |
|
Average ESTOI score: 80.52 for TWELP 1200 bps, versus 81.24 for MELPe 2400 bps and 77.27 for MELPe 1200 bps. | |||
The open-source whisper.cpp ASR engine and the corresponding Whisper large-v3 model can be obtained separately from their official repositories.
The supplied tools and instructions allow the results presented above to be independently reproduced and verified.
Speech Samples (WAV files)
Independent experts compared TWELP 1200 bps with MELPe 2400 bps and 1200 bps in preference listening tests. Listening preferences were closely divided, but TWELP was preferred overall in comparisons with both MELPe bitrates, with listeners describing its speech as more natural and less synthetic.
Listen to the source speech and the corresponding MELPe 2400 bps, MELPe 1200 bps and TWELP 1200 bps samples below.
For more accurate comparison, we recommend using good-quality headphones or speakers.
The complete P.50 sample sets for all languages are also available in the Downloads section at the bottom of this page.
Speech & Beyond
Unlike many low-bitrate vocoders optimized primarily for speech, TWELP also provides improved reproduction of non-speech signals, including alert tones, police, ambulance and fire sirens, music and other audio signals.
Combined with natural speech reproduction, this makes TWELP well suited for digital radio and other communication systems where both speech and non-speech signals must be transmitted over a very low-bitrate channel.
Listen to the comparison below:
| Source signal | MELPe 2400 bps | MELPe 1200 bps | TWELP 1200 bps |
|---|---|---|---|
For this comparison, ESTOI was selected as a single objective intelligibility metric to provide a focused and consistent evaluation of speech under acoustic-noise conditions.
| Language | MELPe 2400 bps | TWELP 1200 bps | MELPe 1200 bps |
|---|---|---|---|
| American | 70.09 | 69.98 | 66.92 |
| Arabic | 72.88 | 71.11 | 69.34 |
| British | 68.02 | 66.98 | 64.32 |
| Chinese | 73.16 | 73.61 | 69.84 |
| Danish | 68.87 | 70.18 | 65.41 |
| Dutch | 68.01 | 67.25 | 64.64 |
| Finnish | 65.96 | 65.97 | 62.45 |
| French | 72.39 | 71.57 | 68.53 |
| German | 68.23 | 68.20 | 65.00 |
| Greek | 71.61 | 71.77 | 68.12 |
| Hindi | 71.61 | 67.23 | 64.62 |
| Hungarian | 71.84 | 71.27 | 68.20 |
| Italian | 70.28 | 68.67 | 66.67 |
| Japanese | 75.04 | 73.08 | 71.71 |
| Norwegian | 73.14 | 73.25 | 69.64 |
| Polish | 71.40 | 70.51 | 67.27 |
| Portuguese | 71.25 | 70.31 | 67.93 |
| Russian | 69.19 | 68.18 | 65.57 |
| Spanish | 71.55 | 72.64 | 68.39 |
| Swedish | 65.02 | 64.40 | 62.05 |
| Average | 70.33 | 69.81 | 65.77 |
|
Average ESTOI score in acoustic noise: 69.81 for TWELP 1200 bps, versus 70.33 for MELPe 2400 bps and 65.77 for MELPe 1200 bps. | |||
The samples below compare heavily noisy English speech (SNR = 10 dB) processed by MELPe 2400 bps, MELPe 1200 bps and TWELP 1200 bps, first with noise preprocessing disabled and then enabled (MELPe NPP / TWELP NCSE).
| NPP NCSE | Input speech (SNR=10dB) | MELPe 2400 bps | MELPe 1200 bps | TWELP 1200 bps |
|---|---|---|---|---|
| Disabled | ||||
| Enabled |
The NCSE integrated into the TWELP vocoder is described in more detail on the webpage for our standalone product, 'NCSE-AGC Preprocessor'.
These vocoders are based on an effective Joint Source-Channel Coding approach. Each vocoder is equipped with a custom-designed FEC, tailored to its specific characteristics and operational conditions.
TWELP Robust vocoders provide high speech quality simultaneously in noisy channel as well as in noiseless channel. FEC can operate with "soft decisions" as well as with "hard decisions" from a modem. "Soft decisions" mode provides much better robustness in comparison with the "hard decisions" mode.
For all users of our non-robust vocoder versions, we offer the following recommendations.
Essentially, the diagram shows by what percentage speech quality is reduced when a specific bit is distorted. The first bits in order cause catastrophic distortions, while the latter bits have significantly less impact on quality.
Additional Functionalities. The following additional functionalities are developed by DSPINI and integrated into TWELP vocoders:
- Noise Cancellation Speech Enhancement (NCSE)
- Automatic Gain Control (AGC),
- Voice Activity Detector (VAD),
- Discontinuous Transmission (DTX),
- Tone Detection/Generation (Single tones and Dual tones). The tones are transmitted by the vocoder facilities.
Note:
The Tone Detector/Generator functionalities are not integrated into the code by default but can be added free of charge upon request.
Each functionality has unique features, performance and characteristics, providing significant superiority over any well-known implementations on the market.
Technical Characteristics And Resource Requirements:
| Bit Rate (bps) | Algorithm | Frame size (ms) | Algorithmic delay (including frame size) (ms) | Sampling rate (kHz) | Signal format | Bit stream format |
|---|---|---|---|---|---|---|
| 1200 | TWELP | 40 | 60 | 8 | Linear 16-bit PCM |
48 |
| Name | Functionality | Technical characteristics | |
|---|---|---|---|
| Name | Value | ||
| AGC | Automatic Gain Control |
Control range: | 0 ... +42 dB |
| NCSE | Noise Canceller - Speech Enhancer |
SNR increasing | 20 dB |
| Speech quality improvement |
> 0.1 PESQ | ||
| Tone Detector |
Single/Dual tones detection |
In accordance with international standards | |
| Tone Generator |
Single/Dual tones generation |
Special generator, kept continuity of signal (phase and amplitude of signal of previous frame) |
|
| DTX | Discontinuous Transmission |
Reduces bit rate down to 110 bps in pauses between active speech regions |
|
| VAD | Voice Activity Detection |
High reliability even with pink noise at an SNR < 0 dB. | |
| CNG | Comfort Noise Generation |
Type of noise | "white" |
| Level | - 60 dB | ||
The NCSE and AGC integrated into the TWELP vocoder are described in more detail on the webpage for our standalone product, 'NCSE-AGC Preprocessor'.
| Module | MIPS* peak | Memory (KBytes) | ||||
|---|---|---|---|---|---|---|
| Program | Data | |||||
| Constants | Channel | Heap | Stack | |||
| Voice Encoder | 96.9 | 35 | 169 | 4.5 | 4.8 | 1.0 |
| NCSE | 6.4 | |||||
| AGC | 0.2 | |||||
| Voice Decoder | 14.0 | |||||
| Voice Encoder + Voice Decoder |
110.9 | |||||
| Total | 117.5 | |||||
| Module | MIPS* peak | Memory (KBytes) | ||||
|---|---|---|---|---|---|---|
| Program | Data | |||||
| Constants | Channel | Heap | Stack | |||
| Voice Encoder | 34.6 | 86 | 169 | 4.5 | 4.8 | 1.0 |
| NCSE | 2.8 | |||||
| AGC | 0.1 | |||||
| Voice Decoder | 4.0 | |||||
| Voice Encoder + Voice Decoder |
38.6 | |||||
| Total | 41.5 | |||||
| Module | MIPS* peak | Memory (KBytes) | ||||
|---|---|---|---|---|---|---|
| Program | Data | |||||
| Constants | Channel | Heap | Stack | |||
| Voice Encoder | 59.0 | 21 | 169 | 4.5 | 4.8 | 1.0 |
| NCSE | 6.7 | |||||
| AGC | 0.2 | |||||
| Voice Decoder | 10.0 | |||||
| Voice Encoder + Voice Decoder |
69 | |||||
| Total | 75.9 | |||||
* DSPINI continues optimization of the TWELP algorithm and code in order to minimize computational complexity of the vocoder.
Software Integrity and Security. DSPINI guarantees the ABSOLUTE integrity of its software, free from any undocumented features, undeclared capabilities, or hidden functions. Our customers can be assured that none of our software/code contains any secret features or functionalities concealed from the user. If necessary, we are ready to provide the source code of our software products for appropriate certification.
Moreover, our software is available in source code form—you simply need to purchase the appropriate license to use it.
Guarantee And Support. DSPINI guarantees a quality and accordance of all technical characteristics of the product to requirement of current specifications. Testing and other method of quality control are used for guarantee support.
Any Platforms. DSPINI can port this vocoder software into any other DSP, RISC or general- purposes platform inshort time: 1-2 months.
Licensing Terms. To use the vocoder, customer should obtain a license from DSPINI only.
Customization. The vocoder can be customized under any specific requirements- other bit rate, frame size, any other robustness to channel errors, etc. Please contact with us for details.
Prospects. DSPINI is impoving and developing continuously a set of new vocoders with range from 300 bps up to 9600 bps, based on TWELP technology.
Related Software. This vocoder may be effectively used in a bundle with other DSPINI's products:
- Linear and acoustic echo cancellers,
- Multichannel noise cancellers (including two-microphone adaptive array),
- Wired or radiomodems for any types of channels and bitrates,
- Other products.
- Datasheet (pdf)
- ITU-T P.50 source speech samples (zip)
- MELPe 2400 bps speech samples (zip)
- MELPe 1200 bps speech samples (zip)
- TWELP 1200 bps speech samples (zip)
- P.862 and STOI/ESTOI utilities
- PC-evaluation package (zip) — on request
- User's Guide document (pdf) — on request