Evolution of Performance Metrics for Accurate Evaluation of Speech-to-Speech Translation Models: A Literature Review

Authors

  • Dr. Gabriel O. Sobola

DOI:

https://doi.org/10.34257/LJER109888UK

Keywords:

Carbon nanotubes (CNTs), Functionalization of CNTs, Chirality, Arc discharge, laser ablation, Biosensors., manufacturing., industry, Technological tasks, metalworking, visualization of the axiom control CNC system, energy storage, hydroelectric power, pumped storage, hydropower station modernisation, infrastructure refurbishment., BERTScore, Bilingual Evaluation Understudy Scores (BLEUS), BLASER, Leaderboards, Mean Opinion Score Naturalness (MOSN), Mean Opinion Score Similarity (MOSS), Recall Oriented Understudy for Gisting Evaluation Longest Common Subsequence (ROUGE-L), Speech-to-speech metrics, Word Error Rate (WER)

Abstract

The translation of speech from a source to speech in a target language with generative artificial intelligence is an area of research that is presently being actively explored. This is aimed at solving global language barriers thereby ensuring seamless communication between the individuals involved. It has been well developed for high-resourced languages like English, Spanish, French and Chinese. Currently, objective evaluation metrics such as Bilingual Evaluation Understudy Scores (BLEUS), and subjective metrics such as Mean Opinion Score Naturalness (MOSN) and Mean Opinion Score Similarity (MOSS) are being used to evaluate the performance of the output of speech-to-speech models. However, low resourced languages are still undeveloped in the area of speech processing applications, especially the African indigenous languages. The output speech in the target language needs to be evaluated to determine the closeness to the ground truth, as well as how natural and intelligible it is to the intended listeners. This paper presents a review of trends from the current metrics to emerging ones such as Recall Oriented Understudy for Gisting Evaluation-L (ROUGE-L) and BLASER. The applications of speech models’ metrics on various leaderboards and modern AI platforms were also discussed. The outcome shows that while BLEU score and MOSN metrics are prevalent for speech models, there is a need to explore metrics such as ROUGE-L, and BERTScore which are machine translation metric due to their benefits.

References

Tjandra Andros, Sakriani Sakti, Satoshi Nakamura (2019) Speech-to-speech translation between untranscribed unknown languages. 593–600.

Elizabeth Salesky, Julian Mäder, Severin Klinger (2021) Assessing evaluation metrics for speech-to-speech translation. 733–740.

Chen Zhang, Tan Xu, Ren Yi, Qin Tao, Zhang Kejun, Liu Tie-Yan (2021) Uwspeech: Speech to speech translation for unwritten languages. 14319–14327.

Ye Jia, Ron J. Weiss, Fadi Biadsy, Wolfgang Macherey, Melvin Johnson, Zhifeng Chen, Yonghui Wu (2019) Direct speech-to-speech translation with a sequence-to-sequence model. 1123-1127.

Takatomo Kano, Sakriani Sakti, Satoshi Nakamura (2021) Transformer-based direct speech-to-speech translation with transcoder. 958–965.

Qianqian Dong, Fengpeng Yue, Tom Ko, Mingxuan Wang, Qibing Bai, Yu Zhang (2022) Leveraging pseudo-labeled data to improve direct speech-to-speech translation. 1781-1785.

Kathrin Blagec, Georg Dorffner, Milad Moradi, Simon Ott, Matthias Samwald (2022) A global analysis of metrics used for measuring performance in natural language processing. 52-63.

R. Jules, Emmanuel Adetiba, Abdultaofeek Abayom, Oluwatobi E. Dare, Ayodele H. Ifijeh (2025) Speech to Speech Translation with Translatotron: A State of the Art Review. arXiv preprint arXiv:2502.05980

Chin-Yew Lin, Franz Josef Och (2004) Orange: a method for evaluating automatic evaluation metrics for machine translation. 501–507.

Hirofumi Inaguma, Sravya Popuri, Ilia Kulikov, Peng-Jen Chen, Changhan Wang, Yu-An Chung, Yun Tang, Ann Lee, Shinji Watanabe, Juan Pino (2023) Unity: Two-pass direct speech-to-speech translation with discrete units.

Daniel Licht, Cynthia Gao, Janice Lam, Francisco Guzman, Mona Diab, Philipp Koehn (2022) Consistent human evaluation of machine translation across language pairs. 309-321.

Loïc Barrault, Yu-An Chung, Mariano Cora Meglioli, David Dale, Ning Dong, Paul-Ambroise Duquenne, Hady Elsahar (2023) SeamlessM4T-Massively Multilingual & Multimodal Machine Translation. https://arxiv.org/abs/2308.11596

David Dale, Marta Costa-jussà (2024) Blaser 2.0: a metric for evaluation and quality estimation of massively multilingual speech and text translation. 16075–16085.

Mingda Chen, Paul-Ambroise Duquenne, Pierre Andrews, Justine Kao, Alexandre Mourachko, Holger Schwenk, Marta R. Costa-jussà (2023) BLASER: A Text-Free Speech-to-Speech Translation Evaluation Metric.

Zairan Gong, Xiaona Xu, Yue Zhao (2025) Tibetan–Chinese speech-to-speech translation based on discrete units. 15(1), 2592.

Eduardo Farófia Medeiros (2023) Deep learning for speech to text transcription for the Portuguese language.

Mishaim Malik, Muhammad Kamran Malik, Khawar Mehmood, Imran Makhdoom (2021) Automatic speech recognition: a survey. 80(6), 9411–9457. https://doi.org/10.1007/s11042-020-10073-7

Max Nelson, Shannon Wotherspoon, Francis Keith, William Hartmann, Matthew Snover (2024) Cross-Lingual Conversational Speech Summarization with Large Language Models. https://arxiv.org/abs/2408.06484

Alejandro Manuel Pérez González de Martos (2022) Deep neural networks for automatic speech-to-speech translation of open educational resources.

Meinard Müller (2007) Dynamic time warping. 69–84.

Ndolu Fajar Henri Erasmus, Ruki Harwahyu (2023) Intrusion Detection System on Nowadays’ Attack using Ensemble Learning. 10(1), 42–50.

Okokpujie Kennedy, Daniella Kalimumbalo, Joke A. Badejo, Emmanuel Adetiba (2022) Congestion Intrusion Detection-based Method for Controller Area Network Bus: A case for KIA SOUL Vehicle. 5.

Hujon, V. Aiusha, Thoudam Doren Singh, Khwairakpam Amitab (2022) Transfer Learning Based Neural Machine Translation of English-Khasi on Low-Resource Settings. 1–8. https://doi.org/10.1016/j.procs.2022.12.396

Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, Ilya Sutskever (2023) Robust speech recognition via large-scale weak supervision. 28492–28518.

Khurana Sameer, Nauman Dawalatabad, Antoine Laurent, Luis Vicente, Pablo Gimeno, Victoria Mingote, James Glass (2023) Improved cross-lingual transfer learning for automatic speech translation. http://arxiv.org/abs/2306.00789

Rubenstein K. Paul, Chulayuth Asawaroengchai, Duc Dung Nguyen, Ankur Bapna, Zalán Borsos, Félix de Chaumont Quitry, Peter Chen (2023) Audiopalm: A large language model that can speak and listen. http://arxiv.org/abs/2306.12925

Jia Ye, Michelle Tadmor Ramanovich, Tal Remez, Roi Pomerantz (2022) Translatotron 2: High-quality direct speech-to-speech translation with voice preservation. 10120-10134.

(2025) Joint speech and text machine translation for up to 100 languages. 637(8046), 587–593.

Nguyen Luan Thanh, Sakriani Sakti (2025) ZeST: A Zero-resourced Speech-to-Speech Translation Approach for Unknown, Unpaired, and Untranscribed Languages.

Shi Jiatong, Yun Tang, Ann Lee, Hirofumi Inaguma, Changhan Wang, Juan Pino, Shinji Watanabe (2023) Enhancing speech-to-speech translation with multiple TTS targets. 1–5.

Anuj Diwan, Anirudh Srinivasan, David Harwath, Eunsol Choi (2023) Textless low-resource speech-to-speech translation with unit language models. arXiv–2305. http://arxiv.org/abs/2305.17732

Kun Song, Ren Yi, Lei Yi, Chunfeng Wang, Wei Kun, Xie Lei, Yin Xiang, Ma Zejun (2023) StyleS2ST: Zero-shot Style Transfer for Direct Speech-to-speech Translation. http://arxiv.org/abs/2305.17732

Ogayo Perez, Graham Neubig, Alan W. Black (2022) Building African Voices. http://arxiv.org/abs/2207.00688

Tolulope Ogunremi, Kola Tubosun, Aremu Anuoluwapo, Orife Iroro, Adelani David Ifeoluwa (2023) Iroyin Speech: A multi-purpose Yoruba Speech Corpus. http://arxiv.org/abs/2307.16071

Alexander Gutkin, Isin Demirsahin, Oddur Kjartansson, Clara Rivera, Kólá Túbòsún (2020) Developing an Open-Source Corpus of Yoruba Speech. 404-408.

Josh Meyer, David Ifeoluwa Adelani, Edresson Casanova, Alp Öktem, Daniel Whitenack, Julian Weber, Salomon Kabongo, Elizabeth Salesky (2022) BibleTTS: a large, high-fidelity, multilingual, and uniquely African speech corpus. 2383-2387.

Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, Yoav Artzi (2019) BERTScore: Evaluating Text Generation with BERT.

Josh Meyer, David Ifeoluwa Adelani, Edresson Casanova, Alp Öktem, Daniel Whitenack Julian Weber, Salomon Kabongo, Elizabeth Salesky et al. (2022) Findings of the first WMT shared task on sign language translation (WMT-SLT22). 744–77.

Deepanjali Singh, Ayush Anand, Abhyuday Chaturvedi, Niyati Baliyan (2024) IWSLT 2024 Indic Track system description paper: Speech-to-Text Translation from English to multiple Low-Resource Indian Languages. 311–316.

U. R. Pol, P. S. Vadar, T. T. Moharekar Hugging Face: Revolutionizing AI and NLP.

Y. Shen, K. Song, X. Tan, D. Li, W. Lu, Y. Zhuang (2023) Hugginggpt: Solving ai tasks with chatgpt and its friends in hugging face. 36, 38154–38180.

Data Science Dojo (2025) Top 5 LLM Leaderboard Platforms for AI Excellence. https://datasciencedojo.com/blog/understanding-llm-leaderboards/

(2025) Speech to Speech Models and Providers Analysis | Artificial Analysis. https://artificialanalysis.ai/models/speech-to-speech

(2025) Speech Generation Evaluation and Leaderboard. https://balacoon.com/blog/tts_leaderboard/

(2025) TTS Arena Legacy. https://huggingface.co/spaces/TTS-AGI/TTS-Arena

(2025) TTS Spaces Arena - a Hugging Face Space by Pendrokar. https://huggingface.co/spaces/Pendrokar/TTS-Spaces-Arena

(2025) Speech generation | Leaderboard. https://labelbox.com/leaderboards/speech-generation/?utm_source=chatgpt.com

(2025) Open ASR Leaderboard. https://huggingface.co/spaces/hf-audio/open_asr_leaderboard?utm_source=chatgpt.com

(2025) TTS Arena V2. https://huggingface.co/spaces/TTS-AGI/TTS-Arena-V2

Y. Hu (2024) GenTranslate: Large language models are generative multilingual speech and machine translators. https://arxiv.org/abs/2402.06894

S. Gandhi, P. Von Platen, A. M. Rush (2022) Esb: A benchmark for multi-domain end-to-end speech recognition. https://arxiv.org/abs/2210.13352

N. K. Nallabala, B. Souprayen, M. Ramasamy, S. K. Penumarti, N. C. Navuluri (2025) An Efficient Speech Synthesizer: A Hybrid Monotonic Architecture for Text-to-speech via VAE & LPC-Net with Independent Sentence Length. 1–25.

L. Santos, N. de Araújo Moreira, R. Sampaio, R. Lima, F. C. M. B. Oliveira (2025) Automatic Speech Recognition: Comparisons Between Convolutional Neural Networks, Hidden Markov Model and Hybrid Architecture. 42(5), e70032.

S. Sultoni, B. Darmawan (2025) Speaker Recognition System Using MFCC and HMM Methods. 7(1), 206–218.

(2024) Perceptual evaluation of speech quality (PESQ): An objective method for end-to-end speech quality assessment of narrow-band telephone networks and speech codecs. https://www.itu.int/rec/T-REC-P.862/en

(2010) A short-time objective intelligibility measure for time-frequency weighted noisy speech. https://doi.org/10.1109/ICASSP.2010.5495701

N. Zheng, X. Wan, K. Liu, Z. Huan (2025) SCDiar: a streaming diarization system based on speaker change detection and speech recognition. https://arxiv.org/abs/2501.16641

M. Tran, et al. (2025) A Domain Adaptation Framework for Speech Recognition Systems with Only Synthetic data. https://arxiv.org/abs/2501.12501

M. Torcoli, M. M. Halimeh, E. A. P. Habets (2025) Navigating PESQ: Up-to-Date Versions and Open Implementations. https://arxiv.org/abs/2505.19760

(2024) ISCA Archive. https://www.isca-archive.org/iberspeech_2024/index.html

Cutler Ross, Ando Saabas, Tanel Pärnamaa, Marju Purin, Evgenii Indenbom, Nicolae-Cătălin Ristea, Jegor Gužvin, Hannes Gamper, Sebastian Braun, Robert Aichner (2024) ICASSP 2023 acoustic echo cancellation challenge.

Evolution of Performance Metrics for Accurate Evaluation of Speech-to-Speech Translation Models: A Literature Review

Downloads

Published

2025-11-07

How to Cite

Evolution of Performance Metrics for Accurate Evaluation of Speech-to-Speech Translation Models: A Literature Review. (2025). London Journal of Engineering Research, 25(4), 65-91. https://doi.org/10.34257/LJER109888UK