AI Voice Quality Assessment via Speech Recognition Correlation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Wireless network operators face challenges in accurately assessing end-user perceived voice quality due to technical, privacy, and legal limitations, which hinders their ability to optimize voice call performance, especially in environments with poor signal propagation or congestion.
Innovation Solution
The implementation of an AI-assisted voice quality assessment system that uses artificial intelligence and automatic speech recognition to generate a mean recognition score by correlating network operational data with key performance indicators, enabling the estimation of voice call quality scores without the need for cumbersome human evaluation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If network operators use traditional methods to assess voice quality, then they can obtain some quality metrics, but they cannot accurately capture end-user perceived voice quality due to technical, privacy, and legal limitations
Solution Approach 1:
The patent introduces an intermediary system consisting of automated call agents and speech recognition technology that mediates between the network operators and end users. This intermediary captures voice calls, performs automated speech recognition, and analyzes the recognized text to determine voice quality metrics, thereby bypassing the limitations of direct network monitoring while maintaining accuracy.
Solution Approach 2:
The patent replaces manual human evaluation of voice quality with automated speech recognition and text analysis systems. The automated call agents conduct structured conversations, the speech recognition system converts spoken words to text, and algorithms automatically analyze the text for quality metrics, eliminating the need for cumbersome human listeners while improving measurement precision.
2Measurement precision
If network operators manually evaluate voice quality, then they can obtain detailed quality assessments, but the process becomes cumbersome and time-consuming
Solution Approach 1:
The system enables self-service automated quality assessment where the call system itself performs the evaluation. Automated call agents conduct the calls, speech recognition systems transcribe the conversations, and algorithms automatically analyze the text to generate quality metrics, allowing the system to evaluate its own performance without external human intervention.
Solution Approach 2:
The patent substitutes manual human evaluation processes with automated computational systems. The automated call management system, speech recognition engines, and text analysis algorithms work together to rapidly process and evaluate voice quality, reducing evaluation time from minutes or hours to seconds while maintaining or improving assessment detail.
3Reliability
If network operators monitor network element operational statistics, then they can obtain objective network data, but these statistics do not fully or accurately reflect end-user perceived voice quality
Solution Approach 1:
The patent introduces an intermediary layer of automated speech recognition and text analysis that bridges the gap between network operational data and end-user perception. This intermediary directly captures what users actually experience by analyzing their spoken words and responses, creating a reliable correlation between objective measurements and subjective quality perception.
Solution Approach 2:
The system combines multiple functions into a unified quality assessment approach: it captures network operational data, conducts automated voice calls, performs speech recognition, analyzes text for quality metrics, and correlates all these data sources. This multi-functional system provides both objective network statistics and accurate end-user quality correlation simultaneously.
Data Source
AI summary
A method, a device, and a non-transitory storage medium for estimating voice call quality include performing automatic speech recognition, for each of a plurality of voice calls, to generate recognized text for both an originating device acoustic signal and a receiving device acoustic signal. The recognized text for both the originating device acoustic signal and the receiving device acoustic signal are compared to the reference text to identified recognition errors and a voice call quality score for each of the originating device acoustic signal and the receiving device acoustic signal are determined. A correlation between the network conditions and the voice call quality scores is then determined.


