Single-Ended Speech Quality Evaluation Using Deep Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech quality evaluation algorithms, such as POLQA, require both input and output speech data, limiting their application as it is difficult to obtain input speech data in certain conditions like residential areas or indoor settings, restricting the scope of speech quality evaluation.
Innovation Solution
A method and device using Deep Learning to determine a speech quality evaluation model based solely on single-ended speech data, allowing for the extraction of evaluation features and computation of quality scores without relying on input speech data, thereby expanding the scope of applications.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional speech quality evaluation algorithms (PESQ, POLQA) are used, then speech quality can be evaluated with input and output speech data, but the application scope is limited because input speech data cannot be obtained in residential areas, malls or other indoor conditions
Solution Approach 1:
The patent extracts the essential quality evaluation function from the traditional two-ended evaluation system and creates a single-ended evaluation model that works with only output speech data. This removes the requirement for input speech data collection, enabling deployment in scenarios where only one end of the communication is accessible.
Solution Approach 2:
The patent introduces an intermediary neural network model trained to predict quality metrics from single-ended speech data. This intermediary model bridges the gap between the available output speech data and the desired quality evaluation, eliminating the need for direct access to input speech data.
2Adaptability or versatility
If single-ended speech data is used for evaluation, then the scope of application is expanded, but traditional algorithms cannot be directly applied
Solution Approach 1:
The patent performs preliminary training of a neural network model using paired speech data (clean and degraded) to learn the relationship between speech characteristics and quality metrics. This pre-trained model can then accurately predict quality from single-ended data without requiring input speech during actual evaluation.
Solution Approach 2:
The patent replaces the traditional mechanical comparison-based evaluation system (requiring direct comparison of input and output speech) with an intelligent system based on neural networks that can infer quality from partial information, achieving similar or better performance with fewer requirements.
Data Source
Figure 1~2
Figure 3~4
AI summary
Proposed are a voice quality evaluation method and apparatus. The voice quality evaluation method comprises: receiving voice data to be evaluated; extracting an evaluation characteristic of the voice data to be evaluated; and according to the evaluation characteristic of the voice data to be evaluated and a constructed voice quality evaluation model, performing quality evaluation on the voice data to be evaluated, wherein the voice quality evaluation model is used for indicating a relationship between the evaluation characteristic of single-ended voice data and quality information about the single-ended voice data. The method can expand the application scope of voice quality evaluation.