Single-Ended Speech Quality Evaluation Using Deep Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech quality evaluation algorithms, such as POLQA, require both input and output speech data, limiting their application as it is difficult to obtain input speech data in certain conditions like residential areas or indoor settings, restricting the scope of speech quality evaluation.

Innovation Solution

A method and device using Deep Learning to determine a speech quality evaluation model based solely on single-ended speech data, allowing for the extraction of evaluation features and computation of quality scores without relying on input speech data, thereby expanding the scope of applications.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional speech quality evaluation algorithms (PESQ, POLQA) are used, then speech quality can be evaluated with input and output speech data, but the application scope is limited because input speech data cannot be obtained in residential areas, malls or other indoor conditions

Engineering Contradiction:
Improveapplication scopeVSAvoiddata collection requirement
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent extracts the essential quality evaluation function from the traditional two-ended evaluation system and creates a single-ended evaluation model that works with only output speech data. This removes the requirement for input speech data collection, enabling deployment in scenarios where only one end of the communication is accessible.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces an intermediary neural network model trained to predict quality metrics from single-ended speech data. This intermediary model bridges the gap between the available output speech data and the desired quality evaluation, eliminating the need for direct access to input speech data.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If single-ended speech data is used for evaluation, then the scope of application is expanded, but traditional algorithms cannot be directly applied

Engineering Contradiction:
Improveapplication scopeVSAvoidevaluation accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent performs preliminary training of a neural network model using paired speech data (clean and degraded) to learn the relationship between speech characteristics and quality metrics. This pre-trained model can then accurately predict quality from single-ended data without requiring input speech during actual evaluation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent replaces the traditional mechanical comparison-based evaluation system (requiring direct comparison of input and output speech) with an intelligent system based on neural networks that can infer quality from partial information, achieving similar or better performance with fewer requirements.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentEP3528250B1Voice quality evaluation method and apparatus
Publication Date: 2022.05.25 IFLYTEK CO LTD
  • EP3528250B1 patent drawingFigure 1~2
  • EP3528250B1 patent drawingFigure 3~4

AI summary

Proposed are a voice quality evaluation method and apparatus. The voice quality evaluation method comprises: receiving voice data to be evaluated; extracting an evaluation characteristic of the voice data to be evaluated; and according to the evaluation characteristic of the voice data to be evaluated and a constructed voice quality evaluation model, performing quality evaluation on the voice data to be evaluated, wherein the voice quality evaluation model is used for indicating a relationship between the evaluation characteristic of single-ended voice data and quality information about the single-ended voice data. The method can expand the application scope of voice quality evaluation.