Voice Quality Evaluation via Frame Loss Event Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice quality evaluation methods in VoIP systems are limited in precision due to their reliance on average distortion information, failing to accurately account for silence periods and varying frame types, leading to inaccurate prediction and evaluation of voice quality.

Innovation Solution

A method that involves parsing voice data packets to identify silence and voice frames, dividing the voice sequence into statements based on frame content characteristics, and extracting non-voice parameters to evaluate voice quality using a preset model, considering frame loss events and their impact on voice quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If average distortion information is used to evaluate voice quality, then the evaluation process is simple, but the prediction precision and accuracy are insufficient

Engineering Contradiction:
Improveevaluation process complexityVSAvoidvoice quality prediction precision
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent segments the voice sequence into multiple statements based on frame content characteristics (silence frames vs. voice frames). Each statement is then divided into frame loss events, allowing detailed analysis of specific loss patterns rather than using a single average distortion value. This segmentation enables precise tracking of where and how packet loss affects voice quality.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality analysis by evaluating different types of frames (silence vs. voice) and different positions in the voice sequence differently. Frame loss events are characterized by location parameters and discrete distribution parameters that capture local impact. This allows the system to weigh the importance of losses in different regions, recognizing that not all packet losses have equal impact on perceived voice quality.

Inventive Principle:
Principle #3Local quality

2Device complexity

If silence frames are not considered in voice quality evaluation, then the evaluation method is simpler, but the accuracy of voice quality assessment deteriorates

Engineering Contradiction:
Improveevaluation method complexityVSAvoidvoice quality assessment accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent explicitly segments frames into silence frames and voice frames based on frame content characteristics. This segmentation allows the system to differentiate between periods when voice quality matters (voice frames) and periods when it less matters (silence frames). The statement division process uses these distinctions to create meaningful evaluation units.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different evaluation criteria to silence frames versus voice frames. By identifying frame content characteristics, the system can assign different weights or evaluation methods to different frame types. This local differentiation improves accuracy by recognizing that packet loss during silence periods has different impact compared to loss during active speech.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS10284712B2Voice quality evaluation method, apparatus, and system
Publication Date: 2019.05.07 HUAWEI TECH CO LTD
  • US10284712B2 patent drawing
  • US10284712B2 patent drawing
  • US10284712B2 patent drawing

AI summary

A voice quality evaluation method, apparatus, and system comprises an obtained voice data packet is parsed, and a frame content characteristic of the data packet is determined according to a parse result, for example, the frame content characteristic is a silence frame and a voice frame. Then, a voice sequence is divided into statements according to the determined frame content characteristic, and the statements are divided into multiple frame loss events; after non-voice parameters are extracted according to the frame loss events, voice quality of each statement is evaluated according to a preset voice quality evaluation model and according to the non-voice parameters. Finally, voice quality of the entire voice sequence is evaluated according to the voice quality of each statement. By using this solution, prediction precision can be improved significantly, and accuracy of an evaluation result can be improved.