Voice Quality Evaluation via Frame Loss Event Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice quality evaluation methods in VoIP systems are limited in precision due to their reliance on average distortion information, failing to accurately account for silence periods and varying frame types, leading to inaccurate prediction and evaluation of voice quality.
Innovation Solution
A method that involves parsing voice data packets to identify silence and voice frames, dividing the voice sequence into statements based on frame content characteristics, and extracting non-voice parameters to evaluate voice quality using a preset model, considering frame loss events and their impact on voice quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If average distortion information is used to evaluate voice quality, then the evaluation process is simple, but the prediction precision and accuracy are insufficient
Solution Approach 1:
The patent segments the voice sequence into multiple statements based on frame content characteristics (silence frames vs. voice frames). Each statement is then divided into frame loss events, allowing detailed analysis of specific loss patterns rather than using a single average distortion value. This segmentation enables precise tracking of where and how packet loss affects voice quality.
Solution Approach 2:
The patent applies local quality analysis by evaluating different types of frames (silence vs. voice) and different positions in the voice sequence differently. Frame loss events are characterized by location parameters and discrete distribution parameters that capture local impact. This allows the system to weigh the importance of losses in different regions, recognizing that not all packet losses have equal impact on perceived voice quality.
2Device complexity
If silence frames are not considered in voice quality evaluation, then the evaluation method is simpler, but the accuracy of voice quality assessment deteriorates
Solution Approach 1:
The patent explicitly segments frames into silence frames and voice frames based on frame content characteristics. This segmentation allows the system to differentiate between periods when voice quality matters (voice frames) and periods when it less matters (silence frames). The statement division process uses these distinctions to create meaningful evaluation units.
Solution Approach 2:
The patent applies different evaluation criteria to silence frames versus voice frames. By identifying frame content characteristics, the system can assign different weights or evaluation methods to different frame types. This local differentiation improves accuracy by recognizing that packet loss during silence periods has different impact compared to loss during active speech.
Data Source
AI summary
A voice quality evaluation method, apparatus, and system comprises an obtained voice data packet is parsed, and a frame content characteristic of the data packet is determined according to a parse result, for example, the frame content characteristic is a silence frame and a voice frame. Then, a voice sequence is divided into statements according to the determined frame content characteristic, and the statements are divided into multiple frame loss events; after non-voice parameters are extracted according to the frame loss events, voice quality of each statement is evaluated according to a preset voice quality evaluation model and according to the non-voice parameters. Finally, voice quality of the entire voice sequence is evaluated according to the voice quality of each statement. By using this solution, prediction precision can be improved significantly, and accuracy of an evaluation result can be improved.


