Voice Frame Reconstruction via Pre-extracted Time-Frequency Parameters
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice over Internet Protocol (VoIP) systems face challenges in maintaining sound quality due to packet loss, with current solutions like Packet Loss Concealment (PLC) being limited in capability and adaptability to diverse transmission scenarios.
Innovation Solution
A voice processing method that determines historical voice frames, acquires frequency-domain and time-domain parameters, and uses these to predict and reconstruct target voice frames, enhancing the voice processing capability and supporting continuous packet loss concealment.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If PLC technology is used to compensate for lost voice frames through signal analysis, then some degree of packet loss concealment is achieved, but the processing capability is limited and not adaptive to diverse transmission scenarios especially sudden packet loss
Solution Approach 1:
The patent extracts and stores multiple characteristics (frequency domain, time domain, and combined time-frequency domain features) from received voice frames before packet loss occurs. When packet loss happens, these pre-extracted characteristics enable rapid reconstruction without real-time analysis, improving both processing capability and adaptability to sudden packet loss scenarios.
Solution Approach 2:
The patent changes the parameter representation by extracting multiple types of characteristics (frequency domain, time domain, and their correlation) from voice frames. This multi-parameter approach allows flexible adaptation to different packet loss scenarios and improves reconstruction accuracy compared to single-parameter methods.
2Reliability
If signal analysis capability is increased to improve packet loss concealment, then better voice frame reconstruction is achieved, but system complexity increases
Solution Approach 1:
The patent performs signal analysis in advance by extracting and storing multiple characteristics from received frames before packet loss occurs. This preliminary action transfers computational complexity from the reconstruction phase to the normal operation phase, enabling simple and fast reconstruction when packet loss happens while maintaining high concealment performance.
3Manufacturing precision
If multiple characteristics are extracted and analyzed to improve reconstruction accuracy, then better sound quality is achieved, but processing time increases
Solution Approach 1:
The patent extracts multiple characteristics (frequency domain, time domain, and correlation features) in advance during normal reception and stores them for later use. When packet loss occurs, the system directly uses these pre-computed characteristics for reconstruction, avoiding time-consuming real-time analysis and achieving both high accuracy and fast processing.
Data Source
AI summary
A voice processing method includes: determining a historical voice frame corresponding to a target voice frame; acquiring a frequency-domain characteristic of the historical voice frame and a time-domain parameter of the historical voice frame; obtaining a parameter set of the target voice frame according to a correlation between the frequency-domain characteristic of the historical voice frame and the time-domain parameter of the historical voice frame, the parameter set including at least two parameters; and reconstructing the target voice frame according to the parameter set.


