Voice Frame Reconstruction via Pre-extracted Time-Frequency Parameters

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing voice over Internet Protocol (VoIP) systems face challenges in maintaining sound quality due to packet loss, with current solutions like Packet Loss Concealment (PLC) being limited in capability and adaptability to diverse transmission scenarios.

Innovation Solution

A voice processing method that determines historical voice frames, acquires frequency-domain and time-domain parameters, and uses these to predict and reconstruct target voice frames, enhancing the voice processing capability and supporting continuous packet loss concealment.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If PLC technology is used to compensate for lost voice frames through signal analysis, then some degree of packet loss concealment is achieved, but the processing capability is limited and not adaptive to diverse transmission scenarios especially sudden packet loss

Engineering Contradiction:
Improvevoice processing capabilityVSAvoidadaptability to diverse transmission scenarios
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent extracts and stores multiple characteristics (frequency domain, time domain, and combined time-frequency domain features) from received voice frames before packet loss occurs. When packet loss happens, these pre-extracted characteristics enable rapid reconstruction without real-time analysis, improving both processing capability and adaptability to sudden packet loss scenarios.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the parameter representation by extracting multiple types of characteristics (frequency domain, time domain, and their correlation) from voice frames. This multi-parameter approach allows flexible adaptation to different packet loss scenarios and improves reconstruction accuracy compared to single-parameter methods.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If signal analysis capability is increased to improve packet loss concealment, then better voice frame reconstruction is achieved, but system complexity increases

Engineering Contradiction:
Improvepacket loss concealment performanceVSAvoidsignal analysis capability
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent performs signal analysis in advance by extracting and storing multiple characteristics from received frames before packet loss occurs. This preliminary action transfers computational complexity from the reconstruction phase to the normal operation phase, enabling simple and fast reconstruction when packet loss happens while maintaining high concealment performance.

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If multiple characteristics are extracted and analyzed to improve reconstruction accuracy, then better sound quality is achieved, but processing time increases

Engineering Contradiction:
Improvereconstruction accuracyVSAvoidprocessing time
Core Design Contradiction:
Manufacturing precisionVSLoss of time

Solution Approach 1:

The patent extracts multiple characteristics (frequency domain, time domain, and correlation features) in advance during normal reception and stores them for later use. When packet loss occurs, the system directly uses these pre-computed characteristics for reconstruction, avoiding time-consuming real-time analysis and achieving both high accuracy and fast processing.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12223972B2Voice processing method and apparatus, electronic device, and computer-readable storage medium
Publication Date: 2025.02.11 TENCENT TECHNOLOGY (SHENZHEN) CO LTD
  • US12223972B2 patent drawing
  • US12223972B2 patent drawing
  • US12223972B2 patent drawing

AI summary

A voice processing method includes: determining a historical voice frame corresponding to a target voice frame; acquiring a frequency-domain characteristic of the historical voice frame and a time-domain parameter of the historical voice frame; obtaining a parameter set of the target voice frame according to a correlation between the frequency-domain characteristic of the historical voice frame and the time-domain parameter of the historical voice frame, the parameter set including at least two parameters; and reconstructing the target voice frame according to the parameter set.