Voice Signal Decoding with Style Fusion at Low Bit Rates

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voice coding algorithms at low bit rates suffer from limited quality in decoded voice signals due to the limitations of DNN-based systems, particularly in neural vocoder-based systems where only decoder parameters can be adjusted, and neural codec systems face poor performance with low bit rate voice outputs.

Innovation Solution

A voice signal decoding method that involves obtaining acoustic and excitation features, performing style fusion processing on these features using a generator, and reconstructing a decoded voice signal to improve quality.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If a neural vocoder-based system is used with low bit rate encoding, then the bit rate is reduced to as low as 1.6 kbps, but the quality of decoded voice is limited because only decoder parameters can be adjusted through big data training

Engineering Contradiction:
Improvebit rateVSAvoiddecoded voice quality
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent segments the voice signal processing into two independent feature streams: excitation features (obtained via excitation network) and style features (extracted from acoustic features). This segmentation allows each feature type to be processed and optimized independently, enabling quality improvement at low bit rates by preserving style information separately from the compressed acoustic features.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent performs preliminary extraction of style features from the acoustic feature encoding results before the final voice synthesis. By pre-extracting and preserving style information (such as tone, pitch, and voice characteristics) before the low bit rate compression fully takes effect, the system ensures that these critical quality-determining features are available for reconstruction even when the overall bit rate is reduced to 1.6 kbps.

Inventive Principle:
Principle #10Preliminary action

2Adaptability or versatility

If an end-to-end neural codec system is used with DNN for both encoder and decoder, then feature extraction flexibility is improved, but the quality of decoded voice signal deteriorates when voice output has low bit rate

Engineering Contradiction:
Improvefeature extraction flexibilityVSAvoiddecoded voice quality at low bit rate
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent applies local quality by treating different feature types differently: excitation features are processed through a dedicated excitation network while style features are extracted and preserved separately from the main acoustic feature stream. This localized processing ensures that style-critical features maintain higher quality even when the overall system operates at low bit rates, addressing the quality deterioration problem in end-to-end neural codec systems.

Inventive Principle:
Principle #3Local quality

3Manufacturing precision

If style fusion processing is performed on style feature and excitation feature, then the quality of decoded voice signal is improved by restoring voice style, but the device complexity increases

Engineering Contradiction:
Improvedecoded voice qualityVSAvoiddecoding system complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent merges the style feature stream and excitation feature stream at the generator input stage. By combining these two feature types through style fusion processing (such as concatenation or element-wise operations) before feeding them into the voice synthesis generator, the system restores voice style information without requiring separate processing pipelines, thus improving decoded voice quality while controlling the increase in device complexity.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentEP4693280A1Voice signal decoding method and apparatus, and electronic device
Publication Date: 2026.02.11 HUAWEI TECH CO LTD
  • EP4693280A1 patent drawingFigure 1a
  • EP4693280A1 patent drawingFigure 1b(1)~2
  • EP4693280A1 patent drawingFigure 3

AI summary

This application provides a voice signal decoding method and apparatus and an electronic device, which are applied to the field of audio coding technologies. In the method, an encoding apparatus encodes an original voice signal, to obtain an encoded bitstream. The encoded bitstream includes an acoustic feature encoding result. A decoding apparatus obtains the acoustic feature encoding result in the encoded bitstream, obtains a style feature from the acoustic feature encoding result, obtains an excitation feature, performs style fusion processing on the excitation feature and the style feature, to obtain a fused voice feature, and reconstructs a decoded voice signal based on the voice feature. Because the style feature indicates a voice style of the original voice signal, a voice feature in the original voice signal can be restored from the voice feature obtained by performing style fusion on the excitation feature and the style feature. In this way, quality of the decoded voice signal reconstructed based on the fused voice feature can be effectively improved in comparison with that of a decoded voice signal obtained without style fusion.