Disentangled Audio Coding with Style Control for Speech Privacy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing neural speech coding technologies face inefficiencies in coding efficiency and lack additional functionality, particularly in separating and transmitting speaker identity and phonetic content effectively.

Innovation Solution

The encoder is subdivided into a style encoder and a content encoder to disentangle long-term and short-term information, allowing separate transmission and manipulation of these streams, enabling higher coding efficiency and privacy features like speaker anonymization.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If end-to-end neural speech coding is used, then speech quality is improved, but coding efficiency deteriorates

Engineering Contradiction:
Improvespeech qualityVSAvoidcoding efficiency
Core Design Contradiction:
Manufacturing precisionVSLoss of information

Solution Approach 1:

The patent segments the speech signal into two distinct components: style information (long-term characteristics like speaker identity and speaking style) and content information (short-term characteristics like phonetic content). By encoding these components separately through dedicated style encoder and content encoder, the system achieves both high speech quality and improved coding efficiency, resolving the contradiction between quality and efficiency

Inventive Principle:
Principle #1Segmentation

2Manufacturing precision

If style information is transmitted frequently, then speech quality is improved, but data transmission rate increases

Engineering Contradiction:
Improvespeech qualityVSAvoiddata transmission rate
Core Design Contradiction:
Manufacturing precisionVSQuantity of substance

Solution Approach 1:

The patent applies periodic action by transmitting style information at a lower frequency (long-term updates) compared to content information (short-term updates). Since style characteristics change slowly over time, frequent retransmission is unnecessary. This periodic transmission strategy maintains speech quality while significantly reducing the overall data transmission rate

Inventive Principle:
Principle #19Periodic action

3Device complexity

If single data stream encoding is used, then device complexity is reduced, but adaptability deteriorates

Engineering Contradiction:
Improveencoder structureVSAvoidtransmission channel adaptability
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent segments the encoding system into separate style encoder and content encoder, each producing independent data streams. This segmentation enables different transmission strategies for different stream types (e.g., error protection levels, transmission frequencies, channel selections), providing high adaptability to various transmission conditions while maintaining manageable device complexity through modular architecture

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces dynamic adaptability by allowing the transmission system to selectively adjust the frequency and method of transmitting style versus content information based on channel conditions. The system can dynamically allocate resources, sending style information less frequently when bandwidth is limited while maintaining content information transmission for immediate reconstruction needs

Inventive Principle:
Principle #15Dynamics

Data Source

PatentEP4600951A1Disentangled audio coding and decoding with style control
Publication Date: 2025.08.13 FRAUNHOFER GESELLSCHAFT ZUR FORDERUNG DER ANGEWANDTEN FORSCHUNG EV
  • EP4600951A1 patent drawingFigure 1~2
  • EP4600951A1 patent drawingFigure 3a~3b
  • EP4600951A1 patent drawing

AI summary

Encoder for encoding an input audio signal comprising: a style encoder (10s) configured to encode style information of the input audio signal to obtain a first data stream (Q1); a content encoder (10c) configured to encode content information to obtain a second data stream (Q2)., the style information is manipulated in response to a user command. Decoder for decoding an encoded input audio signal comprising a conditioning entity (20c) configured to output a control signal being a default control signal and representing a style information, a content decoder (20d) configured to decoded a data stream (Q2) of the encoded input audio signal comprising content information, wherein the content decoder is controlled and/or adapted by the control signal.