Disentangled Audio Coding with Style Control for Speech Privacy
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing neural speech coding technologies face inefficiencies in coding efficiency and lack additional functionality, particularly in separating and transmitting speaker identity and phonetic content effectively.
Innovation Solution
The encoder is subdivided into a style encoder and a content encoder to disentangle long-term and short-term information, allowing separate transmission and manipulation of these streams, enabling higher coding efficiency and privacy features like speaker anonymization.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If end-to-end neural speech coding is used, then speech quality is improved, but coding efficiency deteriorates
Solution Approach 1:
The patent segments the speech signal into two distinct components: style information (long-term characteristics like speaker identity and speaking style) and content information (short-term characteristics like phonetic content). By encoding these components separately through dedicated style encoder and content encoder, the system achieves both high speech quality and improved coding efficiency, resolving the contradiction between quality and efficiency
2Manufacturing precision
If style information is transmitted frequently, then speech quality is improved, but data transmission rate increases
Solution Approach 1:
The patent applies periodic action by transmitting style information at a lower frequency (long-term updates) compared to content information (short-term updates). Since style characteristics change slowly over time, frequent retransmission is unnecessary. This periodic transmission strategy maintains speech quality while significantly reducing the overall data transmission rate
3Device complexity
If single data stream encoding is used, then device complexity is reduced, but adaptability deteriorates
Solution Approach 1:
The patent segments the encoding system into separate style encoder and content encoder, each producing independent data streams. This segmentation enables different transmission strategies for different stream types (e.g., error protection levels, transmission frequencies, channel selections), providing high adaptability to various transmission conditions while maintaining manageable device complexity through modular architecture
Solution Approach 2:
The patent introduces dynamic adaptability by allowing the transmission system to selectively adjust the frequency and method of transmitting style versus content information based on channel conditions. The system can dynamically allocate resources, sending style information less frequently when bandwidth is limited while maintaining content information transmission for immediate reconstruction needs
Data Source
Figure 1~2
Figure 3a~3b
AI summary
Encoder for encoding an input audio signal comprising: a style encoder (10s) configured to encode style information of the input audio signal to obtain a first data stream (Q1); a content encoder (10c) configured to encode content information to obtain a second data stream (Q2)., the style information is manipulated in response to a user command. Decoder for decoding an encoded input audio signal comprising a conditioning entity (20c) configured to output a control signal being a default control signal and representing a style information, a content decoder (20d) configured to decoded a data stream (Q2) of the encoded input audio signal comprising content information, wherein the content decoder is controlled and/or adapted by the control signal.