Speech Coding Apparatus for DTX Interoperability and Flexible Decoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech coding apparatuses with DTX control restrict decoding modes, leading to inefficient information reduction in inactive speech sections and compatibility issues with decoding apparatuses, limiting flexibility and service selection in speech decoding.
Innovation Solution
A speech coding apparatus that generates and synthesizes coded data for active and inactive speech sections, allowing decoding sides to select appropriate modes and embed inactive speech parameters within coded data, enabling flexible decoding and efficient transmission.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If DTX control is implemented in speech coding, then transmission efficiency is improved by reducing information in inactive speech sections, but decoding flexibility is reduced as decoding mode is restricted to match coding mode
Solution Approach 1:
The patent applies dynamics by making the coded data structure adaptable - it can dynamically switch between containing only speech section information (for DTX-compatible decoders) and containing both speech and noise section information (for non-DTX decoders). This dynamic adaptability resolves the contradiction by allowing the same coding apparatus to serve both DTX and non-DTX decoding scenarios.
Solution Approach 2:
The invention achieves universality by creating a coded data format that serves multiple functions: it works with both DTX-controlled and non-DTX-controlled decoding apparatuses. The coded data can universally represent either only speech sections or both speech and noise sections depending on the decoder type, making the system universally compatible across different decoder implementations.
2Loss of information
If coded data format is optimized for DTX control, then information amount is reduced for inactive speech sections, but compatibility with non-DTX decoding apparatuses is lost
Solution Approach 1:
The patent applies parameter changes by modifying the coded data structure based on the decoder type. For non-DTX decoders, the coded data includes parameters for both speech and noise sections. For DTX decoders, it includes only speech section parameters. This parameter adaptation allows the system to maintain compatibility across different decoder types while optimizing information transmission.
3Adaptability or versatility
If speech coding is always performed in active speech mode, then decoding compatibility is maintained, but transmission efficiency is reduced due to continuous transmission
Solution Approach 1:
The patent applies segmentation by dividing the speech signal into active speech sections and inactive speech sections (noise sections). This segmentation allows the system to apply different coding strategies to different segments - using full coding for active sections and reduced coding for inactive sections, thereby improving transmission efficiency while maintaining compatibility through the unified coded data structure.
Data Source
AI summary
There is provided an audio encoding device capable of causing a decoding side to freely select an audio decoding mode corresponding to a control method used for audio encoding and capable of generating data which can be decoded even when the decoding side does not correspond to the control method. The audio encoding device (100) outputs encoded data corresponding to an audio signal containing an audio component and encoded data corresponding to an audio signal containing no audio component. An audio encoding unit (102) encodes the input audio signal in a predetermined section unit and generates encoded data. An audio present/absent judgment unit (106) decides whether the input audio signal contains an audio component for each predetermined section. A bit embedding unit (104) performs synthesis of noise data only for those generated from the input audio signal of the voice absent section in the encoded data generated by the audio encoding unit (102), thereby acquiring encoded data corresponding to an audio signal containing an audio component and encoded data corresponding to an audio signal containing no audio component.


