Speech Processing Using Breathing Waveform Segmentation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech processing technologies are not optimal in scenarios where actual conditions differ from nominal conditions, leading to suboptimal performance, increased complexity, and resource demands, and lack flexibility in adapting to variations in speaker properties and activities.
Innovation Solution
An apparatus and method that utilize a breathing waveform signal to segment audio signals based on a speaker's lung air volume, independent of the audio signal level, to adapt speech processing to the speaker's respiratory pattern, using a trained artificial neural network to determine speech segments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If speech processing algorithms are adapted to compensate for varying speaker properties and conditions, then performance and adaptability are improved, but device complexity and resource usage increase
Solution Approach 1:
The audio signal is segmented into multiple frames, with different processing applied to each frame based on its characteristics. This allows the system to adapt to varying speaker properties without requiring a completely complex adaptive system, as each segment can be processed independently with appropriate parameters.
Solution Approach 2:
The speech processing system dynamically adjusts parameters such as window length and processing intensity based on the characteristics of each audio segment. This dynamic adaptation allows the system to respond to varying speaker conditions without maintaining high complexity across all operating states.
2Productivity
If speech processing is optimized for specific nominal conditions, then processing efficiency is improved, but performance deteriorates when actual conditions differ from expected conditions
Solution Approach 1:
The system changes processing parameters based on the characteristics of each audio segment, including window length, processing intensity, and analysis methods. This allows the system to maintain high efficiency for each specific condition while being adaptable to various overall conditions, resolving the contradiction between optimization for specific conditions and reliability across varying conditions.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An apparatus for speech processing comprises an input (101) arranged to receive an audio signal comprising a speech audio component. A determiner (105) generates a breathing waveform signal which is indicative of a lung air volume of the speaker as a function of time. A segmenter (107) segments the audio signal to generate speech segments of the audio signal in response to the breathing waveform signal. A speech processing circuit (103) performs a speech processing of the audio signal by applying a segment based processing to the speech segments. In some cases, the breathing waveform signal may be generated from the audio signal, e.g. using an artificial neural network. The approach may provide improved speech processing more closely reflecting the current speech properties.