Speech Processing Using Breathing Waveform Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech processing technologies are not optimal in scenarios where actual conditions differ from nominal conditions, leading to suboptimal performance, increased complexity, and resource demands, and lack flexibility in adapting to variations in speaker properties and activities.

Innovation Solution

An apparatus and method that utilize a breathing waveform signal to segment audio signals based on a speaker's lung air volume, independent of the audio signal level, to adapt speech processing to the speaker's respiratory pattern, using a trained artificial neural network to determine speech segments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If speech processing algorithms are adapted to compensate for varying speaker properties and conditions, then performance and adaptability are improved, but device complexity and resource usage increase

Engineering Contradiction:
Improveadaptability to speaker variationsVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The audio signal is segmented into multiple frames, with different processing applied to each frame based on its characteristics. This allows the system to adapt to varying speaker properties without requiring a completely complex adaptive system, as each segment can be processed independently with appropriate parameters.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The speech processing system dynamically adjusts parameters such as window length and processing intensity based on the characteristics of each audio segment. This dynamic adaptation allows the system to respond to varying speaker conditions without maintaining high complexity across all operating states.

Inventive Principle:
Principle #15Dynamics

2Productivity

If speech processing is optimized for specific nominal conditions, then processing efficiency is improved, but performance deteriorates when actual conditions differ from expected conditions

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidperformance under varying conditions
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system changes processing parameters based on the characteristics of each audio segment, including window length, processing intensity, and analysis methods. This allows the system to maintain high efficiency for each specific condition while being adaptable to various overall conditions, resolving the contradiction between optimization for specific conditions and reliability across varying conditions.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP4649481B1Speech processing of audio signal
Publication Date: 2026.04.22 KONINKLIJKE PHILIPS NV
  • EP4649481B1 patent drawingFigure 1
  • EP4649481B1 patent drawingFigure 2
  • EP4649481B1 patent drawingFigure 3

AI summary

An apparatus for speech processing comprises an input (101) arranged to receive an audio signal comprising a speech audio component. A determiner (105) generates a breathing waveform signal which is indicative of a lung air volume of the speaker as a function of time. A segmenter (107) segments the audio signal to generate speech segments of the audio signal in response to the breathing waveform signal. A speech processing circuit (103) performs a speech processing of the audio signal by applying a segment based processing to the speech segments. In some cases, the breathing waveform signal may be generated from the audio signal, e.g. using an artificial neural network. The approach may provide improved speech processing more closely reflecting the current speech properties.