Breathing Waveform Speech Processing for Adaptive Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech processing technologies are not optimal in scenarios where actual conditions deviate from nominal conditions, leading to suboptimal performance, increased complexity, and resource demands, and are not flexible enough to adapt to variations in speaker properties and activities.

Innovation Solution

An apparatus and method for speech processing that generates segments of an audio signal, extracts breathing waveform signals from these segments using a trained artificial neural network, and applies speech processing dependent on the breathing waveform to adapt to variations in speaker conditions, reducing complexity and resource usage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of energy

If speech processing is performed continuously without breathing signal dependency, then speech recognition continuity is maintained, but power consumption increases and unnecessary processing occurs during pauses

Engineering Contradiction:
Improvepower consumptionVSAvoidspeech recognition continuity
Core Design Contradiction:
Loss of energyVSProductivity

Solution Approach 1:

The speech processing system operates periodically based on breathing cycles rather than continuously. The processor activates during inhalation phases when speech is likely to occur and remains inactive during exhalation phases, creating a periodic operation pattern that reduces power consumption while maintaining effective speech recognition coverage

Inventive Principle:
Principle #19Periodic action

Solution Approach 2:

The system uses breathing signal detection as feedback to control speech processing activation. The breathing sensor provides real-time feedback about the user's respiratory state, which the processor uses to dynamically adjust whether speech processing should be active, creating a closed-loop control system that optimizes power consumption based on actual speech likelihood

Inventive Principle:
Principle #23Feedback

2Loss of energy

If speech processing is activated only during inhalation, then power consumption is reduced, but speech detection during exhalation may be missed

Engineering Contradiction:
Improvepower consumptionVSAvoidspeech detection reliability
Core Design Contradiction:
Loss of energyVSReliability

Solution Approach 1:

The system activates speech processing during the inhalation phase, which covers a significant portion of speech production. While not capturing every possible speech moment, this partial action approach achieves acceptable speech recognition performance while dramatically reducing power consumption compared to continuous processing

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system adjusts the timing and duration of speech processing activation based on detected breathing parameters. By analyzing inhalation characteristics such as duration and intensity, the system dynamically modifies when processing occurs, optimizing the balance between power savings and speech detection reliability

Inventive Principle:
Principle #35Parameter changes

3Loss of energy

If breathing signal monitoring is added to control speech processing, then power consumption is optimized, but device complexity increases

Engineering Contradiction:
Improvepower consumptionVSAvoidsystem complexity
Core Design Contradiction:
Loss of energyVSDevice complexity

Solution Approach 1:

The breathing sensor serves multiple functions: it monitors respiratory state for speech processing control, and can potentially be used for other health monitoring applications. This multi-functionality justifies the added component by providing value beyond single-purpose speech activation control

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The breathing sensor acts as an intermediary component that translates physiological state into control signals for the speech processor. This intermediary layer simplifies the control logic by providing a clear, reliable trigger mechanism based on breathing phases, making the overall system easier to manage despite the added hardware

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentEP4662658B1Breathing signal-dependent speech processing of an audio signal
Publication Date: 2026.05.20 KONINKLIJKE PHILIPS NV
  • EP4662658B1 patent drawingFigure 1
  • EP4662658B1 patent drawingFigure 2
  • EP4662658B1 patent drawingFigure 3

AI summary

An apparatus for speech processing comprises an input (101) arranged to receive an audio signal comprising a speech audio component. A segmenter (105) generates segments of the audio signal with a time interval between consecutive segments having an intersegment duration that is less than the segment duration. A first generator (107) generates a fragment of breathing waveform signal for each segment of the audio signal and a second generator (109) generates the breathing waveform signal by combining the fragments of the breathing waveform signal by applying a weighted combination to samples of different fragments for the same time instant with weights being determined from a fragment window function. A speech processor (103) performs speech processing of the audio signal where the speech processing is dependent on the breathing waveform signal. The speech processor (103) may determine one or more breathing parameters, such as a breathing rate, breathing volume, timing of inspiration intervals etc., and the speech processing may be adapted in response to such breathing parameters.