Real-Time Speech Boundary Detection for Accurate Command Capture

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech control systems for medical apparatuses face a trade-off between rapid processing times and accurate capture of operator intentions, leading to potential errors or frustration due to incomplete or incorrect command execution.

Innovation Solution

A method for processing audio signals that recognizes the beginning and end of speech input in real-time, using adaptive time intervals based on continuous analysis, to provide a speech data stream for accurate command recognition, avoiding premature termination or unnecessary waiting.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If speech analysis is performed rapidly to reduce waiting time, then processing speed is improved, but accuracy of capturing operator intention deteriorates

Engineering Contradiction:
Improvespeech analysis speedVSAvoidaccuracy of capturing operator intention
Core Design Contradiction:
SpeedVSMeasurement precision

Solution Approach 1:

The speech analysis process is divided into two distinct segments: a first speech analysis performed rapidly during speech input to enable quick response, and a second, more comprehensive speech analysis performed after speech input ends to ensure complete and accurate capture of the operator's intention. This segmentation allows the system to balance speed and accuracy by applying different analysis depths at different stages.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The first speech analysis is performed as a preliminary action during the speech input phase, providing quick preliminary results that reduce waiting time. This preliminary analysis prepares the system for faster response while the more thorough second analysis ensures complete accuracy, resolving the contradiction between speed and precision.

Inventive Principle:
Principle #10Preliminary action

2Measurement precision

If speech analysis is performed completely and accurately, then capture of operator intention is improved, but processing time increases

Engineering Contradiction:
Improveaccuracy of capturing operator intentionVSAvoidwaiting time for operator
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The comprehensive speech analysis is segmented into two phases: an initial rapid analysis during speech input that provides quick feedback, and a final detailed analysis after speech input completes that ensures complete accuracy. This prevents the operator from waiting for the full comprehensive analysis while still achieving it for final accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The speech analysis process maintains continuity by performing the first analysis continuously during speech input, providing ongoing useful action and quick response. The second analysis then completes the process with comprehensive accuracy, ensuring both speed and completeness without interruption.

Inventive Principle:
Principle #20Continuity of useful action

Data Source

PatentUS12400642B2Method for processing an audio signal, method for controlling an apparatus and associated system
Publication Date: 2025.08.26 SIEMENS HEALTHINEERS AG
  • US12400642B2 patent drawing
  • US12400642B2 patent drawing
  • US12400642B2 patent drawing

AI summary

In a method for processing an audio signal, the audio signal is continuously analyzed substantially in real time from a recognized beginning of the speech input to provide a speech analysis result. The speech analysis result is used to dynamically define an end of the speech input. A speech data stream is provided based on the audio signal between the beginning and the end. The speech data stream may be further analyzed to identify one or more speech commands.