Componentized Voice Server Speech Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional voice servers rely on processor-intensive software-based speech detection routines, which inefficiently use scarce resources and fail to leverage external hardware-based detection mechanisms, leading to suboptimal performance and resource consumption.

Innovation Solution

A pluggable, configurable speech detection component is integrated with internal software-based routines, allowing for selective use of external hardware-based detection mechanisms, such as energy difference detection in telephone channels, to enhance speech detection accuracy and resource management.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If software-based speech detection routines are used to improve speech detection accuracy, then detection precision is improved, but processor and memory consumption increases

Engineering Contradiction:
Improvespeech detection accuracyVSAvoidprocessor and memory consumption
Core Design Contradiction:
Measurement precisionVSUse of energy by moving object

Solution Approach 1:

The speech detection function is segmented into two separate components: hardware-based detection (energy level monitoring) and software-based detection (detailed speech analysis). The hardware component handles initial detection to reduce processor load, while the software component provides accurate speech recognition when needed, thus resolving the contradiction between detection accuracy and resource consumption.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

A hardware-based speech detection component acts as an intermediary between the telephony channel and the software-based speech detection routines. This intermediary performs preliminary energy level monitoring and filtering, reducing the burden on the processor-intensive software routines while maintaining overall detection accuracy.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Productivity

If hardware-based speech detection is used to reduce processor load, then resource efficiency is improved, but speech detection capability is reduced

Engineering Contradiction:
Improveresource efficiencyVSAvoidspeech detection capability
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent merges hardware-based speech detection (energy level monitoring) with software-based speech detection (detailed analysis) into a unified system. The hardware component provides efficient initial detection to reduce processor load, while the software component supplements with accurate speech recognition, achieving both resource efficiency and comprehensive detection capability.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system dynamically adjusts the balance between hardware and software detection based on operational needs. The configurable parameters allow the system to switch between hardware-only, software-only, or combined detection modes, enabling flexible adaptation to different resource availability and detection requirement scenarios.

Inventive Principle:
Principle #15Dynamics

3Reliability

If internal software-based speech detection is used to ensure detection reliability, then detection reliability is improved, but device complexity increases

Engineering Contradiction:
Improvedetection reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The detection system is segmented into independent hardware and software modules with clearly defined interfaces. This segmentation allows each component to be optimized separately for its specific function while maintaining overall system reliability through the coordinated operation of both modules.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The voice server is designed with universal architecture that can accommodate both hardware-based and software-based detection methods through a unified interface. This multi-functionality allows the system to reliably perform speech detection regardless of which detection method is used or combined, reducing overall system complexity despite the added capability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This approach enables efficient use of resources by allowing customers to configure speech detection settings, reducing processor load and improving detection accuracy by combining internal and external detection methods, thereby optimizing voice server performance.

Implementation Method 1

hardware-based techniques can monitor signal energy levels within telephony channels and differentiate speech utterances from silence and/or noise based upon differences in the signal energy levels

Methodology Applied
Scientific EffectSignal energy level detection:

Data Source

PatentUS7925510B2Componentized voice server with selectable internal and external speech detectors
Publication Date: 2011.04.12 MICROSOFT TECHNOLOGY LICENSING LLC
  • US7925510B2 patent drawing
  • US7925510B2 patent drawing
  • US7925510B2 patent drawing

AI summary

A method for detecting speech utterances within a telephone call can include the steps of initializing a componentized voice server having at least one software-based speech detection routine. At least one previously established parameter can be used to discern a speech detection methodology for handling an incoming call. The software-based speech detection routine can be set in accordance with a select one of the parameters. An indicator of particular one of the parameters can be conveyed to an external speech detection component so that the external speech detection component is set to detect speech for the call in accordance with the conveyed indication. The software-based speech detection routine and/or the external speech detection component can detect a speech utterance for the call. The voice server can perform at least one programmatic action responsive to the detecting of the speech utterance.