Speech Recognition Signal Evaluation for Load Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech recognition systems face delays due to high load in signal processing, which occurs before the recognition process, and existing methods do not effectively address this issue, leading to recognition process delays.

Innovation Solution

A speech recognition system that evaluates speech signals to determine if they are from a speech section and only processes those signals, thereby reducing the load on signal processing by limiting processing to necessary sections.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If signal processing is performed on all input signals, then speech recognition performance is improved, but recognition process delay increases due to high processing load

Engineering Contradiction:
Improvespeech recognition performanceVSAvoidrecognition process delay
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent extracts and processes only the necessary portions of speech signals by evaluating sound sections before full signal processing. The evaluation unit identifies sections where speech is likely present, and the signal processing unit processes only those sections, thereby reducing overall processing load and delay while maintaining recognition performance.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent segments the speech signal into sound sections and non-sound sections based on evaluation results. By dividing the continuous signal into discrete processable units and only processing relevant segments, the system reduces processing burden and eliminates unnecessary processing delays.

Inventive Principle:
Principle #1Segmentation

2Reliability

If advanced signal processing is performed, then speech recognition accuracy is improved, but system resource consumption increases

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidCPU resource consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent applies partial action by performing advanced signal processing only on the necessary portions of the signal identified as sound sections, rather than processing the entire signal. This selective processing maintains recognition accuracy for relevant segments while significantly reducing overall CPU resource consumption.

Inventive Principle:
Principle #16Partial or excessive action

3Productivity

If signal processing is limited to sound sections only, then processing load is reduced, but speech recognition completeness may be compromised

Engineering Contradiction:
Improveprocessing speedVSAvoidrecognition completeness
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The patent employs feedback mechanisms where the evaluation unit continuously monitors signal characteristics and provides feedback to the signal processing unit. This feedback loop ensures that processing is dynamically adjusted based on actual sound section detection, maintaining recognition completeness while optimizing processing load reduction.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS8886527B2Speech recognition system to evaluate speech signals, method thereof, and storage medium storing the program for speech recognition to evaluate speech signals
Publication Date: 2014.11.11 NEC CORP
  • US8886527B2 patent drawing
  • US8886527B2 patent drawing
  • US8886527B2 patent drawing

AI summary

A purpose is to suppress recognition process delay generated due to load in signal processing. Included is a speech input means 10 that inputs a speech signal, an output evaluation means 20 that evaluates whether or not the speech signal input by the speech input means 10 is the speech signal in a sound section, which is a speech section assuming that a speaker is speaking, and outputs the speech signal as a speech signal to be processed only when evaluated as the speech signal in the sound section, a signal processing means 30 that performs signal processing to the speech signal, which is output by the output evaluation means 20 as the speech signal to be processed, and a speech recognition processing means 40 that performs a speech recognition process to the speech signal which is signal-processed by the signal processing means 30.