Dynamic Speech Recognition Module Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional hybrid speech recognition systems suffer from high latency and low efficiency due to rigid recognition criteria, lack of adaptability to varying speech conditions, and duplication of computations between devices and servers, leading to suboptimal performance in environments with background noise, echo, or user feedback.

Innovation Solution

A method and system that dynamically select speech recognition modules between devices and servers based on past performance, quality of service, and user feedback, using machine learning algorithms to optimize processing and reduce latency, allowing for fine-grained QoS control and reuse of computations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech recognition is performed using server resources in hybrid architecture, then recognition accuracy is improved, but latency increases

Engineering Contradiction:
Improverecognition accuracyVSAvoidlatency
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent segments the speech recognition system into multiple modules (acoustic model, language model, decoder) that can be independently executed on either the device or server. This allows selective deployment of computation-intensive modules to the server while keeping simpler modules on the device, achieving high accuracy without excessive latency.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically selects which modules to execute on-device and which to send to the server based on real-time conditions such as computational resources available, network status, and speech characteristics. This dynamic adaptation optimizes the balance between accuracy and latency for each specific instance.

Inventive Principle:
Principle #15Dynamics

2Device complexity

If rigid sound recognition criteria are used in hybrid systems, then system complexity is reduced, but adaptability to varying speech conditions deteriorates

Engineering Contradiction:
Improvesystem complexityVSAvoidadaptability to speech conditions
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic module selection where the system automatically adapts its configuration based on detected speech conditions such as background noise, echo, and user feedback. The system can switch between different module combinations and execution locations (device/server) in response to changing conditions, providing high adaptability without requiring complex manual configuration.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system incorporates feedback mechanisms that monitor speech recognition performance and user responses in real-time. Based on this feedback, the system adjusts its operation by selecting different modules or changing execution parameters, enabling continuous adaptation to varying speech conditions while maintaining manageable system complexity through automated decision-making.

Inventive Principle:
Principle #23Feedback

3Reliability

If speech recognition modules are executed independently on device and server, then computational redundancy is increased, but processing reliability is improved

Engineering Contradiction:
Improveprocessing reliabilityVSAvoidcomputational redundancy
Core Design Contradiction:
ReliabilityVSLoss of energy

Solution Approach 1:

The patent merges the execution of speech recognition modules between device and server by allowing both to process the same audio input simultaneously and then combining their results. This approach maintains reliability through multiple processing paths while reducing computational redundancy by utilizing the outputs of both systems in an integrated manner rather than executing completely independent processes.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS11967318B2Method and system for performing speech recognition in an electronic device
Publication Date: 2024.04.23 SAMSUNG ELECTRONICS CO LTD
  • US11967318B2 patent drawing
  • US11967318B2 patent drawing
  • US11967318B2 patent drawing

AI summary

The present subject matter at least describes a method and a system (300, 1200) of performing speech-recognition in an electronic device having an embedded speech recognizer. The method comprises receiving an input-audio comprising speech at a device. In real-time, at-least one speech-recognition module is selected within at least one of the device and a server for recognition of at least a portion of the received speech based on a criteria defined in terms of a) past-performance of speech-recognition modules within the device and server; b) an orator of speech; and c) a quality of service associated with at least one of the device and a networking environment thereof. Based upon the selection of the server, output of the selected speech-recognition modules within the device are selected for processing by corresponding speech-recognition modules of the server. An uttered-speech is determined within the input-audio based on output of the selected speech-recognition modules of the device or the server.