Speech Recognition Model Selection via Multi-Sensor Data Fusion
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech recognition systems face challenges in improving performance, especially in noisy environments, as they rely solely on audio information and lack effective integration of multi-sensor data for enhanced recognition.
Innovation Solution
A speech recognition device and method that utilizes multi-sensor data, including location, image, and proximity information, to select appropriate language and acoustic models, thereby improving speech recognition performance by extracting feature vectors and transmitting recognition results to a terminal.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional speech recognition systems rely solely on audio information, then the system complexity is low, but speech recognition performance deteriorates in noisy environments
Solution Approach 1:
The patent combines audio information with multi-sensor data (location, image, proximity) to create a comprehensive speech recognition system. The server integrates data from multiple sources including microphone audio, camera images, GPS location, and proximity sensor data to improve recognition accuracy in noisy environments while managing system complexity through centralized processing.
Solution Approach 2:
The patent introduces a server as an intermediary between the speech recognition terminal and the processing system. The server receives audio information and multi-sensor data from the terminal, performs model selection and speech recognition processing, then transmits results back to the terminal. This mediator approach improves recognition performance while keeping the terminal device relatively simple.
2Adaptability or versatility
If multiple language and acoustic models are stored on the terminal, then speech recognition adaptability improves, but memory requirements increase
Solution Approach 1:
The patent segments the speech recognition system into terminal and server components. The terminal stores only basic recognition data and transmits audio information to the server, which stores and manages multiple language and acoustic models. This segmentation allows high adaptability through multiple models while keeping terminal memory requirements low.
Solution Approach 2:
The server acts as an intermediary that stores and manages multiple language and acoustic models. Instead of storing all models on the terminal, the system uses the server to provide model data, enabling the terminal to achieve high adaptability without requiring large local memory capacity.
3Speed
If the terminal processes all speech recognition data locally, then processing speed is fast, but device resource consumption increases
Solution Approach 1:
The patent uses a server as an intermediary to perform computationally intensive speech recognition processing. The terminal quickly transmits audio information to the server, which handles the complex model selection and recognition algorithms. This approach maintains fast overall processing speed while reducing the computational resource consumption on the terminal device.
Data Source
AI summary
Described herein is a speech recognition device comprising: a communication module receiving speech data corresponding to speech input from a speech recognition terminal and multi-sensor data corresponding to input environment of the speech; a model selection module selecting a language and acoustic model corresponding to the multi-sensor data among a plurality of language and acoustic models classified according to the speech input environment on the basis of previous multi-sensor data; and a speech recognition module controlling the communication module to apply a feature vector extracted from the speech data to the language and acoustic model and transmit speech recognition result for the speech data to the speech recognition terminal.


