Speech Recognition Model Adaptive Update via Confidence Thresholds
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice recognition technologies face challenges in accurately and efficiently updating speech recognition models based on user input, leading to suboptimal performance and limited adaptability to varying user interactions.
Innovation Solution
An intelligent voice recognizing method and device that obtains a microphone detection signal, recognizes user voice using a pre-learned speech recognition model, generates recognition results, and updates the model based on these results, incorporating post-processing and communication with an AI processor for parameter tuning, while utilizing downlink control information for scheduling and quasi-colocated demodulation-reference signals for efficient data transmission.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If the speech recognition model is updated frequently based on user input, then the adaptability to user interactions is improved, but the device complexity and processing overhead increase
Solution Approach 1:
The system implements a feedback mechanism where speech recognition results are continuously monitored and used to update the speech recognition model. The processor receives speech recognition result information, performs post-processing to determine model update necessity, and updates the pre-learned speech recognition model accordingly, creating a closed-loop adaptive system that improves with usage while maintaining controlled complexity through selective updating based on confidence thresholds and result quality metrics
Solution Approach 2:
The system performs preliminary actions by pre-learning the speech recognition model before deployment and pre-establishing the feedback loop architecture. The model is pre-trained with general speech patterns, and the system pre-configures the update mechanism with confidence thresholds and selection criteria, so that during operation only selective updates based on specific conditions are performed, reducing real-time processing complexity
2Measurement precision
If the speech recognition model is updated in real-time, then the recognition accuracy improves, but the processing time and computational resources increase
Solution Approach 1:
The system applies partial action by selectively updating the speech recognition model only when certain conditions are met, such as when recognition confidence falls below a threshold or when specific types of speech patterns are detected. The processor performs post-processing on recognition results to determine update necessity, performing the computationally intensive model update operation only partially and selectively rather than continuously, thus balancing accuracy improvement with time consumption
Solution Approach 2:
The system implements periodic model updating based on accumulated speech recognition results rather than continuous real-time updates. The processor collects speech recognition result information over time, performs batch post-processing analysis, and updates the model at periodic intervals when sufficient data is accumulated, reducing computational overhead while maintaining accuracy improvement benefits
3Measurement precision
If the speech recognition result information is transmitted to the network for AI processing, then the model parameter tuning is improved, but the communication overhead and data transmission time increase
Solution Approach 1:
The system extracts only the necessary speech recognition result information for transmission to the network, rather than transmitting all raw data. The processor performs post-processing to identify and extract key parameters and metrics that are most valuable for model parameter tuning, such as recognition confidence scores, ambiguous speech patterns, and correction data, reducing communication overhead while maintaining tuning accuracy
Solution Approach 2:
The system uses the network as an intermediary for AI processing of speech recognition results. The processor transmits selected speech recognition result information to the network, which performs AI-based analysis and returns optimized model parameters. This intermediary approach allows complex processing to be distributed to the network while keeping the device lightweight, though it introduces transmission time delays
Data Source
AI summary
Disclosed are an intelligent voice recognizing method, a voice recognizing device, and an intelligent computing device. According to an embodiment of the present invention, an intelligent voice recognizing method of a voice recognizing device may obtain a microphone detection signal, recognize a user's voice from the microphone detection signal based on a pre-learned speech recognition model, output information related to a result of recognition of the user's voice, and update the speech recognition model based on the output speech recognition result information, easily updating the speech recognition model for speech recognition based on the speech recognition result information which is intuitively shown to the user. According to the present invention, one or more of the voice recognizing device, intelligent computing device, and server may be related to artificial intelligence (AI) modules, unmanned aerial vehicles (UAVs), robots, augmented reality (AR) devices, virtual reality (VR) devices, and 5G service-related devices.


