Low Delay Voice Processing System Using Neural Intent Prediction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current voice recognition systems fail to facilitate natural conversations between humans and machines, as they typically require a one-sided utterance followed by a response, lacking the immediacy and interactivity of human dialogue.
Innovation Solution
A low delay voice processing system that generates activation timing information for microphones based on pre-trained neural network models, allowing for real-time intent inference and response generation, enabling natural conversation by optimizing microphone activation and intent extraction from user utterances.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the voice recognition system waits for complete utterance termination before processing, then recognition accuracy is improved, but conversation naturalness deteriorates
Solution Approach 1:
The system performs preliminary actions by predicting user intent during the utterance before it completes. The neural network model analyzes speech patterns and predicts intent in real-time, allowing the system to prepare response strategies early. This preliminary intent recognition enables the system to maintain natural conversation flow while ensuring accurate recognition through progressive analysis of the complete utterance.
2Measurement precision
If the system processes the complete utterance before responding, then response accuracy is improved, but dialogue latency increases
Solution Approach 1:
The system performs preliminary intent recognition during utterance playback using a neural network model that predicts user intent in real-time. This preliminary action allows the system to prepare response strategies before the utterance completes, significantly reducing dialogue latency while maintaining response accuracy through progressive analysis.
Solution Approach 2:
The system implements feedback mechanisms where the predicted intent is continuously refined as the utterance progresses. The neural network model receives feedback from ongoing speech analysis and adjusts predictions dynamically, enabling accurate intent recognition that balances speed and precision in the response generation process.
3Reliability
If the microphone remains activated throughout utterance playback, then response reception reliability is improved, but energy consumption increases
Solution Approach 1:
The system dynamically adjusts microphone activation states based on real-time intent prediction and utterance progress. Rather than maintaining continuous activation, the microphone is strategically activated only when predicted intent indicates a high probability of user response, creating a dynamic balance between response reception reliability and energy consumption.
Solution Approach 2:
The system employs periodic microphone activation synchronized with predicted response timing. Based on neural network predictions of when users are most likely to respond during utterance playback, the microphone is activated in periodic intervals rather than continuously, reducing energy consumption while maintaining reliable response capture during critical moments.
Data Source
AI summary
Disclosed is a speech processing method. The speech processing method controls activation timing of a microphone based on a response pattern of the microphone from a user in order to implement a natural conversation. The speech processing device and the NLP system of the present disclosure may be associated with an artificial intelligence module, a drone (or unmanned aerial vehicle (UAV)), a robot, an augmented reality (AR) device, a virtual reality (VR) device, a device related to 5G service, etc.


