Voice to Text Biasing via Third-Party Context
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice-enabled electronic devices face challenges in accurately converting voice input to text, especially in context-dependent scenarios, leading to potential inaccuracies and increased computational resource usage.
Innovation Solution
The implementation of a voice to text engine within a local agent that dynamically biases voice to text conversions based on contextual parameters provided by a third-party agent, enhancing the accuracy and robustness of the conversion process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If standard voice to text conversion is used without contextual biasing, then the conversion process is simpler and faster, but the accuracy of the text representation decreases
Solution Approach 1:
The system performs preliminary actions by obtaining contextual parameters from the third-party agent before the voice to text conversion occurs. These contextual parameters are used to bias the conversion process, preparing the system in advance to handle the conversion more accurately by having relevant context ready beforehand.
Solution Approach 2:
The system uses feedback from the third-party agent in the form of contextual parameters that indicate potential features of further voice input. This feedback loop allows the voice to text engine to continuously adapt and improve conversion accuracy based on the agent's expectations and the ongoing dialog context.
2Measurement precision
If contextual parameters from third-party agent are used to bias voice to text conversion, then conversion accuracy improves, but computational resource usage increases
Solution Approach 1:
The system applies partial action by using contextual parameters selectively to bias only certain aspects of the voice to text conversion process. Rather than completely reprocessing the voice input, the system partially modifies the conversion based on contextual cues, reducing the overall computational burden while still improving accuracy.
Solution Approach 2:
The system changes parameters of the voice to text model dynamically based on contextual parameters from the third-party agent. By adjusting model parameters rather than retraining or replacing the entire model, the system achieves improved accuracy with minimal additional computational overhead.
3Productivity
If voice to text conversion is performed without contextual biasing, then processing is faster, but inaccuracies in voice input representation increase
Solution Approach 1:
The system performs preliminary actions by obtaining contextual parameters from the third-party agent before the voice to text conversion occurs. These contextual parameters are used to bias the conversion process, preparing the system in advance to handle the conversion more accurately by having relevant context ready beforehand.
Solution Approach 2:
The system uses feedback from the third-party agent in the form of contextual parameters that indicate potential features of further voice input. This feedback loop allows the voice to text engine to continuously adapt and improve conversion accuracy based on the agent's expectations and the ongoing dialog context.
Data Source
AI summary
Implementations relate to dynamically, and in a context-sensitive manner, biasing voice to text conversion. In some implementations, the biasing of voice to text conversions is performed by a voice to text engine of a local agent, and the biasing is based at least in part on content provided to the local agent by a third-party (3P) agent that is in network communication with the local agent. In some of those implementations, the content includes contextual parameters that are provided by the 3P agent in combination with responsive content generated by the 3P agent during a dialog that: is between the 3P agent, and a user of a voice-enabled electronic device; and is facilitated by the local agent. The contextual parameters indicate potential feature(s) of further voice input that is to be provided in response to the responsive content generated by the 3P agent.


