Voice Recognition Device Meta Information Output
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing real-time voice recognition devices struggle to accurately output non-textual information such as speaker conditions and feelings, which are crucial for applications like supporting hearing-impaired individuals and providing realistic information in educational settings.
Innovation Solution
A voice recognition device that generates and outputs meta information, including parameters like volume, utterance speed, and emotional cues, based on voice signals, and determines whether to add this information to the text output by comparing it with a reference presentation vector using similarity calculations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If real-time voice recognition device outputs only text information, then processing speed and simplicity are improved, but the ability to provide contextual and emotional information is lost
Solution Approach 1:
The patent segments the voice recognition output into two distinct components: text information (transcript) and meta information (speaker conditions and feelings). This segmentation allows the system to process and output text rapidly while separately extracting and providing contextual information, thereby maintaining high processing speed while preventing information loss.
Solution Approach 2:
The patent adds another dimension to the traditional text-only output by incorporating meta information that captures speaker conditions and feelings. This dimensional expansion transforms the output from one-dimensional (text only) to two-dimensional (text + contextual information), enabling the system to provide both rapid transcription and enriched contextual data.
2Measurement precision
If real-time voice recognition device outputs detailed meta information about speaker conditions, then the realism and contextual accuracy are improved, but the device complexity increases
Solution Approach 1:
The patent extracts meta information (speaker conditions and feelings) from the voice signal separately from the text recognition process. By taking out this additional information extraction function and implementing it as a dedicated component, the system achieves high measurement precision for speaker condition recognition while managing device complexity through modular architecture.
Solution Approach 2:
The patent implements a universal voice recognition device that can adapt its output based on different application scenarios. The device can selectively output text only, meta information only, or both combinations, making it multi-functional and adaptable to various needs without requiring separate systems for each function.
3Measurement precision
If voice recognition device uses complex models to recognize colloquial expressions and continuous utterances, then recognition accuracy is improved, but the computational resources required increase
Solution Approach 1:
The patent segments the voice recognition process into distinct modules: text recognition, meta information extraction, and similarity calculation. This segmentation allows each module to be optimized independently, using computational resources efficiently while maintaining high recognition accuracy for colloquial expressions and continuous utterances.
Solution Approach 2:
The patent applies partial action by selectively processing voice signals based on the application scenario. Instead of always processing all possible information at full complexity, the system can adjust the level of processing (text only, meta information only, or both) based on actual needs, optimizing the balance between recognition accuracy and computational resource consumption.
Data Source
AI summary
According to an embodiment, a voice recognition device includes one or more processors. The one or more processors are configured to: recognize a voice signal representing a voice uttered by an object speaker, to generate text and meta information representing information that is not included in the text and included in the voice signal; generate an object presentation vector including a plurality of parameters representing a feature of a presentation uttered by the object speaker; calculate a similarity between the object presentation vector and a reference presentation vector including a plurality of parameters representing a feature of a presentation uttered by a reference speaker; and output the text. The one or more processors are further configured to determine whether to output the meta information based on the similarity, and upon determining to output the meta information, add the meta information to the text and output the meta information.


