Anaphora Recognition in Speech AI Using Context
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech recognition systems face challenges in accurately specifying objects referred to by anaphora in speech commands, leading to failures in natural language processing.
Innovation Solution
An artificial intelligence apparatus and method that utilizes a microphone to receive speech commands, an anaphora recognition model to determine anaphora in text data, and a processor to specify objects based on context information, including screen context, to determine responses and control the AI apparatus accordingly.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speech recognition systems use vast amounts of database and cloud server processing, then speech recognition accuracy improves, but system complexity and processing time increase
Solution Approach 1:
The system divides speech recognition into multiple processing stages: voice activity detection, feature extraction, phoneme recognition, and natural language processing. Each stage handles specific aspects independently, reducing overall system complexity while maintaining accuracy through specialized processing at each level.
Solution Approach 2:
The system performs preliminary processing of speech data including voice activity detection and feature extraction before main recognition. Context information is pre-processed and stored for quick retrieval during recognition, reducing real-time processing complexity.
2Speed
If speech recognition systems process speech commands in real-time, then response speed improves, but processing accuracy deteriorates when anaphora is involved
Solution Approach 1:
The system performs preliminary analysis to detect anaphora in speech commands before full processing. Context information is pre-processed and stored in readily accessible formats, enabling quick resolution of anaphoric references without compromising real-time response requirements.
Solution Approach 2:
The system uses context information from previous interactions as feedback to resolve anaphora. The processor continuously updates context based on speech commands and resolves anaphoric references by comparing against stored context, improving accuracy while maintaining real-time processing through efficient feedback loops.
3Measurement precision
If the system uses context information including screen context to specify objects, then object specification accuracy improves, but information processing complexity increases
Solution Approach 1:
The context information processor handles multiple types of context data (conversation history, screen context, device state) through a unified processing mechanism. This multi-functional approach improves object specification accuracy while avoiding the need for separate complex processing systems for each context type.
Solution Approach 2:
The system extracts only the relevant context information needed for resolving anaphora from the overall context data. Rather than processing all available context, the processor identifies and extracts specific elements relevant to object specification, reducing processing complexity while maintaining accuracy.
Data Source
AI summary
Disclosed herein is an artificial intelligence apparatus for performing speech recognition including a microphone configured to receive a speech command of a user, a learning processor configured to determine anaphora included in text data corresponding to the speech command using an anaphora recognition model for determining anaphora included in predetermined text data, and a processor configured to specify an object referred to by the determined anaphora based on context information including information input to or output from the artificial intelligence apparatus, determine a response to the speech command based on the specified object, and control the artificial intelligence apparatus according to the determined response.


