Anaphora Recognition in Speech AI Using Context

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech recognition systems face challenges in accurately specifying objects referred to by anaphora in speech commands, leading to failures in natural language processing.

Innovation Solution

An artificial intelligence apparatus and method that utilizes a microphone to receive speech commands, an anaphora recognition model to determine anaphora in text data, and a processor to specify objects based on context information, including screen context, to determine responses and control the AI apparatus accordingly.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If speech recognition systems use vast amounts of database and cloud server processing, then speech recognition accuracy improves, but system complexity and processing time increase

Engineering Contradiction:
Improvespeech recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system divides speech recognition into multiple processing stages: voice activity detection, feature extraction, phoneme recognition, and natural language processing. Each stage handles specific aspects independently, reducing overall system complexity while maintaining accuracy through specialized processing at each level.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary processing of speech data including voice activity detection and feature extraction before main recognition. Context information is pre-processed and stored for quick retrieval during recognition, reducing real-time processing complexity.

Inventive Principle:
Principle #10Preliminary action

2Speed

If speech recognition systems process speech commands in real-time, then response speed improves, but processing accuracy deteriorates when anaphora is involved

Engineering Contradiction:
Improveresponse speedVSAvoidobject specification accuracy
Core Design Contradiction:
SpeedVSMeasurement precision

Solution Approach 1:

The system performs preliminary analysis to detect anaphora in speech commands before full processing. Context information is pre-processed and stored in readily accessible formats, enabling quick resolution of anaphoric references without compromising real-time response requirements.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system uses context information from previous interactions as feedback to resolve anaphora. The processor continuously updates context based on speech commands and resolves anaphoric references by comparing against stored context, improving accuracy while maintaining real-time processing through efficient feedback loops.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If the system uses context information including screen context to specify objects, then object specification accuracy improves, but information processing complexity increases

Engineering Contradiction:
Improveobject specification accuracyVSAvoidinformation processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The context information processor handles multiple types of context data (conversation history, screen context, device state) through a unified processing mechanism. This multi-functional approach improves object specification accuracy while avoiding the need for separate complex processing systems for each context type.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system extracts only the relevant context information needed for resolving anaphora from the overall context data. Rather than processing all available context, the processor identifies and extracts specific elements relevant to object specification, reducing processing complexity while maintaining accuracy.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS11289074B2Artificial intelligence apparatus for performing speech recognition and method thereof
Publication Date: 2022.03.29 LG ELECTRONICS INC
  • US11289074B2 patent drawing
  • US11289074B2 patent drawing
  • US11289074B2 patent drawing

AI summary

Disclosed herein is an artificial intelligence apparatus for performing speech recognition including a microphone configured to receive a speech command of a user, a learning processor configured to determine anaphora included in text data corresponding to the speech command using an anaphora recognition model for determining anaphora included in predetermined text data, and a processor configured to specify an object referred to by the determined anaphora based on context information including information input to or output from the artificial intelligence apparatus, determine a response to the speech command based on the specified object, and control the artificial intelligence apparatus according to the determined response.