Adaptive Proper Name Entity Recognition in Speech Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current automatic speech recognition (ASR) systems struggle with accurately recognizing proper names, especially uncommon ones, leading to incorrect results in spoken commands that include such names.
Innovation Solution
The system employs a secondary recognizer capable of rapid adaptation to novel proper names, using an adaptation object generated by an adaptation object generation module, which processes acoustic spans identified by a natural language understanding module to provide accurate transcriptions and meanings.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a primary ASR system is used for general speech recognition, then it can handle generic requests accurately, but it fails to recognize proper names especially uncommon ones
Solution Approach 1:
The system segments the speech recognition task into two parts: a primary ASR system for general speech and a secondary specialized recognizer for proper names. The secondary recognizer is specifically adapted to handle proper name recognition, while the primary system continues to handle generic requests, thus resolving the contradiction between general accuracy and specialized adaptability.
Solution Approach 2:
An adaptation object generation module acts as an intermediary between the primary ASR system and the secondary recognizer. It generates adaptation objects that enable the secondary recognizer to rapidly adapt to novel proper names without requiring extensive modification of the primary system, thus bridging the gap between general recognition and specialized adaptability.
2Measurement precision
If the primary ASR system is extensively adapted to recognize novel proper names, then proper name recognition accuracy improves, but system complexity and adaptation time increase
Solution Approach 1:
The system extracts the proper name recognition function from the primary ASR system and implements it in a separate secondary recognizer. This allows the primary system to remain simple and unchanged, while the secondary recognizer handles the complex task of recognizing novel proper names with specialized adaptation mechanisms.
Solution Approach 2:
The adaptation object generation module creates lightweight adaptation objects that are copied to enable the secondary recognizer to adapt to new proper names. This copying mechanism allows rapid adaptation without requiring extensive modification or retraining of the primary ASR system, thus reducing overall system complexity.
3Measurement precision
If a specialized secondary recognizer is introduced for proper names, then proper name recognition accuracy improves, but system complexity increases
Solution Approach 1:
The system merges the primary ASR system and secondary recognizer into a unified architecture where they work together. The adaptation object generation module coordinates between the two systems, allowing them to function as an integrated whole rather than separate complex components, thus managing system complexity while maintaining high proper name recognition accuracy.
Data Source
AI summary
Various embodiments contemplate systems and methods for performing automatic speech recognition (ASR) and natural language understanding (NLU) that enable high accuracy recognition and understanding of freely spoken utterances which may contain proper names and similar entities. The proper name entities may contain or be comprised wholly of words that are not present in the vocabularies of these systems as normally constituted. Recognition of the other words in the utterances in question, e.g. words that are not part of the proper name entities, may occur at regular, high recognition accuracy. Various embodiments provide as output not only accurately transcribed running text of the complete utterance, but also a symbolic representation of the meaning of the input, including appropriate symbolic representations of proper name entities, adequate to allow a computer system to respond appropriately to the spoken request without further analysis of the user's input.


