ASR-NLU Model Alignment via Contribution Bias
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Automatic speech recognition (ASR) and natural language understanding (NLU) applications face laborious and repetitive modeling processes, as each application typically builds its behavior from corpora, leading to inefficiencies in translating input signals into output representations.
Innovation Solution
The alignment process identifies and aligns language model components between ASR and NLU applications based on contribution bias, utilizing pre-trained models, including general and domain-specific components, to configure applications at runtime with explicit directives or identifying tuples, reducing duplicate modeling efforts.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If each application builds its behavior from corpora through traditional modeling, then the application can be customized for specific domains, but the modeling process becomes laborious and repetitive
Solution Approach 1:
The patent pre-trains language model components on common domains and use cases before runtime. These pre-trained components are stored and can be directly applied to multiple applications, eliminating the need to rebuild models from scratch for each application. This preliminary preparation significantly reduces modeling time while maintaining domain adaptability.
Solution Approach 2:
The patent creates language model components with multiple language models that can serve multiple applications across different domains. Instead of building separate models for each application, a single pre-trained language model component can be applied to multiple applications, reducing repetitive modeling work while maintaining customization capabilities.
2Reliability
If traditional separate modeling is used for ASR and NLU, then each component can be optimized independently, but the alignment between ASR and NLU becomes complex and time-consuming
Solution Approach 1:
The patent merges the ASR language model and NLU language model into a single aligned language model component. This unified model simultaneously optimizes both ASR and NLU performance, eliminating the complexity of aligning separate models while maintaining the ability to optimize for both functions together through joint training on aligned corpora.
3Measurement precision
If comprehensive language modeling is performed for each application, then accuracy is improved, but development effort and resource consumption increase
Solution Approach 1:
The patent performs comprehensive language modeling in advance by pre-training language model components on large corpora covering common domains and use cases. These pre-trained models achieve high accuracy and can be directly applied to multiple applications, maintaining recognition accuracy while significantly improving development efficiency by eliminating redundant training work.
Solution Approach 2:
The patent creates reusable language model components that can be copied and applied to multiple applications. Instead of performing comprehensive modeling for each application, the same pre-trained language model component can be copied and adapted, maintaining accuracy while reducing development effort and resource consumption.
Data Source
AI summary
A method, computer program product, and computer system for identifying, by a computing device, at least one language model component of a plurality of language model components in at least one application associated with automatic speech recognition (ASR) and natural language understanding (NLU) usage. A contribution bias may be received for the at least one language model component. The ASR and NLU may be aligned between the plurality of language model components based upon, at least in part, the contribution bias.


