Context-Aware Speech Recognition Syntax Model Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Automatic speech recognition devices face challenges in achieving high recognition rates for extended vocabularies while maintaining low latency and operational relevance, especially in mobile situations with limited computing power, and existing solutions either require significant resources or result in high error rates due to complex syntax model recalculations.

Innovation Solution

An automatic speech recognition device that detects contextual elements and uses multiple syntax models to build candidate phoneme sequences, selecting the model with the highest acoustic and syntax probability product, allowing for intuitive and accurate recognition of oral instructions with minimal computational overhead.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If very reliable acoustic models are used to achieve low error rates, then recognition accuracy is improved, but computing power requirements and database size increase significantly

Engineering Contradiction:
Improverecognition accuracyVSAvoidcomputing power requirements
Core Design Contradiction:
Measurement precisionVSPower

Solution Approach 1:

The patent segments the syntax model into multiple specialized models (e.g., navigation syntax model, communication syntax model, system control syntax model), each optimized for specific domains. This segmentation allows the system to use smaller, more efficient models for each task rather than one large comprehensive model, reducing overall computing requirements while maintaining accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system dynamically selects and switches between different syntax models based on the detected speech context. The context detection unit identifies the current situation (navigation, communication, etc.) and activates the corresponding syntax model, making the system adaptable and efficient without requiring all models to be loaded and processed simultaneously.

Inventive Principle:
Principle #15Dynamics

2Device complexity

If restricted syntax models are used to reduce computing power requirements, then device complexity is reduced, but the number of recognizable instructions is limited

Engineering Contradiction:
Improvecomputational complexityVSAvoidnumber of recognizable instructions
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent creates a universal speech recognition system that handles multiple functions and domains through a set of specialized syntax models. Each model is tailored to a specific function (navigation, communication, system control), but together they provide comprehensive coverage. The context detection unit determines which function is being invoked and activates the appropriate model, achieving versatility without requiring a single overly complex model.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If syntax model recalculation is performed in real-time based on user gaze, then recognition accuracy is improved, but processing time and latency increase

Engineering Contradiction:
Improverecognition accuracyVSAvoidprocessing latency
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary preparation by having multiple syntax models pre-configured and ready for different contexts. Instead of calculating a new syntax model in real-time based on user gaze, the system has pre-defined syntax models for various contexts (navigation, communication, etc.) that can be quickly activated. The context detection unit simply determines which pre-prepared model to use, avoiding time-consuming real-time calculations while maintaining accuracy.

Inventive Principle:
Principle #10Preliminary action

4Adaptability or versatility

If multiple syntax models are maintained for different applications, then versatility is improved, but device complexity increases

Engineering Contradiction:
Improvenumber of recognizable instructionsVSAvoidmodel management complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces a context detection unit as an intermediary between the multiple syntax models and the speech recognition process. This intermediary detects the current speech context (navigation, communication, system control) and automatically selects the appropriate syntax model. This mediation simplifies the management of multiple models by providing a clear selection mechanism based on context, reducing the complexity of model management while maintaining versatility.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS10403274B2Automatic speech recognition with detection of at least one contextual element, and application management and maintenance of aircraft
Publication Date: 2019.09.03 DASSAULT AVIATION SA
  • US10403274B2 patent drawing
  • US10403274B2 patent drawing
  • US10403274B2 patent drawing

AI summary

An automatic speech recognition with detection of at least one contextual element, and application to aircraft flying and maintenance are provided. The automatic speech recognition device comprises a unit for acquiring an audio signal, a device for detecting the state of at least one contextual element, and a language decoder for determining an oral instruction corresponding to the audio signal. The language decoder comprises at least one acoustic model defining an acoustic probability law and at least two syntax models each defining a syntax probability law. The language decoder also comprises an oral instruction construction algorithm implementing the acoustic model and a plurality of active syntax models taken from among the syntax models, a contextualization processor to select, based on the state of the order each contextual element detected by the detection device, at least one syntax model selected from among the plurality of active syntax models, and a processor for determining the oral instruction corresponding to the audio signal.