Service-Oriented Speech Recognition for In-Vehicle Text Input

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current in-vehicle speech recognition systems face challenges in providing safe and efficient hands-free operation for drivers, particularly due to harsh environments like road noise and the need to distinguish between multiple voices, which complicates complex speech tasks and limits the practicality of speech-enabled functionalities.

Innovation Solution

A server-based speech recognition system utilizing a service-oriented architecture (SOA) that leverages multiple specialized speech recognizers, allowing for asynchronous speech recognition and minimizing the need for visual and mechanical interactions, enabling drivers to use speech for tasks like text input without repeating utterances and providing a seamless user experience.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If speech recognition is used for hands-free operation in vehicles, then driver safety is improved by reducing manual interactions, but recognition accuracy deteriorates due to harsh environments like road noise and multiple voices

Engineering Contradiction:
Improvehands-free operationVSAvoidspeech recognition accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The system segments the speech recognition task into multiple specialized recognizers, each optimized for specific speech tasks or conditions. This allows the system to handle different speech patterns and environmental conditions with dedicated recognition models, improving overall accuracy in harsh vehicle environments while maintaining hands-free operation

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary service-oriented architecture that acts as a mediator between the speech input and the recognition processing. This SOA layer manages the complexity of multiple recognizers and environmental variations, coordinating their work to achieve accurate recognition despite road noise and multiple voices

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If multiple specialized speech recognizers are used to improve recognition accuracy, then speech recognition reliability is improved, but system complexity increases

Engineering Contradiction:
Improvespeech recognition reliabilityVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The service-oriented architecture serves as a universal platform that manages multiple specialized recognizers through standardized interfaces. This multi-functional framework allows the system to coordinate various recognizers without requiring separate management systems for each, reducing the overall complexity increase while maintaining high reliability

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The SOA acts as an intermediary layer that abstracts the complexity of multiple recognizers behind a unified interface. This mediator handles the coordination, selection, and integration of different recognizers, shielding the user and application logic from the underlying system complexity while ensuring reliable speech recognition

Inventive Principle:
Principle #24Intermediary (Mediator)

3Measurement precision

If synchronous speech recognition is used, then recognition accuracy can be improved by waiting for complete processing, but user interaction efficiency deteriorates due to delays and requirement to repeat utterances

Engineering Contradiction:
Improverecognition accuracyVSAvoiduser interaction efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs preliminary speech processing and recognition in parallel while the user continues interaction. By initiating recognition processes beforehand and processing speech asynchronously, the system reduces waiting time and eliminates the need for users to repeat utterances, maintaining both accuracy and efficiency

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements dynamic speech recognition that adapts its processing mode based on context. The system can switch between synchronous and asynchronous processing, adjusting the level of confirmation and re-recognition requirements dynamically. This allows the system to maintain high accuracy when needed while improving interaction efficiency in routine scenarios

Inventive Principle:
Principle #15Dynamics

Data Source

PatentEP2411977B1Service oriented speech recognition for in-vehicle automated interaction
Publication Date: 2019.05.08 SIRIUS XM CONNECTED VEHICLE SERVICES INC
  • EP2411977B1 patent drawingFigure 1
  • EP2411977B1 patent drawingFigure 2
  • EP2411977B1 patent drawingFigure 3

AI summary

A system and method for implementing a server-based speech recognition system for multi¬ modal automated interaction in a vehicle includes receiving, by a vehicle driver, audio prompts by an on-board human-to-machine interface and a response with speech to complete tasks such as creating and sending text messages, web browsing, navigation, etc. This service-oriented architecture is utilized to call upon specialized speech recognizers in an adaptive fashion. The human-to-machine interface enables completion of a text input task while driving a vehicle in a way that minimizes the frequency of the driver's visual and mechanical interactions with the interface, thereby eliminating unsafe distractions during driving conditions. After the initial prompting, the typing task is followed by a computerized verbalization of the text. Subsequent interface steps can be visual in nature, or involve only sound.