Context-Aware Voice Control for Autonomous Vehicle Commands

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing vehicle control systems struggle to integrate processor(s) and user-facing interfaces in a way that improves user control without introducing distractions, particularly in understanding and acting on complex, context-specific voice commands.

Innovation Solution

A device and method that utilize a contextual encoder system with machine-learning models to process scene data and speech commands, generating vehicle control signals by aligning embeddings in a shared space and combining them to form a navigation embedding.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional vehicle control systems use conventional interfaces (buttons, knobs, displays), then users can control basic vehicle functions, but the system cannot effectively process complex context-specific voice commands and provides limited user control over autonomous operations

Engineering Contradiction:
Improvecapability to process complex context-specific voice commandsVSAvoidintegration of processor and user-facing interface
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The control system is designed to handle multiple types of inputs (voice commands, scene data from sensors, navigation data) and multiple output functions (vehicle control signals for steering, acceleration, braking) through a unified architecture. The processor integrates these diverse functions into a single system that can adapt to different command types and vehicle operating modes, making the system versatile without requiring separate dedicated systems for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent introduces an intermediary processing layer that receives raw voice commands and scene data, processes them through machine learning models to extract meaningful context, and translates them into vehicle control signals. This intermediary layer acts as a mediator between the user's natural language input and the vehicle's control systems, enabling complex command processing without directly complicating the interface between user and vehicle.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If the system integrates advanced processors and machine learning models for speech processing, then it can understand context-specific commands, but it may introduce distractions or reduce reliability of control

Engineering Contradiction:
Improvenatural language spoken controlVSAvoidcontrol signal accuracy
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The system incorporates feedback mechanisms where the processor continuously monitors both the incoming speech commands and the resulting vehicle responses. Scene data from sensors provides environmental feedback that is cross-referenced with voice commands to verify intended actions. This multi-source feedback loop allows the system to detect and correct potential errors, maintaining reliability while enabling natural language control.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs preliminary processing of voice commands by analyzing speech data against scene context before generating final control signals. Machine learning models pre-process the language input to extract intent and parameters, cross-checking against current vehicle state and environmental conditions. This preliminary action filters out ambiguous or conflicting commands before they reach the execution layer, preventing erroneous control actions.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If the system processes both scene data and speech commands through separate models, then it can handle diverse inputs, but it increases computational complexity and processing time

Engineering Contradiction:
Improveprocessing of multiple data typesVSAvoidmultiple machine learning models
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent merges the processing of scene data and speech commands into a unified machine learning architecture. Rather than treating these as completely separate processing streams, the system integrates them into a single contextual understanding model that processes both inputs simultaneously, sharing computational resources and intermediate representations. This reducing the overall computational complexity while maintaining the ability to handle diverse input types.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250178624A1Speech-based vehicular control
Publication Date: 2025.06.05 QUALCOMM INC
  • US20250178624A1 patent drawing
  • US20250178624A1 patent drawing
  • US20250178624A1 patent drawing

AI summary

A device includes memory configured to store scene data from one or more scene sensors associated with a vehicle. The device also includes one or more processors configured to obtain, via a first machine-learning model of a contextual encoder system, a first embedding based on data representing speech that includes one or more commands for operation of the vehicle. The one or more processors are configured to obtain, via a second machine-learning model of the contextual encoder system, a second embedding based on the scene data and based on state data of the first machine-learning model. The one or more processors are configured to generate one or more vehicle control signals for the vehicle based on the first embedding and the second embedding.