Vehicle Voice Control Using GUI-Aware Semantic Understanding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Voice assistants in vehicles have limited intelligence and require strict voice inputs for recognizable instructions, leading to low efficiency and safety risks during interactions, especially in driving scenarios.

Innovation Solution

A method that combines local speech recognition with server-based semantic understanding using GUI information to generate operation instructions, enabling voice control of vehicle elements by associating voice inputs with graphical user interface (GUI) elements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If strict voice input requirements are imposed for recognizable voice instructions, then voice recognition accuracy is improved, but the level of intelligence and ease of operation deteriorates

Engineering Contradiction:
Improvevoice recognition accuracyVSAvoidlevel of intelligence
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent segments voice recognition into two stages: first, local speech recognition converts voice to text; second, server-based semantic understanding analyzes the text against GUI information. This segmentation allows relaxed initial voice capture while maintaining final recognition accuracy through multi-stage processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces GUI information as an intermediary between voice input and command generation. The server uses GUI element information (positions, attributes, contexts) to bridge the gap between casual voice input and precise command interpretation, enabling intelligent disambiguation without strict voice formatting requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Manufacturing precision

If manual interaction is required for vehicle operations, then control precision is improved, but safety and convenience deteriorate during driving

Engineering Contradiction:
Improvecontrol precisionVSAvoidsafety risks
Core Design Contradiction:
Manufacturing precisionVSObject-affected harmful factors

Solution Approach 1:

The patent replaces manual mechanical interaction with voice-based acoustic input. Drivers can issue commands through speech rather than physical manipulation of controls, eliminating the need for hand-eye coordination during driving while maintaining operational precision through server-based semantic understanding that interprets voice intent against current GUI context.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Speed

If voice assistant processes are performed locally only, then response speed is improved, but semantic understanding capability deteriorates

Engineering Contradiction:
Improveresponse speedVSAvoidsemantic understanding
Core Design Contradiction:
SpeedVSLoss of information

Solution Approach 1:

The patent divides processing between local and remote components: local speech recognition provides rapid initial text conversion, while remote server processing delivers comprehensive semantic understanding by analyzing text against GUI information. This segmentation enables both speed and intelligence by distributing tasks to appropriate processing locations.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary speech-to-text conversion locally before transmitting to the server. This preliminary action reduces the data transmission burden and allows the server to focus computational resources on semantic understanding, thereby maintaining response speed while enhancing interpretation capability.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentEP3955244B1Speech control method, information processing method, vehicle, and server
Publication Date: 2026.03.11 GUANGDONG XIAOPENG MOTORS TECH CO LTD
  • EP3955244B1 patent drawingFigure 1
  • EP3955244B1 patent drawingFigure 2~3
  • EP3955244B1 patent drawingFigure 4~6

AI summary

The present disclosure discloses a voice control method of a vehicle. The voice control method includes: obtaining voice input information; transmitting the voice input information and current Graphical User Interface (GUI) information of the vehicle to a server; receiving a voice control operation instruction generated by the server based on the voice input information, the current GUI information, and voice interaction information corresponding to the current GUI information; and parsing the voice control operation instruction, and performing an action as if being in response to a touch operation corresponding to the voice control operation instruction. In the voice control method for the vehicle according to an embodiment of the present disclosure, in a process of performing voice control, semantic understanding is performed on the voice input information in combination with the current GUI information of the vehicle, such that a voice assistant's ability on sematic understanding is improved, and elements in the GUI can be operated by voice, thereby providing a user with a more convenient way of interaction, a higher level of intelligence, and better user experience. The present disclosure also discloses an information processing method, a vehicle, a server, and a storage medium.