In-Cabin Gesture Recognition With User Grid and Multi-User Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing vehicle control systems face challenges in accurately recognizing non-contact gesture operations, particularly when the driver or passenger is not in a position to interact with the screen, limiting the range of gesture recognition and potentially leading to inaccurate or unrecognized commands.

Innovation Solution

A vehicle control method that constructs a user grid based on in-cockpit images and user position information, using sensors like RGB cameras, IR cameras, and radar to enhance gesture recognition accuracy by providing depth data and identifying multiple users, and determines control permissions based on user positions and facial data to ensure safe and flexible vehicle operation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If an arrayed photoelectric sensor is used to collect gesture operations, then the device structure is simplified, but the recognizable range is limited and gesture operations may not be accurately recognized when users are in certain positions

Engineering Contradiction:
Improvedevice structureVSAvoidgesture recognition accuracy
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent combines multiple sensing approaches (photoelectric sensors, depth cameras, radar) into a unified gesture recognition system. This merging of different sensing technologies expands the recognizable range while maintaining relatively simple device structure, resolving the contradiction between simplicity and accuracy.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system is designed to support multiple users (driver and front passenger) with gesture recognition capabilities. By making the system universal for different users and positions, it overcomes the limitation of restricted recognizable range while keeping the overall device architecture manageable.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Device complexity

If gesture recognition is limited to a single user, then the system complexity is reduced, but the adaptability and user experience are degraded

Engineering Contradiction:
Improvesystem complexityVSAvoidmulti-user support
Core Design Contradiction:
Device complexityVSAdaptability or versatility

Solution Approach 1:

The patent segments the gesture recognition space into multiple user zones (driver zone and front passenger zone) with different control permissions. This segmentation allows multi-user support while managing system complexity through clear spatial division and permission assignment.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system adds a spatial dimension to gesture recognition by determining user positions and assigning different control permissions based on location. This dimensional approach enables multi-user adaptability without proportionally increasing system complexity, as the permission structure follows clear spatial rules.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Adaptability or versatility

If the gesture recognition range is expanded to multiple users, then the adaptability is improved, but the difficulty of detecting and measuring gestures increases

Engineering Contradiction:
Improvemulti-user gesture recognitionVSAvoidgesture detection complexity
Core Design Contradiction:
Adaptability or versatilityVSDifficulty of detecting and measuring

Solution Approach 1:

The patent applies different detection strategies and permission rules to different spatial zones (driver area vs. front passenger area). This local quality approach allows multi-user gesture recognition while managing detection complexity by tailoring the recognition criteria to each user's position and expected gesture patterns.

Inventive Principle:
Principle #3Local quality

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

The method improves gesture recognition accuracy and flexibility, allowing multiple users to control vehicle functions safely and efficiently, expanding the range of gesture recognition beyond the display and ensuring authorized users can perform intended operations.

Implementation Method 1

obtaining an in-cockpit image of a vehicle and user in-position information

Methodology Applied
Scientific EffectLight reflection: Reflection

Implementation Method 2

using sensors like RGB cameras, IR cameras, and radar to enhance gesture recognition accuracy by providing depth data

Methodology Applied
Scientific EffectInfrared radiation: Infrared Radiation

Implementation Method 3

using sensors like RGB cameras, IR cameras, and radar to enhance gesture recognition accuracy

Methodology Applied
Scientific EffectRadar: Radar

Implementation Method 4

using sensors like RGB cameras, IR cameras, and radar to enhance gesture recognition accuracy by providing depth data

Methodology Applied
Scientific EffectTime of flight: Time of Flight

Data Source

PatentEP4679231A1Method and apparatus for vehicle control based on gesture recognition, and vehicle
Publication Date: 2026.01.14 YINWANG INTELLIGENT TECHNOLOGIES CO LTD
  • EP4679231A1 patent drawingFigure 1
  • EP4679231A1 patent drawingFigure 2
  • EP4679231A1 patent drawingFigure 3

AI summary

Embodiments of this application provide a vehicle control method based on gesture recognition, an apparatus, and a vehicle. The method includes: obtaining an in-cockpit image of a vehicle and user in-position information, where the user in-position information indicates positions of M users in the vehicle that are presented in the in-cockpit image; constructing a user grid based on the in-cockpit image and the user in-position information, where the user grid includes depth data of the M users; for each of N target users performing a gesture operation, recognizing a gesture operation intention of the target user based on the user grid, where the M users include the N target users; and separately controlling the vehicle based on gesture operation intentions of the N target users, where both M and N are positive integers, and M is greater than or equal to N. The user grid is constructed based on the user in-position information, to provide refined data of the gesture operation for gesture recognition, thereby improving accuracy of gesture recognition.