Vehicle Data Correlation Using Speech and Gaze Object Labeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Autonomous vehicles face challenges in distinguishing between relevant and irrelevant sensor data, which is crucial for making driving decisions. Existing methods require substantial training for artificial neural networks (ANNs) to differentiate between relevant and irrelevant data.

Innovation Solution

The proposed solution involves using human speech and gaze direction to label relevant objects in sensor data. By analyzing keywords in human speech and correlating them with gaze direction and gestures, the system can identify and label relevant objects in real-time, enhancing the ability of ANNs to discern between relevant and irrelevant data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If artificial neural networks are used to process sensor data for autonomous driving, then the ability to rapidly parse through large quantities of data is improved, but the difficulty of training the network to distinguish between relevant and irrelevant data increases

Engineering Contradiction:
Improvedata processing speedVSAvoidtraining complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The system performs preliminary actions by capturing human driving behavior data (speech, text, images, sensor data) during the driving instruction process before actual driving decisions are made. This pre-captured data is then used to train the artificial neural network, allowing the network to learn relevance distinctions in advance rather than requiring complex trial-and-error training during deployment

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates copies of human cognitive processes by capturing speech, text, and image data from driving instructors and students, then using these copies to train the neural network. The network learns to replicate human ability to distinguish relevant from irrelevant sensor data through these data copies, rather than requiring direct complex training protocols

Inventive Principle:
Principle #26Copying

2Measurement precision

If human speech and gaze direction are used to label relevant objects, then the precision of object identification is improved, but the complexity of the data processing system increases

Engineering Contradiction:
Improveobject identification accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system introduces speech and text data as intermediary elements that mediate between raw sensor data and object identification. The driving instructor's speech and text provide semantic labels that bridge the gap between sensor inputs and meaningful object recognition, simplifying the overall system architecture while improving precision

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system segments the data processing into distinct components: capturing speech separately, capturing text separately, capturing images separately, and capturing sensor data separately. Each component is processed independently and then integrated, making the complex system more manageable and easier to implement while maintaining high identification accuracy

Inventive Principle:
Principle #1Segmentation

Data Source

PatentEP4064222B1Vehicle data relation device and methods therefor
Publication Date: 2025.05.14 INTEL CORP
  • EP4064222B1 patent drawingFigure 1
  • EP4064222B1 patent drawingFigure 2
  • EP4064222B1 patent drawingFigure 3

AI summary

A vehicle data relation device includes an internal audio/image data analyzer, configured to receive first data representing at least one of audio from within the vehicle or an image from within the vehicle; identify within the first data second data representing an audio indicator or an image indicator, wherein the audio indicator is human speech associated with a significance of an object external to the vehicle, and wherein the image indicator is an action of a human within the vehicle associated with a significance of an object external to the vehicle; an external image analyzer, configured to receive third data representing an image of a vicinity external to the vehicle; identify within the third data an object corresponding to at least one of the audio indicator or the video indicator; and an object data generator, configured to generate data corresponding to the object.