Vehicle Data Correlation Using Speech and Gaze Object Labeling
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Autonomous vehicles face challenges in distinguishing between relevant and irrelevant sensor data, which is crucial for making driving decisions. Existing methods require substantial training for artificial neural networks (ANNs) to differentiate between relevant and irrelevant data.
Innovation Solution
The proposed solution involves using human speech and gaze direction to label relevant objects in sensor data. By analyzing keywords in human speech and correlating them with gaze direction and gestures, the system can identify and label relevant objects in real-time, enhancing the ability of ANNs to discern between relevant and irrelevant data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If artificial neural networks are used to process sensor data for autonomous driving, then the ability to rapidly parse through large quantities of data is improved, but the difficulty of training the network to distinguish between relevant and irrelevant data increases
Solution Approach 1:
The system performs preliminary actions by capturing human driving behavior data (speech, text, images, sensor data) during the driving instruction process before actual driving decisions are made. This pre-captured data is then used to train the artificial neural network, allowing the network to learn relevance distinctions in advance rather than requiring complex trial-and-error training during deployment
Solution Approach 2:
The system creates copies of human cognitive processes by capturing speech, text, and image data from driving instructors and students, then using these copies to train the neural network. The network learns to replicate human ability to distinguish relevant from irrelevant sensor data through these data copies, rather than requiring direct complex training protocols
2Measurement precision
If human speech and gaze direction are used to label relevant objects, then the precision of object identification is improved, but the complexity of the data processing system increases
Solution Approach 1:
The system introduces speech and text data as intermediary elements that mediate between raw sensor data and object identification. The driving instructor's speech and text provide semantic labels that bridge the gap between sensor inputs and meaningful object recognition, simplifying the overall system architecture while improving precision
Solution Approach 2:
The system segments the data processing into distinct components: capturing speech separately, capturing text separately, capturing images separately, and capturing sensor data separately. Each component is processed independently and then integrated, making the complex system more manageable and easier to implement while maintaining high identification accuracy
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A vehicle data relation device includes an internal audio/image data analyzer, configured to receive first data representing at least one of audio from within the vehicle or an image from within the vehicle; identify within the first data second data representing an audio indicator or an image indicator, wherein the audio indicator is human speech associated with a significance of an object external to the vehicle, and wherein the image indicator is an action of a human within the vehicle associated with a significance of an object external to the vehicle; an external image analyzer, configured to receive third data representing an image of a vicinity external to the vehicle; identify within the third data an object corresponding to at least one of the audio indicator or the video indicator; and an object data generator, configured to generate data corresponding to the object.