Adaptive Video Sampling for Hidden Context Prediction in Traffic Entities
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional systems fail to accurately predict the motion of non-stationary traffic objects like pedestrians and bicyclists, leading to unnatural vehicle movements, as they rely on kinematic methods that do not account for the intentions and awareness of these entities, and the generation of training data for machine learning models is time-consuming and expensive.
Innovation Solution
A system that collects video data from autonomous vehicles, samples video frames to create training datasets with user-annotated hidden context attributes, evaluates model performance, and adapts sampling to improve predictions by retraining the model with specific datasets for dimension attributes where performance is poor.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If kinematic methods are used to predict motion of non-stationary objects, then the system is simple to implement, but the prediction accuracy is poor
Solution Approach 1:
The patent replaces traditional kinematic methods with machine learning models that analyze video data to predict hidden context and future positions of traffic entities. The system uses neural networks trained on sampled video frames to infer intentions and awareness of pedestrians and bicyclists, substituting mechanical prediction algorithms with data-driven intelligent systems.
Solution Approach 2:
The patent introduces an intermediary machine learning model that processes video data and predicts hidden context attributes between the observed video frames and the final motion prediction. This intermediary layer enables the system to infer unobservable states (intentions, awareness) that directly influence future motion, bridging the gap between current observations and future predictions.
2Measurement precision
If user input is collected to generate labeled training datasets, then the model accuracy improves, but the time and cost increase
Solution Approach 1:
The patent performs preliminary action by pre-sampling video frames and pre-processing data before full model training. The system selectively samples video frames based on criteria such as frame rate, temporal intervals, and scene changes, preparing training data in advance to reduce the overall training time and computational resources required.
Solution Approach 2:
The patent applies partial action by using selective sampling of video frames rather than processing complete video sequences. The system samples only necessary frames that provide sufficient information for training, reducing the amount of data processing while maintaining model accuracy. This partial processing approach significantly reduces time and computational costs.
3Reliability
If complete video data is used for training, then the training comprehensiveness improves, but the data processing time and computational resources increase
Solution Approach 1:
The patent extracts only the essential and representative video frames from complete video sequences for training purposes. The sampling mechanism identifies and extracts key frames that capture critical moments and scenarios, removing redundant data while preserving training comprehensiveness. This extraction process maintains reliability while improving processing efficiency.
Solution Approach 2:
The patent segments video data into discrete, sampled frames rather than processing continuous video streams. By dividing the video data into selective temporal segments and processing only those segments that meet sampling criteria, the system achieves efficient processing while maintaining comprehensive training coverage across different scenarios and conditions.
Data Source
AI summary
A vehicle collects video data of an environment surrounding the vehicle including traffic entities, e.g., pedestrians, bicyclists, or other vehicles. The captured video data is sampled and presented to users to provide input on a traffic entity's state of mind. The user responses on the captured video data is used to generate a training dataset. A machine learning based model configured to predict a traffic entity's state of mind is trained with the training dataset. The system determines input video frames and associated dimension attributes for which the model performs poorly. The dimension attributes characterize stimuli and/or an environment shown in the input video frames. The system generates a second training dataset based on video frames that have the dimension attributes for which the model performed poorly. The model is retrained using the second training dataset and provided to an autonomous vehicle to assist with navigation in traffic.


