Road-Scene Human Intent Prediction Beyond Motion Vectors

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current systems, such as autonomous driving vehicles, struggle to accurately predict human behavior beyond motion vectors, leading to inferior results in anticipating pedestrian, cyclist, and motorist actions, especially in complex scenarios like crowded areas.

Innovation Solution

A system that uses a computing device to generate stimulus data from images or videos of road scenes, aggregates user responses to create statistical data, and trains machine learning models like neural networks to predict human behavior, allowing for more accurate anticipation of actions and movements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If motion vector extrapolation methods are used to predict human behavior, then the prediction process is simple and fast, but the prediction accuracy is inferior

Engineering Contradiction:
Improveprediction accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system segments human behavior prediction into multiple independent analysis dimensions: motion vectors, eye gaze direction, head orientation, and body posture. Each dimension is processed separately to extract specific behavioral cues, which are then integrated to form a comprehensive prediction. This segmentation allows the system to capture nuanced human intent that simple motion extrapolation misses, improving prediction accuracy without overwhelming system complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system transitions from analyzing only spatial motion dimensions to incorporating temporal and directional dimensions through eye tracking and posture analysis. By adding these new dimensions of observation, the system can predict human behavior more accurately, particularly for detecting intent to cross streets or change lanes, going beyond what motion vectors alone can provide.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Reliability

If comprehensive human behavior analysis is implemented, then prediction accuracy improves, but computational resources and processing time increase

Engineering Contradiction:
Improvebehavior prediction reliabilityVSAvoidcomputational energy consumption
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The system performs preliminary analysis by continuously tracking eye gaze direction and head orientation in real-time, preparing these data streams in advance. When a pedestrian is detected, the system already has accumulated eye and head position data ready for immediate integration with motion analysis, reducing the computational burden during critical prediction moments and improving response reliability.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system introduces intermediate processing layers that aggregate and pre-process visual data from multiple sources (eye tracking, head tracking, body posture) before final behavior prediction. These intermediaries organize raw data into meaningful features, reducing the computational complexity of the final prediction step while maintaining high reliability through comprehensive analysis.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS11987272B2System and method of predicting human interaction with vehicles
Publication Date: 2024.05.21 PERCEPTIVE AUTOMATA INC
  • US11987272B2 patent drawing
  • US11987272B2 patent drawing
  • US11987272B2 patent drawing

AI summary

Systems and methods for predicting user interaction with vehicles. A computing device receives an image and a video segment of a road scene, the first at least one of an image and a video segment being taken from a perspective of a participant in the road scene and then generates stimulus data based on the image and the video segment. Stimulus data is transmitted to a user interface and response data is received, which includes at least one of an action and a likelihood of the action corresponding to another participant in the road scene. The computing device aggregates a subset of the plurality of response data to form statistical data and a model is created based on the statistical data. The model is applied to another image or video segment and a prediction of user behavior in the another image or video segment is generated.