Intersection Scenario Retrieval Using Action Unit Annotation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing computer image analysis systems for classifying and labeling objects and their motions in video streams are inefficient and data-intensive, as they rely on finite offline-trained object labels.
Innovation Solution
A computer-implemented method for intersection scenario retrieval that trims video streams into clips, annotates ego vehicles and dynamic objects with action units describing their motions, and retrieves relevant scenarios from an electronic dataset using a neural network to control the presentation of intersection scenario video clips.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If existing systems use finite offline-trained object labels for classification, then the system structure is simple, but the system becomes inefficient and data-intensive
Solution Approach 1:
The patent segments video streams into discrete video clips representing specific intersection scenarios. Each clip is independently annotated with action units, allowing the system to process and retrieve specific scenarios without handling entire video datasets, thereby improving efficiency while reducing data intensity.
Solution Approach 2:
The patent introduces action units as an intermediary layer between raw video data and classification labels. These action units (combining nouns for objects and verbs for motions) serve as a structured representation that enables efficient querying and retrieval without requiring extensive raw training data for each specific scenario.
2Reliability
If the system stores and processes complete video streams, then all intersection scenarios are captured, but the retrieval time and processing load increase significantly
Solution Approach 1:
The patent divides continuous video streams into discrete video clips, each representing a specific intersection scenario. This segmentation allows the system to store only relevant scenario clips rather than complete video streams, enabling fast retrieval based on action unit queries while maintaining comprehensive scenario coverage.
Solution Approach 2:
The patent performs preliminary annotation of video clips with action units during the indexing phase. This preliminary action enables the system to quickly retrieve relevant clips by querying action units without needing to analyze entire video streams at retrieval time, significantly reducing retrieval time while maintaining complete scenario coverage.
3Measurement precision
If the system uses detailed action units to describe ego vehicles and dynamic objects, then the classification precision improves, but the complexity of the annotation system increases
Solution Approach 1:
The patent introduces action units as a standardized intermediary notation system that combines nouns (for ego vehicles and dynamic objects) and verbs (for motions and interactions). This structured intermediary layer enables precise classification by capturing specific intersection scenarios without requiring complex custom annotations for each case, as the same action unit vocabulary can be reused across different scenarios.
Solution Approach 2:
The patent creates a universal action unit vocabulary that can describe multiple types of objects (ego vehicles, dynamic objects) and their various motions and interactions. This universal system reduces annotation complexity by using the same set of nouns and verbs across different scenarios, eliminating the need for scenario-specific annotation schemes while maintaining high classification precision.
Data Source
AI summary
A system and method for performing intersection scenario retrieval that includes receiving a video stream of a surrounding environment of an ego vehicle. The system and method also include analyzing the video stream to trim the video stream into video clips of an intersection scene associated with the travel of the ego vehicle. The system and method additionally include annotating the ego vehicle, dynamic objects, and their motion paths that are included within the intersection scene with action units that describe an intersection scenario. The system and method further include retrieving at least one intersection scenario based on a query of an electronic dataset that stores a combination of action units to operably control a presentation of at least one intersection scenario video clip that includes the at least one intersection scenario.


