Vehicle Video Scene Extraction Using Rules and ML Estimation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing technologies do not effectively extract videos of specific scenes, such as interruption scenes, from accumulated video data.
Innovation Solution
An information processing apparatus that specifies periods within videos captured by vehicle-installed image capturing apparatuses where recognized information matches rules for specific scenes, acquires datasets combining recognized information and correct labels, and generates trained models to estimate whether videos in specific periods are of the specific scenes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If rule-based specification is used to identify specific scenes, then the extraction process becomes systematic and reproducible, but the accuracy of scene identification is insufficient
Solution Approach 1:
The patent combines rule-based specification with machine learning-based estimation to create a hybrid system. The rule-based unit provides systematic and reproducible scene candidate identification, while the machine learning unit enhances accuracy by learning from labeled video data. This merging resolves the contradiction by integrating the reliability of rule-based methods with the precision of data-driven approaches.
Solution Approach 2:
The system creates a composite approach by integrating two different methodologies: rule-based logic and machine learning models. Just as composite materials combine different substances to achieve superior properties, this composite system combines the structured reliability of rules with the adaptive accuracy of machine learning to achieve both systematic processing and high identification accuracy.
2Measurement precision
If machine learning is applied to extract specific scenes from accumulated videos, then the accuracy of scene identification improves, but the computational load increases
Solution Approach 1:
The patent segments the video processing task into two distinct stages: rule-based specification that identifies candidate scenes with low computational cost, and machine learning-based estimation that refines identification accuracy only for the specified candidates. This segmentation reduces overall computational load by limiting the application of intensive machine learning to a subset of videos rather than processing all accumulated videos through the full ML pipeline.
Solution Approach 2:
The rule-based specification acts as a preliminary action that pre-filters and identifies candidate specific scenes before applying machine learning. This preliminary step reduces the volume of data requiring computationally intensive processing, thereby lowering the overall computational load while maintaining high accuracy through subsequent ML-based estimation on the pre-selected candidates.
3Productivity
If only rule-based methods are used for scene extraction, then the processing speed is maintained, but the ability to accurately identify complex scenes is insufficient
Solution Approach 1:
The system segments scene identification into rule-based detection for straightforward cases and machine learning-based estimation for complex scenes. This segmentation allows the system to maintain high processing speed for simple scenes using fast rule-based methods while accurately handling complex scenes through machine learning, thus resolving the contradiction between speed and accuracy for different scene types.
Solution Approach 2:
The patent applies different processing qualities to different scenes: rule-based methods are used for simple, easily identifiable scenes to maintain processing speed, while machine learning-based estimation is applied to complex scenes requiring higher identification accuracy. This local differentiation of processing quality optimizes both overall processing speed and complex scene identification accuracy.
Data Source
AI summary
Provided is an information processing apparatus (10) including: a specification means (11) for specifying, from a period within a video captured by an image capturing apparatus installed in a vehicle, a period in which information being recognized from a video matches with a rule according to a specific scene; an acquisition means (12) for acquiring a data set being a combination of information being recognized from a video in the period specified by the specification means and a correct label indicating whether the video in the period is a video of the specific scene; and a generation means (13) for executing learning, based on the data set acquired by the acquisition means, and generating a trained model for estimating whether a video in a specific period is a video of the specific scene.


