Work Machine Video Frame Masking for Object Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing work machine controllers struggle to accurately identify and exclude components of the work machine from video frames, particularly those associated with moving implements, which can interfere with object detection and lead to false alerts or missed detections of nearby objects, potentially causing collisions.

Innovation Solution

A method and device that process video frames to determine apparent motion, generate a composite video frame, and create a mask to exclude work machine components, allowing for accurate object detection and alerting operators or adjusting the work machine's path to avoid collisions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If video frames are processed to include all visible components, then complete scene information is captured, but object detection accuracy deteriorates due to interference from work machine components

Engineering Contradiction:
Improveobject detection accuracyVSAvoidscene information completeness
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The video frame is segmented into multiple regions: a mask region containing work machine components that are excluded from detection, and a non-mask region where object detection is performed. This segmentation allows the system to process complete scene information while preventing interference from machine components by spatially separating their influence.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent extracts and isolates the mask region containing work machine components from the overall video frame. By extracting this interfering element and applying a mask, the system removes its negative impact on object detection while preserving the rest of the scene information for accurate detection.

Inventive Principle:
Principle #2Taking out (Extraction)

2Measurement precision

If motion processing is applied to identify moving implements, then dynamic components can be excluded, but processing complexity increases

Engineering Contradiction:
Improvecomponent identification accuracyVSAvoidvideo processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary motion processing on video frames to identify moving implements before the main object detection process. By pre-identifying dynamic components and incorporating them into the mask region, the system simplifies the subsequent detection process and avoids false positives from moving parts.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements dynamic mask updating by processing video frames to detect motion and automatically adjusting the mask region to include newly identified moving implements. This dynamic adaptation allows the system to handle changing scenes without increasing structural complexity.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS10949685B2Excluding a component of a work machine from a video frame based on motion information
Publication Date: 2021.03.16 CATERPILLAR INC
  • US10949685B2 patent drawing
  • US10949685B2 patent drawing
  • US10949685B2 patent drawing

AI summary

A controller may process a plurality of video frames to determine an apparent motion of each pixel of one or more pixels or each group of pixels of one or more groups of pixels of each video frame of the plurality of video frames. The controller may select one or more processed video frames, of the plurality of processed video frames, that correspond to a duration of time and may generate a composite video frame based on the one or more processed video frames. The controller may generate a video frame mask based on the composite video frame and may obtain additional video data that includes at least one additional video frame. The controller may cause the video frame mask to be applied to the at least one additional video frame and may cause the at least one additional video frame to be processed using an object detection technique.