Image Network Learning With Dynamic Masked Pose Refinement
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing self-supervised depth estimation methods for autonomous driving in dynamic environments suffer from inaccurate pose estimation due to image sequence matching issues, leading to trajectory drifting and reduced learning performance.
Innovation Solution
A method and device that generate a dynamic mask to filter out dynamic regions from image features, refining the pose estimation by training a synthetic image model using a depth network and pose network, which improves depth and pose accuracy in dynamic environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of manufacture
If self-supervised depth estimation uses image sequence matching in dynamic environments, then depth information can be estimated without large amounts of GT depth maps, but pose estimation accuracy deteriorates due to dynamic regions causing trajectory drifting
Solution Approach 1:
The patent segments the image sequence into static and dynamic regions using optical flow analysis. By dividing the scene into these two components, the system can selectively process only static regions for pose estimation, thereby eliminating the negative impact of dynamic objects while preserving the self-supervised learning approach that reduces training costs.
Solution Approach 2:
The patent extracts and removes dynamic regions from the image sequence before performing pose estimation. By taking out the harmful dynamic components that cause trajectory drifting, the system maintains accurate pose estimation while still utilizing the cost-effective self-supervised learning methodology.
2Loss of information
If pose model accumulates results for trajectory estimation in dynamic environments, then complete trajectory information can be obtained, but trajectory drifting occurs between predicted trajectory and GT trajectory
Solution Approach 1:
The patent segments the trajectory estimation process by separately handling static and dynamic regions. Optical flow analysis divides the scene, and only static region matches are accumulated for trajectory construction. This segmentation ensures complete trajectory information from static structures while preventing dynamic regions from causing accuracy degradation.
Solution Approach 2:
The patent introduces optical flow analysis as an intermediary step between image matching and trajectory accumulation. This intermediary filters out dynamic regions before they can contaminate the trajectory data, allowing complete trajectory information to be gathered from static elements while maintaining high accuracy.
3Productivity
If depth network and pose network are trained simultaneously using synthetic images, then learning efficiency is improved, but accuracy deteriorates when dynamic regions are not properly filtered
Solution Approach 1:
The patent segments the training data by identifying and masking dynamic regions before feeding images to the joint depth-pose network. This segmentation allows efficient simultaneous training while ensuring that only reliable static region information contributes to accuracy, resolving the contradiction between learning efficiency and precision.
Solution Approach 2:
The patent performs preliminary filtering of dynamic regions using optical flow analysis before the main training process. By removing harmful dynamic content in advance, the joint training of depth and pose networks can proceed efficiently without being corrupted by dynamic region mismatches, thereby maintaining high accuracy.
Data Source
AI summary
A method for controlling autonomous driving of a vehicle is introduced. The method may comprise, outputting, by a depth network, an inference depth from a sequence image, outputting, by a pose network and based on the sequence image, an initial inference pose, generating, based on a synthetic depth, a dynamic mask, wherein the synthetic depth is generated based on the inference depth and the initial inference pose, generating, by the pose network and based on the sequence image and the dynamic mask, a refined inference pose, based on the sequence image, the inference depth, and the refined inference pose, training a synthetic image model may comprise the depth network and the pose network to generate a synthetic image, outputting a signal associated with the synthetic image, and controlling, based on the signal, autonomous driving of the vehicle.


