Visual Localization Model Training for Dynamic 3D Environments

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional learning-based visual localization systems are limited by static training samples that fail to reflect dynamic environments, provide partial coverage, require sensor modality and configuration similarity, and lack reliable confidence estimation, leading to limited deployment and portability.

Innovation Solution

A neural network model trained using 3D models derived from diverse data sources, including crowd-sourced feedback, to enhance environmental representation and adapt to dynamic changes, with confidence-aware predictions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional learning-based visual localization uses static training samples obtained during specific periods, then the model can be trained and deployed, but the training samples fail to reflect dynamic evolution of the target environment

Engineering Contradiction:
Improvereliability of pose estimationVSAvoidadaptability to dynamic environment
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent applies dynamics by implementing continuous retraining mechanisms where the neural network model is periodically retrained with newly collected training samples from the evolving environment. This transforms the static training process into a dynamic one, allowing the model to adapt to environmental changes while maintaining reliable pose estimation through ongoing updates.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system implements feedback by continuously collecting new training samples from the target environment and using them to retrain the model. This closed-loop feedback mechanism ensures the model learns from actual environmental evolution, improving both reliability and adaptability by incorporating real-world changes back into the training process.

Inventive Principle:
Principle #23Feedback

2Ease of manufacture

If training samples are collected from limited locations and time periods, then data collection is feasible, but the training samples only provide partial coverage of the target environment

Engineering Contradiction:
Improveease of data collectionVSAvoidcoverage of target environment
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The patent applies universality by designing a multi-functional data collection system that gathers training samples through multiple channels: initial comprehensive collection, continuous ongoing collection, and crowd-sourced feedback from multiple mobile devices. This multi-functional approach ensures complete environmental coverage while maintaining ease of collection through automated processes.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system performs preliminary action by conducting an initial comprehensive data collection phase before deployment, establishing a baseline training set. This preliminary collection, combined with subsequent continuous collection, ensures complete environmental coverage is achieved before the model goes into production use.

Inventive Principle:
Principle #10Preliminary action

3Manufacturing precision

If conventional visual localization requires similarity in modality and sensor configuration between training and input data, then training can proceed, but deployment flexibility and portability are limited

Engineering Contradiction:
Improveprecision of training-data matchingVSAvoiddeployment flexibility
Core Design Contradiction:
Manufacturing precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies parameter changes by implementing data augmentation techniques that systematically vary parameters such as lighting conditions, camera angles, and environmental configurations in training samples. This exposes the model to diverse parameter variations during training, enabling it to maintain precision across different sensor configurations and deployment scenarios.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system achieves universality by collecting and training on data from diverse sensor configurations and modalities. The multi-source training approach creates a universal model that can adapt to different camera types, sensor setups, and input modalities, significantly improving deployment flexibility while maintaining training precision.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Device complexity

If conventional learning-based visual localization lacks confidence estimation, then the system is simpler, but practical usage is limited without reliable confidence prediction

Engineering Contradiction:
Improvecomplexity of localization systemVSAvoidconfidence estimation reliability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent applies segmentation by separating the neural network into distinct functional components: one branch for pose estimation and another for confidence prediction. This modular segmentation allows the system to maintain relative simplicity while adding reliable confidence estimation, as each component can be optimized independently.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The confidence estimation module acts as an intermediary that evaluates the reliability of pose predictions before they are used for navigation decisions. This intermediary layer provides reliable confidence information with minimal additional complexity, enabling practical usage by filtering or weighting predictions based on their confidence scores.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12475590B2Devices and methods for visual localization
Publication Date: 2025.11.18 HUAWEI TECH CO LTD
  • US12475590B2 patent drawing
  • US12475590B2 patent drawing
  • US12475590B2 patent drawing

AI summary

A computing device is provided for supporting a mobile device to perform a visual localization in a target environment. The computing device is configured to obtain and modify one or more 3D models of the target environment to generate a set of 3D models. The computing device is further configured to determine a set of training samples based on the set of 3D models. Each training sample comprises an image derived from the set of the 3D models and corresponding pose information. The computing device is further configured to train a neural network model based on the set of training samples, so that the trained neural network model is configured to predict pose information for an image captured in the target environment and to output a confidence value associated with the predicted pose information.