An adaptive hierarchical robot localization method and device for degenerate scenarios

CN122174210APending Publication Date: 2026-06-09TIANFU YONGXING LAB
View PDF 4 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
TIANFU YONGXING LAB
Filing Date
2026-05-06
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Traditional robot localization systems struggle to effectively address insufficient observation information in degraded scenarios, leading to ill-posed pose estimation. Furthermore, sensor switching strategies lack predictive capabilities, which can easily cause pose jumps and information waste.

Method used

An adaptive hierarchical robot localization method is adopted. By monitoring the data quality of heterogeneous sensors in real time, four degradation levels are classified: visual, geometric, hybrid, and complete. For each level, a special strategy is designed, including feature matching, singular value decomposition, performance prediction network, and the use of inertial odometry, to achieve dynamic scheduling and self-improvement of sensors.

Benefits of technology

It improves the system's robustness in complex environments, avoids overreaction or underreaction of sensors, slows down the system's slide towards complete degradation, achieves smooth mode switching and continuous self-improvement, and enhances positioning accuracy and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122174210A_ABST
    Figure CN122174210A_ABST
Patent Text Reader

Abstract

This invention belongs to the field of robot localization and relates to an adaptive hierarchical robot localization method and apparatus for degraded scenarios. The method includes: real-time monitoring of the data quality acquired by heterogeneous sensors, mapping the states of the heterogeneous sensors to a unified quantization space, and outputting unique degradation level identifiers D1 to D4; for the D1 environment, uncertainty quantification is performed on feature matching, and the most contributing feature subset is selected; for the D2 environment, the degraded subspace is located through singular value decomposition; for the D3 environment, the expected error of the heterogeneous sensor engine is estimated through a performance prediction network, and continuous confidence weighting is used instead of hard switching; for the D4 environment, a cross-platform generalized pure inertial odometry is provided through a three-level architecture of pre-training-fine-tuning-online adaptation; and historical events are stored as empirical data. This achieves refined differentiation and directional processing of degraded scenarios and constructs an experience-driven closed-loop optimization mechanism.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot localization technology, and in particular to an adaptive hierarchical robot localization method and apparatus for degraded scenarios. Background Technology

[0002] Simultaneous Localization and Mapping (SLAM) technology is a core support for mobile robots to achieve autonomous navigation in unknown environments. However, traditional localization systems face severe challenges when robots operate in degraded scenarios—that is, scenarios where the perceptual information provided by the environment is insufficient to support stable pose estimation. The essence of the degradation problem is that the information entropy of the observed information is lower than the minimum threshold required for pose estimation. When the environment cannot provide sufficient geometric constraints, the pose estimation problem transforms from a well-posed problem into an ill-posed problem, and the optimization equation exhibits infinitely many solutions or no solution in the degradation direction.

[0003] Traditional multi-sensor fusion methods typically assume that sensors operate in ideal conditions over a long period. Once degradation occurs, they often resort to overall degradation (e.g., discarding vision sensors altogether) or simply switching to a backup sensor, ignoring the potential for directional, localized, and gradual degradation. Existing methods, when introducing inertial measurement units (IMUs), often impose prior constraints on all degrees of freedom, suppressing environmental observations even in non-degrading directions, resulting in information waste. Traditional strategies often switch instantaneously when sensor quality falls below a threshold, easily causing pose jumps, and the switching decisions lack the ability to predict the future state of the sensors. Existing pure inertial odometry relies heavily on supervised training using data from specific platforms, requiring re-acquisition and labeling across devices, which is extremely costly; furthermore, the model remains unchanged during degradation operations, unable to learn and improve from degraded data. Summary of the Invention

[0004] To address the aforementioned technical problems, this invention provides an adaptive hierarchical robot localization method for degraded scenarios, employing the following technical solution, including the following steps:

[0005] The system monitors the data quality acquired by heterogeneous sensors in real time, maps the state of the heterogeneous sensors to a unified quantization space, and outputs unique degradation level identifiers D1~D4, representing visual degradation, geometric degradation, mixed degradation, and complete degradation, respectively.

[0006] For visual degradation level D1 environments, uncertainty quantification is performed on feature matching, and the most contributing feature subset is selected.

[0007] For geometrically degraded D2 level environments, the degraded subspace is located by singular value decomposition, and prior constraints of the inertial measurement unit are introduced only in the degraded direction, while the original sensor confidence observations are retained in the non-degraded direction.

[0008] For hybrid degradation D3 level environments, the expected error of the heterogeneous sensor engine is estimated through a performance prediction network, continuous confidence weighting is used instead of hard switching, and the sensor operating mode is scheduled.

[0009] For fully degraded D4 level environments, a three-level architecture of pre-training-fine-tuning-online adaptation is used to provide a cross-platform generalizable pure inertial odometry, and to achieve unsupervised continuous self-improvement during the degradation process;

[0010] The degradation events encountered in historical operations, the strategies adopted, and the resulting effects are structured and stored as empirical data, and the hyperparameters of each level of strategy are optimized offline using the empirical data.

[0011] Preferably, the step of real-time monitoring of the data quality acquired by heterogeneous sensors, mapping the state of the heterogeneous sensors to a unified quantization space, and outputting unique degradation level identifiers D1~D4, representing visual degradation, geometric degradation, mixed degradation, and complete degradation respectively, specifically includes:

[0012] Real-time monitoring of the data quality acquired by the heterogeneous sensor, and output of the health score and degradation type label of the heterogeneous sensor;

[0013] Based on the health score and the degradation type label, the current environment is classified into one of four degradation levels through decision tree logic: visual degradation level D1, geometric degradation level D2, mixed degradation level D3 and complete degradation level D4, and a unique activated intervention level identifier is output.

[0014] Spatiotemporal alignment calibration is performed on the heterogeneous sensor.

[0015] Preferably, the step of quantifying the uncertainty of feature matching and selecting the most contributing feature subset for visual degradation level D1 environments specifically includes:

[0016] For visual degradation level D1 environments, feature uncertainty is quantified based on covariance estimation to generate a feature candidate set and covariance matrix;

[0017] Based on the feature candidate set and the covariance matrix, select the preset N features with the largest information gain to participate in pose optimization;

[0018] Adaptive threshold dynamic adjustment and feature activation are performed.

[0019] Preferably, the steps for locating the degenerate subspace through singular value decomposition for a geometrically degraded D2 level environment, introducing prior constraints for the inertial measurement unit only in the degenerate direction, and retaining the original sensor confidence observations in the non-degenerate direction specifically include:

[0020] For D2 level geometrically degraded environments, the direction of degradation is identified through singular value decomposition.

[0021] Actively select to introduce additional constraint sources only in the degradation direction, while retaining the full confidence of the original sensor in the non-degradation direction;

[0022] Degradation release detection and reinforcement factor withdrawal were performed.

[0023] Preferably, the steps of estimating the expected error of the heterogeneous sensor engine through a performance prediction network, replacing hard switching with continuous confidence weighting, and scheduling the sensor operating modes for a hybrid degradation D3 level environment specifically include:

[0024] Real-time performance prediction of sensor engine for hybrid degradation D3 level environments;

[0025] Based on the performance prediction results and the actual residual statistics within the current sliding window, preset sensor factors are dynamically activated / frozen in the factor graph.

[0026] Actively scheduling the operating modes of the sensors enables them to complement each other in both time and space.

[0027] Preferably, the steps for providing a cross-platform generalized pure inertial odometry through a three-level architecture of pre-training-fine-tuning-online adaptation for a fully degraded D4 level environment, and achieving unsupervised continuous self-improvement during the degradation process, specifically include:

[0028] For fully degraded D4 level environments, cross-platform heterogeneous IMU base model pre-training is performed;

[0029] Adapt the cross-platform heterogeneous IMU base model to the specific dynamic characteristics of the current robot;

[0030] Continuously optimize the IMU odometry under unsupervised conditions so that it can operate, learn, and improve while in the degradation range.

[0031] Preferably, the step of structurally storing the degradation events encountered in historical operations, the strategies adopted, and the resulting effects as empirical data, and using the empirical data to offline optimize the hyperparameters of strategies at each level specifically includes:

[0032] Build a library of experience replays of degradation events;

[0033] Based on the degradation event experience replay library, using historical data as the training set, the adjustable parameters in the hierarchical strategies at each level are optimized.

[0034] Perform hot loading of strategy parameters and rolling version updates.

[0035] To address the aforementioned technical problems, this invention also provides an adaptive hierarchical robot localization device for degraded scenarios, employing the following technical solution, including:

[0036] The mapping module is used to monitor the data quality acquired by heterogeneous sensors in real time, map the state of heterogeneous sensors to a unified quantization space, and output unique degradation level identifiers D1~D4, representing visual degradation, geometric degradation, mixed degradation and complete degradation, respectively.

[0037] The quantization module is used to perform uncertainty quantization on feature matching and select the most contributing feature subset for visual degradation D1 level environments.

[0038] The positioning module is used to locate the degenerate subspace in a geometrically degenerate D2 level environment through singular value decomposition, and introduces inertial measurement unit prior constraints only in the degenerate direction, while retaining the original sensor confidence observations in the non-degenerate direction.

[0039] The scheduling module is used to estimate the expected error of the heterogeneous sensor engine through a performance prediction network for hybrid degradation D3 level environments, replace hard switching with continuous confidence weighting, and schedule the sensor operating modes.

[0040] The fine-tuning module is designed for fully degraded D4 level environments. Through a three-level architecture of pre-training-fine-tuning-online adaptation, it provides cross-platform generalized pure inertial odometry and achieves unsupervised continuous self-improvement during the degradation process.

[0041] The optimization module is used to structure and store the degradation events encountered in historical operations, the strategies adopted, and the resulting effects as empirical data, and to use the empirical data to optimize the hyperparameters of strategies at all levels offline.

[0042] To address the aforementioned technical problems, the present invention also provides a computer device that employs the technical solution described below, comprising a memory and a processor. The memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the aforementioned adaptive hierarchical robot localization method for degraded scenarios.

[0043] To address the aforementioned technical problems, the present invention also provides a computer-readable storage medium, which employs the technical solution described below. The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the aforementioned adaptive hierarchical robot localization method for degraded scenarios.

[0044] Compared with the prior art, the present invention has the following main advantages:

[0045] (1) Degradation is clearly divided into four levels: visual, geometric, mixed, and complete. Specialized coping strategies are designed for each level to achieve targeted measures, avoid over- or under-response, improve the robustness of the system in complex environments, and realize the fine differentiation and targeted processing of degradation scenarios.

[0046] (2) For visual degradation, the most contributing features are retained instead of being discarded entirely; for geometric degradation, constraints are introduced only in the direction of degradation instead of accepting inertial priors in their entirety. This approach of partial reliability and partial constraint effectively preserves the still valuable environmental information, delays the system's slide toward complete degradation, avoids the simple discarding of unreliable sensors, and instead selectively utilizes their effective information.

[0047] (3) In the hybrid degradation scenario, continuous confidence weighting is used instead of hard switching, eliminating pose jump caused by sudden changes in sensor state; at the same time, a performance prediction network is introduced to estimate the expected error, so that the fusion weight is forward-looking rather than remedial after the fact. The three-level architecture of pure inertial odometry solves the cross-platform generalization problem and continuously improves itself in the degradation process, getting rid of the dependence on expensive labeled data and realizing the smoothness and learnability of mode switching.

[0048] (4) The historical degradation events, response strategies and effects are stored in a structured manner and used for offline hyperparameter tuning, so that the system has the ability to remember and evolve. As the running data accumulates, the strategies at all levels can continuously approach the optimal configuration, realize the upgrade from rule-driven to data and rule hybrid-driven, and build an experience-driven closed-loop optimization mechanism. Attached Figure Description

[0049] To more clearly illustrate the solutions in this invention, the accompanying drawings used in the description of the embodiments of this invention will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0050] Figure 1 This is a flowchart of an embodiment of the adaptive hierarchical robot localization method for degraded scenarios according to the present invention;

[0051] Figure 2 This is a schematic diagram of a structure of an embodiment of the adaptive hierarchical robot localization device for degraded scenarios of the present invention;

[0052] Figure 3 This is a schematic diagram of the structure of an embodiment of the computer device of the present invention. Detailed Implementation

[0053] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains; the terminology used herein in the specification is for the purpose of describing particular embodiments only and is not intended to limit the invention; the terms "comprising" and "having," and any variations thereof, in the specification, claims, and foregoing drawings are intended to cover non-exclusive inclusion. The terms "first," "second," etc., in the specification, claims, or foregoing drawings are used to distinguish different objects and not to describe a particular order.

[0054] In this document, the term "embodiment" means that a particular feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of the invention. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a separate or alternative embodiment mutually exclusive with other embodiments. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described herein can be combined with other embodiments.

[0055] To enable those skilled in the art to better understand the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings.

[0056] It should be noted that the adaptive hierarchical robot localization method for degradation scenarios provided in the embodiments of the present invention is generally executed by a server / terminal device, and correspondingly, the adaptive hierarchical robot localization device for degradation scenarios is generally set in the server / terminal device.

[0057] It should be understood that the number of terminal devices, networks, and servers is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be used.

[0058] Degradation scenarios include, but are not limited to: low-texture walls, drastic lighting changes, and smoke obstruction in visual SLAM; long straight corridors, symmetrical structures, and dynamic crowds in laser SLAM; and hybrid degradation caused by sensor failure in multi-sensor fusion systems.

[0059] Example 1

[0060] Please refer to Figure 1 The diagram illustrates a flowchart of an embodiment of the adaptive hierarchical robot localization method for degraded scenarios according to the present invention. The adaptive hierarchical robot localization method for degraded scenarios includes the following steps:

[0061] Step S1: Monitor the data quality acquired by the heterogeneous sensors in real time, map the state of the heterogeneous sensors to a unified quantization space, and output unique degradation level identifiers D1~D4, representing visual degradation, geometric degradation, mixed degradation, and complete degradation, respectively.

[0062] In this embodiment, the electronic device (e.g., a server / terminal device) on which the adaptive hierarchical robot localization method for degradation scenarios runs can receive adaptive hierarchical robot localization requests for degradation scenarios via wired or wireless connections. It should be noted that the aforementioned wireless connection methods may include, but are not limited to, 3G / 4G / 5G connections, WiFi connections, Bluetooth connections, WiMAX connections, Zigbee connections, UWB (ultra wideband) connections, and other currently known or future-developed wireless connection methods.

[0063] In this embodiment, the heterogeneous sensors include, but are not limited to, three core types of sensors: cameras, LiDAR, and inertial measurement units (IMU).

[0064] In this embodiment, step S1 may specifically include the following steps:

[0065] S11 monitors the data quality acquired by heterogeneous sensors in real time and outputs the health score and degradation type label of the heterogeneous sensors.

[0066] A multi-index weighted fusion evaluation model is adopted to score each frame of sensor data independently.

[0067] For visual sensors, a three-channel quality evaluation vector is constructed: .

[0068] in: The visual sensor quality evaluation vector consists of three dimensions. The method for calculating the lighting quality score is as follows: convert the RGB image to grayscale and calculate the entropy value of the pixel brightness histogram. and normalized to interval, , The reference entropy value under full illumination.

[0069] For sharpness scoring, the variance of the Laplacian operator is used for calculation: , Let Laplace's response variance be... The threshold for determining clarity / fuzziness. When... When the score is below the threshold, it approaches 0.

[0070] To score feature richness, the number of ORB (OrientedFAST and Rotated BRIEF) feature points extracted in the current frame is calculated. Compared with the preset expected quantity The ratio and truncation: .

[0071] The overall health of a visual sensor is defined as the weighted product of three factors:

[0072] . Overall health status of visual sensors.

[0073] Weighting coefficient Obtained through offline calibration, meeting the requirements The default setting is The significance of using a product form instead of a weighted average is that when any score approaches 0, the overall health score quickly drops to zero, which aligns with the physical reality that "visual sensors fail as a whole under severe degradation."

[0074] For LiDAR sensors, a geometric degradation detection model is constructed. The environmental structure factor for the current point cloud frame is defined: . : Environmental structural factor, used to quantify the degree of geometric degradation.

[0075] in These are the three eigenvalues ​​of the covariance matrix of the currently scanned point cloud. This ratio measures the degree of concentration of the point cloud distribution along its principal directions. In structure-rich environments, the three eigenvalues ​​are close to each other. In geometrically degraded environments such as corridors and open plazas, , .

[0076] Lidar health status is defined as: . Lidar sensor health status. The closer it is to 1, the more degenerate the geometric structure (such as a corridor). Approaching 0. This design allows for performance even with severe geometric degradation. Approaching 0, it accurately reflects the unreliability of Lidar odometers in this environment.

[0077] For IMU sensors, the primary function is to detect the degree of motion excitation. Constructing IMU excitation factors: .in As the IMU motion excitation factor, This is the measured value of angular velocity. The acceleration measurement value, This is an estimate of the gravity vector. and This is a normalized reference value. IMU health status, and Positive correlation; the more intense the exercise, the higher the score. IMU Health It is lower when at rest (the integral is prone to drift) and higher when there is sufficient motion stimulation.

[0078] Step S11 establishes a unified quantization framework, mapping the degradation states of the three types of heterogeneous sensors to the same... The scoring space provides comparable and computable quantitative inputs for subsequent steps, avoiding the difficulty of subjectively judging which sensor is more reliable.

[0079] The purpose of step S11 is to monitor the raw data quality of the three core sensors—camera, LiDAR, and IMU—in real time, and output the health score and degradation type label of each sensor to provide input basis for subsequent hierarchical decision-making.

[0080] S12, based on health scores and degradation type labels, classifies the current environment into one of four degradation levels through decision tree logic: visual degradation level D1, geometric degradation level D2, mixed degradation level D3, and complete degradation level D4, and outputs a unique activated intervention level identifier.

[0081] Construct a four-output decision tree classifier, using a priority masking rule for the decision logic:

[0082] Define binary degenerate vector Initialize it as a vector of all zeros.

[0083] Rule 1 (Complete Degradation D4 Decision): If and and ,but Set all other levels to 0 and output D4 directly.

[0084] Rule 2 (D3 determination of mixed degradation): If Rule 1 is not satisfied, and ( and (simultaneously established) or ( and If the alternating fluctuations exceed the set frequency threshold, then Output D3.

[0085] Rule 3 (Geometric Degeneration D2 Decision): If Rules 1-2 are not satisfied, and and ,but Output D2.

[0086] Rule 4 (Visual Degradation D1 Judgment): If Rules 1-3 are not satisfied, and and ,but Output D1.

[0087] Rule 5 (No Degeneracy): If none of the above conditions are met, then Output D0 (nominal operating condition).

[0088] in, : Degradation level identifier vector. : The threshold that triggers complete degradation of camera / Lidar health. The health threshold for determining whether the IMU is still usable. The rest... The symbols correspond to the threshold values ​​for different levels of degradation.

[0089] threshold Statistical calibration was performed through offline experiments. For example, , wait.

[0090] Step S12 discretizes the continuous health scores into distinct, mutually exclusive degradation levels. This hard classification strategy provides a clear basis for mode switching in hierarchical localization, avoiding frequent oscillations between multiple strategies. The tree structure of the decision tree naturally conforms to the degradation screening logic from heavy to light and from global to local.

[0091] The function of step S12 is to classify the current environment into one of the four degradation levels based on the sensor health vector output in step S11 through decision tree logic: visual degradation level (D1), geometric degradation level (D2), mixed degradation level (D3), and complete degradation level (D4), and output a uniquely activated intervention level identifier.

[0092] S13 performs spatiotemporal alignment calibration on heterogeneous sensors.

[0093] It adopts a hybrid architecture of hardware triggering and software soft synchronization.

[0094] Time Synchronization: The master computer runs the PTPv2 (Precision Time Protocol Version 2, corresponding to the IEEE 1588-2008 standard) precision time protocol to establish a master-slave clock synchronization mechanism. All sensors that support hardware timestamps (such as industrial cameras and PTP-enabled LiDARs) directly use the MAC layer timestamp, achieving sub-microsecond accuracy; for sensors that do not support hardware timestamps, a linear interpolation soft synchronization algorithm is used.

[0095] For IMU data streams, establish an asynchronous sampling timestamp alignment model:

[0096] Given a camera or Lidar frame timestamp We need to interpolate the IMU attitude pre-integration term at that moment. Let... and Distance The two most recent IMU sampling times correspond to the angular velocity measurements. , The acceleration measurement value is , Linear interpolation is used to obtain... Virtual measurement of time: , .

[0097] in, : Reference timestamp (e.g., image frame time). :distance The two most recent IMU sampling times. : The measured angular velocity value at the corresponding moment. : The acceleration measurement value at the corresponding moment. : obtained by interpolation Real-time virtual IMU measurement values.

[0098] Spatial synchronization (extrinsic parameter calibration): For camera-Lidar joint calibration, the mutual information maximization method is adopted. The objective function for mutual information between the point cloud reflectance intensity map and the image brightness map is constructed as follows:

[0099] .

[0100] in Let be the extrinsic transformation matrix to be determined (the optimal extrinsic transformation matrix from camera to LiDAR). Images taken by the camera, For Lidar point clouds, To convert Lidar point clouds based on external parameters The function projected onto the camera plane. The mutual information metric is determined iteratively using the gradient descent method.

[0101] Spatiotemporal alignment is a physical prerequisite for multi-sensor fusion. Without rigorous data synchronization, even the most advanced fusion algorithms cannot achieve unbiased estimates. Step S13 employs a hardware-software combined approach, striking a balance between accuracy and real-time performance. Mutual information calibration does not rely on manual calibration boards and can be continuously optimized online in natural environments, solving the problem of potential failure of traditional calibration methods in degraded scenarios.

[0102] The purpose of step S13 is to ensure that all sensor data participating in the fusion have a unified spatiotemporal reference, eliminating inconsistencies in observations caused by hardware trigger delays, processing delays, and clock drift. This step is a prerequisite for all subsequent fusion steps.

[0103] Step S2: For visual degradation level D1 environment, uncertainty quantification is performed on feature matching, and the most contributing feature subset is selected.

[0104] In this embodiment, step S2 may specifically include the following steps:

[0105] S21, for visual degradation level D1 environment, quantifies feature uncertainty based on covariance estimation, and generates feature candidate set and covariance matrix.

[0106] A learning-based covariance prediction network, named ThermalPoint-Lite, is employed. This network accepts a pair of matched feature points (pixel coordinates) from two consecutive image frames. , ) and local image patches centered on them , As input, output the 2×2 covariance matrix of the matched pair. .

[0107] Network architecture: A lightweight Siamese convolutional structure is adopted, with two branches sharing weights to extract image patch features. The feature maps are processed by a correlation layer to calculate the similarity tensor, and then the lower triangular elements of the covariance matrix are regressed through three fully connected layers.

[0108] Define the prediction loss function as the negative log-likelihood loss: .in, Let be the loss function for the covariance prediction network; , which are high-confidence true pixel coordinates obtained through high-precision equipment (such as motion capture systems) or subsequent optimization (such as BA (Bundle Adjustment)); , where is the location of the feature points predicted by the network; , where is the matching uncertainty covariance matrix output by the network; The Mahalanobis distance; The determinant of the covariance matrix represents the total volume of uncertainty. The significance of this loss function is that it both penalizes bias in predicted locations and encourages the network to output a larger covariance for unreliable matches (corresponding to a larger...). This enables the ability to quantify the uncertainty of something.

[0109] Traditional feature extraction algorithms (ORB, SIFT, etc.) treat all feature points equally, failing to distinguish between accidental matches in weakly textured regions and reliable matches in corner regions. Step S21 adds an uncertainty measure to each feature match, enabling the subsequent optimizer to automatically weight them based on the amount of information. This is the theoretical basis for achieving lightweight filtering—eliminating a large number of outliers based solely on uncertainty without relying on complex geometric verification.

[0110] The purpose of step S21 is to address the issue that in visually degraded (D1 level) environments, traditional feature extractors blindly track a large number of low-quality features, leading to decreased pose estimation accuracy or even optimization failure. The core function of this step is to calculate the uncertainty covariance for each potential feature matching pair, serving as the basis for subsequent screening.

[0111] S22, based on the feature candidate set and covariance matrix, select the preset N features with the largest information gain to participate in pose optimization.

[0112] This step constructs a combinatorial optimization problem. Let there be a total of... Each feature matches a candidate, and each match The associated covariance matrix is Furthermore, based on the pixel coordinates and inverse depth of the feature points, the contribution of each feature point to the Fisher information matrix for pose estimation can be pre-calculated through offline calibration. (6-DoF pose). The information matrix and the covariance matrix are inversely related.

[0113] Objective: From Select from the features Features ( This maximizes the determinant of the overall information matrix, i.e., maximizes: .

[0114] in , is the prior information matrix (from IMU pre-integration); For the selected feature subset; , for the first Fisher information matrix of the contribution of each feature to 6-DoF pose estimation; : Matrix determinant, used here to quantify the total amount of information.

[0115] This is an NP-hard combinatorial optimization problem. This step uses a greedy sequence selection algorithm to approximate the solution:

[0116] initialization Information matrix ;

[0117] for to :

[0118] Iterate through all unselected features Calculate marginal gain ;

[0119] Select the feature with the largest marginal gain. ;

[0120] Will join in ,renew ;

[0121] Output feature subset .

[0122] Step S22 redefines feature selection from an information theory perspective—selecting the most useful features, rather than the easiest to track. In visually degraded scenarios (such as darkness or smoke), the number of trackable features is already scarce, and blindly discarding low-quality features may result in no features being available. This method, however, quantifies the contribution of each feature to pose estimation through the Fisher information matrix, maximizing information utilization with limited computational resources.

[0123] The purpose of step S22 is to select N features with the largest information gain from hundreds of features based on the feature candidate set and covariance matrix generated in step S21 to participate in pose optimization, thereby significantly reducing the computational load while ensuring accuracy.

[0124] S23 performs adaptive threshold dynamic adjustment and feature activation.

[0125] Establish a piecewise linear adjustment function:

[0126] .

[0127] in, This represents the number of target features to be extracted in the current frame. For low / high health thresholds, , This is an empirical threshold; For the maximum / minimum number of target features, , .

[0128] Meanwhile, the initial response threshold of the covariance network It also adjusts according to health level:

[0129] .in As the baseline threshold, This is the adjustment coefficient. When visual health deteriorates, the threshold decreases, allowing more potential matches to enter the candidate pool. Quality is ensured through subsequent information gain filtering rather than hard threshold truncation.

[0130] Step S23 enables the system to respond flexibly to the severity of environmental conditions. In cases of slight visual degradation (such as slightly dim lighting), the system proactively reduces computational load to conserve energy; in cases of severe visual degradation (such as smoke), the system allocates more computational resources and relaxes access criteria, trading computation for perception. This dynamic adjustment mechanism is a microscopic manifestation of the hierarchical positioning concept within a single step.

[0131] The purpose of step S23 is to: based on the visual health score output in step S11 The number of target features in step S22 is dynamically adjusted. The initial matching threshold in step S21 achieves a balance between saving power when the operating conditions are good and maintaining accuracy when the operating conditions are poor.

[0132] Step S3: For a geometrically degraded D2 level environment, the degraded subspace is located by singular value decomposition, and prior constraints of the inertial measurement unit are introduced only in the degraded direction, while the original sensor confidence observations are retained in the non-degraded direction.

[0133] In this embodiment, step S3 may specifically include the following steps:

[0134] S31 identifies the direction of degradation through singular value decomposition for D2 level environments with geometric degradation.

[0135] This step is based on the null space analysis of the observation Jacobian matrix of the current sliding window.

[0136] Constructing the system observation matrix within the sliding window (15-dimensional states: position 3, attitude 3, velocity 3, gyroscope bias 3, accelerator bias 3). Matrix It is composed of visual reprojection Jacobian and Lidar registration Jacobian stacked together.

[0137] For matrix Perform singular value decomposition: .in, : The system observation Jacobian matrix within the sliding window (state dimension 15); The left singular matrix, singular value diagonal matrix, and right singular matrix obtained from SVD decomposition. , Singular values ​​are arranged in descending order.

[0138] Define degenerate set To meet Singular value index, This is a relative threshold (typically 1e-3). The right singular vectors corresponding to these minimal singular values... Zhang Cheng's degenerate subspace This refers to the state direction that cannot be constrained by current observations. : No. There are 1 singular value, and .

[0139] Furthermore, the degenerate subspace is projected onto specific state variables. A directional degeneracy degree vector is defined. , its first The elements are: .in, For the first The degree of degradation of each state component; A degenerate set containing indices of minimal singular values; For the first One right singular vector; For vectors In the Projected weights on each state component. The larger the value, the higher the value. The higher the degree of degradation of a state component.

[0140] Step S31 refines the macroscopic concept of geometric degradation to a specific state dimension (e.g., translational freedom along the corridor direction is severely degraded, but translational freedom perpendicular to the wall direction remains good). This refined understanding is a prerequisite for subsequent directional reinforcement rather than overall suppression, avoiding the erroneous practice of discarding globally useful information due to local degradation.

[0141] The purpose of step S31 is to address the issue that in a geometrically degraded (D2 level) environment, the LiDAR point cloud fails to provide constraints along certain directions (such as the axis of a corridor), causing pose estimation to diverge in those directions. The core of this step is to accurately identify which state dimensions are degrading, rather than simply determining that the LiDAR odometry is unreliable.

[0142] S32 actively selects to introduce additional constraint sources only in the degradation direction, while retaining the full confidence of the original sensor in the non-degradation direction.

[0143] This step implements the dynamic injection of adaptive prior factors within the factor graph optimization framework.

[0144] Suppose that the degenerate subspace identified in step S31 is composed of the basis vector matrix. This indicates that a degradation direction reinforcement factor is constructed, and its residual is defined as: .in The current state vector to be estimated is... This is the state predicted by the IMU forward propagation. The residuals are the prior factors, and their dimensions are equal to the dimensions of the degenerate subspace. .

[0145] The dimension of the residual is equal to the dimension of the degenerate subspace. Its physical meaning is: to project the state vector onto the degradation direction and constrain the projection to be as consistent as possible with the projection of the IMU prediction.

[0146] Corresponding information matrix Set to: .

[0147] in, The constraint strength in the degradation direction is taken as an empirical value of 100-1000 (much greater than the information content of ordinary observations). This means that in the degradation direction, the optimizer will highly trust the short-time integral prediction of the IMU; while in the non-degradation direction, the IMU is only used as a weak prior (or completely dominated by the external sensor).

[0148] Step S32 achieves anisotropic fusion weight allocation for the first time. Traditional fusion methods apply the same weights to all directions (e.g., loose coupling) or approximate the fusion by adjusting the diagonal elements of the overall covariance (while still assuming independence of each axis). This method directly manipulates the geometry of the state space through singular vectors, forcibly aligning IMU predictions in the degenerate direction, fundamentally suppressing drift divergence caused by degradation.

[0149] The purpose of step S32 is to actively select and introduce additional constraint sources only in these degradation directions after identifying the degradation direction, while retaining the full confidence of the original sensor in the non-degradation direction. This achieves the goal of compensating for weaknesses without dismantling strengths.

[0150] S33, perform degradation release detection and revocation of reinforcement factors.

[0151] This sub-step employs a sliding window singular value evolution monitoring method.

[0152] Maintenance length is Frame degradation history queue ,in The environmental structure factor defined in step 1.1. Simultaneously monitor the degenerate subspace dimension output in step S31. .

[0153] The condition for determining release is that three consecutive frames satisfy any of the following rules:

[0154] Rule A: and .in, : Environmental structural factors at any given time; : Determine the release threshold for structural recovery; : The dimension of the degenerate subspace at any given moment.

[0155] This condition indicates that the environmental structure has been restored and the degenerate subspace has shrunk to a low dimension (typically, only the gravity vector remains unobservable).

[0156] Rule B: And it continues to rise. This condition indicates that Lidar's health has returned to a reliable level.

[0157] Once release is determined, execute:

[0158] Remove all degradation direction enhancement factors from the factor plot;

[0159] Will Clear and reset the degenerate subspace;

[0160] Record the duration of this degradation event as input for the parameter self-evolution in step six.

[0161] Step S33 embodies the concept of closed-loop control—interventions must be withdrawn promptly after the disturbance disappears. If reinforcing factors persist, they can hinder the system from learning the correct state from real-world observations, turning treatment into a burden. By monitoring the temporal evolution of environmental structural factors and health, this step achieves precise management of the initiation and termination of degenerative events.

[0162] The purpose of step S33 is to ensure that when the robot leaves the geometrically degenerate environment (such as walking out of a corridor and into an open hall), the previously missing constraint directions are observed again. At this time, the artificial prior factors injected in step S32 should be promptly removed to avoid unnecessary bias to the actual observation.

[0163] Step S4: For hybrid degradation D3 level environment, the expected error of heterogeneous sensor engine is estimated by performance prediction network, continuous confidence weighting is used instead of hard switching, and sensor working mode is scheduled.

[0164] In this embodiment, step S4 may specifically include the following steps:

[0165] S41 provides real-time performance prediction for sensor engines in hybrid degradation D3 level environments.

[0166] This step constructs a lightweight performance prediction network, with the sensor health vector output from step S11 as the input. The output is the predicted trajectory error (ATE) values ​​of the three candidate engines. , , .

[0167] Network architecture: A three-layer fully connected network (MLP) is used, with a 3-dimensional input layer, 64-dimensional and 32-dimensional hidden layers respectively, and a 3-dimensional output layer. The activation function is ReLU. The loss function is mean squared error (MSE). .

[0168] in, The mean squared error loss function of the performance prediction network; The first prediction of network prediction Trajectory error (ATE) of sensor-like engines; The reference trajectory error obtained through backend optimization (including loop closure) on a real dataset is used as a supervision signal.

[0169] To achieve real-time prediction, the network must operate with extremely low latency at the edge. Therefore, this step employs model quantization techniques: compressing the weights from FP32 to INT8, deploying using TensorRT, with a single inference time of <0.5ms.

[0170] Step S41 enables proactive understanding of the sensor engine's performance. Traditional methods often react after problems arise—passively switching over only when the current engine has already developed significant drift. This method, by predicting the accuracy level the engine is about to reach, can proactively initiate a switchover before the drift accumulates to an unacceptable level, achieving a leap from reactive to predictive approaches.

[0171] The purpose of step S41 is to address the alternating or intermittent failure of vision and LiDAR in a hybrid degradation (D3 level) environment. The core function of this step is to predict the pose accuracy achievable at each currently enabled sensor engine at every moment, providing a quantitative basis for subsequent engine switching.

[0172] S42, based on the performance prediction results and the actual residual statistics within the current sliding window, dynamically activate / freeze preset sensor factors in the factor graph.

[0173] This step abandons the binary switching (0 or 1) and adopts a continuous confidence weighting method.

[0174] For each type of sensor factor Define its confidence weight The weights consist of two parts:

[0175] Prior weights The performance prediction from step S41 is calculated using the following formula:

[0176] .in, Based on performance prediction error Calculated prior weights; For prediction error, , This is the preset error normalization range.

[0177] The posterior weights are calculated based on the chi-square test pass rate within the current sliding window. Consistency test of residuals within the current sliding window. Define the chi-square test pass rate:

[0178] .

[0179] in For the first The residual vector of each factor Its information matrix, The chi-square threshold represents the 95% confidence level. This metric measures the degree of agreement between current sensor measurements and model predictions.

[0180] The final fusion weight is: . This is the balance coefficient, with a typical value of 0.4.

[0181] In factor graph optimization, the information matrix of the sensor factors is multiplied by the weights. :

[0182] .in, : Original information matrix of sensor factors; The information matrix that actually participated in the optimization after weighting.

[0183] Step S43 achieves true flexible fusion. When the sensor begins to degrade but has not yet completely failed, its information weight is continuously reduced rather than instantaneously reduced to zero, avoiding state jumps at the moment of switching; at the same time, the posterior consistency check provides on-site verification of performance prediction, preventing erroneous switching caused by prediction bias.

[0184] The purpose of step S42 is to enable the flexible switching of trusted and reliable sensors and suppress unreliable sensors, rather than abruptly enabling or disabling them.

[0185] S43, actively schedules the working mode of the sensor, making the two redundant and complementary in time and space.

[0186] This step constructs constraints that satisfy the optimization problem.

[0187] Define the decision variable: camera frame rate Hz, LiDAR scan frequency Hz, camera resolution level (1=VGA, 2=HD, 3=4K).

[0188] Objective function: Maximize effective observation coverage.

[0189] .

[0190] in, : Indicator function, returns 1 if the condition is true, otherwise returns 0; : The health status of the time camera and LiDAR; : The health threshold for determining the effectiveness and usability of a sensor.

[0191] Constraints:

[0192] Computational resource constraints: , CPU / GPU utilization For available budget.

[0193] Power consumption constraints: .

[0194] Minimum performance constraint: When at level D3, at least guarantee .

[0195] Solution method: Lightweight dynamic programming is adopted. Since the decision space is small (5×4×3=60 combinations) and the power consumption and computation model can be calibrated offline, it is entirely feasible to perform exhaustive search in each scheduling cycle (1 second).

[0196] Step S43 treats the sensors as configurable resources rather than black boxes. In traditional fusion methods, sensors operate in a fixed mode (e.g., a camera always operates at 30fps), either discarding data or receiving it all during degradation. This step, by actively adjusting the sensors' own operating parameters, maximizes the total available time of at least one sensor within the resource budget, fundamentally improving the system's survivability in mixed degradation environments.

[0197] The purpose of step S43 is to address the complementary failure characteristics of the camera and LiDAR in mixed degradation scenarios. When the camera fails due to overexposure, the LiDAR may be usable; conversely, when the LiDAR experiences a significant increase in noise due to rain or snow, the camera may function normally. This step aims to proactively schedule the sensor's operating modes (such as frame rate, resolution, and ROI) to create temporal and spatial redundancy, maximizing the overall effective observation time of the system.

[0198] Step S5: For fully degraded D4 level environments, a cross-platform generalized pure inertial odometry is provided through a three-level architecture of pre-training-fine-tuning-online adaptation, and unsupervised continuous self-improvement is achieved during the degradation process.

[0199] In this embodiment, step S5 may specifically include the following steps:

[0200] S51 is used for cross-platform heterogeneous IMU base model pre-training for fully degraded D4 level environments.

[0201] This step uses the Tartan IMU architecture to build a three-level framework of pre-training, fine-tuning, and online adaptation.

[0202] Pre-training phase: Collect over 100 hours of IMU data covering four types of robot platforms (wheeled, legged, drone, and handheld) to build a heterogeneous training set. The core innovation lies in the heterogeneous shared Backbone + multi-head output architecture.

[0203] Network input: continuous Frame IMU measurements, each frame contains 6 axes of data ( , ), Input tensor size is .

[0204] Backbone: Employs a variant of ResNet18 (1D convolution adapted to time series) to extract spatiotemporal features, which are then connected to an LSTM layer to capture long-range dependencies.

[0205] Multi-head output: On top of a shared Backbone, for each platform type (Cars, drones, legged, handheld) Set up independent regression heads to output the velocity estimate at the current moment. and its covariance .

[0206] The loss function formula is: .

[0207] in, : Robot platform type index (wheeled, legged, drone, handheld); : True velocity values ​​at any given moment (provided by a high-precision Lidar-INS integrated system); Network targeting platform Predicted Speed ​​in a given moment; The velocity estimation covariance matrix of network prediction; Mahalanobis distance represents the velocity error weighted by uncertainty.

[0208] Data augmentation: Introducing rotational isovariant augmentation—applying arbitrary rotations to the input IMU sequence. This requires the network output speed to also rotate accordingly: .

[0209] This enhancement forces the network to learn the geometric invariance of motion itself, rather than memorizing the mounting orientation of specific sensors, greatly improving cross-platform portability.

[0210] Step S51 breaks through the bottleneck of dedicated IMU models' inability to generalize. Traditional learning-based IMU odometry is trained on a single platform, and its performance drops sharply on unseen robot forms. This step, through heterogeneous multi-task learning, enables the model to abstract the underlying common features of motion and sensors, achieving one-time training and universal applicability across multiple platforms, providing a feasible guarantee for complete degradation mitigation.

[0211] The purpose of step S51 is to address the issue that in a fully degraded (D4 level) environment, both the camera and LiDAR are completely inoperable, leaving only the IMU as the system's operational component. The core of this step is to provide a pure inertial odometry system with cross-platform generalization capabilities, enabling low-drift positioning within minutes across various robot forms (wheeled, legged, drone, handheld), thus providing a solution for the system to survive the fully degraded range.

[0212] S52 adapts the cross-platform heterogeneous IMU base model to the specific dynamic characteristics of the current robot.

[0213] This step employs Low-Rank Adaptation (LoRA). For the weight matrix in the pre-trained network... Do not modify directly Instead, incremental updates are introduced. It is decomposed into the product of two low-rank matrices:

[0214] .

[0215] During forward propagation: .in, The original weight matrix frozen in the pre-trained network; The incremental update matrix to be learned; Low-rank decomposition matrix ; Input features of the network layer; The output characteristics of the network layer after adaptation.

[0216] Training strategy: Freeze Update only and The number of trainable parameters has increased from [previous figure]. Down to .Pick For the fully connected layer of ResNet18 ( The number of trainable parameters is reduced by 97%.

[0217] Adaptation process: Acquire approximately 1-2 minutes of standard motion IMU data on the target robot (no ground truth pose required; weak supervision signals can be provided by the pre-degradation vision / Lidar odometry), and run LoRA fine-tuning for approximately 60 epochs. Experiments show that this strategy can reduce trajectory error by 35%-55%, and the fine-tuned model shows almost no performance degradation on the source domain task (no catastrophic forgetting).

[0218] Step S52 resolves the contradiction between general models and individual characteristics. Pre-trained models are knowledge-rich but not tailored enough, and full fine-tuning is costly and easily forgotten. LoRA achieves rapid personalized adaptation with very few additional parameters, reducing the cold start time of IMU odometry in fully degraded scenarios from hours to minutes, making it feasible for engineering deployment.

[0219] The purpose of step S52 is to quickly adapt the general IMU base model from step S51 to the specific dynamic characteristics of the current robot. Although the pre-trained model possesses general motion knowledge, it still requires fine-tuning for the specific robot's suspension characteristics, actuator vibration modes, installation deviations, etc. This step employs efficient parameter fine-tuning technology to complete the adaptation after minutes of data acquisition, without forgetting the original knowledge.

[0220] S53 continuously optimizes the IMU odometry under unsupervised conditions, enabling it to operate, learn, and improve while in the degradation range.

[0221] This step constructs a dual-buffer dynamic learning architecture.

[0222] Motion pattern perception: A Gaussian Mixture Model (GMM) is used to perform online clustering of the raw IMU data stream. Feature vectors are defined. It has 6 dimensions. The GMM is updated online (using the incremental EM algorithm) and automatically identifies the current motion mode categories (such as stationary, constant speed straight, rapid acceleration, left turn, right turn, bumpy, etc.).

[0223] Dynamic training buffer: Maintains a circular buffer of fixed capacity. ,capacity Sample. When a new IMU segment arrives:

[0224] Determine its motion pattern category using GMM;

[0225] If the pattern accounts for less than 60% of the average in the buffer, then the sample is forcibly inserted;

[0226] Otherwise, by probability Randomly replace old samples with the same pattern in the buffer.

[0227] This strategy ensures that the buffer always maintains a diverse distribution of motion patterns, avoiding overfitting of the model to a single motion pattern (such as prolonged straight driving leading to a degradation in steering prediction ability).

[0228] Supervision signal generation: In a fully degenerate scenario, there is no external truth value. This step uses a sliding window smoothing constraint as a self-supervised signal.

[0229] .

[0230] in, For self-supervised loss functions in online continuous learning, For the current position estimate output by the IMU odometry, This method enables linear smoothing prediction of displacement within the same time period using a sliding window. The loss function forces the trajectory output by the IMU odometry to have temporal consistency, suppressing high-frequency oscillations and instantaneous jumps.

[0231] Step S53 introduces online continuous learning into pure inertial odometry for the first time. Traditional IMU odometry uses static models—the parameters are fixed after training, and the performance drops sharply when encountering motion patterns not covered by the training set. This step, through a diversified buffer guided by a GMM and a self-supervised smooth loss, achieves the ability to learn while flying and become more accurate over time, enabling the system to maintain reliable positioning within a complete degradation range of several minutes.

[0232] The purpose of step S53 is to address the drastic changes in the robot's movement pattern during the ongoing fully degraded scenario (e.g., from slow walking on flat ground to rapid sprinting across terrain). The core of this step is to continuously optimize the IMU odometry under unsupervised conditions, enabling it to operate, learn, and improve simultaneously within the degradation range.

[0233] Step S6: The degradation events encountered in the historical operation, the strategies adopted, and the resulting effects are structured and stored as empirical data, and the hyperparameters of each level of strategy are optimized offline using the empirical data.

[0234] In this embodiment, step S6 may specifically include the following steps:

[0235] S61, Build a library for replaying degraded event experiences.

[0236] Define the data structure for the degradation event tuple: .in, : Data structure for a single degradation event; : Start and end times of the event; : Time series sequence of sensor health status during the event; The time sequence of degradation levels activated during the event; The system employs a hierarchical sequence of intervention actions (including feature selection thresholds, state orientation reinforcement strength, factor weights, and IMU online adaptive hyperparameters). : The final pose estimation trajectory during the event (generated by the system at that time); The trajectory error distribution corrected by backend optimization (including loop closure detection) after the event ends serves as a post-event evaluation of this strategy combination.

[0237] Index event tuples using a vector database (such as FAISS) to identify degradation mode feature vectors. The statistical measures (mean, variance, rate of change) are used as keys to enable quick retrieval of similar historical scenarios.

[0238] Step S61 endows the system with long-term memory capabilities. Traditional adaptive systems only respond to the current state, handling each degradation event independently, and cannot learn from historical experience which strategy combinations are most effective in which scenarios. The experience replay library transforms each degradation response into a training sample, providing supervision signals for subsequent parameter optimization.

[0239] The purpose of step S61 is to structurally store the various levels of degradation events encountered by the robot during operation, the intervention strategies adopted, and the intervention effects (trajectory error, positioning continuity), thus constructing a long-term experience playback library. This library is the data foundation for the system to achieve memory and evolution.

[0240] S62 optimizes the adjustable parameters in each level of the hierarchical strategy based on the experience replay library of degradation events and using historical data as the training set.

[0241] A two-layer optimization problem is constructed: the inner layer is for generating localization trajectories under specific degradation events, and the outer layer is for global optimization of policy parameters.

[0242] Define the parameter vector to be optimized Includes but is not limited to:

[0243] Step S23: Number of target features for feature selection Confidence threshold ;

[0244] Step S32: Coefficients of the Directional Reinforcement Factor Information Matrix ;

[0245] Step S33: Degradation Release Threshold ;

[0246] Step S42: Weighted fusion coefficients ;

[0247] Step S53: Online learning rate, buffer capacity .

[0248] Optimization goal: .in, The vector of policy parameters to be optimized (including thresholds, weights, coefficients, etc. at each level); For the training subset of the experience base, For parameters Configure the system to handle events The trajectory obtained by replaying the simulation. This is a high-precision reference trajectory (approximation of the true value) obtained through global optimization after the event ends.

[0249] Solution method: Bayesian optimization is used. Since the objective function is non-convex and not differentiable with respect to parameters, and each evaluation requires running a complete localization system simulation, the computational cost is high. Bayesian optimization uses a Gaussian process surrogate model to find a better solution within a small number of evaluations.

[0250] Step S62 introduces meta-learning into the robot localization system for the first time. Traditional methods often rely on manual tuning of parameters such as thresholds and weights, which is time-consuming, labor-intensive, and difficult to handle diverse degradation scenarios. This step, based on the simple idea that past failures are the best teachers, allows the system to automatically learn optimal policy parameters from its own historical operational data, enabling the localization system to evolve and become smarter with use.

[0251] The purpose of step S62 is to utilize the experience base of degradation events accumulated in step S61, using historical data as the training set, to optimize the adjustable parameters in the hierarchical strategies at each level, so that the system can adopt better intervention strategies when similar scenarios reappear.

[0252] S63 performs hot loading of strategy parameters and rolling version updates.

[0253] This step implements a double-buffered parameter hot-switching mechanism.

[0254] Two copies of the parameters are maintained in memory:

[0255] Active version: The strategy parameters actually used by the current positioning thread;

[0256] Staging version: The latest optimized parameters are loaded in the background (initially the same as the Active version).

[0257] Update process:

[0258] Step S62 produces a new set of parameters. Verify its physical legitimacy (e.g., within a threshold). (interval, non-negative weights, etc.)

[0259] Will Loaded into the Staging area, but not switched immediately;

[0260] During the state prediction phase before the next frame of data processing begins, it is determined whether the system is currently in a stable period:

[0261] The degradation level remains D0 or D1 for more than 3 seconds;

[0262] The health status of each sensor did not fluctuate significantly.

[0263] The covariance of pose estimation is below a set threshold.

[0264] If the above conditions are met, parameter swapping is performed during the idle time window between the two frames, and the atomic operation points the Active pointer to the Staging area.

[0265] Record the switchover moment and archive the old version parameters for rollback.

[0266] If the health or estimated covariance deteriorates abnormally for 5 consecutive frames after the switch, the rollback mechanism will be automatically triggered, and the Active pointer will be pointed back to the old version.

[0267] Step S63 addresses the final mile problem of the disconnect between learning and execution in adaptive systems. Many systems with learning capabilities can only update parameters offline, requiring system downtime for new model deployment. This step, through double buffering and hot-switching techniques, enables continuous online evolution without downtime, allowing the system to absorb new experiences while running and continuously optimize its degradation response strategies, truly achieving a lifelong learning positioning system.

[0268] The purpose of step S63 is to smoothly deploy the optimized new strategy parameters from step S62 to the running system, achieving rolling version updates without restarting or interrupting the location service. This is a crucial bridge connecting offline learning and online operation.

[0269] The beneficial effects of implementing this embodiment are:

[0270] (1) Degradation is clearly divided into four levels: visual, geometric, mixed, and complete. Specialized coping strategies are designed for each level to achieve targeted measures, avoid over- or under-response, improve the robustness of the system in complex environments, and realize the fine differentiation and targeted processing of degradation scenarios.

[0271] (2) For visual degradation, the most contributing features are retained instead of being discarded entirely; for geometric degradation, constraints are introduced only in the direction of degradation instead of accepting inertial priors in their entirety. This approach of partial reliability and partial constraint effectively preserves the still valuable environmental information, delays the system's slide toward complete degradation, avoids the simple discarding of unreliable sensors, and instead selectively utilizes their effective information.

[0272] (3) In the hybrid degradation scenario, continuous confidence weighting is used instead of hard switching, eliminating pose jump caused by sudden changes in sensor state; at the same time, a performance prediction network is introduced to estimate the expected error, so that the fusion weight is forward-looking rather than remedial after the fact. The three-level architecture of pure inertial odometry solves the cross-platform generalization problem and continuously improves itself in the degradation process, getting rid of the dependence on expensive labeled data and realizing the smoothness and learnability of mode switching.

[0273] (4) The historical degradation events, response strategies and effects are stored in a structured manner and used for offline hyperparameter tuning, so that the system has the ability to remember and evolve. As the running data accumulates, the strategies at all levels can continuously approach the optimal configuration, realize the upgrade from rule-driven to data and rule hybrid-driven, and build an experience-driven closed-loop optimization mechanism.

[0274] This invention can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This invention can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This invention can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0275] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by instructing related hardware through computer-readable instructions. These computer-readable instructions can be stored in a computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. The aforementioned storage medium can be a non-volatile storage medium such as a magnetic disk, optical disk, or read-only memory (ROM), or random access memory (RAM).

[0276] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0277] Example 2

[0278] Further reference Figure 2 As a response to the above Figure 1 The present invention provides an embodiment of an adaptive hierarchical robot localization device for degraded scenarios, which is implemented in accordance with the method shown. Figure 1 Corresponding to the method embodiments shown, this device can be specifically applied to various electronic devices.

[0279] like Figure 2 As shown, the adaptive hierarchical robot localization device 70 for degraded scenarios described in this embodiment includes: a mapping module 71, a quantization module 72, a localization module 73, a scheduling module 74, a fine-tuning module 75, and an optimization module 76. Wherein:

[0280] The mapping module 71 is used to monitor the data quality acquired by heterogeneous sensors in real time, map the state of heterogeneous sensors to a unified quantization space, and output unique degradation level identifiers D1~D4, which represent visual degradation, geometric degradation, mixed degradation and complete degradation, respectively.

[0281] Quantization module 72 is used to perform uncertainty quantization on feature matching and select the most contributing feature subset for visual degradation D1 level environments;

[0282] The positioning module 73 is used to locate the degenerate subspace through singular value decomposition for a geometrically degenerate D2 level environment, and introduces inertial measurement unit prior constraints only in the degenerate direction, while retaining the original sensor confidence observations in the non-degenerate direction.

[0283] The scheduling module 74 is used to estimate the expected error of the heterogeneous sensor engine through a performance prediction network for hybrid degradation D3 level environments, replace hard switching with continuous confidence weighting, and schedule the sensor operating modes.

[0284] The fine-tuning module 75 is designed for fully degraded D4 level environments. Through a three-level architecture of pre-training-fine-tuning-online adaptation, it provides cross-platform generalized pure inertial odometry and achieves unsupervised continuous self-improvement during the degradation process.

[0285] The optimization module 76 is used to structure and store the degradation events encountered in historical operation, the strategies adopted, and the resulting effects as empirical data, and to use the empirical data to optimize the hyperparameters of each level of strategy offline.

[0286] The beneficial effects of implementing this embodiment are: it enables refined differentiation and targeted processing of degradation scenarios; it avoids simply discarding unreliable sensors, but selectively utilizes their effective information; it achieves smooth and learnable mode switching; and it constructs an experience-driven closed-loop optimization mechanism.

[0287] Example 3

[0288] To address the aforementioned technical problems, embodiments of the present invention also provide a computer device. Please refer to [link / reference needed]. Figure 3 , Figure 3 This is a basic structural block diagram of the computer device in this embodiment.

[0289] The aforementioned computer device 8 includes a memory 81, a processor 82, and a network interface 83 that are interconnected via a system bus. It should be noted that only the computer device 8 with components 81, 82, and 83 is shown in the figure; however, it should be understood that it is not required to implement all the shown components, and more or fewer components can be implemented alternatively. Those skilled in the art will understand that the computer device described herein is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions, and its hardware includes, but is not limited to, microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.

[0290] The aforementioned computer devices can be desktop computers, laptops, handheld computers, and cloud servers, among other computing devices. These devices can facilitate human-computer interaction with users through keyboards, mice, remote controls, touchpads, or voice-activated devices.

[0291] The aforementioned memory 81 includes at least one type of readable storage medium, including flash memory, hard disk, multimedia card, card-type memory (e.g., SD or DX memory), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the aforementioned memory 81 may be an internal storage unit of the aforementioned computer device 8, such as the hard disk or memory of the computer device 8. In other embodiments, the aforementioned memory 81 may also be an external storage device of the aforementioned computer device 8, such as a plug-in hard disk, smart media card (SMC), secure digital (SD) card, flash card, etc., equipped on the computer device 8. Of course, the aforementioned memory 81 may also include both the internal storage unit and its external storage device of the aforementioned computer device 8. In this embodiment, the aforementioned memory 81 is typically used to store the operating system and various application software installed on the aforementioned computer device 8, such as computer-readable instructions for an adaptive hierarchical robot localization method for degradation scenarios. In addition, the aforementioned memory 81 can also be used to temporarily store various types of data that have been output or will be output.

[0292] In some embodiments, the processor 82 described above may be a central processing unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chip. The processor 82 is typically used to control the overall operation of the computer device 8. In this embodiment, the processor 82 is used to execute computer-readable instructions stored in the memory 81 or to process data, for example, to execute the computer-readable instructions of the adaptive hierarchical robot localization method for degraded scenarios.

[0293] The network interface 83 may include a wireless network interface or a wired network interface, which is typically used to establish a communication connection between the computer device 8 and other electronic devices.

[0294] The beneficial effects of implementing this embodiment are: it enables refined differentiation and targeted processing of degradation scenarios; it avoids simply discarding unreliable sensors, but selectively utilizes their effective information; it achieves smooth and learnable mode switching; and it constructs an experience-driven closed-loop optimization mechanism.

[0295] Example 4

[0296] The present invention also provides another embodiment, namely, providing a computer-readable storage medium storing computer-readable instructions that can be executed by at least one processor to cause the at least one processor to perform the steps of the adaptive hierarchical robot localization method for degraded scenarios as described above.

[0297] The beneficial effects of implementing this embodiment are: it enables refined differentiation and targeted processing of degradation scenarios; it avoids simply discarding unreliable sensors, but selectively utilizes their effective information; it achieves smooth and learnable mode switching; and it constructs an experience-driven closed-loop optimization mechanism.

[0298] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of the present invention.

[0299] Obviously, the embodiments described above are merely some embodiments of the present invention, not all embodiments. The accompanying drawings show preferred embodiments of the present invention, but do not limit the patent scope of the present invention. The present invention can be implemented in many different forms; rather, these embodiments are provided to provide a more thorough and complete understanding of the disclosure of the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing specific embodiments, or make equivalent substitutions for some of the technical features. Any equivalent structures made using the content of this specification and drawings, directly or indirectly applied to other related technical fields, are similarly within the patent protection scope of this invention.

Claims

1. An adaptive hierarchical robot localization method for degraded scenarios, characterized in that, Includes the following steps: The system monitors the data quality acquired by heterogeneous sensors in real time, maps the state of the heterogeneous sensors to a unified quantization space, and outputs unique degradation level identifiers D1~D4, representing visual degradation, geometric degradation, mixed degradation, and complete degradation, respectively. For visual degradation level D1 environments, uncertainty quantification is performed on feature matching, and the most contributing feature subset is selected. For geometrically degraded D2 level environments, the degraded subspace is located by singular value decomposition, and prior constraints of the inertial measurement unit are introduced only in the degraded direction, while the original sensor confidence observations are retained in the non-degraded direction. For hybrid degradation D3 level environments, the expected error of the heterogeneous sensor engine is estimated through a performance prediction network, continuous confidence weighting is used instead of hard switching, and the sensor operating mode is scheduled. For fully degraded D4 level environments, a three-level architecture of pre-training-fine-tuning-online adaptation is used to provide a cross-platform generalizable pure inertial odometry, and to achieve unsupervised continuous self-improvement during the degradation process; The degradation events encountered in historical operations, the strategies adopted, and the resulting effects are structured and stored as empirical data, and the hyperparameters of each level of strategy are optimized offline using the empirical data.

2. The adaptive hierarchical robot localization method for degraded scenarios according to claim 1, characterized in that, The steps of real-time monitoring of the data quality acquired by heterogeneous sensors, mapping the state of the heterogeneous sensors to a unified quantization space, and outputting unique degradation level identifiers D1~D4, representing visual degradation, geometric degradation, mixed degradation, and complete degradation respectively, specifically include: Real-time monitoring of the data quality acquired by the heterogeneous sensor, and output of the health score and degradation type label of the heterogeneous sensor; Based on the health score and the degradation type label, the current environment is classified into one of four degradation levels through decision tree logic: visual degradation level D1, geometric degradation level D2, mixed degradation level D3 and complete degradation level D4, and a unique activated intervention level identifier is output. Spatiotemporal alignment calibration is performed on the heterogeneous sensor.

3. The adaptive hierarchical robot localization method for degraded scenarios according to claim 1, characterized in that, The steps for quantifying the uncertainty of feature matching and selecting the most contributing feature subset for visual degradation level D1 environments specifically include: For visual degradation level D1 environments, feature uncertainty is quantified based on covariance estimation to generate a feature candidate set and covariance matrix; Based on the feature candidate set and the covariance matrix, select the preset N features with the largest information gain to participate in pose optimization; Adaptive threshold dynamic adjustment and feature activation are performed.

4. The adaptive hierarchical robot localization method for degraded scenarios according to claim 1, characterized in that, The steps for locating the degenerate subspace through singular value decomposition for geometrically degraded D2-level environments, introducing prior constraints for inertial measurement units only in the degenerate direction, and retaining the original sensor confidence observations in the non-degenerate direction specifically include: For D2 level geometrically degraded environments, the direction of degradation is identified through singular value decomposition. Actively select to introduce additional constraint sources only in the degradation direction, while retaining the full confidence of the original sensor in the non-degradation direction; Degradation release detection and reinforcement factor withdrawal were performed.

5. The adaptive hierarchical robot localization method for degraded scenarios according to claim 1, characterized in that, The steps for estimating the expected error of the heterogeneous sensor engine through a performance prediction network, replacing hard switching with continuous confidence weighting, and scheduling the sensor operating modes for hybrid degradation D3 level environments specifically include: Real-time performance prediction of sensor engine for hybrid degradation D3 level environments; Based on the performance prediction results and the actual residual statistics within the current sliding window, preset sensor factors are dynamically activated / frozen in the factor graph. Actively scheduling the operating modes of the sensors enables them to complement each other in both time and space.

6. The adaptive hierarchical robot localization method for degraded scenarios according to claim 1, characterized in that, The steps for providing a cross-platform generalized pure inertial odometry through a three-level architecture of pre-training-fine-tuning-online adaptation for fully degraded D4 level environments, and achieving unsupervised continuous self-improvement during degradation, specifically include: For fully degraded D4 level environments, cross-platform heterogeneous IMU base model pre-training is performed; Adapt the cross-platform heterogeneous IMU base model to the specific dynamic characteristics of the current robot; Continuously optimize the IMU odometry under unsupervised conditions so that it can operate, learn, and improve while in the degradation range.

7. The adaptive hierarchical robot localization method for degraded scenarios according to any one of claims 1-6, characterized in that, The step of structuring and storing the degradation events encountered in historical operations, the strategies adopted, and the resulting effects as empirical data, and then using the empirical data to optimize the hyperparameters of each level of strategy offline, specifically includes: Build a library of experience replays of degradation events; Based on the degradation event experience replay library, using historical data as the training set, the adjustable parameters in the hierarchical strategies at each level are optimized. Perform hot loading of strategy parameters and rolling version updates.

8. An adaptive hierarchical robot localization device for degraded scenarios, characterized in that, include: The mapping module is used to monitor the data quality acquired by heterogeneous sensors in real time, map the state of heterogeneous sensors to a unified quantization space, and output unique degradation level identifiers D1~D4, representing visual degradation, geometric degradation, mixed degradation and complete degradation, respectively. The quantization module is used to perform uncertainty quantization on feature matching and select the most contributing feature subset for visual degradation D1 level environments. The positioning module is used to locate the degenerate subspace in a geometrically degenerate D2 level environment through singular value decomposition, and introduces inertial measurement unit prior constraints only in the degenerate direction, while retaining the original sensor confidence observations in the non-degenerate direction. The scheduling module is used to estimate the expected error of the heterogeneous sensor engine through a performance prediction network for hybrid degradation D3 level environments, replace hard switching with continuous confidence weighting, and schedule the sensor operating modes. The fine-tuning module is designed for fully degraded D4 level environments. Through a three-level architecture of pre-training-fine-tuning-online adaptation, it provides cross-platform generalized pure inertial odometry and achieves unsupervised continuous self-improvement during the degradation process. The optimization module is used to structure and store the degradation events encountered in historical operations, the strategies adopted, and the resulting effects as empirical data, and to use the empirical data to optimize the hyperparameters of strategies at all levels offline.

9. A computer device, characterized in that, The method includes a memory and a processor, wherein the memory stores computer-readable instructions, and the processor executes the computer-readable instructions to implement the steps of the adaptive hierarchical robot localization method for degraded scenarios as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-readable instructions, which, when executed by a processor, implement the steps of the adaptive hierarchical robot localization method for degraded scenarios as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Feature degradation scene detection method, robot positioning method, equipment and medium

    CN117168433A

  • Multi-sensor pose estimation method and device considering perceptual degradation

    CN117745821A

  • Mobile robot positioning method in typical laser radar degradation scene

    CN118129733A

  • Autonomous robot with deep learning environment recognition and sensor calibration

    EP4202866A1