Method, device and medium for determining sensor weight model of unmanned equipment

By acquiring the actual pose and sensor information of the unmanned equipment, determining the acquisition residuals and training the target weight coefficients, the problem of positioning and mapping accuracy caused by fixed sensor weights in unmanned equipment is solved, achieving higher environmental adaptability and accuracy.

CN121234046BActive Publication Date: 2026-03-31BEIJING YICHEN TIMES TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-16
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

The fixed weighting coefficients of visual sensors and lidar in unmanned equipment cannot adapt to dynamic and complex environments, resulting in poor positioning and mapping accuracy.

Method used

By acquiring the actual pose information of the unmanned equipment and the sensor acquisition information, the acquisition residual of the sensor at the moment to be processed is determined, and the initial weight coefficients are traversed. The target weight coefficients are obtained by training a neural network model, generating a sensor weight model, and the sensor weights are dynamically adjusted to adapt to environmental changes.

Benefits of technology

It improves the positioning and mapping accuracy of unmanned equipment in dynamic and complex environments and enhances its adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121234046B_ABST
    Figure CN121234046B_ABST
Patent Text Reader

Abstract

The application provides a method and device for determining a sensor weight model of an unmanned device, equipment and a medium. Actual pose information and sensor collection information of the unmanned device at a historical time are obtained. Sensor collection residuals of the sensor at a to-be-processed time after a processable time of the unmanned device are determined according to a surrounding map and the sensor collection information at the historical time. At least one initial weight coefficient of the sensor is obtained, and the initial weight coefficient is iterated. The inference pose information of the unmanned device at the to-be-processed time is determined by using the currently iterated initial weight coefficient and the sensor collection residuals. The inference error of the currently iterated initial weight coefficient is determined according to the inference pose information and the actual pose information. The target weight coefficient is determined according to the inference error of the initial weight coefficient. The sensor weight model is obtained by training the neural network model using the sensor collection information and the target weight coefficient, and the sensor weight of the unmanned device in different environments is generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of unmanned equipment technology, and in particular to methods, apparatus, equipment and media for determining sensor weight models for unmanned equipment. Background Technology

[0002] In related technologies, unmanned devices such as robots and drones typically integrate unmanned systems. Unmanned devices refer to machines or platforms that autonomously complete tasks without direct and continuous intervention from human drivers or operators. Unmanned devices often employ tightly coupled LIVO (Lidar-Inertial-Visual Odometry) technology for high-precision positioning and mapping. Based on this positioning and mapping information, the unmanned device can perform autonomous navigation and path planning.

[0003] The unmanned equipment is equipped with sensors such as visual sensors, lidar, and IMU (Inertial Measurement Unit). The unmanned equipment can acquire measurement information of the surrounding environment from these sensors, and based on this measurement information and the corresponding weighting coefficients of the sensors, it can obtain the pose of the unmanned equipment and a map of the surrounding environment.

[0004] However, in related technologies, the weighting coefficients of visual sensors and LiDAR are fixed during the operation of unmanned equipment, and cannot change with dynamic and complex environments. This results in poor adaptability, leading to poor localization and mapping accuracy of unmanned equipment in dynamic and complex environments. For example, in scenarios where images are blurred due to sparse textures, drastic changes in lighting, or high-speed movement, the measurement reliability of visual sensors decreases significantly. If the fixed high weights of the visual sensors do not change, erroneous constraints will be introduced, affecting the overall state estimation of the unmanned equipment and leading to a decrease in the accuracy of localization and mapping. Summary of the Invention

[0005] This application provides a method, apparatus, electronic device, and computer-readable storage medium for determining the sensor weight model of an unmanned device, in order to solve the problem that the weight coefficients corresponding to visual sensors and lidar are fixed during the operation of the unmanned device and cannot change with the dynamic and complex environment, resulting in poor adaptability and poor positioning and mapping accuracy of the unmanned device in dynamic and complex environments.

[0006] This application discloses a method for determining a sensor weight model for an unmanned device, applicable to the unmanned device, which includes at least one sensor. The method includes:

[0007] Acquire the actual pose information of the unmanned device and / or the sensor-collected information of the unmanned device by the sensor at at least one historical moment, and take any one of the at least one historical moments as the moment to be processed, and any moment before the moment to be processed as the moment that can be processed.

[0008] Based on the surrounding map of the unmanned equipment at the processable time, the ideal acquisition information of the sensor at the unprocessed time is determined, and based on the ideal acquisition information and the sensor acquisition information, the sensor acquisition residual of the sensor at the unprocessed time is determined.

[0009] At least one initial weight coefficient of the sensor is obtained, and the initial weight coefficients are traversed. The inference pose information of the unmanned device at the time to be processed is determined by using the currently traversed initial weight coefficients and the sensor acquisition residuals of the sensor.

[0010] Based on the inferred pose information and the actual pose information, determine the inference error corresponding to the initial weight coefficient of the current traversal;

[0011] Based on the inference error corresponding to the at least one initial weight coefficient, a target weight coefficient is determined from the at least one initial weight coefficient;

[0012] The sensor weight model is obtained by training a preset neural network model using the sensor acquisition information at the time to be processed and the target weight coefficient corresponding to the time to be processed.

[0013] Optionally, the method includes:

[0014] Acquire the actual information collected by the sensor about the environment surrounding the unmanned device at the at least one historical moment;

[0015] Based on the actual pose information and / or actual collected information corresponding to the processable moment, a map of the unmanned device's surroundings at the processable moment is constructed.

[0016] Optionally, the sensor includes at least one of a lidar, a vision sensor, and an inertial measurement unit; the initial weighting coefficient includes the initial weighting coefficient of the lidar and / or the initial weighting coefficient of the vision sensor; the sensor acquisition residual includes at least one of the acquisition residual of the lidar, the acquisition residual of the vision sensor, and the acquisition residual of the inertial measurement unit.

[0017] The step of determining the inference pose information of the unmanned device at the time to be processed by using the initial weight coefficients of the current traversal and the sensor acquisition residuals of the sensor includes:

[0018] An initial function is constructed using the initial weight coefficients of the current traversal and the sensor acquisition residuals of the sensor;

[0019] The inference pose information is determined based on the initial function; the initial function is:

[0020]

[0021] Among them, F new (x) is the initial function, x includes the inference pose information of the unmanned device, r imu r is the acquisition residual of the inertial measurement unit. vision Let r be the acquisition residual of the visual sensor. lidar Σ represents the acquisition residual of the lidar. imu Let Σ be the covariance matrix of the inertial measurement unit. vision Let Σ be the covariance matrix of the visual sensor. lidar Let w be the covariance matrix of the lidar. v w represents the initial weighting coefficients of the visual sensor. l The initial weighting coefficients for the lidar are given.

[0022] Optionally, the sensor acquisition information includes image information from the visual sensor and / or point cloud information from the lidar; the step of training a preset neural network model using the sensor acquisition information at the time to be processed and the target weight coefficients corresponding to the time to be processed to obtain a sensor weight model includes:

[0023] Extract the image feature vector from the image information and extract the point cloud feature vector from the point cloud information;

[0024] The image feature vector and / or the point cloud feature vector are input into the neural network model to obtain the training weight coefficients of the sensor;

[0025] The mean square error of the neural network model is determined based on the training weight coefficients and the target weight coefficients.

[0026] Based on the mean square error, the neural network model is adjusted, and it is determined whether the adjusted neural network model meets the preset model training objective.

[0027] If the conditions are not met, the steps of inputting the image feature vector and / or the point cloud feature vector into the neural network model to obtain the training weight coefficients of the sensor, determining the mean square error of the neural network model based on the training weight coefficients and the target weight coefficients, and adjusting the neural network model based on the mean square error are repeated until the adjusted neural network model meets the model training objective.

[0028] Optionally, the method includes:

[0029] Acquire the current information collected by the at least one sensor on the unmanned device at the current moment;

[0030] Based on the currently acquired information, determine the current acquisition residual of the sensor;

[0031] Based on the currently collected information, the current weight coefficients of the at least one sensor are determined using the sensor weight model;

[0032] Using the current weight coefficients and the current acquisition residuals, construct the objective function corresponding to the initial function;

[0033] The current pose of the unmanned device is determined based on the objective function.

[0034] Optionally, the currently acquired information includes at least one of the current image information of the visual sensor, the current inertial information of the inertial measurement unit, and the current point cloud information of the lidar; determining the current acquisition residual of the sensor based on the current acquired information includes:

[0035] Based on the current image information, determine the current image residual of the visual sensor; and / or,

[0036] Based on the current inertial information, determine the current inertial residual of the inertial measurement unit; and / or,

[0037] Based on the current point cloud information, determine the current point cloud residual of the lidar.

[0038] Optionally, determining the current weight coefficients of the at least one sensor based on the currently acquired information using the sensor weight model includes:

[0039] The current image information is input into the backbone network of a preset convolutional neural network to obtain the current image feature vector corresponding to the current image information;

[0040] The current point cloud information is input into a preset point cloud encoder to obtain the current point cloud feature vector corresponding to the current point cloud information;

[0041] The current image feature vector and / or the current point cloud feature vector are input into the sensor weight model to obtain the current weight coefficients of the at least one sensor.

[0042] This application also discloses a device for determining a sensor weight model for an unmanned device, applied to the unmanned device, which includes at least one sensor, and the device includes:

[0043] The information acquisition module is used to acquire the actual pose information of the unmanned device and / or the sensor data collected by the sensor on the unmanned device at at least one historical moment, and to take any one of the at least one historical moments as the moment to be processed, and any moment before the moment to be processed as the moment that can be processed.

[0044] The residual determination module is used to determine the ideal acquisition information of the sensor at the unprocessed time based on the surrounding map of the unmanned equipment at the processable time, and to determine the sensor acquisition residual of the sensor at the unprocessed time based on the ideal acquisition information and the sensor acquisition information.

[0045] The inference pose information determination module is used to obtain at least one initial weight coefficient of the sensor, and traverse the initial weight coefficients, and use the currently traversed initial weight coefficients and the sensor acquisition residuals of the sensor to determine the inference pose information of the unmanned device at the time to be processed.

[0046] The inference error determination module is used to determine the inference error corresponding to the initial weight coefficient of the current traversal based on the inference pose information and the actual pose information.

[0047] A target weight coefficient determination module is used to determine a target weight coefficient from the at least one initial weight coefficient based on the inference error corresponding to the at least one initial weight coefficient.

[0048] The training module is used to train a preset neural network model using the sensor acquisition information at the time to be processed and the target weight coefficients corresponding to the time to be processed, so as to obtain a sensor weight model.

[0049] Optionally, the device includes:

[0050] The actual data acquisition module is used to acquire the actual data collected by the sensor on the environment surrounding the unmanned device at the at least one historical moment.

[0051] The map building module is used to build a map of the unmanned device's surroundings at the processable time based on the actual pose information and / or actual collected information corresponding to the processable time.

[0052] Optionally, the sensor includes at least one of a lidar, a vision sensor, and an inertial measurement unit; the initial weighting coefficient includes the initial weighting coefficient of the lidar and / or the initial weighting coefficient of the vision sensor; the sensor acquisition residual includes at least one of the acquisition residual of the lidar, the acquisition residual of the vision sensor, and the acquisition residual of the inertial measurement unit.

[0053] The inference pose information determination module includes:

[0054] The initial function construction submodule is used to construct the initial function using the initial weight coefficients of the current traversal and the sensor acquisition residuals of the sensor;

[0055] The inference pose information determination submodule is used to determine the inference pose information based on the initial function; the initial function is:

[0056]

[0057] Among them, F new (x) is the initial function, x includes the inference pose information of the unmanned device, r imu r is the acquisition residual of the inertial measurement unit. vision Let r be the acquisition residual of the visual sensor. lidar Σ represents the acquisition residual of the lidar. imu Let Σ be the covariance matrix of the inertial measurement unit. vision Let Σ be the covariance matrix of the visual sensor. lidar Let w be the covariance matrix of the lidar. v w represents the initial weighting coefficients of the visual sensor. l The initial weighting coefficients for the lidar are given.

[0058] Optionally, the sensor-acquired information includes image information from the visual sensor and / or point cloud information from the lidar; the training module includes:

[0059] The feature vector extraction submodule is used to extract image feature vectors from the image information and point cloud feature vectors from the point cloud information.

[0060] The training weight coefficients submodule is used to input the image feature vector and / or the point cloud feature vector into the neural network model to obtain the training weight coefficients of the sensor;

[0061] The mean squared error determination submodule is used to determine the mean squared error of the neural network model based on the training weight coefficients and the target weight coefficients.

[0062] The adjustment submodule is used to adjust the neural network model according to the mean square error, and to determine whether the adjusted neural network model meets the preset model training objective.

[0063] The repeat submodule is used to repeat the steps of inputting the image feature vector and / or the point cloud feature vector into the neural network model to obtain the training weight coefficients of the sensor, determining the mean square error of the neural network model based on the training weight coefficients and the target weight coefficients, and adjusting the neural network model based on the mean square error, until the adjusted neural network model meets the model training objective.

[0064] Optionally, the device includes:

[0065] The current information acquisition module is used to acquire the current information collected by the at least one sensor on the unmanned device at the current moment;

[0066] The current acquisition residual determination module is used to determine the current acquisition residual of the sensor based on the current acquisition information.

[0067] The current weight coefficient determination module is used to determine the current weight coefficient of the at least one sensor based on the currently collected information and using the sensor weight model.

[0068] The objective function construction module is used to construct the objective function corresponding to the initial function using the current weight coefficients and the current acquisition residuals;

[0069] The current pose determination module is used to determine the current pose of the unmanned device based on the objective function.

[0070] Optionally, the currently acquired information includes at least one of the current image information of the visual sensor, the current inertial information of the inertial measurement unit, and the current point cloud information of the lidar; the currently acquired residual determination module includes:

[0071] The residual determination submodule is used to determine the current image residual of the visual sensor based on the current image information; and / or,

[0072] Based on the current inertial information, determine the current inertial residual of the inertial measurement unit; and / or,

[0073] Based on the current point cloud information, determine the current point cloud residual of the lidar.

[0074] Optionally, the current weight coefficient determination module includes:

[0075] The image feature vector acquisition submodule is used to input the current image information into the backbone network of a preset convolutional neural network to obtain the current image feature vector corresponding to the current image information;

[0076] The point cloud feature vector acquisition submodule is used to input the current point cloud information into a preset point cloud encoder to obtain the current point cloud feature vector corresponding to the current point cloud information;

[0077] The current weight coefficient acquisition submodule is used to input the current image feature vector and / or the current point cloud feature vector into the sensor weight model to obtain the current weight coefficient of the at least one sensor.

[0078] This application also discloses an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;

[0079] The memory is used to store computer programs;

[0080] When the processor executes a program stored in the memory, it implements the method described in the embodiments of this application.

[0081] This application also discloses one or more computer-readable media storing instructions that, when executed by one or more processors, cause the processors to perform the methods described in this application.

[0082] The embodiments of this application have the following advantages:

[0083] In this embodiment, the unmanned device includes at least one sensor to acquire the actual pose information of the unmanned device and / or sensor acquisition information of the unmanned device at at least one historical moment, and to take any moment in the at least one historical moment as the moment to be processed, and any moment before the moment to be processed as the moment that can be processed; to determine the ideal acquisition information of the sensor at the moment to be processed based on the surrounding map of the unmanned device at the moment that can be processed, and to determine the sensor acquisition residual of the sensor at the moment to be processed based on the ideal acquisition information and the sensor acquisition information; to acquire at least one initial weight coefficient of the sensor, and to traverse the initial weight coefficients, and to determine the inferred pose information of the unmanned device at the moment to be processed using the currently traversed initial weight coefficients and the sensor acquisition residual of the sensor; and to determine the inferred pose information of the unmanned device based on the inferred pose information. Based on pose information and actual pose information, the inference error corresponding to the initial weight coefficients of the current traversal is determined. Based on the inference error corresponding to at least one initial weight coefficient, the target weight coefficient is determined from at least one initial weight coefficient. Using the sensor acquisition information at the time to be processed and the target weight coefficient corresponding to the time to be processed, a preset neural network model is trained to obtain a sensor weight model. Using the sensor weight model, sensor weights for unmanned equipment in different environments can be generated. This solves the problem that the weight coefficients corresponding to the residual terms of visual sensors and lidar are fixed during the operation of unmanned equipment and cannot change with dynamic and complex environments, resulting in poor adaptability and poor positioning and mapping accuracy of unmanned equipment in dynamic and complex environments. Attached Figure Description

[0084] Figure 1 This is a flowchart illustrating the steps of a method for determining the sensor weight model of an unmanned device provided in an embodiment of this application.

[0085] Figure 2 This is a flowchart illustrating a method for determining the pose of an unmanned device provided in an embodiment of this application.

[0086] Figure 3 This is a flowchart illustrating a training method for a neural network model provided in an embodiment of this application;

[0087] Figure 4 This is a flowchart illustrating a method for constructing a training set for a neural network model provided in an embodiment of this application;

[0088] Figure 5 This is a structural block diagram of a device for determining the sensor weight model of an unmanned device provided in an embodiment of this application;

[0089] Figure 6 This is a block diagram of an electronic device provided in an embodiment of this application;

[0090] Figure 7 This is a schematic diagram of a computer-readable medium provided in an embodiment of this application. Detailed Implementation

[0091] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0092] Reference Figure 1 This document illustrates a flowchart of the steps involved in determining a sensor weight model for an unmanned device, as provided in an embodiment of this application. The method is applied to an unmanned device, which includes at least one sensor, and specifically includes the following steps:

[0093] Step 101: Obtain the actual pose information of the unmanned device and / or the sensor data collected by the sensor on the unmanned device at at least one historical moment, and take any one of the at least one historical moments as the moment to be processed, and any one of the moments before the moment to be processed as the moment that can be processed.

[0094] In this embodiment, the unmanned device is equipped with at least one sensor, including a vision sensor, a lidar sensor, and an IMU (Inertial Measurement Unit). The unmanned device can acquire measurement information of the surrounding environment from these sensors, and then use a tightly coupled LIVO method to determine the pose of the unmanned device and the mapping of its surrounding environment. Here, "tight coupling" means simultaneously utilizing the raw measurement information of all sensors to directly solve for the pose of the unmanned device and the map of its surrounding environment within a unified optimization framework, rather than independently determining the pose and map corresponding to each sensor using the measurement information of each sensor, and then fusing the poses and maps of each sensor to determine the final pose and map.

[0095] In this embodiment, within the tightly coupled LIVO method, the unified optimization framework for the unmanned device can be a function. The unmanned device can first construct this function, and then minimize it using a nonlinear optimization algorithm to accurately obtain the unmanned device's pose and a map of its surrounding environment. The function corresponding to the optimization framework can be:

[0096]

[0097] Where F(x) represents the function corresponding to the optimization framework; x is the state vector to be optimized in the function, including: the position, attitude, and velocity of the unmanned device; r imu For the residual term of the inertial measurement unit; r vision For the residual term of the vision sensor; r lidar For the residual terms of the lidar; Σimu Σ is the covariance matrix of the inertial measurement unit; vision Σ is the covariance matrix of the visual sensor; lidar Let be the covariance matrix of the lidar; A, B, and C are the weighting coefficients of the inertial measurement unit, the vision sensor, and the lidar, respectively.

[0098] Therefore, the function corresponding to the optimization framework is a weighted sum of the residuals of the inertial measurement unit, LiDAR, and vision sensor. The sensor residuals refer to the difference between the sensor's predicted and actual observations; the sensor's covariance matrix is ​​used to weight the residuals of different sensors, so that the more reliable sensor has a greater impact on the function.

[0099] If we represent the sensor's residual term with α, then The calculation formula is:

[0100]

[0101] In this embodiment, the information collected by the inertial measurement unit includes information such as the acceleration and angular velocity of the unmanned device. In the function corresponding to the optimization framework, the information collected by the inertial measurement unit is usually less affected by the surrounding environment of the unmanned device, and the information collected by the inertial measurement unit is usually the basis for the motion constraints of the unmanned device. Therefore, the weight coefficient of the inertial measurement unit is kept unchanged.

[0102] In one example, the weighting coefficient of the inertial measurement unit can be set to 1.

[0103] In this embodiment, an offline training method can be used to construct a sensor weight model for an unmanned device. This model can then be run online to obtain the target weight coefficients of the unmanned device's visual sensor and LiDAR. The offline training step can be performed on a server equipped with a high-performance GPU (Graphics Processing Unit).

[0104] In this embodiment, sensor information collected by at least one sensor of the unmanned device at at least one historical moment can be obtained. This sensor information may include image information collected by a vision sensor, inertial information collected by an inertial measurement unit, and point cloud information collected by a lidar sensor.

[0105] In one example, sensor-acquired information could be a large-scale dataset encompassing various scenes and challenging environments, such as the nuScenes dataset. The nuScenes dataset provides time-stamped, strictly synchronized image information, LiDAR point cloud information, and IMU data. Strictly synchronized timestamps mean that the data acquisition timestamps of image information acquired by visual sensors (such as cameras), point cloud information acquired by LiDAR, and inertial information acquired by IMU are strictly aligned.

[0106] In this embodiment of the application, the actual pose information of the unmanned device at at least one historical moment can also be obtained. The actual pose information is a high-precision 6-DOF actual pose, including 3 translational degrees of freedom and 3 rotational degrees of freedom.

[0107] In this embodiment of the application, any moment from at least one historical moment can be used as the moment to be processed, and any moment before the moment to be processed can be used as the moment that can be processed. The moment that can be processed can be the moment before the moment to be processed.

[0108] In the embodiments of this application, the data acquisition timestamps of image information acquired by visual sensors (such as cameras), point cloud information acquired by lidar, and inertial information acquired by IMU are strictly aligned. Therefore, the time to be processed is represented by the image frame to be processed, and the processable time is represented by the processable image frame preceding the image frame to be processed.

[0109] In some embodiments of this application, the method includes:

[0110] Acquire the actual information collected by the sensor about the environment surrounding the unmanned device at the at least one historical moment;

[0111] Based on the actual pose information and / or actual collected information corresponding to the processable moment, a map of the unmanned device's surroundings at the processable moment is constructed.

[0112] In the embodiments of this application, the actual information collected by the sensors on the environment surrounding the unmanned device at at least one historical moment can be obtained, such as: actual image information collected by the visual sensor on the environment surrounding the unmanned device, and actual point cloud information collected by the lidar on the environment surrounding the unmanned device.

[0113] In this embodiment, the unmanned device includes a front-end and a back-end. The back-end of the unmanned device can utilize functions corresponding to the optimization framework to determine information such as the device's pose and velocity. Furthermore, the back-end can construct a map of the unmanned device's surroundings at the processable moment based on the actual pose information at that moment and / or the actual information collected by sensors from the surrounding environment.

[0114] In one example, the unmanned device can acquire the actual pose information (Pose) corresponding to a processable image frame. i-1_gt Then, the backend of the unmanned device can pose the image based on the actual pose information corresponding to the processable image frames. i-1_gt The actual information collected by sensors about the environment around the unmanned device is used to determine the true value of map points of the unmanned device in the processable image frame, which is equivalent to the map of the unmanned device around the processable image frame.

[0115] Step 102: Based on the surrounding map of the unmanned equipment at the processable time, determine the ideal acquisition information of the sensor at the unprocessed time, and based on the ideal acquisition information and the sensor acquisition information, determine the sensor acquisition residual of the sensor at the unprocessed time.

[0116] In this embodiment, the processable time can be the time preceding the time to be processed. Based on the map surrounding the unmanned device at the processable time, the ideal data acquisition information from the unmanned device's sensors at the time to be processed can be determined.

[0117] In this embodiment, based on the ideal acquisition information of the sensor at the time of processing and the sensor acquisition information at the time of processing, the sensor acquisition residual at the time of processing can be determined. The sensor acquisition residual includes at least one of the acquisition residual of the vision sensor, the acquisition residual of the lidar, and the acquisition residual of the inertial measurement unit.

[0118] Step 103: Obtain at least one initial weight coefficient of the sensor, and iterate through the initial weight coefficients. Using the currently iterated initial weight coefficients and the sensor acquisition residuals of the sensor, determine the inference pose information of the unmanned device at the time to be processed.

[0119] In this embodiment, at least one initial weighting coefficient of the sensor can be obtained. Since the weighting coefficient of the inertial measurement unit remains unchanged, and the weighting coefficient of the inertial measurement unit can be set to 1, an initial weighting coefficient of the sensor may include the initial weighting coefficient of the vision sensor and the initial weighting coefficient of the lidar.

[0120] In this embodiment, the initial weight coefficients of the sensors can be traversed, and the initial weight coefficients of the currently traversed sensors and the sensor acquisition residuals of the sensors can be used to determine the inference pose information of the unmanned equipment at the moment to be processed.

[0121] In some embodiments of this application, the sensor includes at least one of a lidar, a vision sensor, and an inertial measurement unit; the initial weighting coefficient includes the initial weighting coefficient of the lidar and / or the initial weighting coefficient of the vision sensor; the sensor acquisition residual includes at least one of the acquisition residual of the lidar, the acquisition residual of the vision sensor, and the acquisition residual of the inertial measurement unit.

[0122] The step of determining the inference pose information of the unmanned device at the time to be processed by using the initial weight coefficients of the current traversal and the sensor acquisition residuals of the sensor includes:

[0123] An initial function is constructed using the initial weight coefficients of the current traversal and the sensor acquisition residuals of the sensor;

[0124] The inference pose information is determined based on the initial function; the initial function is:

[0125]

[0126] Among them, F new (x) is the initial function, x includes the inference pose information of the unmanned device, r imu r is the acquisition residual of the inertial measurement unit. vision Let r be the acquisition residual of the visual sensor. lidar Σ represents the acquisition residual of the lidar. imu Let Σ be the covariance matrix of the inertial measurement unit. vision Let Σ be the covariance matrix of the visual sensor. lidar Let w be the covariance matrix of the lidar. v w represents the initial weighting coefficients of the visual sensor. l The initial weighting coefficients for the lidar are given.

[0127] In this embodiment, the sensors of the unmanned device include at least one of a lidar, a vision sensor, and an inertial measurement unit. An initial weighting coefficient for the sensors includes the initial weighting coefficient of the lidar and / or the initial weighting coefficient of the vision sensor. The sensor acquisition residual at the time to be processed includes at least one of the acquisition residuals of the lidar, the vision sensor, and the inertial measurement unit at the time to be processed.

[0128] In the embodiments of this application, an initial function can be constructed using the initial weight coefficients of the currently traversed sensors and the sensor acquisition residuals at the time to be processed.

[0129] Based on the initial function, the initial function is solved, and the initial function is minimized by a nonlinear optimization algorithm to obtain the inference pose information of the unmanned device at the time to be processed.

[0130] The initial function is:

[0131]

[0132] Among them, F new (x) is the initial function, x includes the inference pose information of the unmanned device, r imu r represents the acquisition residual of the inertial measurement unit. vision r represents the acquisition residual of the vision sensor. lidar Σ represents the acquisition residual of the lidar. imu Let Σ be the covariance matrix of the inertial measurement unit. vision Let Σ be the covariance matrix of the visual sensor. lidar Let w be the covariance matrix of the lidar. v w represents the initial weighting coefficients for the visual sensor. l These are the initial weighting coefficients for the lidar.

[0133] Step 104: Determine the inference error corresponding to the initial weight coefficient of the current traversal based on the inference pose information and the actual pose information;

[0134] In this embodiment of the application, the inference error corresponding to the initial weight coefficient of the current traversal can be determined based on the inference pose information of the unmanned device at the time to be processed and the actual pose information of the unmanned device at the time to be processed.

[0135] Step 105: Determine the target weight coefficient from the at least one initial weight coefficient based on the inference error corresponding to the at least one initial weight coefficient;

[0136] In this embodiment, a target weight coefficient can be determined from at least one initial weight coefficient based on the inference error corresponding to at least one initial weight coefficient. For example, the initial weight coefficient with the smallest inference error can be used as the target weight coefficient.

[0137] In this embodiment of the application, the sensor-collected information at the time to be processed and the target weight coefficient at the time to be processed can be used as training data for the neural network model.

[0138] In this embodiment of the application, the sensor-acquired information includes image information from a visual sensor and / or point cloud information from a lidar. Image feature vectors can be extracted from the image information, and point cloud feature vectors can be extracted from the point cloud information. The image feature vectors at the time to be processed, the point cloud feature vectors at the time to be processed, and the target weight coefficients at the time to be processed are used as training data for the neural network model.

[0139] In one example, the initial weighting coefficients w of the visual sensor can be set first. v The initial weighting coefficient w of the lidar l The relation parameterization is w v =α, w l = 1 - α. Where α is a scalar between 0 and 1. The continuous search space of α is discretized into a finite set, which can be set as S = 0.0, 0.1, 0.2, ..., 1.0, with a total of 11 candidate weight values. This simplifies the search process, defines the weight search space, and obtains 11 initial weight coefficients for the sensor, that is, 11 initial weight coefficients for the vision sensor and 11 initial weight coefficients for the LiDAR.

[0140] Then, each moment in at least one historical moment can be used as the moment to be processed, and the following steps are performed to generate a data pair containing {feature, label}, obtain a pseudo-label training set, and train the neural network model.

[0141] Since the time to be processed can be represented by the image frame to be processed, and the processable time can be represented by the image frame that can be processed, refer to... Figure 4 The diagram illustrates a flowchart of a method for constructing a training set for a neural network model provided in an embodiment of this application.

[0142] First, a traversal search is performed, iterating through all 11 α values ​​in the weight search space S for the image frame i to be processed. Then, single-frame optimization is performed, constructing a simplified single-frame state estimation optimization problem for the currently traversed α value. The optimization variable in this problem is only the pose of the image frame i to be processed. i Specifically, based on the map surrounding the unmanned equipment in the processable image frame, the ideal acquisition information of the sensors in the image frame to be processed can be determined. Then, based on the ideal acquisition information and the sensor acquisition information corresponding to the image frame to be processed, the sensor acquisition residual corresponding to the image frame to be processed can be determined. An initial function is constructed based on the currently traversed α value and the sensor acquisition residual corresponding to the image frame to be processed. A nonlinear optimization solver (such as GN or LM) is run to solve the initial function to obtain the inferred pose corresponding to the currently traversed α value. ik GN refers to the Gauss-Newton method, and LM refers to the Levenberg-Marquardt method.

[0143] Next, error evaluation is performed. The calculated inference pose is then determined. ik The actual pose of the image frame to be processed i_gtThe inference error between the two. Since the actual pose information is a high-precision 6-DOF actual pose, including 3 translational degrees of freedom and 3 rotational degrees of freedom, the inference error can be defined as the weighted sum of the Euclidean distance of the translational part and the angle difference of the rotational part of the unmanned device. The inference error can be expressed as:

[0144]

[0145] Among them, Error k λ represents the inference error and is a preset balance coefficient.

[0146] Finally, for the image frame i to be processed, after traversing the 11 α values, the inference errors corresponding to all α values ​​are compared, and the target weight coefficient of the sensor corresponding to the image frame to be processed is determined based on the α value with the smallest inference error.

[0147] Step 106: Using the sensor acquisition information at the time to be processed and the target weight coefficient corresponding to the time to be processed, train the preset neural network model to obtain the sensor weight model.

[0148] In this embodiment of the application, the sensor weight model can be trained using the sensor acquisition information at the time to be processed and the target weight coefficient corresponding to the time to be processed.

[0149] In some embodiments of this application, the sensor acquisition information includes image information from the visual sensor and / or point cloud information from the lidar; the step of training a preset neural network model using the sensor acquisition information at the time to be processed and the target weight coefficients corresponding to the time to be processed to obtain a sensor weight model includes:

[0150] Extract the image feature vector from the image information and extract the point cloud feature vector from the point cloud information;

[0151] The image feature vector and / or the point cloud feature vector are input into the neural network model to obtain the training weight coefficients of the sensor;

[0152] The mean square error of the neural network model is determined based on the training weight coefficients and the target weight coefficients.

[0153] Based on the mean square error, the neural network model is adjusted, and it is determined whether the adjusted neural network model meets the preset model training objective.

[0154] If the conditions are not met, the steps of inputting the image feature vector and / or the point cloud feature vector into the neural network model to obtain the training weight coefficients of the sensor, determining the mean square error of the neural network model based on the training weight coefficients and the target weight coefficients, and adjusting the neural network model based on the mean square error are repeated until the adjusted neural network model meets the model training objective.

[0155] In this embodiment, image feature vectors are extracted from image information, and point cloud feature vectors are extracted from point cloud information. These image feature vectors and / or point cloud feature vectors are then input into a neural network model to obtain the sensor's training weight coefficients. Based on the training weight coefficients and target weight coefficients, the mean squared error (MSE) of the neural network model can be determined. The neural network model is then adjusted based on the MSE, and it is determined whether the adjusted model meets the preset model training objective. If it does, the adjusted neural network model is the sensor weight model for the unmanned device. If it does not, the process of inputting image feature vectors and / or point cloud feature vectors into the neural network model to obtain the sensor's training weight coefficients, determining the MSE of the neural network model based on the training weight coefficients and target weight coefficients, and adjusting the neural network model based on the MSE is repeated until the adjusted neural network model meets the model training objective. The model training objective can be: loss convergence on the validation set.

[0156] Reference Figure 3 The diagram illustrates a flowchart of a training method for a neural network model provided in an embodiment of this application. First, a training dataset is prepared, for example, nuScenes. Then, each frame i in the dataset is traversed to generate the optimal weight label for that single frame. Extracting image and point cloud features f vi f li Construct training samples {(f vi ,f li ),α i *}, thus obtaining the training set. Use the training set for supervised learning to train WANN. Set {(f vi ,f li ),α i * Inputting data into the WANN model, the WANN model can output a_pred and calculate the loss, Loss = MSE(a_pred). pred Using backpropagation, the WANN model is optimized. After training, the trained WANN model is saved.

[0157] In one example, the image information of the image frame to be processed is input into a MobileNetV2 network pre-trained on ImageNet, and a 1280-dimensional image feature vector f is obtained from its global average pooling layer. vi The laser point cloud information of the image frame to be processed (e.g., randomly sampled or sampled from the farthest point to 2048 points) is input into a PointNet encoder, which obtains a 1024-dimensional point cloud feature vector f from its global feature extraction module. li The feature vector pairs (f) of the image frames to be processed vi ,f li ) and α corresponding to the image frame to be processed * As a data pair, it is stored in the training set.

[0158] Then, the neural network model is trained using this training set. The neural network model can be a WANN (weight adjustment network). The training objective of the WANN is to train a neural network model that can predict the optimal weights based on sensor features.

[0159] WANN's network architecture is a multilayer perceptron (MLP) structure with two input branches, specifically including input branch 1 (vision), input branch 2 (laser), and a fusion and main body. Input branch 1 (vision) receives a 1280-dimensional image feature vector f. v A fully connected layer (linear projection layer) maps it to 256 dimensions. Input branch 2 (laser) receives a 1024-dimensional point cloud feature vector f. l The fully connected layer (linear projection layer) maps the vectors to 256 dimensions. The fusion and main body concatenates the 256-dimensional vectors output from the two branches, forming a 512-dimensional fusion vector. This vector then passes through two hidden layers (each with 256 neurons and using ReLU as the activation function), and finally through a single-neuron output layer using the Sigmoid activation function to ensure that the output training weight coefficients α are consistent. pred Between 0 and 1.

[0160] The WANN training process involves using deep learning frameworks such as PyTorch or TensorFlow to train {(f vi ,f li ),α * The dataset is used as training data. (f) vi ,f li As network input, α * As a label for supervised learning, mean squared error (MSE) is chosen as the loss function Loss = (α) pred -α * ) 2The Adam optimizer is used for multiple rounds of iterative training with an appropriate learning rate until the loss converges on the validation set.

[0161] After training is complete, the network parameters of WANN and the two linear projection layers are saved as a whole as a model file for use in the online running phase.

[0162] In yet another example, the steps for generating the training set for the neural network model are as follows:

[0163] (1) Data preparation: Select a dataset containing high-precision 6-DOF pose ground truth.

[0164] (2) Define the weight search space: Discretize the weights. For example, set the weights to be controlled by a single parameter α∈[0,1], w v =α,w l =1-α. The search space for α is set as a discrete set, such as S = 0.0, 0.1, 0.2, ..., 1.0.

[0165] (3) Traversal search to generate single-frame labels: For each frame in the dataset, traverse the discrete candidate weights and substitute them into the state estimation optimization problem to solve. Select the candidate weight that minimizes the error between the current frame pose and the true value as the optimal weight pseudo-label for that frame.

[0166] (4) Constructing the training set: Repeat the above steps for all frames in the dataset to generate an image feature vector f for each frame. vi and point cloud feature vector f li This ultimately forms a large-scale training pair set {((f)} vi ,f li ),α i * )}.

[0167] The steps for supervised learning training of WANN are as follows: using feature vector pairs (f vi ,f li () is the network input, with corresponding pseudo-tags. To monitor the signal, the mean squared error (MSE) loss function is used to train the parameters of WANN.

[0168] In some embodiments of this application, the method includes:

[0169] Acquire the current information collected by the at least one sensor on the unmanned device at the current moment;

[0170] Based on the currently acquired information, determine the current acquisition residual of the sensor;

[0171] Based on the currently collected information, the current weight coefficients of the at least one sensor are determined using the sensor weight model;

[0172] Using the current weight coefficients and the current acquisition residuals, construct the objective function corresponding to the initial function;

[0173] The current pose of the unmanned device is determined based on the objective function.

[0174] In this embodiment, the current data collected by at least one sensor from the unmanned device at the current moment is obtained. Then, based on the current data collected, the current acquisition residual of the sensor is determined. Based on the current data collected, the current weight coefficients of at least one sensor are determined using a sensor weighting model. Using the current weight coefficients and the current acquisition residuals, an objective function corresponding to the initial function is constructed. Based on the objective function, the current pose of the unmanned device is determined.

[0175] The objective function determines the current pose of the unmanned device, which involves using a nonlinear least squares algorithm, such as the GN (Gauss-Newton) or LM (Levenberg-Marquardt) algorithm, to find the optimal pose estimate. First, for ease of calculation and representation, the weighted residual terms are stacked into a total residual vector r(x), then the objective function can be simplified as follows:

[0176]

[0177] Taking the Gauss-Newton (GN) method as an example, it approximates the optimal solution by iteratively linearizing the problem at each step. Its core step is to... (The sentence is incomplete and requires more context to translate accurately.) k Near the same area, find an optimal increment Δx.

[0178] First, for the residual vector r(x) in the current estimate x k Perform a first-order Taylor expansion at this point:

[0179] r(x k +Δx)≈r(x k )+J(x k )Δx

[0180] in, The Jacobian matrix describes the first-order partial derivatives of the total residuals with respect to the state variable x. Substituting the above linearized model into the objective function aims to minimize:

[0181]

[0182] To find the Δx that minimizes this equation, we differentiate it and set the derivative to zero, obtaining the following core equation of the Gauss-Newton method:

[0183] (J T (x k )J(x k ))Δx=-J T (x k )r(x k )

[0184] The system iteratively solves for the increment Δx and uses the calculated increment to update the state until the convergence condition is met, thus obtaining the optimal estimate x*.

[0185] In some embodiments of this application, the currently acquired information includes at least one of the current image information of the visual sensor, the current inertial information of the inertial measurement unit, and the current point cloud information of the lidar; determining the current acquisition residual of the sensor based on the current acquired information includes:

[0186] Based on the current image information, determine the current image residual of the visual sensor; and / or,

[0187] Based on the current inertial information, determine the current inertial residual of the inertial measurement unit; and / or,

[0188] Based on the current point cloud information, determine the current point cloud residual of the lidar.

[0189] In the embodiments of this application, the unmanned device can determine the current image residual of the visual sensor based on the current image information; and / or, based on the current inertial information, it can determine the current inertial residual of the inertial measurement unit; and / or, based on the current point cloud information, it can determine the current point cloud residual of the lidar.

[0190] In one example, the visual measurement residuals generated by the visual front end are used to calculate photometric errors, for example, based on the sparse direct method:

[0191] The system selects a previously observed map point from the map; the coordinates of this point in the global coordinate system are... In addition to the coordinates, this point is accompanied by an 8×8 pixel reference image block Q. i The visual residual r corresponding to this map point. c That is, the current image I k The pixel block at the projection position of this point and the reference block Q i The difference lies in brightness. The calculation formula can be expressed as:

[0192]

[0193] Where: x k It is the current frame state (including pose) to be optimized; T I-1 (x k The matrix T represents the current pose, i.e., the transformation matrix from the global coordinate system to the IMU coordinate system (using the IMU coordinate system as the body coordinate system); C -1 It is the transformation matrix (extrinsic parameter) from the IMU coordinate system to the camera coordinate system; π(...) is the camera projection function, which projects 3D points onto the pixel plane; I k (...) involves sampling pixel values ​​for the projected points on the current image; the visual measurement residual is the set of photometric errors for all map points, i.e.

[0194]

[0195] The lidar measurement residuals generated by the laser front end, for example, are based on the distance error from a point to a local plane.

[0196] For a point p currently scanned by the lidar j L The system uses the current pose estimate to transform it to the world coordinate system. Then, it finds the local plane fitted by its nearest neighbors on the global map (defined by the normal vector u). j A little q on the flat surface j (Definition). LiDAR residual r l This is the perpendicular distance from the point to the plane. Its calculation formula can be expressed as:

[0197]

[0198] in: It is the original measurement point in the lidar coordinate system;

[0199] T L It is the transformation matrix (extrinsic parameter) from the lidar coordinate system to the IMU coordinate system; T I (x k ) represents the transformation matrix from the IMU coordinate system to the global coordinate system;

[0200] This represents the dot product operation, which calculates the distance from a point to a plane.

[0201] The residual of lidar measurement is the set of point-to-area distance errors between all lidar scan points and the global map, i.e.:

[0202]

[0203] IMU measurement residuals generated by the IMU pre-integration module, for example:

[0204] r imu =x k -x^k

[0205] Specifically, from the previous frame to the current frame, the system pre-integrates all high-frequency IMU measurements (angular velocity and acceleration) to obtain a prediction of the system state, namely x^k. The IMU residual r... imu This refers to the difference between the "state independently predicted by the IMU" x^k and the "state estimated by the optimizer" x. k The differences between them include the residuals of each item in the state, such as rotation residuals, translation residuals, velocity residuals, etc.

[0206] In some embodiments of this application, determining the current weight coefficient of the at least one sensor based on the currently acquired information using the sensor weight model includes:

[0207] The current image information is input into the backbone network of a preset convolutional neural network to obtain the current image feature vector corresponding to the current image information;

[0208] The current point cloud information is input into a preset point cloud encoder to obtain the current point cloud feature vector corresponding to the current point cloud information;

[0209] The current image feature vector and / or the current point cloud feature vector are input into the sensor weight model to obtain the current weight coefficients of the at least one sensor.

[0210] In this embodiment, the current image information is input into the backbone network of a preset convolutional neural network to obtain the current image feature vector corresponding to the current image information; the current point cloud information is input into a preset point cloud encoder to obtain the current point cloud feature vector corresponding to the current point cloud information; the current image feature vector and / or the current point cloud feature vector are input into a sensor weight model to obtain the current weight coefficient of at least one sensor.

[0211] Reference Figure 2 This document illustrates a flowchart of a pose determination method for an unmanned device provided in an embodiment of this application. The LIVO system is run online, and the offline-trained WANN model is deployed to a specific mobile robot / drone or autonomous vehicle platform. During the operation of a standard LIVO system, when the system triggers a state estimation, the following steps are performed:

[0212] (1) Obtain the input required for optimization: Obtain all the information required for this state estimation from the front-end module of the LIVO system. This mainly includes: the image and lidar point cloud of the current frame, the visual measurement residuals generated by the visual front-end (e.g., reprojection error), the lidar measurement residuals generated by the laser front-end (e.g., point-to-surface distance error), and the IMU measurement residuals generated by the IMU pre-integration module.

[0213] (2) Real-time generation of dynamic weights: The image of the current frame is input into the pre-trained CNN backbone network to extract the image feature vector f. v Simultaneously, the corresponding point cloud is input into a pre-trained point cloud encoder to extract the point cloud feature vector f. l The feature vector (f) v ,f l The input is fed into the loaded WANN model for a fast forward inference. The WANN's sigmoid output layer immediately provides a scalar α between 0 and 1. The system uses this output to set the dynamic weights for the current frame: w v =α,w l =1-α.

[0214] (3) State estimation: Construct an optimization problem and build an objective function based on dynamic weights:

[0215]

[0216] Among them, the visual residual term and the laser residual term are respectively determined by the weights w output by WANN. v and w l Weighting is applied, and a nonlinear optimization library (such as Ceres Solver or g2o) is called to iteratively optimize the dynamically weighted factor graph to find the optimal estimate of the pose.

[0217] (4) Complete and continue: The optimized pose will be used in LIVO to update the system map and other subsequent modules. The system then continues to process new sensor data and waits for the next state estimation to be triggered.

[0218] In one example, the online operation method of the LIVO system is as follows:

[0219] (1) Multimodal feature extraction:

[0220] a. Image Feature Extraction: The image of the current frame is input into the backbone of a lightweight convolutional neural network (CNN) pre-trained on a large image dataset (such as ImageNet), such as MobileNetV2 or EfficientNet-B0. Leveraging its powerful general feature extraction capabilities, after the network's global average pooling layer, a high-dimensional image feature vector f is output, representing information such as image texture richness and illumination intensity. v (For example, 1280 dimensions).

[0221] b. Point Cloud Feature Extraction: The laser point cloud of the current frame is input into a pre-trained point cloud feature extraction network, such as an encoder based on the PointNet architecture. This network directly processes the raw point cloud, aggregating the features of all points through max pooling, and outputting a global point cloud feature vector f that can characterize the sparseness of the 3D structure and the strength of geometric features of the point cloud. l (For example, 1024 dimensions).

[0222] (2) Feature fusion and weight inference:

[0223] The extracted high-dimensional image feature vector f v and point cloud feature vector f l Simultaneously, the input is fed into a pre-trained weighted neural network (WANN). The internal structure of the WANN first performs dimensional alignment and fusion of input features of different dimensions through parallel processing branches, and then performs inference through subsequent network layers to quickly output the weight w representing the visual residual term of the current frame. v The weight w of the laser residual term l .

[0224] (3) Dynamically weighted backend optimization:

[0225] Construct an objective function that includes the IMU measurement residual, the visual measurement residual, and the laser measurement residual:

[0226]

[0227] Among them, the visual residual term and the laser residual term are respectively determined by the weights w output by WANN. v and w l Weighting is applied, with the weights of the IMU residuals remaining unchanged because they are typically less affected by environmental factors and serve as the basis for system motion constraints. Finally, a nonlinear optimization algorithm is used to solve the objective function, yielding the optimized system state.

[0228] In this embodiment, the unmanned device includes at least one sensor to acquire the actual pose information of the unmanned device and / or sensor acquisition information of the unmanned device at at least one historical moment, and to take any moment in the at least one historical moment as the moment to be processed, and any moment before the moment to be processed as the moment that can be processed; to determine the ideal acquisition information of the sensor at the moment to be processed based on the surrounding map of the unmanned device at the moment that can be processed, and to determine the sensor acquisition residual of the sensor at the moment to be processed based on the ideal acquisition information and the sensor acquisition information; to acquire at least one initial weight coefficient of the sensor, and to traverse the initial weight coefficients, and to determine the inferred pose information of the unmanned device at the moment to be processed using the currently traversed initial weight coefficients and the sensor acquisition residual of the sensor; and to determine the inferred pose information of the unmanned device based on the inferred pose information. Based on pose information and actual pose information, the inference error corresponding to the initial weight coefficients of the current traversal is determined. Based on the inference error corresponding to at least one initial weight coefficient, the target weight coefficient is determined from at least one initial weight coefficient. Using the sensor acquisition information at the time to be processed and the target weight coefficient corresponding to the time to be processed, a preset neural network model is trained to obtain a sensor weight model. Using the sensor weight model, sensor weights for unmanned equipment in different environments can be generated. This solves the problem that the weight coefficients corresponding to the residual terms of visual sensors and lidar are fixed during the operation of unmanned equipment and cannot change with dynamic and complex environments, resulting in poor adaptability and poor positioning and mapping accuracy of unmanned equipment in dynamic and complex environments.

[0229] In this embodiment, a lightweight weighted neural network (WANN) is introduced. This network can dynamically evaluate the "confidence" of various sensors based on the input real-time image and laser point cloud data, and output the corresponding weights. This solves the technical problem in the tightly coupled LIVO method, where the use of fixed or static weights to fuse visual and LiDAR residuals leads to poor adaptability and insufficient robustness of the system in complex dynamic environments. It also solves the problem that the setting of weights largely depends on the experience of engineers for manual adjustment, making it difficult to find a parameter combination that performs optimally in all scenarios and resulting in poor generalization ability.

[0230] In the embodiments of this application, the system perceives environmental changes in real time and intelligently relies on the more reliable sensor, which has strong environmental adaptability. By dynamically suppressing the influence of unreliable sensors, the system effectively prevents tracking failure in scenarios where a single sensor degrades, and obtains higher positioning accuracy, significantly improving robustness and accuracy. The training scheme proposed in this application, as well as the feature fusion and weight inference methods based on the pre-trained model, constitute a complete, specific, and feasible technical route, which has both theoretical innovation and engineering practicality.

[0231] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of this application are not limited to the described order of actions, because according to the embodiments of this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of this application.

[0232] Reference Figure 5 This diagram illustrates a structural block diagram of a sensor weight model determination device for an unmanned device provided in an embodiment of this application. The device is applied to an unmanned device, which includes at least one sensor and may specifically include the following modules:

[0233] The information acquisition module 501 is used to acquire the actual pose information of the unmanned device and / or the sensor acquisition information of the unmanned device by the sensor at at least one historical moment, and to take any one of the at least one historical moments as the moment to be processed, and any moment before the moment to be processed as the moment that can be processed.

[0234] The residual determination module 502 is used to determine the ideal acquisition information of the sensor at the unprocessed time based on the surrounding map of the unmanned equipment at the processable time, and to determine the sensor acquisition residual of the sensor at the unprocessed time based on the ideal acquisition information and the sensor acquisition information.

[0235] The inference pose information determination module 503 is used to obtain at least one initial weight coefficient of the sensor, and traverse the initial weight coefficients, and use the currently traversed initial weight coefficients and the sensor acquisition residual of the sensor to determine the inference pose information of the unmanned device at the time to be processed.

[0236] The inference error determination module 504 is used to determine the inference error corresponding to the initial weight coefficient of the current traversal based on the inference pose information and the actual pose information.

[0237] The target weight coefficient determination module 505 is used to determine a target weight coefficient from the at least one initial weight coefficient based on the inference error corresponding to the at least one initial weight coefficient.

[0238] The training module 506 is used to train a preset neural network model using the sensor acquisition information at the time to be processed and the target weight coefficients corresponding to the time to be processed, so as to obtain a sensor weight model.

[0239] In one optional embodiment of this application, the apparatus includes:

[0240] The actual data acquisition module is used to acquire the actual data collected by the sensor on the environment surrounding the unmanned device at the at least one historical moment.

[0241] The map building module is used to build a map of the unmanned device's surroundings at the processable time based on the actual pose information and / or actual collected information corresponding to the processable time.

[0242] In one optional embodiment of this application, the sensor includes at least one of a lidar, a vision sensor, and an inertial measurement unit; the initial weighting coefficient includes the initial weighting coefficient of the lidar and / or the initial weighting coefficient of the vision sensor; the sensor acquisition residual includes at least one of the acquisition residual of the lidar, the acquisition residual of the vision sensor, and the acquisition residual of the inertial measurement unit.

[0243] The inference pose information determination module includes:

[0244] The initial function construction submodule is used to construct the initial function using the initial weight coefficients of the current traversal and the sensor acquisition residuals of the sensor;

[0245] The inference pose information determination submodule is used to determine the inference pose information based on the initial function; the initial function is:

[0246]

[0247] Among them, F new (x) is the initial function, x includes the inference pose information of the unmanned device, r imu r is the acquisition residual of the inertial measurement unit. vision Let r be the acquisition residual of the visual sensor. lidar Σ represents the acquisition residual of the lidar. imu Let Σ be the covariance matrix of the inertial measurement unit. vision Let Σ be the covariance matrix of the visual sensor. lidar Let w be the covariance matrix of the lidar. v w represents the initial weighting coefficients of the visual sensor. l The initial weighting coefficients for the lidar are given.

[0248] In one optional embodiment of this application, the sensor-acquired information includes image information from the visual sensor and / or point cloud information from the lidar; the training module includes:

[0249] The feature vector extraction submodule is used to extract image feature vectors from the image information and point cloud feature vectors from the point cloud information.

[0250] The training weight coefficients submodule is used to input the image feature vector and / or the point cloud feature vector into the neural network model to obtain the training weight coefficients of the sensor;

[0251] The mean squared error determination submodule is used to determine the mean squared error of the neural network model based on the training weight coefficients and the target weight coefficients.

[0252] The adjustment submodule is used to adjust the neural network model according to the mean square error, and to determine whether the adjusted neural network model meets the preset model training objective.

[0253] The repeat submodule is used to repeat the steps of inputting the image feature vector and / or the point cloud feature vector into the neural network model to obtain the training weight coefficients of the sensor, determining the mean square error of the neural network model based on the training weight coefficients and the target weight coefficients, and adjusting the neural network model based on the mean square error, until the adjusted neural network model meets the model training objective.

[0254] In one optional embodiment of this application, the apparatus includes:

[0255] The current information acquisition module is used to acquire the current information collected by the at least one sensor on the unmanned device at the current moment;

[0256] The current acquisition residual determination module is used to determine the current acquisition residual of the sensor based on the current acquisition information.

[0257] The current weight coefficient determination module is used to determine the current weight coefficient of the at least one sensor based on the currently collected information and using the sensor weight model.

[0258] The objective function construction module is used to construct the objective function corresponding to the initial function using the current weight coefficients and the current acquisition residuals;

[0259] The current pose determination module is used to determine the current pose of the unmanned device based on the objective function.

[0260] In one optional embodiment of this application, the currently acquired information includes at least one of the current image information of the visual sensor, the current inertial information of the inertial measurement unit, and the current point cloud information of the lidar; the currently acquired residual determination module includes:

[0261] The residual determination submodule is used to determine the current image residual of the visual sensor based on the current image information; and / or,

[0262] Based on the current inertial information, determine the current inertial residual of the inertial measurement unit; and / or,

[0263] Based on the current point cloud information, determine the current point cloud residual of the lidar.

[0264] In one optional embodiment of this application, the current weight coefficient determination module includes:

[0265] The image feature vector acquisition submodule is used to input the current image information into the backbone network of a preset convolutional neural network to obtain the current image feature vector corresponding to the current image information;

[0266] The point cloud feature vector acquisition submodule is used to input the current point cloud information into a preset point cloud encoder to obtain the current point cloud feature vector corresponding to the current point cloud information;

[0267] The current weight coefficient acquisition submodule is used to input the current image feature vector and / or the current point cloud feature vector into the sensor weight model to obtain the current weight coefficient of the at least one sensor.

[0268] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0269] In addition, embodiments of this application also provide an electronic device, such as... Figure 6 As shown, it includes a processor 601, a communication interface 602, a memory 603, and a communication bus 604, wherein the processor 601, the communication interface 602, and the memory 603 communicate with each other through the communication bus 604.

[0270] Memory 603 is used to store computer programs;

[0271] When processor 601 executes a program stored in memory 603, it performs the following steps:

[0272] Acquire the actual pose information of the unmanned device and / or the sensor-collected information of the unmanned device by the sensor at at least one historical moment, and take any one of the at least one historical moments as the moment to be processed, and any moment before the moment to be processed as the moment that can be processed.

[0273] Based on the surrounding map of the unmanned equipment at the processable time, the ideal acquisition information of the sensor at the unprocessed time is determined, and based on the ideal acquisition information and the sensor acquisition information, the sensor acquisition residual of the sensor at the unprocessed time is determined.

[0274] At least one initial weight coefficient of the sensor is obtained, and the initial weight coefficients are traversed. The inference pose information of the unmanned device at the time to be processed is determined by using the currently traversed initial weight coefficients and the sensor acquisition residuals of the sensor.

[0275] Based on the inferred pose information and the actual pose information, determine the inference error corresponding to the initial weight coefficient of the current traversal;

[0276] Based on the inference error corresponding to the at least one initial weight coefficient, a target weight coefficient is determined from the at least one initial weight coefficient;

[0277] The sensor weight model is obtained by training a preset neural network model using the sensor acquisition information at the time to be processed and the target weight coefficient corresponding to the time to be processed.

[0278] In one optional embodiment of this application, the method includes:

[0279] Acquire the actual information collected by the sensor about the environment surrounding the unmanned device at the at least one historical moment;

[0280] Based on the actual pose information and / or actual collected information corresponding to the processable moment, a map of the unmanned device's surroundings at the processable moment is constructed.

[0281] In one optional embodiment of this application, the sensor includes at least one of a lidar, a vision sensor, and an inertial measurement unit; the initial weighting coefficient includes the initial weighting coefficient of the lidar and / or the initial weighting coefficient of the vision sensor; the sensor acquisition residual includes at least one of the acquisition residual of the lidar, the acquisition residual of the vision sensor, and the acquisition residual of the inertial measurement unit.

[0282] The step of determining the inference pose information of the unmanned device at the time to be processed by using the initial weight coefficients of the current traversal and the sensor acquisition residuals of the sensor includes:

[0283] An initial function is constructed using the initial weight coefficients of the current traversal and the sensor acquisition residuals of the sensor;

[0284] The inference pose information is determined based on the initial function; the initial function is:

[0285]

[0286] Among them, F new (x) is the initial function, x includes the inference pose information of the unmanned device, r imu r is the acquisition residual of the inertial measurement unit. vision Let r be the acquisition residual of the visual sensor. lidar Σ represents the acquisition residual of the lidar. imu Let Σ be the covariance matrix of the inertial measurement unit. vision Let Σ be the covariance matrix of the visual sensor. lidar Let w be the covariance matrix of the lidar. v w represents the initial weighting coefficients of the visual sensor. l The initial weighting coefficients for the lidar are given.

[0287] In one optional embodiment of this application, the sensor acquisition information includes image information from the visual sensor and / or point cloud information from the lidar; the step of training a preset neural network model using the sensor acquisition information at the time to be processed and the target weight coefficients corresponding to the time to be processed to obtain a sensor weight model includes:

[0288] Extract the image feature vector from the image information and extract the point cloud feature vector from the point cloud information;

[0289] The image feature vector and / or the point cloud feature vector are input into the neural network model to obtain the training weight coefficients of the sensor;

[0290] The mean square error of the neural network model is determined based on the training weight coefficients and the target weight coefficients.

[0291] Based on the mean square error, the neural network model is adjusted, and it is determined whether the adjusted neural network model meets the preset model training objective.

[0292] If the conditions are not met, the steps of inputting the image feature vector and / or the point cloud feature vector into the neural network model to obtain the training weight coefficients of the sensor, determining the mean square error of the neural network model based on the training weight coefficients and the target weight coefficients, and adjusting the neural network model based on the mean square error are repeated until the adjusted neural network model meets the model training objective.

[0293] In one optional embodiment of this application, the method includes:

[0294] Acquire the current information collected by the at least one sensor on the unmanned device at the current moment;

[0295] Based on the currently acquired information, determine the current acquisition residual of the sensor;

[0296] Based on the currently collected information, the current weight coefficients of the at least one sensor are determined using the sensor weight model;

[0297] Using the current weight coefficients and the current acquisition residuals, construct the objective function corresponding to the initial function;

[0298] The current pose of the unmanned device is determined based on the objective function.

[0299] In one optional embodiment of this application, the currently acquired information includes at least one of the current image information of the visual sensor, the current inertial information of the inertial measurement unit, and the current point cloud information of the lidar; determining the current acquisition residual of the sensor based on the current acquired information includes:

[0300] Based on the current image information, determine the current image residual of the visual sensor; and / or,

[0301] Based on the current inertial information, determine the current inertial residual of the inertial measurement unit; and / or,

[0302] Based on the current point cloud information, determine the current point cloud residual of the lidar.

[0303] In one optional embodiment of this application, determining the current weight coefficient of the at least one sensor based on the currently acquired information and using the sensor weight model includes:

[0304] The current image information is input into the backbone network of a preset convolutional neural network to obtain the current image feature vector corresponding to the current image information;

[0305] The current point cloud information is input into a preset point cloud encoder to obtain the current point cloud feature vector corresponding to the current point cloud information;

[0306] The current image feature vector and / or the current point cloud feature vector are input into the sensor weight model to obtain the current weight coefficients of the at least one sensor.

[0307] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0308] The communication interface is used for communication between the aforementioned terminal and other devices.

[0309] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0310] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0311] like Figure 7 As shown, in another embodiment provided in this application, a computer-readable storage medium 701 is also provided, which stores instructions that, when executed on a computer, cause the computer to execute a method for determining a sensor weight model for an unmanned device as described in the above embodiments.

[0312] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to execute a method for determining a sensor weight model for an unmanned device as described in the above embodiments.

[0313] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0314] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0315] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0316] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.

Claims

1. A method for determining a sensor weight model of an unmanned device, the method comprising: The method is applied to an unmanned device including at least one sensor, and comprises the following steps: acquiring actual pose information of the unmanned device and sensor acquisition information of the unmanned device by the sensor at at least one historical moment, taking any moment in the at least one historical moment as a to-be-processed moment, and taking any moment before the to-be-processed moment as a processable moment; determining ideal acquisition information of the sensor at the to-be-processed moment according to a surrounding map of the unmanned device at the processable moment, and determining sensor acquisition residual of the sensor at the to-be-processed moment according to the ideal acquisition information and the sensor acquisition information; acquiring at least one initial weight coefficient of the sensor, and traversing the initial weight coefficient, and determining inferred pose information of the unmanned device at the to-be-processed moment by using the currently traversed initial weight coefficient and the sensor acquisition residual of the sensor; determining an inference error corresponding to the currently traversed initial weight coefficient according to the inferred pose information and the actual pose information; determining a target weight coefficient in the at least one initial weight coefficient according to the inference error corresponding to the at least one initial weight coefficient; training a preset neural network model by using the sensor acquisition information at the to-be-processed moment and the target weight coefficient corresponding to the to-be-processed moment, and obtaining a sensor weight model.

2. The method of claim 1, wherein, The method comprises the following steps: acquiring actual acquisition information of a surrounding environment of the unmanned device by the sensor at the at least one historical moment; constructing a surrounding map of the unmanned device at the processable moment according to the actual pose information and the actual acquisition information corresponding to the processable moment.

3. The method of claim 1, wherein, The sensor comprises at least one of a laser radar, a visual sensor and an inertial measurement unit; the initial weight coefficient comprises an initial weight coefficient of the laser radar and / or an initial weight coefficient of the visual sensor; the sensor acquisition residual comprises at least one of acquisition residual of the laser radar, acquisition residual of the visual sensor and acquisition residual of the inertial measurement unit; The method comprises the following steps: constructing an initial function by using the currently traversed initial weight coefficient and the sensor acquisition residual of the sensor; determining the inferred pose information according to the initial function; the initial function is: wherein F new (x) is the initial function, x includes the inferred pose information of the unmanned device, r imu is the collection residual error of the inertial measurement unit, r vision is the collection residual error of the visual sensor, r lidar is the collection residual error of the laser radar, Σ imu is the covariance matrix of the inertial measurement unit, Σ vision is the covariance matrix of the visual sensor, Σ lidar is the covariance matrix of the laser radar, is the initial weight coefficient of the visual sensor, is the initial weight coefficient of the laser radar.

4. The method of claim 3, wherein, The sensor acquisition information comprises picture information of the visual sensor and / or point cloud information of the laser radar; the method comprises the following steps: extracting an image feature vector in the picture information and extracting a point cloud feature vector in the point cloud information; inputting the image feature vector and / or the point cloud feature vector into the neural network model to obtain a training weight coefficient of the sensor. determine a mean square error of the neural network model according to the training weight coefficient and the target weight coefficient; adjust the neural network model according to the mean square error, and determine whether the adjusted neural network model meets a preset model training target; if not, repeat the steps of inputting the image feature vector and / or the point cloud feature vector into the neural network model to obtain the training weight coefficient of the sensor, determining a mean square error of the neural network model according to the training weight coefficient and the target weight coefficient, and adjusting the neural network model according to the mean square error until the adjusted neural network model meets the model training target.

5. The method of claim 3, wherein, The method comprises: obtaining current collection information of the unmanned device by the at least one sensor at a current time; determining a current collection residual error of the sensor according to the current collection information; determining a current weight coefficient of the at least one sensor by using the sensor weight model based on the current collection information; constructing a target function corresponding to the initial function by using the current weight coefficient and the current collection residual error; determining a current pose of the unmanned device based on the target function.

6. The method of claim 5, wherein, The current collection information comprises at least one of current image information of the visual sensor, current inertial information of the inertial measurement unit, and current point cloud information of the laser radar; and the determination of the current collection residual error of the sensor according to the current collection information comprises: determining a current image residual error of the visual sensor according to the current image information; and / or determining a current inertial residual error of the inertial measurement unit according to the current inertial information; and / or determining a current point cloud residual error of the laser radar according to the current point cloud information.

7. The method of claim 6, wherein, The determination of the current weight coefficient of the at least one sensor by using the sensor weight model based on the current collection information comprises: inputting the current image information into a backbone network of a preset convolutional neural network to obtain a current image feature vector corresponding to the current image information; inputting the current point cloud information into a preset point cloud encoder to obtain a current point cloud feature vector corresponding to the current point cloud information; inputting the current image feature vector and / or the current point cloud feature vector into the sensor weight model to obtain the current weight coefficient of the at least one sensor.

8. A device for determining the sensor weight model of an unmanned device, characterized in that, The device is applied to an unmanned device, and the unmanned device comprises at least one sensor, and the device comprises: an information acquisition module, configured to acquire actual pose information of the unmanned device and sensor collection information of the sensor on the unmanned device at at least one historical time, and take any time in the at least one historical time as a to-be-processed time and take any time before the to-be-processed time as a processable time; a residual determination module, configured to determine ideal collection information of the sensor at the to-be-processed time according to a surrounding map of the unmanned device at the processable time, and determine sensor collection residual of the sensor at the to-be-processed time according to the ideal collection information and the sensor collection information; an inference pose information determination module, configured to obtain at least one initial weight coefficient of the sensor, and traverse the initial weight coefficient, and determine inference pose information of the unmanned device at the to-be-processed time by using the currently traversed initial weight coefficient and the sensor collection residual of the sensor; an inference error determination module, configured to determine an inference error corresponding to the currently traversed initial weight coefficient according to the inference pose information and the actual pose information; a target weight coefficient determination module, configured to determine a target weight coefficient from the at least one initial weight coefficient according to the inference error corresponding to the at least one initial weight coefficient; a training module, configured to train a preset neural network model by using the sensor collection information at the to-be-processed time and the target weight coefficient corresponding to the to-be-processed time, and obtain a sensor weight model.

9. An electronic device, comprising: comprising a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory complete mutual communication through the communication bus; the memory is used to store a computer program; the processor is used to execute the program stored on the memory, and implement the method in any one of claims 1-7.

10. One or more computer readable media having instructions stored thereon that, when executed by one or more processors, cause the processors to perform the method of any one of claims 1-7.

Citation Information

Patent Citations

  • Unmanned aerial vehicle pose estimation method based on multi-sensor fusion

    CN113432602A

  • Double-sensor target detection fusion method and system based on evidence reasoning rule

    CN115049996A