Multi-sensor fusion-based depth estimation model training method, multi-sensor fusion method and related equipment

By constructing a multi-sensor fusion depth estimation model, using the training methods of encoders, decoders and fusion decoders, and combining shallow feature credibility judgment, the problem of decreased model reliability caused by sensor degradation is solved, and stable depth estimation is achieved in complex environments.

CN120672819APending Publication Date: 2025-09-19SHENZHEN INST OF ARTIFICIAL INTELLIGENCE & ROBOTICS FOR SOC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510760056.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing multi-sensor depth estimation methods have difficulty adapting to real-world scenarios when faced with noise interference, occlusion, distortion, and hardware failures, resulting in reduced reliability of model fusion results and a lack of diagnostic capabilities for sensor degradation.

Method used

A depth estimation model based on multi-sensor fusion is constructed, including an encoder, a decoder and a fusion decoder. The model parameters are trained through a preset loss function. Combined with the credibility judgment of shallow features, degraded sensors are identified and rejected, and a reliable sensor combination is selected for depth estimation.

Benefits of technology

It improves the robustness and generalization ability of the multi-sensor fusion depth estimation model in complex environments, can achieve stable and accurate depth estimation when sensor data is degraded, and has the ability to autonomously diagnose and respond to sensor failures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120672819A_ABST
    Figure CN120672819A_ABST
Patent Text Reader

Abstract

The invention discloses a training method of a depth estimation model based on multi-sensor fusion, a multi-sensor fusion method and related equipment, which are used for realizing a reliable depth estimation result when degradation interference exists in a sensor. The method specifically comprises the following steps: constructing a depth estimation model based on multi-sensor fusion, wherein the model comprises an encoder, a decoder and a plurality of fusion decoders set by each type of sensor; inputting observation data of each sensor into an encoder and a decoder of each sensor to generate a depth map corresponding to each sensor, and updating network parameters of the encoder and the decoder of each sensor according to an error between the depth map and a real depth map; and inputting observation data of any sensor combination into an encoder to extract multi-level feature data, sending the multi-level feature data into a fusion decoder to generate a depth map based on multi-sensor fusion, updating network parameters of the fusion decoder according to an error between the depth map and a real depth map, and finally obtaining a pre-trained depth estimation model based on multi-sensor fusion.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of sensor information processing, and in particular to a training method for a depth estimation model based on multi-sensor fusion, a multi-sensor fusion method, and related equipment. Background Art

[0002] With the increasing demand for depth estimation accuracy, multi-sensor fusion (MSF) technology has become an important means to improve its performance. However, existing multi-sensor depth estimation (MSDE) methods generally rely on ideal sensor inputs in training data, making them difficult to adapt to data disturbances caused by noise interference, occlusion, distortion, and hardware failures in real-world scenarios. As a result, they are prone to introducing low-quality data into depth estimation models based on multi-sensor fusion, affecting the reliability of the model fusion results. Summary of the Invention

[0003] Based on the above problems, the embodiments of the present application provide a training method, a multi-sensor fusion method and related equipment for a depth estimation model based on multi-sensor fusion, with the aim of improving the robustness and generalization ability of the depth estimation model based on multi-sensor fusion in complex environments, and achieving reliable depth estimation results when there is degradation, interference or uncertainty in the sensor observation data.

[0004] In a first aspect, an embodiment of the present application provides a training method for a depth estimation model based on multi-sensor fusion, including:

[0005] Constructing a depth estimation model based on multi-sensor fusion, the depth estimation model based on multi-sensor fusion comprising an encoder set for each type of sensor, a decoder set for each type of sensor, and at least one fusion decoder for fusing inputs of different types of sensors;

[0006] Input the observation data of each type of sensor into its corresponding encoder and decoder to generate the depth map corresponding to each sensor;

[0007] Based on the error between the depth map corresponding to each sensor and the true depth map at the same viewing angle, the encoder and decoder of the sensor are trained according to a preset loss function to minimize the error, thereby obtaining updated network parameters of the encoder and decoder; the network parameters include at least convolution weight values ​​and convolution bias values;

[0008] For any sensor combination consisting of different types of sensors, the observation data of each sensor in the sensor combination is input into its corresponding encoder to obtain multi-level feature data;

[0009] Inputting the multi-level feature data into a fusion decoder corresponding to the sensor combination to generate a depth map based on multi-sensor fusion;

[0010] According to the error between the depth map based on multi-sensor fusion and the real depth map at the same viewing angle, the fusion decoder is trained according to the preset loss function to minimize the error, thereby obtaining the network parameters of the updated fusion decoder;

[0011] A pre-trained depth estimation model based on multi-sensor fusion is obtained according to the updated network parameters of each encoder and the updated network parameters of each fusion encoder.

[0012] In one embodiment, the encoder and decoder of the sensor are trained according to a preset loss function to minimize the error between the depth map corresponding to each sensor and the true depth map at the same viewing angle. The encoder and decoder of the sensor are trained according to the preset loss function to minimize the prediction error, which is implemented by the following function:

[0013]

[0014] Among them, g i represents the encoder of the i-th sensor; f i represents the decoder of the i-th sensor; Indicates three sensor types, v, l, and t represent vision, lidar, and thermal imaging sensors respectively; K represents the total number of training samples involved in the training; k represents the sample index, indicating the kth training sample; represents the depth map corresponding to the i-th sensor generated by the k-th training sample through the i-th sensor encoder and the i-th sensor decoder; y k represents the true depth map of the kth training sample; Represents the preset loss function that measures the error between the depth map corresponding to the i-th sensor and the true depth map under the same training sample.

[0015] In one embodiment, the training of the fusion decoder according to the preset loss function to minimize the error between the depth map based on multi-sensor fusion and the real depth map at the same viewing angle is specifically implemented by the following function:

[0016]

[0017] Among them, f j represents the fusion decoder for processing the j-th sensor combination; Represents different types of sensor combinations, vl, vt, lt, vlt respectively represent vl represents the sensor combination consisting of visual sensor and lidar sensor, vt represents the sensor combination consisting of visual sensor and thermal imaging sensor, lt represents the sensor combination consisting of lidar sensor and thermal imaging sensor, and vlt represents the sensor combination consisting of visual sensor, lidar sensor, and infrared sensor; K represents the total number of training samples participating in the training; k represents the sample index, which represents the kth training sample; represents the depth map based on multi-sensor fusion generated by the encoder corresponding to the j-th sensor combination and the fusion decoder corresponding to the j-th sensor combination for the k-th training sample; y k represents the true depth map of the kth training sample; represents a preset loss function for measuring the error between the depth map based on multi-sensor fusion and the true depth map under the same training sample.

[0018] In one embodiment, the specific form of the preset loss function is:

[0019]

[0020] Where e represents a natural constant; K represents the total number of training samples involved in the training; k represents the sample index, which represents the kth training sample; s k Represents the uncertainty value given by the prediction result of the k-th training sample at the pixel level based on the depth estimation model based on multi-sensor fusion; represents the uncertainty weighting term, which is used to suppress the impact of high uncertainty pixels on the loss; Represents the depth map corresponding to the i-th sensor generated by the k-th training sample through the encoder and decoder of the i-th sensor Or the depth map based on multi-sensor fusion generated by the k-th training sample through the encoder and fusion decoder of the j-th sensor combination y k represents the true depth map of the kth training sample.

[0021] In a second aspect, an embodiment of the present application further provides a multi-sensor fusion method, which is applied to a depth estimation model based on multi-sensor fusion obtained by the method described in the first aspect of the embodiment of the present application or any specific implementation manner of the first aspect, wherein the depth estimation model based on multi-sensor fusion includes an encoder set for each type of sensor, a decoder set for each type of sensor, and at least one fusion decoder for fusing inputs of different types of sensors;

[0022] The multi-sensor fusion method comprises:

[0023] Acquiring observation data of at least one sensor combination composed of different types of sensors;

[0024] Inputting the observation data of each sensor in the sensor combination into its corresponding encoder respectively to extract shallow features of each of the observation data;

[0025] Performing a credibility judgment on each sensor based on shallow features extracted from each sensor to determine a credible sensor;

[0026] Based on the trusted sensor, a fusion decoder corresponding to the trusted sensor is selected, multi-level feature data corresponding to each of the trusted sensors is fused, and a fusion result is output as a predicted depth map of the current scene.

[0027] In one embodiment, the performing credibility judgment on each sensor based on the shallow features extracted from each sensor to determine the credible sensor includes:

[0028] Calculating an average value of the shallow features extracted by each of the sensors;

[0029] Calculating the cosine similarity between the average value of the shallow features of each sensor and the average value of the shallow features in the standard sample set;

[0030] Sensors whose cosine similarity is higher than a preset similarity threshold are determined as trustworthy sensors.

[0031] In a third aspect, an embodiment of the present application further provides a training device for a depth estimation model based on multi-sensor fusion, comprising:

[0032] A model construction unit, configured to construct a depth estimation model based on multi-sensor fusion, wherein the depth estimation model based on multi-sensor fusion includes an encoder set for each type of sensor, a decoder set for each type of sensor, and at least one fusion decoder for fusing inputs of different types of sensors;

[0033] A depth map generation unit, configured to input observation data of each type of sensor into its corresponding encoder and decoder to generate a depth map corresponding to each sensor;

[0034] A single-modal training unit is configured to train the encoder and decoder of the sensor according to a preset loss function based on the error between the depth map corresponding to each sensor and the true depth map at the same viewing angle to minimize the error, thereby obtaining updated network parameters of the encoder and decoder; the network parameters include at least convolution weight values ​​and convolution bias values;

[0035] A feature extraction unit is configured to input observation data of each sensor in any sensor combination composed of different types of sensors into its corresponding encoder to obtain multi-level feature data;

[0036] The depth map generating unit is further configured to input the multi-level feature data into a fusion decoder corresponding to the sensor combination to generate a depth map based on multi-sensor fusion;

[0037] a multimodal training unit, configured to train the fusion decoder according to the preset loss function to minimize the error between the depth map based on multi-sensor fusion and the real depth map at the same viewing angle, and obtain network parameters of the updated fusion decoder;

[0038] The model output unit is used to obtain a pre-trained depth estimation model based on multi-sensor fusion according to the updated network parameters of each encoder and the updated network parameters of each fusion encoder.

[0039] In a fourth aspect, an embodiment of the present application further provides a multi-sensor fusion apparatus, which is applied to a depth estimation model based on multi-sensor fusion obtained by the method described in the first aspect of the embodiment of the present application or any specific implementation manner of the first aspect, wherein the depth estimation model based on multi-sensor fusion includes an encoder set for each type of sensor, a decoder set for each type of sensor, and at least one fusion decoder for fusing inputs of different types of sensors;

[0040] The multi-sensor fusion device comprises:

[0041] a data acquisition unit, configured to acquire observation data from at least one sensor combination consisting of sensors of different types;

[0042] A shallow feature extraction unit, configured to input the observation data of each sensor in the sensor combination into its corresponding encoder to extract shallow features of each of the observation data;

[0043] a credibility judgment unit, configured to judge the credibility of each sensor based on the shallow features extracted by each sensor, and determine a credible sensor;

[0044] A depth map generation unit is configured to select a fusion decoder corresponding to the trusted sensor based on the trusted sensor, perform fusion processing on the multi-level feature data corresponding to each of the trusted sensors, and output the fusion result as a predicted depth map of the current scene.

[0045] In a fifth aspect, an embodiment of the present application further provides a computer device, including:

[0046] CPU, memory, input and output interfaces;

[0047] The memory is a transient storage memory or a persistent storage memory;

[0048] The central processing unit is configured to communicate with the memory and execute instruction operations in the memory to perform the method described in the first aspect, any specific implementation of the first aspect, the second aspect, or any specific implementation of the second aspect of the embodiment of the present application.

[0049] In a sixth aspect, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the method described in the first aspect of the embodiment of the present application, any specific implementation of the first aspect, the second aspect, or any specific implementation of the second aspect is executed.

[0050] It can be seen from the above technical solutions that the embodiments of the present application have the following advantages:

[0051] By introducing encoder and decoder decoupling training, as well as credibility judgment based on shallow features, the robustness of the depth estimation model based on multi-sensor fusion in complex environments is effectively improved; in the case of degradation, occlusion, noise or failure of sensor data, it is still possible to screen reliable sensors through the shallow features extracted from the observation data of each sensor, avoiding the interference of the observation data of unreliable sensors on the final fusion result, thereby achieving more stable and accurate depth estimation. In addition, the training method of this application can generate a pre-trained model with adaptability and generalization capabilities, which is convenient for rapid deployment in practical applications and can cope with the changing quality of sensor input. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on the provided drawings without paying any creative work.

[0053] Figure 1 A flowchart of a method for training a depth estimation model based on multi-sensor fusion provided in an embodiment of the present application;

[0054] Figure 2 A schematic diagram of a multi-sensor fusion method flow chart provided in an embodiment of the present application;

[0055] Figure 3 A schematic diagram of the structure of a training device for a depth estimation model based on multi-sensor fusion provided in an embodiment of the present application;

[0056] Figure 4 A schematic diagram of the computer device structure provided in an embodiment of the present application. DETAILED DESCRIPTION

[0057] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0058] Ranging is an essential function for animals and robots to perceive their environment and navigate autonomously. Within the animal kingdom, humans rely on their biological visual system to perceive distance; bats and dolphins locate their prey by emitting sound waves and receiving echoes; and snakes use heat signatures to estimate the distance of their prey. Similarly, in robotics, the key to improving ranging capabilities lies in accurately estimating the depth of a scene through data-driven learning methods using sensors such as monocular or binocular cameras, radar, lidar, or thermal imagers.

[0059] However, single-sensor depth estimation (SSDE) faces numerous challenges in practical applications. Depth information estimated by visual cameras is often limited by insufficient precision and scale ambiguity. While lidar (LiDAR) offers high measurement accuracy, its sampling points are extremely sparse, making it difficult to generate detailed depth maps. This results in poor performance in applications requiring high-precision depth information, such as 3D obstacle avoidance, dense simultaneous localization and mapping, small target detection and tracking, and navigation in complex environments. In contrast, multi-sensor fusion (MSF) technology can fully leverage the strengths of different sensor data modalities, significantly improving depth estimation accuracy. For example, a proposed multi-sensor depth estimation (MSDE) method can extract rich semantic features from RGB images and effectively interpolate them using precise lidar scan data. Extensive experiments have demonstrated that methods that fuse camera and lidar data far outperform those relying solely on a single sensor. Therefore, depth estimation techniques based on multi-sensor fusion exhibit great potential in practical applications such as autonomous driving and robotic navigation.

[0060] However, previous research has mostly focused on improving the accuracy of depth estimation, with insufficient attention paid to the frequent sensor interference issues in practical applications. This has made existing MSDE methods difficult to effectively deploy in real-world scenarios. This lack of robustness stems from several key factors: First, existing MSDE methods can only output accurate results when the input data distribution is consistent with the model training dataset. Once faced with out-of-distribution (OOD) degraded observations, model performance deteriorates dramatically. For example, when encountering data from scenarios different from the model training dataset or data containing anomalies or degradation (such as sensor failures and abnormal interference), the model cannot generalize, and the accuracy of the model output results drops significantly, forming a significant performance bottleneck. Second, the lack of diagnostic capabilities for sensor degradation and the inability to take effective countermeasures when sensor failures occur further exacerbate the system's vulnerability.

[0061] More specifically, with the increasing requirements for depth estimation accuracy, multi-sensor fusion MSF technology has become an important means to improve its performance. However, as mentioned above, the existing multi-sensor fusion MSF methods often show significant deficiencies in accuracy and robustness when facing sensor degradation. Although some common degradation problems, such as insufficient lighting and severe weather conditions, can be alleviated to a certain extent through relevant data collected in data-driven learning, existing methods are still unable to effectively deal with out-of-distribution OOD sensor degradation. To solve this problem, this application proposes a training method and multi-sensor fusion method for a depth estimation model based on Combinable and Separable Multi-Sensor Fusion (CSMSF), which aims to significantly enhance the robustness of depth estimation in various sensor degradation scenarios. The design of the CSMSF method is based on the following core principles:

[0062] 1) Composability: As the number of effective sensors increases, the performance of depth estimation should improve accordingly;

[0063] 2) Separability: The model can decouple degraded sensors to prevent them from interfering with model performance;

[0064] 3) Diagnosability: The system should have the ability to autonomously diagnose whether the data from each sensor is degraded, and promptly identify and eliminate the impact of degraded sensors.

[0065] In the case of a multi-sensor fusion-based depth estimation model trained using a multi-sensor fusion-based depth estimation model training method, the multi-sensor fusion method implemented using this model can effectively identify and reject degraded sensors, and autonomously select a reliable sensor combination for scene depth estimation. Experimental results show that the embodiments of this application demonstrate excellent robustness under various complex environmental conditions, fully demonstrating their effectiveness in addressing challenges related to sensor degradation and laying a solid foundation for the widespread promotion of depth estimation technology in practical applications.

[0066] The following is a further detailed description of the various embodiments of the present application in conjunction with the accompanying drawings.

[0067] The present invention provides a method for training a depth estimation model based on multi-sensor fusion. Figure 1 As shown, the method includes steps S101-S107.

[0068] S101: Build a depth estimation model based on multi-sensor fusion.

[0069] In the embodiments of this application, the core purpose of constructing a depth estimation model based on multi-sensor fusion is to achieve feature expression and fusion processing of observation data from different types of sensors. For ease of description, the "depth estimation model based on multi-sensor fusion" is referred to as the "model" below. The model includes three main structures:

[0070] First, an encoder is set up for each type of sensor to extract multi-level feature data from the observation data of this type of sensor;

[0071] Second, a decoder is set up for each type of sensor to restore the multi-level feature data extracted by the encoder of the sensor to its corresponding predicted depth map;

[0072] Third, at least one fusion decoder is set to process the fusion feature information formed after the observation data of two or more different types of sensors are extracted by the encoder, and output the fused predicted depth map.

[0073] In an embodiment of the present application, the structures of the encoder and decoder can be constructed using a convolutional neural network, and the fusion decoder can be set separately according to different sensor modal combinations, such as a dual-modal fusion decoder (such as a visual camera + lidar) or a tri-modal fusion decoder (such as a visual camera + lidar + infrared thermal sensor) to ensure that multimodal data can be flexibly fused.

[0074] S102: Input observation data of each type of sensor into its corresponding encoder and decoder to generate a depth map corresponding to each sensor.

[0075] The observation data of each type of sensor is input into the encoder and decoder paths corresponding to the sensor to complete the forward calculation of the single-modal prediction training phase for the single sensor. Specifically, for the i-th sensor (such as a visual camera, lidar or infrared thermal sensor), its observation data is first input into the encoder of the sensor type (denoted as g i ), the encoder of this type of sensor performs multi-level convolution feature extraction on the observation data and outputs multi-scale intermediate feature maps (also known as "multi-level feature data"). Subsequently, the multi-level feature data is passed to the decoder of the same type of sensor (denoted as f i ), the decoder performs upsampling or deconvolution operations on the multi-level feature data to reconstruct the depth map corresponding to the type of sensor that is consistent with the original image size (denoted as ). Step S102 of this application is mainly the initialization training stage of the single-modal sensor in the model training process, which provides an effective single-modal sensor feature basis for the subsequent multi-modal sensor fusion training stage.

[0076] S103: Based on the error between the depth map corresponding to each sensor and the true depth map under the same viewing angle, the encoder and decoder of the sensor are trained according to a preset loss function to minimize the error, and the updated network parameters of the encoder and the decoder are obtained.

[0077] To effectively train single-modal sensors, a supervisory signal is constructed based on the difference between the depth map output by the model for each sensor type and the ground-truth depth map corresponding to that sensor at the same viewing angle. This signal is then used to update the learnable network parameters within the model. These network parameters include at least the convolution weights and biases.

[0078] Specifically, a preset loss function can be used to measure the error between the predicted value and the true value of each pixel. By minimizing the preset loss function, the encoder g i and decoder f i Gradient backpropagation and parameter updates are performed on the learnable network parameters to optimize the sensor's representation capabilities. After training, the updated encoder and decoder network parameters for the individual sensor are obtained, serving as the basis for subsequent shared feature extraction for multimodal sensor fusion.

[0079] S104: For any sensor combination consisting of different types of sensors, input observation data of each sensor in the sensor combination into its corresponding encoder to obtain multi-level feature data.

[0080] In an embodiment of the present application, for any sensor combination composed of different types of sensors (such as a visual camera and a lidar, a visual camera and an infrared thermal imaging sensor, etc.), the observation data of each sensor in the sensor combination is input into the encoder corresponding to each sensor to obtain the feature expression of the observation data collected under different modal sensors. Since step S103 has completed the pre-training of each encoder, the network parameters of each sensor encoder in this step remain fixed and are only used as a feature extractor. After the observation data of each sensor is extracted by its encoder, a shallow or middle-level feature map (also known as "multi-level feature data") under the modality is obtained; the multi-level feature data includes the structure, texture, edge or heat information collected by the modal sensor in the current observation scene. Extracting high-quality features from the inputs of different modal sensors and providing alignable and diversified modal inputs for the subsequent fusion decoder can help the subsequent multimodal feature fusion processing.

[0081] S105: Inputting the multi-level feature data into a fusion decoder corresponding to the sensor combination to generate a depth map based on multi-sensor fusion.

[0082] The multi-level feature data of different modalities obtained in step S104 are used as input and passed to the fusion decoder corresponding to the sensor combination to generate a depth map based on multi-sensor fusion. The structure of the fusion decoder can be flexibly selected according to the modality type of the sensor combination. For example, for a combination of a visual camera and a lidar, a dual-modal fusion decoder f is used. vl , f vl That is, the decoder corresponding to the sensor combination composed of a visual camera and a lidar; for a sensor combination of three modalities, such as a sensor combination of a visual camera + lidar + infrared thermal sensor (thermal), the corresponding three-modal fusion decoder f is used. vlt After receiving the multi-level feature data of the multi-channel modal sensor, the fusion decoder realizes the effective fusion of the multi-modal feature data through feature splicing, weighted fusion, attention mechanism or cross convolution, and reconstructs the depth map based on multi-sensor fusion with the same size as the original input. ) integrates the complementary perception capabilities of different modal sensors to the scene, improving the accuracy and robustness of depth prediction.

[0083] S106: According to the error between the depth map based on multi-sensor fusion and the real depth map at the same viewing angle, the fusion decoder is trained according to the preset loss function to minimize the error, and the network parameters of the updated fusion decoder are obtained.

[0084] In step S106 of the present application, the depth map based on multi-sensor fusion and the real depth map under the same perspective need to be evaluated pixel by pixel, a supervisory signal is constructed, and the network parameters of the fusion decoder are optimized. Similar to single-modal training, a preset loss function can be selected to construct a robust fusion training objective function. By minimizing the loss function, only the network parameters of the corresponding fusion decoder are updated (the encoder network parameters remain unchanged), thereby improving the prediction performance of the fusion path under various modal combination inputs.

[0085] S107: Obtain a pre-trained depth estimation model based on multi-sensor fusion according to the updated network parameters of each encoder and the updated network parameters of each fusion encoder.

[0086] The network parameters of all encoders (from step S103) and the network parameters of all fusion decoders (from step S106) obtained through previous training are unified and integrated to form a pre-trained depth estimation model based on multi-sensor fusion. The pre-trained depth estimation model based on multi-sensor fusion has a modular architecture that can flexibly support input configurations of different sensor combinations, and can still make predictions by switching to the fusion path of the available sensor combination when the subsequent sensor input is degraded, occluded or missing. Ultimately, the model can be used as the basic model in the inference stage for online multimodal fusion depth estimation tasks, with good generalization and robustness, providing support for high-precision three-dimensional perception applications in a variety of practical scenarios.

[0087] In one embodiment, based on the error between the depth map corresponding to each sensor and the true depth map at the same viewing angle, the encoder and decoder of the sensor are trained according to a preset loss function to minimize the error, which is implemented by the following function:

[0088]

[0089] Among them, g i represents the encoder of the i-th sensor; f i represents the decoder of the i-th sensor; Indicates three sensor types, v, l, and t represent vision, lidar, and thermal imaging sensors respectively; K represents the total number of training samples involved in the training; k represents the sample index, indicating the kth training sample; represents the depth map corresponding to the i-th sensor generated by the k-th training sample through the i-th sensor encoder and the i-th sensor decoder; y k represents the true depth map of the kth training sample; represents a preset loss function that measures the error between the depth map corresponding to the i-th sensor of the k-th training sample at the same viewing angle and the true depth map.

[0090] Using the above function and then back-propagating to update the convolutional neural network parameters of each type of sensor path, including the convolution weights and convolution bias values ​​of the encoder and decoder, can ensure that each single-modal sensor path has independent feature learning and depth estimation capabilities, and provide a high-quality input foundation for subsequent multimodal fusion.

[0091] In one embodiment, based on the error between the depth map based on multi-sensor fusion and the real depth map at the same viewing angle, the fusion decoder is trained according to the preset loss function to minimize the error, which is specifically implemented by the following function:

[0092]

[0093] Among them, f j represents the fusion decoder for processing the j-th sensor combination; Represents different types of sensor combinations, vl, vt, lt, vlt respectively represent vl represents the sensor combination consisting of visual sensor and lidar sensor, vt represents the sensor combination consisting of visual sensor and thermal imaging sensor, lt represents the sensor combination consisting of lidar sensor and thermal imaging sensor, and vlt represents the sensor combination consisting of visual sensor, lidar sensor, and infrared sensor; K represents the total number of training samples participating in the training; k represents the sample index, which represents the kth training sample; represents the depth map based on multi-sensor fusion generated by the encoder corresponding to the j-th sensor combination and the fusion decoder corresponding to the j-th sensor combination for the k-th training sample; y k represents the true depth map of the kth training sample; represents a preset loss function that measures the error between the multi-sensor fusion-based depth map and the ground-truth depth map for the kth training sample at the same viewpoint. This function is used to train fusion decoders, enabling each fusion decoder to effectively integrate feature data extracted from multiple modal sensors and learn semantic complementarity and redundancy between modalities.

[0094] Furthermore, in one embodiment, the specific form of the preset loss function is:

[0095]

[0096] Where e represents a natural constant; K represents the total number of training samples involved in the training; k represents the sample index, which represents the kth training sample; s k Represents the uncertainty value given by the model at the pixel level for the prediction result of the kth training sample; represents the uncertainty weighting term, which is used to suppress the impact of high uncertainty pixels on the loss; Represents the depth map corresponding to the i-th sensor generated by the k-th training sample through the encoder and decoder of the i-th sensor Or the depth map based on multi-sensor fusion generated by the k-th training sample through the encoder and fusion decoder of the j-th sensor combination y k represents the true depth map of the kth training sample.

[0097] This function can be controlled by uncertainty values, dynamically reducing the model's sensitivity to abnormal pixels or degraded modal inputs during training.

[0098] Specifically, when the model predicts a depth value of a certain pixel with a high confidence level, it indicates that the prediction result is more certain. It can be understood that the lower the uncertainty, the corresponding s k →-∞, thus This means that the weight of the pixel prediction error will be increased in the loss function, prompting the model to focus on learning the information of these high-confidence pixels, thereby improving depth estimation accuracy and prediction accuracy.

[0099] On the contrary, when the model predicts a pixel with high uncertainty, for example, because the area is affected by occlusion, blur or modal degradation, s k →+∞, at this time The model will automatically reduce the impact of the pixel error in the overall loss. Even if some pixel predictions are inaccurate, it will not cause significant interference to the entire model. This can effectively suppress the damage of outliers to the training process and prevent the model from overfitting or mislearning.

[0100] The s added later k It is actually a regularization term, which is used to suppress the model from increasing uncertainty aimlessly to avoid loss. It can be regarded as a self-balancing mechanism to prevent the model from raising s k make The loss is reduced to a very low level, close to zero, which has the effect of limiting the model's "lazy behavior" and ensuring that the uncertainty output is meaningful. Therefore, the entire loss function is essentially a variation of the negative log-likelihood loss, which ensures that depth estimation models based on multi-sensor fusion can both express uncertainty and avoid abusing it.

[0101] Accordingly, an embodiment of the present application further provides a multi-sensor fusion method, which is applied to a depth estimation model based on multi-sensor fusion obtained by the method of the first aspect of the embodiment of the present application or any specific implementation manner of the first aspect, wherein the depth estimation model based on multi-sensor fusion includes an encoder set for each type of sensor, a decoder set for each type of sensor, and at least one fusion decoder for fusing inputs of different types of sensors;

[0102] Please refer to Figure 2 , the multi-sensor fusion method includes steps S210-S240:

[0103] S210: Acquire observation data of at least one sensor combination composed of different types of sensors;

[0104] In step S210, the system first obtains observation data of a sensor combination composed of multiple different types of sensors from the target scene to be processed, such as images, depth maps or heat maps, etc. The observation data is the original input for subsequent feature extraction and fusion processing.

[0105] Specifically, let the target domain be D = {X v ,X l ,X t ,Y}, where X v ,X l ,X t They represent the observation datasets of visual cameras, lidars, and infrared thermal imaging sensors respectively; Y represents the real depth map label datasets corresponding to visual cameras, lidars, and infrared thermal imaging sensors.

[0106] In each prediction task, the system selects any non-empty sensor combination input from the target domain, which is recorded as and satisfy

[0107] In a specific implementation, the sensor combination may be:

[0108] 1. Single modal input (such as only x v );

[0109] 2. Dual-mode combination input (such as x v ,x l , i.e. visual camera + lidar);

[0110] 3. Trimodal fusion input (i.e. x v ,x l ,x t all at once).

[0111] The input combination is constrained by the symbol Given, it means selecting a non-empty subset from all possible modal combinations as the actual input. Here, P{·} is a power set, and after removing the empty set, it represents all legal modal combinations.

[0112] Each selected modal input will serve as the starting basis for multimodal encoding, credibility judgment and fusion processing in subsequent steps, thus ensuring that the system is adaptable to different modal combination inputs.

[0113] S220: Inputting the observation data of each sensor in the sensor combination into its corresponding encoder to extract shallow features of each observation data;

[0114] The observation data from each sensor type obtained in step S210 is fed into its corresponding encoder to extract the corresponding shallow-layer feature information. The encoder is a pre-trained neural network model from the model training phase. Its structure may include multiple convolutional layers, normalization layers, and activation function layers. In this step, the encoder parameters remain fixed and are used only for forward propagation to extract features.

[0115] Shallow features, on the other hand, typically contain local structural information such as edges, textures, contours, or thermal intensity, and possess higher spatial resolution and modal discrimination. While intuitively, deep features are more task-specific, shallow features capture more of the original signal. Therefore, shallow features are more sensitive to perturbations than deep features, making them suitable for use as discriminant evidence in the credibility assessment stage.

[0116] S230: performing a credibility judgment on each sensor based on the shallow features extracted by each sensor, and determining a credible sensor;

[0117] Based on the shallow features of each type of sensor extracted in step S220, the reliability of the current observation data of the sensor is evaluated. The reliability judgment is used to automatically filter out modal inputs that are blocked, degraded, abnormal or invalid before fusion, and improve the robustness of multi-level feature data fusion processing. This reliability judgment corresponds to Figure 2 The TA module (Trustworthiness Assessment) in the algorithm is implemented by statistical modeling and feature comparison of the extracted shallow features to determine whether the modal features are within the credible distribution range learned during training. If they are beyond the credible distribution range, they are regarded as degradation or failure signals, corresponding to Figure 2 The Drop! operation in the [Drop!] does not participate in the subsequent multi-level feature data fusion processing. Step S230 is primarily used to address disturbances caused by out-of-distribution sensor degradation and is a key technical step in improving the accuracy of the predicted depth map.

[0118] S240: Based on the trusted sensor, select a fusion decoder corresponding to the trusted sensor, perform fusion processing on the multi-level feature data corresponding to each trusted sensor, and output the fusion result as the predicted depth map of the current scene.

[0119] Specifically, for any (x v ,x l ,x t ,y)∈D, the model aims to infer an accurate depth map by the following formula:

[0120]

[0121] in, F(x) represents the selection and calling of the corresponding fusion decoder function according to the input combination x to realize the joint decoding of multimodal features and the output of depth map. Function F is a set of decoders (including decoders corresponding to a single sensor and fusion decoders corresponding to sensor combinations), which is defined as:

[0122] F={f v , f l , f t , f vl , f vt , f lt , f vlt}

[0123] Among them, f v , f l , f t Single-modal decoders corresponding to visual cameras, lidar, and infrared thermal imaging sensors respectively;

[0124] f vl , f vt , f lt are the fusion decoders corresponding to the combination of any two modal sensors;

[0125] f vlt The fusion decoder corresponding to the combination of three modal sensors.

[0126] At runtime, the system automatically selects the decoder / fusion decoder corresponding to the currently selected credible modal combination x from F, and inputs the multi-level feature data corresponding to these modalities into the decoder / fusion decoder for splicing, weighting or attention mechanism for feature integration, and then the corresponding fusion decoder performs feature decoding, and the final output is the predicted depth map That is the estimation result of the current scene based on multi-sensor fusion.

[0127] In one embodiment, step 230 specifically includes:

[0128] S231: Calculate the average value of the shallow features extracted by each sensor;

[0129] The shallow feature maps extracted by each sensor are statistically processed to calculate their mean vector in the feature space. Specifically, for each shallow feature map extracted by a sensor, an average operation is performed on its channel dimension or spatial dimension to obtain a mean vector used to represent the feature distribution. This mean vector can well reflect the overall structural distribution characteristics of the current input image in that modality. Because shallow features have strong modality independence, their mean information can be used as a basis for determining whether the current observation data is within the training set distribution.

[0130] S232: Calculate the cosine similarity between the average value of the shallow features of each sensor and the average value of the shallow features in the standard sample set;

[0131] The cosine similarity between the average shallow feature value of each sensor obtained in step S231 and the shallow feature mean value of the preset standard sample set is further calculated. The standard sample set is the set of shallow feature means of credible observation samples collected during the model training phase. Cosine similarity measures the degree of consistency between the directions of two vectors, with a value range of [-1, 1]. A higher similarity indicates that the current observation sample is closer to the training set distribution, i.e., more reliable. This method does not rely on label information, but only on feature statistics to determine distribution differences, resulting in efficient and robust computation.

[0132] The specific calculation method of the shallow feature mean of the preset standard sample set is as follows:

[0133]

[0134] Where k represents the number of training samples; To use the shallow features extracted by the i-th encoder for the k-th credible observation sample, is the average feature of the shallow features extracted by the i-th encoder on all credible observation samples. In practical applications, let It is the shallow feature mean of the preset standard sample set, which serves as the reference standard for normal sensor signals.

[0135] The specific cosine similarity calculation method can refer to the following formula:

[0136]

[0137] θ represents cosine similarity; cos() represents cosine similarity calculation, z i Characterize the shallow features extracted from the i-th sensor.

[0138] S233: Determine the sensor whose cosine similarity is higher than a preset similarity threshold as a trusted sensor.

[0139] Based on the cosine similarity value obtained in the previous stage, it is compared with a preset similarity threshold (such as 0.8, which can be limited according to the experimental environment and specific requirements). If the cosine similarity value between the shallow feature mean and the standard feature mean corresponding to a certain sensor is higher than the preset similarity threshold, its observation data is considered to be a credible modal input at the current moment, and the sensor corresponding to the data will be marked as a credible sensor; otherwise, the observation data input of the modal sensor will be temporarily eliminated and will not participate in the current round of multimodal fusion process, thereby achieving dynamic screening at the sensor input level and significantly reducing the impact of sensor degradation on the final depth estimation result.

[0140] In order to implement the training method of the depth estimation model based on multi-sensor fusion of the embodiment of the present application, the embodiment of the present application also provides a training device for the depth estimation model based on multi-sensor fusion, which includes:

[0141] A model construction unit, configured to construct a depth estimation model based on multi-sensor fusion, wherein the depth estimation model based on multi-sensor fusion includes an encoder set for each type of sensor, a decoder set for each type of sensor, and at least one fusion decoder for fusing inputs of different types of sensors;

[0142] A depth map generation unit, configured to input observation data of each type of sensor into its corresponding encoder and decoder to generate a depth map corresponding to each sensor;

[0143] A single-modal training unit is configured to train the encoder and decoder of the sensor according to a preset loss function based on the error between the depth map corresponding to each sensor and the true depth map at the same viewing angle to minimize the error, thereby obtaining updated network parameters of the encoder and decoder; the network parameters include at least convolution weight values ​​and convolution bias values;

[0144] A feature extraction unit is configured to input observation data of each sensor in any sensor combination composed of different types of sensors into its corresponding encoder to obtain multi-level feature data;

[0145] The depth map generating unit is further configured to input the multi-level feature data into a fusion decoder corresponding to the sensor combination to generate a depth map based on multi-sensor fusion;

[0146] a multimodal training unit, configured to train the fusion decoder according to the preset loss function to minimize the error between the depth map based on multi-sensor fusion and the real depth map at the same viewing angle, and obtain network parameters of the updated fusion decoder;

[0147] The model output unit is used to obtain a pre-trained depth estimation model based on multi-sensor fusion according to the updated network parameters of each encoder and the updated network parameters of each fusion encoder.

[0148] In one embodiment, based on the error between the depth map corresponding to each sensor and the true depth map at the same viewing angle, the encoder and decoder of the sensor are trained according to a preset loss function to minimize the error, which is implemented by the following function:

[0149]

[0150] Among them, g i represents the encoder of the i-th sensor; f i represents the decoder of the i-th sensor; Indicates three sensor types, v, l, and t represent vision, lidar, and thermal imaging sensors respectively; K represents the total number of training samples involved in the training; k represents the sample index, indicating the kth training sample; represents the depth map corresponding to the i-th sensor generated by the k-th training sample through the i-th sensor encoder and the i-th sensor decoder; y k represents the true depth map of the kth training sample; Represents the preset loss function that measures the error between the depth map corresponding to the i-th sensor and the true depth map under the same training sample.

[0151] In one embodiment, the training of the fusion decoder according to the preset loss function to minimize the error between the depth map based on multi-sensor fusion and the real depth map at the same viewing angle is specifically implemented by the following function:

[0152]

[0153] Among them, f j represents the fusion decoder for processing the j-th sensor combination; Represents different types of sensor combinations, vl, vt, lt, vlt respectively represent vl represents the sensor combination consisting of visual sensor and lidar sensor, vt represents the sensor combination consisting of visual sensor and thermal imaging sensor, lt represents the sensor combination consisting of lidar sensor and thermal imaging sensor, and vlt represents the sensor combination consisting of visual sensor, lidar sensor, and infrared sensor; K represents the total number of training samples participating in the training; k represents the sample index, which represents the kth training sample; represents the depth map based on multi-sensor fusion generated by the encoder corresponding to the j-th sensor combination and the fusion decoder corresponding to the j-th sensor combination for the k-th training sample; yk represents the true depth map of the kth training sample; represents a preset loss function for measuring the error between the depth map based on multi-sensor fusion and the true depth map under the same training sample.

[0154] In one embodiment, the specific form of the preset loss function is:

[0155]

[0156] Where e represents a natural constant; K represents the total number of training samples involved in the training; k represents the sample index, which represents the kth training sample; s k Represents the uncertainty value given by the prediction result of the k-th training sample at the pixel level based on the depth estimation model based on multi-sensor fusion; represents the uncertainty weighting term, which is used to suppress the impact of high uncertainty pixels on the loss; Represents the depth map corresponding to the i-th sensor generated by the k-th training sample through the encoder and decoder of the i-th sensor Or the depth map based on multi-sensor fusion generated by the k-th training sample through the encoder and fusion decoder of the j-th sensor combination y k represents the true depth map of the kth training sample.

[0157] To implement the multi-sensor fusion method of the embodiment of the present application, the embodiment of the present application further provides a multi-sensor fusion device, which is applied to a multi-sensor fusion-based depth estimation model obtained by the method described in the first aspect of the embodiment of the present application or any specific implementation manner of the first aspect, wherein the multi-sensor fusion-based depth estimation model includes an encoder set for each type of sensor, a decoder set for each type of sensor, and at least one fusion decoder for fusing inputs of different types of sensors;

[0158] The device includes:

[0159] a data acquisition unit, configured to acquire observation data from at least one sensor combination consisting of sensors of different types;

[0160] A shallow feature extraction unit, configured to input the observation data of each sensor in the sensor combination into its corresponding encoder to extract shallow features of each of the observation data;

[0161] a credibility judgment unit, configured to judge the credibility of each sensor based on the shallow features extracted by each sensor, and determine a credible sensor;

[0162] The depth map generation unit is used to select a fusion decoder corresponding to the trusted sensor based on the trusted sensor, fuse the multi-level feature data corresponding to each of the trusted sensors, and output the fusion result as the predicted depth map of the current scene.

[0163] In one embodiment, the credibility judgment unit is specifically configured to:

[0164] Calculating an average value of the shallow features extracted by each of the sensors;

[0165] Calculating the cosine similarity between the average value of the shallow features of each sensor and the average value of the shallow features in the standard sample set;

[0166] Sensors whose cosine similarity is higher than a preset similarity threshold are determined as trustworthy sensors.

[0167] It should be noted that: the above embodiment provides a training device for a depth estimation model based on multi-sensor fusion. When training a depth estimation model based on multi-sensor fusion, only the division of the above-mentioned program modules is used as an example. In actual applications, the above-mentioned processing can be assigned to different program modules as needed, that is, the internal structure of the device can be divided into different program modules to complete all or part of the processing described above. In addition, the training device for a depth estimation model based on multi-sensor fusion provided in the above embodiment and the training method embodiment of a depth estimation model based on multi-sensor fusion belong to the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0168] Similarly, the above embodiments provide a multi-sensor fusion device, using the aforementioned division of program modules as an example to illustrate multi-sensor fusion. In actual applications, the aforementioned processing can be assigned to different program modules as needed, i.e., the internal structure of the device can be divided into different program modules to complete all or part of the aforementioned processing. Furthermore, the multi-sensor fusion device and the multi-sensor fusion method embodiments provided in the above embodiments share the same concept. The specific implementation process is detailed in the method embodiments and will not be further elaborated here.

[0169] Based on the hardware implementation of the above program modules, and in order to implement a training method for a depth estimation model based on multi-sensor fusion and a multi-sensor fusion method provided in an embodiment of the present application, an embodiment of the present application further provides a computer device, such as Figure 4 As shown, the computer device 400 includes:

[0170] CPU 401, memory 402 and input / output interface 403;

[0171] The memory 402 is a temporary storage memory or a permanent storage memory;

[0172] The central processing unit 401 is configured to communicate with the memory 402 and execute instruction operations in the memory 402 to perform the training method of the depth estimation model based on multi-sensor fusion described in the first aspect of the embodiment of the present application or any specific implementation of the first aspect, and implement the multi-sensor fusion method described in the second aspect of the embodiment of the present application or any specific implementation of the second aspect.

[0173] Of course, in actual application, the various components in the computer device 400 are coupled together through the bus system 404. It can be understood that the bus system 404 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 404 also includes a power bus, a control bus, and a status signal bus. However, for the sake of clarity, Figure 4 Various buses are labeled as bus system 404 .

[0174] The memory 402 in the embodiment of the present application is used to store various types of data to support the operation of the computer device 400. Examples of such data include: any computer program used to operate on the computer device 400.

[0175] It is understandable that when the processor in the computer device described above executes the computer program, it can also implement the functions of the various units in the corresponding device embodiments described above, which will not be repeated here. For example, the computer program can be divided into one or more modules / units, one or more modules / units are stored in the memory and executed by the processor to complete the various embodiments of the present application. One or more modules / units can be a series of computer program instruction segments that can perform specific functions, and the instruction segments are used to describe the execution process of the computer program in the computer device. For example, the computer program can be divided into the various units in the above-mentioned computer device, and each unit can implement the specific functions described in the above-mentioned corresponding computer device.

[0176] A computer device may be a desktop computer, laptop, PDA, cloud server, or other computing device. A computer device may include, but is not limited to, a processor and memory. Those skilled in the art will appreciate that processors and memory are merely examples of computer devices and do not constitute a limitation of computer devices. Computer devices may include more or fewer components, or combinations of certain components, or different components. For example, a computer device may also include input / output devices, network access devices, buses, and the like.

[0177] The processor can be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor. The processor is the control center of the computer device and connects the various parts of the entire computer device using various interfaces and lines.

[0178] The memory can be used to store computer programs and / or modules. The processor implements various functions of the computer device by running or executing the computer programs and / or modules stored in the memory, and calling the data stored in the memory. The memory can mainly include a program storage area and a data storage area. The program storage area can store an operating system, at least one application required for a function, etc.; the data storage area can store data created based on the use of the terminal, etc. In addition, the memory can include high-speed random access memory and non-volatile memory, such as a hard disk, internal memory, a plug-in hard disk, a smart memory card (SMC, Smart Media Card), a secure digital (SD, Secure Digital) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state storage device.

[0179] An embodiment of the present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, it executes the method described in the first aspect of the embodiment of the present application, any specific implementation of the first aspect, the second aspect, or any specific implementation of the second aspect.

[0180] An embodiment of the present application also provides a computer program product having a computer program / instruction stored thereon, which, when executed by a processor, is used to implement the method described in the first aspect, any specific implementation of the first aspect, the second aspect, or any specific implementation of the second aspect of the embodiment of the present application.

[0181] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0182] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0183] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0184] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0185] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

Claims

1. A training method for a depth estimation model based on multi-sensor fusion, characterized in that: include: Constructing a depth estimation model based on multi-sensor fusion, the depth estimation model based on multi-sensor fusion comprising an encoder set for each type of sensor, a decoder set for each type of sensor, and at least one fusion decoder for fusing inputs of different types of sensors; Input the observation data of each type of sensor into its corresponding encoder and decoder to generate the depth map corresponding to each sensor; Based on the error between the depth map corresponding to each sensor and the true depth map at the same viewing angle, the encoder and decoder of the sensor are trained according to a preset loss function to minimize the error, thereby obtaining updated network parameters of the encoder and decoder; The network parameters include at least: a convolution weight value and a convolution bias value; For any sensor combination consisting of different types of sensors, the observation data of each sensor in the sensor combination is input into its corresponding encoder to obtain multi-level feature data; Inputting the multi-level feature data into a fusion decoder corresponding to the sensor combination to generate a depth map based on multi-sensor fusion; According to the error between the depth map based on multi-sensor fusion and the real depth map at the same viewing angle, the fusion decoder is trained according to the preset loss function to minimize the error, thereby obtaining the network parameters of the updated fusion decoder; A pre-trained depth estimation model based on multi-sensor fusion is obtained according to the updated network parameters of each encoder and the updated network parameters of each fusion encoder.

2. The method according to claim 1, characterized in that Based on the error between the depth map corresponding to each sensor and the true depth map at the same viewing angle, the encoder and decoder of the sensor are trained according to a preset loss function to minimize the error, which is achieved by the following function: Among them, g i represents the encoder of the i-th sensor; f i represents the decoder of the i-th sensor; Indicates three sensor types, v, l, and t represent vision, lidar, and thermal imaging sensors respectively; K represents the total number of training samples involved in the training; k represents the sample index, indicating the kth training sample; represents the depth map corresponding to the i-th sensor generated by the k-th training sample through the i-th sensor encoder and the i-th sensor decoder; y k represents the true depth map of the kth training sample; Represents a preset loss function that measures the error between the depth map corresponding to the i-th sensor and its corresponding true depth map.

3. The method according to claim 1, characterized in that The error between the depth map based on multi-sensor fusion and the real depth map under the same viewing angle is trained on the fusion decoder according to the preset loss function to minimize the error, which is specifically implemented by the following function: Among them, f j represents the fusion decoder for processing the j-th sensor combination; Represents different types of sensor combinations, vl, vt, lt, vlt respectively represent vl represents the sensor combination consisting of visual sensor and lidar sensor, vt represents the sensor combination consisting of visual sensor and thermal imaging sensor, lt represents the sensor combination consisting of lidar sensor and thermal imaging sensor, and vlt represents the sensor combination consisting of visual sensor, lidar sensor, and infrared sensor; K represents the total number of training samples participating in the training; k represents the sample index, which represents the kth training sample; represents the depth map based on multi-sensor fusion generated by the encoder corresponding to the j-th sensor combination and the fusion decoder corresponding to the j-th sensor combination for the k-th training sample; y k represents the true depth map of the kth training sample; represents a preset loss function for measuring the error between the depth map based on multi-sensor fusion and its corresponding real depth map.

4. The method according to any one of claims 1 to 3, characterized in that The specific form of the preset loss function is: Where e represents a natural constant; K represents the total number of training samples involved in the training; k represents the sample index, which represents the kth training sample; s k Represents the uncertainty value given by the prediction result of the k-th training sample at the pixel level based on the depth estimation model based on multi-sensor fusion; represents the uncertainty weighting term, which is used to suppress the impact of high uncertainty pixels on the loss; Represents the depth map corresponding to the i-th sensor generated by the k-th training sample through the encoder and decoder of the i-th sensor Or the depth map based on multi-sensor fusion generated by the k-th training sample through the encoder and fusion decoder of the j-th sensor combination y k represents the true depth map of the kth training sample.

5. A multi-sensor fusion method, characterized in that: Applied to a depth estimation model based on multi-sensor fusion obtained by the depth estimation model training method based on multi-sensor fusion according to any one of claims 1 to 4, the depth estimation model based on multi-sensor fusion comprising an encoder set for each type of sensor, a decoder set for each type of sensor, and at least one fusion decoder for fusing inputs of different types of sensors; The multi-sensor fusion method comprises: Acquiring observation data of at least one sensor combination composed of different types of sensors; Inputting the observation data of each sensor in the sensor combination into its corresponding encoder respectively to extract shallow features of each of the observation data; Performing a credibility judgment on each sensor based on shallow features extracted from each sensor to determine a credible sensor; Based on the trusted sensor, a fusion decoder corresponding to the trusted sensor is selected, multi-level feature data corresponding to each of the trusted sensors is fused, and a fusion result is output as a predicted depth map of the current scene.

6. The method according to claim 5, characterized in that The performing credibility judgment on each sensor based on the shallow features extracted from each sensor to determine a credible sensor includes: Calculating an average value of the shallow features extracted by each of the sensors; Calculating the cosine similarity between the average value of the shallow features of each sensor and the average value of the shallow features in the standard sample set; Sensors whose cosine similarity is higher than a preset similarity threshold are determined as trustworthy sensors.

7. A training device for a depth estimation model based on multi-sensor fusion, characterized in that: include: A model construction unit, configured to construct a depth estimation model based on multi-sensor fusion, wherein the depth estimation model based on multi-sensor fusion includes an encoder set for each type of sensor, a decoder set for each type of sensor, and at least one fusion decoder for fusing inputs of different types of sensors; A depth map generation unit, configured to input observation data of each type of sensor into its corresponding encoder and decoder to generate a depth map corresponding to each sensor; a single-modal training unit, configured to train the encoder and decoder of the sensor according to a preset loss function based on the error between the depth map corresponding to each sensor and the true depth map at the same viewing angle to minimize the error, thereby obtaining updated network parameters of the encoder and decoder; The network parameters include at least: a convolution weight value and a convolution bias value; A feature extraction unit is configured to input observation data of each sensor in any sensor combination composed of different types of sensors into its corresponding encoder to obtain multi-level feature data; The depth map generating unit is further configured to input the multi-level feature data into a fusion decoder corresponding to the sensor combination to generate a depth map based on multi-sensor fusion; a multimodal training unit, configured to train the fusion decoder according to the preset loss function to minimize the error between the depth map based on multi-sensor fusion and the real depth map at the same viewing angle, and obtain network parameters of the updated fusion decoder; The model output unit is used to obtain a pre-trained depth estimation model based on multi-sensor fusion according to the updated network parameters of each encoder and the updated network parameters of each fusion encoder.

8. A multi-sensor fusion device, characterized in that: Applied to a depth estimation model based on multi-sensor fusion obtained by the depth estimation model training method based on multi-sensor fusion according to any one of claims 1 to 4, the depth estimation model based on multi-sensor fusion comprising an encoder set for each type of sensor, a decoder set for each type of sensor, and at least one fusion decoder for fusing inputs of different types of sensors; The multi-sensor fusion device comprises: a data acquisition unit, configured to acquire observation data from at least one sensor combination consisting of sensors of different types; A shallow feature extraction unit, configured to input the observation data of each sensor in the sensor combination into its corresponding encoder to extract shallow features of each of the observation data; a credibility judgment unit, configured to judge the credibility of each sensor based on the shallow features extracted by each sensor, and determine a credible sensor; A depth map generation unit is configured to select a fusion decoder corresponding to the trusted sensor based on the trusted sensor, perform fusion processing on the multi-level feature data corresponding to each of the trusted sensors, and output the fusion result as a predicted depth map of the current scene.

9. A computer device, characterized in that: include: CPU, memory and input / output interfaces; The memory is a transient storage memory or a persistent storage memory; The central processing unit is configured to communicate with the memory and execute instructions in the memory to perform the method according to any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is performed.