Millimeter wave radar and camera fusion dense depth estimation method based on self-supervision

By using a self-supervised millimeter wave radar and camera fusion dense depth estimation method in harsh environments, combined with uncertainty and depth code optimization technology, the sparse point cloud and noise problems of 4D millimeter wave radar in harsh environments is solved, and efficient and accurate depth estimation is achieved.

CN120070532AActive Publication Date: 2025-05-30WUHAN UNIV
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510061441.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-15
Publication Date
2025-05-30
Estimated Expiration
2045-01-15

AI Technical Summary

Technical Problem

The prior art is difficult to accurately identify obstacles under harsh environmental conditions. The point clouds acquired by 4D millimeter wave radar are sparse, have obvious noise and narrow detection field, resulting in incomplete point cloud information and serious noise interference when the fusion image is used to achieve dense depth estimation, making it difficult to accurately reconstruct scene details.

Method used

The dense depth estimation method based on self-supervised millimeter wave radar and camera fusion is adopted. By acquiring the camera image sequence and the 4D millimeter wave radar point cloud sequence, the trained dense depth estimation model is input after alignment. The uncertainty branch and dense depth estimation branch are combined with the depth code and convolutional features to perform dense depth estimation, and the depth code is optimized to improve the estimation accuracy through the digital-analog-combined depth code optimization method.

Benefits of technology

Efficient and accurate depth estimation under harsh environmental conditions is achieved, and the dependence on dense depth annotation is reduced through self-supervised learning. The depth code is optimized by multi-frame common viewing information, which improves the accuracy and efficiency of dense depth estimation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120070532A_ABST
    Figure CN120070532A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a millimeter wave radar and camera fusion dense depth estimation method based on self-supervision, and relates to the technical field of scene perception. The method comprises the following steps: inputting an image sequence of a camera and a point cloud sequence of a 4D millimeter wave radar which are aligned in time into a trained dense depth estimation model, and processing an image and a point cloud pair through an uncertainty branch to obtain convolution features and uncertainty; and obtaining a depth code corresponding to the current frame, inputting the depth code and the convolution feature into the dense depth estimation branch, decoding to obtain the dense depth of each pixel unit on the image of the corresponding frame, and obtaining a dense depth estimation graph. According to the method, the model is trained through a self-supervision method, so that the training cost caused by manual marking of dense depth values is avoided; through constructing a geometric optimization problem between a plurality of images and point cloud pairs, a depth code is adjusted in a prediction process, so that good depth estimation precision can be achieved in a complex application scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of scene perception technology, and particularly to a self-supervised millimeter-wave radar and camera fusion dense depth estimation. Background Art

[0002] Depth estimation can provide key three-dimensional perception capabilities for positioning and navigation systems. In application scenarios such as autonomous driving and drone navigation, depth estimation can help the system accurately understand the three-dimensional structure of the environment, including information such as obstacles and roads, so as to achieve precise positioning and path planning, and improve the safety and reliability of the system.

[0003] Currently, depth estimation can be performed by methods based on image data or the fusion of images and lidar. However, environmental factors such as rain, snow, fog, dust, and other harsh conditions may affect the performance of traditional sensors. In harsh environmental conditions, the sensors may be severely affected and unable to accurately identify obstacles. Since the millimeter-wave wavelength is relatively long, 4D millimeter-wave radar can penetrate smoke and dust and is not affected by light. Therefore, the method of fusing 4D millimeter-wave radar point cloud and image data is more suitable for depth estimation under all-weather conditions. However, compared with cameras and lidar, the point cloud obtained by 4D millimeter-wave radar is sparse, has obvious noise, and has a narrow detection field of view (generally within 120°). These disadvantages make 4D millimeter-wave radar face the problems of incomplete point cloud information and noise interference when fusing images to achieve dense depth estimation, resulting in difficulty in accurately reconstructing the details of the scene. The current dense depth estimation methods for 4D millimeter-wave radar camera fusion have high computational consumption in the inference network and relatively slow inference speed, which may limit the performance of real-time applications. Moreover, through deep learning models for calculation, it often depends on real dense depth values, and real dense depth values require a large amount of manual annotation, with high labor costs and difficult acquisition. Related technologies are usually limited to the processing of single-frame radar point cloud and camera image pairs when fusing 4D millimeter-wave radar point cloud and image data for dense depth estimation, and do not effectively utilize temporal co-visible information. Therefore, there is still a lack of an efficient and accurate depth estimation method suitable for complex scenarios. Summary of the Invention

[0004] An embodiment of this application provides a self-supervised millimeter-wave radar and camera fusion dense depth estimation method to solve the defects in the above-mentioned related technologies. The technical solution is as follows:

[0005] In a first aspect, an embodiment of this application provides a self-supervised millimeter-wave radar and camera fusion dense depth estimation method, including:

[0006] Obtain an image sequence of a camera and a point cloud sequence of a 4D millimeter-wave radar;

[0007] Align the image sequence and the point cloud sequence, and output the aligned image sequence and point cloud sequence;

[0008] Input the aligned image sequence and point cloud sequence into the uncertainty branch of the trained dense depth estimation model. Through the uncertainty branch, perform feature encoding on each frame of the image in the aligned image sequence and the corresponding frame of the point cloud in the aligned point cloud sequence to obtain convolutional features, and decode the convolutional features to obtain the uncertainty of each pixel unit in the corresponding frame of the image;

[0009] Obtain the depth code corresponding to the current frame, input the depth code and the convolutional features into the dense depth estimation branch of the dense depth estimation model, and through the dense depth estimation branch, combine the depth code and the convolutional features to perform decoding to obtain the dense depth of each pixel unit on the corresponding frame of the image, and output the dense depth estimation map.

[0010] In an alternative scheme of the first aspect, obtaining the depth code corresponding to the current frame includes:

[0011] If the current frame is the first frame in the image sequence, obtain a preset initial depth code;

[0012] Otherwise, optimize the depth code, including:

[0013] Obtain the uncertainties, depth codes, and dense depth estimation maps corresponding to multiple key frame images before the current frame in the image sequence;

[0014] Construct an objective function for optimizing the depth code based on the differences between the current frame and each key frame image before the current frame; the objective function includes the photometric error, reprojection error of matching feature points, and reprojection error of the millimeter-wave radar of each key frame image before the current frame;

[0015] Through a least squares solver, combine the objective function and the prior distribution constraint of the depth code to calculate the optimized camera poses of the current frame and each key frame image before it and the optimized depth code.

[0016] In an alternative scheme of the first aspect, the constructing of the objective function for optimizing the depth code includes:

[0017] Project the current frame onto each key frame image before the current frame respectively, calculate the light intensity difference on the corresponding pixel units, and construct the photometric error;

[0018] After projecting the current frame onto each key frame image before the current frame respectively, at least a pair of matching feature points are selected between two frames, the difference in pixel coordinates of the matching feature points is calculated, and the reprojection error of the matching feature points is constructed.

[0019] Project the sparse points of the radar of the current frame onto the dense depth estimation map corresponding to each key frame image before the current frame, compare the depth value of the pixel unit where the radar sparse point is located with the depth value of the corresponding pixel unit on the dense depth estimation map, and construct the reprojection error of the millimeter-wave radar.

[0020] In an alternative scheme of the first aspect, the aligning the image sequence and the point cloud sequence and inputting the aligned image sequence and point cloud sequence into the trained dense depth estimation model includes:

[0021] Obtain the timestamps of the image sequence and the point cloud sequence, and align the time sequences of the image sequence and the point cloud sequence according to the timestamps;

[0022] Obtain the external parameters of the camera and the 4D millimeter-wave radar, perform external parameter transformation on the point cloud of the corresponding frame in the point cloud sequence according to the external parameters, and obtain the point cloud depth map in the camera coordinate system to spatially align the image sequence and the point cloud sequence;

[0023] Crop the point cloud sequence after the time sequence alignment and the spatial alignment, and unify the sizes of each frame of image and the point cloud of the corresponding frame;

[0024] Input the cropped image sequence and point cloud sequence into the trained dense depth estimation model;

[0025] Among them, each pixel unit on each frame of point cloud can correspond to the pixel unit on the corresponding frame of image. Note that some pixel units of the point cloud map have no valid depth and can be marked with Inf or 0.

[0026] In an alternative scheme of the first aspect, the dense depth estimation model is trained through the following steps, including:

[0027] Train the teacher network model, specifically including:

[0028] Use the sample image sequence of the camera and the sample point cloud sequence of the 4D millimeter-wave radar as training samples, use the sparse depth map of each frame of point cloud in the sample point cloud sequence as the supervision sample of the teacher network model, and train the teacher network model so that the teacher network model learns the relationship between the depth information and the image features on each pixel unit of each frame of sample image in the sample image sequence, and output the trained teacher network model;

[0029] Train the student network model, specifically including:

[0030] Use the sample image sequence of the camera and the sample point cloud sequence of the 4D millimeter-wave radar as training samples, and use the teacher dense depth estimation map of each frame of sample image predicted by the trained teacher network model as the supervision sample for the student network model to train the student network model;

[0031] Output the trained student network model as the dense depth estimation model.

[0032] In an alternative scheme of the first aspect, the loss function for training the teacher network model includes:

[0033] Obtain the teacher dense depth estimation map of each frame of sample image predicted by the teacher network model, obtain the sparse depth map of the point cloud corresponding to the sample image in the sample point cloud sequence, and construct a sparse depth loss function based on the difference between the effective depth value on each pixel unit in the sparse depth map and the depth value of the corresponding pixel unit in the teacher dense depth estimation map of each frame of sample image;

[0034] Based on the preset pose transformation and the dense depth estimation map of the current frame of sample image, transform the next adjacent frame of sample image of the current frame to the current frame to obtain the transformed image of the current frame of sample image; construct a photometric consistency loss function based on the transformed image and the current frame of sample image;

[0035] According to the sparse depth loss and the photometric consistency loss, train the teacher network model and output the trained teacher network model.

[0036] In an alternative scheme of the first aspect, training the student network model and outputting the trained student network model as the dense depth estimation model includes:

[0037] Obtain the sample depth code obtained by the dense depth estimation branch of the student network model for feature encoding each frame of sample image in the sample image sequence, and construct a KL divergence loss function based on the gap between the sample depth code corresponding to each frame of sample image and the standard normal distribution;

[0038] Obtain the student dense depth estimation map output by the dense depth estimation branch of the student network model, and construct a dense depth loss function for all pixel units based on the difference between the depth value on each pixel unit in the student dense depth estimation map output by the student network model and the depth value of the corresponding pixel unit in the teacher dense depth estimation map;

[0039] Obtain the sparse depth map of the point cloud corresponding to the student's dense depth estimation map in the sample point cloud sequence, and construct a sparse depth loss function based on the difference between the depth values of each pixel unit in the sparse depth map and the depth values of the corresponding pixel units in the student's dense depth estimation map;

[0040] Based on the KL divergence loss, the dense depth prediction loss, and the sparse depth loss, train the student network model, and output the converged student network model as the dense depth estimation model.

[0041] In a second aspect, an embodiment of the present application further provides a self-supervised millimeter-wave radar and camera fusion dense depth estimation device, including:

[0042] A data module for obtaining an image sequence of a camera and a point cloud sequence of a 4D millimeter-wave radar;

[0043] The data module is further configured to align the image sequence and the point cloud sequence, output the aligned image sequence and point cloud sequence, and input the aligned image sequence and point cloud sequence into the uncertainty branch of the trained dense depth estimation model;

[0044] An uncertainty calculation module, through the uncertainty branch in the dense depth estimation model, performs feature encoding on each frame of the image in the aligned image sequence and the point cloud of the corresponding frame in the aligned point cloud sequence to obtain convolution features; then, by decoding these convolution features, calculate the uncertainty of each pixel in the corresponding image frame;

[0045] A dense depth estimation module for obtaining the depth code corresponding to the current frame, inputting the depth code and the convolution features into the dense depth estimation branch of the dense depth estimation model, and combining the depth code and the convolution features through the dense depth estimation branch to decode the dense depth of each pixel unit on the image of the corresponding frame, and output a dense depth estimation map.

[0046] A digital-analog joint depth code optimization module for using multi-frame co-visibility information to construct a least squares loss function to optimize the depth codes of historical key frames and the current frame, so as to provide more accurate depth codes for the dense depth estimation module. This loss function includes the photometric error between the current frame and multiple previous key frame images, the reprojection error of matching feature points, the reprojection error of the millimeter-wave radar, and also includes the Gaussian prior factor of the depth code.

[0047] In a third aspect, an embodiment of the present application further provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the method provided in the first aspect or any implementation manner of the first aspect of the embodiments of the present application.

[0048] In a fourth aspect, the present application also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method provided in the first aspect or any implementation manner of the first aspect of the embodiments of the present application is implemented.

[0049] The beneficial effects brought by the technical solutions provided in the embodiments of the present application at least include:

[0050] A self-supervised millimeter-wave radar and camera fusion dense depth estimation method provided in an embodiment of the present application can combine an image sequence and a point cloud sequence to achieve dense depth prediction that fuses sparse depth and image information. Aiming at the problem that existing depth estimation methods require dense depth supervision, a self-supervised method of using input data is used to train a dense depth estimation model. The teacher network model is trained with the sparse depth information collected by a 4D millimeter-wave radar in a limited number of training samples, and the student network model is trained with the predicted dense depth output by the teacher network model, thereby avoiding the defect that dense depth needs to be manually annotated in related technologies. On the other hand, based on a digital-analog joint depth code least squares optimization method, spatio-temporal co-visibility constraints can be formed by using multi-frame information in time series to optimize the estimation of the depth code, so as to obtain a more accurate dense depth estimation of the current frame. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are the embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0052] Figure 1 is a schematic flowchart of a self-supervised millimeter-wave radar and camera fusion dense depth estimation method according to an embodiment of the present application;

[0053] Figure 2 is a schematic structural diagram of a dense depth estimation student network model according to an embodiment of the present application;

[0054] Figure 3 is a schematic structural diagram of a dense network estimation teacher network model according to an embodiment of the present application;

[0055] Figure 4 is a self-supervised training flowchart of a teacher network according to an embodiment of the present application;

[0056] Figure 5 is a self-supervised training flowchart of a student network according to an embodiment of the present application;

[0057] Figure 6It is a schematic structural diagram of a self-supervised millimeter-wave radar and camera fusion dense depth estimation device provided by an embodiment of the present application;

[0058] Figure 7 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0059] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions in the present application will be clearly and completely described below with reference to the accompanying drawings in the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art without making creative efforts based on the embodiments in the present application belong to the scope of protection of the present application.

[0060] The terms "including" and "having" and any variations thereof in the specification and claims of the present application and the above-mentioned accompanying drawings are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or modules is not limited to the listed steps or modules, but optionally further includes steps or modules not listed, or optionally further includes other steps or modules inherent to these processes, methods, products, or devices.

[0061] Related depth network-based technologies usually need to rely on true dense depth values to train a dense depth estimation model. However, obtaining the true dense depth value corresponding to each pixel unit on an image often requires a large amount of manual work, and it is difficult to generate enough supervised training samples in a short time, thus limiting the realization of large-scale data training. Although reducing the number of samples or the number of pixel units of the image can reduce the workload, this will limit the training process of the dense depth estimation model, and the prediction accuracy of the trained dense depth estimation model is difficult to reach the expectation.

[0062] Furthermore, in the training process of the current dense depth estimation model, generally only the information of the current frame is considered, that is, only the case of a pair of radar point clouds and camera frames is limited, and the correlation information between each frame and adjacent frames in a multi-frame sequence is not considered. In complex scenarios, such as when an autonomous vehicle enters a tunnel from the outside, the light intensity changes greatly. If only the information of the current frame is considered and the correlation information between each frame and adjacent frames is not considered, it will result in a large difference in the output between the current frame and adjacent frames of the dense depth estimation model. Therefore, the current dense depth estimation method has poor robustness in complex application scenarios, and the accuracy of the output result is limited.

[0063] The present application will be described in detail below with reference to specific embodiments.

[0064] Next, in combination withFigure 1 and Figure 2 , introduce the self-supervised millimeter-wave radar and camera fusion dense depth estimation method provided by the embodiments of the present application. For details, please refer to Figure 2 , Figure 2 shows a schematic structural diagram of the dense depth estimation student network model provided by the embodiments of the present application. As Figure 1 shown, the method includes the following steps:

[0065] S101, obtain the image sequence of the camera and the point cloud sequence of the 4D millimeter-wave radar;

[0066] S102, align the image sequence and the point cloud sequence, and output the aligned image sequence and point cloud sequence;

[0067] S103, input the aligned image sequence and point cloud sequence into the uncertainty branch of the trained dense depth estimation model, and perform feature encoding on each frame of the image in the aligned image sequence and the point cloud corresponding to the corresponding frame in the aligned point cloud sequence through the uncertainty branch to obtain convolutional features, and decode the convolutional features to obtain the uncertainty of each pixel unit in the corresponding frame of the image;

[0068] S104, obtain the depth code corresponding to the current frame, input the depth code and the convolutional features into the dense depth estimation branch of the dense depth estimation model, and combine the depth code and the convolutional features through the dense depth estimation branch to decode the dense depth of each pixel unit on the corresponding frame of the image, and output the dense depth estimation map.

[0069] S105, use the multi-frame co-visibility information to construct a least squares loss function to optimize the depth codes of the historical key frames and the current frame, so as to provide more accurate depth codes for subsequent frames.

[0070] In the embodiments of the present application, the image sequence of the camera obtained in S101 includes multiple frames of images taken by the camera of the area to be estimated according to the preset shooting time, and the point cloud sequence of the 4D millimeter-wave radar includes multiple frames of point clouds output by the 4D millimeter-wave radar after detecting the same area to be estimated according to the preset detection frequency.

[0071] It can be understood that during the data acquisition process, since the parameters such as the acquisition start time, end time, and acquisition frequency of the camera and the 4D millimeter-wave radar may be different, therefore, before predicting the dense depth by fusing the image and the corresponding radar point cloud frame, it is necessary to perform the steps of S102 to perform preprocessing of time and space consistency on the sequences collected by the two sensors.

[0072] In the embodiment of the present application, in S102, the timestamps of the image sequence and the point cloud sequence can be obtained, and the time sequences of the image sequence and the point cloud sequence can be aligned according to the timestamps, so as to ensure that each frame of image and each frame of point cloud are aligned at the acquisition moment.

[0073] In the embodiment of the present application, in some cases, the acquisition frequencies of the image sequence and the point cloud sequence may be different, and there may be unaligned images and point clouds at some moments. Motion compensation calculation can be performed on the missing point cloud data to fill the missing images or point clouds at the corresponding moments.

[0074] In the embodiment of the present application, in S102, the external parameters of the camera and the 4D millimeter-wave radar can be obtained, and the point clouds of the corresponding frames in the point cloud sequence can be subjected to external parameter transformation according to the external parameters, projected into the camera coordinate system, and then cropped, so that the image sequence and the point cloud sequence are spatially aligned, and a spatially aligned image sequence and a point cloud map sequence can be obtained, so that each pixel unit on the projection map of each frame of point cloud can correspond to the pixel unit on the corresponding frame of image. Since the point cloud of the millimeter-wave radar is sparse, some depth values on the point cloud projection map may be invalid values.

[0075] Exemplarily, 224×224 images of uniform size and the point cloud of the millimeter-wave radar can be obtained, where the point cloud of the millimeter-wave radar is stored in the format of a sparse depth map.

[0076] It should be noted that the point cloud of the millimeter-wave radar gives the sparse depth map of the corresponding frame, that is, the depth detection values of some pixel units are given. For those without depth detection values at the positions of the corresponding pixel units in the sparse depth map, the depth values can be set to 0 or Inf.

[0077] Furthermore, the aligned image sequence and point cloud sequence are input into the trained dense depth estimation model to perform the steps of S103, which specifically include:

[0078] Exemplarily, the trained dense depth estimation model is as Figure 2 shown. Taking the image type in the input image sequence as an RGB image as an example, the RGB image sequence and the point cloud sequence are input into the uncertainty branch of the dense depth estimation model.

[0079] Specifically, a frame of the RGB image sequence and the point cloud of the same frame in the point cloud sequence are input into the uncertainty branch of the dense depth estimation model in pairs. First, the aligned data is processed through an encoding-decoding network. Among them, the encoder uses the lightweight convolutional network MobileNet to perform feature encoding on the image frame from the camera and the sparse depth data from the millimeter-wave radar (i.e., the point cloud corresponding to the image frame), extracts the convolutional features, and then transfers the information in the encoder to the decoder through skip connections to determine the uncertainty of each pixel unit in the image of the corresponding frame.

[0080] Understandably, the uncertainty is used to characterize the accuracy of the depth value on each pixel unit in the image. The higher the uncertainty, the lower the accuracy of the depth value corresponding to the pixel unit.

[0081] Furthermore, the RGB image sequence and the convolutional features extracted by the uncertainty branch are input into the dense depth estimation branch of the dense depth estimation model to perform the steps of S104, specifically including:

[0082] According to the depth code corresponding to the image of the current frame, combined with the convolutional features of the image of the current frame extracted by the uncertainty branch, the decoder of the dense depth estimation branch decodes to obtain the depth of each pixel unit on the image of the corresponding frame, and outputs a dense depth estimation map.

[0083] Understandably, the dense depth estimation model learns the mapping relationship and its probability distribution related to the depth information of the input image and point cloud during the training process, so as to predict the depth value on each pixel unit.

[0084] It should be noted that the dense depth estimation model converts the input camera image and radar data into a low-dimensional latent variable representation, that is, a depth code, through the encoder part, such as through the lightweight convolutional network MobileNet. The depth code can be regarded as a compression or abstraction of the input data, containing features related to the depth information, that is, the dense depth is a function of the depth code.

[0085] It should be noted that for the first frame image in the image sequence to be processed, a preset initial depth code is selected. The initial depth code can be obtained during the training process of the dense depth estimation model or can be set according to the actual scenario. The embodiments of the present application do not limit this.

[0086] In order to further improve the accuracy of depth estimation in the embodiments of the present application, the temporal frame information can be used to further optimize the depth code, so as to achieve dense depth estimation driven by both data and model.

[0087] Thus, after S104, the following steps can also be performed:

[0088] S105. Use multi-frame co-visibility information to construct a least-squares loss function to optimize the depth codes of historical key frames and the current frame.

[0089] Specifically, for images after the first frame in the image sequence to be processed, the depth codes can be optimized based on the differences between the current frame image and the images of some previous key frames in terms of depth codes, dense depth estimation maps, etc. Specifically, it includes:

[0090] Obtain the uncertainty, depth code, and dense depth estimation map corresponding to each key frame image before the current frame in the image sequence;

[0091] Construct an objective function for optimizing the depth code and camera pose based on the differences between the current frame and each key frame image before the current frame;

[0092] Through a least-squares solver, combined with the objective function and the prior distribution constraint of the depth code, calculate the optimized camera poses of the current frame and each key frame image before, and the optimized depth codes.

[0093] Specifically, the geometry of the scene can be expressed by the depth map D of the camera key frames m set G, G = {D m-k , D m-k+1 ,.., D m}, and the key frames can be understood as several image frames containing scene information extracted from the image sequence. The depth map corresponding to the key frame can be generated by a dense depth estimation model and the depth code to be optimized. Thus, the variables of the multi-frame dense depth optimization problem only include the poses of k + 1 camera key frames and the related depth code c i .

[0094] Among them, the key frames can be selected according to some rules. For example, if the camera moves or rotates beyond a certain limit, then a key frame is selected. For example, if the translation distance of the current frame relative to the previous key frame exceeds 0.3 m, or the rotation angle exceeds 20°, then the current frame is selected as the new key frame.

[0095] Among them, the objective function for optimizing the depth code includes the following items:

[0096] The photometric error, reprojection error of matching feature points, and reprojection error of the millimeter-wave radar of each key frame image before the current frame, and also include the Gaussian prior factor of the depth code.

[0097] Specifically, the photometric error includes:

[0098] Project the current frame onto each key-frame image before the current frame respectively, calculate the light intensity difference on the corresponding pixel units, and construct an objective function for photometric error, using the formula:

[0099]

[0100] where the photometric reprojection error constrains the light intensity difference between the pixel I i of the current frame image and the pixel I j of the target image transformed to frame i, and Ω i is the set of pixels on the i-th frame image;

[0101]

[0102] where ω ji is the transformation function that transforms the pixel point x in the i-th frame to the pixel point in the j-th frame, and π and π -1 are the camera projection and inverse projection functions respectively, is the dense depth estimation map decoded from the depth code c i , T ji is the pose transformation from frame i to frame j, and c i is the depth code of the i-th frame. Note that D i (x) also depends on the image I i and the convolutional features of the millimeter-wave radar sparse depth map .

[0103] Specifically, the reprojection error of the matching feature points includes:

[0104] Project the current frame onto each key-frame image before the current frame respectively, select at least one pair of matching feature points between the two frames, calculate the difference in the pixel coordinates of the matching feature points, and construct an objective function for the reprojection error of the matching feature points. The reprojection error measures the difference between the feature points observed in the j-th frame image and the pixel coordinates predicted from the i-th frame image. Use the formula:

[0105]

[0106] where M ij is the set of pairs of matching feature points in the i-th and j-th frame images, and y is the coordinate of the matching feature point in the previous frame j image;

[0107] It should be noted that the matching feature points can be detected and determined by the BRISK algorithm.

[0108] Specifically, the reprojection error of the millimeter-wave radar includes:

[0109] For the millimeter-wave radar depth information frame i and the frame j containing the network-predicted depth map D j For the pixel x in frame i, after projection, it corresponds to the pixel in frame j The sparse depth reprojection error is the difference between the millimeter-wave radar observed depth and the network-predicted depth of two corresponding pixels:

[0110] Project the sparse radar points of the current frame onto the dense depth estimation map corresponding to each key frame image before the current frame, compare the depth values of the pixel units where the sparse radar points are located with the depth values of the corresponding pixel units on the dense depth estimation map, and construct an objective function for the reprojection error of the millimeter-wave radar, and apply the formula:

[0111]

[0112] where is the sparse depth reprojection error, and [x] z represents the z component of the x vector.

[0113] Specifically, the Gaussian prior factor of the depth code ensures that the depth code distribution of each millimeter-wave radar point cloud and image frame pair i is close to a Gaussian distribution:

[0114]

[0115] Based on the above constraints for optimization, the optimized camera pose and depth code of the current frame and several historical key frames can be calculated.

[0116] The decoder of the dense depth estimation model in the embodiments of the present application can use the updated depth code to more accurately predict the dense depth of the current frame.

[0117] In the embodiments of the present application, the dense depth estimation model can be trained by setting a teacher network model and a student network model, specifically including:

[0118] Train the teacher network model, using the sample image sequence of the camera and the sample point cloud sequence of the 4D millimeter-wave radar as training samples, using the sample point cloud sequence as the supervision sample, training the teacher network model, and outputting the trained teacher network model;

[0119] Train the student network model, using the sample image sequence of the camera and the sample point cloud sequence of the 4D millimeter-wave radar as training samples, using the teacher dense depth estimation map output by the trained teacher network model as the supervision sample, training the student network model, and outputting the trained student network model as the dense depth estimation model.

[0120] Specifically, the teacher network model is used to predict the dense depth as the dense depth supervision signal for training the student network. The teacher network model is such asFigure 3 As shown, the teacher network adopts an encoder-decoder structure. The encoder consists of residual blocks of ResNet-34, gradually reducing the spatial resolution of the feature map through the convolution process, while the decoder gradually increases the spatial resolution of the feature map through transposed convolutional layers. The input sparse depth and RGB images respectively obtain shallow features through initial convolution, and then are concatenated into a single tensor as the input of the ResNet-34 residual blocks. The output of the encoding layer is passed to the corresponding decoding layer through skip connections for fusion to prevent the loss of shallow information.

[0121] Specifically, a self-supervised framework as Figure 4 shown is used to train the teacher network model. This framework takes the aligned camera sample image sequence and the 4D millimeter-wave radar sample point cloud sequence as input training samples, and takes the point cloud sequence as the supervised sample, that is, uses the sparse depth information provided by the point cloud as the supervised information to train the teacher network model, so that the teacher network model learns the relationship between the depth information and the image features on each pixel unit in each frame of the sample image in the sample image sequence, and outputs the trained teacher network model.

[0122] Specifically, the loss function used in the self-supervised training of the teacher network model consists of three parts: sparse depth prediction loss, edge-aware smoothing loss, and adjacent RGB image reprojection photometric consistency loss.

[0123] Specifically, obtain the teacher dense depth estimation map of each frame of the sample image predicted by the teacher network model, obtain the sparse depth map of the point cloud corresponding to the sample image in the sample point cloud sequence, and construct a sparse depth loss function based on the difference between the depth values on each pixel unit in the sparse depth map and the depth values of the corresponding pixel units in the teacher dense depth estimation map:

[0124]

[0125] where represents considering only depth-valid points, is the L2 norm.

[0126] To smooth the generated inverse depth map, the edge-aware smoothing loss is defined as:

[0127]

[0128] where represents the inverse depth normalized by the mean to avoid the shrinkage of the estimated depth.

[0129] Specifically, based on the preset pose transformation T 1→2 and the dense depth estimation map pred of the sample image of the current frame 11 Transform the sample image of the next adjacent frame 2 of the current frame 1 to the current frame. Specifically, given the intrinsic matrix K of the camera, for any pixel p in the current frame 1 1 has a corresponding projection in the next adjacent frame 2:

[0130] p 2 = K T 1→2 pred 1 (p 1 ) K -1 p 1 ;

[0131] Therefore, by performing bilinear interpolation using the four adjacent pixels of p 2 , an RGB image corresponding to the current frame 1 can be synthesized to obtain the transformed image of the current frame 1:

[0132] warped 1 (p 1 ) = bilinear(RGB 2 (K T 1→2 pred 1 (p 1 ) K -1 p 1 ));

[0133] Construct a photometric consistency loss function based on the transformed image and the sample image of the current frame 1. When the environment remains static and the occlusion caused by the perspective change is limited, the image warped 1 obtained by inverse transformation from the next adjacent frame 2 is similar to the current frame 1, and the photometric consistency loss is:

[0134]

[0135] where S is the scale factor for all pixels, (·) (s) indicates that the image is scaled by the factor s. It should be noted that only for points without depth supervision, that is, only for pixel units {d = 0} without sparse depth information, the photometric loss is considered.

[0136] Specifically, the model can be trained according to the sparse depth loss, smooth loss, and photometric consistency loss to obtain an optimized teacher network model.

[0137] In the embodiments of the present application, using the teacher dense depth estimation map of each frame of sample image predicted by the trained teacher network model as a supervision sample, the student network model is supervised and trained, including:

[0138] Specifically, as Figure 5As shown, the sample image sequence of the camera and the sample point cloud sequence of the 4D millimeter-wave radar are used as training sample inputs for the self-supervised depth reconstruction network, which is the teacher network model, to obtain the teacher dense depth estimation map predicted by the teacher network model. The sample image sequence of the camera, the sample point cloud sequence of the 4D millimeter-wave radar, and the teacher dense depth estimation map are used as training sample inputs for the student network model, and a training loss function is constructed based on the output of the student network model and the output of the teacher network model, including KL divergence loss, dense depth prediction loss, and sparse depth prediction loss.

[0139] In the embodiment of the present application, the sample depth codes obtained by encoding the features of each frame of the sample image sequence by the dense depth estimation branch of the student network model can be obtained, and the KL divergence loss function is constructed based on the gap between each sample depth code corresponding to each frame of the sample image and the standard normal distribution. The KL divergence is used to narrow the gap between the depth code distribution and the standard normal distribution, and the formula is applied:

[0140]

[0141] where n is the dimension of the depth code, μ and σ represent the mean and variance of the depth code, μ j and σ j represent the j-th element in the two vectors.

[0142] In the embodiment of the present application, the uncertainty output by the uncertainty branch of the student network model can be obtained, and the student dense depth estimation map output by the dense depth estimation branch of the student network model can be obtained. Based on the uncertainty and the difference between the depth value on each pixel unit of the teacher dense depth estimation map and the depth value on the corresponding pixel unit of the student dense depth estimation map, the dense depth prediction loss function for all pixel units is constructed, and the formula is applied:

[0143]

[0144] where x represents a pixel in the output feature map, D(·) is the value of the teacher dense depth estimation map, is the logarithm of the inverse depth value predicted from the student network model, is the logarithm of the depth uncertainty.

[0145] It should be noted that during training, the loss value of each batch will be averaged, and the predicted depth value will also be thresholded to prevent infinite logarithmic values.

[0146] In the embodiment of the present application, a sparse depth map of the point cloud corresponding to the student dense depth estimation map in the sample point cloud sequence is obtained, and a sparse depth loss function is constructed based on the difference between the depth value on each pixel unit in the sparse depth map and the depth value of the corresponding pixel unit in the student dense depth estimation map. The formula is applied:

[0147]

[0148] where D R is the sparse depth value of the millimeter-wave radar point cloud, is the predicted depth value output by the student network model at the same pixel unit position.

[0149] Furthermore, based on the KL divergence loss function, the dense depth prediction loss function, and the sparse depth loss function, the student network model is trained by an optimization method such as Adam, and its output is used as the dense depth estimation model.

[0150] The following is the device embodiment of the present application, which can be used to execute the method embodiment of the present application. For the details not disclosed in the device embodiment of the present application, please refer to the method embodiment of the present application.

[0151] Next, please refer to Figure 6 , which is a schematic structural diagram of a self-supervised millimeter-wave radar and camera fusion dense depth estimation device provided by an exemplary embodiment of the present application. This device can be implemented as all or part of a terminal through software, hardware, or a combination of both, and can also be integrated as an independent module on a server. The self-supervised millimeter-wave radar and camera fusion dense depth estimation device 60 in the embodiment of the present application includes a data module 601, an uncertainty calculation module 602, a dense depth estimation module 603, and a depth code optimization module 604, where:

[0152] The data module 601 is used to obtain the image sequence of the camera and the point cloud sequence of the 4D millimeter-wave radar;

[0153] The data module 601 is further used to align the image sequence and the point cloud sequence, output the aligned image sequence and point cloud sequence, and input the aligned image sequence and point cloud sequence into the uncertainty branch of the trained dense depth estimation model;

[0154] The uncertainty calculation module 602 performs feature encoding on each frame of the image in the aligned image sequence and the point cloud of the corresponding frame in the aligned point cloud sequence through the uncertainty branch of the dense depth estimation model to obtain convolution features, and decodes the convolution features to obtain the uncertainty of each pixel unit in the corresponding frame of the image;

[0155] The dense depth estimation module 603 is used to obtain the depth code corresponding to the current frame, input the depth code and the convolutional features into the dense depth estimation branch of the dense depth estimation model, combine the depth code and the convolutional features through the dense depth estimation branch, decode to obtain the dense depth of each pixel unit on the image of the corresponding frame, and output a dense depth estimation map.

[0156] The depth code optimization module 604 is used to utilize the multi-frame co-visibility information to construct a least squares loss function to optimize the depth codes of the historical key frames and the current frame, so as to provide a more accurate depth code for the dense depth estimation module.

[0157] It should be noted that when the device 60 provided in the above embodiment executes the method for dense depth estimation by fusing a millimeter-wave radar and a camera based on self-supervision, only the above division of each functional module is used for illustration. In practical applications, the above functions can be assigned to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the device provided in the above embodiment and the embodiment of the method for dense depth estimation by fusing a millimeter-wave radar and a camera based on self-supervision belong to the same concept. The implementation process is shown in detail in the method embodiment and will not be elaborated here.

[0158] The embodiment of the present application also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the steps of the method in any of the above embodiments.

[0159] Please refer to Figure 7 , which is a structural block diagram of an electronic device provided by the embodiment of the present application.

[0160] As Figure 7 shown, the electronic device 700 includes a processor 701 and a memory 702.

[0161] In the embodiment of the present application, the processor 701 is the control center of the computer system, which can be the processor of a physical machine or the processor of a virtual machine. The processor 701 can include one or more processing cores, such as a 4-core processor or an 8-core processor. The processor 701 can be implemented in at least one of the following hardware forms: DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array).

[0162] The processor 701 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state.

[0163] The memory 702 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 702 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In the embodiments of the present application, the non-transitory computer-readable storage medium in the memory 702 is used to store at least one instruction, and the at least one instruction is used to be executed by the processor 701 to implement the method in the embodiments of the present application.

[0164] In the embodiments of the present application, the electronic device 700 further includes: a peripheral device interface 703 and at least one peripheral device 704. The processor 701, the memory 702, and the peripheral device interface 703 may be connected through a bus or signal lines. Each peripheral device 704 may be connected to the peripheral device interface 703 through a bus, signal lines, or a circuit board. Specifically, the peripheral device 704 includes: a display screen, a camera, and a millimeter-wave radar. The peripheral device interface 703 may be used to connect at least one peripheral device related to I / O (Input / Output) to the processor 701 and the memory 702.

[0165] In the embodiments of the present application, the processor 701, the memory 702, and the peripheral device interface 703 are integrated on the same chip or circuit board; in some other embodiments of the present application, any one or two of the processor 701, the memory 702, and the peripheral device interface 703 may be implemented on a separate chip or circuit board. The embodiments of the present application do not make specific limitations on this.

[0166] The block diagram of the electronic device shown in the embodiments of the present application does not constitute a limitation on the electronic device 700. The electronic device 700 may include more or fewer components than shown in the figure, or combine certain components, or adopt a different component layout.

[0167] The embodiments of the present application also provide a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the steps of the method in any of the foregoing embodiments are implemented. Among them, the computer-readable storage medium may include, but is not limited to, any type of disk, including floppy disks, optical disks, DVDs, CD-ROMs, microdrives, and magneto-optical disks, ROMs, RAMs, EPROMs, EEPROMs, DRAMs, VRAMs, flash memory devices, magnetic cards or optical cards, nanosystems (including molecular memory ICs), or any type of medium or device suitable for storing instructions and / or data.

[0168] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solution, or the part that contributes to the related technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disks, optical disks, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0169] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A self-supervised millimeter-wave radar and camera fusion dense depth estimation method, characterized in that: include: Get the camera's image sequence and the 4D millimeter-wave radar's point cloud sequence; Searching for image frames and point cloud frames that are close in time for the image sequence and the point cloud sequence, and transforming the millimeter-wave radar point cloud into an image coordinate system based on external parameters to obtain a sequence of aligned image and point cloud frame pairs; Input the aligned image sequence and point cloud sequence into the uncertainty branch of the trained dense depth estimation model, perform feature encoding on each frame of the aligned image sequence and the point cloud of the corresponding frame in the aligned point cloud sequence through the uncertainty branch to obtain convolution features, and obtain the uncertainty of each pixel unit in the image of the corresponding frame by decoding the convolution features; Obtain the depth code corresponding to the current frame, input the depth code and the decoded convolutional features into the dense depth estimation branch of the dense depth estimation model, decode to obtain the dense depth of each pixel unit on the image of the corresponding frame, and output a dense depth estimation map.

2. The method for dense depth estimation based on self-supervision of millimeter-wave radar and camera fusion according to claim 1, characterized in that: The obtaining of the depth code corresponding to the current frame includes: If the current frame is the first frame in the image sequence, obtaining a preset initial depth code; Otherwise, the deep code is optimized, including: Obtaining uncertainty, depth code, and dense depth estimation map corresponding to a plurality of key frame images before a current frame in the image sequence; Constructing an objective function for optimizing the depth code based on the difference between the current frame and each key frame image before the current frame; the objective function includes a photometric error of each key frame image before the current frame, a reprojection error of matching feature points, and a reprojection error of a millimeter wave radar; The optimized camera poses of the current frame and each previous key frame image and the optimized depth code are calculated by combining the objective function and the prior distribution constraints of the depth code with a least squares solver.

3. The method for dense depth estimation based on self-supervision of millimeter-wave radar and camera fusion according to claim 1, characterized in that: The constructing of the objective function for optimizing the deep code comprises: The current frame is projected onto each key frame image before the current frame, and the light intensity difference on the corresponding pixel unit is calculated to construct a photometric error; After projecting the current frame onto each key frame image before the current frame, at least one pair of matching feature points between two frames is selected, and the difference of the pixel coordinates of the matching feature points is calculated to construct a reprojection error of the matching feature points; The radar sparse points of the current frame are projected onto a dense depth estimation map corresponding to each key frame image before the current frame, and the depth values ​​of the pixel units where the radar sparse points are located are compared with the depth values ​​of the corresponding pixel units on the dense depth estimation map to construct a reprojection error of the millimeter-wave radar.

4. The method for dense depth estimation based on self-supervision of millimeter-wave radar and camera fusion according to claim 1, characterized in that: Finding image frames and point cloud frames that are close in time for the image sequence and the point cloud sequence, transforming the point cloud into an image coordinate system, and then inputting the sequence of aligned image and point cloud pairs into a trained dense depth estimation model, including: Acquire timestamps of the image sequence and the point cloud sequence, and align the time sequence of the image sequence and the point cloud sequence according to the timestamps; Acquire external parameters of the camera and the 4D millimeter-wave radar, perform external parameter transformation on the point cloud in the point cloud sequence according to the external parameters, and project the point cloud into a camera coordinate system, so as to spatially align the image sequence and the point cloud sequence; Cropping the point cloud projection images after the temporal alignment and the spatial alignment to unify the sizes of each frame image and the point cloud of the corresponding frame; Inputting the cropped image sequence and point cloud sequence into the trained dense depth estimation model; Among them, each pixel unit on each frame of point cloud can correspond to a pixel unit on the corresponding frame image.

5. The method for dense depth estimation based on self-supervision of millimeter-wave radar and camera fusion according to claim 1, characterized in that: The dense depth estimation model is trained by the following steps, including: Training the teacher network model includes: The sample image sequence of the camera and the sample point cloud sequence of the 4D millimeter wave radar are used as training samples, and the sparse depth map of each frame of the point cloud in the sample point cloud sequence is used as a supervision sample of the teacher network model, and the teacher network model is trained so that the teacher network model learns the relationship between the depth information and the image features of each pixel unit in each frame of the sample image in the sample image sequence, and outputs the trained teacher network model; Training the student network model includes: The sample image sequence of the camera and the sample point cloud sequence of the 4D millimeter wave radar are used as training samples, and the teacher dense depth estimation map of each frame of the sample image predicted by the trained teacher network model is used as the supervision sample of the student network model to train the student network model; The trained student network model is output as the dense depth estimation model.

6. The method for dense depth estimation based on self-supervision of millimeter-wave radar and camera fusion according to claim 5, characterized in that: The steps of training the teacher network model include: Obtain a teacher dense depth estimation map of each frame of sample image predicted by the teacher network model, obtain a sparse depth map of the point cloud corresponding to the sample image in the sample point cloud sequence, and construct a sparse depth loss function based on the difference between the depth value of each valid pixel unit in the sparse depth map and the depth value of the corresponding pixel unit in the teacher dense depth estimation map; Based on a preset pose transformation and a dense depth estimation map of a sample image of a current frame, a sample image of a next adjacent frame of the current frame is transformed to a current frame to obtain a transformed image of the adjacent frame; and a photometric consistency loss function is constructed based on the transformed image and the sample image of the current frame; The teacher network model is trained by minimizing the sparse depth loss and the photometric consistency loss, and the trained teacher network model is output.

7. The method for dense depth estimation based on self-supervision of millimeter-wave radar and camera fusion according to claim 5, characterized in that: Training the student network model and outputting the trained student network model as the dense depth estimation model includes: Obtain a sample depth code obtained by performing feature encoding on each frame of the sample image in the sample image sequence by the dense depth estimation branch of the student network model, and construct a KL divergence loss function based on the gap between the sample depth code corresponding to each frame of the sample image and a standard normal distribution; Obtain the uncertainty of the uncertainty branch output of the student network model, obtain the student dense depth estimation map output by the dense depth estimation branch of the student network model, and construct a dense depth loss function for all pixel units based on the uncertainty output by the student network model and the difference between the depth value on each pixel unit on the student dense depth estimation map output by the student network model and the depth value on the corresponding pixel unit on the teacher dense depth estimation map; Obtain a sparse depth map of the point cloud corresponding to the student dense depth estimation map in the sample point cloud sequence, and construct a sparse depth loss function based on the difference between the depth value of each pixel unit in the sparse depth map and the depth value of the corresponding pixel unit in the student dense depth estimation map; The student network model is trained by minimizing the KL divergence loss, the dense depth loss, and the sparse depth loss, and the converged student network model is output as the dense depth estimation model.

8. A dense depth estimation device based on self-supervision of millimeter-wave radar and camera fusion, characterized in that: include: Data module, used to obtain the camera's image sequence and the 4D millimeter-wave radar's point cloud sequence; The data module is also used to align the image sequence and the point cloud sequence, output the aligned image sequence and point cloud sequence, and input the aligned image sequence and point cloud sequence into the trained dense depth estimation network and the digital-analog joint depth code optimization method; The uncertainty calculation module performs feature encoding on each frame of the aligned image sequence and the corresponding point cloud frame in the aligned point cloud sequence through the uncertainty branch in the dense depth estimation model, thereby obtaining convolution features; then, the convolution features are decoded to calculate the uncertainty of the depth of each pixel in the corresponding frame image; A dense depth estimation module is used to obtain a depth code corresponding to the current frame, input the depth code and the convolution feature into a dense depth estimation branch of the dense depth estimation model, decode the depth code and the convolution feature through the dense depth estimation branch to obtain a dense depth of each pixel unit on the image of the corresponding frame, and output a dense depth estimation map; The digital-analog joint depth code optimization module is used to utilize multi-frame common view information to construct a least squares loss function to optimize the depth codes of historical key frames and the current frame, thereby providing an optimized depth code for the dense depth estimation module. The least squares loss function includes the photometric error between the current frame and multiple previous key frame images, the reprojection error of matching feature points, the reprojection error of the millimeter-wave radar, and the Gaussian prior factor of the depth code.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Self-supervised monocular depth estimation method and device

    CN114022799A

  • Unsupervised monocular depth estimation method based on point cloud learning

    CN116912207A

  • Visual mileage calculation method and device, electronic equipment and storage medium

    CN117392228A

  • Multi-source fusion positioning method and system for digital-analog hybrid estimation

    CN117451043A

  • Depth estimation method based on millimeter wave radar-camera multi-stage fusion

    CN117830775A