Depth Estimation Method Based on the Fusion of Visual Images and Millimeter-Wave Radar Feature Matrices
By integrating millimeter wave radar feature matrices with visual images through a novel network architecture, the method enhances depth estimation accuracy, addressing the limitations of existing methods and improving performance in challenging conditions.
Patent Information
- Application Number
- CN202310303285.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-27
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-03-27
AI Technical Summary
The existing depth estimation methods are difficult to improve the accuracy of monocular depth estimation, especially in bad weather conditions, and insufficient utilization of millimeter-wave radar point cloud data information.
By fusing the visual image with the millimeter-wave radar feature matrix, the radar data is processed using a windowed 4-dimensional fast Fourier transform to generate a four-dimensional feature matrix, and aligning the two modal features with the residual network and the InfoNCE loss function. The feature fusion is performed using the hierarchical window attention module to finally generate a depth map.
It improves the accuracy of depth estimation, especially in severe weather conditions such as rain, fog, ice and snow, and fully utilizes the rich feature information of millimeter wave radar, enhancing the robustness of the method.
Smart Images

Figure CN116363186B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision, and particularly relates to a depth estimation method based on the fusion of visual images and millimeter-wave radar feature matrices. Background Art
[0002] Depth estimation is one of the key tasks in the field of computer vision and an important method for three-dimensional environmental perception. It has been widely applied to various scenarios, such as autonomous driving, robotics, intelligent spaces, intelligent transportation, augmented reality, etc. Currently, there are mainly active and passive methods to obtain the depth information of a scene.
[0003] Active methods use optical emitters, such as structured light, lasers, and microwaves, to estimate depth. These sensors have better resolution and low-light performance. However, their applications are restricted by the power and characteristics of the emitters and are difficult to use in various application scenarios. Passive methods use passive camera sensors and special algorithms to estimate depth. They have relatively low costs and theoretically have infinite ranging capabilities, but are easily affected by low-light conditions, complex algorithms, or small baseline distances.
[0004] Due to the emergence of low-cost computing power and high-resolution cameras, some people are committed to improving the performance of passive sensors, among which monocular depth estimation has attracted wide attention. Monocular depth estimation attempts to estimate the distance of corresponding scene obstacles using the information collected by a single camera, and some advanced learning algorithms are used to solve the problem of modeling the relationship between the depth map and the visual image of the same scene. However, since the depth information has been lost during the imaging process, these algorithms are easily affected by outliers that are difficult or less considered during training. In order to better converge and predict, special loss functions have also been designed.
[0005] Some methods further explore depth estimation methods using multi-modal data fusion, mostly the fusion of camera and lidar data or the fusion of camera and radar data. Among them, there are more methods using the fusion of camera and lidar. Lidar is an active high-resolution three-dimensional perception sensor. Some researchers have managed to reconstruct the depth map using sparse lidar points and the RGB image of the same scene, which is called depth completion and has also achieved good results. However, lidar is easily affected by bad weather conditions, and this sensor is prone to wear and tear and expensive, which is not conducive to practical popularization and use.
[0006] Methods for fusing cameras and millimeter-wave radars for depth estimation have not been fully explored. Millimeter-wave radar is a sensor for high-resolution environmental perception, which is highly robust to harsh weather conditions such as rain, fog, and dust and can also work properly in low-light conditions. In addition, millimeter-wave radar can measure relative radial velocity, which provides an alternative source of information for perceiving the surrounding environment. In recent years, millimeter-wave radar has attracted more attention, and some depth estimation methods based on millimeter-wave radar data have been proposed, which can be roughly divided into depth estimation using millimeter-wave radar in combination with images or using millimeter-wave radar alone. However, most researchers only focus on the point cloud detected by the radar, which is the result of detecting the transformed radar feature matrix and is much sparser and contains less information than lidar point clouds.
[0007] "A Fusion Depth Estimation Method Based on Vision and Millimeter-Wave Radar" (Patent Application Publication No. CN114627351A) proposes a method for fusing visual images and millimeter-wave radar point clouds for depth estimation. A sparse-rough coding network is used to extract features to obtain a first fusion feature map. Subsequently, a sparse-rough decoding network is used to decode to obtain a rough depth map, which is fed into the sparse-rough decoding network in the second stage. A feature fusion module is used to fuse the features of the sparse-rough decoding network in the first stage into the sparse-rough decoding network in the second stage, and the final feature map is decoded. This method uses filtering-interpolation with a binary mask to filter the lidar measurement results, removing outliers and enhancing the density of the labeled data. However, although this method uses millimeter-wave radar data to improve the depth estimation task, it only considers the point cloud data of the millimeter-wave radar. The point cloud data is the result of object detection on the radar feature matrix, and there is information loss in this process, and the rich features obtained from radar measurements are not fully utilized.
[0008] "A Method for Warning of Obstacles around Vehicles Based on Monocular Depth Estimation" (Patent Application Publication No. CN114495064A) proposes first using a monocular depth estimation network to detect obstacles around vehicles and then cooperating with the Kalman filtering algorithm for obstacle tracking. This method does not consider the accuracy problem of monocular depth estimation in harsh weather.
[0009] At the 2018 Conference on Computer Vision and Pattern Recognition (CVPR), Guan et al. proposed a method for depth estimation using the three-dimensional feature matrix generated by millimeter-wave radar. This method uses a conditional generative adversarial network to generate a depth map of vehicles in the scene with the help of the three-dimensional feature matrix of millimeter-wave radar. This method uses a millimeter-wave radar matrix with relatively rich features, but does not use images to improve the accuracy of depth estimation.
[0010] In summary, the accuracy of existing depth estimation methods still needs to be improved. Summary of the Invention
[0011] The present invention provides a depth estimation method based on the fusion of visual images and millimeter-wave radar feature matrices to solve the technical problem of difficult improvement in accuracy in monocular depth estimation.
[0012] To solve the above technical problems, the present invention provides the following technical solutions:
[0013] On the one hand, the present invention provides a depth estimation method based on the fusion of visual images and millimeter-wave radar feature matrices, and the depth estimation method based on the fusion of visual images and millimeter-wave radar feature matrices includes:
[0014] Obtain the original millimeter-wave radar data and visual images of the scene to be measured;
[0015] Process the original millimeter-wave radar data to obtain a radar feature matrix;
[0016] Based on the radar feature matrix and the visual image, use a preset network model to obtain an estimated depth map.
[0017] Further, processing the original millimeter-wave radar data to obtain a radar feature matrix includes:
[0018] Perform signal transformation on the original millimeter-wave radar data using a windowed 4D fast Fourier transform, and the transformation result is a four-dimensional feature matrix; wherein, the four dimensions of the four-dimensional feature matrix respectively represent: distance, horizontal angle, elevation angle, and Doppler velocity;
[0019] Take the mean of the Doppler velocity dimension to obtain a first three-dimensional feature matrix with dimensions of distance, horizontal angle, and elevation angle, and the eigenvalue is positively correlated with the electromagnetic wave energy value returned by the obstacle; filter out the energy of low radial velocity in the Doppler velocity dimension through weighted summation to obtain a second three-dimensional feature matrix with the same dimensions, and the eigenvalue is positively correlated with the electromagnetic wave energy value returned by the obstacle with relative velocity;
[0020] Concatenate the first three-dimensional feature matrix and the second three-dimensional feature matrix along a new dimension to obtain a four-dimensional radar feature matrix.
[0021] Further, the obtaining of the estimated depth map using the preset network model includes;
[0022] Use a first encoder to extract features from the radar feature matrix;
[0023] Use a second encoder to extract features from the visual image;
[0024] After the radar feature matrix and the visual image are fed into their respective encoders to extract fine features, the InfoNCE loss is used to maximize the mutual information between the partial features of the two modalities, that is, to align the information of the partial features;
[0025] The features after information alignment are fused through a hierarchical window attention module, and then concatenated with the remaining features in the channel dimension to obtain the concatenated features;
[0026] A fully convolutional decoder is used to decode the concatenated features to obtain the estimated depth map.
[0027] Further, the first encoder adopts a 3D convolutional network based on a residual network.
[0028] Further, the second encoder adopts a model composed of a ResNet-101 and a perception module.
[0029] Further, during the training process of the network model, the learning rates of the first encoder and the second encoder are successively suppressed to 1% of the normal learning rate.
[0030] Further, the hierarchical window attention module achieves global attention to the input features by performing hierarchical window attention on the input feature matrix; between each layer, the features are downsampled by two-dimensional convolution, and then the attention matrix is calculated using the downsampled features in the next layer.
[0031] Further, the perception module extracts multi-scale image features through multiple groups of parallel network structures; among them, the multiple groups of parallel network structures include: a first network structure composed of global pooling, a fully connected layer with an activation function, and a 1×1 two-dimensional convolutional layer, a second network structure composed of two 1×1 two-dimensional convolutional layers with activation, a third network structure composed of a 6-fold dilated convolution with activation and a 1×1 two-dimensional convolutional layer, a fourth network structure composed of a 12-fold dilated convolution with activation and a 1×1 two-dimensional convolutional layer, and a fifth network structure composed of an 18-fold dilated convolution with activation and a 1×1 two-dimensional convolutional layer.
[0032] On the other hand, the present invention also provides an electronic device, which includes a processor and a memory; wherein, at least one instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the above method.
[0033] On another hand, the present invention also provides a computer-readable storage medium, in which at least one instruction is stored, and the instruction is loaded and executed by the processor to implement the above method.
[0034] The beneficial effects brought by the technical solution provided by the present invention at least include:
[0035] 1. The present invention innovatively integrates the four-dimensional feature matrix of millimeter-wave radar and visual images for depth estimation. The four-dimensional feature matrix of millimeter-wave radar has richer information than radar point clouds and stronger robustness than images.
[0036] 2. The present invention proposes a new method for processing radar raw signals, which is beneficial to expressing dynamic features in the scene.
[0037] 3. The present invention uses the InfoNCE loss to align the information of two data modalities, maximizing the mutual information of the two data modalities, which can effectively improve the performance of the method.
[0038] 4. The present invention proposes a hierarchical window attention module, which can further improve the performance of the method. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0040] Figure 1 is a schematic execution flow diagram of the depth estimation method based on the fusion of visual images and millimeter-wave radar feature matrices provided by the embodiments of the present invention;
[0041] Figure 2 is a schematic diagram of the data format used in the present invention; wherein, (a) is a visual image, (b) is the ground truth, and (c) is the visualization of the four-dimensional radar matrix;
[0042] Figure 3 is an overall view of the network model structure provided by the embodiments of the present invention;
[0043] Figure 4 is a flow chart of the hierarchical window attention provided by the embodiments of the present invention;
[0044] Figure 5 is a schematic diagram of the hierarchical window attention provided by the embodiments of the present invention;
[0045] Figure 6 is a structural diagram of the perception module provided by the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0046] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will further describe the embodiments of the present invention in detail with reference to the drawings.
[0047] First Embodiment
[0048] This embodiment provides a depth estimation method based on the fusion of visual images and millimeter-wave radar feature matrices, which can be implemented by an electronic device. The execution process is as follows Figure 1 shown, including the following steps:
[0049] S1. Obtain the millimeter-wave radar raw data and visual images of the scene to be measured;
[0050] Specifically, in this embodiment, the raw data is collected by building a hardware environment including a multiple-input multiple-output (MIMO) millimeter-wave radar and a binocular depth camera; the visual images and depth maps of the same scene can be collected using a binocular depth camera, or a lidar can be used to collect depth maps. Among them, the radar adopted in this embodiment has three transmitting antennas and four receiving antennas. The radar senses the environment by transmitting chirp waves with continuous frequency hopping. Its starting frequency is 77 GHz, the period is 32 μs, the frequency growth rate is 80 MHz / μs, the number of ADC sampling points is 256, the sampling rate is 12000 ksps, and each group has 128 chirp waves. The hardware environment of the binocular depth camera adopts a ZED2i binocular depth camera.
[0051] In this embodiment, a program is written to control the MIMO millimeter-wave radar and the binocular depth camera to collect data, so that the data collected by the MIMO millimeter-wave radar and the binocular depth camera are synchronized in time.
[0052] S2. Process the obtained millimeter-wave radar raw data to obtain a four-dimensional millimeter-wave radar feature matrix;
[0053] Among them, the method for processing the obtained millimeter-wave radar raw data in this embodiment is: perform signal transformation on the collected millimeter-wave radar raw data, reduce the data volume by weighted summation of specific dimensions of the transformation result, and retain the original information as much as possible, and save the processed radar feature matrix. For visual images and depth maps, the visual images of the same scene can be collected using a binocular camera, and depth maps can be generated using stereo vision. A data set is constructed using the results of the foregoing processing.
[0054] Specifically, in this embodiment, the method for performing signal transformation on the millimeter-wave radar raw data adopts a windowed 4D fast Fourier transform, and the transformation result is a four-dimensional feature matrix. The four dimensions respectively represent: distance, horizontal angle, elevation angle, and Doppler velocity.
[0055] The method of reducing the data volume by weighted summation of specific dimensions of the transformation result is as follows: calculate the mean value of the Doppler velocity dimension to obtain a three-dimensional feature matrix with dimensions of distance, horizontal angle, and elevation angle, and eigenvalues positively correlated with the electromagnetic wave energy value returned by the obstacle. Filter out the energy of low radial velocity in the Doppler velocity dimension through weighted summation to obtain a three-dimensional feature matrix with the same dimensions and eigenvalues positively correlated with the electromagnetic wave energy value returned by the obstacle with relative velocity.
[0056] Concatenate the two three-dimensional matrices along the new dimension to obtain the required four-dimensional feature matrix of the millimeter-wave radar.
[0057] Based on the above, the calculation method of the four-dimensional feature matrix of the millimeter-wave radar is as follows:
[0058] For a radar system with N T transmitting antennas and N R receiving antennas, when the transmitting antennas transmit N loop chirp waves, rearrange the data received by the receiving antennas to obtain the original data matrix, where N sample is the number of ADC sampling points of the radar.
[0059] After performing windowed fast Fourier transform on the four dimensions, the matrix can be obtained, where N range , N azimuth , N elevation , N doppler are the number of Fourier transform points for the distance, horizontal angle, elevation angle, and Doppler dimensions respectively. Then the calculation method of the four-dimensional feature matrix of the millimeter-wave radar is:
[0060]
[0061]
[0062] where r, a, e, d correspond to the four dimensions of matrix M, and W dop is a matrix of the same size as M. The values of the matrix along the Doppler dimension are as follows:
[0063]
[0064] where the vector w is any vector along the Doppler dimension of matrix W dop , n is the velocity threshold, and k satisfies the following relationship:
[0065]
[0066] Specifically, in this embodiment, combined with the actual data, rearrange the data received by the receiving antennas to obtain Mraw ∈R 256×4×3×128 The original data matrix ∈R. After performing windowed fast Fourier transform on the four dimensions, the matrix M' ∈R can be obtained. 128×64×40×128 . Among them, when performing fast Fourier transform on the horizontal angle and elevation angle dimensions, the Taylor window is used. When performing fast Fourier transform on the range and Doppler dimensions, the Hanning window is used. Since the detection range of the radar in the azimuth angle dimension is 180°, and the visible angle of the binocular camera is 120°, the matrix M' is cropped to obtain M' ∈R. 128×56×40×128 The four-dimensional feature matrix of the millimeter-wave radar is calculated by the method described in formula (1) to obtain the matrix H ∈R. 2 ×128×64×40 , and the obtained four-dimensional radar feature matrix is as Figure 2 shown.
[0067] Furthermore, for the method of generating a depth map using stereo vision, in this embodiment, the ZED software development environment is configured in the computer, and the depth map is generated according to the collected images using the baseline length and camera parameters of the binocular camera. The resolution of the visual images collected by the camera is 1280x720. To save video memory and facilitate calculation, the image resolution is adjusted to 640x352.
[0068] S3. Based on the four-dimensional feature matrix of the millimeter-wave radar and the visual image, an estimated depth map is obtained using a preset network model.
[0069] Specifically, in this embodiment, the process of obtaining the estimated depth map is as Figure 3 shown, including:
[0070] By constructing appropriate encoders for the radar data and visual image data respectively, the radar data and visual image data are sent into their respective encoders to extract fine features, and the InfoNCE loss is used to maximize the mutual information between the partial features of the two modalities, that is, information alignment of the partial features.
[0071] Among them, the encoder used for the radar data adopts a 3D convolutional network based on the residual network. The encoder used for the visual image data uses a model composed of a residual network-101 and a specially designed perception module. Among them, the perception module is as Figure 6 shown. This module extracts multi-scale image features through multiple groups of parallel network structures. These structures include: 1. Downsample the features through global pooling, and then pass through a fully connected layer with an activation function and a 1×1 two-dimensional convolutional layer, and obtain the feature F through interpolation. i,0 ∈R 512×44×80 . 2. Obtain the feature F through two layers of 1×1 two-dimensional convolutional layers with activation. i,1 ∈R 512×44×803. The feature F is obtained through 6-fold dilated convolution with activation and a 1×1 two-dimensional convolutional layer. i,2 ∈R 512×44×80 4. The feature F is obtained through 12-fold dilated convolution with activation and a 1×1 two-dimensional convolutional layer. i,3 ∈R 512×44×80 5. The feature F is obtained through 18-fold dilated convolution with activation and a 1×1 two-dimensional convolutional layer. i,4 ∈R 512×44×80 Finally, the feature matrix F after visual image encoding i ∈R 2560×44×80 , where F iu ∈R 256×44×80 , F ic ∈R 2304×44×80 . The feature matrix F after radar data encoding r ∈R 512×44×80 , where F ru ∈R 256×44×80 , F rc ∈R 256×44×80 .
[0072] The InfoNCE loss calculation method with a queue is adopted to align the information of the two data modalities. The loss calculation method is as follows:
[0073]
[0074]
[0075]
[0076] If F i ={F iu , F ic}, F r ={F ru , F rc} respectively represent the results of the encoder encoding the two data modalities, then the vector q i =linear(F ic ), and the vector q r =linear(F rc ). k i is a queue for storing K vectors q i , and k r is the same. Through training with a large amount of data, the L nce loss can maximize the mutual information between F ic and F rc to align the information of the features.
[0077] Maximize the mutual information between the partial features of the two modalities using the InfoNCE loss, that is, calculate the InfoNCE using formula (4) and use the gradient descent algorithm to reduce this loss, which maximizes F ic and F rc Align the information by maximizing the mutual information of the two data modalities existing between them.
[0078] Furthermore, to better use the InfoNCE loss for information alignment, this embodiment innovatively adopts an optimized learning rate adjustment strategy. That is, during the training process, the learning rates of the visual image encoder and the millimeter-wave radar encoder will be successively suppressed to 1% of the normal learning rate. This can enable the encoders with suppressed learning rates to encode more stable features, which is beneficial to the queue learning of the InfoNCE loss.
[0079] Fuse the features after information alignment through a self-designed hierarchical window attention module, then concatenate them with the remaining features in the channel dimension, and use a fully convolutional decoder to decode the concatenated features to obtain the estimated depth map. Among them, the hierarchical window attention module achieves global attention on the input features by performing hierarchical window attention on the input feature matrix. Between each layer, the features are downsampled through two-dimensional convolution, and then the attention matrix is calculated using the downsampled features in the next layer. Specifically, the calculation process of the hierarchical window attention module is as Figure 4 and Figure 5 shown. The hierarchical window attention module calculates the attention matrix for the input features through a multi-layer window attention mechanism, and then applies the attention matrix to the original features. This can calculate the attention for the features with relatively small memory and computational overhead.
[0080] Based on the above, this embodiment finally constructs a dataset including 9624 groups of data, and uses the InfoNCE loss and the ordinal loss to train the model on two Nvidia RTX3090s. Finally, the performance of this method in the test set is shown in the following table. This proves that the performance of this method is relatively leading.
[0081] Table 2 Test results of this method in the test set
[0082]
[0083] In summary, this embodiment proposes a depth estimation method that fuses visual images and millimeter-wave radar feature matrices by virtue of the new radar data format, mainly involving the following two aspects:
[0084] 1. Propose a new calculation method for the radar feature matrix
[0085] This method performs a windowed four-dimensional fast Fourier transform on the original ADC signals of the MIMO radar to obtain a four-dimensional data matrix. The four dimensions of this matrix represent range, azimuth angle, elevation angle, and Doppler velocity respectively. Each value in the matrix is positively correlated with the energy of the radar electromagnetic wave reflected by the obstacle at a specific range, azimuth angle, elevation angle, and Doppler velocity. Due to the large amount of data in this matrix, the present invention reduces the dimension of the Doppler velocity dimension in two ways. The first way is to calculate the mean value of the four-dimensional matrix in the Doppler velocity dimension, reducing it to a three-dimensional matrix containing the energy information of all obstacles within the radar detection range. The second way is to calculate the four-dimensional matrix in the Doppler velocity dimension using formula (1), reducing the dimension to obtain a three-dimensional matrix containing only the energy information of obstacles with a relatively high radial velocity relative to the radar within the detection range. These two matrices are concatenated along the new dimension to obtain the four-dimensional feature matrix of the millimeter-wave radar used in the invention.
[0086] 2. A new depth estimation method that fuses millimeter-wave radar and visual images is proposed
[0087] This method is the first solution to fuse the millimeter-wave radar feature matrix and visual images for depth estimation. This method also proposes a new radar feature matrix, which expresses more information about the scene with less space occupancy. This method uses two encoders to encode the four-dimensional feature matrix of the millimeter-wave radar and the visual image respectively. Among them, the encoder for millimeter-wave radar data uses a 3D convolutional network based on the residual network, and the encoder for visual images uses a model composed of the residual network-101 and a specially designed perception module. For the feature matrices obtained after respective encoding, this method innovatively uses the InfoNCE loss with a queue to align the features of the two data modalities, which is beneficial to improving the performance of depth estimation. In order to better utilize the InfoNCE loss, this method also innovatively uses a strategy of alternating learning rate suppression, suppressing the learning rates of the two encoders to 1% of the original during the training process to make their encoding results more stable. In addition, this method also innovatively proposes a hierarchical window attention module, which uses two-dimensional convolution to reduce the dimension to obtain a multi-level feature matrix, then uses window attention to obtain the corresponding attention matrix, and applies it to the original feature matrix to complete the attention operation, which also helps to improve the performance of depth estimation.
[0088] Based on the above, the method provided in this embodiment can effectively improve the accuracy of full-scene depth estimation, especially in environments with interference such as rain, fog, ice, snow, or low-light conditions, and has broad application prospects.
[0089] Second Embodiment
[0090] This embodiment provides an electronic device, which includes a processor and a memory; wherein, at least one instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the method of the first embodiment.
[0091] The electronic device may vary greatly due to different configurations or performances, and may include one or more processors (central processing units, CPUs) and one or more memories. Among them, at least one instruction is stored in the memory, and the instruction is loaded and executed by the processor to implement the above method.
[0092] The third embodiment
[0093] This embodiment provides a computer-readable storage medium, in which at least one instruction is stored, and the instruction is loaded and executed by the processor to implement the method of the above first embodiment. Among them, the computer-readable storage medium may be a ROM, a random access memory, a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc. The instructions stored therein can be loaded and executed by the processor in the terminal to implement the above method.
[0094] In addition, it should be noted that the present invention can be provided as a method, a device or a computer program product. Therefore, the embodiments of the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the embodiments of the present invention can take the form of a computer program product implemented on one or more computer-usable storage media containing computer-usable program code.
[0095] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, embedded processors, or other programmable data processing terminal devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing terminal devices generate for implementing in the process Figure 1 one process or multiple processes and / or blocks Figure 1 a device for implementing the functions specified in one block or multiple blocks.
[0096] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing terminal device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured product including an instruction device, and the instruction device implements in the process Figure 1 one process or multiple processes and / or blocksFigure 1 The functions specified in one or more boxes. These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable terminal device provide steps for implementing the functions specified in one Figure 1 process or multiple processes and / or boxes Figure 1 or more boxes.
[0097] It should also be noted that in this text, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or terminal device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or terminal device. Without further limitation, an element defined by the statement "comprising one..." does not exclude the presence of additional identical elements in the process, method, article or terminal device comprising the said element.
[0098] Finally, it should be noted that the above is the preferred embodiment of the present invention. It should be pointed out that although the preferred embodiments of the present invention have been described, for those skilled in the art of this technology, once the basic creative concept of the present invention is known, several improvements and refinements can be made without departing from the principle of the present invention. These improvements and refinements should also be regarded as the protection scope of the present invention. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications falling within the scope of the embodiments of the present invention.
Claims
1. A depth estimation method based on the fusion of visual images and millimeter-wave radar feature matrices, characterized in that The depth estimation method based on the fusion of visual images and millimeter-wave radar feature matrices includes: Obtain the original millimeter-wave radar data and visual images of the scene to be measured; Process the original millimeter-wave radar data to obtain a radar feature matrix; Based on the radar feature matrix and the visual image, use a preset network model to obtain an estimated depth map; Processing the original millimeter-wave radar data to obtain a radar feature matrix includes: Perform signal transformation on the original millimeter-wave radar data using a windowed 4D fast Fourier transform, and the transformation result is a four-dimensional feature matrix; where the four dimensions of the four-dimensional feature matrix respectively represent: distance, horizontal angle, elevation angle, and Doppler velocity; Calculate the mean value of the Doppler velocity dimension to obtain a first three-dimensional feature matrix with dimensions of distance, horizontal angle, and elevation angle, and the eigenvalue is positively correlated with the electromagnetic wave energy value returned by the obstacle; filter the energy of the low radial velocity in the Doppler velocity dimension through weighted summation to obtain a second three-dimensional feature matrix with the same dimensions, and the eigenvalue is positively correlated with the electromagnetic wave energy value returned by the obstacle with relative velocity; Stitch the first three-dimensional feature matrix and the second three-dimensional feature matrix along a new dimension to obtain a four-dimensional radar feature matrix; The obtaining of the estimated depth map by using the preset network model includes: Use the first encoder to extract features from the radar feature matrix; Use the second encoder to extract features from the visual image; After the radar feature matrix and the visual image are sent to their respective encoders to extract fine features, use the InfoNCE loss to maximize the mutual information between the partial features of the two modalities, that is, align the partial features; Fuse the features after information alignment through a hierarchical window attention module, and then stitch them with the remaining features in the channel dimension to obtain the stitched features; Use a fully convolutional decoder to decode the stitched features to obtain an estimated depth map; The hierarchical window attention module achieves global attention to the input features by performing hierarchical window attention on the input feature matrix; between each layer, the features are downsampled through two-dimensional convolution, and then the attention matrix is calculated using the downsampled features in the next layer.
2. The depth estimation method based on the fusion of visual images and millimeter-wave radar feature matrices according to claim 1, wherein, The first encoder uses a 3D convolutional network based on a residual network.
3. The depth estimation method based on the fusion of visual images and millimeter-wave radar feature matrices according to claim 1, characterized in that The second encoder uses a model composed of a ResNet-101 and a perception module.
4. The depth estimation method based on the fusion of visual images and millimeter-wave radar feature matrices according to claim 1, characterized in that During the training process of the network model, the learning rates of the first encoder and the second encoder will be successively suppressed to 1% of the normal learning rate.
5. The depth estimation method based on the fusion of visual images and millimeter-wave radar feature matrices according to claim 3, wherein, The perception module extracts multi-scale image features through multiple groups of parallel network structures. Among them, the multiple groups of parallel network structures include: a first network structure composed of global pooling, a fully connected layer with an activation function, and a 1×1 two-dimensional convolutional layer; a second network structure composed of two layers of 1×1 activated two-dimensional convolutional layers; a third network structure composed of an activated 6-fold dilated convolution and a 1×1 two-dimensional convolutional layer; a fourth network structure composed of an activated 12-fold dilated convolution and a 1×1 two-dimensional convolutional layer; and a fifth network structure composed of an activated 18-fold dilated convolution and a 1×1 two-dimensional convolutional layer.
Citation Information
Patent Citations
Early warning method for obstacles around vehicle based on monocular depth estimation
CN114495064A
Fusion depth estimation method based on vision and millimeter wave radar
CN114627351A