Method, device and system for acquiring environment information in unmanned driving and storage medium

By combining the stereo matching network of binocular camera and lidar with the three-branch pyramid grafting network, the problem of insufficient sensor accuracy and robustness in driverless car environment perception is solved, and efficient acquisition of dense depth maps is achieved.

CN120279529APending Publication Date: 2025-07-08GUILIN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510410685.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-02
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

In the prior art, it is difficult for a single sensor to directly acquire high-precision and robust depth images. Traditional lidar performance is degraded and costly under severe weather conditions. The depth data accuracy of stereo cameras is limited, making it difficult to meet the environmental perception needs of driverless cars.

Method used

Using a combination of binocular cameras and lidar, a three-branch pyramid grafting network is achieved through a stereo matching network and a three-branch pyramid grafting network, including horizontal-vertical cross attention feature extraction, multi-head cross attention differential connection aggregation module and iterative refinement perception global aggregation module, combined with Swin Transformer, Mobilev3 and Xception branch networks, the fusion of binocular vision and laser data is achieved.

Benefits of technology

The feature extraction capability and matching accuracy are improved, and the limited focus parallax refinement of occlusion and edge areas is achieved, a good balance between matching accuracy and speed is achieved, and dense depth maps are obtained efficiently.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279529A_ABST
    Figure CN120279529A_ABST
Patent Text Reader

Abstract

The invention discloses a method, a device and a system for acquiring environment information in unmanned driving, and a storage medium. The method comprises the following steps: acquiring an RGB image acquired by a binocular camera and point cloud data acquired by a laser radar; inputting the RGB image into a stereo matching network to obtain binocular stereo image data; the stereo matching network comprises a feature extraction layer with horizontal-vertical cross attention, a multi-head cross attention differential connection aggregation module and an iterative refinement perception global aggregation module; inputting the point cloud data and the binocular stereo image data into a three-branch pyramid-based grafting network to obtain surrounding environment information of the pilotless automobile; the grafting network based on the three-branch pyramid comprises a cross-model transplantation module based on an attention mechanism, a Swin Transform branch network, a Mobiv3 branch network and an Xception branch network. By adopting the technical scheme of the invention, a high-precision and high-robustness depth image can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of driving technology, and in particular relates to a method and device, a system, and a storage medium for acquiring environmental information in unmanned driving. Background Art

[0002] Environmental perception is a crucial guiding technology and core technology in driverless cars. It can efficiently and accurately obtain information about the surrounding environment. It is the basic premise and key link to ensure the safe and efficient driving of driverless cars. Perception is equivalent to the "eyes" of driverless cars. It can obtain and understand the scene information around the car body, including obstacle detection, target detection and recognition, synchronous positioning and mapping. The accuracy of perception directly affects the quality of subsequent control, decision-making and estimation work, and is related to the stability and safety of driverless cars. As an important part of environmental perception, depth information acquisition technology is the key to accurately perceive, identify and understand the three-dimensional scene information of the objective world. It can intuitively and accurately reflect, describe and reconstruct the real world, and has extremely important application value.

[0003] In recent years, the use of multiple complementary sensors to obtain high-resolution images and high-precision depth information has gradually become a research focus. In particular, with the rapid development of deep learning technology, the performance of camera and lidar fusion algorithms has been significantly improved, providing a new solution for depth information acquisition. The vision-based depth information acquisition system can obtain the color and semantic information of the perceived target at a low cost, but a single camera is difficult to provide reliable three-dimensional geometric structure, which is crucial for unmanned driving technology. In contrast, stereo cameras can capture three-dimensional geometric information, but the accuracy of their depth data is limited and easily affected by the environment. As an active sensor, lidar can directly obtain accurate three-dimensional geometric information of scene point clouds. It has a long perception distance, high measurement accuracy, and is not affected by light conditions. However, lidar also has shortcomings such as low resolution, slow refresh rate, performance degradation under bad weather conditions, and high cost. Especially in vehicle-mounted applications, traditional scanning lidar often has motion distortion problems such as point cloud confusion, blurring, and tailing when capturing high-speed moving targets. The emergence of array laser technology provides a new idea for solving the above problems. This technology can achieve instant imaging of high-speed moving targets, effectively eliminating the mechanical scanning process and significantly speeding up the ranging and imaging speed. However, even though array lasers have many advantages, single sensors are still limited by their own imaging principles and it is difficult to directly obtain high-precision, high-robustness depth images. Summary of the invention

[0004] The technical problem to be solved by the present invention is to provide a method and device, system and storage medium for obtaining environmental information during unmanned driving.

[0005] To achieve the above object, the present invention adopts the following technical solutions:

[0006] An environmental information acquisition method for driverless driving, comprising:

[0007] Step S1, acquiring RGB images collected by a binocular camera and point cloud data collected by a lidar;

[0008] Step S2, inputting the RGB images into a stereo matching network to obtain binocular stereo image data; wherein, the stereo matching network includes: a feature extraction layer with horizontal-vertical cross-attention, a multi-head cross-attention differential connection aggregation module, and an iterative refinement perception global aggregation module;

[0009] Step S3, inputting the point cloud data and the binocular stereo image data into a three-branch pyramid grafting network to obtain the surrounding environmental information of the driverless vehicle; wherein, the three-branch pyramid grafting network includes: a cross-model transplantation module based on an attention mechanism, a Swin Transformer branch network, a Mobilev3 branch network, and an Xception branch network.

[0010] Preferably, in step S2, the feature extraction layer with horizontal-vertical cross-attention is used to extract global context information and high-level features at the pixel level; the multi-head cross-attention differential connection aggregation module is used to aggregate high-level semantic and low-level texture features; the iterative refinement perception global aggregation module utilizes global spatial correlation and context guidance to gradually optimize the disparity map of the limited focus range in the occluded and edge regions.

[0011] Preferably, in step S3, the Swin Transformer branch network extracts FPN-style binocular features from the binocular stereo images, and transmits the global semantic information of the binocular features to the Mobilev3 branch network to achieve feature grafting; the cross-model transplantation module based on the attention mechanism guides the feature grafting process through a guidance loss, and at the same time uses a CAM matrix for supervision; the Xception branch network fuses the extracted point cloud features with the binocular features supervised by the self-attention mechanism.

[0012] The present invention also provides an environmental information acquisition device for driverless driving, comprising:

[0013] An acquisition module, configured to acquire RGB images collected by a binocular camera and point cloud data collected by a lidar;

[0014] A first processing module, configured to input the RGB images into a stereo matching network to obtain binocular stereo image data; wherein, the stereo matching network includes: a feature extraction layer with horizontal-vertical cross-attention, a multi-head cross-attention differential connection aggregation module, and an iterative refinement perception global aggregation module;

[0015] A second processing module, configured to input the point cloud data and the binocular stereo image data into a three-branch pyramid grafting network to obtain the surrounding environment information of the driverless vehicle; wherein, the three-branch pyramid grafting network includes: a cross-model transplantation module based on an attention mechanism, a Swin Transformer branch network, a Mobilev3 branch network, and an Xception branch network.

[0016] Preferably, a feature extraction layer with horizontal-vertical cross-attention is used to extract global context information and high-level features at the pixel level; a multi-head cross-attention differential connection aggregation module is used to aggregate high-level semantic and low-level texture features; an iterative refinement perception global aggregation module uses global spatial correlation and context guidance to gradually optimize the disparity map of the limited focus range in the occluded and edge regions.

[0017] Preferably, the Swin Transformer branch network extracts FPN-style binocular features from the binocular stereo image, and transmits the global semantic information of the binocular features to the Mobilev3 branch network to achieve feature grafting; the cross-model transplantation module based on the attention mechanism guides the feature grafting process through a guiding loss, and at the same time uses a CAM matrix for supervision; the Xception branch network fuses the extracted point cloud features with the binocular features supervised by the self-attention mechanism.

[0018] The present invention also provides an in-driverless environment information acquisition system, including: a memory and a processor, where a computer program run by the processor is stored on the memory, and the computer program executes the in-driverless environment information acquisition method when run by the processor.

[0019] The present invention also provides a storage medium, where a computer program is stored on the storage medium, and the computer program executes the in-driverless environment information acquisition method when running.

[0020] The present invention has the following technical effects:

[0021] 1. By introducing a multi-head cross-attention mechanism and a differential connection aggregation module, the present invention effectively improves the feature extraction ability and matching accuracy. At the same time, through the iterative refinement perception global aggregation module, the limited focus disparity refinement of the occluded and edge regions is realized, thus achieving a good balance between matching accuracy and speed.

[0022] 2. The present invention adopts a three-branch network of Mobilev3, Swin Transformer and Xception, and realizes the efficient acquisition of a dense depth map by fusing binocular vision and laser data. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings described below are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on the provided drawings.

[0024] Figure 1 It is a flowchart of a method for obtaining environmental information in the driverless of the embodiments of the present invention. Specific Embodiments

[0025] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0026] To make the above objects, features, and advantages of the present invention more clearly understood, the present invention will be further described in detail below with reference to the drawings and specific embodiments.

[0027] Embodiment 1:

[0028] As Figure 1 shown, the embodiments of the present invention provide a method for obtaining environmental information in driverless, including:

[0029] Step S1, obtaining the RGB images collected by the binocular camera and the point cloud data collected by the lidar;

[0030] Step S2, inputting the RGB images into the stereo matching network to obtain binocular stereo image data; wherein, the stereo matching network includes: a feature extraction layer with horizontal-vertical cross-attention, a multi-head cross-attention differential connection aggregation module, and an iterative refinement perception global aggregation module;

[0031] Step S3, inputting the point cloud data and the binocular stereo image data into the three-branch pyramid grafting network to obtain the surrounding environmental information of the driverless vehicle; wherein, the three-branch pyramid grafting network includes: a cross-model transplantation module based on the attention mechanism, a Swin Transformer branch network, a Mobilev3 branch network, and an Xception branch network.

[0032] As an implementation manner of an embodiment of the present invention, in step S2, the feature extraction layer with horizontal-vertical cross-attention is used to extract global context information and high-level features at the pixel level; the multi-head cross-attention differential connection aggregation module is used to aggregate high-level semantics and low-level texture features; the iterative refinement perception global aggregation module uses global spatial correlation and context guidance to gradually optimize the disparity map of the limited focus range in the occluded and edge regions.

[0033] Further, the RGB images collected by the binocular camera are input into the feature extraction layer with horizontal-vertical cross-attention, which is the primary step of stereo matching. Representative features are extracted from the input left and right images, aiming to reduce the data dimension, reduce the computational complexity, and retain the necessary feature information for stereo matching. Then the features are passed into the multi-head cross-attention differential connection aggregation module, which calculates the cost volume based on the left and right image features and performs a channel fusion operation on the cost volume to enhance the feature expression ability of the cost volume.

[0034] Cost volume calculation: Let the left and right features be f l , f r ∈R C×W×H . For each pixel (x, y) in the left feature map, the corresponding positions in the right image within the disparity range are (x - d, y). The matching cost of the left-view pixel at different disparities with the corresponding position in the right view is calculated using the inner product. Finally, the cost volume C(x, y, d) can be obtained, and the calculation formula is as follows:

[0035] C(x,y,d)=1 / N c <f l (x,y),f r (x-d,y)〉

[0036] where <,> represents the inner product operation, and N c represents the number of feature channels.

[0037] Finally, an enhanced cost volume is obtained. Finally, the cost volume is passed into the iterative refinement perception global aggregation module, which uses an iterative mechanism to calculate the disparity of each pixel point from the cost volume. After the serial processing of the three modules, the disparity map is finally obtained from the binocular images.

[0038] Iterative mechanism: Receive the disparity d t-1 at the t - 1 stage and the left-view neighborhood pixel information F′ L extracted from the CNN as inputs. The cross-attention mechanism is used for the left and right feature maps respectively to obtain the cross-attention matrix Then, the multi-head cross-attention is used again to sample the cross-attention, and local cross-attention is constructed to measure d t-1The similarity between the left and right surrounding features. Then, the local cross-attention and the corresponding disparity are passed to the disparity encoder to obtain the matching features Meanwhile, the information F′ of the left view neighborhood pixels L is concatenated with the matching features to supplement the local features. In smooth regions, this information is sufficient for disparity optimization. However, for occluded and edge regions, the global spatial correlation of the left image is calculated using the self-attention module, and the self-global attention matrix is obtained By querying the attention map, the correlation between a specific point (i, j) and all other pixels in the left view can be obtained and feature aggregation with differential connection is performed to obtain the global features Subsequently, using the iterative refinement-aware global aggregation mechanism, local features are retained in smooth regions, while global features are retained in occluded and edge regions to generate the adaptive feature F′ ada , for overall disparity refinement. The calculation formula is as follows:

[0039]

[0040] Iterative optimization uses a GRU network, and the calculation process is as follows:

[0041] Input data:

[0042] 1. Input data preparation:

[0043] The input vector at the current time step

[0044] The hidden state h at the previous time step t-1 , with an initial value of 0.

[0045] 2. Calculate the gate values:

[0046] Calculate the update gate z t = σ(W z x t + U z h t-1 + b z ).

[0047] Calculate the reset gate r t = σ(W r x t + U r h t-1 + b r ).

[0048] 3. Generate the candidate hidden state:

[0049] Use the reset gate to filter the old hidden state information r t ⊙ h t-1 .

[0050] Calculate candidate hidden states

[0051] 4. Update the hidden state:

[0052] Calculate the final hidden state according to the weights of the update gate

[0053] 5. Output the hidden state:

[0054] Output h t , which is also used as the input for the next time step.

[0055] As an implementation manner of the embodiment of the present invention, in step S3, the Swin Transformer branch network extracts binocular features in the FPN style from the binocular stereo image, and transfers the global semantic information of the binocular features to the Mobilev3 branch network to achieve feature grafting; the cross-model transplantation module based on the attention mechanism guides the feature grafting process through the guiding loss, and at the same time uses the CAM matrix for supervision; the Xception branch network fuses the extracted point cloud features with the binocular features supervised by the self-attention mechanism.

[0056] Cross-model transplantation process:

[0057] Cross-model transplantation aims to effectively fuse the features f M5 and f S2 extracted by two different encoders. First, map the feature f M5 ∈V H×W×C to the same spatial dimension f′ M ∈υ 1×C×H×W , and normalize and linearly project f S2 to obtain After that, calculate Z using matrix multiplication, and this process can be expressed as a mathematical formula as follows

[0058]

[0059]

[0060] Then, input Z into the linear projection layer and reshape it back to the size V H×W×C before feeding it into the convolutional layer. Two shortcut connections are made during the process. In addition, during the cross-attention process, a cross-attention matrix is generated based on Y, and the calculation formula is as follows:

[0061] CAM = ReLU(BN(Conv(Y + Y T )))

[0062] The point cloud data and the binocular stereo images are respectively input into three branch networks arranged in parallel. Each branch network processes one type of data. Among them, the Xception branch network processes the point cloud data and adopts the idea of depthwise separable convolution. This design not only reduces the number of parameters in the network and improves the computational efficiency, but also first efficiently extracts the spatial features and structural information of the point cloud data, and then decodes the depth information from the features; the Mobilev3 branch network processes the left view data of the binocular stereo images and adopts depthwise separable convolution and inverted residual structure, which can reduce the computational complexity and the number of parameters while maintaining high performance. It can extract rich detail information when processing high-resolution inputs, providing necessary local features for depth estimation; the Swin Transformer branch network processes the right view data of the binocular stereo images, and it uses the attention mechanism to extract image features. Then, the left view features output by the Mobilev3 branch network and the right view features output by the Swin Transformer branch network are merged into a new feature map. Next, the decoder operation is used to regress the depth information of the binocular stereo images from the merged features. Finally, the depth information and the depth information output by the Xception are fused to finally obtain the depth map after the fusion of the radar map and the binocular stereo images.

[0063] Depthwise separable convolution: Depthwise separable convolution is a variant of the standard convolution. By dividing the standard convolution into two steps: depthwise convolution and pointwise convolution, it significantly reduces the computational amount and the number of parameters. Its processing flow is as follows:

[0064] 1. Depthwise convolution

[0065] Input: The size of the feature map is H×W×C in (height × width × number of input channels).

[0066] Operation: Perform spatial convolution on each input channel separately (using C in independent convolution kernels), and each convolution kernel only acts on the corresponding single channel.

[0067] Convolution kernel size: K h ×K w ×1

[0068] Output: The size of the feature map is H′×W′×C in , where H′ and W′ are determined by the stride and padding.

[0069] 2. Pointwise convolution

[0070] Input: The output of the depthwise convolution H′×W′×C in .

[0071] Operation: Use a 1×1 convolution kernel to fuse the channel information and combine Cin One channel is mapped to the number of target output channels C out 。

[0072] Convolution kernel size: 1×1×C in ×C out 。

[0073] Output: The size of the feature map is H′×W′×C out 。

[0074] Embodiment 2:

[0075] The embodiment of the present invention further provides an environmental information acquisition device for driverless driving, including:

[0076] An acquisition module, configured to acquire RGB images collected by a binocular camera and point cloud data collected by a lidar;

[0077] A first processing module, configured to input the RGB image into a stereo matching network to obtain binocular stereo image data; wherein, the stereo matching network includes: a feature extraction layer with horizontal-vertical cross-attention, a multi-head cross-attention differential connection aggregation module, and an iterative refinement perception global aggregation module;

[0078] A second processing module, configured to input the point cloud data and the binocular stereo image data into a three-branch pyramid grafting network to obtain the surrounding environmental information of the driverless vehicle; wherein, the three-branch pyramid grafting network includes: a cross-model transplantation module based on an attention mechanism, a Swin Transformer branch network, a Mobilev3 branch network, and an Xception branch network.

[0079] As an implementation manner of the embodiment of the present invention, the feature extraction layer with horizontal-vertical cross-attention is used to extract global context information and high-level features at the pixel level; the multi-head cross-attention differential connection aggregation module is used to aggregate high-level semantic and low-level texture features; the iterative refinement perception global aggregation module uses global spatial correlation and context guidance to gradually optimize the disparity map of the limited focus range in the occluded and edge regions.

[0080] As an implementation manner of the embodiment of the present invention, the Swin Transformer branch network extracts FPN-style binocular features from the binocular stereo image and transmits the global semantic information of the binocular features to the Mobilev3 branch network to achieve feature grafting; the cross-model transplantation module based on the attention mechanism guides the feature grafting process through a guiding loss and simultaneously uses the CAM matrix for supervision; the Xception branch network fuses the extracted point cloud features with the binocular features supervised by the self-attention mechanism.

[0081] Embodiment 3:

[0082] An embodiment of the present invention further provides an environmental information acquisition system for driverless driving, including: a memory and a processor, where a computer program run by the processor is stored on the memory, and the computer program executes an environmental information acquisition method for driverless driving when run by the processor.

[0083] Example 4:

[0084] An embodiment of the present invention further provides a storage medium, where a computer program is stored on the storage medium, and the computer program executes an environmental information acquisition method for driverless driving when running.

[0085] The above-described embodiments are only descriptions of the preferred embodiments of the present invention, and do not limit the scope of the present invention. Without departing from the design spirit of the present invention, various deformations and improvements made by those of ordinary skill in the art to the technical solution of the present invention shall fall within the protection scope determined by the claims of the present invention.

Claims

1. An environmental information acquisition method in unmanned driving, characterized in that, Including: Step S1: Obtain the RGB image collected by the binocular camera and the point cloud data collected by the lidar; Step S2: Input the RGB image into the stereo matching network to obtain binocular stereo image data; Among them, the stereo matching network includes: a feature extraction layer with horizontal-vertical cross-attention, a multi-head cross-attention differential connection aggregation module, and an iterative refinement perception global aggregation module; Step S3: Input the point cloud data and the binocular stereo image data into the three-branch pyramid grafting network to obtain the surrounding environment information of the driverless vehicle; Among them, the three-branch pyramid grafting network includes: a cross-model transplantation module based on the attention mechanism, a Swin Transformer branch network, a Mobilev3 branch network, and an Xception branch network.

2. The method for obtaining environmental information in driverless driving according to claim 1, wherein In Step S2, the feature extraction layer with horizontal-vertical cross-attention is used to extract global context information and high-level features at the pixel level; the multi-head cross-attention differential connection aggregation module is used to aggregate high-level semantic and low-level texture features; the iterative refinement perception global aggregation module uses global spatial correlation and context guidance to gradually optimize the disparity map in the limited focus range of the occluded and edge regions.

3. The method for obtaining environmental information in driverless driving according to claim 2, characterized in that, In Step S3, the Swin Transformer branch network extracts FPN-style binocular features from the binocular stereo image, and transfers the global semantic information of the binocular features to the Mobilev3 branch network to achieve feature grafting; the cross-model transplantation module based on the attention mechanism guides the feature grafting process through the guidance loss, and at the same time uses the CAM matrix for supervision; the Xception branch network fuses the extracted point cloud features with the binocular features supervised by the self-attention mechanism.

4. An environmental information acquisition device in driverless driving, characterized in that, Including: An acquisition module, used to obtain the RGB image collected by the binocular camera and the point cloud data collected by the lidar; A first processing module, used to input the RGB image into the stereo matching network to obtain binocular stereo image data; Among them, the stereo matching network includes: a feature extraction layer with horizontal-vertical cross-attention, a multi-head cross-attention differential connection aggregation module, and an iterative refinement perception global aggregation module; A second processing module, used to input the point cloud data and the binocular stereo image data into the three-branch pyramid grafting network to obtain the surrounding environment information of the driverless vehicle; Among them, the three-branch pyramid grafting network includes: a cross-model transplantation module based on the attention mechanism, a Swin Transformer branch network, a Mobilev3 branch network, and an Xception branch network.

5. The environmental information acquisition device in the driverless vehicle according to claim 4, characterized in that, The feature extraction layer with horizontal-vertical cross-attention is used to extract global context information and high-level features at the pixel level; the multi-head cross-attention differential connection aggregation module is used to aggregate high-level semantic and low-level texture features; the iterative refinement perception global aggregation module uses global spatial correlation and context guidance to gradually optimize the disparity map in the limited focus range of the occluded and edge regions.

6. The environmental information acquisition device in the driverless vehicle according to claim 5, wherein The Swin Transformer branch network extracts FPN-style binocular features from binocular stereo images and transfers the global semantic information of the binocular features to the Mobilev3 branch network to achieve feature grafting. The cross-model transplantation module based on the attention mechanism guides the feature grafting process through the guidance loss and simultaneously uses the CAM matrix for supervision. The Xception branch network fuses the extracted point cloud features with the binocular features supervised by the self-attention mechanism.

7. An environmental information acquisition system in unmanned driving, characterized in that, Including: A memory and a processor, wherein a computer program is stored on the memory and run by the processor, and the computer program, when run by the processor, executes the method for obtaining environmental information in unmanned driving according to any one of claims 1 to 3.

8. A storage medium, characterized in that, A computer program is stored on the storage medium, and the computer program, when running, executes the method for obtaining environmental information in unmanned driving according to any one of claims 1 to 3.