Three-dimensional sensing method and device and computer equipment
Through the multi-scale top-view feature map fusion method, multi-scale feature maps are obtained using binocular cameras and 4D millimeter wave radars, and through deep learning and transformer processing, the problem of insufficient fineness and comprehensiveness of 3D environment perception in the prior art is solved, and stable and high-precision perception in severe weather conditions is achieved.
Patent Information
- Application Number
- CN202510238257.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-28
- Publication Date
- 2025-07-04
AI Technical Summary
The existing 3D environment perception method that integrates 4D millimeter wave radar and binocular stereo cameras is not very fine and comprehensive when processing 3D information, and is sensitive to weather conditions, and is prone to performance degradation in bad weather.
The multi-scale top-view feature map fusion method is adopted, and the multi-scale top-view feature map is obtained through binocular cameras and 4D millimeter wave radars, and the data is aligned using timestamp synchronization and external parameter calibration technology, and features are extracted in combination with deep learning and voxel convolution networks, and the transformer's 3D detection head and occupancy head are fused to achieve the precise fusion of multi-scale features.
It improves the precision and comprehensiveness of environmental perception, enhances the system's stability and anti-interference ability in harsh weather conditions, and provides more reliable 3D environmental information.
Smart Images

Figure CN120259996A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and in particular, to a three-dimensional perception method, apparatus, and computer device. Background Art
[0002] With the rapid development of autonomous driving technology, the environmental perception system, as a core part thereof, has become the focus of research and application. Traditional environmental perception systems mainly rely on lidar and multiple surround-view cameras to obtain three-dimensional (3D) information, but there are problems such as high cost and susceptibility to weather conditions. The fusion of 4D millimeter-wave radar and binocular stereo cameras (also referred to as "binocular cameras" in this article) can make up for these deficiencies and provide more comprehensive and reliable environmental perception capabilities.
[0003] However, in the 3D environmental perception method that fuses 4D millimeter-wave radar and binocular stereo cameras, after extracting the 3D information of the point cloud data and the image, the processing process of the 3D information is relatively simple, and fusion detection is performed after simple coordinate transformation, resulting in low fineness and comprehensiveness of perception. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a three-dimensional perception method, apparatus, and computer device, which can capture and process environmental information at different scales, and improve the fineness and comprehensiveness of perception.
[0005] In a first aspect, an embodiment of the present invention provides a three-dimensional perception method, the method comprising: Obtaining image information collected by a binocular camera, and obtaining a first multi-scale top-down feature map according to the image information; Obtaining point cloud data collected by a 4D millimeter-wave radar, and obtaining a second multi-scale top-down feature map according to the point cloud data; Fusing the features in the first multi-scale top-down feature map and the second multi-scale top-down feature map to obtain a fused top-down view feature; Performing three-dimensional perception on the fused top-down view feature to obtain a perception result.
[0006] Optionally, the method further comprises: Using timestamp synchronization technology to align the data collected by the binocular camera and the 4D millimeter-wave radar in time; Using extrinsic parameter calibration technology to align the data collected by the binocular camera and the 4D millimeter-wave radar in space.
[0007] Optionally, the obtaining a first multi-scale top-down feature map according to the image information comprises: Obtain first 3D information based on the image information, where the first 3D information includes a 3D voxel volume with semantic information; Divide the first 3D information into first 3D voxel volumes of different scales; Extract the features of the first 3D voxel volumes of different scales; Convert the first 3D voxel volumes of different scales and their features to the top view space to obtain the first multi-scale top view feature map.
[0008] Optionally, the obtaining of the first 3D information based on the image information includes Using a depth prediction network to perform depth estimation on the image information to generate a depth map and converting the depth map into the first 3D information.
[0009] Optionally, the obtaining of the second multi-scale top view feature map based on the point cloud data includes: Obtain second 3D information based on the point cloud data, where the second 3D information includes a 3D voxel volume with spatial information; Divide the second 3D information into second 3D voxel volumes of different scales; Extract the features of the second 3D voxel volumes of different scales; Convert the second 3D voxel volumes of different scales and their features to the top view space to obtain the second multi-scale top view feature map.
[0010] Optionally, the obtaining of the second 3D information based on the point cloud data includes: Using a voxelization method to convert the point cloud data into the second 3D information.
[0011] Optionally, the performing of 3D perception on the fused top view features to obtain a perception result includes: Using a 3D detection head based on a transformer to perform object detection on the fused top view features to obtain an object detection result; and / or Using a 3D occupancy head based on a transformer to perform environmental occupancy perception on the fused top view features to obtain the environmental occupancy situation.
[0012] On the other hand, an embodiment of the present invention provides a 3D perception device, and the device includes: An acquisition module, configured to acquire image information collected by a binocular camera and point cloud data collected by a 4D millimeter-wave radar; A first feature generation module, configured to obtain a first multi-scale top view feature map based on the image information; A second feature generation module, configured to obtain a second multi-scale top view feature map based on the point cloud data; A feature fusion module, configured to fuse the features in the first multi-scale top-down feature map and the second multi-scale top-down feature map to obtain a fused top-down view feature; A 3D perception module, configured to perform 3D perception on the fused top-down view feature to obtain a perception result.
[0013] On the other hand, an embodiment of the present invention provides a storage medium, which includes a stored program. When the program runs, it controls the device where the storage medium is located to execute the above method.
[0014] On the other hand, an embodiment of the present invention provides a computer device, including a memory and a processor. The memory is used to store information including program instructions, and the processor is used to control the execution of the program instructions. When the program instructions are loaded and executed by the processor, the steps of the above method are implemented.
[0015] In the technical solutions of the 3D perception method, device, and computer device provided by the embodiments of the present invention, the method includes: obtaining image information collected by a binocular camera, and obtaining a first multi-scale top-down feature map according to the image information; obtaining point cloud data collected by a 4D millimeter-wave radar, and obtaining a second multi-scale top-down feature map according to the point cloud data; fusing the features in the first multi-scale top-down feature map and the second multi-scale top-down feature map to obtain a fused top-down view feature; performing 3D perception on the fused top-down view feature to obtain a perception result. It can capture and process environmental information at different scales, improving the fineness and comprehensiveness of perception. Description of the Drawings
[0016] Figure 1 It is a flowchart of a 3D perception method provided by an embodiment of the present invention; Figure 2 It is a flowchart of obtaining a first multi-scale top-down feature map according to image information provided by an embodiment of the present invention; Figure 3 It is a flowchart of obtaining a second multi-scale top-down feature map according to point cloud data provided by an embodiment of the present invention; Figure 4 It is a flowchart of 3D environment perception based on image information and point cloud data in an embodiment of the present invention; Figure 5 It is a schematic structural diagram of a 3D perception device provided by an embodiment of the present invention; Figure 6 It is a schematic diagram of a computer device provided by an embodiment of the present invention. Detailed Embodiments
[0017] To better understand the technical solutions of the present invention, the embodiments of the present invention will be described in detail below with reference to the drawings.
[0018] It should be clear that the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments in the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.
[0019] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "the" and "said" used in the embodiments of the present invention and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0020] It should be understood that the term " / and" used herein is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this article generally represents an "or" relationship between the associated objects before and after.
[0021] In recent years, with the rapid development of autonomous driving technology, the environmental perception system, as its core component, has received extensive attention and research. Autonomous driving requires the vehicle to be able to perceive the surrounding environment, including obstacles such as roads, pedestrians, and vehicles.
[0022] The environmental perception system needs to obtain 3D information of the surrounding environment in real time and accurately to support functions such as path planning, obstacle detection, and obstacle avoidance.
[0023] Traditional environmental perception systems mainly rely on sensors such as lidar (Light Detection and Ranging, LiDAR) and cameras to obtain 3D environmental information. However, the cost of lidar is relatively high, and it is easily affected by weather conditions (such as rain, fog, snow) and lighting conditions (such as strong light, low light). Cameras obtain 2D and 3D information through image processing technology, but perform poorly under lighting changes and occlusion; although cameras have high resolution, it is difficult to obtain depth information.
[0024] As an emerging sensor technology, 4D millimeter-wave radar can provide distance, speed, and angle information, and has the advantages of strong penetration, strong anti-interference ability, and all-weather operation. In contrast, 4D millimeter-wave radar is not as high in resolution and accuracy as lidar, but it has a lower cost and performs well in adverse weather conditions. Binocular stereo cameras obtain 3D depth information by calculating the disparity between two images and can provide high-resolution visual data, but there are certain limitations in detecting distant targets and in adverse weather conditions. Therefore, fusing 4D millimeter-wave radar and binocular stereo cameras can make full use of the advantages of both sensors, make up for their respective deficiencies, and thus provide a more comprehensive and reliable environmental perception ability.
[0025] However, after extracting the point cloud data and 3D information of the images, the 3D environmental perception method that fuses 4D millimeter-wave radar and binocular stereo cameras has a relatively simple process for processing the 3D information and performs fusion detection after simple coordinate transformation, resulting in low fineness and comprehensiveness of perception.
[0026] Based on the above technical problems, an embodiment of the present invention provides a three-dimensional perception method.
[0027] Figure 1 The flowchart of a three-dimensional perception method provided by an embodiment of the present invention is as Figure 1 shown, and the method includes: Step 101, obtain the image information collected by the binocular camera, and obtain the first multi-scale top-down feature map according to the image information.
[0028] In an embodiment of the present invention, the image information includes the left image and the right image captured by the binocular camera.
[0029] In some possible embodiments, as Figure 2 shown, obtaining the first multi-scale top-down feature map according to the image information includes: Step 1011, obtain the first 3D information according to the image information, and the first 3D information includes a 3D voxel volume with semantic information.
[0030] In some possible embodiments, step 1011 includes: using a depth prediction network to perform depth estimation on the image information to generate a depth map and converting the depth map into the first 3D information.
[0031] For the data of the binocular stereo camera, use a depth prediction network to perform depth estimation, generate a dense depth map, and convert it into a 3D voxel volume.
[0032] In an embodiment of the present invention, depth estimation adopts a learnable method, is trained and predicted through a depth prediction network, and calculates the disparity between two images to obtain 3D depth information.
[0033] For example, the depth prediction network can be a Convolutional Neural Network (CNN).
[0034] For example, the learnable depth prediction formula is as follows: Wherein, represents the predicted depth, represents the depth value obtained at the predefined depth level , is the depth level interval, represents the probability that each feature pixel is at the depth level. After obtaining the predicted depth, supervised learning is performed using the true depth of the dataset to continuously correct the network parameters to obtain more accurate depth information.
[0035] Step 1012: Divide the first 3D information into first 3D voxel volumes of different scales.
[0036] The obtained 3D information is divided into 3D voxel volumes of different scales. The 3D voxel volumes of larger scales are used to describe the macroscopic structure in the environment, and the 3D voxel volumes of smaller scales are used to capture detailed information.
[0037] Step 1013: Extract the features of the first 3D voxel volumes of different scales.
[0038] The 3D voxel volumes of different scales are subjected to voxel segmentation to extract the key features (such as shape, size, density, etc.) in the voxel blocks. A voxel convolutional neural network (VCNN) is used to extract the features of the voxel volumes to extract semantic features.
[0039] Wherein, the voxel convolution formula is as follows: Wherein, represents the feature after convolution, represents the convolutional kernel weight, represents adjacent voxels.
[0040] Step 1014: Convert the first 3D voxel volumes of different scales and their features to the top view space to obtain a first multi-scale top view feature map.
[0041] Design and implement an algorithm for converting 3D voxel volumes to the top view space. By projecting the multi-scale 3D voxel volumes onto the top view plane, a multi-scale top view feature map is generated.
[0042] Wherein, the top view space conversion formula is as follows: Among them, represents the coordinates of the 3D voxel, represents the projected coordinates on the top view plane.
[0043] Step 102: Obtain the point cloud data collected by the 4D millimeter-wave radar, and obtain the second multi-scale top view feature map according to the point cloud data.
[0044] The 4D millimeter-wave radar obtains distance, speed, and angle information by transmitting millimeter waves and receiving reflected signals, and generates preliminary 3D point cloud data.
[0045] In some possible embodiments, such as Figure 3 shown, obtaining the second multi-scale top view feature map according to the point cloud data includes: Step 1021: Obtain the second 3D information according to the point cloud data, and the second 3D information includes a 3D voxel body with spatial information.
[0046] In some possible embodiments, step 1021 includes: converting the point cloud data into the second 3D information by using a voxelization method.
[0047] For the data of the 4D millimeter-wave radar, the point cloud data is converted into a 3D voxel body by using a voxelization method. The voxelization formula is as follows: Among them, represents the th voxel grid.
[0048] Step 1022: Divide the second 3D information into second 3D voxel bodies of different scales.
[0049] The obtained 3D information is divided into 3D voxel bodies of different scales. The 3D voxel bodies of larger scales are used to describe the macroscopic structure in the environment, and the 3D voxel bodies of smaller scales are used to capture detailed information.
[0050] Step 1023: Extract the features of the second 3D voxel bodies of different scales.
[0051] Perform voxel segmentation on the 3D voxel bodies of different scales, and extract the key features (such as shape, size, density, etc.) in the voxel blocks. Use a voxel convolutional neural network to extract features from the voxel body and extract spatial features.
[0052] Step 1024: Convert the second 3D voxel bodies of different scales and their features to the top view space to obtain the second multi-scale top view feature map.
[0053] In the embodiments of the present invention, the timestamp synchronization technology can be used to align the data collected by the binocular camera and the 4D millimeter-wave radar in time, ensuring the temporal consistency of the data of the two sensors. The extrinsic calibration technology is used to align the data collected by the binocular camera and the 4D millimeter-wave radar in space, ensuring the spatial alignment of the data of the two sensors and laying a foundation for subsequent fusion.
[0054] In the embodiments of the present invention, the generation and processing method of the multi-scale 3D voxel volume included in step 101 and step 102 enhances the multi-scale information processing ability of the perception system, enabling the perception system to capture and process environmental information at different scales, and improving the fineness and comprehensiveness of perception. The large-scale voxel volume is used for macroscopic structure description, and the small-scale voxel volume is used for detail capture. The combination of the two provides a complete environmental description.
[0055] Step 103: Fuse the features in the first multi-scale top-down feature map and the second multi-scale top-down feature map to obtain the fused top-down view feature.
[0056] In the top-down view space, the features of 3D voxel volumes at different scales are fused. A multi-scale feature fusion algorithm (such as a pyramid convolutional neural network) is used to combine information at different scales, extract and enhance key environmental features. The multi-scale feature fusion formula is as follows: Among them, represents the fused feature, represents the feature of the s-th scale, represents the fusion weight.
[0057] Step 104: Perform three-dimensional perception on the fused top-down view feature to obtain a perception result.
[0058] In some possible embodiments, step 104 includes: step 1041 and / or step 1042.
[0059] Step 1041: Use a 3D detection head based on a transformer to perform object detection on the fused top-down view feature to obtain an object detection result.
[0060] Design a 3D detection head based on a transformer, and use the self-attention mechanism of the transformer to perform object detection on the fused top-down view feature. The self-attention mechanism can capture long-range dependencies and improve the accuracy of object detection. The self-attention formula is as follows: Among them, represents the detection head query matrix, represents the key matrix from the top-down feature map, Represents the value matrix from the top-down feature map, Represents the dimension of the key vector, Represents the transpose of the matrix. Then, through the linear transformation layer, the detection results are output as follows: Step 1042: Use a 3D occupancy head based on the transformer to perform environmental occupancy perception on the fused top-down view features to obtain the environmental occupancy situation.
[0061] Design a 3D occupancy head based on the transformer. By further analyzing the fused top-down view features, it perceives the occupancy situation in the environment. The self-attention mechanism is used to judge the drivable areas and obstacle distributions in the space, improving the accuracy of occupancy perception. The occupancy perception formula is similar to that of the detection head, and feature extraction and occupancy judgment are performed through the self-attention mechanism and linear transformation.
[0062] In summary, Figure 4 This is the flowchart of 3D environmental perception based on image information and point cloud data in the embodiments of the present invention. As Figure 4 shown, first, use the timestamp synchronization technology to ensure the temporal consistency of the data from the two sensors; adopt the extrinsic calibration technology to ensure the spatial alignment of the data from the two sensors, laying a foundation for subsequent fusion. The binocular stereo camera captures two left and right images. The 4D millimeter-wave radar obtains distance, speed, and angle information by emitting millimeter waves and receiving reflected signals, generating preliminary 3D point cloud data, that is, radar point cloud. Then, it is divided into two branches to process the left and right images and the radar point cloud respectively to obtain multi-scale top-down view features.
[0063] As Figure 4 shown, in the left and right image branches, first, perform data preprocessing on the left and right images, then input them into the deep learnable network to obtain a multi-scale 3D voxel volume with semantic information, and then obtain multi-scale top-down view features based on the multi-scale 3D voxel volume with semantic information.
[0064] Among them, the data preprocessing of the left and right images can include pixel and size processing of the images.
[0065] Among them, after obtaining the predicted depth through the deep learnable network, a depth supervision head can be used to perform supervised learning through the true depth of the dataset, continuously correcting the network parameters to obtain more accurate depth information.
[0066] As Figure 4 shown, in the radar point cloud branch, first, perform data preprocessing on the radar point cloud, then input it into the point cloud voxelization network to obtain a multi-scale 3D voxel volume with spatial information, and then obtain multi-scale top-down view features based on the multi-scale 3D voxel volume with spatial information.
[0067] Among them, the data preprocessing of the radar point cloud may include filtering out invalid point cloud data through velocity filtering.
[0068] As Figure 4 shown, the multi-scale top view features obtained from the left and right image branches and the radar point cloud branch are feature fused, and then input into a multi-task 3D perception head including a 3D detection head based on a transformer and a 3D occupancy head based on a transformer to obtain a perception result.
[0069] Currently, multi-sensor based 3D environmental perception methods rely heavily on costly lidar and surround-view cameras, are very sensitive to weather conditions, and the perception effect is greatly reduced in bad weather; moreover, the failure of one of the sensors will lead to the collapse and paralysis of the entire perception system. However, the embodiment of the present invention utilizes a 4D millimeter-wave radar and a binocular camera. By separately processing the point cloud and image data streams, the system cost is well reduced, the ability to cope with bad weather is improved, and the anti-interference performance of the system is enhanced. By comprehensively utilizing the advantages of the millimeter-wave radar and the stereo camera, the embodiment of the present invention can work stably under complex weather and lighting conditions, cope with the situation of the failure of one of the sensors, and improve the practicality and adaptability of the system. The millimeter-wave radar can still provide reliable information under bad weather conditions, while the stereo camera provides high-precision data under normal weather conditions.
[0070] The embodiment of the present invention provides a multi-task 3D perception method based on dual-stream multi-scale top view representation, including: original sensor alignment, data acquisition, multi-modal data processing, etc. The embodiment of the present invention can provide more reliable and accurate 3D environmental perception under various environmental conditions by integrating the advantages of a 4D millimeter-wave radar and a binocular stereo camera; effectively fusing the 3D information obtained by the 4D millimeter-wave radar and the binocular stereo camera, combining the all-weather working ability of the millimeter-wave radar and the high-resolution characteristics of the stereo camera, enables the system to cope with complex environmental changes and improves the accuracy and robustness of environmental perception.
[0071] In the embodiment of the present invention, the conversion of the top view space and the multi-scale feature fusion enable the system to perform environmental analysis and target detection more efficiently. The multi-scale feature fusion algorithm improves the detection speed and accuracy and can achieve a balance between real-time performance and accuracy.
[0072] An embodiment of the present invention proposes a 3D object detection and 3D occupancy method based on a 4D millimeter-wave radar and a binocular stereo camera, which can obtain more accurate, robust, and comprehensive environmental perception results. The embodiment of the present invention designs and implements a 3D detection head and a 3D occupancy head based on a transformer to complete efficient 3D object detection and environmental occupancy perception tasks.
[0073] The embodiment of the present invention proposes to fuse a 4D millimeter-wave radar and a binocular stereo camera to achieve more comprehensive and reliable environmental perception. Utilize the all-weather working ability of the millimeter-wave radar and the high-resolution characteristics of the stereo camera to provide more accurate 3D environmental information. Innovatively adopt a method for generating and processing multi-scale 3D voxel volumes, and capture environmental features at different scales through hierarchical processing and feature extraction to improve the fineness and comprehensiveness of perception. Propose and implement an algorithm for converting from a 3D voxel volume to a top-view space, and perform multi-scale feature fusion in the top-view space, retaining and enhancing important environmental information during the fusion process. Adopt deep learning technology to process and enhance the top-view features, improving the efficiency and accuracy of environmental analysis and object detection.
[0074] The embodiment of the present invention extracts multi-scale 3D voxel volumes from sensor data and processes and fuses them without losing key information.
[0075] In the embodiment of the present invention, the method for generating and processing multi-scale 3D voxel volumes enhances the multi-scale information processing ability, enabling the system to capture and process environmental information at different scales, and improving the fineness and comprehensiveness of perception. Large-scale voxel volumes are used for macroscopic structure description, and small-scale voxel volumes are used for detail capture. The combination of the two provides a complete environmental description.
[0076] In the technical solution of a three-dimensional perception method provided by the embodiment of the present invention, the method includes: obtaining image information collected by a binocular camera, and obtaining a first multi-scale top-view feature map according to the image information; obtaining point cloud data collected by a 4D millimeter-wave radar, and obtaining a second multi-scale top-view feature map according to the point cloud data; fusing the features in the first multi-scale top-view feature map and the second multi-scale top-view feature map to obtain a fused top-view feature; performing three-dimensional perception on the fused top-view feature to obtain a perception result. It can capture and process environmental information at different scales, and improves the fineness and comprehensiveness of perception.
[0077] Figure 5 As shown in the structural schematic diagram of a three-dimensional perception device provided by the embodiment of the present invention, Figure 5 The device includes: An acquisition module 11, configured to acquire image information collected by a binocular camera and point cloud data collected by a 4D millimeter-wave radar; The first feature generation module 12 is configured to obtain a first multi-scale top-down feature map based on the image information; The second feature generation module 13 is configured to obtain a second multi-scale top-down feature map based on the point cloud data; The feature fusion module 14 is configured to fuse the features in the first multi-scale top-down feature map and the second multi-scale top-down feature map to obtain the fused top-down view feature; The 3D perception module 15 is configured to perform 3D perception on the fused top-down view feature to obtain a perception result.
[0078] In some possible embodiments, the apparatus further includes: a spatio-temporal alignment module 16; the spatio-temporal alignment module 16 is configured to align the data collected by the binocular camera and the 4D millimeter-wave radar in time by using a timestamp synchronization technique; and align the data collected by the binocular camera and the 4D millimeter-wave radar in space by using an extrinsic calibration technique.
[0079] In some possible embodiments, the first feature generation module 12 is specifically configured to obtain first 3D information based on the image information, where the first 3D information includes a 3D voxel volume with semantic information; divide the first 3D information into first 3D voxel volumes of different scales; extract the features of the first 3D voxel volumes of different scales; and convert the first 3D voxel volumes of different scales and their features to the top-down view space to obtain a first multi-scale top-down feature map.
[0080] In some possible embodiments, the first feature generation module 12 is specifically configured to perform depth estimation on the image information by using a depth prediction network to generate a depth map and convert the depth map into first 3D information.
[0081] In some possible embodiments, the second feature generation module 13 is specifically configured to obtain second 3D information based on the point cloud data, where the second 3D information includes a 3D voxel volume with spatial information; divide the second 3D information into second 3D voxel volumes of different scales; extract the features of the second 3D voxel volumes of different scales; and convert the second 3D voxel volumes of different scales and their features to the top-down view space to obtain a second multi-scale top-down feature map.
[0082] In some possible embodiments, the second feature generation module 13 is specifically configured to convert the point cloud data into second 3D information by using a voxelization method.
[0083] In some possible embodiments, the 3D perception module 15 is specifically configured to perform object detection on the fused top-down view feature by using a 3D detection head based on a transformer to obtain an object detection result; and / or perform environment occupancy perception on the fused top-down view feature by using a 3D occupancy head based on a transformer to obtain the environment occupancy situation.
[0084] In the technical solution provided by the embodiment of the present invention, image information collected by a binocular camera is obtained, and a first multi-scale top-down feature map is obtained according to the image information; point cloud data collected by a 4D millimeter-wave radar is obtained, and a second multi-scale top-down feature map is obtained according to the point cloud data; features in the first multi-scale top-down feature map and the second multi-scale top-down feature map are fused to obtain a fused top-down view feature; and a three-dimensional perception is performed on the fused top-down view feature to obtain a perception result. It can capture and process environmental information at different scales, improving the fineness and comprehensiveness of perception.
[0085] An embodiment of the present application provides a storage medium, which includes a stored program. When the program runs, it controls the device where the storage medium is located to execute the above method.
[0086] An embodiment of the present application provides a computer device, including a memory and a processor. The memory is used to store information including program instructions, and the processor is used to control the execution of the program instructions. When the program instructions are loaded and executed by the processor, the steps of the above method are implemented.
[0087] Figure 6 It is a schematic structural diagram of a computer device provided by an embodiment of the present application. As Figure 6 shown, the computer device 20 includes: a processor 21, a memory 22, and a computer program 23 stored in the memory 22 and executable on the processor 21. When the computer program 23 is executed by the processor 21, it implements the application to the energy management method in the embodiment. To avoid repetition, it will not be elaborated here one by one.
[0088] The computer device 20 includes, but is not limited to, a processor 21 and a memory 22. Those skilled in the art can understand that Figure 6 it is only an example of the computer device 20 and does not constitute a limitation on the computer device 20. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the computer device 20 may further include input / output devices, network access devices, buses, etc.
[0089] The so-called processor 21 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0090] The memory 22 may be an internal storage unit of the computer device 20, such as the hard disk or memory of the computer device 20. The memory 22 may also be an external storage device of the computer device 20, such as a plug-in hard disk, Smart Media Card (SMC), Secure Digital (SD) card, Flash Card, etc. equipped on the computer device 20. Further, the memory 22 may also include both the internal storage unit and the external storage device of the computer device 20. The memory 22 is used to store computer programs and other programs and data required by the computer device 20. The memory 22 may also be used to temporarily store data that has been output or is to be output.
[0091] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.
[0092] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces, and the indirect coupling or communication connection of the devices or units may be in an electrical, mechanical, or other form.
[0093] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0094] In addition, each functional unit in various embodiments of the present application may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a hardware plus software functional unit.
[0095] The above-mentioned integrated unit implemented in the form of a software functional unit can be stored in a computer-readable storage medium. The above-mentioned software functional unit stored in a storage medium includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute some steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.
[0096] The foregoing is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included within the scope of protection of the present application.
Claims
1. A three-dimensional perception method, characterized in that, The method includes: Obtaining the image information collected by the binocular camera, and obtaining the first multi-scale top-down feature map according to the image information; Obtaining the point cloud data collected by the 4D millimeter-wave radar, and obtaining the second multi-scale top-down feature map according to the point cloud data; Fusing the features in the first multi-scale top-down feature map and the second multi-scale top-down feature map to obtain the fused top-down view feature; Performing three-dimensional perception on the fused top-down view feature to obtain a perception result.
2. The method according to claim 1, characterized in that The method further includes: Using a timestamp synchronization technology to align the data collected by the binocular camera and the 4D millimeter-wave radar in time; Using an extrinsic calibration technology to align the data collected by the binocular camera and the 4D millimeter-wave radar in space.
3. The method according to claim 1 or 2, characterized in that, The obtaining the first multi-scale top-down feature map according to the image information includes: Obtaining the first 3D information according to the image information, where the first 3D information includes a 3D voxel volume with semantic information; Dividing the first 3D information into first 3D voxel volumes of different scales; Extracting the features of the first 3D voxel volumes of different scales; Converting the first 3D voxel volumes of different scales and their features to the top-down view space to obtain the first multi-scale top-down feature map.
4. The method according to claim 3, wherein The obtaining the first 3D information according to the image information includes Using a depth prediction network to perform depth estimation on the image information to generate a depth map and converting the depth map to the first 3D information.
5. The method according to claim 1 or 2, characterized in that The obtaining the second multi-scale top-down feature map according to the point cloud data includes: Obtaining the second 3D information according to the point cloud data, where the second 3D information includes a 3D voxel volume with spatial information; Dividing the second 3D information into second 3D voxel volumes of different scales; Extracting the features of the second 3D voxel volumes of different scales; Converting the second 3D voxel volumes of different scales and their features to the top-down view space to obtain the second multi-scale top-down feature map.
6. The method according to claim 5, wherein The obtaining the second 3D information according to the point cloud data includes: Using a voxelization method to convert the point cloud data to the second 3D information.
7. The method according to claim 1, wherein The performing three-dimensional perception on the fused top-down view feature to obtain a perception result includes: Using a 3D detection head based on a transformer to perform object detection on the fused top-down view feature to obtain an object detection result; and / or Using a 3D occupancy head based on a transformer to perform environment occupancy perception on the fused top-down view feature to obtain the environment occupancy situation.
8. A three-dimensional sensing device, characterized in that, It includes: An acquisition module, configured to acquire the image information collected by the binocular camera and the point cloud data collected by the 4D millimeter-wave radar; A first feature generation module, configured to obtain the first multi-scale top-down feature map according to the image information; A second feature generation module, configured to obtain the second multi-scale top-down feature map according to the point cloud data; A feature fusion module, configured to fuse the features in the first multi-scale top-down feature map and the second multi-scale top-down feature map to obtain the fused top-down view feature; A three-dimensional perception module, configured to perform three-dimensional perception on the fused top-down view feature to obtain a perception result.
9. A storage medium, characterized in that, The storage medium includes a stored program, wherein when the program runs, it controls the device where the storage medium is located to execute the three-dimensional perception method according to any one of claims 1 to 7.
10. A computer device, comprising a memory and a processor, the memory being used for storing information including program instructions, and the processor being used for controlling the execution of the program instructions, characterized in that, When the program instructions are loaded and executed by a processor, the steps of the three-dimensional perception method according to any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Multi-view 3D perception method based on space-time modeling and context enhancement
CN121354061A
Three-dimensional perception method and apparatus, and computer device
WO2026179677A1