Multi-modal 3D detection method and system fused with depth completion, and medium

By generating dense depth maps and combining weighted fusion strategies with image edge adjustment weights, the problem of sparse depth information of low-cost radar is solved, improving the accuracy and accuracy of 3D target detection, especially providing reliable information support in complex environments.

CN120356171APending Publication Date: 2025-07-22NINGBO INST OF MATERIALS TECH & ENG CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510424937.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

In the existing multimodal 3D object detection methods, the depth information obtained by low-cost radar is too sparse, resulting in the lack of depth information and inaccurate depth estimation of edge parts, which affects the accuracy and accuracy of 3D object detection, and cannot provide sufficient information support especially in complex environments.

Method used

By acquiring image data for depth estimation, a depth estimation map is generated, and a dense depth map is generated using sparse depth map constraints. A weighted fusion strategy is used to adjust the weight with the image edge distance to generate a high-quality fusion depth map, and feature fusion with the sparse depth map feature map, and finally object detection is performed.

Benefits of technology

The accuracy of 3D object detection is improved, the depth errors and edge blur defects of sparse depth maps are compensated, and efficient 3D object detection is achieved in complex environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120356171A_ABST
    Figure CN120356171A_ABST
Patent Text Reader

Abstract

The invention provides a multi-modal 3D detection method and system fused with depth completion, and a medium. The method comprises the following steps: acquiring image data and sparse depth data; obtaining a depth estimation map and a sparse depth map; constraining the depth estimation map to obtain a dense depth map; carrying out weighted fusion to obtain a fused depth map; in the weighted fusion process, the weight of a pixel point in the depth estimation image is inversely correlated with the distance of the edge of the object in the distance image; acquiring a first feature map by using the fused depth map, and acquiring a second feature map by using the sparse depth map; fusing to obtain a fused feature map; and performing target detection based on the fused feature map to obtain 3D environment information. According to the technical scheme provided by the invention, visible light depth estimation is applied to depth completion of the sparse depth map, the defects of regional depth errors and edge blurring existing in direct fusion of the visible light image and the sparse depth map are overcome, and the 3D target detection effect is improved by introducing feature space fusion and fully utilizing multi-modal information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image recognition, and in particular, to a multi-modal 3D detection method, system and medium integrating depth completion. Background Art

[0002] Environmental perception is a core research area in autonomous driving systems and a key link to ensure driving safety, and its performance directly affects path planning, decision-making, and control effects. In environmental perception, 3D object detection, as an important task, aims to accurately identify key objects (such as obstacles) encountered by an autonomous driving vehicle during driving and extract key information such as their categories, shapes, and positions to provide necessary support for the vehicle's behavior decision-making. The multi-modal 3D object detection method aims to improve the performance of downstream tasks such as object classification and recognition, pose prediction, and speed estimation by fusing perceptual data from different sensors. By comprehensively utilizing the advantages of different sensors, this method realizes information complementarity, thereby enhancing the robustness and accuracy of detection. Therefore, fully integrating and utilizing multi-modal perceptual data is crucial for improving the overall performance of 3D object detection.

[0003] Currently, the research on multi-modal 3D object detection mainly focuses on the fusion detection of lidar and camera data. A monocular visible light camera can obtain rich color, texture and other information in the scene, while lidar can provide three-dimensional information such as depth. The depth information obtained by low-cost lidar is often sparse. If denser depth information is required, a lidar with a higher cost is needed, which greatly limits its popular application in intelligent driving vehicles. Therefore, the existing 3D object detection methods have the problem that the sparse depth information obtained by low-cost radar restricts the accuracy of 3D object detection.

[0004] To address the problem of overly sparse radar depth information, in the prior art, there are some depth completion networks based on deep learning. These networks usually directly fuse only visible light images with sparse depth maps. However, sparse depth maps often have significant information losses, especially in the distant regions of the scene or the edge parts of complex objects, and these missing information cannot be effectively completed by simple visible light images. Due to the large blanks or errors in the depth values of the depth map itself in certain regions, this direct fusion method may not be able to accurately restore the depth information, leading to the following problems: First, the missing information in the sparse depth map may not be accurately predicted, resulting in depth errors; Second, the edge parts of objects in the scene may not be accurately estimated in depth, affecting the precise positioning and recognition of objects. The above defects significantly reduce the accuracy and precision of the 3D object detection task, especially in complex environments and high-dynamic scenarios, and cannot provide sufficient reliable information support, thus affecting the overall performance and safety of the autonomous driving system. Therefore, there is a large room for improvement in the prior art in the multi-modal 3D object detection that fuses depth completion. Summary of the Invention

[0005] Aiming at the deficiencies of the prior art, the purpose of the present invention is to provide a multi-modal 3D detection method, system and medium that fuse depth completion.

[0006] To achieve the aforementioned invention purpose, the technical solutions adopted by the present invention include:

[0007] In the first aspect, the present invention provides a multi-modal 3D detection method that fuses depth completion, which includes:

[0008] Obtain image data and sparse depth data;

[0009] Perform depth estimation on the image data to obtain a depth estimation map, and perform coordinate projection on the sparse depth data to obtain a sparse depth map;

[0010] Use the sparse depth map to constrain the depth estimation map to obtain a dense depth map;

[0011] Weightedly fuse the dense depth map and the depth estimation map to obtain a fused depth map; during the weighted fusion process, the weight of the pixel point in the depth estimation map is inversely correlated with the distance of the pixel point from the image edge;

[0012] Use the fused depth map to obtain a first feature map, and use the sparse depth map to obtain a second feature map;

[0013] Fuse the first feature map and the second feature map to obtain a fused feature map;

[0014] Perform object detection based on the fused feature map to obtain 3D environment information.

[0015] In a second aspect, the present invention further provides a multimodal 3D detection system with depth completion fusion, which includes:

[0016] A data acquisition module for acquiring image data and sparse depth data;

[0017] A data preprocessing module for performing depth estimation on the image data to obtain a depth estimation map, and performing coordinate projection on the sparse depth data to obtain a sparse depth map;

[0018] A dense depth module for using the sparse depth map to constrain the depth estimation map to obtain a dense depth map;

[0019] A fused depth module for weighted fusion of the dense depth map and the depth estimation map to obtain a fused depth map; during the weighted fusion process, the weight of a pixel point in the depth estimation map is inversely correlated with the distance of the pixel point from the image edge;

[0020] A feature map module for obtaining a first feature map using the fused depth map and obtaining a second feature map using the sparse depth map;

[0021] A fused feature module for fusing the first feature map and the second feature map to obtain a fused feature map;

[0022] An object detection module for performing object detection based on the fused feature map to obtain 3D environment information.

[0023] In a third aspect, the present invention further provides a readable storage medium, in which a computer program is stored, and when the computer program is run, the steps of the above multimodal 3D detection method are executed.

[0024] Based on the above technical solutions, compared with the prior art, the beneficial effects of the present invention at least include:

[0025] The multimodal 3D detection method provided by the present invention improves the accuracy of multimodal 3D object detection by introducing depth estimation information of image data to complete the sparse depth map; specifically, first, a scale-free dense depth estimation map is generated from a visible light image, and it is fused with an absolute scale sparse depth map in a designed fusion network to obtain an absolute scale dense depth map, thereby generating a high-quality bird's-eye view feature map of the visible light image and efficiently fusing it with the feature map of the lidar point cloud, and finally achieving accurate 3D object detection.

[0026] Compared with the prior art, the technical solution proposed by the present invention applies visible light depth estimation to the depth completion of sparse depth maps, making up for the defects of direct fusion of visible light images and sparse depth maps, such as regional depth errors and edge blurring. Moreover, by introducing feature space fusion, multi-modal information is fully utilized to improve the 3D object detection effect.

[0027] The above description is only an overview of the technical solution of the present invention. In order to enable those skilled in the art to more clearly understand the technical means of the present application and implement it in accordance with the content of the specification, the following is a detailed description of the preferred embodiments of the present invention in conjunction with the detailed drawings. Brief Description of the Drawings

[0028] Figure 1 It is a schematic diagram of the overall process of the multi-modal 3D detection method provided by a typical embodiment of the present invention;

[0029] Figure 2 It is an architecture diagram of the depth completion system provided by a typical embodiment of the present invention;

[0030] Figure 3 It is a schematic diagram of the module architecture and process of the object detection method provided by a typical embodiment of the present invention. Detailed Embodiments

[0031] In view of the deficiencies in the prior art, the inventors of this case have proposed the technical solution of the present invention through long-term research and a large number of practices. The following will further explain the technical solution, its implementation process, principles, etc.

[0032] In the following description, many specific details are set forth in order to provide a thorough understanding of the present invention. However, the present invention may be implemented in other ways different from those described herein. Therefore, the scope of protection of the present invention is not limited by the specific embodiments disclosed below.

[0033] Moreover, relational terms such as "first" and "second" are only used to distinguish one component or method step with the same name from another, and do not necessarily require or imply any actual relationship or order between these components or method steps.

[0034] The main objective of the present invention is to provide a multimodal 3D object detection method integrating depth completion, which solves the problems of missing depth information and insufficient estimation accuracy in the prior art by using the depth estimation information of visible light images to complete sparse depth maps. In a specific example, it is preferably to generate a high-quality bird's-eye view (BEV) feature map of the visible light image from the completed dense depth map and efficiently fuse it with the BEV feature map of the point cloud, significantly improving the accuracy of 3D object detection. It overcomes the defect in the prior art that the depth information obtained by low-cost radar is too sparse, thus restricting the accuracy of 3D object detection.

[0035] Based on the above objective, as shown in Figure 1 the embodiments of the present invention first provide a multimodal 3D detection method integrating depth completion, which includes the following steps:

[0036] Obtain image data and sparse depth data;

[0037] Perform depth estimation on the image data to obtain a depth estimation map, and perform coordinate projection on the sparse depth data to obtain a sparse depth map;

[0038] Use the sparse depth map to constrain the depth estimation map to obtain a dense depth map;

[0039] Weightedly fuse the dense depth map and the depth estimation map to obtain a fused depth map; during the weighted fusion process, the weight of a pixel point in the depth estimation map is inversely correlated with the distance of the pixel point from the image edge;

[0040] Use the fused depth map to obtain a first feature map, and use the sparse depth map to obtain a second feature map;

[0041] Fuse the first feature map and the second feature map to obtain a fused feature map;

[0042] Perform object detection based on the fused feature map to obtain 3D environment information.

[0043] In the above technical solution, preferably, the image data is visible light data obtained by a camera, but it is not limited thereto, and image capture data in other bands can also be used as image data; the sparse depth data refers to sparse point clouds obtained by lidar or other radars with similar principles; and the finally obtained 3D environment information can include various information such as the types, positions, postures, sizes, etc. of the objects around the camera and the radar.

[0044] Typical applications of the above technical solution lie in autonomous / assisted driving. For example, it is applied to the autonomous / assisted driving system of a vehicle, obtaining data through the combination of a camera and a lidar, identifying targets around the vehicle using the above detection method, and then achieving autonomous driving or assisted driving, etc. Of course, it is not limited to the above application scenarios, and any mechanical equipment that requires path planning and autonomous driving, such as drones, unmanned vehicles, and floor-sweeping robots, can be applied.

[0045] Regarding the specific steps, in some embodiments, the acquisition process of the depth estimation map and the sparse depth map can be expressed as:

[0046] D est = Dep(I)

[0047] P Lidar = {(x i , y i , z i )}

[0048] P cam = (x cam , y cam , z cam ) = T CL P Lidar

[0049]

[0050] D sparse (u i , v i ) = z cam

[0051] Wherein, D est represents the depth estimation map, Dep(·) represents the depth estimation model, I represents the image data, P Lidar represents the sparse depth data, x i , y i represents the coordinates of the pixel points in the sparse depth data, z i represents the depth value in the sparse depth data, i represents the pixel index, P cam represents the projection point, x cam , y cam represents the coordinates of the projection point, z cam represents the depth value of the projection point, T CL represents the transformation matrix that projects the pixel points of the sparse depth data into the coordinate system of the image data, u i , v i represents the coordinates of the pixel points in the sparse depth map, f x and f yrespectively represent the ratio of the focal length of the camera that captures the image data to the physical size of a single pixel along the x-axis or y-axis, i.e., where f is the physical focal length of the lens, d x and d y are the physical sizes of a single pixel of the camera sensor in the x and y directions, c x and c y represent the optical center coordinates of the camera, D sparse represents the sparse depth map.

[0052] In some embodiments, the generation process of the dense depth map may specifically include the following processes:

[0053] For the pixel points existing in the sparse depth map,

[0054] D dense (u i , v i ) = D sparse (u i , v i )

[0055] For the pixel points not existing in the sparse depth map,

[0056]

[0057] ΔD est (u, v) = D est (u, v) - D est (u′, v′)

[0058]

[0059] where D d e nse (·) represents the dense depth map, u, v represent the coordinates in the dense depth map, (u′, v′) represents the adjacent coordinate points that are the nearest neighbors to the coordinate point (u, v) in the dense depth map and have completed the constraint, ΔD est (u, v) represents the offset of the depth estimation map corresponding to the coordinate point (u, v), S = {(u i , v i ) ≠ 0} represents the set of non-empty sparse pixel points, and ||·||2 represents the Euclidean distance.

[0060] In the above steps, first perform the nearest neighbor search, then calculate the depth offset of (u, v) relative to the nearest neighbor point (u′, v′) based on the local geometric consistency of the depth estimation map D est , and finally use the absolute depth value D sparse(u′, v′), convert the relative offset to the absolute depth. Through this step, ensure that the completed depth map D dense has an accurate absolute scale globally.

[0061] In some embodiments, the process of obtaining the fused depth map can be expressed as:

[0062] D dep = αD dense +(1 - α)D est

[0063] where D dep represents the fused feature map, α represents the weight of the dense depth map, and 1 - α represents the weight of the depth estimation map.

[0064] According to the multi-modal 3D detection method described in claim 4, characterized in that the weight adjustment process of the weighted fusion includes:

[0065] Calculate the gradient magnitude G(u, v) of the depth estimation map D est :

[0066]

[0067] Use the Canny edge detection algorithm to extract the edge region:

[0068] E = {(u e , v e )|G(u e , v e ) > τ}

[0069] where τ represents a preset gradient threshold. In the embodiments of the present invention, τ = 0.1, but is not limited thereto. This parameter can be appropriately determined through conditional experiments. E represents the edge region, and (u e , v e ) represents the pixel coordinates in the edge region;

[0070] Dynamic weight calculation:

[0071] For each position (u, v), calculate the pixel distance d(u, v) to the nearest edge:

[0072]

[0073] Dynamically adjust the weight α(u, v) according to the distance:

[0074]

[0075] where is the Sigmoid function, and γ represents the attenuation coefficient, which controls the rate of change of the weight with distance. In the embodiments of the present invention, γ = 0.5, but it is not limited thereto.

[0076] In some embodiments, the generation process of the first feature map can be expressed as:

[0077]

[0078] Z = D dep (u, v)

[0079] I bev = f(X, Y, Z)

[0080] where I bev represents the first feature map, f(·) represents the conversion function for projecting the three-dimensional point cloud onto the BEV plane, and X, Y, Z represent the coordinates of the three-dimensional point cloud;

[0081] The generation process of the first feature map is expressed as:

[0082] P bev = V(P Lidar )

[0083] where P bev represents the second feature map, and V(·) represents the function for discretizing the three-dimensional point cloud into a two-dimensional grid and extracting the feature information within each grid.

[0084] In some embodiments, the generation process of the fused feature map can be expressed as:

[0085] F bev = T(I bev , P bev )

[0086] where F bev represents the fused feature map, and T(·) represents the fusion function with a Transformer structure.

[0087] In some embodiments, the process of feature detection can be expressed as:

[0088] Result = Det(F bev )

[0089] where Result represents the 3D environmental information, and Det(·) represents the object detection network.

[0090] Corresponding to the above method, the second aspect of the embodiments of the present invention further provides a multi-modal 3D detection system for fusion depth completion, which includes:

[0091] A data acquisition module for acquiring image data and sparse depth data;

[0092] A data preprocessing module for performing depth estimation on the image data to obtain a depth estimation map, and performing coordinate projection on the sparse depth data to obtain a sparse depth map;

[0093] A dense depth module for using the sparse depth map to constrain the depth estimation map to obtain a dense depth map;

[0094] A fused depth module for weighted fusion of the dense depth map and the depth estimation map to obtain a fused depth map; during the weighted fusion process, the weight of a pixel point in the depth estimation map is inversely correlated with the distance of the pixel point from the image edge;

[0095] A feature map module for obtaining a first feature map using the fused depth map and obtaining a second feature map using the sparse depth map;

[0096] A fused feature module for fusing the first feature map and the second feature map to obtain a fused feature map;

[0097] An object detection module for performing object detection based on the fused feature map to obtain 3D environment information.

[0098] An embodiment of the present invention also provides a readable storage medium, in which a computer program is stored, and when the computer program is run, it executes the steps of the multi-modal 3D detection method provided in any of the above embodiments.

[0099] The technical solution of the present invention will be further described in detail below through several embodiments in conjunction with the drawings. However, the selected embodiments are only used to illustrate the present invention and do not limit the scope of the present invention.

[0100] Embodiment 1

[0101] As Figure 1 shown, the multi-modal 3D object detection method for fused depth completion provided in this embodiment includes the following steps:

[0102] (1) Preprocess the visible light image data and the lidar point cloud data.

[0103] Before fused depth completion, the following preparatory data was created: a depth estimation map, using a large model-based monocular depth estimation method to perform depth estimation on a single visible light image I to obtain an unscaled dense depth image D est . The formula is as follows:

[0104] D est = Dep(I)

[0105] Where, is the input image, is the depth estimation map, where H×W are the height and width of the image respectively, and Dep is the depth estimation large model.

[0106] Sparse depth map: The transformation matrix T obtained by the joint calibration of the camera and lidar CL , projects the lidar point cloud P Lidar ={(x i , y i , z i )} into the camera coordinate system to obtain the projected point P cam =T CL P Lidar . Using the perspective relationship, the projected point is converted into the image pixel coordinates (u i , v i ), corresponding to the depth Z i , to obtain the sparse depth map D sparse . The conversion formula is as follows:

[0107] P cam =T CL P Lidar

[0108]

[0109] D spar s e (u i , v i ) = z cam

[0110] where f x and f y are the focal lengths of the camera, c x and c y are the optical center coordinates of the camera, and (x cam , y cam , z cam ) are the coordinates of the projected point in the camera coordinate system.

[0111] (2) Fusing the sparse depth map with depth estimation to generate a dense depth map.

[0112] As Figure 2 shown, the sparse depth map D sparse is used to constrain the depth estimation map D est , to generate a dense depth map D dense with absolute scale. The specific method is as follows:

[0113] For the positions (u i , v i ) where there are sparse depth points, directly assign values:

[0114] D dense (u i ,v i ) = D sparse (u i ,v i )

[0115] For the position (u, v) without sparse depth points, use the offset ΔD of the depth estimation map est (u, v) plus the recently constrained point D dense (u′, v′) for calculation:

[0116] D dense (u, v) = D dense (u′, v′) + ΔD est (u, v)

[0117] where, ΔD est (u, v) refers to the embodiment shown above and is determined by the nearest neighbor interpolation method to ensure a smooth transition of the offset of depth estimation.

[0118] In the edge region of the depth map, increase the confidence in the depth estimation map D est . Adopt a weighted fusion strategy to perform a second feature fusion on the depth estimation map and the preliminarily fused dense depth map to obtain the fused depth map D dep , to improve the depth accuracy in the edge region. The formula is as follows:

[0119] D dep = αD dense + (1 - α)D est

[0120] where, α is the weight coefficient, which is dynamically adjusted according to the characteristics of the edge region. The specific adjustment process is:

[0121] 1. Edge region detection:

[0122] Calculate the gradient magnitude of the depth estimation map D est :

[0123]

[0124] Use the Canny edge detection algorithm to extract the edge region:

[0125] E = {(u e , v e ) | G(u e , v e ) > T}

[0126] where, τ = 0.1 (preset gradient threshold)

[0127] 2. Dynamic weight calculation:

[0128] For each position (u, v), calculate the pixel distance d(u, v) to the nearest edge:

[0129]

[0130] Dynamically adjust the weight α(u, v) according to the distance:

[0131]

[0132] Among them, is the Sigmoid function, which is used for smooth transition of weights. γ = 0.5 is the attenuation coefficient, which controls the rate of weight change with distance.

[0133] (3) Extract features from the fused depth map and convert them into BEV features.

[0134] Utilize the fused depth map D dep , combined with the camera intrinsic matrix K. Convert the fused depth map into the first feature map I bev in the BEV view. The conversion process is as follows:

[0135]

[0136]

[0137] Z = D dep (u, v)

[0138] I bev = f(X, Y, Z)

[0139] Among them, f is the conversion function that projects the 3D point cloud onto the BEV plane to generate the BEV feature map.

[0140] (4) Extract features from the LiDAR point cloud and convert them into BEV features.

[0141] As Figure 3 shown, downsample and filter the LiDAR point P Lidar to extract high-quality point cloud. Convert the point cloud data into the second feature map P bev in the BEV view through the rasterization conversion method. The formula is as follows:

[0142] P bev = V(P Lidar )

[0143] Among them, V discretizes the 3D point cloud into a 2D grid and extracts the feature information within each grid.

[0144] (5) Fuse the BEV features of the fused image and the LiDAR point cloud.

[0145] Continue to refer to Figure 3 As shown, the Transformer structure is used to perform feature fusion on the first image BEV feature map I bev and the second lidar BEV feature map P bev to generate the fused BEV fused feature map F bev . The specific process is as follows:

[0146] F bev = T(I bev , P bev )

[0147] The function T is existing. Refer to the existing BEVfuison, including the multi-head self-attention mechanism and the feed-forward neural network. The parameters include the query weight matrix W q , the key weight matrix W k , the value weight matrix W v , etc. Through the self-attention mechanism, T can effectively capture the correlation between the image and the point cloud BEV features and achieve efficient fusion of multi-source features.

[0148] (6) Perform BEV 3D object detection on the fused BEV features.

[0149] Continue to refer to Figure 3 As shown, using the fused BEV feature F bev , target classification and bounding box regression are performed through an improved 3D object detection network. The specific formula is as follows:

[0150] Result = Det(F bev )

[0151] Among them, the Det network includes several convolutional layers, fully connected layers and activation functions, adopts a multi-task learning strategy, optimizes the cross-entropy loss and the regression loss, and realizes accurate 3D object detection.

[0152] Based on the above embodiments and comparative examples, it can be clear that the multi-modal 3D detection method provided by the embodiments of the present invention improves the accuracy of multi-modal 3D object detection by introducing the depth estimation information of the image data to complement the sparse depth map; specifically, first, a scale-free dense depth estimation map is generated from the visible light image, and it is fused with the absolute scale sparse depth map in the designed fusion network to obtain the absolute scale dense depth map, thereby generating a high-quality bird's-eye view feature map of the visible light image, and efficiently fusing it with the feature map of the lidar point cloud, and finally realizing accurate 3D object detection.

[0153] Compared with the prior art, the technical solution proposed by the present invention applies visible light depth estimation to the depth completion of sparse depth maps, making up for the defects of regional depth errors and edge blurring in the direct fusion of visible light images and sparse depth maps, and also fully utilizes multi-modal information by introducing feature space fusion to improve the 3D object detection effect.

[0154] It should be understood that the above embodiments are only used to illustrate the technical concept and characteristics of the present invention, and their purpose is to enable those who are familiar with this technology to understand the content of the present invention and implement it accordingly, and cannot be used to limit the protection scope of the present invention. Any equivalent changes or modifications made according to the spirit and essence of the present invention should be covered within the protection scope of the present invention.

Claims

1. A multi-modal 3D detection method integrating depth completion, characterized in that, Including: Obtain image data and sparse depth data; Perform depth estimation on the image data to obtain a depth estimation map, and perform coordinate projection on the sparse depth data to obtain a sparse depth map; Use the sparse depth map to constrain the depth estimation map to obtain a dense depth map; Weightedly fuse the dense depth map and the depth estimation map to obtain a fused depth map; During the weighted fusion process, the weight of a pixel point in the depth estimation map is inversely correlated with the distance of the pixel point from the edge of the object in the image; Use the fused depth map to obtain a first feature map, and use the sparse depth map to obtain a second feature map; Fuse the first feature map and the second feature map to obtain a fused feature map; Perform object detection based on the fused feature map to obtain 3D environment information.

2. The multimodal 3D detection method according to claim 1, wherein The acquisition processes of the depth estimation map and the sparse depth map are expressed as: D est = Dep(I) P Lidar = {(x i , y i , z i )} P cam =(x cam , y cam , z cam ) = T CL P Lidar D sparse (u i ,v i ) = z cam Among them, D est represents the depth estimation map, Dep(·) represents the depth estimation model, I represents the image data, and P Lidar represents the sparse depth data, and x i , y i represent the coordinates of the pixel points in the sparse depth data, z i represents the depth value in the sparse depth data, i represents the pixel index, and P cam represents the projection point, and x cam , y cam represent the coordinates of the projection point, z cam represents the depth value of the projection point, T CL represents the transformation matrix for projecting the pixel points of the sparse depth data onto the coordinate system of the image data, u i , v i represent the coordinates of the pixel points in the sparse depth map, f x and f y respectively represent the ratio of the focal length of the camera that captures the image data to the physical size of a single pixel on the x-axis or y-axis, c x and c y represent the optical center coordinates of the camera, and D sparse represents the sparse depth map.

3. The multimodal 3D detection method according to claim 1, wherein The generation process of the dense depth map specifically includes: For the pixel points existing in the sparse depth map, D dense (u i ,v i ) = D sparse (u i ,v i ) For the pixel points not existing in the sparse depth map, ΔD est (u, v) = D est (u, v) - D est (u′, v′) Among them, D dense (·) represents the dense depth map, u and v represent coordinates in the dense depth map, (u′, v′) represents the adjacent coordinate points that are the nearest neighbors to the coordinate point (u, v) in the dense depth map and have completed constraints, and ΔD est (u, v) represents the offset of the depth estimation map corresponding to the coordinate point (u, v), S = {(u i , v i ) ≠ 0} represents the set of non-empty sparse pixel points, and ||·||2 represents the Euclidean distance.

4. The multimodal 3D detection method according to claim 1, wherein The acquisition process of the fused depth map is expressed as: D dep = αD dense + (1 - α)D est Among them, D dep represents the fused feature map, α represents the weight of the dense depth map, and 1-α represents the weight of the depth estimation map.

5. The multimodal 3D detection method according to claim 4, wherein The weight adjustment process of the weighted fusion includes: Calculate the gradient magnitude G(u, v) of the depth estimation map D est : Use the Canny edge detection algorithm to extract the edge region: E = {(u e , v e ) | G(u e , v e ) > T} Among them, τ represents a preset gradient threshold, E represents the edge region, and (u e , v e ) represents the pixel coordinates in the edge region; Dynamic weight calculation: For each position (u, v), calculate its pixel distance d(u, v) to the nearest edge: Dynamically adjust the weight α(u, v) according to the distance: Among them, is the Sigmoid function, and γ represents the attenuation coefficient.

6. The multimodal 3D detection method according to claim 1, wherein, The generation process of the first feature map is expressed as: Z = D dep (u, v) I bev = f(X, Y, Z) Among them, I bev represents the first feature map, f(·) represents the conversion function for projecting the 3D point cloud onto the BEV plane, and X, Y, Z represent the coordinates of the 3D point cloud; The generation process of the first feature map is expressed as: P bev = V(P Lidar ) Among them, P bev represents the second feature map, and V(·) represents a function that discretizes the three-dimensional point cloud into a two-dimensional grid and extracts the feature information within each grid.

7. The multimodal 3D detection method according to claim 6, wherein The generation process of the fused feature map is expressed as: F bev = T(I bev , P bev ) Among them, F bev represents the fused feature map, and T(·) represents a fusion function with a Transformer structure.

8. The multimodal 3D detection method according to claim 1, wherein The process of the feature detection is expressed as: Result = Det(F bev ) Wherein, Result represents the 3D environment information, and Det(·) represents the object detection network.

9. A multi-modal 3D detection system integrating depth completion, characterized in that, Including: A data acquisition module for obtaining image data and sparse depth data; A data preprocessing module for performing depth estimation on the image data to obtain a depth estimation map, and performing coordinate projection on the sparse depth data to obtain a sparse depth map; A dense depth module for using the sparse depth map to constrain the depth estimation map to obtain a dense depth map; A fused depth module for weightedly fusing the dense depth map and the depth estimation map to obtain a fused depth map; During the weighted fusion process, the weight of a pixel point in the depth estimation map is inversely correlated with the distance of the pixel point from the image edge; A feature map module for using the fused depth map to obtain a first feature map and using the sparse depth map to obtain a second feature map; A fused feature module for fusing the first feature map and the second feature map to obtain a fused feature map; An object detection module for performing object detection based on the fused feature map to obtain 3D environment information.

10. A readable storage medium, characterized in that, A computer program is stored in the readable storage medium, and when the computer program is run, it executes the steps of the multi-modal 3D detection method according to any one of claims 1-8.