Three-dimensional object detection method and system based on multi-modal feature fusion under cross-view angle

By constructing a multimodal feature fusion network with cross-view perspectives in intelligent vehicles, the problem of limited fusion effect caused by the heterogeneity of radar point cloud and camera image information is solved, achieving more accurate 3D target detection and improving the safety of intelligent vehicle systems.

CN115965847BActive Publication Date: 2026-01-13TSINGHUA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310076916.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-17
Publication Date
2026-01-13
Estimated Expiration
2043-01-17

AI Technical Summary

Technical Problem

Existing methods for fusing radar point clouds and camera images fail to effectively consider millimeter-wave radar point cloud information and camera image information as two spatially heterogeneous types of information, resulting in limited fusion performance in 3D target perception tasks and making it difficult to meet the advanced autonomous driving requirements of intelligent vehicles.

Method used

By constructing a multimodal feature fusion network under cross-view perspectives, we can extract features from different perspectives of camera images and millimeter-wave radar data and perform cross-view transformation. Combined with a deep fusion network, we can regress target category and 3D position information to achieve adaptive fusion of spatial characteristics of multimodal data.

Benefits of technology

It improves the accuracy and reliability of target detection, enhances the safety of intelligent vehicle systems, adapts to the spatial characteristics of different modal data, and achieves more accurate three-dimensional target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115965847B_ABST
    Figure CN115965847B_ABST
Patent Text Reader

Abstract

The application relates to a three-dimensional target detection method and system for multi-modal feature fusion under cross-view angles, which comprises the following steps: feature extraction of camera image data and millimeter wave radar data under different view angles, cross-view angle conversion, and obtaining feature information under cross-view angles; constructing a fusion network based on cross-view multi-modal data, deeply fusing the obtained feature information under cross-view angles, extracting features, simultaneously performing regression of target categories and three-dimensional position information, and obtaining complete three-dimensional target detection information. The application fully considers spatial features of camera image information under a front view angle and spatial features of millimeter wave radar point cloud information under a bird's eye view angle, can effectively fuse spatial characteristics of different sensors, improves fusion performance, effectively improves accuracy, and is convenient for subsequent algorithm processing. The application can be widely applied to the environment perception field of intelligent automobiles.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of environmental perception for intelligent vehicles, and in particular to a three-dimensional target detection method and system that utilizes multimodal feature fusion from cross-view perspectives. Background Technology

[0002] Intelligent vehicles need to utilize observational information provided by onboard sensors to perceive and understand the driving environment. Through algorithms such as target detection and tracking, semantic segmentation, and scene understanding, the perception results are used for driving tasks such as path planning and obstacle avoidance. However, because the driving environment of intelligent vehicles is typically highly complex and dynamic, it places high demands on the accuracy, stability, and reliability of the vehicle's perception system. Single sensors have limitations in perception range, accuracy, and information richness, making it difficult to meet the perception requirements of advanced autonomous driving. Therefore, using multi-sensor information for fusion perception has become an effective means of perception enhancement.

[0003] Cameras and millimeter-wave radar are two common automotive perception sensors. Camera-captured images contain dense semantic information, while millimeter-wave radar can directly observe the relative position and velocity of targets. Furthermore, the all-weather capability of millimeter-wave radar allows it to withstand harsh weather conditions. Therefore, information fusion algorithms from these two sensors have been widely applied in mass-produced driver assistance systems. However, because camera-captured images lack depth information, and millimeter-wave radar point clouds lack height information, and because millimeter-wave radar point clouds are sparse and cluttered, information from both cameras and millimeter-wave radar lacks a complete description of the three-dimensional environment, making them difficult to directly apply in three-dimensional target perception tasks.

[0004] Existing methods for fusing radar point clouds and camera images typically project millimeter-wave radar point cloud data onto camera images using a spatial calibration matrix and design corresponding feature representation rules to ensure consistent representation of the radar point cloud in the image space before proceeding with subsequent fusion feature extraction for target detection or tracking. However, this approach fails to consider the significant spatial differences between millimeter-wave radar point cloud and camera image information as two spatially heterogeneous datasets. Forcibly unifying them to the front-view perspective of the camera space during multimodal data fusion results in poor adaptation to the features of millimeter-wave radar data, leading to limited fusion effectiveness and requiring further algorithm performance improvement. Summary of the Invention

[0005] To address the aforementioned problems, the present invention aims to provide a three-dimensional target detection method and system based on multimodal feature fusion from cross-perspectives. This method can be applied to perception algorithms that fuse multimodal spatial heterogeneous information from multiple sensors at the data level. It can fully adapt to the spatial characteristics of different modal data, improve fusion performance, obtain more accurate and reliable target detection information, and enhance the safety of intelligent vehicle systems.

[0006] To achieve the above objectives, the present invention adopts the following technical solution:

[0007] In a first aspect, the present invention provides a three-dimensional target detection method based on multimodal feature fusion from cross-viewpoints, comprising the following steps:

[0008] Feature extraction is performed on camera image data and millimeter-wave radar data from different perspectives, and cross-perspective transformation is performed to obtain feature information from the cross-perspective.

[0009] A fusion network based on cross-view multimodal data is constructed to deeply fuse and extract features from the obtained cross-view feature information. At the same time, regression of target category and 3D position information is performed to obtain complete 3D target detection information.

[0010] Furthermore, the method for extracting features from camera image data and millimeter-wave radar data from different perspectives and performing cross-perspective transformation to obtain feature information from the cross-perspective includes the following steps:

[0011] A feature extractor is constructed based on the camera image viewpoint to extract features from the camera image, thereby obtaining two-dimensional multi-scale convolutional features and their corresponding 2D target detection locations from the front viewpoint.

[0012] A feature extractor from the perspective of millimeter-wave radar is constructed to perform multi-frame point cloud accumulation processing on millimeter-wave radar point cloud data and obtain a feature distribution map of radar point cloud from a bird's-eye view.

[0013] A cross-view feature converter is constructed to convert the two-dimensional multi-scale convolutional features from the front view perspective and the radar point cloud feature distribution map from the bird's-eye view perspective to obtain feature information from the cross-view perspective.

[0014] Furthermore, the method for constructing a feature extractor from the camera image perspective to extract features from the camera image and obtain two-dimensional multi-scale convolutional features and their corresponding 2D target detection locations from the front view perspective includes the following steps:

[0015] By using a convolutional neural network to extract features from camera images, two-dimensional multi-scale convolutional features from the front view perspective are obtained.

[0016] Based on the obtained two-dimensional multi-scale convolutional features, for each pixel in the camera image with coordinates (h, w)... Estimate its depth distribution D and category distribution Simultaneously, a preliminary regression of the target's 2D position is performed to obtain the 2D target detection position in the camera image.

[0017] Furthermore, the method for constructing a feature extractor from the perspective of millimeter-wave radar, performing multi-frame point cloud accumulation processing on millimeter-wave radar data, and obtaining a radar point cloud feature distribution map from a bird's-eye view includes the following steps:

[0018] Multi-frame point cloud accumulation processing is performed on millimeter-wave radar point cloud data to obtain the current frame millimeter-wave radar point cloud observation results.

[0019] Based on the current frame of millimeter-wave radar point cloud observation results, a feature distribution map of radar point cloud from a bird's-eye view is constructed using a Gaussian probability distribution model.

[0020] Furthermore, the current frame millimeter-wave radar point cloud observation results are as follows:

[0021] Z radar (t)=T c_from_r T c_from_g (t)T g_from_c (tk)T r_from_c Z radar (tk)

[0022] Among them, Z radar (t) represents the current frame's millimeter-wave radar point cloud observation result; T c_from_r T represents the transition matrix from the millimeter-wave radar coordinate system to the vehicle coordinate system; c_from_g (t) represents the transition matrix from the global coordinate system to the vehicle coordinate system in the current frame; T g_from_c (tk) represents the transition matrix from the vehicle coordinate system to the global coordinate system in the k-th frame; T r_from_c Z represents the transition matrix from the vehicle to the millimeter-wave radar coordinate system; radar (t) and Z radar (tk) represents the millimeter-wave radar point cloud observation results of the current frame and the millimeter-wave radar point cloud observation results of the previous k frames.

[0023] Furthermore, the method for constructing a cross-view feature converter, which transforms the two-dimensional multi-scale convolutional features from the front view perspective and the radar point cloud feature distribution map from the bird's-eye view perspective to obtain feature information from the cross-view perspective, includes the following steps:

[0024] A front-view feature converter is constructed based on the intrinsic and extrinsic parameters of the image to convert the radar point cloud feature distribution map under the bird's-eye view to the front-view view, and obtain the radar Gaussian feature fusion result under the front view.

[0025] Based on the intrinsic and extrinsic parameters of the image, a bird's-eye view feature converter is constructed to transform the two-dimensional multi-scale convolutional features from the front view to the bird's-eye view, thus obtaining the image convolutional feature fusion result from the bird's-eye view.

[0026] Furthermore, the method of constructing a front-view feature converter based on image intrinsic and extrinsic parameter information to convert the radar point cloud feature distribution map from the bird's-eye view to the front-view view, thereby obtaining the radar Gaussian feature fusion result in the front view, includes the following steps:

[0027] First, utilize the spatial transformation relationship T from the bird's-eye view to the front view coordinate system. f_from_b Project the radar point cloud onto the front view;

[0028] Secondly, based on the 2D target detection position in the camera image, the radar point cloud that falls within the two-dimensional bounding box of the target in the front view is preserved, and the corresponding pixel positions are filled to obtain the radar Gaussian feature fusion result in the front view.

[0029] Furthermore, based on the calibration information between millimeter-wave radar and the image, a bird's-eye view feature converter is constructed to convert the two-dimensional multi-scale convolutional features from the front view perspective to the bird's-eye view perspective, obtaining the image convolutional feature fusion result from the bird's-eye view perspective. This includes the following steps:

[0030] First, determine the spatial size (d) represented by each pixel in three-dimensional space. x d y d z ), and construct the corresponding view cone;

[0031] Secondly, using the intrinsic parameters from the camera to the image coordinate system and the extrinsic parameters from the camera space to the bird's-eye view space coordinate system, the position distribution of each pixel with image coordinates (u, v) in the bird's-eye view space is obtained. The above transition matrix is ​​denoted as T. b_from_f ;

[0032] Next, the corresponding image frustum is transformed using the aforementioned position transformation relationship T. b_from_f Projected onto the bird's-eye view;

[0033] Finally, the convolutional features of the two-dimensional multi-scale image are filled into the corresponding view frustum under the bird's-eye view using an interpolation function to obtain the image convolutional feature fusion result under the bird's-eye view perspective.

[0034] Furthermore, the construction of a fusion network based on cross-view multimodal data involves deep fusion and feature extraction of the obtained cross-view feature information, while simultaneously regressing target category and 3D position information to obtain complete 3D target detection information, including:

[0035] A multimodal data fusion network based on cross-viewpoints is constructed, which connects the two-dimensional multi-scale convolutional features under the front view perspective and the projected radar Gaussian features together, and connects the radar point cloud feature distribution map under the bird's-eye view perspective and the projected image convolutional features together.

[0036] The obtained feature fusion information is subjected to deep fusion feature extraction. After unifying the scale, all feature information is connected together, and convolutional layers are used to achieve the corresponding information weight allocation.

[0037] By using a 3D object detection head to regress object category and pose information, calculating the corresponding loss function, training and optimizing the network, complete 3D object detection information is obtained.

[0038] Secondly, the present invention provides a three-dimensional target detection system based on multimodal feature fusion under cross-view perspectives, comprising:

[0039] The cross-view feature extraction module is used to extract features from camera image data and millimeter-wave radar data from different viewpoints, and to perform cross-viewpoint transformation to obtain feature information from the cross-viewpoint.

[0040] The 3D target detection module is used to construct a fusion network based on cross-view multimodal data, further extract features from the obtained cross-view feature information, and regress the target category and 3D position information to obtain complete 3D target detection information.

[0041] The present invention has the following advantages due to the adoption of the above technical solutions: The present invention fully considers the spatial characteristics of camera image information in the front view and the spatial characteristics of millimeter-wave radar point cloud information in the bird's-eye view, and can adapt to the spatial characteristics of different sensors to further perform effective fusion, improve fusion performance, effectively improve accuracy, and facilitate subsequent algorithm processing. Attached Figure Description

[0042] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts. In the drawings:

[0043] Figure 1This is a schematic diagram of the method for three-dimensional target detection based on multimodal feature fusion from a cross-view perspective in an embodiment of the present invention;

[0044] Figure 2 This is a schematic diagram of the cross-view feature converter in an embodiment of the present invention. Detailed Implementation

[0045] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the described embodiments of the present invention are within the scope of protection of the present invention.

[0046] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0047] In some embodiments of this invention, a 3D target detection method based on multimodal feature fusion under cross-view perspectives is provided. This method extracts features from camera images using a convolutional neural network; extracts millimeter-wave radar point cloud features based on an understanding of millimeter-wave radar characteristics; and fuses camera image data features with millimeter-wave radar data features from both the front view and bird's-eye view perspectives (the bird's-eye view coordinate system can be directly selected as the coordinate system of the millimeter-wave radar, or, for subsequent application considerations, the vehicle coordinate system; in this invention, it is uniformly a 3D spatial coordinate system under the bird's-eye view). Utilizing this multimodal feature fusion, further feature extraction and regression of target category and 3D position information are performed, ultimately outputting complete 3D target detection information. This invention can be applied to perception algorithms that fuse multimodal spatial heterogeneous information from multiple sensors at the data level. It can fully adapt to the spatial characteristics of different modal data, improve fusion performance, obtain more accurate and reliable target detection information, and enhance the safety of intelligent vehicle systems.

[0048] Correspondingly, in other embodiments of the present invention, a three-dimensional target detection system, device, and medium for multimodal feature fusion under cross-view perspectives are provided.

[0049] Example 1

[0050] like Figure 1As shown, this embodiment provides a three-dimensional target detection method based on multimodal feature fusion under cross-view perspectives, including the following steps:

[0051] 1) Extract features from camera image data and millimeter-wave radar data from different perspectives, and perform cross-perspective transformation to obtain feature information from the cross-perspective.

[0052] 2) Construct a fusion network based on cross-view multimodal data, deeply fuse the feature information obtained from the cross-view and extract features, and simultaneously regress the target category and 3D position information to obtain complete 3D target detection information.

[0053] Preferably, in step 1) above, the method for extracting features from camera image data and millimeter-wave radar data from different perspectives and performing cross-perspective transformation to obtain feature information from the cross-perspective includes the following steps:

[0054] 1.1) Construct a feature extractor from the camera image perspective to extract features from the camera image and obtain two-dimensional multi-scale convolutional features and their corresponding 2D target detection locations from the front view perspective.

[0055] 1.2) Construct a feature extractor from the perspective of millimeter-wave radar, perform multi-frame point cloud accumulation processing on millimeter-wave radar point cloud data, and obtain a feature distribution map of radar point cloud from the perspective of bird's-eye view.

[0056] 1.3) Construct a cross-view feature converter to convert the two-dimensional multi-scale convolutional features from the front view perspective and the radar point cloud feature distribution map from the bird's-eye view perspective to obtain feature information from the cross-view perspective.

[0057] Preferably, in step 1.1) above, the method of constructing a feature extractor from the camera image viewpoint to extract features from the camera image and obtain two-dimensional multi-scale convolutional features and their corresponding 2D target detection positions from the front viewpoint includes the following steps:

[0058] 1.1.1) Use convolutional neural networks to extract features from camera images to obtain two-dimensional multi-scale convolutional features from the front view perspective.

[0059] Specifically, in this embodiment, a deep residual convolutional neural network (ResNet101) is used as the main network framework to extract deep multi-channel convolutional features of camera images. Then, the deep multi-channel convolutional features are further extracted through a feature pyramid network (FPN) to obtain two-dimensional multi-scale convolutional feature information.

[0060] 1.1.2) Based on the obtained two-dimensional multi-scale convolution features, for each pixel in the camera image with coordinates (h, w) Estimate its depth distribution D and category distribution Simultaneously, a preliminary regression of the target's 2D position is performed to obtain the 2D target detection position in the camera image.

[0061] Specifically, in this embodiment, the softmax function is used to predict the classification probability results of each discrete depth from the two-dimensional multi-scale convolutional features, and the depth distribution D = {d0, d0+Δ, ..., d0+kΔ} is obtained. D contains k+1 discrete points, d0 is the minimum value of the predicted depth distribution, and the unit is m; Δ is the interval distance value of the discrete depth distribution, and the unit is m.

[0062] Preferably, in step 1.2) above, the method for constructing a feature extractor from the millimeter-wave radar perspective, performing multi-frame point cloud accumulation processing on millimeter-wave radar data, and obtaining a radar point cloud feature distribution map from a bird's-eye view includes the following steps:

[0063] 1.2.1) Perform multi-frame point cloud accumulation processing on the millimeter-wave radar point cloud data to obtain the current frame millimeter-wave radar point cloud observation results.

[0064] In this embodiment, location information and timestamp information are used to accumulate radar point cloud information from the past 5 frames (this invention uses this as an example but is not limited to this) to the current frame, thereby further increasing the point cloud density. The calculation formula is as follows:

[0065] Z radar (t)=T c_from_r T c_from_g (t)T g_from_c (tk)T r_from_c Z radar (tk) (1)

[0066] Among them, Z radar (t) represents the current frame's millimeter-wave radar point cloud observation result; T c_from_r T represents the transition matrix from the millimeter-wave radar coordinate system to the vehicle coordinate system; c_from_g (t) represents the transition matrix from the global coordinate system to the vehicle coordinate system in the current frame; T g_from_c (tk) represents the transition matrix from the vehicle coordinate system to the global coordinate system in the k-th frame; T r_from_c Z represents the transition matrix from the vehicle to the millimeter-wave radar coordinate system; radar (t) and Z radar (tk) represents the millimeter-wave radar point cloud observation results of the current frame and the millimeter-wave radar point cloud observation results of the previous k frames.

[0067] As can be seen from equation (1), in order to obtain the millimeter-wave radar point cloud observation results Z from k frames ago... radar (tk) Projected onto the current frame Z radar (t), Z needs to beradar (tk) using T r_from_c Projected onto the vehicle's coordinate system, and then using the relationship T between the vehicle and the global coordinate system provided by the positioning system for k frames prior. g_from_c (tk) and the relationship between the vehicle and the global coordinate system at the current moment T c_from_g (t) is used to consider the motion of the vehicle coordinate system within k frames, and finally T is used to... c_from_r Projected back to the current time Z radar (t) space.

[0068] Among them, Z radar (t) is represented as:

[0069]

[0070] Among them, dis lat Indicates the lateral position of the target measured by radar; dis long Indicates the longitudinal position measured by radar; vel lat Indicates the relative lateral velocity measured by radar; vel long It represents the longitudinal velocity measured by the radar; RCS represents the reflection area of ​​the radar measurement signal, reflecting the reflection intensity of the radar echo.

[0071] Each transition matrix consists of a rotation matrix R and a translation vector t, satisfying the following relationship:

[0072]

[0073] 1.2.2) Based on the current frame millimeter-wave radar point cloud observation results, a radar point cloud feature distribution map from a bird's-eye view is constructed using a Gaussian probability distribution model.

[0074] Based on the relative position, velocity, and radar cross section (RCS, which reflects radar reflection intensity) provided by radar point cloud information, a feature distribution map of radar point cloud data under bird's-eye view is constructed using a Gaussian probability distribution model.

[0075] Specifically, in this embodiment, the radar point cloud data feature distribution map includes 5 channels, each representing the target's lateral position (dis). lat ), vertical position (dis) long ), lateral velocity (vel) lat Longitudinal velocity (vel) long Based on the radar cross-section (RCS) and reflection intensity information, a feature distribution map of the radar point cloud is designed using a binary linear Gaussian normal distribution from a bird's-eye view.

[0076]

[0077] Where (x, y) represents the rasterized space under the bird's-eye view, and the position coordinates of any cell are determined based on the position information of the target observed laterally and radially under the bird's-eye view by the millimeter-wave radar; μ represents the position information of the target observed laterally and radially by the radar [μ1 = dis long μ2 = dis lat The value of σ is determined based on the measurement values ​​of each variable by the millimeter-wave radar; σ represents the uncertainty of the target position value measured by the radar, which is determined based on the measurement accuracy of each variable by the millimeter-wave radar.

[0078] Preferably, in step 1.3) above, such as Figure 2 As shown, a method for constructing a cross-view feature converter to transform the 2D multi-scale convolutional features from the front view perspective and the radar point cloud feature distribution map from the bird's-eye view perspective to obtain feature information from the cross-view perspective includes the following steps:

[0079] 1.3.1) Construct a front view feature converter based on the intrinsic and extrinsic parameters of the image to convert the radar point cloud feature distribution map under the bird's-eye view to the front view, and obtain the radar Gaussian feature fusion result under the front view.

[0080] The intrinsic parameters of the image include the transformation matrix from the camera coordinate system to the image plane pixel coordinate system; the extrinsic parameters include the spatial transformation matrix from the camera coordinate system to the reference coordinate system. Specifically, the feature transformation under the front view includes:

[0081] First, utilize the spatial transformation relationship T from the bird's-eye view to the front view coordinate system. f_from_b First, the radar point cloud is projected onto the front view. Second, based on the predicted 2D target detection position of the camera image in step 1.1.2), the radar point clouds falling within the target's 2D bounding box in the front view are retained (in practice, considering the influence of errors, some radar point clouds that are closer to the target's bounding box may be retained, so the threshold is flexibly adjusted by taking an α (α≥1) value), and the corresponding pixel positions are filled to obtain the radar Gaussian feature fusion result under the front view:

[0082]

[0083] in, Indicates the radar target position (dis) long dis lat The pixel position projected onto the image plane using intrinsic and extrinsic parameters; j h j () represents the length and width of the two-dimensional rectangular frame formed after the radar target is projected onto the image plane; This indicates the pixel value after the radar is projected onto the image plane.

[0084] 1.3.2) Based on the intrinsic and extrinsic parameters of the image (including the spatial transformation relationship from radar to camera and the transformation relationship from camera to image plane), a bird's-eye view feature converter is constructed to convert the two-dimensional multi-scale convolutional features under the front view perspective to the bird's-eye view perspective, and obtain the image convolutional feature fusion result under the bird's-eye view perspective.

[0085] Specifically, image feature transformation from a bird's-eye view perspective includes:

[0086] First, determine the spatial size (d) represented by each pixel in three-dimensional space. x d y d z First, the corresponding view frustum is constructed. Second, using the intrinsic parameters from the camera to the image coordinate system and the extrinsic parameters from the camera space to the bird's-eye view space coordinate system, the position distribution of each pixel with image coordinate position (u, v) in the bird's-eye view space is obtained. The above transition matrix is ​​denoted as T. bb_fromb_f Next, the corresponding image frustum is transformed using the aforementioned position transformation relationship T. b_from_f The image is then projected onto the bird's-eye view. Finally, the previously generated image convolutional features are filled into the corresponding view frustum under the bird's-eye view using an interpolation function to obtain the image convolutional feature fusion result from the bird's-eye view perspective.

[0087] Preferably, in step 2) above, a fusion network based on cross-view multimodal data is constructed, and further feature extraction is performed based on the cross-view feature information obtained in step 1) to obtain the regression of target category information and 3D position information, including the following steps:

[0088] 2.1) Construct a multimodal data fusion network based on cross-viewpoints, connecting the two-dimensional multi-scale convolutional features from the front view perspective and the projected radar Gaussian features together, while connecting the radar point cloud feature distribution map from the bird's-eye view perspective and the projected image convolutional features together.

[0089] 2.2) Using a multilayer perceptron (MLP) and a convolutional neural network, the feature fusion information obtained in step 2.1) under the cross-perspective is deeply fused to extract features. After unifying the scale, all feature information is connected together, and the corresponding information weights are allocated using convolutional layers.

[0090] In this embodiment, the convolutional neural network uses a deep residual neural network (ResNet18) and a feature pyramid network, but is not limited to these.

[0091] 2.3) Use the 3D target detection head to regress the target category and target pose information, calculate the corresponding loss function, train and optimize the network, and then obtain complete 3D target detection information.

[0092] In summary, this invention, through the design of a cross-view feature converter, achieves feature representation of image information and millimeter-wave radar information across spatial perspectives; and through a deep fusion neural network, it completes the fusion and regression of multimodal data. It can fully adapt to the spatial representation of data under different perspectives, realize 3D target detection through multi-source information fusion, and can be further applied to single-vehicle multi-sensor fusion perception tasks and multi-vehicle joint perception tasks.

[0093] Example 2

[0094] The above-described embodiment 1 provides a 3D target detection method based on multimodal feature fusion under cross-viewpoints. Correspondingly, this embodiment provides a 3D target detection system based on multimodal feature fusion under cross-viewpoints. The system provided in this embodiment can implement the 3D target detection method based on multimodal feature fusion under cross-viewpoints of embodiment 1. This system can be implemented through software, hardware, or a combination of both. For example, the system may include integrated or separate functional modules or units to execute the corresponding steps in the methods of embodiment 1. Since the system in this embodiment is basically similar to the method embodiment, the description process in this embodiment is relatively simple. For relevant details, please refer to the description of embodiment 1. The system embodiment provided in this embodiment is merely illustrative.

[0095] The 3D target detection system based on multimodal feature fusion under cross-view perspectives provided in this embodiment includes:

[0096] The cross-view feature extraction module is used to extract features from camera image data and millimeter-wave radar data from different viewpoints, and to perform cross-viewpoint transformation to obtain feature information from the cross-viewpoint.

[0097] The 3D target detection module is used to construct a fusion network based on cross-view multimodal data, further extract features from the obtained cross-view feature information, and regress the target category and 3D position information to obtain complete 3D target detection information.

[0098] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A three-dimensional object detection method of multi-modal feature fusion under cross-view angle, characterized in that, The method comprises the following steps: The camera image data and the millimeter wave radar data are subjected to feature extraction under different visual angles, cross-visual angle conversion is performed, and feature information under the cross-visual angle is obtained; A fusion network based on cross-visual angle multi-modal data is constructed, the obtained feature information under the cross-visual angle is subjected to deep fusion and feature extraction, and regression of target categories and three-dimensional position information is performed, and complete three-dimensional target detection information is obtained; The method for performing feature extraction under different visual angles on the camera image data and the millimeter wave radar data and performing cross-visual angle conversion to obtain feature information under the cross-visual angle comprises the following steps: A feature extractor under the camera image visual angle is constructed, camera image feature extraction is performed, two-dimensional multi-scale convolution features under the front view visual angle and corresponding 2D target detection positions are obtained; A feature extractor under the millimeter wave radar visual angle is constructed, multi-frame point cloud accumulation processing is performed on the millimeter wave radar point cloud data, and a radar point cloud feature distribution map under the bird's eye view visual angle is obtained; A cross-visual angle feature converter is constructed, the two-dimensional multi-scale convolution features under the front view visual angle and the radar point cloud feature distribution map under the bird's eye view visual angle are subjected to visual angle conversion, and feature information under the cross-visual angle is obtained; The method for constructing the cross-visual angle feature converter, performing visual angle conversion on the two-dimensional multi-scale convolution features under the front view visual angle and the radar point cloud feature distribution map under the bird's eye view visual angle, and obtaining feature information under the cross-visual angle comprises the following steps: A front view visual angle feature converter is constructed based on image internal and external parameter information, the radar point cloud feature distribution map under the bird's eye view visual angle is converted to the front view visual angle, and a radar Gaussian feature fusion result under the front view is obtained; A bird's eye view visual angle feature converter is constructed based on image internal and external parameter information, the two-dimensional multi-scale convolution features under the front view visual angle are converted to the bird's eye view visual angle, and an image convolution feature fusion result under the bird's eye view visual angle is obtained; The method for constructing the fusion network based on cross-visual angle multi-modal data, performing deep fusion on the obtained feature information under the cross-visual angle, extracting features, simultaneously performing regression of target categories and three-dimensional position information, and obtaining complete three-dimensional target detection information comprises the following steps: A multi-modal data fusion network based on the cross-visual angle is constructed, the two-dimensional multi-scale convolution features under the front view visual angle and the projected radar Gaussian features are connected together, and the radar point cloud feature distribution map under the bird's eye view visual angle and the projected image convolution features are connected together; Deep fusion feature extraction is performed on the obtained feature fusion information, all the feature information is connected together after unification of scales, and corresponding information weight distribution is realized by using a convolution layer; Three-dimensional target detection heads are used to perform regression of target categories and target pose information, a corresponding loss function is calculated, network training and optimization are performed, and complete three-dimensional target detection information is obtained.

2. The three-dimensional object detection method of claim 1, wherein, The method for constructing the feature extractor under the camera image visual angle, performing feature extraction on the camera image, and obtaining two-dimensional multi-scale convolution features under the front view visual angle and corresponding 2D target detection positions comprises the following steps: The camera image is subjected to feature extraction by using a convolutional neural network to obtain two-dimensional multi-scale convolutional features under the front view perspective; Based on the obtained two-dimensional multi-scale convolutional features, the coordinates of each pixel in the camera image are... pixels Estimate its depth distribution and category distribution Simultaneously, a preliminary regression of the target's 2D position is performed to obtain the 2D target detection position in the camera image.

3. The method of claim 1, wherein the method further comprises: The method comprises the following steps: The multi-frame point cloud accumulation processing is performed on the millimeter wave radar point cloud data to obtain a current frame millimeter wave radar point cloud observation result; Based on the current frame millimeter wave radar point cloud observation result, a radar point cloud feature distribution map under the bird's eye view perspective is constructed by using a Gaussian probability distribution model.

4. The three-dimensional object detection method of claim 3, wherein, The current frame millimeter wave radar point cloud observation result is: wherein, is the current frame millimeter wave radar point cloud observation result; denotes the transfer matrix from the millimeter wave radar coordinate system to the ego vehicle coordinate system; denotes the transfer matrix from the global coordinate system to the ego vehicle coordinate system in the current frame; denotes the transfer matrix from the ego vehicle coordinate system to the global coordinate system in the previous kth frame; denotes the transfer matrix from the ego vehicle to the millimeter wave radar coordinate system; and denotes the current frame millimeter wave radar point cloud observation result and the millimeter wave radar point cloud observation result in the frame before.

5. The method of claim 1, wherein, The front view perspective feature converter is constructed based on the image-based internal and external parameter information, the radar point cloud feature distribution map under the bird's eye view perspective is converted to the front view perspective, and a radar Gaussian feature fusion result under the front view perspective is obtained. First, the spatial transformation relationship from the bird's eye view to the front view coordinate system is used projecting the radar point cloud to the front view; Secondly, according to the 2D target detection position of the camera image, the radar point cloud falling within the two-dimensional space bounding box of the target in the front view perspective is reserved, and the corresponding pixel position is filled, so as to obtain the radar Gaussian feature fusion result under the front view perspective.

6. The three-dimensional object detection method of claim 1, wherein, Based on the calibration information between the millimeter wave radar and the image, the bird's eye view perspective feature converter is constructed, the two-dimensional multi-scale convolutional features under the front view perspective are converted to the bird's eye view perspective, and an image convolutional feature fusion result under the bird's eye view perspective is obtained. First, determine the spatial size of each pixel in three-dimensional space , and construct the corresponding view frustum; Secondly, the position distribution of each image coordinate position as pixel in the bird's eye view space is obtained by using the camera internal parameter to the image coordinate system and the camera space to the bird's eye view space coordinate system, and the position conversion relationship is recorded as . ; Again, the corresponding image frustum is projected to the bird's eye view using the position conversion relationship to the bird's eye view; Finally, the two-dimensional multi-scale image convolutional features are filled into the corresponding view cone under the bird's eye view by using an interpolation function, so as to obtain the image convolutional feature fusion result under the bird's eye view perspective.

7. A three-dimensional object detection system under cross-view multi-modal feature fusion, characterized in that, It comprises: The cross-perspective feature extraction module is used for performing feature extraction on camera image data and millimeter wave radar data under different perspectives, and performing cross-perspective conversion to obtain feature information under the cross-perspective; The three-dimensional target detection module is used for constructing a fusion network based on cross-perspective multi-modal data, further extracting the feature information obtained under the cross-perspective, and simultaneously performing regression on target categories and three-dimensional position information to obtain complete three-dimensional target detection information. The method for performing feature extraction on camera image data and millimeter wave radar data under different perspectives and performing cross-perspective conversion to obtain feature information under the cross-perspective comprises the following steps: The camera image perspective feature extractor is constructed to perform feature extraction on the camera image to obtain two-dimensional multi-scale convolutional features under the front view perspective and corresponding 2D target detection positions; The millimeter wave radar perspective feature extractor is constructed to perform multi-frame point cloud accumulation processing on millimeter wave radar point cloud data and obtain a radar point cloud feature distribution map under the bird's eye view perspective; The cross-perspective feature converter is constructed to perform perspective conversion on the two-dimensional multi-scale convolutional features under the front view perspective and the radar point cloud feature distribution map under the bird's eye view perspective to obtain feature information under the cross-perspective. The method for constructing the cross-perspective feature converter to perform perspective conversion on the two-dimensional multi-scale convolutional features under the front view perspective and the radar point cloud feature distribution map under the bird's eye view perspective to obtain feature information under the cross-perspective comprises the following steps: An image-based internal and external parameter information is used to construct a front view perspective feature converter, and a radar point cloud feature distribution map under a bird's eye view perspective is converted to a front view perspective to obtain a radar high-dimensional feature fusion result under the front view perspective; An image-based internal and external parameter information is used to construct a bird's eye view perspective feature converter, and a two-dimensional multi-scale convolution feature under a front view perspective is converted to a bird's eye view perspective to obtain an image convolution feature fusion result under the bird's eye view perspective; The fusion network based on the cross-perspective multi-modal data is constructed, the obtained feature information under the cross-perspective is deeply fused and features are extracted, and regression of target categories and three-dimensional position information is simultaneously performed to obtain complete three-dimensional target detection information, including: The fusion network based on the cross-perspective multi-modal data is constructed, the obtained feature information under the cross-perspective is deeply fused and features are extracted, and regression of target categories and three-dimensional position information is simultaneously performed to obtain complete three-dimensional target detection information, including: The obtained feature fusion information is deeply fused and features are extracted, all the feature information is connected together after unification of scales, and corresponding information weight distribution is realized by using a convolution layer; Regression of target categories and target pose information is performed by using a three-dimensional target detection head, a corresponding loss function is calculated, the network is trained and optimized, and complete three-dimensional target detection information is obtained.

Citation Information

Patent Citations

  • Millimeter wave radar and vision fused three-dimensional target detection method based on attention mechanism

    CN114708585A

  • 3D target detection method and device based on multi-view fusion

    CN114913506A