3D target detection method based on cascade view cone image fusion
The cascaded view cone image fusion method improves 3D target detection by integrating 2D image features with point cloud data, addressing inefficiencies in existing methods and enhancing semantic information for accurate 3D object detection.
Patent Information
- Application Number
- CN202510383119.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-15
AI Technical Summary
The existing 3D object detection technology relies on lidar point cloud data, and has problems such as low processing efficiency and lack of apparent information.
Using a method based on cascading cone image fusion, the initial two-dimensional image and point cloud data are feature extraction and transformation, cone features and image features are fused, and the attention mechanism is used to perform refined granularity fusion, and finally input the detection head for 3D object detection.
It improves the processing efficiency of point cloud data, enhances semantic information, and realizes more accurate 3D object detection.
Smart Images

Figure CN120318490A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of object detection, and in particular to a 3D object detection method based on cascaded frustum image fusion. Background Art
[0002] In existing 3D object detection technologies, lidar point cloud data is usually relied on. The point cloud data has the characteristics of large scale, disorder and irregularity, which restricts the processing efficiency of the point cloud data and lacks appearance information. Therefore, the existing 3D object detection technologies have problems of difficult processing of point cloud data and lack of appearance information. Summary of the Invention
[0003] Based on this, the purpose of this application is to provide a 3D object detection method based on cascaded frustum image fusion, which can overcome the deficiencies of the prior art.
[0004] To achieve the above purpose, the technical solution adopted by this application is as follows:
[0005] A 3D object detection method based on cascaded frustum image fusion, comprising:
[0006] Performing object extraction on an initial two-dimensional image to obtain an object image;
[0007] Performing two-dimensional feature extraction and feature transformation processing on the object image to obtain a first three-dimensional point feature;
[0008] Performing feature transformation processing on the frustum feature obtained by mapping the point cloud in the initial frustum to the frustum axis to obtain a second three-dimensional point feature;
[0009] Performing feature channel number transformation processing on the point cloud data in the initial frustum to obtain a first transformed point feature;
[0010] Fusing the first three-dimensional point feature, the second three-dimensional point feature and the first transformed point feature to obtain a first fused point feature;
[0011] Decomposing the first fused feature and extracting object features to obtain an object image feature and an object frustum feature;
[0012] Performing feature channel number transformation processing on the first fused feature to obtain a second transformed feature;
[0013] Fusing the second transformed feature, the object image feature and the object frustum feature to obtain a second fused feature;
[0014] To refine the fusion granularity of the object image features and the object frustum features, the object image features and the object frustum features are fused by an attention mechanism, and the output features and the second transformed features are input into a point fusion module to obtain the second fusion features;
[0015] The second fusion features are input into a detection head to obtain 3D object detection boxes; the 3D object detection boxes are used to indicate 3D objects in the world coordinate system.
[0016] In one embodiment, the step of performing two-dimensional feature extraction and feature transformation processing on the object image to obtain the first three-dimensional point features includes:
[0017] Extract two-dimensional features of the object image through a 2D FCN;
[0018] Perform feature transformation processing on the two-dimensional features through a G2P module to obtain the first three-dimensional point features.
[0019] In one embodiment, the step of performing feature transformation processing on the frustum features obtained by mapping the point cloud in the initial frustum to the frustum axis to obtain the second three-dimensional point features includes:
[0020] Perform feature channel number upsampling processing on the frustum features obtained by mapping the point cloud data in the initial frustum to the frustum axis through a multi-layer perceptron to obtain upsampled features;
[0021] Perform feature transformation processing on the upsampled features through a G2P module to obtain the second three-dimensional point features.
[0022] In one embodiment, the step of performing feature channel number transformation processing on the point cloud data in the initial frustum to obtain the first transformed point features includes:
[0023] Perform feature channel number transformation processing on the frustum features obtained by mapping the point cloud data in the initial frustum to the frustum axis through a multi-layer perceptron to obtain the first transformed point features.
[0024] In one embodiment, the step of fusing the first three-dimensional point features, the second three-dimensional point features, and the first transformed point features to obtain the first fused point features includes:
[0025] Input the first three-dimensional point features, the second three-dimensional point features, and the first transformed point features into a point fusion module for feature fusion processing to obtain the first fused point features.
[0026] In one embodiment, the step of decomposing the first fusion features and extracting object features to obtain object image features and object frustum features includes:
[0027] Decompose the first fusion feature through the P2G module to obtain an image view and a frustum view;
[0028] Extract object features from the image view through a 2D FCN to obtain the object image features;
[0029] Extract object features from the frustum view through a 1D FCN to obtain the object frustum features.
[0030] In one embodiment, a fusion method using an attention mechanism is adopted to refine the fusion granularity of the object image features and the object frustum features. Therefore, the step of fusing the second transformed feature, the object image features, and the object frustum features to obtain a second fusion feature includes:
[0031] Use the object frustum features as the Query, the object image features as the Key and Value. Concatenate the Query and the Key and input them into the attention head. Concatenate the feature output by the attention head with the Value to obtain an output feature. Input the output feature and the second transformed feature into a point fusion module to obtain the second fusion feature.
[0032] In one embodiment, the step of decomposing the first fusion feature through the P2G module to obtain an image view and a frustum view includes:
[0033] Obtain the image view through the following formula:
[0034]
[0035] where represents the image view, represents the floor operation, u k represents the abscissa, v k represents the ordinate, ks.t. represents the condition, c represents the number of channels, and k represents the index of the point whose corresponding two-dimensional position falls on the two-dimensional grid;
[0036] Obtain the frustum view through the following formula:
[0037]
[0038] where represents the frustum view, n k represents the frustum view coordinate.
[0039] In one embodiment, the model formula of the G2P module is:
[0040]
[0041] Among them, and represent the integers of the 2D grid coordinates of the projection of the k-th 3D point. And represents four adjacent neighbors, w i,j,k represents the weight of bilinear interpolation;
[0042]
[0043] In one embodiment, the step of using the object frustum feature as the Query, the object image feature as the Key and Value, concatenating the Query and the Key and then inputting them into the attention head, and concatenating the feature output by the attention head with the Value to obtain the second fusion feature is shown in the following formula:
[0044]
[0045] Among them, represents the output feature, q represents the query, k represents the key, v represents the value, N q represents the cross-modal neighbor, and d represents the vector dimension.
[0046] Compared with the traditional technology, the beneficial effects of this application are:
[0047] This application uses a frustum model to filter the point cloud data under the same viewing angle, and introduces the corresponding 2D image to make up for the defect that the point cloud lacks appearance information. For the problem of efficient fusion of point cloud and image, the present invention adopts a two-stage point fusion method, and obtains point features with richer semantic information through modules such as P2G, G2P, and point fusion. For the problem of finer-grained fusion of image and frustum features in the second-stage point fusion process, the present invention adopts a method based on the attention mechanism to efficiently fuse and interact the frustum and the image. Finally, the fused point features are input into the frustum convolution detection head for processing to obtain accurate predicted 3D bounding boxes and object categories. Common frustum-based 3D object detection methods are only based on single-modal point cloud features. The present invention innovatively introduces image features under the same viewing angle, and adopts a two-stage point fusion and attention mechanism to obtain point features with richer semantic information, which helps to achieve more accurate 3D object detection.
[0048] For better understanding and implementation, the present application will be described in detail below with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0049] Figure 1 is a step diagram of a 3D object detection method based on cascaded frustum image fusion according to an embodiment of the present application;
[0050] Figure 2Schematic diagram of frustum generation for the 3D object detection method based on cascaded frustum image fusion according to an embodiment of the present application;
[0051] Figure 3 Flow schematic diagram of the frustum axis for the 3D object detection method based on cascaded frustum image fusion according to an embodiment of the present application;
[0052] Figure 4 Schematic diagram of the attention mechanism for the 3D object detection method based on cascaded frustum image fusion according to an embodiment of the present application;
[0053] Figure 5 Effect schematic diagram of the 3D object detection box mapped to the RGB image for the 3D object detection method based on cascaded frustum image fusion according to an embodiment of the present application. Detailed implementation manners
[0054] To make the objectives, technical solutions, and advantages of the present application clearer, the following will further describe the embodiments of the present application in detail in conjunction with the accompanying drawings.
[0055] It should be clear that the described embodiments are only part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts belong to the scope protected by the embodiments of the present application.
[0056] When the following description involves the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. In the description of the present application, it should be understood that the terms "first", "second", "third", etc. are only used to distinguish similar objects, and do not have to be used to describe a specific order or sequence, nor can they be understood as indicating or implying relative importance. For those of ordinary skill in the art, the specific meanings of the above terms in the present application can be understood according to specific circumstances. The singular forms of "a", "the", and "said" used in the present application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. The word "if" / "when" used herein can be interpreted as "when" or "when" or "in response to a determination".
[0057] In addition, in the description of the present application, unless otherwise specified, "a plurality of" means two or more. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. These three situations. The character " / " generally represents an "or" relationship between the associated objects before and after.
[0058] Please refer to Figures 1-5, which is a 3D object detection method based on cascaded frustum image fusion in this application, includes:
[0059] S1: Extract object images from the initial two-dimensional images.
[0060] S2: Perform two-dimensional feature extraction and feature transformation on the object images to obtain the first three-dimensional point features.
[0061] S3: Perform feature transformation on the frustum features obtained by mapping the point cloud in the initial frustum onto the frustum axis to obtain the second three-dimensional point features.
[0062] Among them, the initial frustum is obtained by using the projection transformation matrix to project the 2D detection box on the initial two-dimensional image into the 3D space according to the 2D detection box on the initial two-dimensional image, and then according to the optical characteristics of the camera, by connecting the rays between the corner points of the 2D bounding box and the camera imaging plane, the near plane and the far plane are delimited along the camera optical axis, and finally a geometric body similar to a quadrangular pyramid is obtained, which is called the initial frustum.
[0063] S4: Perform feature channel number transformation on the point cloud data in the initial frustum to obtain the first transformed point features.
[0064] S5: Fuse the first three-dimensional point features, the second three-dimensional point features, and the first transformed point features to obtain the first fused point features.
[0065] S6: Decompose the first fusion feature and extract object features to obtain object image features and object frustum features.
[0066] S7: Perform feature channel number transformation on the first fusion feature to obtain the second transformed feature.
[0067] S8: Fuse the second transformed feature, the object image features, and the object frustum features to obtain the second fusion feature.
[0068] S9: Input the second fusion feature into the detection head to obtain a 3D object detection box; the 3D object detection box is used to indicate the 3D object in the world coordinate system.
[0069] Among them, the world coordinate system is used to describe the positions of objects or points in the real world. The 2D image is located in the image coordinate system. The world coordinate system and the image coordinate system can be mutually transformed through the projection matrix. The schematic diagram of the effect of the 3D object detection box of the 3D object detection method based on cascaded frustum image fusion in this embodiment mapped to the RGB image is as Figure 5 shown.
[0070] In a feasible embodiment, the step of performing two-dimensional feature extraction and feature transformation processing on the object image to obtain the first three-dimensional point features includes:
[0071] Extract the two-dimensional features of the object image through 2D FCN;
[0072] Perform feature transformation processing on the two-dimensional features through the G2P module to obtain the first three-dimensional point features.
[0073] In a feasible embodiment, the step of performing feature transformation processing on the cone features obtained by mapping the point cloud in the initial frustum to the frustum axis to obtain the second three-dimensional point features includes:
[0074] Perform dimensionality increase processing on the feature channels of the cone features obtained by mapping the initial two-dimensional image to the frustum axis through a multi-layer perceptron to obtain the dimensionality-increased features;
[0075] Perform feature transformation processing on the dimensionality-increased features through the G2P module to obtain the second three-dimensional point features.
[0076] In a feasible embodiment, the step of performing feature channel number transformation processing on the point cloud data in the initial frustum to obtain the first transformed point features includes:
[0077] Perform feature channel number transformation processing on the cone features obtained by mapping the initial two-dimensional image to the frustum axis through a multi-layer perceptron to obtain the first transformed point features.
[0078] In a feasible embodiment, the step of fusing the first three-dimensional point features, the second three-dimensional point features, and the first transformed point features to obtain the first fused point features includes:
[0079] Input the first three-dimensional point features, the second three-dimensional point features, and the first transformed point features into a point fusion module for feature fusion processing to obtain the first fused point features.
[0080] In a feasible embodiment, the step of decomposing the first fusion feature and extracting object features to obtain object image features and object frustum features includes:
[0081] Decompose the first fusion feature through the P2G module to obtain an image view and a frustum view;
[0082] Extract object features from the image view through 2D FCN to obtain the object image features;
[0083] Extract object features from the frustum view through 1D FCN to obtain the object frustum features.
[0084] In a feasible embodiment, an attention mechanism-based fusion method is adopted to refine the fusion granularity of the object image feature and the object frustum feature. Therefore, the step of fusing the second transformed feature, the object image feature, and the object frustum feature to obtain a second fusion feature includes:
[0085] Using the object frustum feature as the Query, the object image feature as the Key and Value, concatenating the Query and the Key and inputting them into the attention head, concatenating the feature output by the attention head with the Value to obtain an output feature, and inputting the output feature and the second transformed feature into a point fusion module to obtain the second fusion feature.
[0086] In a feasible embodiment, the step of decomposing the first fusion feature by the P2G module to obtain an image view and a frustum view includes:
[0087] The image view is obtained through the following formula:
[0088]
[0089] where, represents the image view, represents the floor operation, u k represents the abscissa, v k represents the ordinate, ks.t. represents the condition, c represents the number of channels, and k represents the index of the point whose corresponding two-dimensional position falls on the two-dimensional grid;
[0090] The frustum view is obtained through the following formula:
[0091]
[0092] where, represents the frustum view, n k represents the frustum view coordinate.
[0093] In a feasible embodiment, the model formula of the G2P module is:
[0094]
[0095] where, and represent the integers of the 2D grid coordinates of the projection of the k-th 3D point. While represents four adjacent neighbors, w i,j,k represents the weight of bilinear interpolation;
[0096]
[0097] In a feasible embodiment, the step of using the object frustum feature as the Query, using the object image feature as the Key and Value, cascading the Query and the Key and then inputting them into the attention head, and cascading the feature output by the attention head with the Value to obtain the second fusion feature is shown by the following formula:
[0098]
[0099] Wherein, represents the output feature, q represents the query, k represents the key, v represents the value, N q represents the cross-modal neighbor, and d represents the vector dimension.
[0100] Compared with the traditional technology, the beneficial effects of this application are as follows:
[0101] This application uses a frustum model to screen point cloud data under the same viewing angle and introduces the corresponding 2D image to make up for the defect that the point cloud lacks appearance information. Aiming at the problem of efficient fusion of point cloud and image, the present invention adopts a two-stage point fusion method, and obtains point features with richer semantic information through modules such as P2G, G2P, and point fusion. Aiming at the problem of finer-grained fusion of image and frustum features in the second-stage point fusion process, the present invention adopts a method based on the attention mechanism to efficiently fuse and interact the frustum and the image. Finally, the fused point features are input into the frustum convolution detection head for processing to obtain accurate predicted 3D bounding boxes and object categories. Common frustum-based 3D object detection methods are only based on single-modal point cloud features. The present invention innovatively introduces image features under the same viewing angle and adopts a two-stage point fusion and attention mechanism to obtain point features with richer semantic information, which helps to achieve more accurate 3D object detection.
[0102] The device embodiments described above are merely illustrative. The components described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this application. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0103] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.
[0104] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the selected functions in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks. These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the selected functions in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.
[0105] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide steps for implementing the selected functions in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.
[0106] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0107] The memory may include non-permanent memory in the form of computer-readable media, random access memory (RAM), and / or non-volatile memory such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.
[0108] Computer readable media include permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. Information can be computer readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disk read-only memory (CD-ROM), digital versatile disk (DVD) or other optical storage, magnetic cassettes, magnetic tape disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer readable media does not include temporary computer readable media (transitory media), such as modulated data signals and carrier waves.
[0109] It should also be noted that the terms "include", "comprises" or any other variations thereof are intended to cover non-exclusive inclusion, so that a process, method, commodity or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity or device. In the absence of more restrictions, the elements defined by the sentence "comprises a ..." do not exclude the existence of other identical elements in the process, method, commodity or device including the elements.
[0110] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application should be included in the scope of the claims of the present application.
Claims
1. A 3D object detection method based on cascaded frustum image fusion, characterized in that Including: Performing object extraction on an initial two-dimensional image to obtain an object image; Performing two-dimensional feature extraction and feature transformation processing on the object image to obtain first three-dimensional point features; Performing feature transformation processing on the cone features obtained by mapping the point cloud in the initial view frustum to the view frustum axis to obtain second three-dimensional point features; Performing feature channel number transformation processing on the point cloud data in the initial view frustum to obtain first transformed point features; Fusing the first three-dimensional point features, the second three-dimensional point features, and the first transformed point features to obtain first fused point features; Decomposing the first fused feature and extracting object features to obtain object image features and object view frustum features; Performing feature channel number transformation processing on the first fused feature to obtain second transformed features; Fusing the second transformed features, the object image features, and the object view frustum features to obtain second fused features; Inputting the second fused features into a detection head to obtain 3D object detection boxes; the 3D object detection boxes are used to indicate 3D objects in the world coordinate system.
2. The 3D object detection method based on cascaded frustum image fusion according to claim 1, wherein, The step of performing two-dimensional feature extraction and feature transformation processing on the object image to obtain first three-dimensional point features includes: Extracting two-dimensional features of the object image through a 2D FCN; Performing feature transformation processing on the two-dimensional features through a G2P module to obtain the first three-dimensional point features.
3. The 3D object detection method based on cascaded frustum image fusion according to claim 1, wherein The step of performing feature transformation processing on the cone features obtained by mapping the point cloud in the initial view frustum to the view frustum axis to obtain second three-dimensional point features includes: Performing feature channel number upsampling processing on the cone features obtained by mapping the point cloud data in the initial view frustum to the view frustum axis through a multi-layer perceptron to obtain upsampled features; Performing feature transformation processing on the upsampled features through a G2P module to obtain the second three-dimensional point features.
4. The 3D object detection method based on cascaded frustum image fusion according to claim 1, characterized in that, The step of performing feature channel number transformation processing on the point cloud data in the initial view frustum to obtain first transformed point features includes: Performing feature channel number transformation processing on the cone features obtained by mapping the point cloud data in the initial view frustum to the view frustum axis through a multi-layer perceptron to obtain the first transformed point features.
5. The 3D object detection method based on cascaded frustum image fusion according to claim 1, wherein The step of fusing the first three-dimensional point features, the second three-dimensional point features, and the first transformed point features to obtain first fused point features includes: Inputting the first three-dimensional point features, the second three-dimensional point features, and the first transformed point features into a point fusion module for feature fusion processing to obtain the first fused point features.
6. The 3D object detection method based on cascaded frustum image fusion according to claim 1, wherein The step of decomposing the first fused feature and extracting object features to obtain object image features and object view frustum features includes: Decomposing the first fused feature through a P2G module to obtain an image view and a view frustum view; Extracting object features from the image view through a 2D FCN to obtain the object image features; Extracting object features from the view frustum view through a 1D FCN to obtain the object view frustum features.
7. The 3D object detection method based on cascaded frustum image fusion according to claim 1, characterized in that The step of fusing the second transformed features, the object image features, and the object view frustum features to obtain second fused features includes: Taking the object frustum feature as the Query, taking the object image feature as the Key and Value, concatenating the Query and the Key and inputting the result into the attention head, concatenating the feature output by the attention head with the Value to obtain the output feature, and inputting the output feature and the second transformed feature into the point fusion module to obtain the second fused feature.
8. The 3D object detection method based on cascaded frustum image fusion according to claim 6, characterized in that The step of decomposing the first fused feature into an image view and a frustum view through the P2G module includes: Obtaining the image view through the following formula: Among them, represents an image view, represents a floor operation, u k represents the abscissa, v k represents the ordinate, ks.t. represents the condition, c represents the number of channels, and k represents the index of the point whose corresponding two-dimensional position falls on the two-dimensional grid; Obtaining the frustum view through the following formula: Among them, represents the frustum view, and n k represents the frustum view coordinates.
9. The 3D object detection method based on cascaded frustum image fusion according to claim 2 or 3, characterized in that The model formula of the G2P module is: Among them, and represent the integers of the 2D grid coordinates of the projection of the k-th 3D point; while represents four adjacent neighbors, and w i,j,k represents the weights of bilinear interpolation.
10. The 3D object detection method based on cascaded frustum image fusion according to claim 8, wherein The step of taking the object frustum feature as the Query, taking the object image feature as the Key and Value, concatenating the Query and the Key and inputting the result into the attention head, concatenating the feature output by the attention head with the Value to obtain the second fused feature is shown in the following formula: Among them, represents the output feature, q represents the query, k represents the key, v represents the value, and N q represents the cross-modal neighbor, and d represents the vector dimension.