3D multi-target detection method based on global feature enhancement and false negative correction

By adopting global feature enhancement and false negative correction methods in 3D object detection, using sliding window attention module and multi-stage heat map encoder, combined with adaptive dynamic area relocation and cropping and APM system, the problem of false negative problems and insufficient utilization of global features in complex environments is solved, and more efficient and accurate 3D multi-object detection is achieved.

CN119942076AInactive Publication Date: 2025-05-06GUILIN UNIV OF ELECTRONIC TECH

Patent Information

Application Number
CN202510012161.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing 3D object detection technology is prone to false negative problems in complex environments, and lacks an effective global feature learning mechanism, which affects the detection effect.

Method used

Using a 3D multi-objective detection method based on global feature enhancement and false negative correction, the sliding window attention module and multi-stage heat map encoder are combined with an adaptive dynamic area relocation and cropping and cumulative positive sample mask (APM) system to enhance local feature representation and global context fusion to reduce false negative errors.

Benefits of technology

It improves the accuracy and robustness of 3D multi-object detection, reduces false positives and missed detection, enhances the understanding of complex scenarios, and improves detection performance and generalization capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942076A_ABST
    Figure CN119942076A_ABST
Patent Text Reader

Abstract

The invention discloses a 3D multi-target detection method based on global feature enhancement and false negative correction, and the method comprises the steps: firstly obtaining point cloud data from a LiDAR sensor, converting the point cloud data into a regular voxel grid, and extracting the features of a bird's-eye view (BEV); a sliding window attention module is designed, and self-adaptive dynamic region relocation cutting is combined, so that the features of each region interact with other regions, the local feature representation capability is enhanced, and global context information fusion is promoted. The method specifically comprises the steps of region division, self-adaptive dynamic region relocation cutting, region self-attention and application of a sliding region attention mechanism so as to capture interaction among different regions; a parallel multi-stage heat map encoder is constructed to decode a center heat map from the BEV features and project the center heat map to a BEV view. The peak value of the heat map corresponds to a potential target position, the first k most significant target features are identified by analyzing intensity distribution, and accurate positioning is ensured; and meanwhile, an accumulated false positive management (APM) system is introduced, a mask pattern is generated on the basis of each layer of heat map, a detection result is updated by combining an upper-layer mask pattern and a current heat map, new first k highest peak instance features are selected, misinformation and missing detection are reduced, and the detection precision is improved. And finally, enhancing the characterization capability of the instance in the global situation through a multi-head self-attention and local cross attention mechanism, and finally optimizing BEV features to predict a 3D bounding box.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of 3D target detection, and in particular to a 3D multi-target detection method based on global feature enhancement and false negative correction. Background Art

[0002] In today's computer vision field, 3D object detection tasks play a vital role in many key application scenarios such as autonomous driving, robot navigation, augmented reality, etc. With the continuous development of technology, although many excellent 3D object detection models have been proposed, there are still some challenging problems that need to be solved, among which the false negative problem and the effective extraction and utilization of global features are particularly prominent. These problems affect the accuracy and reliability of the inspection system in practical applications.

[0003] 1. False Negative Problem

[0004] In a factory environment, unmanned vehicles need to identify various fixed or moving obstacles, such as equipment, tool boxes, personnel, etc., to ensure safe navigation. However, in the complex and changing factory environment, traditional 3D object detection methods, including some widely used models, often have false negatives when faced with these complex situations, that is, they fail to detect the actual target objects. This is mainly attributed to the following factors:

[0005] 1. Occlusion problem: When the target object is partially or completely occluded by other objects, its visible feature information is greatly reduced, making it difficult for the model to accurately identify it. For example, in an autonomous driving scenario, a vehicle may be occluded by pedestrians, buildings, or other vehicles, making it impossible for the model based on traditional feature extraction and detection methods to fully capture the features of the target vehicle, thereby missing the target and generating false negative results, which poses a serious threat to the safety of the autonomous driving system.

[0006] 2. Limitations of feature extraction: Existing feature extraction methods may not be able to fully exploit the unique features of target objects in complex environments, especially for those targets that are similar to the background or in low-contrast areas. For example, in some cases where the light is dim or the surface texture of the target object is not obvious, it is difficult for the model to extract discriminative features from point cloud data or image data, causing some target objects to be misjudged as background during the detection process, leading to false negative errors.

[0007] 3. Singleness of detection algorithm: Most traditional detection algorithms use relatively fixed detection strategies and threshold settings when processing target objects, and lack the ability to adapt to different scenarios and target types. This makes it impossible for the model to flexibly adjust the detection parameters in some special cases, such as when the target object is deformed, has a rare posture, or is at a long distance, so it is impossible to accurately detect the target, increasing the probability of false negatives.

[0008] 2. Insufficient use of global features

[0009] In order to achieve efficient and reliable inspections, unmanned vehicles not only need to accurately locate a single target, but also need to understand the layout of the entire scene, including the relative position relationship between devices. Current methods have the following limitations:

[0010] 1. Local feature-dominated detection: Many existing 3D object detection models rely too much on local features. For example, in point cloud processing, they focus on feature analysis of a single point or a local point cloud area, while ignoring the global relationship between different target objects in the scene and between the target and the environment. This local feature-dominated detection method makes it difficult for the model to grasp the semantic information and spatial layout of the scene as a whole when facing complex scenes, thus affecting the accurate classification and positioning of the target object.

[0011] 2. Lack of effective global feature learning mechanism: The current model lacks an effective mechanism for learning global features and cannot fully tap into the potential semantic information and spatial relationships in the scene. For example, when processing large-scale point cloud data, it is difficult for the model to automatically learn the association patterns between point clouds in different regions, as well as the intrinsic relationship between these patterns and the target objects, making it impossible to use global features to assist in target detection, reducing the detection performance and generalization ability of the model in complex scenes.

[0012] In summary, the false negative problem and insufficient utilization of global features are the key factors affecting the effect of 3D target detection. Therefore, it is particularly urgent to develop a new method that can enhance the feature representation capability and effectively reduce false negatives. By improving feature extraction technology, optimizing detection algorithms, and introducing smarter learning mechanisms, the safety and efficiency of unmanned vehicle inspections can be significantly improved, thereby promoting the development of intelligent manufacturing. Therefore, it is necessary to provide a 3D multi-target detection method based on global feature enhancement and false negative correction to solve the above problems. Summary of the invention

[0013] The purpose of the present invention is to provide a 3D multi-target detection method based on global feature enhancement and false negative correction, which solves the false negative problem and insufficient scene fusion existing in the existing 3D target detection technology.

[0014] To achieve the above object, the present invention provides a 3D multi-target detection method based on global feature enhancement and false negative correction, comprising:

[0015] S1: Obtain the original point cloud data of the target scene from the LiDAR camera sensor, ensuring the integrity and accuracy of the data, covering the target object information under different angles, distances and lighting conditions.

[0016] S2: Convert sparse point cloud data into a regular voxel grid and use the VoxelNext sparse convolutional network to efficiently extract features. Subsequently, these voxel features are compressed to form a bird's-eye view (BEV) feature representation.

[0017] S3: Design a sliding window attention module, combined with adaptive dynamic region relocation cropping, so that the features of each region interact with the features of other regions, enhance the representation ability of local features, and promote the fusion of global context information.

[0018] S4: Construct a parallel multi-stage heatmap encoder to decode the central heatmap from the bird's-eye view (BEV) features and project it to the BEV view. The heatmap peak corresponds to the potential target location, and the top k most significant target features are identified by analyzing the intensity distribution to ensure accurate positioning. A mask map is generated based on each layer of heatmap and accumulated to the APM system. The detection results are updated by combining the upper layer mask map and the current heatmap, and the new top K highest peak instance features are selected to improve detection accuracy and reduce false positives and missed detections.

[0019] S5: The instance features generated by the multi-stage heat map encoder are fused through the multi-head self-attention mechanism to learn the correlation and difference between instances. The enhanced instance features are combined with the global bird's-eye view features as a query group, and their global representation ability is enhanced through the local cross-attention mechanism. Finally, the bird's-eye view (BEV) features are optimized to predict the 3D bounding box.

[0020] Preferably, in step S3, the following steps are specifically included:

[0021] S3.1: Divide the EBV feature map into multiple non-overlapping local regions, each of which consists of M×M grid cells The subset described represents a part of the BEV feature. Each region is considered as a whole and the regional self-attention f is set. SA (·) Information fusion is performed on all regions. The formula is as follows:

[0022]

[0023] in, Yes MSA (·) The resulting grid cell.

[0024] S3.2: Design adaptive dynamic region relocation cropping to redefine the region to facilitate the subsequent capture of the interaction between different regions. Specifically, after self-attention processing, The whole is moved horizontally and vertically by M / 2 grid units to generate a new moving area. in Indicates the new grid index after displacement.

[0025] S3.3: Introducing sliding area attention f S-MSA (·), feature learning is performed on each new moving area containing M×M grid features after movement to capture the features in the area and the interaction with features in other areas, thus enhancing the global feature representation capability of BEV. The formula is as follows:

[0026]

[0027] This makes each mesh interact with meshes from different regions before sliding, thus capturing long-range dependencies. To obtain the rich BEV feature map B F ∈R W ×H×C .

[0028] Preferably, in step S4, the following steps are specifically included:

[0029] S4.1: A multi-level heat map detection head is used to predict objects on multi-level BEV features, where each heat map peak represents the center position of an object, i.e., the instance candidate position. The generation of multi-level BEV features is completed in a cascade manner, using lightweight reverse residual blocks between levels.

[0030] S4.2: A multi-stage hard instance detection head is introduced to identify missed objects through a series of filtering operations such as matching instance candidates on the heat map with real objects, and guide the model to focus on mining hard-to-detect objects. Specifically, firstly, the real instance is annotated as O = {o i , i=1,2,…}, using the thermal map to detect the head in B F The center position of the instance is predicted and some instance candidates are generated. Then the instance candidate P detected in the kth stage is k ={p i , i=1,2,…} and the real annotated objects are matched σ(·,·), and the matched object candidates are called true samples (TP), and the unmatched objects are called false negatives (FN). The object matching formula is as follows:

[0031]

[0032] S4.3: Using an Accumulated Positive Mask (APM) The positions of the true samples in the previous stage are masked on the heat map, thereby ignoring objects that are easy to detect and focusing on false negative objects. The specific steps include:

[0033] S4.3.1: Design a zero matrix of the same size as the BEV feature map as the initial APM, and then use a pool-based mask to predict the first k objects p in this stage i ∈P k Set the mask value to get the mask matrix M i , where the area where the successfully matched instance candidate object is located is marked as 1, i.e. M (x,y,c) =1 (the center point of the smaller instance is filled with 1, while the larger instance is filled with a mask value of 1 in the 3x3 area outward from the center point of the instance).

[0034] S4.3.2: Simply accumulate the prior mask matrix to get the following:

[0035]

[0036] By shielding BEV heat map S k and

[0037]

[0038] in, The k instance sample features obtained by this cascade are then formed into a feature group Q list .

[0039] Preferably, in step S5, the following steps are specifically included:

[0040] S5.1: Enhance the candidate groups of k instances in the first N stages through the multi-head self-attention mechanism The instance correlation and difference between features. The multi-head self-attention mechanism can be designed as:

[0041] SA(Q list )=Self_Attention(q1,q2,…,q i×j ,…,q N×k ) (7)

[0042] Among them, q i×j represents the jth instance feature in the i-th stage. Then the result SA(Q list ) that is, Q as query, multi-level BEV feature B F As key and value input into the local variable attention mechanism, to enhance the feature representation ability of all instances in the world. The local variable attention mechanism can be designed as:

[0043]

[0044] S5.2: The final BEV features are input into the feed-forward neural network FFN to predict the specific attributes of each instance, including center point, offset, size, orientation, etc., and perform category and bounding box prediction to achieve 3D multi-target detection.

[0045] S5.3: Design loss function L lloss , respectively calculate the loss from the three aspects of bounding box and category information, and update the weights in model training through gradient backpropagation, thereby achieving high detection accuracy and comprehensive detection. The loss function is designed as follows:

[0046]

[0047] Among them, L class represents the instance category loss, L bbox Represents the 3D detection box loss of the instance.

[0048] Therefore, the present invention adopts the above-mentioned 3D multi-target detection method based on global feature enhancement and false negative correction, which has the following beneficial effects:

[0049] (1) The adaptive dynamic region relocation cropping designed in the present invention enhances the local feature representation and global context fusion. Through the sliding window attention module, the understanding of instance targets of various forms and scales is enhanced, thereby improving the robustness and adaptability of the model.

[0050] (2) The present invention designs multi-stage heat map encoding and false negative correction, adopts mask mechanism and APM system, reduces false alarms and missed detections, prevents the occurrence of false negatives, and improves detection precision and accuracy.

[0051] The technical solution of the present invention is further described in detail below through the accompanying drawings and embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The drawings constituting a part of the present application are used to provide a further understanding of the present application. The illustrative embodiments and descriptions of the present application are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0053] Figure 1 It is a framework diagram of a 3D multi-target detection method based on global feature enhancement and false negative correction of the present invention;

[0054] Figure 2 It is a processing block diagram of the sliding window attention module of the present invention. DETAILED DESCRIPTION

[0055] The technical solution of the present invention is further described below through the accompanying drawings and embodiments.

[0056] Unless otherwise defined, technical or scientific terms used in the present invention shall have the common meanings understood by one having ordinary skills in the field to which the present invention belongs.

[0057] The words "include" or "comprises" and the like used in the present invention mean that the elements before the word include the elements listed after the word, and do not exclude the possibility of also including other elements. The orientation or position relationship indicated by the terms "inside", "outside", "upper", "lower", etc. is based on the orientation or position relationship shown in the drawings, which is only for the convenience of describing the present invention and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, it cannot be understood as a limitation of the present invention. When the absolute position of the described object changes, the relative position relationship may also change accordingly. In the present invention, unless otherwise clearly specified and limited, the terms "attachment" and the like should be understood in a broad sense, for example, it can be a fixed connection, a detachable connection, or an integral body; it can be directly connected, or indirectly connected through an intermediate medium, and it can be the internal connection of two elements or the interaction relationship between two elements. For ordinary technicians in this field, the specific meanings of the above terms in the present invention can be understood according to specific circumstances.

[0058] like Figure 1 , Figure 2 As shown, the present invention provides a 3D multi-target detection method based on global feature enhancement and false negative correction, which mainly includes the following steps:

[0059] S1: Obtain the original point cloud data of the target scene from the LiDAR camera sensor, ensuring the integrity and accuracy of the data, covering the target object information under different angles, distances and lighting conditions.

[0060] S2: Convert sparse point cloud data into a regular voxel grid and use the VoxelNext sparse convolutional network to efficiently extract features. Subsequently, these voxel features are compressed to form a bird's-eye view (BEV) feature representation.

[0061] S3: Design a sliding window attention module, combined with adaptive dynamic region relocation cropping, so that the features of each region interact with the features of other regions, enhance the representation ability of local features, and promote the fusion of global context information.

[0062] S4: Construct a parallel multi-stage heatmap encoder to decode the central heatmap from the BEV features and project it to the BEV view. The heatmap peak corresponds to the potential target location, and the top k most significant target features are identified by analyzing the intensity distribution to ensure accurate positioning. A mask map is generated based on each layer of heatmap and accumulated to the APM system. The detection results are updated by combining the upper layer mask map and the current heatmap, and the new top K highest peak instance features are selected to improve detection accuracy and reduce false positives and missed detections.

[0063] S5: The instance features generated by the multi-stage heat map encoder are fused through the multi-head self-attention mechanism to learn the correlation and difference between instances. The enhanced instance features are combined with the global bird's-eye view features as a query group, and their global representation ability is enhanced through the local cross-attention mechanism. Finally, the bird's-eye view (BEV) features are optimized to predict the 3D bounding box.

[0064] Preferably, in step S3, the following steps are specifically included:

[0065] S3.1: Divide the EBV feature map into multiple non-overlapping local regions, each of which consists of M×M grid cells The subset described represents a part of the BEV feature. Each region is considered as a whole and the regional self-attention f is set. SA (·) Information fusion is performed on all regions. The formula is as follows:

[0066]

[0067] in, Yes MSA (·) The resulting grid cell.

[0068] S3.2: Design adaptive dynamic region relocation cropping to redefine the region to facilitate the subsequent capture of the interaction between different regions. Specifically, after self-attention processing, The whole is moved horizontally and vertically by M / 2 grid units to generate a new moving area. in Indicates the new grid index after displacement.

[0069] S3.3: Introducing sliding area attention f S-MSA (·), feature learning is performed on each new moving area containing M×M grid features after movement to capture the features in the area and the interaction with features in other areas, thus enhancing the global feature representation capability of BEV. The formula is as follows:

[0070]

[0071] This makes each mesh interact with meshes from different regions before sliding, thus capturing long-range dependencies. To obtain the rich BEV feature map B F ∈R W ×H×C .

[0072] Preferably, in step S4, the following steps are specifically included:

[0073] S4.1: A multi-level heat map detection head is used to predict objects on multi-level BEV features, where each heat map peak represents the center position of an object, i.e., the instance candidate position. The generation of multi-level BEV features is completed in a cascade manner, using lightweight reverse residual blocks between levels.

[0074] S4.2: A multi-stage hard instance detection head is introduced to identify missed objects through a series of filtering operations such as matching instance candidates on the heat map with real objects, and guide the model to focus on mining hard-to-detect objects. Specifically, firstly, the real instance is annotated as O = {o i , i=1,2,…}, using the thermal map to detect the head in B F The center position of the instance is predicted and some instance candidates are generated. Then the instance candidate P detected in the kth stage is k ={p i , i=1,2,…} and the real annotated objects are matched σ(·,·), and the matched object candidates are called true samples (TP), and the unmatched objects are called false negatives (FN). The object matching formula is as follows:

[0075]

[0076] S4.3: Using an Accumulated Positive Mask (APM) The positions of the true samples in the previous stage are masked on the heat map, thereby ignoring objects that are easy to detect and focusing on false negative objects. The specific steps include:

[0077] S4.3.1: Design a zero matrix of the same size as the BEV feature map as the initial APM, and then use a pool-based mask to predict the first k objects p in this stage i ∈P k Set the mask value to get the mask matrix M i , where the area where the successfully matched instance candidate object is located is marked as 1, i.e. M (x,y,c) =1 (the center point of the smaller instance is filled with 1, while the larger instance is filled with a mask value of 1 in the 3x3 area outward from the center point of the instance).

[0078] S4.3.2: Simply accumulate the prior mask matrix to get the following:

[0079]

[0080] By shielding BEV heat map S k and

[0081]

[0082] in, The k instance sample features obtained by this cascade are then formed into a feature group Q list .

[0083] Preferably, in step S5, the following steps are specifically included:

[0084] S5.1: Enhance the candidate groups of k instances in the first N stages through the multi-head self-attention mechanism The instance correlation and difference between features. The multi-head self-attention mechanism can be designed as:

[0085] SA(Q list )=Self_Attention(q1,q2,…,q i×j ,…,q N×k ) (7)

[0086] Among them, q i×j represents the jth instance feature in the i-th stage. Then the result SA(Q list ) that is, Q as query, multi-level BEV feature B F As key and value input into the local variable attention mechanism, to enhance the feature representation ability of all instances in the world. The local variable attention mechanism can be designed as:

[0087]

[0088] S5.2: The final BEV features are input into the feed-forward neural network FFN to predict the specific attributes of each instance, including center point, offset, size, orientation, etc., and perform category and bounding box prediction to achieve 3D multi-target detection.

[0089] S5.3: Design loss function L lloss , respectively calculate the loss from the three aspects of bounding box and category information, and update the weights in model training through gradient backpropagation, thereby achieving high detection accuracy and comprehensive detection. The loss function is designed as follows:

[0090]

[0091] Among them, L class represents the instance category loss, L bbox Represents the 3D detection box loss of the instance.

[0092] Therefore, the present invention adopts the above-mentioned 3D multi-target detection method based on global feature enhancement and false negative correction, and enhances the instance scene information representation capability by using a sliding window area grid attention module in the bird's-eye view data processing process; at the same time, by constructing a parallel multi-stage heat map encoder and introducing an accumulated false positive management (APM) system, a mask mechanism is used to achieve high-precision 3D multi-target detection. This method not only enhances the accuracy of target positioning, but also effectively solves the false negative problem by iteratively updating the mask map, thereby improving the detection performance in complex scenes as a whole.

[0093] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that they can still modify or replace the technical solution of the present invention with equivalents, and these modifications or equivalent replacements cannot cause the modified technical solution to deviate from the spirit and scope of the technical solution of the present invention.

Claims

1. A 3D multi-target detection method based on global feature enhancement and false negative correction, characterized in that: The following steps are involved: S1: Obtain the original point cloud data of the target scene from the LiDAR camera sensor, ensuring the integrity and accuracy of the data, covering the target object information under different angles, distances and lighting conditions. S2: Convert sparse point cloud data into a regular voxel grid and use the VoxelNext sparse convolutional network to efficiently extract features. Subsequently, these voxel features are compressed to form a bird's-eye view (BEV) feature representation. S3: Design a sliding window attention module, combined with adaptive dynamic region relocation cropping, so that the features of each region interact with the features of other regions, enhance the representation ability of local features, and promote the fusion of global context information. S4: Construct a parallel multi-stage heatmap encoder to decode the central heatmap from the bird's-eye view (BEV) features and project it to the BEV view. The heatmap peak corresponds to the potential target location, and the top k most significant target features are identified by analyzing the intensity distribution to ensure accurate positioning. A mask map is generated based on each layer of heatmap and accumulated to the APM system. The detection results are updated by combining the upper layer mask map and the current heatmap, and the new top K highest peak instance features are selected to improve detection accuracy and reduce false positives and missed detections. S5: The instance features generated by the multi-stage heat map encoder are fused through the multi-head self-attention mechanism to learn the correlation and difference between instances. The enhanced instance features are combined with the global bird's-eye view features as a query group, and their global representation ability is enhanced through the local cross-attention mechanism. Finally, the bird's-eye view (BEV) features are optimized to predict the 3D bounding box.

2. A 3D multi-target detection method with global feature enhancement and false negative correction according to claim 1, characterized in that: The specific method of step S3 is as follows: S31: Divide the EBV feature map into multiple non-overlapping local regions, each of which is composed of M×M grid cells. The subset described represents a part of the BEV feature. Each region is considered as a whole and the regional self-attention f is set. SA (·) Information fusion is performed on all regions. The formula is as follows: in, Yes MSA (·) The resulting grid cell. S32: Design adaptive dynamic region relocation cropping to redefine the region to facilitate the subsequent capture of the interaction between different regions. Specifically, after self-attention processing, The whole is moved horizontally and vertically by M / 2 grid units to generate a new moving area. where {Δ1,Δ2,…,Δ M2 } represents the new grid index after displacement. S33: Introducing sliding area attention f S-MSA (·), feature learning is performed on each new moving area containing M×M grid features after movement to capture the features in the area and the interaction with features in other areas, thus enhancing the global feature representation capability of BEV. The formula is as follows: This makes each mesh interact with meshes from different regions before sliding, thus capturing long-range dependencies. To obtain the rich BEV feature map B F ∈R W×H×C .

3. The 3D multi-target detection method with global feature enhancement and false negative correction according to claim 1, characterized in that: The specific method of step S4 is as follows: S41: A multi-level heat map detection head is used to predict objects on multi-level BEV features, where each heat map peak represents the center position of an object, i.e., the instance candidate position. The generation of multi-level BEV features is completed in a cascade manner, using lightweight reverse residual blocks between levels. S42: A multi-stage hard instance detection head is introduced to identify missed objects through a series of filtering operations such as matching instance candidates on the heat map with real objects, and guide the model to focus on mining hard-to-detect objects. Specifically, firstly, the real instance is annotated as O = {o i , i=1,2,…}, using the thermal map to detect the head in B F The center position of the instance is predicted and some instance candidates are generated. Then the instance candidate P detected in the kth stage is k ={p i , i=1,2,…} and the real annotated objects are matched σ(·,·), and the matched object candidates are called true samples (TP), and the unmatched objects are called false negatives (FN). The object matching formula is as follows: S43: Using a cumulative positive sample mask The positions of the true samples in the previous stage are masked on the heat map, thereby ignoring objects that are easy to detect and focusing on false negative objects. The specific steps include: S431: Design a zero matrix of the same size as the BEV feature map as the initial APM, and then use a pool-based mask to predict the first k objects p in this stage i ∈P k Set the mask value to get the mask matrix M i , where the area where the successfully matched instance candidate object is located is marked as 1, i.e. M (x,y,c) =1 (the center point of the smaller instance is filled with 1, while the larger instance is filled with a mask value of 1 in the 3x3 area outward from the center point of the instance). S432: Simply accumulate the prior mask matrix to obtain the following: By shielding BEV heat map S k and in, The k instance sample features obtained by this cascade are then formed into a feature group Q list .

4. The 3D multi-target detection method with global feature enhancement and false negative correction according to claim 1, characterized in that: The specific method of step S5 is as follows: S51: Enhance the candidate groups of k instances in the first N stages through the multi-head self-attention mechanism The instance correlation and difference between features. The multi-head self-attention mechanism can be designed as: SA(Q list )=Self_Attention(q1,q2,…,q i×j ,…,q N×k ) (7) Among them, q i×j represents the jth instance feature in the i-th stage. Then the result SA(Q list ) that is, Q as query, multi-level BEV feature B F As key and value input into the local variable attention mechanism, it can enhance the feature representation ability of all instances in the world. The local variable attention mechanism can be designed as: S52: The final BEV features are input into the feed-forward neural network FFN to predict the specific attributes of each instance, including the center point, offset, size, orientation, etc., and perform category and bounding box prediction to achieve 3D multi-target detection. S53: Design loss function L lloss , respectively calculate the loss from the three aspects of bounding box and category information, and update the weights in model training through gradient backpropagation, thereby achieving high detection accuracy and comprehensive detection. The loss function is designed as follows: Among them, L class represents the instance category loss, L bbox Represents the 3D detection box loss of the instance.

Citation Information

Patent Citations

  • Three-dimensional point cloud single target tracking method based on regional self-attention mechanism

    CN115909010A

  • Automatic driving three-dimensional target detection method based on FPN Swin Transformer and Pointnet + +

    CN116403186A

  • Multi-modal 3D target detection algorithm based on pseudo point cloud enhancement and multi-stage feature fusion

    CN118314426A

  • Top-down object detection from lidar point clouds

    US20240410981A1

Cited By

  • Construction site safety detection method, device, equipment and medium

    CN120853023A