Multi-modal target part detection method based on point cloud diversity representation and PointRCNN
By combining point cloud diversity characterization and PointRCNN, the advantages of cylinder representation and graph representation are utilized, and the manifold self-attention mechanism are combined, the problem of poor target detection effect when point cloud data is sparse or large amounts of noise is solved, and more efficient and accurate multimodal target part detection is achieved.
Patent Information
- Application Number
- CN202510106625.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-01-23
AI Technical Summary
The prior art uses poor object detection effect when point cloud data is sparse or there is a lot of noise, and the processing and feature fusion of multimodal data requires a large amount of computing resources.
By combining point cloud diversity characterization and PointRCNN, the advantages of cylinder representation and graph representation are utilized, and the relative position and structural information between points in the point cloud are constructed in combination with the manifold self-attention mechanism to perform multimodal target parts detection.
It improves the accuracy and efficiency of multimodal target parts detection, overcomes the problem of low detection rate of traditional methods under sparse point cloud data, and reduces the consumption of computing resources.
Smart Images

Figure CN120198352A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of mechanical engineering, intelligent manufacturing technology, etc. Specifically, it is a multi-modal target part detection method based on point cloud diversity representation and PointRCNN. Background Art
[0002] Multi-modal target detection integrates data from multiple sensors such as lidar, cameras, and millimeter-wave radars to achieve the detection and recognition of targets in a three-dimensional environment. Compared with single-modal target detection, this method provides higher accuracy and robustness. By combining data from different modalities, it can effectively overcome the limitations of a single modality in specific scenarios, thereby improving the accuracy and stability of object detection in practical applications. Early studies, such as MV3D and AVOD, usually represent objects as multiple two-dimensional images from different perspectives and perform target detection through multi-view fusion techniques. However, due to the differences in angles and resolutions of images from different perspectives, the fusion process is often affected by inconsistent information. In addition, methods based on the fusion of images and LiDAR, such as Transfusion, 3D-CVF, and CLOCs, utilize the semantic information of images and the spatial features of point cloud data to improve the accuracy and robustness of target detection. However, most of these methods rely on simple feature fusion techniques, such as concatenation or weighted summation, and fail to fully explore the correlation and complementarity between different feature types. Therefore, enhancing the fusion process to better utilize the synergistic effect between different modal features has become one of the important research directions. In addition, point cloud multi-modal target detection also faces some challenges. Different data representation methods have their own advantages and disadvantages, but there is relatively little research on combining multiple representation methods; the processing and feature fusion of multi-modal data require a large amount of computing resources; and single multi-modal data fusion methods, such as direct splicing or using attention mechanisms, become the bottleneck of multi-modal target detection, limiting its adaptability in diverse scenarios and application requirements.
[0003] The PointRCNN algorithm can retain more original information without any processing on the point cloud data, thus performing better in terms of object detection accuracy. However, this algorithm has a strong dependence on the quality and density of the input point cloud data. When the point cloud data is sparse, incomplete, or contains a large amount of noise, the detection effect is often unsatisfactory. This problem mainly stems from the fact that PointRCNN directly uses the point cloud data in the feature extraction stage, making it vulnerable to interference from irrelevant points and noise. In addition, the PointNet++ network highly depends on the quality and density of the point cloud data, resulting in difficulty in effectively capturing key features in a sparse environment. The workflow of PointRCNN can be divided into two main stages: bottom-up 3D candidate box generation and normalized 3D detection box refinement. The specific steps are as follows: First, use a point cloud encoding network based on the PointNet++ network to extract the feature vector of each point from the three-dimensional point cloud; then, generate a segmentation mask through a foreground point segmentation network to mark the foreground points in the point cloud, and then generate a small number of high-quality 3D candidate boxes in a bottom-up manner based on these foreground points. Second, after obtaining the 3D candidate boxes, expand each box to encode its context information and retain all the points located within the expanded box. Finally, transform the aggregated points of each candidate box into a canonical coordinate system, learn the local features of the target, and combine the local features with the global semantic features to achieve more accurate bounding box refinement and confidence prediction. Summary of the Invention
[0004] The purpose of the present invention is to provide a multi-modal target part detection method based on point cloud diversity representation and PointRCNN, which combines the cylindrical representation and graph representation of point cloud data, makes full use of their respective advantages to improve the overall feature expression ability of the algorithm, and combines the manifold self-attention mechanism to construct the relative position and structural information between points in the point cloud, ultimately improving the accuracy and efficiency of multi-modal target part detection.
[0005] The present invention is realized through the following technical solutions: A multi-modal target part detection method based on point cloud diversity representation and PointRCNN includes the following steps:
[0006] 1) Through data preprocessing, convert the original point cloud data into the data format required by the 2D backbone network;
[0007] 2) Input the data obtained in step 1) into the point cloud diversity representation module for processing, so as to utilize the advantages of the cylindrical representation and graph representation to improve the accuracy and efficiency of point cloud target part detection and obtain diverse features;
[0008] 3) Input the data obtained in step 1) into the manifold grouping feature sampling module to effectively calculate the mutual relationship between the internal points of the point cloud, and obtain multi-layer image features, enabling the detection method to more accurately capture the context information and global features in the point cloud data. The manifold grouping feature sampling module not only improves the algorithm's understanding ability of complex space structures but also significantly enhances the algorithm's generalization ability;
[0009] 4) After step 3), input the output features (diverse features) of the point cloud diversity representation module and the output features (multi-layer image features) of the manifold grouping feature sampling module into the double-layer feature fusion module for fusion and enhancement processing; that is, fuse the point cloud and image features in the double-layer feature fusion module to obtain a comprehensive feature representation, and on the basis of the preliminary fusion, enhance the point cloud semantic features to obtain a more abundant feature expression, thereby improving the detection efficiency and outputting multi-modal fusion features;
[0010] 5) Use a 3D backbone network to perform target part detection on the multi-modal fusion features output by the double-layer feature fusion module, and obtain the detection results of the target parts through the 3D point cloud detection head in the 3D backbone network.
[0011] To better implement the multi-modal target part detection method based on point cloud diversity representation and PointRCNN of the present invention, the following setting method is particularly adopted: The 2D backbone network in step 1) uses a pre-trained Swin-Transformer model as the detection model (backbone network) of the 2D backbone network for image feature extraction to give full play to the advantages of self-attention; An important difference between RGB images and lidar point cloud data is that the former can provide rich color and texture information, which helps to distinguish different target part categories; In this method, when the 2D backbone network preprocesses the data, the multi-level semantic features in the RGB image are integrated into the point cloud features. Compared with traditional convolutional neural networks, the self-attention mechanism can more effectively capture the long-range dependencies between any two positions in the image, and by exploring the correlation of the overall image, the global information and context of the image are pooled; To further enrich the feature expression, a feature pyramid network structure is introduced in the 2D backbone network, and an additional pooling layer is added on the basis of the original four layers, thereby generating five-layer image features, denoted as F image ={f I 1 ,f I 2 ,f I 3 ,f I 4 ,f I 5}, the fifth - layer features are obtained by performing pooling operations on the top - layer features, which can introduce more features at different scales, thereby further enhancing the detection algorithm's perception ability of target parts; in the feature pyramid stage, the number of input channels for each layer is set to 96, 192, 384, and 768, while the number of output channels is fixed at 96. This configuration allows the algorithm to effectively extract features at different scales, enabling it to capture fine - grained and deep - level information from the input image. The resolution of the image is set to 2048×618 pixels.
[0012] To better implement the multi - modal target part detection method based on point - cloud diversity representation and PointRCNN of the present invention, the following setting method is specifically adopted: The point - cloud diversity representation module in step 2) combines the advantages of voxel representation and graph representation, providing flexible resolution while retaining detailed information, and processes the data output from step 1) using dynamic voxel feature encoding and dynamic graph network encoding; dynamic voxel feature encoding can effectively reduce the data dimension, thus saving computing resources. Different from traditional methods, the dynamic voxel feature encoding does not sample the point cloud into a fixed number of voxels with a fixed volume, but retains the complete mapping relationship between points and voxels: The original point - cloud data is represented as P = {p1, p2, p3,..., p n}, where p i = {(x i , y i , z i )|i = 1, 2, 3,..., n}, (x i , y i , z i ) represents the three - dimensional coordinates of the i - th point, r represents the reflectivity intensity. The point - cloud space is divided into a three - dimensional voxel space, and each voxel has a specified size [Δx, Δy, Δz]. Then a three - dimensional voxel can be represented by the following formula:
[0013]
[0014] where, represents the rounding operation, (i, j, k) represents the center - point coordinates of the three - dimensional voxel. For the center - point of each voxel, the point - cloud data is mapped into the corresponding voxel, and the average value of the point cloud within each voxel is calculated to obtain the corresponding voxel value At the same time, this process generates the voxel features F pillar = {f p1 , f p2 , f p3 …f pn}.
[0015] Dynamic graph network encoding can retain more point cloud information without losing details, and can process point cloud data of any shape, increasing the diversity of shapes. In addition, the dynamic graph network has stronger representation ability and scalability, and can extract higher-level semantic information, thereby improving the accuracy of target part detection. The dynamic graph network encoding in the point cloud diversity representation module is based on the cylinder feature F pillar , and through three feature extraction layers, further extracts and aggregates to obtain the point cloud graph encoding feature F g ′ raph . Multi-layer feature extraction can help the module gradually learn the hierarchical feature representation of point cloud data. Each layer focuses on capturing different levels of abstract features in the point cloud, making the final representation more rich and expressive; as the number of layers increases, the receptive field of each layer also expands accordingly, enabling the module to better understand the global and local point cloud structures during the learning process, and helping to better capture the overall shape and context information of the target part.
[0016] In addition, the multi-layer feature extraction in the dynamic graph network encoding also introduces more non-linear transformations, which helps the module learn the complex relationships in the point cloud data. For the target part detection task, a deeper feature extraction network is often required to better simulate the complexity of the point cloud data. For each cylinder, the graph encoding feature F graph =(f g1 , f g2 , f g3 ,..., f gj ) is extracted using the dynamic graph network, where f gj represents the graph-based feature of the j-th cylinder. Let the parameters of the neural network be Θ, then the feature extraction process can be represented by the following formula:
[0017] f gj =f Θ (F pj , {f pi} i∈N(j) );
[0018] Among them, F pj represents all the point cloud data contained in the j-th cylinder, {f pi} i∈N(j) represents the feature representations of other cylinders adjacent to the j-th cylinder, and N(j) is the neighboring set of the j-th cylinder. The k-nearest neighbor algorithm is used to capture the local relationships and similarities between nodes in the graph structure. The present invention combines the discriminative k-nearest neighbor and fuzzy k-nearest neighbor algorithms at different levels to improve the adaptability of the algorithm to different data distributions and features, and a feature aggregation layer is added in the third layer to aggregate the graph encoding features F graph of each cylinder, so as to obtain the aggregated point cloud graph encoding feature F' graph。
[0019] Preferably, the column encoding input of the point cloud diversity representation module consists of 17,600 points sampled from four dimensions. The range of the point cloud is [0, -40, -3, 70.4, 40, 1]. Dynamic column feature encoding alignment is used for encoding, and the size of the column is set to [0.05, 0.05, 0.1]. This module incorporates a feature fusion layer that utilizes five layers of image feature maps, where the image channels are 256, the point cloud channels are 64, and the output channels are 128.
[0020] Preferably, the graph encoding of the point cloud diversity representation module extracts features from four dimensions of the point cloud. Three graph feature extraction layers are used, each layer containing a grouping operation, with 20 samples in each group. The first layer uses the discriminative k-nearest neighbor algorithm, while the other two layers use the fuzzy k-nearest neighbor algorithm. The hyperparameter of the activation function is set to 0.2.
[0021] To better implement the multi-modal target part detection method based on point cloud diversity characterization and PointRCNN described in the present invention, the following setting method is specifically adopted: The manifold grouping feature sampling module in step 3) can unit-vectorize the input feature vector f of each point in the point cloud to obtain a feature representation mapped into the manifold space, so as to use the self-attention mechanism in the manifold space. The calculation formula is described as follows:
[0022]
[0023] where ∥·∥ represents the Frobenius norm, means dividing the input feature vector f by its Frobenius norm, thereby realizing the unitization of the input feature vector f, so that the length of each vector is 1. Through this process, the input feature vector of each point in the point cloud can be mapped to the unit-length sphere, realizing the mapping of the point cloud in the manifold space;
[0024] In the manifold grouping feature sampling module, when the point cloud is mapped to the manifold space, the shortest path between them can be determined by the included angle represented by the two-point vector. By calculating the included angle values between these unit vectors and using the softmax function for normalization, the manifold self-attention score can finally be obtained. The formed formula is as follows:
[0025]
[0026] In the formula, E(p i ) represents the set of points around point p i , and x j is E(p i) is a point in, n represents the dimension of the attention query vector Q and the attention key vector K, V represents the attention value vector, diag(·) represents extracting the diagonal values of a matrix, ⊙ represents the element-wise product of two matrices, and f i attn represents the input point p i is the output feature obtained after the manifold attention weighted calculation.
[0027] To better implement the multi-modal target part detection method based on point cloud diversity representation and PointRCNN of the present invention, the following setting method is specifically adopted: The manifold grouping feature sampling module in step 3) introduces a grouped self-attention mechanism. By grouping the feature channels and sharing the attention weights within the groups, the number of parameters and the computational complexity of the module are reduced, and the overfitting phenomenon is alleviated. The grouped self-attention mechanism inherits the advantages of both vector attention and multi-head attention at the same time, and can learn deeper feature representations more effectively, thereby enhancing the processing ability and generalization ability of the module for complex data. In grouped self-attention, the channels of the value vector are evenly divided into k groups, and 1 ≤ k ≤ c. Therefore, the weight encoding layer outputs not a complete attention vector of c channels, but a grouped attention vector of k channels, where the channels of each group share a scalar attention weight. The grouped linear transformation function τ is:
[0028]
[0029] where rel is the relationship vector between the attention query vector Q and the attention key vector K, p1, p2, …, p k are the parameters of the grouped linear transformation, where each point p i is a grouped matrix used to independently transform the input relationship vector rel; the zero matrix is used to isolate different groups in the transformation. To enhance the role of position information in the attention mechanism, an additional position encoding multiplier is introduced to improve the influence of position information on weight encoding, so that the module can better learn and understand the relative positions between points in three-dimensional space, thereby enhancing its understanding and processing ability of complex spatial structures.
[0030] The manifold grouping feature sampling module promotes information exchange between groups and enhances the non-linear expression ability of the module by normalizing, activating, and further linearly transforming each group of relationship vectors rel passing through the grouped linear transformation function τ. Finally, the grouped weight encoding function of the relationship vector rel is formed as:
[0031] ω(rel) = Norm(Relu(Linear(ζ mul (p i - p j )⊙τ(rel) + ζbias (p i -p j ));
[0032] Among them, p i and p j are the spatial point coordinates and adjacent point coordinates respectively, ζ mul and ζ bias are position encoding functions implemented by two fully connected layers, which are used to enhance the module's understanding and representation ability of the complex spatial relationships in the point cloud data. Linear(·) is a linear transformation layer implemented by a fully connected layer, which is responsible for further feature transformation between groups; Relu(·) is an activation function, which is used to introduce non-linearity to better learn complex features; Norm(·) is a normalization layer, which is used to accelerate the training speed and improve the stability of the module. Finally, the calculation formula of the manifold grouping self-attention feature vector is as follows:
[0033]
[0034] Among them, l represents the group index, and V j lc / k+m represents the m-th feature in the l-th group of the value vector grouping. By applying the grouped attention mechanism, the weighted and aggregation of the context information of the spatial point coordinates p i are realized. f i attn not only contains the information of the spatial point coordinates p i itself, but also integrates the information of the surrounding neighborhoods, significantly improving the efficiency and expression ability of the module, so as to more effectively extract features in the sparse point cloud.
[0035] Furthermore, to better implement the multi-modal target part detection method based on point cloud diversity representation and PointRCNN described in the present invention, the following setting method is particularly adopted: The double-layer feature fusion module in the step 4) is divided into a point cloud branch and an image branch. Among them, the point cloud branch transforms and aggregates the diverse features in the step 2), and the image branch performs alignment transformation and aggregation on the multi-layer image features in the step 3), and finally fuses the features aggregated by the point cloud branch and the image branch;
[0036] In the point cloud branch, the diverse features obtained through the point cloud diversity representation module are transformed. Using a fully connected layer and normalization processing, unified point cloud features are obtained and represented as Among them represents the 3D feature of the n-th point; Subsequently, the unified point cloud features are grouped and aggregated, and the K-dimensional tree method is used to perform spatial partitioning on the point cloud features F pc based on its local area; For the point cloud feature F pcFor each local region, the feature vectors within that region are averaged to obtain a single feature vector representing the region; each point cloud feature F pc The new feature vector of is determined by the feature vectors of the corresponding local regions to which it belongs, and the calculation process is shown in the following formula:
[0037] F p ′ c = Mean(KD(F pc ));
[0038] where F p ′ c is the point cloud feature after grouped aggregation, and Mean(·) represents the mean aggregation method;
[0039] In the image branch, the multi-layer image features obtained from the manifold grouping feature sampling module are aligned and transformed to obtain an image feature representation with the same dimension as the point cloud feature F pc in the point cloud branch; the multi-layer image features are aligned with the point cloud feature F pc through the alignment sub-module, the extracted multi-layer image features and the point cloud feature F pc are normalized, and the iterative closest point algorithm is used to align the multi-layer image features with the point cloud feature F pc , and their spatial relationship is optimized. The aligned multi-layer image features are obtained through convolution operations and enhancement transformations to obtain the image feature F img ;
[0040] Then, the image feature F img is grouped and aggregated. For each image feature F img , its surrounding image features are found using radius neighborhood search, the found surrounding image features are aggregated, and pooling operations are used to aggregate the features within the neighborhood into a representative feature. The aggregated feature is fused with the initial image feature F img to form a new image feature F i ′ mg ;
[0041] Finally, the grouped and aggregated point cloud feature F p ′ c is fused with the new image feature F i ′ mg to obtain the multi-modal fusion feature F fusion .
[0042] The double-layer feature fusion module uses a sparse 3D-U-shaped network for the preliminarily fused multi-modal fusion feature F fusionPerform finer-grained processing and enhance it using semantic features to obtain better target part detection results. The multi-modal fusion feature F fusion is input into this 3D-U-shaped network to further obtain the spatial feature F spatial and the semantic feature F semantic . These two features are combined to obtain the final fusion feature F final , as shown in the following formula:
[0043] F final = Concat(F spatial , Conv(F semantic ));
[0044] Among them, Concat(·) represents the feature concatenation operation, and Conv(·) is used to expand the features to the same dimension using a sparse 3D-U-shaped network for concatenation.
[0045] To further better implement the multi-modal target part detection method based on point cloud diversity characterization and PointRCNN described in the present invention, the following setting method is particularly adopted: The 3D backbone network in step 5) uses the concept of spatial autocorrelation as a measure of the degree of spatial dispersion of the point cloud. Spatial autocorrelation describes the influence of the values of random variables at different positions in space by the values of random variables at nearby positions, showing a certain spatial trend. Data with high spatial autocorrelation indicates a certain degree of aggregation and trend; on the contrary, data with low spatial autocorrelation is more dispersed and may be interfering noise. The 3D backbone network calculates the spatial correlation between each point and its nearby points, determines whether the point cloud is discrete, and eliminates the point cloud with low spatial autocorrelation to reduce the number of point clouds while retaining key information.
[0046] Each point cloud data is represented in the Cartesian coordinate system as (x, y, z, r). The spatial autocorrelation algorithm converts the point cloud data set from the Cartesian coordinate system to the spherical coordinate system to obtain the distance r and the pitch angle θ, azimuth angle for calculating the weight value and spatial autocorrelation value between points; the specific calculation formula is:
[0047]
[0048] After inputting the original point cloud data, the entire point cloud space will be divided into multiple grids, and the spatial autocorrelation value of each grid will be calculated separately.
[0049] Since different points in space contribute differently to the calculation of the spatial autocorrelation value, the angular information between points can be used as the relative position information of the data points in the three-dimensional space. The 3D backbone network calculates the spatial angular distance w ij, to measure the correlation between two points; the greater the correlation between the two points, the greater the assigned weight, and the spatial angular distance w ij The calculation formula is:
[0050]
[0051] Among them, i represents the processing point, and j represents the surrounding points of the processing point. In the spatial autocorrelation measurement, autocorrelation is used to evaluate the spatial relationship between a point and other points, and the relative position of a point to itself is always zero. Therefore, when i = j, w ij is set to zero, and the calculation formula for the spatial autocorrelation value I in each grid is:
[0052]
[0053] Among them, is the average value of all distances, is the sum of all spatial angular distances w ij After obtaining the spatial autocorrelation value of each grid, according to the set threshold, the grid point cloud with I less than the threshold is removed, so as to obtain the point cloud data after removing noise points and irrelevant points, so that the subsequent 3D point cloud detection head can better focus on the feature detection of the target part.
[0054] Furthermore, to better implement the multi-modal target part detection method based on point cloud diversity representation and PointRCNN of the present invention, the following setting method is particularly adopted: the 3D point cloud detection head in the 3D backbone network in step 5) uses the following bounding box encoding function for the detection object:
[0055]
[0056] Among them, x, y, z are the center coordinates; w, l, h are the width, length, and height respectively; θ is the yaw angle around the z-axis; x gt and x a are the ground truth and the anchor box respectively; is the diagonal of the bottom edge of the anchor box.
[0057] Furthermore, to better implement the multi-modal target part detection method based on point cloud diversity representation and PointRCNN of the present invention, the following setting method is particularly adopted: the loss function of the 3D backbone network in step 5) is set with: foreground and background classification loss function, interval-based localization loss function, and total loss function in the refinement stage;
[0058] In point cloud processing, since the number of foreground points (target object points) and background points (non-target object points) is unbalanced, and the number of background points is much larger than that of foreground points, the foreground and background point classification loss function adopts the focal loss function to reduce the weight of negative samples in training:
[0059] where p represents the foreground prediction probability of each point, and α and γ are hyperparameters of the focal loss function. In this algorithm, these parameters are set to be the same as those in the standard model PointRCNN;
[0060] The interval-based localization loss function is:
[0061] where pos represents the set of positive sample points; N pos represents the number of positive sample points; is the interval-based target bounding box position regression loss, and is the interval-based target bounding box size regression loss, and is the discrete interval predicted by the foreground point in the dimension u ∈ {x, z, θ}, and where u p is the true center coordinate of the object, and u (p) is the true coordinate of the corresponding foreground point, is the search range on the X and Z axes, and η is the unified interval length; is the residual between the discrete interval predicted by the foreground point and the true position, and where C is a constant used for normalization to ensure that the residual calculation is not affected by different interval sizes. Considering that the central position of most objects changes little in the vertical direction (Y-axis), the residual on the Y-axis is calculated by the difference between the true center coordinate y p of the object and the coordinate y (p) of the foreground point: The cross-entropy loss function is used to measure the consistency between the predicted discrete interval and the true interval; is the smooth L1 loss function, which is used to reduce the difference between the predicted residual and the true residual;
[0062] In order to obtain a more accurate target part detection box for the 3D backbone network, the candidate box refinement stage needs to ensure that the network can correctly classify each detected object and accurately predict the position and size of the object's bounding box. The total loss function of the refinement stage is:
[0063]
[0064] where, represents the total number of sample points for a processing batch, represents the total number of positive sample points in a processing batch, prob is the predicted label, and label represents the true label. and correspond to the position regression and size regression of the refined generated candidate boxes respectively. This method allows the 3D backbone network to process continuous attributes and discretize them, which can utilize the simplicity of the classification task and retain the accuracy of the regression task, thus better predicting the detection boxes of target parts.
[0065] To better implement the multi-modal target part detection method based on point cloud diversity representation and PointRCNN described in the present invention, the following setting method is specifically adopted: The positioning loss of the 3D point cloud detection head is defined as:
[0066] L loc = ∑ b∈(x,y,z,w,l,h,θ) SmoothL1(Δb);
[0067] Since the angle positioning loss cannot distinguish flipped bounding boxes, the Softmax classification loss L loc is used for learning the discrete direction, and the object classification loss using focal loss is:
[0068] L cls = -α a (1 - p a ) γ log p a ;
[0069] where p a represents the class probability of the anchor point, α takes the default value of 0.25, γ takes the default value of 3, the AdamW optimizer is used to optimize the loss function, and it decays by 0.1 times every 15 epochs.
[0070] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0071] The present invention combines the cylindrical representation and graph representation of point cloud data, makes full use of their respective advantages to improve the overall feature expression ability of the algorithm, and combines the manifold self-attention mechanism to construct the relative position and structural information between points in the point cloud, overcoming the inherent limitation of traditional point cloud target detection methods relying on single representation features, solving the problem of low target detection rate in the case of sparse point cloud data or the existence of a large number of noise points, and finally improving the accuracy and efficiency of multi-modal target part detection. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 It is a structural diagram of the point cloud diversity representation module.
[0073] Figure 2 It is a flowchart of the manifold grouped self-attention mechanism.
[0074] Figure 3 It is a structural diagram of the manifold grouped feature sampling module.
[0075] Figure 4 It is a flowchart of the point cloud branch and the image branch.
[0076] Figure 5 It is a structural diagram of the double-layer feature fusion module. Specific implementation manners
[0077] The present invention will be further described in detail below in conjunction with embodiments, but the implementation manners of the present invention are not limited thereto.
[0078] To make the purpose, technical solutions and advantages of the implementation manners of the present invention clearer, the technical solutions in the implementation manners of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the implementation manners of the present invention. Obviously, the described implementation manners are part of the implementation manners of the present invention, rather than all of the implementation manners. Based on the implementation manners in the present invention, all other implementation manners obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention. Therefore, the following detailed description of the implementation manners of the present invention provided in the drawings is not intended to limit the scope of the present invention to be protected, but merely represents the selected implementation manners of the present invention. Based on the implementation manners in the present invention, all other implementation manners obtained by those of ordinary skill in the art without making creative efforts belong to the scope of protection of the present invention.
[0079] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of such features. In the description of the present invention, "a plurality" means two or more unless otherwise specifically defined.
[0080] Glossary:
[0081] PointRCNN: It is a point cloud-based object detection algorithm that draws on the idea of the region convolutional neural network (RCNN) in 2D object detection and applies it to point cloud data. The core idea of PointRCNN is to regard point cloud data as a set of points, generate proposal boxes by sampling and aligning the set of points, and finally identify the target objects in the point cloud.
[0082] Swin-Transformer Model: The Swin-Transformer model is a deep learning model designed specifically for computer vision tasks. It gradually reduces the spatial resolution of the feature map through multiple stages while increasing the number of channels, enabling the model to effectively capture features at different scales. It introduces a windowed multi-head self-attention mechanism. The input feature map is divided into several non-overlapping small windows, and self-attention calculations are performed within each window. Using the shifted window strategy, elements in adjacent windows have the opportunity to interact with each other, which helps to construct deeper image feature representations.
[0083] MV3D: MV3D (Multi-View 3D) is a deep learning framework for 3D object detection. It extracts features from lidar point clouds and RGB images respectively. For lidar data, a voxelization method is used to convert the 3D point cloud into a 2D or 3D grid. For images, a convolutional neural network is used to extract visual features. The features from these two sources are fused to obtain a more rich and representative feature description, ensuring the effective combination of features from different modalities. Based on the fused feature map, a series of candidate regions are generated, and a regressor is applied within each proposed region to determine the specific 3D bounding box parameters. At the same time, a classifier is used to judge whether there is actually an object in the region and its category.
[0084] tAVOD: tAVOD (Temporal Awareness Vehicle Object Detection) is a deep learning model for 3D object detection. This model takes into account the information in the time dimension. It compares and correlates the features of the current frame with those of historical frames to construct a feature representation that includes the time dimension. By effectively integrating the information on the time axis, the model can predict the positions and development trends of objects in the next few seconds to a certain extent, enhancing the model's adaptability in a rapidly changing environment.
[0085] Transfusion: Transfusion is a deep learning framework for multi-modal data fusion. By combining data from different sensors (such as lidar, cameras, etc.), it improves the accuracy and robustness of 3D object detection and tracking. It uses a soft association mechanism to replace the hard association mechanism in previous fusion methods, making the model more robust to degraded image quality and sensor misalignment. A detection head based on the Transformer decoder is used to achieve adaptive feature fusion between images and point clouds. The spatial range of cross-attention is restricted around the initial bounding box to enable the model to better access relevant positions. This model provides richer image features for object detection and is more robust to poor image conditions.
[0086] 3D-CVF: 3D-CVF (3D Camera-LiDAR Fusion) is a multi-sensor fusion technology for autonomous driving that combines three-dimensional data from cameras and LiDAR to improve the ability to understand the surrounding environment and the accuracy of object detection. The visual information of the camera and the three-dimensional spatial information of LiDAR complement each other, making the object detection and scene understanding of the model more accurate. In complex environments or poor lighting conditions, a single sensor may be limited. 3D-CVF enhances the robustness and reliability of the model by combining data from different sensors, enabling more precise positioning, environmental perception, and decision-making.
[0087] CLOCs: CLOCs (Camera-LiDAR Object Candidates) combines the color and texture information provided by the camera with the depth and spatial position information provided by LiDAR. It uses the output results of 3D and 2D detectors generated before non-maximum suppression, and also utilizes its geometric and semantic consistency to achieve more accurate detection accuracy. Adopting a late fusion strategy, it integrates the detection results of different modalities at the decision-making level, significantly improving the performance of the model's three-dimensional object detection and enhancing the model's ability to perceive complex environments and recognize objects.
[0088] LiDAR: LiDAR (Light Detection and Ranging) refers to "Light Detection and Ranging", which is a remote sensing technology that determines distance by emitting laser pulses and measuring the time it takes for these pulses to reflect back from the target object, and combines angle information and other auxiliary data to generate detailed three-dimensional spatial information, including features such as terrain and buildings, facilitating subsequent data analysis and processing. LiDAR systems can create high-resolution digital surface models and three-dimensional point cloud maps, and are widely used in fields such as geographic information systems, autonomous driving, and three-dimensional object detection.
[0089] PointNet++ network: The PointNet++ network is a deep neural network specifically designed for processing point cloud data. It extracts features from local regions at different scales to capture geometric information at different levels, dynamically selects neighbor points based on the distance between points to form local regions, thus more flexibly adapting to different geometric structures. To effectively organize point cloud data and construct local regions, PointNet++ uses a sampling and grouping strategy based on query spheres. For each center point, a sphere with a fixed radius is defined, and neighbor points are searched within this range to form local regions, ensuring that the points within the local region have similar spatial positions and significantly enhancing the network's ability to understand point cloud data.
[0090] The present invention is obtained based on the following theoretical basis:
[0091] Compared with other point cloud object detection methods, the PointRCNN algorithm can directly process the original point cloud data, avoiding information loss that may occur during data conversion. This algorithm adopts a bottom-up strategy to generate high-quality 3D candidate boxes, significantly reducing the search space, and fine-tuning the bounding boxes in the second stage, which makes it particularly suitable for dealing with complex three-dimensional environments. However, due to the sparsity of point cloud data and the presence of a large amount of noise, PointRCNN faces challenges in extracting effective point cloud features. This algorithm introduces a spatial self-correlation method, which not only reduces the scale of the input data but also eliminates the interference of noise points and irrelevant points on feature learning during model training, thus helping the model to converge better. Compared with the two-dimensional convolutional module used in the backbone network of PointRCNN, the proposed multi-modal object part detection algorithm based on point cloud diversity representation and PointRCNN can effectively calculate the mutual relationship between points inside the point cloud using the self-attention mechanism. This method can capture the context information and global features in point cloud data more accurately, thereby enhancing the algorithm's ability to understand complex spatial structures and significantly improving the algorithm's generalization ability. Compared with traditional convolutional networks, this self-attention-based strategy shows higher efficiency and accuracy in dealing with the disorder and inhomogeneity of point cloud data, providing a more efficient and accurate detection method for the field of point cloud object part detection.
[0092] In sparse scenarios, PointRCNN fails to extract sufficient effective point cloud features, resulting in insufficient feature information and thus affecting the accuracy of object detection. To improve the detection performance, the proposed multi-modal object part detection method based on point cloud diversity representation and PointRCNN improves the point cloud encoding network of PointRCNN and introduces a spatial self-correlation algorithm to preprocess the point cloud during the training stage. Although the point cloud encoding network of PointRCNN uses PointNet to directly extract features, traditional convolutional neural networks are mainly designed to process regularly arranged image data, while three-dimensional point cloud data is essentially an irregular set of points embedded in a continuous space, which makes convolutional neural networks perform poorly when directly applied to point cloud data processing. Especially when extracting local features in complex three-dimensional spatial structures, due to its fixed-size convolutional kernel and dependence on regular grid structures, convolutional neural networks often cannot fully capture rich spatial relationships. In addition, the receptive field limitation of convolutional neural networks makes them less adaptable to density changes and irregular distributions in point cloud data. To overcome these limitations, this method introduces the self-attention mechanism in Transformer into the point cloud encoding network of PointRCNN. Different from convolutional neural networks, the self-attention mechanism can naturally process irregular and disordered data structures, dynamically adjust the aggregation weights according to data features and relative positions, and thus capture complex spatial relationships more effectively.
[0093] Traditional dot - product self - attention mainly calculates the relationships between features in Euclidean space. However, for point - cloud data with complex structures distributed in non - Euclidean space, traditional dot - product self - attention may not be able to effectively capture its intrinsic geometric structure. Therefore, to better adapt to the spatial characteristics of point - cloud data, this method transforms the dot - product operation in the self - attention mechanism into the calculation of the shortest path in the manifold space, which can more accurately construct the relative positions and structural information between points in the point cloud, thereby helping the model better understand the spatial relationships and geometric properties in the point - cloud data. By introducing the grouped self - attention mechanism, the number of parameters for weight encoding is reduced, effectively preventing overfitting, and thus enhancing the model's ability to capture data features. The manifold - grouped feature sampling module in the multi - modal object part detection method based on point - cloud diversity representation and PointRCNN combines the manifold self - attention mechanism and the grouped self - attention mechanism, improving the point - cloud encoding network of PointRCNN and providing a more powerful 3D spatial feature extraction ability for the PointRCNN algorithm.
[0094] Embodiment 1:
[0095] The present invention designs a multi - modal object part detection method based on point - cloud diversity representation and PointRCNN, which combines the cylindrical representation and the graph representation of point - cloud data, fully utilizes their respective advantages to improve the overall feature expression ability of the algorithm, and combines the manifold self - attention mechanism to construct the relative positions and structural information between points in the point cloud, ultimately improving the accuracy and efficiency of multi - modal object part detection. The method includes the following steps:
[0096] 1) Through data pre - processing, convert the original point - cloud data into the data format required by the 2D backbone network;
[0097] 2) Input the data obtained in step 1) into the point - cloud diversity representation module for processing to utilize the advantages of the cylindrical representation and the graph representation to improve the accuracy and efficiency of point - cloud object part detection and obtain diverse features;
[0098] 3) Input the data obtained in step 1) into the manifold - grouped feature sampling module to effectively calculate the mutual relationships between points inside the point cloud, obtaining multi - layer image features, enabling the detection method to more accurately capture the context information and global features in the point - cloud data. The manifold - grouped feature sampling module not only improves the algorithm's ability to understand complex spatial structures but also significantly enhances the algorithm's generalization ability;
[0099] 4) After step 3), the output features (diverse features) of the point cloud diversity representation module and the output features (multi-layer image features) of the manifold grouping feature sampling module are input into the double-layer feature fusion module for fusion and enhancement processing; that is, the point cloud and image features are fused in the double-layer feature fusion module to obtain a comprehensive feature representation, and on the basis of the preliminary fusion, by enhancing the point cloud semantic features, a richer feature expression is obtained, thereby improving the detection efficiency and outputting multi-modal fusion features;
[0100] 5) Use the 3D backbone network to detect the target parts for the multi-modal fusion features output by the double-layer feature fusion module, and obtain the detection results of the target parts through the 3D point cloud detection head in the 3D backbone network.
[0101] Example 2:
[0102] This embodiment is further optimized on the basis of the above embodiment. The same parts as the foregoing technical solutions will not be elaborated here. To better implement the multi-modal target part detection method based on point cloud diversity characterization and PointRCNN of the present invention, the following setting method is particularly adopted: the 2D backbone network in step 1) uses a pre-trained Swin-Transformer model as the detection model (backbone network) of the 2D backbone network for image feature extraction to give full play to the advantages of self-attention; an important difference between RGB images and lidar point cloud data is that the former can provide rich color and texture information, which helps to distinguish different target part categories; in this method, when the 2D backbone network preprocesses the data, the multi-level semantic features in the RGB image are integrated into the point cloud features. Compared with traditional convolutional neural networks, the self-attention mechanism can more effectively capture the long-range dependencies between any two positions in the image, and by exploring the correlation of the overall image, the global information and context of the image are aggregated; to further enrich the feature expression, this method introduces a feature pyramid network structure in the 2D backbone network, and on the basis of the original four layers, an additional pooling layer is added to generate five layers of image features, denoted as F image ={f I 1 ,f I 2 ,f I 3 ,f I 4 ,f I 5}, the fifth - layer features are obtained by performing pooling operations on the top - layer features, which can introduce more scale features and further improve the perception ability of the detection algorithm for target parts. In the feature pyramid stage, the number of input channels for each layer is set to 96, 192, 384, and 768, while the number of output channels is fixed at 96. This configuration allows the algorithm to effectively extract features at different scales, enabling it to capture fine - grained and deep - level information from the input image. The resolution of the image is set to 2048×618 pixels.
[0103] Embodiment 3:
[0104] This embodiment is further optimized based on any of the above - mentioned embodiments. The same parts as the foregoing technical solutions will not be elaborated here. To better implement the multi - modal target part detection method based on point - cloud diversity representation and PointRCNN of the present invention, the following setting method is particularly adopted: The point - cloud diversity representation module in step 2) combines the advantages of voxel representation and graph representation. As Figure 1 shown, while retaining detailed information, it provides flexible resolution. The output data of step 1) is processed using dynamic voxel feature encoding and dynamic graph network encoding. The dynamic graph network encoding consists of a dynamic graph convolutional network. The graph - encoded features are further extracted through three feature extraction layers. Each feature extraction layer consists of a graph convolutional module. A feature aggregation layer is added in the third feature extraction layer to aggregate the graph - encoded features of each voxel, thereby obtaining the aggregated point - cloud graph - encoded features.
[0105] The dynamic voxel feature encoding can effectively reduce the data dimension, thereby saving computing resources. Different from traditional methods, the dynamic voxel feature encoding does not sample the point cloud into a fixed number of voxels with a fixed volume, but retains the complete mapping relationship between points and voxels: The original point - cloud data is represented as P = {p1, p2, p3,..., p n}, where p i ={(x i , y i , z i )|i = 1, 2, 3,..., n}, (x i , y i , z i ) represents the three - dimensional coordinates of the i - th point, and r represents the reflectivity intensity. The point - cloud space is divided into a three - dimensional voxel space, and each voxel has a specified size [Δx, Δy, Δz]. Then a three - dimensional voxel can be represented by the following formula:
[0106]
[0107] where, Denotes the rounding operation, and (i, j, k) represents the center point coordinates of the three-dimensional cylinder. For the center point of each cylinder, the point cloud data is mapped into the corresponding cylinder, and the average value of the point cloud within each cylinder is calculated to obtain the corresponding cylinder value. Meanwhile, this process generates the cylinder feature F of the point cloud. pillar ={f p1 , f p2 , f p3 … f pn}.
[0108] The dynamic graph network encoding can retain more point cloud information without losing details, and can process point cloud data of any shape, increasing the shape diversity. In addition, the dynamic graph network has stronger representation ability and scalability, and can extract higher-level semantic information, thus improving the accuracy of target part detection. The dynamic graph network encoding in the point cloud diversity representation module, based on the cylinder feature F pillar , further extracts and aggregates through three feature extraction layers to obtain the point cloud graph encoding feature F g ′ raph . Multi-layer feature extraction can help the module gradually learn the hierarchical feature representation of the point cloud data. Each layer focuses on capturing different levels of abstract features in the point cloud, making the final representation more rich and expressive; as the number of layers increases, the receptive field of each layer also expands accordingly, enabling the module to better understand the global and local point cloud structures during the learning process, which helps to better capture the overall shape and context information of the target part.
[0109] In addition, the multi-layer feature extraction in the dynamic graph network encoding also introduces more non-linear transformations, which helps the module learn the complex relationships in the point cloud data. For the target part detection task, a deeper feature extraction network is often required to better simulate the complexity of the point cloud data. For each cylinder, the graph encoding feature F graph =(f g1 , f g2 , f g3 ,..., f gj ) is extracted using the dynamic graph network, where f gj represents the graph-based feature of the j-th cylinder. Let the parameters of the neural network be Θ, then the feature extraction process can be represented by the following formula:
[0110] f gj = f Θ (F pj , {f pi} i∈N(j) );
[0111] where, F pj represents all the point cloud data contained in the j-th cylinder, {fpi} i∈N(j) It represents the feature representation of other cylinders adjacent to the j-th cylinder. N(j) is the neighboring set of the j-th cylinder, and the k-nearest neighbor algorithm is used to capture the local relationships and similarities between nodes in the graph structure. The present invention combines the discriminative k-nearest neighbor and fuzzy k-nearest neighbor algorithms at different levels to improve the adaptability of the algorithm to different data distributions and features, and a feature aggregation layer is added in the third layer to aggregate the graph-encoded features F of each cylinder graph to obtain the aggregated point cloud graph-encoded feature F g ′ raph .
[0112] Preferably, the point cloud diversity representation module introduces a residual structure to reduce network parameters, accelerate network training, and improve the performance and efficiency of the module. The dynamic graph network encoding can, to a certain extent, make up for the limitations of voxelization processing. However, the graph-encoding representation may introduce higher computational complexity and memory requirements. Therefore, this method introduces a residual connection in the point cloud diversity representation module to reduce the number of parameters and memory occupancy of the module, thereby alleviating the problem of gradient disappearance and accelerating the network training speed. Through the residual connection, the cylinder feature F pillar and the point cloud graph-encoded feature F g ′ raph are processed, and finally the diversified feature F output by the point cloud diversity representation module is obtained diversity , as shown in the following formula:
[0113] F diversity = Res(F pillar , F g ′ raph );
[0114] where Res(·) represents the residual connection operation.
[0115] Preferably, the cylinder encoding input of the point cloud diversity representation module consists of 17,600 points sampled in four dimensions. The range of the point cloud is [0, -40, -3, 70.4, 40, 1], and dynamic cylinder feature encoding alignment is used for encoding. The size of the cylinder is set to [0.05, 0.05, 0.1]. This module integrates a feature fusion layer and utilizes five layers of image feature mapping, where the image channel is 256, the point cloud channel is 64, and the output channel is 128.
[0116] Preferably, the graph encoding of the point cloud diversity representation module extracts features from four dimensions of the point cloud, uses three graph feature extraction layers, each layer contains a grouping operation, each group contains 20 samples, the first layer uses the discriminative k-nearest neighbor algorithm, while the other two layers use the fuzzy k-nearest neighbor algorithm, and the hyperparameter of the activation function is set to 0.2.
[0117] Example 4:
[0118] This embodiment is further optimized on the basis of any of the above embodiments. The same parts as the foregoing technical solutions will not be described herein again. To better implement the multi-modal target part detection method based on point cloud diversity representation and PointRCNN of the present invention, the following setting method is particularly adopted: the manifold grouping feature sampling module in step 3) can unit-vectorize the input feature vector f of each point in the point cloud to obtain a feature representation mapped into the manifold space, so as to use the self-attention mechanism in the manifold space. The calculation formula is described as follows:
[0119]
[0120] where, ∥·∥ represents the Frobenius norm, means dividing the input feature vector f by its Frobenius norm, thereby realizing the unitization of the input feature vector f, so that the length of each vector is 1. Through this process, the input feature vector of each point in the point cloud can be mapped to the sphere with unit length, realizing the mapping of the point cloud in the manifold space;
[0121] In the manifold grouping feature sampling module, when the point cloud is mapped to the manifold space, the shortest path between them can be determined by the included angle represented by the two-point vector. By calculating the included angle values between these unit vectors and normalizing them using the softmax function, the manifold self-attention score can be finally obtained. The formed formula is as follows:
[0122]
[0123] In the formula, E(p i ) represents the set of points around point p i , x j is a point in E(p i ), n represents the dimension of the attention query vector Q and the attention key vector K, V represents the attention value vector, diag(·) represents extracting the diagonal value of the matrix, ⊙ represents the element-wise product of two matrices, and f i attn represents the output feature obtained after the manifold attention weighted calculation of the input point p i .
[0124] Example 5:
[0125] This embodiment is a further optimization based on any of the above embodiments. The same parts as the foregoing technical solutions will not be elaborated here. To better implement the multi-modal target part detection method based on point cloud diversity representation and PointRCNN of the present invention, the following setting method is specifically adopted: The manifold grouping feature sampling module in step 3) introduces a grouped self-attention mechanism. By grouping the feature channels and sharing the attention weights within the groups, the number of parameters and computational complexity of the module are reduced, and the overfitting phenomenon is alleviated. The grouped self-attention mechanism simultaneously inherits the advantages of vector attention and multi-head attention, and can more effectively learn deep feature representations, thereby enhancing the module's processing ability and generalization ability for complex data. In grouped self-attention, the channels of the value vector are evenly divided into k groups, and 1 ≤ k ≤ c. Therefore, the weight encoding layer outputs not a complete attention vector of c channels, but a grouped attention vector of k channels, where the channels of each group share a scalar attention weight. The grouped linear transformation function τ is:
[0126]
[0127] where rel is the relationship vector between the attention query vector Q and the attention key vector K, p1, p2,..., p k are the parameters of the grouped linear transformation, where each point p i is a grouped matrix for independently transforming the input relationship vector rel; the zero matrix is used to isolate different groups in the transformation. To enhance the role of position information in the attention mechanism, an additional position encoding multiplier is introduced to improve the influence of position information on weight encoding, enabling the module to better learn and understand the relative positions between points in three-dimensional space, thereby enhancing its understanding and processing ability for complex spatial structures.
[0128] The manifold grouping feature sampling module promotes information exchange between groups and enhances the module's non-linear expression ability by normalizing, activating, and further linearly transforming each group of relationship vectors rel passing through the grouped linear transformation function τ. Finally, the grouped weight encoding function of the relationship vector rel is formed as:
[0129] ω(rel) = Norm(Relu(Linear(ζ mul (p i -p j )⊙τ(rel)+ζ bias (p i -p j ));
[0130] where p i and p jare the spatial point coordinates and the coordinates of adjacent points, ζ mul and ζ bias are position encoding functions implemented by two fully connected layers, which are used to enhance the module's understanding and representation ability of the complex spatial relationships in point cloud data. Linear(·) is a linear transformation layer implemented by a fully connected layer, which is responsible for further feature transformation between groups; Relu(·) is an activation function used to introduce non-linearity to better learn complex features; Norm(·) is a normalization layer used to accelerate the training speed and improve the stability of the module. Finally, the calculation formula of the feature vector based on manifold grouped self-attention is as follows:
[0131]
[0132] where l represents the group index, and V j lc / k+m represents the m-th feature in the l-th group of the value vector grouping. By applying the grouped attention mechanism, the weighted aggregation of the context information of the spatial point coordinates p i is realized. f i attn not only contains the information of the spatial point coordinates p i itself, but also integrates the information of the surrounding neighborhoods, significantly improving the efficiency and expression ability of the module, and thus more effectively extracting features in sparse point clouds.
[0133] The process of the manifold grouped self-attention mechanism in the manifold grouped feature sampling module is as Figure 2 shown. f i and f j represent the spatial point coordinate features and the adjacent point coordinate features respectively, which are converted into query vector Q, key vector K and value vector V through convolution operations. p i and p j are the spatial point coordinates and the adjacent point coordinates respectively. Δp represents the spatial coordinate difference between two points. The included angle values between these unit vectors are calculated and normalized using the softmax function. Finally, the manifold self-attention score can be obtained; the grouped operation sub-module is used to perform grouped self-attention operations on the value vector V to make it correspond one-to-one with the similarly grouped relationship vectors, and finally the feature vector based on manifold grouped self-attention is calculated.
[0134] The overall structure of the manifold grouped feature sampling module is as Figure 3As shown, this module receives the original point cloud data, embeds the original features into a higher-dimensional feature space through the patch embedding method, fuses the scattered information in the original point cloud data, and generates high-level and discriminative features for more effective learning by subsequent modules. The high-dimensional features after patch embedding capture the relationships between points in the point cloud through the multi-layer manifold grouping self-attention sub-module, encode the spatial structure features for each point, and extract the global features through the max pooling layer, thereby obtaining the overall context information of the point cloud data. The global features will be combined with the local features and encoded again through the manifold grouping self-attention sub-module to obtain the comprehensive features of each point. This combination method not only preserves the inherent attributes of each point but also fuses the spatial information and context information learned from the data, significantly enhancing the algorithm's feature expression ability for point cloud data.
[0135] The manifold grouping feature sampling module uses the manifold grouping self-attention mechanism to map the features to the manifold space, obtain the spatial correlation of the point cloud, thereby improving the method's learning ability for point cloud features. At the same time, the grouping self-attention mechanism reduces the number of parameters for weight encoding, effectively preventing overfitting and enhancing the generalization ability of the algorithm. The point cloud encoding network in the manifold grouping feature sampling module is mainly composed of multiple downsampling layers and upsampling layers. Each downsampling layer consists of a point cloud sampling layer and a manifold grouping self-attention sub-module. In the point cloud sampling layer, the farthest point sampling algorithm is used to obtain the key points in the point set, and then local neighborhoods are formed based on a fixed radius, and other points close to the sampling points are grouped together. The upsampling layer performs precise origin mapping through 3D linear interpolation to retain the detailed features of the point cloud data to the greatest extent. To effectively maintain this detailed information, skip connections are used to cascade the output of the upsampling layer with the corresponding downsampling layer. This structural design enables the manifold grouping feature sampling module to effectively combine low-level features and high-level features, thereby enhancing the module's learning and processing ability for detailed information.
[0136] Example 6:
[0137] This embodiment is further optimized based on any of the above embodiments. The same parts as the foregoing technical solutions will not be described herein again. To better implement the multi-modal target part detection method based on point cloud diversity representation and PointRCNN of the present invention, the following setting method is particularly adopted: The double-layer feature fusion module in step 4) is divided into a point cloud branch and an image branch, as Figure 4 shown, where the point cloud branch transforms and aggregates the diverse features in step 2), the image branch performs alignment transformation and aggregation on the multi-layer image features in step 3), and finally fuses the features aggregated by the point cloud branch and the image branch.
[0138] In the point cloud branch, a transformation operation is performed on the diverse features obtained through the point cloud diversity representation module. Using a fully connected layer and normalization processing, unified point cloud features are obtained and represented as where represents the 3D feature of the nth point; subsequently, the unified point cloud features are grouped and aggregated. Using the KD-tree method, the point cloud features F pc are spatially partitioned based on their local regions; for each local region of the point cloud feature F pc , the feature vectors within the region are averaged to obtain a single feature vector representing the region; the new feature vector of each point cloud feature F pc is determined by the feature vectors of the corresponding local regions to which it belongs. The calculation process is shown in the following formula:
[0139] F p ′ c = Mean(KD(F pc ));
[0140] where, F p ′ c is the point cloud feature after grouped aggregation, and Mean(·) represents the mean aggregation method;
[0141] In the image branch, the multi-layer image features obtained from the manifold grouping feature sampling module are aligned and transformed to obtain an image feature representation with the same dimension as the point cloud feature F pc in the point cloud branch; the multi-layer image features are aligned with the point cloud feature F pc through the alignment sub-module. The extracted multi-layer image features and the point cloud feature F pc are normalized, and the iterative closest point algorithm is used to align the multi-layer image features with the point cloud feature F pc , and their spatial relationship is optimized. The aligned multi-layer image features are obtained as the image feature F img through convolution operations and enhancement transformations;
[0142] Then, the grouped aggregation is performed on the image feature F img . For each image feature F img , its surrounding image features are found using radius neighborhood search, and the found surrounding image features are aggregated. The pooling operation is used to aggregate the features within the neighborhood into a representative feature, and the aggregated feature is fused with the initial image feature F img to form a new image feature F i ′ mg ;
[0143] Finally, the grouped and aggregated point cloud feature F p′ c With the new image feature F i ′ mg Fusion is performed to obtain the multi-modal fusion feature F fusion .
[0144] The double-layer feature fusion module is as follows Figure 5 As shown, the features aggregated by the point cloud branch and the image branch are fused. A sparse 3D-U-shaped network is used to perform a finer-grained processing on the preliminarily fused multi-modal fusion feature F fusion . The low-level features are directly passed to the high-level decoder part by using skip connection features to accelerate the training process. At the same time, the previous layer features are directly passed to retain the rich local information and spatial structure in the features, and the semantic features and spatial features are used to enhance them to obtain better target part detection results. The multi-modal fusion feature F fusion is input into this 3D-U-shaped network to further obtain the spatial feature F spatial and the semantic feature F semantic . These two features are combined to obtain the final fusion feature F final , as shown in the following formula:
[0145] F final = Concat(F spatial , Conv(F semantic ));
[0146] Among them, Concat(·) represents the feature concatenation operation, and Conv(·) is used to expand the features to the same dimension using a sparse 3D-U-shaped network for concatenation.
[0147] Example 7:
[0148] This embodiment is further optimized on the basis of any of the above embodiments. The same parts as the foregoing technical solutions will not be described herein again. To better implement the multi-modal target part detection method based on point cloud diversity representation and PointRCNN of the present invention, the following setting method is particularly adopted: The 3D backbone network in step 5) uses the concept of spatial autocorrelation as a measure of the degree of spatial dispersion of the point cloud. Spatial autocorrelation describes that the values of random variables at different positions in space are affected by the values of random variables at nearby positions, showing a certain spatial trend. Data with high spatial autocorrelation indicates a certain degree of aggregation and trend; on the contrary, data with low spatial autocorrelation is more dispersed and may be interfering noise. The 3D backbone network calculates the spatial correlation between each point and its nearby points, determines whether the point cloud is discrete, and eliminates the point cloud with low spatial autocorrelation to reduce the number of point clouds while retaining key information.
[0149] Each point cloud data is represented as (x, y, z, r) using the Cartesian coordinate system. The spatial autocorrelation algorithm converts the point cloud data set from the Cartesian coordinate system to the spherical coordinate system, obtaining the distance r, the pitch angle θ, and the azimuth angle , which are used to calculate the weight values and spatial autocorrelation values between points; the specific calculation formula is:
[0150]
[0151] After inputting the original point cloud data, the entire point cloud space will be divided into multiple grids, and the spatial autocorrelation values of each grid will be calculated separately.
[0152] Since different points in space contribute differently to the calculation of the spatial autocorrelation value, the angular information between points can be used as the relative position information of the data points in three-dimensional space. The 3D backbone network calculates the spatial angular distance w ij between two points to measure the correlation between the two points; the greater the correlation between the two points, the greater the weight assigned, and the spatial angular distance w ij The calculation formula is:
[0153]
[0154] where i represents the processing point and j represents the surrounding points of the processing point. In the spatial autocorrelation metric, autocorrelation is used to evaluate the spatial relationship between a point and other points, and the relative position of a point to itself is always zero. Therefore, when i = j, w ij is set to zero. The calculation formula for the spatial autocorrelation value I in each grid is:
[0155]
[0156] where is the average value of all distances, is the sum of all spatial angular distances w ij . N represents the number of point clouds in a grid. After obtaining the spatial autocorrelation values of each grid, the grid point clouds with I less than the threshold are removed according to the set threshold, so as to obtain the point cloud data after removing noise points and irrelevant points, enabling the subsequent 3D point cloud detection head to better focus on the feature detection of the target part.
[0157] Example 8:
[0158] This embodiment is further optimized on the basis of any of the above embodiments. The same parts as the foregoing technical solutions will not be described herein again. To better implement the multi-modal target part detection method based on point cloud diversity representation and PointRCNN of the present invention, the following setting method is specifically adopted: The 3D point cloud detection head in the 3D backbone network in step 5) uses the following bounding box encoding function for the detection object:
[0159]
[0160] where x, y, and z are the center coordinates; w, l, and h are the width, length, and height respectively; θ is the yaw angle around the z-axis; x gt and x a are the ground truth and the anchor box respectively; is the diagonal of the bottom side of the anchor box.
[0161] The loss function of the 3D backbone network in step 5) is set as: foreground and background classification loss function, interval-based localization loss function, and total loss function in the refinement stage;
[0162] Since in point cloud processing, the number of foreground points (target object points) and background points (non-target object points) is unbalanced, and the number of background points is much larger than that of foreground points, the foreground and background classification loss function uses the focal loss function to reduce the weight of negative samples in training:
[0163] where p represents the foreground prediction probability of each point, and α and γ are hyperparameters of the focal loss function. In this algorithm, these parameters are set to be the same as those of the standard model PointRCNN;
[0164] In order for the 3D backbone network to more accurately capture the position of the object center from the discrete point cloud, each point in the point cloud is divided into a series of discrete intervals along the X and Z axes. This method can more accurately locate the center of the object. The search ranges on the X and Z axes are evenly divided into intervals with a unified length of η, which are used to represent the positions of different object centers on the X-Z plane. Calculate the discrete intervals of the foreground points in the dimension u ∈ {x, z, θ} where u p is the true center coordinate of the object, and u (p) is the true coordinate of the corresponding foreground point; Through the above process, the continuous position is converted into discrete category information, and it is determined that the foreground points belong to a certain interval on the X and Z axes respectively.
[0165] To refine the exact position of the foreground points in the interval, and calculate the residual through the following formula:
[0166]
[0167] Among them, C is a constant used for normalization to ensure that the residual calculation is not affected by different interval sizes. Considering that the central positions of most objects change little in the vertical direction (Y-axis), the residual on the Y-axis is calculated by the difference between the true center coordinate y of the object p and the foreground point coordinate y (p) :
[0168]
[0169] Based on the above formula, the regression loss of the target bounding box position and the regression loss of the size based on the interval can be obtained:
[0170]
[0171] Among them, is the discrete interval predicted by the foreground point in the dimension u ∈ {x, z, θ}, and the cross-entropy loss function is used to measure the consistency between the predicted interval and the true interval, and the smooth L1 loss function is used to reduce the difference between the predicted residual and the true residual. Combining and , the final localization loss function based on the interval is obtained as:
[0172]
[0173] Among them, pos represents the set of positive sample points, and N pos represents the number of positive sample points.
[0174] For the 3D backbone network to obtain a more accurate target part detection box, the candidate box refinement stage needs to ensure that the network can correctly classify each detected object and accurately predict the position and size of the object's bounding box. The total loss function of the refinement stage is:
[0175]
[0176] Among them, represents the total number of sample points in a processing batch, represents the total number of positive sample points in a processing batch, prob is the predicted label, and label represents the true label. and correspond to the position regression and size regression after refinement of the generated candidate box respectively. This method allows the 3D backbone network to process continuous attributes and discretize them, which can not only utilize the simplicity of the classification task but also retain the accuracy of the regression task, so as to better predict the detection box of the target part.
[0177] Example 9:
[0178] This embodiment is further optimized on the basis of any of the above embodiments. The same parts as the foregoing technical solutions will not be described herein again. To better implement the multi-modal target part detection method based on point cloud diversity representation and PointRCNN of the present invention, the following setting method is specifically adopted: The positioning loss of the 3D point cloud detection head is defined as:
[0179] L loc =∑ b∈(x,y,z,w,l,h,θ) SmoothL1(Δb);
[0180] Since the angular positioning loss cannot distinguish flipped bounding boxes, the Softmax classification loss L loc is used to learn discrete directions, and the object classification loss using focal loss is:
[0181] L cls =-α a (1 - p a ) γ log p a ;
[0182] where p a represents the class probability of the anchor point. α takes the default value of 0.25, γ takes the default value of 3, and the AdamW optimizer is used to optimize the loss function, which is decayed by 0.1 times every 15 epochs.
[0183] The above is only a preferred embodiment of the present invention, and does not impose any form of limitation on the present invention. Any simple modification or equivalent change made to the above embodiments based on the technical essence of the present invention falls within the protection scope of the present invention.
Claims
1. A multimodal target part detection method based on point cloud diversity representation and PointRCNN, characterized by: The steps include: 1) Through data preprocessing, the original point cloud data is converted into the data format required by the 2D backbone network; 2) Inputting the data obtained in step 1) into the point cloud diversity representation module for processing to obtain diversified features; 3) Inputting the data obtained in step 1) into the manifold grouping feature sampling module to calculate the relationship between the internal points of the point cloud to obtain multi-layer image features; 4) After step 3), the output features of the point cloud diversity representation module and the output features of the manifold grouping feature sampling module are input into the double-layer feature fusion module for fusion and enhancement processing, and the multimodal fusion features are output; 5) The 3D backbone network is used to detect target parts using the multimodal fusion features output by the double-layer feature fusion module, and the detection results of the target parts are obtained through the 3D point cloud detection head in the 3D backbone network.
2. The multimodal target part detection method based on point cloud diversity representation and PointRCNN according to claim 1, characterized in that: The 2D backbone network in step 1) uses the pre-trained Swin-Transformer model as the detection model of the 2D backbone network to extract image features; the 2D backbone network integrates the multi-level semantic features in the RGB image into the point cloud features when preprocessing the data; a feature pyramid network structure is introduced into the 2D backbone network, and an additional pooling layer is added on the basis of the original four layers, thereby generating five layers of image features, represented as F image ={f I 1 ,f I 2 ,f I 3 ,f I 4 ,f I 5 }, the fifth layer features are obtained by performing a pooling operation on the top layer features.
3. The multimodal target part detection method based on point cloud diversity representation and PointRCNN according to claim 1, characterized in that: The point cloud diversity representation module in step 2) combines cylinder representation and graph representation, and uses dynamic cylinder feature coding and dynamic graph network coding to process the data output from step 1); the dynamic cylinder feature coding represents the original point cloud data as P = {p1, p2, p3, ..., p n }, where p i ={(x i ,y i ,z i )|i=1,2,3,...,n},(x i ,y i ,z i ) represents the three-dimensional coordinates of the i-th point, r represents the reflectivity intensity, and the point cloud space is divided into a three-dimensional cylinder space. Each cylinder has a specified size [Δx, Δy, Δz]. Then a three-dimensional cylinder is represented by the following formula: in, Indicates the rounding operation, (i, j, k) represents the coordinates of the center point of the three-dimensional cylinder. For the center point of each cylinder, the point cloud data is mapped to the corresponding cylinder, and the average value of the point cloud in each cylinder is calculated to obtain the corresponding cylinder value. At the same time, this process generates the cylindrical feature F of the point cloud. pillar ={f p1 ,f p2 ,f p3 …f pn }; Dynamic graph network coding on column feature F pillar Based on this, the point cloud coding feature F′ is further extracted and aggregated through three feature extraction layers. graph .
4. The multimodal target part detection method based on point cloud diversity representation and PointRCNN according to claim 3 is characterized in that: The manifold grouping feature sampling module in step 3) can unit-vectorize the input feature vector f of each point in the point cloud to obtain a feature representation mapped to the manifold space, so as to use the self-attention mechanism in the manifold space. The calculation formula is described as follows: Among them, ∥·∥ represents the Frobenius norm, It means dividing the input feature vector f by its Frobenius norm, so as to normalize the input feature vector f so that the length of each vector is 1; In the manifold grouping feature sampling module, when the point cloud is mapped to the manifold space, the shortest path between the two point vectors is determined by the angle between them. The angle between these unit vectors is calculated and normalized using the softmax function to finally obtain the manifold self-attention score. The formula is as follows: Where E(p i ) represents point p i The surrounding point set, x j is E(p i ), n represents the dimension of the attention query vector Q and the attention key vector K, V represents the attention value vector, diag(·) represents the diagonal value of the extraction matrix, ⊙ represents the element-by-element product of the two matrices, Represents the input point p i The output features obtained after manifold attention weighted calculation.
5. The multimodal target part detection method based on point cloud diversity representation and PointRCNN according to claim 4 is characterized in that: The manifold grouping feature sampling module in step 3) introduces a grouping self-attention mechanism. In the grouping self-attention, the channels of the value vector are evenly divided into k groups, and 1≤k≤c, where the channels of each group share a scalar attention weight, and the grouping linear transformation function τ is: Where rel is the relationship vector between the attention query vector Q and the attention key vector K; p1,p2,…,p k are the parameters of the grouped linear transformation, where each point p i are all grouping matrices used to independently transform the input relation vector rel; the zero matrix Used to isolate different groups in a transformation.
6. The multimodal target part detection method based on point cloud diversity representation and PointRCNN according to claim 1, characterized in that: The double-layer feature fusion module in step 4) is divided into a point cloud branch and an image branch, wherein the point cloud branch transforms and aggregates the diversified features in step 2), and the image branch aligns, transforms and aggregates the multi-layer image features in step 3), and finally fuses the aggregated features of the point cloud branch and the image branch; In the point cloud branch, the diverse features obtained by the point cloud diversity representation module are transformed, and a fully connected layer and normalization process are used to obtain a unified point cloud feature and represented as in Represents the 3D features of the nth point; then the unified point cloud features are grouped and aggregated, and the point cloud features F are grouped based on their local areas using the K-dimensional tree method. pc Perform spatial division; for point cloud feature F pc For each local area, the feature vectors in the area are averaged to obtain a single feature vector representing the area; each point cloud feature F pc The new feature vector of is determined by the feature vector of the corresponding local area to which it belongs. The calculation process is shown in the following formula: F′ pc =Mean(KD(F pc )); Among them, F′ pc is the point cloud feature after grouping and aggregation, Mean(·) represents the mean aggregation method; In the image branch, the multi-layer image features obtained from the manifold grouping feature sampling module are aligned and transformed to obtain the point cloud features F in the point cloud branch. pc Image feature representation of the same dimension; multi-layer image features and point cloud features F are aligned through the alignment submodule pc Alignment, extracting multi-layer image features and point cloud features F pc Normalize and use the nearest point iteration algorithm to align the multi-layer image features with the point cloud features F pc , and optimize their spatial relationship, and obtain the image feature F through convolution operation and enhancement transformation of the aligned multi-layer image features. img ; Then, the image feature F img Perform group aggregation, for each image feature F img , use the radius neighborhood search to find the surrounding image features, aggregate the surrounding image features found, use the pooling operation to aggregate the features in the neighborhood into a representative feature, and compare the aggregated features with the initial image features F img Fusion is performed to form a new image feature F′ img ; Finally, the grouped and aggregated point cloud features F′ pc With the new image feature F′ img Fusion, get the multimodal fusion feature F fusion .
7. The multimodal target part detection method based on point cloud diversity representation and PointRCNN according to claim 1, characterized in that: The 3D backbone network in step 5) uses the concept of spatial autocorrelation as a measure of the degree of spatial dispersion of the point cloud. The 3D backbone network determines whether the point cloud is discrete by calculating the spatial correlation between each point and its nearby points, and removes the point cloud with low spatial autocorrelation. Each point cloud data is represented by a Cartesian coordinate system as (x, y, z, r). The spatial autocorrelation algorithm converts the point cloud data set from the Cartesian coordinate system to the spherical coordinate system to obtain the distance r, the pitch angle θ, and the azimuth angle Used to calculate the weight value and spatial autocorrelation value between points; the specific calculation formula is:
8. The multimodal target part detection method based on point cloud diversity representation and PointRCNN according to claim 1, characterized in that: The loss function of the 3D backbone network in step 5) includes: a front and back scenic spot classification loss function, an interval-based positioning loss function, and a total loss function in the refinement stage; The foreground and background scene classification loss function uses a focal loss function to reduce the weight of negative samples in training: Among them, p represents the foreground prediction probability of each point, α and γ are the hyperparameters of the focal loss function; The interval-based positioning loss function is: Among them, pos represents the set of positive sample points; N pos Represents the number of positive sample points; is the interval-based target bounding box position regression loss, and is the interval-based target bounding box size regression loss, and is the discrete interval predicted by the foreground point in dimension u∈{x,z,θ}, and Among them, u p is the real center coordinate of the object, u (p) are the real coordinates of the corresponding foreground point, is the search range on the X and Z axes, and η is the uniform interval length; is the residual between the predicted discrete interval of the foreground point and the true position, and where C is a constant used for normalization and is the real center coordinate y of the object p and the foreground point coordinate y (p) The residual on the Y axis is calculated by the difference between: Cross Entropy Loss Function Used to measure the consistency between the predicted discrete interval and the true interval; It is a smooth L1 loss function used to reduce the difference between the predicted residual and the true residual; The total loss function of the refinement stage is: in, Represents the total number of sample points in a processing batch, Represents the total number of positive sample points in a processing batch, prob is the predicted label, label represents the true label, and They correspond to the position regression and size regression after the generated candidate box is refined.
9. The multimodal target part detection method based on point cloud diversity representation and PointRCNN according to claim 1, characterized in that: The positioning loss of the 3D point cloud detection head is defined as: L loc =∑ b∈(x,y,z,w,l,h,θ) SmoothL1(Δb); And use Softmax classification loss L loc For learning in discrete directions, use the object classification loss with focal loss: THE cls =-α a (1-p a ) γ log a ; Among them, p a Represents the class probability of the anchor point, a takes the default value of 0.25, γ takes the default value of 3, and the loss function is optimized using the AdamW optimizer, and decays by 0.1 times every 15 cycles.
Citation Information
Patent Citations
Multi-modal medical image rapid detection method
CN115619768A
Three-dimensional target detection method based on multi-modal fusion and deep attention mechanism
CN116612468A
Three-dimensional target detection method based on multi-modal fusion and deformable attention
CN117975436A
Method of extracting three-dimensional point group feature based on multi-modal attention drive
JP2023133087A
Three-dimensional object detection method based on multi-modal fusion and deep attention mechanism
WO2024217115A1
Cited By
Workpiece parallelism detection method and device, medium and equipment
CN120388017A
Part manufacturing-oriented lattice structure mechanical property prediction method
CN120597564A
Unmanned monitoring ship based on sonar vision meteorological water quality fusion and monitoring method thereof
CN121613461A
Unmanned monitoring ship based on fusion of sonar vision and meteorological water quality and monitoring method thereof
CN121613461B