A three-dimensional point cloud object classification method and device based on multi-view projection, a terminal device, and a storage medium

By employing multi-view projection and feature fusion methods, the problem of low accuracy in 3D point cloud classification of substations was solved, achieving high-precision classification of individual substation equipment.

CN119360092BActive Publication Date: 2025-10-24GUANGDONG POWER GRID CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411380510.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-30
Publication Date
2025-10-24
Estimated Expiration
2044-09-30

AI Technical Summary

Technical Problem

In existing technologies, the classification accuracy of 3D point clouds in substations is low, making it difficult to meet the needs of automatic equipment identification and classification in complex scenarios.

Method used

A 3D point cloud object classification method based on multi-view projection is adopted. By acquiring 3D point cloud data of substations, individual object extraction and multi-view projection are performed to obtain multi-channel 2D projection data. Feature fusion is performed using channel attention and spatial attention formulas to extract multi-view global feature representations and finally determine the category of individual equipment.

Benefits of technology

It improves the accuracy of 3D point cloud classification in substations, enabling better differentiation of individual equipment, enriching the features of point cloud data, and enhancing classification accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119360092B_ABST
    Figure CN119360092B_ABST
Patent Text Reader

Abstract

The application discloses a three-dimensional point cloud object classification method and device based on multi-view projection, a terminal equipment and a storage medium. Three-dimensional point cloud data of a substation to be measured is acquired. Single-body extraction is performed on the three-dimensional point cloud data to obtain single-body point cloud data of a plurality of single-body devices. Multi-view projection is performed on each single-body point cloud data to obtain multi-channel two-dimensional projection data corresponding to each single-body point cloud data. Feature fusion is performed on each multi-channel two-dimensional projection data to obtain multi-view global feature expression corresponding to each single-body point cloud data. The category corresponding to each single-body point cloud data is determined based on each multi-view global feature expression. The application fuses the features captured by multi-view projection, and then enriches the features of single-body point cloud data, so that the substation single-body can be better classified, and the three-dimensional point cloud classification precision of the substation is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of three-dimensional point cloud, and particularly to a three-dimensional point cloud object classification method and device based on multi-view projection, a terminal device and a storage medium. BACKGROUND

[0002] With the rapid development of laser scanning technology, SLAM mapping and photogrammetry technology, three-dimensional point cloud data has been widely applied in production scenarios such as autonomous driving, digital twinning and slope monitoring. Since the environmental perception range of point cloud collection devices is limited, and the quality of point clouds collected by different devices is limited by hardware and physical factors, in order to obtain large-scale fused point clouds that meet practical requirements, it is usually necessary to register and fuse three-dimensional point clouds of multiple stations to obtain large-scene point clouds with geometric consistency. However, the existing large-scene point cloud semantic segmentation model usually has fewer identifiable categories and limited segmentation accuracy, which cannot meet the actual needs of automatic identification and classification of complex three-dimensional point clouds in substations.

[0003] Therefore, there is an urgent need for a three-dimensional point cloud object classification strategy to solve the problem of low classification accuracy of three-dimensional point clouds in substations. SUMMARY

[0004] The embodiments of the present application provide a three-dimensional point cloud object classification method and device based on multi-view projection, a terminal device and a storage medium to solve the problem of low classification accuracy of three-dimensional point clouds in substations.

[0005] To solve the above problems, an embodiment of the present application provides a three-dimensional point cloud object classification method based on multi-view projection, comprising:

[0006] Obtaining three-dimensional point cloud data of a to-be-tested substation; wherein the to-be-tested substation comprises a plurality of single devices;

[0007] Single device extraction is performed on the three-dimensional point cloud data to obtain single point cloud data of the plurality of single devices;

[0008] Multi-view projection is performed on each single point cloud data to obtain multi-channel two-dimensional projection data corresponding to each single point cloud data;

[0009] Feature fusion is performed on each multi-channel two-dimensional projection data to obtain multi-view global feature expression corresponding to each single point cloud data;

[0010] Based on each multi-view global feature expression, the category corresponding to each single point cloud data is determined, and then the category of each single device in the to-be-tested substation is obtained.

[0011] As an improvement to the above solution, performing multi-view projection on each of the individual point cloud data to obtain multi-channel two-dimensional projection data corresponding to each of the individual point cloud data includes:

[0012] Normalizing each of the single point cloud data to obtain normalized single point cloud data corresponding to each of the single point cloud data;

[0013] Determine the segmentation angle based on the preset projection viewpoint;

[0014] In each of the normalized monomer point cloud data, based on the segmentation angle and the rotation matrix, each of the normalized monomer point cloud data is rotated to obtain a number of rotated monomer point cloud data, and each of the rotated monomer point cloud data is projected and spliced ​​to obtain multi-channel two-dimensional projection data corresponding to each of the monomer point cloud data; wherein the number of the rotated monomer point cloud data corresponds to the number of segmentations.

[0015] As an improvement to the above solution, the determining of the segmentation angle based on a preset projection viewpoint includes:

[0016] Selecting a quarter sphere of normalized single point cloud data, performing isotropic segmentation on the quarter sphere to obtain a trajectory of a projection viewpoint of the normalized single point cloud data;

[0017] Determine a number of cut angles on each track.

[0018] As an improvement to the above solution, the normalized single point cloud data is rotated based on the segmentation angle and the rotation matrix to obtain a plurality of rotated single point cloud data, including:

[0019] The segmentation angle and the rotation matrix are input into a preset rotation formula to obtain the rotated single point cloud data corresponding to different segmentation angles; wherein the rotation formula includes:

[0020]

[0021] Where, represents the normalized point cloud, Represents the point cloud after rotation, is the rotation θ around the x-axis m The rotation matrix of the angle, R n,κ is the rotation θ around the z axis n Angle rotation matrix, dθ is the segmentation angle, n splitNum is the number of divided tracks, d is the pixel size of the two-dimensional projection image, and m and n are rotation parameters.

[0022] As an improvement of the above scheme, the projection and splicing of each rotating monomer point cloud data to obtain the multi-channel two-dimensional projection data corresponding to each monomer point cloud data comprises:

[0023] Each rotating monomer point cloud data is input into a two-dimensional statistical frequency projection formula to obtain a two-dimensional statistical frequency projection image corresponding to each rotating monomer point cloud data; each rotating monomer point cloud data is input into a two-dimensional visual depth projection formula to obtain a two-dimensional visual depth projection image corresponding to each rotating monomer point cloud data; wherein the two-dimensional statistical frequency projection formula is specifically:

[0024]

[0025] In the formula, The number of points on the statistical image pixel (i, j) is N, d is, x is the x-axis coordinate, y is the y-axis coordinate, and z is the z-axis coordinate;

[0026] The two-dimensional visual depth projection formula is specifically:

[0027]

[0028] In the formula, The minimum distance from the point cloud to the image pixel (i, j) is recorded;

[0029] The two-dimensional statistical frequency projection image and the two-dimensional visual depth projection image corresponding to each rotating monomer point cloud data are spliced to obtain a multi-channel two-dimensional projection image corresponding to each rotating monomer point cloud data.

[0030] All the multi-channel two-dimensional projection images corresponding to the rotating monomer point cloud data are summarized to obtain multi-channel two-dimensional projection data corresponding to each monomer point cloud data.

[0031] As an improvement of the above scheme, the feature fusion of each multi-channel two-dimensional projection data is performed to obtain a multi-view global feature expression corresponding to each monomer point cloud data, comprising:

[0032] In each multi-channel two-dimensional projection data, each multi-channel two-dimensional projection image is calculated based on a preset channel attention formula to obtain a channel attention weight, and each multi-channel two-dimensional projection image is calculated based on a preset spatial attention formula to obtain a spatial attention weight; wherein the channel attention formula is specifically:

[0033]

[0034] In the formula, f C is the channel attention weight, MLP is a multi-layer perception function, and sigmoid is an activation function, is average pooling of the two-dimensional image, is maximum pooling of the two-dimensional image, M 2D is a multi-channel two-dimensional projection image;

[0035] The spatial attention formula is specifically:

[0036]

[0037] In the formula, f S is a spatial attention weight; [·] represents a matrix concatenation operation; conv 7×7 is a 7X7 convolutional layer;

[0038] In each of the multi-channel two-dimensional projection data, according to the channel attention weight and the spatial attention weight, an enhanced image feature matrix of each multi-channel two-dimensional image is calculated;

[0039] In each of the multi-channel two-dimensional projection data, each enhanced image feature matrix is input into a neural network to extract a multi-view global feature expression corresponding to each of the multi-channel two-dimensional projection data.

[0040] As an improvement of the above scheme, the calculation of the enhanced image feature matrix of each multi-channel two-dimensional image according to the channel attention weight and the spatial attention weight comprises:

[0041] The channel attention weight and each multi-channel two-dimensional image are substituted into a channel attention enhancement formula to perform pixel-by-pixel multiplication, to obtain an image feature enhanced by channel attention corresponding to each multi-channel two-dimensional image; wherein the channel attention enhancement formula is specifically:

[0042] M′ 2D = f C (M 2D )⊙M 2D

[0043] In the formula, M′ 2D is an image feature enhanced by channel attention, and is pixel-by-pixel multiplication;

[0044] The spatial attention weight and each image feature enhanced by channel attention are substituted into a spatial attention enhancement formula to perform pixel-by-pixel multiplication, to obtain an enhanced image feature matrix of each multi-channel two-dimensional image; wherein the spatial attention enhancement formula is specifically:

[0045] M″ 2D = f S (M′ 2D )⊙M′ 2D

[0046] In the formula, M″2D to enhance the image feature matrix.

[0047] Accordingly, an embodiment of the present application also provides a three-dimensional point cloud object classification device based on multi-view projection, comprising a data acquisition module, a single body extraction module, a projection module, a feature fusion module and a classification module.

[0048] The data acquisition module is configured to acquire three-dimensional point cloud data of a to-be-tested transformer substation, wherein the to-be-tested transformer substation comprises a plurality of single body devices.

[0049] The single body extraction module is configured to extract single body point cloud data of the plurality of single body devices from the three-dimensional point cloud data.

[0050] The projection module is configured to perform multi-view projection on each single body point cloud data to obtain multi-channel two-dimensional projection data corresponding to each single body point cloud data.

[0051] The feature fusion module is configured to perform feature fusion on each multi-channel two-dimensional projection data to obtain multi-view global feature expression corresponding to each single body point cloud data.

[0052] The classification module is configured to determine a category corresponding to each single body point cloud data based on each multi-view global feature expression, and further obtain the category of each single body in the to-be-tested transformer substation.

[0053] Accordingly, an embodiment of the present application also provides a computer terminal device, comprising a processor, a memory and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement a three-dimensional point cloud object classification method based on multi-view projection as described in the present application.

[0054] Accordingly, an embodiment of the present application also provides a computer readable storage medium, comprising a stored computer program, wherein the computer readable storage medium controls a device where the computer readable storage medium is located to execute a three-dimensional point cloud object classification method based on multi-view projection as described in the present application when the computer program runs.

[0055] As can be seen from the above, the present application has the following advantages:

[0056] The application provides a three-dimensional point cloud object classification method based on multi-view projection, three-dimensional point cloud data of a to-be-tested transformer substation is acquired; wherein the to-be-tested transformer substation comprises a plurality of single devices; single point cloud data of the plurality of single devices is obtained by single extraction on the three-dimensional point cloud data; multi-channel two-dimensional projection data corresponding to each single point cloud data is obtained by multi-view projection on each single point cloud data; multi-view global feature expression corresponding to each single point cloud data is obtained by feature fusion on each multi-channel two-dimensional projection data; and the category corresponding to each single point cloud data is determined based on each multi-view global feature expression. The application splits the transformer substation into single devices, captures the interaction relationship between different view point cloud features based on multi-view projection on the point cloud data corresponding to the single devices, fuses the features captured by multi-view projection, and further enriches the features of the single point cloud data, so as to facilitate better category division of the single devices of the transformer substation and improve the three-dimensional point cloud classification precision of the transformer substation. BRIEF DESCRIPTION OF DRAWINGS

[0057] Figure 1 FIG. 1 is a flowchart of a three-dimensional point cloud object classification method based on multi-view projection provided by an embodiment of the application;

[0058] Figure 2 FIG. 2 is a structural diagram of a three-dimensional point cloud object classification device based on multi-view projection provided by an embodiment of the application;

[0059] Figure 3 FIG. 3 is a structural diagram of a terminal device provided by an embodiment of the application;

[0060] Figure 4 FIG. 4 is a multi-view projection diagram provided by an embodiment of the application;

[0061] Figure 5 FIG. 5 is a flowchart of two-dimensional image generation of multi-view projection provided by an embodiment of the application;

[0062] Figure 6 FIG. 6 is a feature fusion diagram provided by an embodiment of the application;

[0063] Figure 7 FIG. 7 is a flowchart of a three-dimensional point cloud object classification method based on multi-view projection provided by another embodiment of the application. DETAILED DESCRIPTION

[0064] With reference to the drawings and embodiments of the present application, the technical solutions in the embodiments of the present application will be described clearly and completely. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of the present application.

[0065] Embodiment one

[0066] Reference Figure 1 , Figure 1 is a flowchart of a three-dimensional point cloud object classification method based on multi-view projection provided by an embodiment of the present application, as shown in Figure 1 , the embodiment includes steps 101 to 105, and each step is specifically as follows:

[0067] Step 101: obtaining three-dimensional point cloud data of a to-be-tested transformer substation; wherein the to-be-tested transformer substation includes a plurality of single devices.

[0068] In the embodiment, the three-dimensional point cloud data of the to-be-tested transformer substation is the three-dimensional point cloud corresponding to a plurality of single devices in the to-be-tested transformer substation scene.

[0069] Step 102: performing single-body extraction on the three-dimensional point cloud data to obtain single-body point cloud data of the plurality of single devices.

[0070] In the embodiment, the single-body point cloud object automatic extraction technology can be used to perform single-body extraction on the three-dimensional point cloud data.

[0071] Step 103: performing multi-view projection on each single-body point cloud data to obtain multi-channel two-dimensional projection data corresponding to each single-body point cloud data.

[0072] In the embodiment, the multi-view projection on each single-body point cloud data to obtain multi-channel two-dimensional projection data corresponding to each single-body point cloud data includes:

[0073] performing normalization processing on each single-body point cloud data to obtain normalized single-body point cloud data corresponding to each single-body point cloud data;

[0074] determining a split angle based on a preset projection view point;

[0075] in each normalized single-body point cloud data, rotating each normalized single-body point cloud data based on the split angle and a rotation matrix to obtain a plurality of rotated single-body point cloud data, and projecting and splicing each rotated single-body point cloud data to further obtain multi-channel two-dimensional projection data corresponding to each single-body point cloud data; wherein the number of the rotated single-body point cloud data corresponds to the number of times of splitting.

[0076] In a specific embodiment, each of the single point cloud data is centered, the coordinate values thereof are normalized, and a two-dimensional image size and a unit size are specified according to a minimum bounding box of the normalized point cloud.

[0077] In the embodiment, the split angle is determined based on the preset projection viewpoint, and the method comprises:

[0078] A quarter spherical surface of the normalized single point cloud data is selected, the quarter spherical surface is equiangularly split to obtain a trajectory of the projection viewpoint of the normalized single point cloud data;

[0079] A plurality of split angles on each trajectory are determined.

[0080] For better illustration, refer to Figure 4 , considering the symmetry of the point cloud projection, only a quarter spherical surface of the point cloud is selected as the projection viewpoint, and specifically, as shown in Figure 4 , the quarter spherical surface is equiangularly sliced to form the trajectory of each viewpoint (as shown by the red line), and each projection viewpoint is further determined along the trajectory of each viewpoint at the same split angle.

[0081] In the embodiment, each of the normalized single point cloud data is rotated based on the split angle and the rotation matrix to obtain a plurality of rotated single point cloud data, and the method comprises:

[0082] The split angle and the rotation matrix are input into a preset rotation formula to obtain the rotated single point cloud data corresponding to different split angles; wherein the rotation formula comprises:

[0083]

[0084] In the formula, P represents the point cloud after normalization, represents the point cloud after rotation, is a rotation matrix of rotating θ m angle around the x-axis, R n,κ is a rotation matrix of rotating θ n angle around the z-axis, dθ is the split angle, n splitNum is the number of split tracks, d is the pixel size of the two-dimensional projection image, and m and n are rotation parameters.

[0085] In the embodiment, each of the rotated single point cloud data is projected and spliced to further obtain the multi-channel two-dimensional projection data corresponding to each of the single point cloud data, and the method comprises:

[0086] inputting each of the rotation monomer point cloud data into a two-dimensional statistical frequency projection formula to obtain a two-dimensional statistical frequency projection image corresponding to each of the rotation monomer point cloud data; inputting each of the rotation monomer point cloud data into a two-dimensional visual depth projection formula to obtain a two-dimensional visual depth projection image corresponding to each of the rotation monomer point cloud data; wherein the two-dimensional statistical frequency projection formula is specifically:

[0087]

[0088] In the formula, N is the number of points on the image pixel (i, j), d is the distance from the image pixel (i, j) to the nearest point in the point cloud, x is the x-axis coordinate, y is the y-axis coordinate, and z is the z-axis coordinate.

[0089] The two-dimensional visual depth projection formula is specifically:

[0090]

[0091] In the formula, N is the number of points on the image pixel (i, j), d is the distance from the image pixel (i, j) to the nearest point in the point cloud, x is the x-axis coordinate, y is the y-axis coordinate, and z is the z-axis coordinate. record the minimum distance from the point cloud to the image pixel (i, j);

[0092] stitching the two-dimensional statistical frequency projection image and the two-dimensional visual depth projection image corresponding to each of the rotation monomer point cloud data to obtain a multi-channel two-dimensional projection image corresponding to each of the rotation monomer point cloud data;

[0093] aggregate the multi-channel two-dimensional projection images corresponding to all the rotation monomer point cloud data to obtain multi-channel two-dimensional projection data corresponding to each of the monomer point cloud data.

[0094] In a specific embodiment, the point cloud of each view is projected by using the two-dimensional statistical frequency projection formula and the two-dimensional visual depth projection formula, thereby generating two two-dimensional projection images, as shown in Figure 5 It is noted that the two two-dimensional images will be further stitched to form a multi-channel two-dimensional projection image in order to extract global multi-view features.

[0095] Step 104: performing feature fusion on each of the multi-channel two-dimensional projection data to obtain a multi-view global feature expression corresponding to each of the monomer point cloud data.

[0096] In the embodiment, the feature fusion on each of the multi-channel two-dimensional projection data to obtain a multi-view global feature expression corresponding to each of the monomer point cloud data includes:

[0097] ​In each of the multi-channel two-dimensional projection data, each multi-channel two-dimensional projection image is calculated based on a preset channel attention formula to obtain a channel attention weight, and each multi-channel two-dimensional projection image is calculated based on a preset spatial attention formula to obtain a spatial attention weight; wherein the channel attention formula is specifically:

[0098]

[0099] wherein f C (·) is the channel attention weight, MLP is a multi-layer perceptron function, sigmoid is an activation function, is an average pooling of a two-dimensional image, is a maximum value pooling of a two-dimensional image, M 2D is a multi-channel two-dimensional projection image;

[0100] The spatial attention formula is specifically:

[0101]

[0102] wherein f S (·) is the spatial attention weight; [·] represents a matrix splicing operation; conv 7×7 is a 7X7 convolution layer;

[0103] In each of the multi-channel two-dimensional projection data, an enhanced image feature matrix of each multi-channel two-dimensional image is calculated according to the channel attention weight and the spatial attention weight;

[0104] In each of the multi-channel two-dimensional projection data, each enhanced image feature matrix is input into a neural network to extract a multi-view global feature expression corresponding to each of the multi-channel two-dimensional projection data.

[0105] In a specific embodiment, before the feature fusion of each of the multi-channel two-dimensional projection data, it further includes: using a one-dimensional convolution layer to perform normalization processing on each of the multi-channel two-dimensional projection data.

[0106] It can be understood that the channel attention mechanism and the spatial attention mechanism are used to calculate the attention weights between different channel images and the attention weights on different pixel positions of the same image, as shown in the channel attention formula and the spatial attention formula. Wherein f C (·) performs average pooling and maximum pooling at the image or channel level, and uses a multi-layer perceptron and an activation function to learn the attention weights of different channels; f S(·) Perform average pooling and max pooling inside the same image at the pixel level, and use a 7x7 convolution operation and an activation function to learn the attention weights of different pixel positions in the entire image, and [·] represents the matrix concatenation operation.

[0107] In the embodiment, the calculation of the enhanced image feature matrix of each multi-channel two-dimensional image according to the channel attention weight and the spatial attention weight comprises:

[0108] The channel attention weight is substituted into the channel attention enhancement formula to perform pixel-by-pixel multiplication with each multi-channel two-dimensional image, to obtain a channel attention enhanced image feature corresponding to each multi-channel two-dimensional image; wherein the channel attention enhancement formula is specifically:

[0109] M′ 2D = f C (M 2D )⊙M 2D

[0110] In the formula, M′ 2D is the channel attention enhanced image feature, and is pixel-by-pixel multiplication;

[0111] The spatial attention weight is substituted into the spatial attention enhancement formula to perform pixel-by-pixel multiplication with each channel attention enhanced image feature, to obtain an enhanced image feature matrix of each multi-channel two-dimensional image; wherein the spatial attention enhancement formula is specifically:

[0112] M″ 2D = f S (M′ 2D )⊙M′ 2D

[0113] In the formula, M″ 2D is the enhanced image feature matrix.

[0114] It can be understood that the enhanced feature of the normalized multi-channel two-dimensional projection image is calculated based on the channel attention weight and the spatial attention weight, that is, the channel attention weight is first used to perform pixel-by-pixel multiplication with the multi-channel two-dimensional projection image to obtain a channel attention enhanced image feature, and the spatial attention weight is further used to perform pixel-by-pixel multiplication with the channel attention enhanced image feature to obtain a spatial attention enhanced image feature, as shown in the channel attention enhancement formula and the spatial attention enhancement formula. Wherein, represents pixel-by-pixel multiplication, M′ 2D is the image feature matrix obtained using channel attention, and M″ 2D is the image feature matrix obtained using spatial attention.

[0115] To better illustrate feature fusion, refer to the schematic diagram shown in Figure 6 .

[0116] Step 105: Based on each of the multi-view global feature expression, determine the category corresponding to each of the single point cloud data, and then obtain the category of each single in the substation to be tested.

[0117] In the embodiment, the three-dimensional point cloud object classification module mainly consists of the following steps:

[0118] Firstly, the multi-view global feature expression is input into the multi-layer perception graph layer to realize linear transformation and full connection operation;

[0119] Secondly, the maximum pooling operation is used to extract the category information of the multi-view global feature expression after transformation, and the probability vector of each type to which the single point cloud data belongs is obtained;

[0120] Finally, the maximum value in the type probability vector and the corresponding category are selected as the final predicted category of the single point cloud data, and thus the classification task of the single point cloud data is completed.

[0121] To better illustrate, in order to evaluate the three-dimensional point cloud object classification method designed in this embodiment, the model classification results are compared with the most advanced methods in the literature, including PointNet, PointNet++, PointNeXt, PointVector and SPoTr. As shown in Table 1, the present experiment lists the comparison results of the overall accuracy (OA) and the average accuracy (MA) of the present model relative to different methods when evaluated on three data sets. The OA and MA are calculated as shown in equations (1) and (2) below. In addition, the three data sets used in the experiment are ModelNet40, ModelNet40-C and our dataset of substation point clouds. ModelNet40 is a three-dimensional shape classification data set based on CAD drawing and three-dimensional face sampling, which has a total of 40 types and 12318 noise-free and well-shaped point cloud samples, and is the most widely used benchmark in point cloud analysis and three-dimensional shape classification tasks. ModelNet40-C is a challenging three-dimensional point cloud object classification data set, which is constructed based on the regular shape point cloud of ModelNet40 by adding damage disturbance to the data set to form a total of 185000 point cloud classification data samples with noise interference. The two data sets are divided into 80% training set and 20% test set to complete the training and testing of the model, and the training set is further divided into 80% for model training and 20% for model verification. The substation point cloud data set includes BIM three-dimensional models and real scene point clouds scanned by ground-based laser radar scanners. By analyzing the label information of the BIM model, the equipment in the substation scene is classified and arranged to form a training data set of a total of 6372 training samples; the ground-based laser radar collected real scene point cloud is ground point free, and the point cloud is automatically processed and the object type label is manually annotated to form a total of 222 single real scene equipment objects as a test data set.

[0122] For ModelNet40 and ModelNet40-C datasets, the OA values of all methods are greater than 90% and the MA values are greater than 85%, which indicates that these methods are all effective. However, the proposed model performs the best in terms of both OA and MA values, while SPoTr or PointNet performs the worst. For example, on the ModelNet40 dataset, the OA values of the proposed model are 1.75%, 0.49%, 0.09%, 0.04% and 2.51% higher than those of PointNet, PointNet++, PointNeXt, PointVector and SPoTr, respectively, and the MA values are 3.15%, 0.08%, 1.15%, 0.19% and 2.92% higher than those of PointNet, PointNet++, PointNeXt, PointVector and SPoTr, respectively. The proposed model also has good performance on the ModelNet40-C dataset relative to other models. For example, in terms of OA values, the accuracy of the proposed model is improved by 2.77%, 0.88%, 0.59%, 0.24% and 2.74%, respectively; and in terms of MA values, the accuracy of the proposed model is improved by 2.96%, 1.05% and 2.97% compared with PointNet, PointNet++ and SPoTr, respectively, but is slightly lower than PointNeXt and PointVector. This result again proves the effectiveness and robustness of the proposed model.

[0123] In addition, the proposed model is evaluated using our dataset, and the results show that the proposed model performs the best among the state-of-the-art models, but PointNeXt performs the worst. Specifically, in terms of OA values, the performance of the proposed model is 4.03%, 0.36%, 7.98%, 3.40% and 0.36% higher than that of PointNet, PointNet++, PointNeXt, PointVector and SPoTr, respectively; and in terms of MA values, the performance of the proposed model is 16.35%, 18.88% and 5.86% higher than that of PointNet, PointNeXt and PointVector, respectively, but the accuracy is slightly lower than that of SPoTr and PointNet++. It is worth noting that the classification accuracy of PointNet++ ranks first in MA because it can capture the hierarchical features or multi-level organization of three-dimensional objects well, which is particularly suitable for electrical equipment with complex structures. In this regard, it has similarities with the proposed model and has a comparable level in terms of accuracy. Therefore, through quantitative comparison with the state-of-the-art methods, this finding again confirms the effectiveness of the proposed model in three-dimensional point cloud object classification of electrical equipment with complex structures.

[0124]

[0125] Num Correct = Num Correct + 1 correct Num Correct represents the total number of samples whose labels are correctly predicted by the model classifier, Num total Num Total represents the total number of samples used for prediction by the model; Num Correct = Num Correct + 1 Num Correct represents the total number of samples whose labels are correctly predicted by the model classifier, Num cls Num Total represents the total number of samples used for prediction by the model;

[0126] Table 1. Comparison of classification accuracy results of the present model and other advanced models

[0127]

[0128] In a specific embodiment, the flowchart of the three-dimensional point cloud object classification can be seen from Figure 7 .

[0129] Referring to Figure 2 , Figure 2 is a structural schematic diagram of a three-dimensional point cloud object classification device based on multi-view projection provided by an embodiment of the present application, comprising a data acquisition module 201, a single extraction module 202, a projection module 203, a feature fusion module 204 and a classification module 205.

[0130] The data acquisition module is configured to acquire three-dimensional point cloud data of a to-be-tested substation; wherein the to-be-tested substation comprises a plurality of single devices.

[0131] The single extraction module is configured to perform single extraction on the three-dimensional point cloud data to obtain single point cloud data of the plurality of single devices.

[0132] The projection module is configured to perform multi-view projection on each single point cloud data to obtain multi-channel two-dimensional projection data corresponding to each single point cloud data.

[0133] The feature fusion module is configured to perform feature fusion on each multi-channel two-dimensional projection data to obtain multi-view global feature expression corresponding to each single point cloud data.

[0134] The classification module is configured to determine the category corresponding to each single point cloud data based on each multi-view global feature expression, and further obtain the category of each single device in the to-be-tested substation.

[0135] It can be understood that the above system item embodiments correspond to the method item embodiments of the present application, and can realize the multi-view projection based three-dimensional point cloud object classification method provided by any one of the above method item embodiments of the present application.

[0136] The embodiment obtains three-dimensional point cloud data of a to-be-tested transformer substation, wherein the to-be-tested transformer substation includes a plurality of single devices; single device extraction is performed on the three-dimensional point cloud data to obtain single device point cloud data of the plurality of single devices; multi-view projection is performed on each single device point cloud data to obtain multi-channel two-dimensional projection data corresponding to each single device point cloud data; feature fusion is performed on each multi-channel two-dimensional projection data to obtain multi-view global feature expression corresponding to each single device point cloud data; and the category corresponding to each single device point cloud data is determined based on each multi-view global feature expression. The embodiment splits the transformer substation into single devices, captures the interaction relationship between different view point cloud features based on the multi-view projection of the point cloud data corresponding to the single devices, fuses the features captured by the multi-view projection, and further enriches the features of the single device point cloud data, thereby facilitating better category distinction of the single devices of the transformer substation and improving the three-dimensional point cloud classification accuracy of the transformer substation.

[0137] Embodiment two

[0138] Reference is made to Figure 3 , Figure 3 is a terminal device structure diagram provided by an embodiment of the present application.

[0139] The terminal device of the embodiment includes a processor 301, a memory 302, and a computer program stored in the memory 302 and executable on the processor 301. The processor 301 implements the steps of the above various three-dimensional point cloud object classification methods based on multi-view projection in the embodiments when executing the computer program, for example, all steps of the three-dimensional point cloud object classification method based on multi-view projection as shown in Figure 1 . Alternatively, the processor implements the functions of the modules in the above various device embodiments when executing the computer program, for example, all modules of the three-dimensional point cloud object classification device based on multi-view projection as shown in Figure 2 .

[0140] In addition, the embodiment of the present application also provides a computer readable storage medium including a stored computer program, wherein when the computer program runs, the device where the computer readable storage medium is located performs the three-dimensional point cloud object classification method based on multi-view projection as described in any of the above embodiments.

[0141] Those skilled in the art can understand that the diagram is only an example of the terminal device and does not constitute a limitation on the terminal device, and can include more or fewer components than the diagram, or combine certain components, or different components, for example, the terminal device can also include an input / output device, a network access device, a bus, etc.

[0142] The processor 301 can be a central processing unit (CPU), and can also be other general-purpose processors, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor. The processor 301 is a control center of the terminal device, and is connected with various parts of the terminal device through various interfaces and lines.

[0143] The memory 302 can be used to store computer programs and / or modules, and the processor 301 realizes various functions of the terminal device by running or executing the computer programs and / or modules stored in the memory and calling data stored in the memory 302. The memory 302 can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application program required by a function (such as a sound playing function, an image playing function, etc.), etc.; and the data storage area can store data created according to use of the terminal device (such as audio data, a phone book, etc.), etc. In addition, the memory can include a high-speed random access memory, and can also include a nonvolatile memory, for example, a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash storage device, or other volatile solid-state storage device.

[0144] The modules / units integrated in the terminal device, if in the form of software function units and sold or used as independent products, can be stored in a computer readable storage medium. Based on this understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware. The computer program can be stored in a computer readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. The computer program includes computer program code, which can be in the form of source code, object code, executable files or some intermediate forms, etc. The computer readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal and software distribution medium, etc.

[0145] It should be noted that the above-described device embodiments are only schematic, and the units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, i.e. they can be located in one place or distributed on multiple network units. Part or all of the modules can be selected to achieve the purpose of the embodiment scheme according to actual needs. In addition, the connection relationship between the modules in the device embodiment provided by the present application indicates that there is a communication connection between them, which can be realized as one or more communication buses or signal lines. Those skilled in the art can understand and implement it without creative labor.

[0146] The above is the preferred embodiment of the present application. It should be noted that for those skilled in the art, without departing from the principles of the present application, a number of improvements and refinements can be made, which are also considered within the scope of protection of the present application.

Claims

1. A method for three-dimensional point cloud object based on multi-view projection, characterized in that, The application relates to a method for classifying substation equipment, and belongs to the technical field of substation equipment classification. The method comprises the following steps: acquiring three-dimensional point cloud data of a to-be-tested substation; wherein the to-be-tested substation comprises a plurality of single devices; extracting the three-dimensional point cloud data to obtain single point cloud data of the plurality of single devices; performing multi-view projection on each single point cloud data to obtain multi-channel two-dimensional projection data corresponding to each single point cloud data; wherein the multi-view projection on each single point cloud data to obtain multi-channel two-dimensional projection data corresponding to each single point cloud data comprises the following steps: normalizing each single point cloud data to obtain normalized single point cloud data corresponding to each single point cloud data; 2. The multi-view projection based three-dimensional point cloud object method of claim 1, wherein, determining a split angle based on a preset projection viewpoint; rotating each normalized single point cloud data in each normalized single point cloud data based on the split angle and a rotation matrix to obtain a plurality of rotated single point cloud data, and projecting and splicing each rotated single point cloud data to obtain multi-channel two-dimensional projection data corresponding to each single point cloud data; wherein the number of the rotated single point cloud data corresponds to the number of times of splitting; performing feature fusion on each multi-channel two-dimensional projection data to obtain multi-view global feature expression corresponding to each single point cloud data; 3. The multi-view projection based three-dimensional point cloud object method of claim 2, wherein, determining the category of each single point cloud data based on each multi-view global feature expression to obtain the category of each single device in the to-be-tested substation. The method comprises the following steps: wherein, denotes the point cloud after normalization, denotes the point cloud after rotation, is a rotation matrix around the x-axis by an angle of θ m R n,k is a rotation matrix around the z-axis by an angle of dθ n is a rotation matrix around the z-axis by an angle of dθ splitNum is the number of cuts of the trajectory, d is the pixel size of the two-dimensional projection image, and m, n are rotation parameters.

4. The multi-view projection based three-dimensional point cloud object method of claim 3, wherein, selecting a quarter spherical surface of the normalized single point cloud data, and performing equiangular split on the quarter spherical surface to obtain the trajectory of the projection viewpoint of the normalized single point cloud data; determining a plurality of split angles on each trajectory. wherein counting the number of points on the image pixel (i,j), d is the pixel size of the two-dimensional projection image, x is the x-axis coordinate, y is the y-axis coordinate, and z is the z-axis coordinate, is the point cloud after rotation; The method comprises the following steps: In the formula, record the minimum distance of the point cloud to the image pixel (i, j); inputting the split angle and the rotation matrix into a preset rotation formula to obtain rotated single point cloud data corresponding to different split angles; wherein the rotation formula comprises: The method comprises the following steps: inputting each rotated single point cloud data into a two-dimensional statistical frequency projection formula to obtain a two-dimensional statistical frequency projection image corresponding to each rotated single point cloud data; and inputting each rotated single point cloud data into a two-dimensional visual depth projection formula to obtain a two-dimensional visual depth projection image corresponding to each rotated single point cloud data; wherein the two-dimensional statistical frequency projection formula is specifically as follows: The two-dimensional visual depth projection formula is specifically as follows: splicing the two-dimensional statistical frequency projection image and the two-dimensional visual depth projection image corresponding to each rotated single point cloud data to obtain multi-channel two-dimensional projection images corresponding to each rotated single point cloud data; collecting the multi-channel two-dimensional projection images corresponding to all the rotated single point cloud data to obtain multi-channel two-dimensional projection data corresponding to each single point cloud data.

5. The multi-view projection based three-dimensional point cloud object method of claim 4, wherein, The feature fusion of each multi-channel two-dimensional projection data obtains a multi-view global feature expression corresponding to each single point cloud data, including: In each multi-channel two-dimensional projection data, the channel attention weight of each multi-channel two-dimensional projection image is calculated based on a preset channel attention formula, and the spatial attention weight of each multi-channel two-dimensional projection image is calculated based on a preset spatial attention formula; wherein the channel attention formula is specifically: where f C (·) is the channel attention weight, MLP is a multi-layer perceptron function, sigmoid is an activation function, is the average pooling of the two-dimensional image, is the maximum value pooling of the two-dimensional image, M 2D is a multi-channel two-dimensional projection image; The spatial attention formula is specifically: where f S (·) is the spatial attention weight; [·] denotes the matrix concatenation operation; conv 7×7 is a 7X7 convolutional layer; In each multi-channel two-dimensional projection data, the enhanced image feature matrix of each multi-channel two-dimensional image is calculated according to the channel attention weight and the spatial attention weight; In each multi-channel two-dimensional projection data, each enhanced image feature matrix is input into a neural network to extract a multi-view global feature expression corresponding to each multi-channel two-dimensional projection data.

6. The multi-view projection based three-dimensional point cloud object method of claim 5, wherein, The calculation of the enhanced image feature matrix of each multi-channel two-dimensional image according to the channel attention weight and the spatial attention weight includes: The channel attention weight and each multi-channel two-dimensional image are substituted into the channel attention enhancement formula to perform pixel-by-pixel multiplication to obtain an image feature corresponding to each multi-channel two-dimensional image after channel attention enhancement; wherein the channel attention enhancement formula is specifically: M' 2D = f C (M 2D )⊙M 2D Where M′ 2D is the image feature after channel attention enhancement, ⊙ is pixel-by-pixel multiplication; The spatial attention weight and each channel attention enhanced image feature are substituted into the spatial attention enhancement formula to perform pixel-by-pixel multiplication to obtain the enhanced image feature matrix of each multi-channel two-dimensional image; wherein the spatial attention enhancement formula is specifically: M" 2D = f S (M' 2D )⊙M' 2D In the formula, M" 2D to enhance the image feature matrix.

7. A multi-view projection based three-dimensional point cloud object device, characterized by, including: Data acquisition module, single extraction module, projection module, feature fusion module and classification module; The data acquisition module is configured to acquire three-dimensional point cloud data of a to-be-tested transformer substation; wherein the to-be-tested transformer substation includes a plurality of single devices; The single extraction module is configured to perform single extraction on the three-dimensional point cloud data to obtain single point cloud data of the plurality of single devices; The projection module is configured to perform multi-view projection on each single point cloud data to obtain multi-channel two-dimensional projection data corresponding to each single point cloud data; wherein the multi-view projection on each single point cloud data to obtain multi-channel two-dimensional projection data corresponding to each single point cloud data includes: performing normalization processing on each single point cloud data to obtain normalized single point cloud data corresponding to each single point cloud data; determining a split angle based on a preset projection view point; in each normalized single point cloud data, rotating each normalized single point cloud data based on the split angle and a rotation matrix to obtain a plurality of rotated single point cloud data, and projecting and splicing each rotated single point cloud data to obtain multi-channel two-dimensional projection data corresponding to each single point cloud data; wherein the number of rotated single point cloud data corresponds to the number of splits; The feature fusion module is configured to perform feature fusion on each multi-channel two-dimensional projection data to obtain a multi-view global feature expression corresponding to each single point cloud data; The classification module is configured to determine the category corresponding to each single point cloud data based on each multi-view global feature expression, and further obtain the category of each single body in the substation to be measured.

8. A computer terminal device, characterized by A computer program product including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, the processor implementing the method of claim 1-6 when executing the computer program.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium includes a stored computer program, wherein the computer readable storage medium controls the device where the computer readable storage medium is located to execute the method of claim 1-6 when the computer program is running.

Citation Information

Patent Citations

  • Substation equipment object segmentation method and device, equipment and storage medium

    CN116912532A

  • Three-dimensional point cloud identification device, learning device, three-dimensional point cloud identification method, learning method and program

    US20230040195A1