A point cloud data processing method, system and computer device based on multi-channel feature extraction
By using multi-channel feature extraction and learnable parameter matrix processing, the problem of information redundancy in 3D point cloud data processing is solved, achieving more efficient and accurate point cloud data segmentation. In particular, it significantly improves the accuracy of component segmentation in applications such as autonomous driving and intelligent perception.
Patent Information
- Application Number
- CN202310915767.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-25
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2043-07-25
AI Technical Summary
Existing technologies for 3D point cloud data processing suffer from problems such as information redundancy caused by single-channel sampling, low processing accuracy, and slow efficiency. In particular, it is difficult to accurately identify the type of point cloud components in the fields of autonomous driving and intelligent perception.
A multi-channel feature extraction method is adopted, which extracts and fuses features from point cloud sub-data by setting multiple KNN sampling channels. Redundancy is reduced by combining a learnable parameter matrix, and the weight parameters are optimized by using the Log Softmax function and the Negative Log-Likelihood Loss function to achieve efficient segmentation of point cloud data.
It improves the accuracy and speed of point cloud data processing, enhances feature sparsity, and improves the accuracy of point cloud segmentation, especially significantly improving the accuracy of component segmentation in applications such as autonomous driving and intelligent perception.
Smart Images

Figure CN117058372B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of point cloud data intelligent analysis, and particularly relates to a point cloud data processing method and system based on multi-channel feature extraction and a computer device. BACKGROUND
[0002] With the rapid development of artificial intelligence and machine learning technology, the application of deep learning algorithm, target point detection as one of the important research directions in the field of computer vision has been widely applied in the fields of automatic driving, intelligent perception, robot positioning, etc.
[0003] 3D point cloud data, as an important data source for application scenarios such as automatic driving and intelligent perception, provides original geometric information and rich shape ratio information, but the original 3D point cloud data includes various data information, and the components to be detected and analyzed need to be segmented from the point cloud data to more accurately extract the shape and structure information of the required components.
[0004] When the existing data algorithm model performs component segmentation on 3D point cloud data, since the coordinate vector of the point cloud has only three dimensions of x, y and z, in order to fully utilize the three dimensions, the data algorithm model generally performs dimensionality increasing on a 1x3 point cloud coordinate through convolution, such as increasing the dimension to 1x1024. After dimensionality increasing, more useful features can be extracted to distinguish the component category of the current point cloud. However, dimensionality increasing of the point cloud vector will also bring another problem: information redundancy. Information redundancy will make each dimension of the 1x1024 point cloud vector have little difference (i.e. the numerical values of each dimension are not much different), so that the model cannot accurately distinguish which component the current point cloud belongs to.
[0005] Further, when the existing PointNet++ algorithm model is used for sampling, it is a single-channel sampling method, and the information extracted by feature sampling is relatively small. In the down-sampling operation, the detail information is easily lost, and in the convolution coding process, the multi-scale features cannot be well extracted, so that the processing precision is low and the efficiency is slow. SUMMARY
[0006] In order to overcome the defects of the prior art, the purpose of the present application is to provide a point cloud data processing method, system and computer device based on multi-channel feature extraction, which is mainly used to solve the defects of low processing precision and slow efficiency caused by single-channel sampling and data redundancy when the 3D point cloud data is intelligently analyzed in the prior art.
[0007] To solve the above problems, the technical scheme adopted by the present application is as follows:
[0008] In a first aspect, the present application provides a point cloud data processing method based on multi-channel feature extraction, comprising:
[0009] collecting 3D point cloud data, the 3D point cloud data comprising N groups of point cloud sub-data with different sampling attributes;
[0010] According to the number of sampling attributes, M KNN sampling channels are set, each KNN sampling channel extracts features from a group of point cloud sub-data to obtain M sampling results, and the M sampling results are fused to obtain point cloud fusion features J, wherein M is less than or equal to N;
[0011] The point cloud fusion features J are up-sampled in combination with the M down-sampled results to obtain a first point cloud feature vector H with a length of E;
[0012] The first point cloud feature vector H is enhanced in feature, a learnable parameter matrix with a dimension of E×E is established, the first point cloud feature vector H is multiplied by the learnable parameter matrix, and the weight parameters in the learnable parameter matrix are used to weaken redundant data to obtain a second point cloud feature vector W;
[0013] The second point cloud feature vector W is reduced in dimension by a convolution layer L, and probability scores of the current point cloud under different sampling attributes are obtained by processing the second point cloud feature vector W by a Log Softmax function, a loss value is obtained by using a Negative Log-Likelihood Loss loss function, and the weight parameters are updated by back propagation;
[0014] The segmentation result of the 3D point cloud data based on different sampling attributes in multiple channels is output.
[0015] In some embodiments, M sampling points are set, each sampling point has different sampling attributes, each sampling point corresponds to a group of point cloud sub-data which are independent of each other or have intersections, and M = 3;
[0016] For the first sampling point, in the first sampling channel, 16 unit point clouds around the first sampling point are selected as the first point cloud sub-data, one-dimensional convolution is used to up-sample the three-dimensional point cloud to 64 dimensions, and the first sampling result is obtained;
[0017] For the second sampling point, in the second sampling channel, 32 unit point clouds around the second sampling point are selected as the second point cloud sub-data, one-dimensional convolution is used to up-sample the three-dimensional point cloud to 128 dimensions, and the second sampling result is obtained;
[0018] For the third sampling point, in the third sampling channel, 128 unit point clouds around the third sampling point are selected as the third point cloud sub-data, one-dimensional convolution is used to up-sample the three-dimensional point cloud to 128 dimensions, and the third sampling result is obtained.
[0019] In some embodiments, the first sampling result, the second sampling result and the third sampling result are sequentially spliced, encoded by three convolutional layers, and dimensioned to 1024 dimensions to obtain a point cloud fusion feature J.
[0020] In some embodiments, a set of learnable parameter matrices S, T, V with E×E dimensions are set, and the first point cloud feature vector H includes a first sub-vector a, a second sub-vector b, and a third sub-vector c.
[0021] The first sub-vector a, the second sub-vector b, and the third sub-vector c are multiplied by S, T, and V to obtain a feature correlation vector, a feature inhibition vector, and an information vector, the feature correlation vector includes Sa, Sb, and Sc, the feature inhibition vector includes Ta, Tb, and Tc, and the information vector includes Va, Vb, and Vc.
[0022] The inner product of each feature correlation vector Sa, Sb, and Sc with all feature inhibition vectors Ta, Tb, and Tc is calculated to obtain corresponding weight scores Qa`, Qb`, and Qc`, and the weight scores Qa`, Qb`, and Qc` are weighted and summed with the corresponding information vectors Va, Vb, and Vc to obtain a weighted vector.
[0023] The weighted vector is batch normalized and activated by a ReLU function, and all weighted vectors are spliced into a data sequence (a`, b`, c`) to obtain a second point cloud feature vector W.
[0024] In some embodiments, a residual network is established, and the first sub-vector a, the second sub-vector b, and the third sub-vector c are added to the data sequence (a`, b`, c`) to obtain the second point cloud feature vector W.
[0025] In some embodiments, the point cloud fusion feature J and the down-sampling result are input into a two-layer cascaded feature propagation module to obtain a first point cloud feature vector H with a data format of 512×128×2048.
[0026] Wherein, 512 represents the number of training samples taken from the training set in each training of the first point cloud feature vector H, 128 represents the dimension of the first point cloud feature vector H, and 2048 represents the number of point clouds of the first point cloud feature vector H.
[0027] In some embodiments, the learnable parameter matrices S, T, and V are realized by a 1×1 convolutional layer with an input dimension of 128 and an output dimension of 128, the information vectors Va, Vb, and Vc are used to quantize the important features of the first sub-vector a, the second sub-vector b, and the third sub-vector c, respectively, and based on the numerical values of the information vectors Va, Vb, and Vc, the dimension value of the sub-vector with the larger numerical information vector is enhanced in the second point cloud feature vector W, and the dimension value of the sub-vector with the smaller numerical information vector tends to 0.
[0028] In some embodiments, the second point cloud feature vector W is dimensionally reduced by the convolution layer L, and the dimension number of the obtained result is equal to the number of component categories, and a Log Softmax function is used to increase the probability value distinguishing degree of the point cloud belonging to different component categories.
[0029] In a second aspect, the present application provides a point cloud data processing system based on multi-channel feature extraction, which applies the point cloud data processing method described above, and comprises:
[0030] A data acquisition module is configured to acquire 3D point cloud data, wherein the 3D point cloud data comprises N groups of point cloud sub-data with different sampling attributes.
[0031] A multi-channel sampling module is configured to set M KNN sampling channels according to the number of sampling attributes, and each KNN sampling channel is configured to perform feature extraction on a group of point cloud sub-data to obtain M sampling results, wherein M is less than or equal to N.
[0032] A feature fusion module is configured to perform feature fusion on the M sampling results to splice a point cloud fusion feature J.
[0033] A feature propagation module is configured to combine the M down-sampling results to perform up-sampling processing on the point cloud fusion feature J to obtain a first point cloud feature vector H with a length of E.
[0034] A feature enhancement module is configured to perform feature enhancement on the first point cloud feature vector H, establish a learnable parameter matrix with a dimension of E×E, multiply the first point cloud feature vector H by the learnable parameter matrix, weaken redundant data through weight parameters in the learnable parameter matrix, and obtain a second point cloud feature vector W.
[0035] A data output module is configured to perform dimension reduction on the second point cloud feature vector W by a convolution layer L, perform Log Softmax function processing to obtain probability scores of a current point cloud under different sampling attributes, use a Negative Log-LikelihoodLoss loss function to obtain a loss value, update weight parameters through back propagation, and finally output a segmentation result of 3D point cloud data based on different sampling attributes under multi-channels.
[0036] In a third aspect, the present application provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the above method when executing the computer program.
[0037] Compared with the prior art, the present application has at least the following beneficial effects:
[0038] (1) By setting multiple KNN sampling channels, feature extraction based on different sampling points is performed on the original 3D point cloud data, and then M sampling results are fused to realize multi-channel KNN sampling, expand the point cloud sampling range, increase multi-scale information, provide more data for the decision layer, and speed up the processing speed;
[0039] (2) In order to reduce redundancy, the first point cloud feature vector H is enhanced, multiplied by an E*E dimension learnable parameter matrix, and the weight parameter is used to weaken the redundant data, increase the feature sparsity, and improve the point cloud segmentation accuracy.
[0040] The application will be further described in detail below in combination with the drawings and specific embodiments. BRIEF DESCRIPTION OF DRAWINGS
[0041] The application will be further described in detail below in combination with the drawings and specific embodiments.
[0042] Figure 1 is a flowchart of the point cloud data processing method provided by the embodiment.
[0043] Figure 2 is a detailed flowchart of the point cloud data processing method provided by the embodiment.
[0044] Figure 3 is a schematic diagram of the multi-channel sampling module provided by the embodiment.
[0045] Figure 4 is a schematic diagram of the feature fusion module provided by the embodiment.
[0046] Figure 5 is a schematic diagram of the feature enhancement module provided by the embodiment.
[0047] Figure 6 is a schematic diagram of the point cloud data processing system provided by the embodiment. DETAILED DESCRIPTION
[0048] The technical solutions of the application will be described in detail below in combination with the drawings. Obviously, the described embodiments are part of the embodiments of the application, rather than all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the application.
[0049] In the description of the present application, it should be noted that the terms "center", "upper", "lower", "left", "right", "vertical", "horizontal", "inner", "outer" and the like indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, and are only for the convenience of describing the present application and simplifying the description, and do not indicate or imply that the devices or elements referred to must have a particular orientation, be constructed and operated in a particular orientation, and therefore cannot be understood as limiting the present application. In addition, the terms "first", "second", "third" are only for descriptive purposes and cannot be understood as indicating or implying relative importance.
[0050] In the description of the present application, when it is described that a specific device is located between a first device and a second device, there can be an intervening device between the specific device and the first device or the second device, or there can be no intervening device. When it is described that a specific device is connected to other devices, the specific device can be directly connected to the other devices without an intervening device, or can not be directly connected to the other devices with an intervening device.
[0051] Techniques, methods, and equipment known to those of ordinary skill in the relevant art can not be discussed in detail, but in appropriate cases, the techniques, methods, and equipment should be considered part of the specification.
[0052] The inventors have found that:
[0053] The existing PointNet++ is a mature and commonly used deep learning algorithm in the field of point cloud processing, mainly used to process point cloud data and segment the objects to be detected from the point cloud. This algorithm mainly performs single-channel KNN sampling on the entire frame of point cloud data, and after convolutional coding and dimensionality increasing for feature extraction, it enters the second layer of single-channel KNN sampling, and sequentially passes through four layers of operation before entering subsequent feature propagation and convolution processing. This approach results in less sampling data information and slow inference speed, and dimensionality increasing processing causes a large amount of information redundancy.
[0054] Especially in the field of autonomous driving or intelligent perception, when estimating the pose of an object, due to the current single-channel sampling and information redundancy problem, the algorithm model cannot accurately determine which component the current point cloud belongs to, resulting in low segmentation accuracy.
[0055] In view of this, in order to solve the above existing problems, in a first aspect, with reference to Figures 1 to 5 The embodiments of the present application provide a point cloud data processing method based on multi-channel feature extraction, comprising:
[0056] Collect 3D point cloud data, and the 3D point cloud data includes N groups of point cloud sub-data with different sampling attributes; wherein, for a scene or a whole 3D point cloud data, the point cloud sub-data in different regions are divided into different point cloud sub-data, and different point cloud sub-data have different sampling attributes, for example, in the field of automatic driving, some point cloud sub-data represent buildings, some point cloud sub-data represent cars, and some point cloud sub-data represent pedestrians, different objects have different sampling attributes, so different point cloud sub-data are divided into different groups according to different sampling attributes;
[0057] According to the number of sampling attributes, M KNN sampling channels are set, each KNN sampling channel extracts features of a group of point cloud sub-data to obtain M sampling results, and the M sampling results are fused to obtain point cloud fusion features J, wherein M is less than or equal to N; it should be noted that in some embodiments, for example, in the scene of object pose estimation, the nose and wings of the aircraft need to be segmented, so by setting, the sampling attribute of the target is selected, and the KNN sampling channel is configured for the point cloud sub-data with the target sampling attribute, and one KNN sampling channel is used to extract features of a group of point cloud sub-data, so that the geometric information of the nose and wings, such as shape, size, position, orientation and curvature, can be accurately extracted;
[0058] The M down-sampled down-sampled results are combined to up-sample the point cloud fusion features J to obtain a first point cloud feature vector H with a length of E;
[0059] The first point cloud feature vector H is enhanced in feature, a learnable parameter matrix with a dimension of E×E is established, the first point cloud feature vector H is multiplied by the learnable parameter matrix, and the weight parameters in the learnable parameter matrix are used to weaken the redundant data to obtain a second point cloud feature vector W; wherein, the weight parameters in the learnable parameter matrix can be obtained through model training, which is used to highlight the features of the parts that need to be accurately extracted, and suppress the remaining useless features, so as to remove the redundant information generated in the middle process of the model and highlight the key features;
[0060] The second point cloud feature vector W is reduced in dimension through a convolution layer L, and the probability score of the current point cloud under different sampling attributes is obtained through a Log Softmax function, and the loss value is obtained by using a Negative Log-Likelihood Loss loss function, and the weight parameters are updated through back propagation;
[0061] After the algorithm model is trained, the 3D point cloud data is input, and the segmentation result of the 3D point cloud data based on different sampling attributes under multiple channels is finally output.
[0062] In combination with Figure 3In this embodiment, taking M=3 as an example, three sampling points and three sampling channels are set, each sampling point has different sampling attributes with the sampling channel, and each sampling channel corresponds to a group of point cloud sub-data which are independent of each other or have intersections; wherein, the point cloud sub-data has two modes of being independent of each other and having intersections, for example, when detecting the attitude of an airplane, different point cloud sub-data are independent of each other, but in the field of automatic driving, the point cloud data of a car and a building may have intersections, so in this embodiment, only the sampling attributes of the point cloud sub-data are classified, and the sampling attributes are used as information identifiers to set corresponding sampling points;
[0063] For the first sampling point, 16 unit point clouds around the first sampling point are selected as the first point cloud sub-data in the first sampling channel, and one-dimensional convolution is used to upgrade the three-dimensional point cloud to 64 dimensions to obtain the first sampling result;
[0064] For the second sampling point, 32 unit point clouds around the second sampling point are selected as the second point cloud sub-data in the second sampling channel, and one-dimensional convolution is used to upgrade the three-dimensional point cloud to 128 dimensions to obtain the second sampling result; preferably, at least a part of the important sampling points are defined as the second sampling points;
[0065] For the third sampling point, 128 unit point clouds around the third sampling point are selected as the third point cloud sub-data in the third sampling channel, and one-dimensional convolution is used to upgrade the three-dimensional point cloud to 128 dimensions to obtain the third sampling result; preferably, at least a part of the important sampling points are defined as the third sampling points.
[0066] In combination Figure 4 The first sampling result, the second sampling result and the third sampling result are sequentially spliced, and are encoded by three convolution layers to upgrade to 1024 dimensions to obtain the point cloud fusion feature J.
[0067] It should be noted that after splicing the three sampling results, a 320-dimensional data sequence is obtained, which is first processed by the first convolution layer, normalization and ReLU function to obtain a 256-dimensional data sequence, then processed by the second convolution layer, normalization and ReLU function to obtain a 512-dimensional data sequence, and finally processed by the third convolution layer, normalization and ReLU function to obtain a 1024-dimensional data sequence, and the 1024-dimensional data sequence is taken as the point cloud fusion feature J.
[0068] Compared with the traditional processing mode, the 1024-dimensional point cloud fusion feature J obtained by splicing mode fuses the features of multiple sampling points, and for important sampling points, the more dimensions they occupy, the more conducive to subsequent data analysis.
[0069] As an implementation, the point cloud fusion feature J and the down-sampling result are input into a two-level cascaded feature propagation module to obtain a first point cloud feature vector H with a data format of 512x128x2048.
[0070] More specifically, the three sampling results are down-sampled by the farthest point sampling algorithm, so that the sampling results can uniformly cover the entire point cloud, and each sampled point is different, unlike random sampling which may sample the same point. The point cloud fusion feature J and the three down-sampled down-sampling results are input as data into the first layer feature propagation module, and the output of the first layer feature propagation module is input into the second layer feature propagation module. The first point cloud feature vector H is obtained after output up-sampling of the second layer feature propagation module.
[0071] Among them, 512 represents the number of training samples taken in the training set each time the first point cloud feature vector H is trained, 128 represents the dimension of the first point cloud feature vector H, and 2048 represents the number of point clouds of the first point cloud feature vector H.
[0072] In combination with Figure 5 As an implementation, when the first point cloud feature vector H is enhanced, a set of learnable parameter matrices S, T, and V with ExE dimensions are set, S, T, and V are three different parameter matrices, and the specifications of S, T, and V are all ExE. The first point cloud feature vector H includes a first sub-vector a, a second sub-vector b, and a third sub-vector c.
[0073] The first sub-vector a, the second sub-vector b, and the third sub-vector c are multiplied by S, T, and V to obtain a feature correlation vector, a feature inhibition vector, and an information vector. The feature correlation vector includes Sa, Sb, and Sc, the feature inhibition vector includes Ta, Tb, and Tc, and the information vector includes Va, Vb, and Vc.
[0074] respectively, and the inner product of Sa and Tc is Qc1; similarly, Sb and Sc, after the inner product of Ta, Tb, Tc respectively, Qa2, Qb2, Qc2 are obtained, and after the inner product of Ta, Tb, Tc respectively, Qa3, Qb3, Qc3 are obtained; Qa1, Qa2, Qa3 are specific numerical values, collectively referred to as Qa`, and similarly for Qb`, Qc`;
[0075] The weighted score Qa`, Qb`, Qc` is weighted and summed with the corresponding information vector Va, Vb, Vc to obtain a weighted vector, which is used to represent the corresponding input feature;
[0076] The weighted vector is batch normalized and activated by a ReLU function, and all the weighted vectors are spliced into a data sequence (a`, b`, c`) to obtain a second point cloud feature vector W; wherein, the batch normalization is used to process the data to speed up the inference and reduce the internal covariance movement of the data. The batch normalized ReLU function sets the output of negative input value to 0 and keeps the positive input value unchanged, realizes nonlinear mapping, and also makes the data sparse, thereby improving the accuracy.
[0077] Preferably, a residual network is established, and the first sub-vector a, the second sub-vector b and the third sub-vector c are added to the data sequence (a`, b`, c`), which is equivalent to adding the input data of the feature enhancement module to its data output, thereby constructing a residual network to obtain the second point cloud feature vector W.
[0078] As an implementation manner, the parameter matrices S, T and V can be learned by a 1x1 convolution layer with an input dimension of 128 and an output dimension of 128, the information vectors Va, Vb and Vc are used to quantize the important features of the first sub-vector a, the second sub-vector b and the third sub-vector c respectively, and based on the numerical value of the information vectors Va, Vb and Vc, the dimension value of the information vector with larger numerical value is enhanced in the second point cloud feature vector W, and the dimension value of the information vector with smaller numerical value tends to 0.
[0079] Since the first sub-vector a, the second sub-vector b and the third sub-vector c have embodied the importance after being processed by the feature propagation module, in this step, the feature enhancement can be further performed, and for more important sampling points, the values derived from the sub-vectors and V will be larger, so the dimension values of the sub-vectors with larger numerical information vectors in the second point cloud feature vector W are enhanced; otherwise, the same reasoning applies. Since each component category corresponds to several dimensions of 1x1024, by the above method, the dimension values of a certain component category can be increased in the second point cloud feature vector W, and the dimension values irrelevant to the current component category can be reduced and tend to 0, thereby realizing accurate feature highlighting of different component categories in the 3D point cloud data and improving the accuracy of component segmentation.
[0080] wherein the dimension values obtained after the second point cloud feature vector W is reduced in dimension by the convolution layer L are equal to the number of component categories, and the Log Softmax function is used to increase the probability value difference of the point cloud belonging to different component categories, so that the decision layer can more easily determine the category to which the current point cloud belongs, and the gradient calculation of the Log Softmax function is very simple, which can accelerate the convergence speed.
[0081] Reference Figure 6 In a second aspect, the embodiments of the present application provide a point cloud data processing system based on multi-channel feature extraction, which applies the point cloud data processing method in the above embodiments, and comprises:
[0082] A data acquisition module is configured to acquire 3D point cloud data, and the 3D point cloud data comprises N groups of point cloud sub-data with different sampling attributes; optionally, the data acquisition module is a laser radar module, and the laser radar module is used to acquire and generate 3D point cloud data;
[0083] A multi-channel sampling module is configured to set M KNN sampling channels according to the number of sampling attributes, each KNN sampling channel is configured to extract features from a group of point cloud sub-data to obtain M sampling results, and M is less than or equal to N;
[0084] A feature fusion module is configured to fuse the M sampling results to splice point cloud fusion features J;
[0085] A feature propagation module is configured to combine the M down-sampling results to perform up-sampling on the point cloud fusion features J to obtain a first point cloud feature vector H with a length of E;
[0086] A feature enhancement module is configured to perform feature enhancement on the first point cloud feature vector H, establish a learnable parameter matrix with a dimension of Ex E, multiply the first point cloud feature vector H by the learnable parameter matrix, weaken redundant data through the weight parameters in the learnable parameter matrix, and obtain a second point cloud feature vector W;
[0087] The data output module is configured to reduce dimensionality of the second point cloud feature vector W through the convolution layer L, obtain probability scores of the current point cloud under different sampling attributes through a Log Softmax function, obtain a loss value by using a Negative Log-Likelihood Loss loss function, update weight parameters through back propagation, and finally output a segmentation result of the 3D point cloud data based on different sampling attributes of the multi-channel.
[0088] Embodiment 1
[0089] In this embodiment, the above point cloud data processing method and system are applied to component segmentation of autonomous driving. 3D point cloud data of different objects is collected by a laser radar, and the 3D point cloud data includes point cloud sub-data of objects such as buildings, cars, and pedestrians, i.e., is divided into different categories. By setting the importance or prominence of different component categories, corresponding KNN sampling channels are used to extract features of corresponding point cloud sub-data, and then the sampling results are spliced. Through the feature enhancement of the feature enhancement module, the dimension values of the required component categories can be highlighted, and the component segmentation accuracy is improved. Experiments show that after feature enhancement by this embodiment, the segmentation accuracy of the model is improved from 0.817 to 0.82.
[0090] Embodiment 2
[0091] In this embodiment, the above point cloud data processing method and system are applied to intelligent perception. Point cloud component segmentation mainly divides different parts of an object, such as the head, wings, and tail of an airplane. Through the segmentation result, the geometric information (shape, size, position, orientation, and curvature, etc.) of different parts of the current object can be known, which can be used to guide industrial production. For example, in the test process of an airplane leaving the factory, the point cloud of the entire airplane is collected by a radar, and the curvature of the airplane head calculated by point cloud segmentation does not meet the design requirements, which can remind the manufacturer to modify the curvature of the airplane head.
[0092] Embodiment 3
[0093] In the embodiment, the point cloud data processing method and system are applied to object posture estimation. By segmenting the point cloud into different components, the shape and structure information of the object can be easily extracted. For example, a frame of point cloud records the point cloud data of a person at a certain time. By finding the relative positions of different parts through component segmentation, it can be known that the person is in what posture at that time. For example, the relative positions of the hands, head and feet of a person running are different from those of a person squatting. When running, the head and feet are farther apart; when squatting, the head and feet are closer. The point cloud component segmentation determines the posture of the person at a certain time by finding the relative positions of the hands, head and feet.
[0094] In a third aspect, an embodiment of the present application provides a computer device, including a memory and a processor. The memory stores a computer program. When the processor executes the computer program, the steps of the above method are implemented.
[0095] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium. The computer readable storage medium can be a non-volatile computer readable storage medium. The computer readable storage medium stores a computer program. When the computer program is executed by a processor, the processor executes the order management method as described above.
[0096] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices, apparatuses and units can refer to the corresponding processes in the foregoing method embodiments, which will not be described here. Those skilled in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized by electronic hardware, computer software or a combination of both. In order to clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been described in a general manner in the foregoing description. Whether the functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0097] In summary, compared with the prior art, the above embodiment provides a point cloud data processing method, system and computer device based on multi-channel feature extraction. By setting multiple KNN sampling channels, the original 3D point cloud data is subjected to feature extraction based on different sampling points, and then M sampling results are subjected to feature fusion, so as to realize multi-channel KNN sampling, expand the point cloud sampling range, increase multi-scale information, provide more data for the decision layer, and speed up the processing speed.
[0098] In order to reduce redundancy, the first point cloud feature vector H is enhanced by multiplication with a learnable parameter matrix of dimension E*E, which reduces the redundancy of the data by using weight parameters, increases the sparsity of the features, and improves the accuracy of point cloud segmentation.
[0099] The above embodiments are only preferred embodiments of the present application, and cannot be used to limit the scope of protection of the present application, and any non-essential changes and substitutions made by those skilled in the art on the basis of the present application are within the scope of protection of the present application.
Claims
1. A point cloud data processing method based on multi-channel feature extraction, characterized in that, The application relates to a method for segmenting 3D point cloud data based on different sampling attributes in multiple channels. The method comprises the following steps: Collecting 3D point cloud data, wherein the 3D point cloud data comprises N groups of point cloud sub-data with different sampling attributes; According to the number of sampling attributes, M KNN sampling channels are set, each KNN sampling channel extracts features from a group of point cloud sub-data, M sampling results are obtained, and the M sampling results are fused to obtain point cloud fusion features J, wherein M is less than or equal to N; M down-sampling results are combined to perform up-sampling on the point cloud fusion features J to obtain a first point cloud feature vector H with a length of E; The first point cloud feature vector H is subjected to feature enhancement, a learnable parameter matrix with a dimension of E*E is established, the first point cloud feature vector H is multiplied by the learnable parameter matrix, the weight parameters in the learnable parameter matrix are used to weaken redundant data, and a second point cloud feature vector W is obtained; The second point cloud feature vector W is subjected to dimension reduction through a convolution layer L, and probability scores of a current point cloud under different sampling attributes are obtained through a Log Softmax function, a loss value is obtained by using a Negative Log-Likelihood Loss loss function, and the weight parameters are updated through back propagation; 2. The point cloud data processing method based on multi-channel feature extraction according to claim 1, characterized in that, The method outputs a segmentation result of the 3D point cloud data based on different sampling attributes in multiple channels. M sampling points are set, each sampling point has different sampling attributes, each sampling point corresponds to a group of point cloud sub-data which are independent or have intersections, and M=3; For the first sampling point, 16 unit point clouds around the first sampling point are selected as first point cloud sub-data in the first sampling channel, one-dimensional convolution is used to upgrade the three-dimensional point cloud to 64 dimensions, and a first sampling result is obtained; For the second sampling point, 32 unit point clouds around the second sampling point are selected as second point cloud sub-data in the second sampling channel, one-dimensional convolution is used to upgrade the three-dimensional point cloud to 128 dimensions, and a second sampling result is obtained; 3. The method of claim 2, wherein, For the third sampling point, 128 unit point clouds around the third sampling point are selected as third point cloud sub-data in the third sampling channel, one-dimensional convolution is used to upgrade the three-dimensional point cloud to 128 dimensions, and a third sampling result is obtained.
4. The point cloud data processing method based on multi-channel feature extraction of claim 3, wherein, The first sampling result, the second sampling result and the third sampling result are sequentially spliced, three convolution layers are used for coding, and the dimension is upgraded to 1024 dimensions, and point cloud fusion features J are obtained. A group of learnable parameter matrices S, T and V with a dimension of E*E are set, and the first point cloud feature vector H comprises a first sub-vector a, a second sub-vector b and a third sub-vector c; The first sub-vector a, the second sub-vector b and the third sub-vector c are multiplied by S, T and V to obtain a feature correlation vector, a feature inhibition vector and an information vector, the feature correlation vector comprises Sa, Sb and Sc, the feature inhibition vector comprises Ta, Tb and Tc, and the information vector comprises Va, Vb and Vc. Respectively, each feature correlation vector Sa, Sb, Sc and all feature inhibition vectors Ta, Tb, Tc are inner product, and the corresponding weight score Qa`, Qb`, Qc` is obtained. The weight score Qa`, Qb`, Qc` and the corresponding information vector Va, Vb, Vc are weighted and summed to obtain a weighted vector; The weighted vector is batch normalized and activated by the ReLU function, and all weighted vectors are spliced into a data sequence (a`, b`, c`), and a second point cloud feature vector W is obtained.
5. The method of claim 4, wherein, A residual network is established, and the first sub-vector a, the second sub-vector b and the third sub-vector c are added to the data sequence (a`, b`, c`) to obtain the second point cloud feature vector W.
6. The method of claim 5, wherein, The point cloud fusion feature J and the down-sampling result are input into two levels of cascaded feature propagation modules to obtain a first point cloud feature vector H with a data format of 512×128×2048. Among them, 512 represents the number of training samples taken in the training set each time the first point cloud feature vector H is trained, 128 represents the dimension of the first point cloud feature vector H, and 2048 represents the number of point clouds of the first point cloud feature vector H.
7. The method of claim 6, wherein, The learnable parameter matrix S, T and V are all realized by a 1×1 convolution layer with an input dimension of 128 and an output dimension of 128. The information vectors Va, Vb and Vc respectively quantify the important features of the first sub-vector a, the second sub-vector b and the third sub-vector c. Based on the numerical value of the information vectors Va, Vb and Vc, the dimension value of the sub-vector with the larger numerical information vector is enhanced in the second point cloud feature vector W, and the dimension value of the sub-vector with the smaller numerical information vector tends to 0.
8. The method of claim 7, wherein, After the second point cloud feature vector W is reduced in dimension by the convolution layer L, the obtained dimension value is equal to the number of component categories, and the Log Softmax function is used to increase the probability value difference of the point cloud belonging to different component categories.
9. A point cloud data processing system based on multi-channel feature extraction, applying the point cloud data processing method according to any one of claims 1 to 8, characterized in that, It comprises: A data acquisition module is used to acquire 3D point cloud data, and the 3D point cloud data comprises N groups of point cloud sub-data with different sampling attributes; A multi-channel sampling module is used to set M KNN sampling channels according to the number of sampling attributes, each KNN sampling channel extracts features from a group of point cloud sub-data to obtain M sampling results, and M is less than or equal to N; A feature fusion module is used to fuse the M sampling results to obtain a point cloud fusion feature J; A feature propagation module is used to combine the M down-sampled down-sampling results to up-sample the point cloud fusion feature J to obtain a first point cloud feature vector H with a length of E; A feature enhancement module is used to enhance the features of the first point cloud feature vector H, establish a learnable parameter matrix with a dimension of E×E, multiply the first point cloud feature vector H by the learnable parameter matrix, weaken the redundant data through the weight parameters in the learnable parameter matrix, and obtain a second point cloud feature vector W. The data output module is configured to perform dimension reduction on the second point cloud feature vector W through the convolution layer L, obtain probability scores of the current point cloud under different sampling attributes through a Log Softmax function, obtain a loss value by using a Negative Log-Likelihood Loss loss function, update weight parameters through back propagation, and finally output a segmentation result of 3D point cloud data based on different sampling attributes under multi-channel. 10.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-9. The processor, when executing the computer program, implements the steps of the method of any one of claims 1 to 8.
Citation Information
Patent Citations
Feature extraction method and device based on point cloud segmentation and computer equipment
CN113706708A
Three-dimensional point cloud semantic segmentation method and apparatus, and device and medium
WO2022088676A1