Point cloud efficient classification and precise segmentation method based on high-frequency polyhedron representation

By constructing a high-frequency multihedral representation and self-attention mechanism, the problem of local and global feature capture in point cloud processing is solved, efficient classification and precise segmentation are achieved, computing efficiency and accuracy are improved, and it is suitable for applications such as autonomous driving and robot navigation.

CN120495716APending Publication Date: 2025-08-15YUNNAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510363850.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

When existing point cloud processing methods process irregularity, sparseness and uneven density point cloud data, it is difficult to efficiently capture local and global features while retaining high-frequency details, resulting in large consumption of computing resources and loss of information.

Method used

By constructing high-frequency polyhedral representation, neighborhood sampling and triangular face plate normal vector calculation are used, multi-frequency trigonometric function coding and self-attention mechanism are combined, global and local context features are dynamically integrated, and lightweight MLP network is designed to realize point cloud classification and semantic segmentation.

Benefits of technology

It significantly improves the description ability of local geometric structures, improves classification accuracy and segmentation accuracy, reduces calculation complexity, is suitable for practical applications such as autonomous driving and robot navigation, and has real-time processing capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495716A_ABST
    Figure CN120495716A_ABST
Patent Text Reader

Abstract

The invention discloses a point cloud efficient classification and precise segmentation method based on high-frequency polyhedron representation, and relates to the technical field of point cloud processing, and the method comprises the following steps: constructing local geometric features of a high-frequency polyhedron through neighborhood sampling and triangular patch normal vector calculation, and carrying out the aggregation to form umbrella-shaped polyhedron representation; a multi-frequency trigonometric function is adopted to carry out high-frequency coding on point coordinates, edge and corner information is enhanced, and the detail expression ability is improved; a local self-attention and self-positioning attention mechanism is designed, and global and local context features are dynamically fused in combination with spatial-semantic weights; an encoding-decoding network architecture is constructed, and point cloud classification and semantic segmentation tasks are respectively realized through progressive down-sampling and feature interpolation; according to the method, the triangular neighborhood and umbrella-shaped surface features are constructed, high-frequency trigonometric function coding is combined, and the description capability of a local geometric structure is remarkably enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of point cloud processing, and in particular relates to a method for efficient classification and precise segmentation of point clouds based on high-frequency polyhedron representation. Background Art

[0002] Currently, the processing and analysis of three-dimensional data are of great significance in the field of computer vision, especially point cloud data has become a research hotspot because it can retain the original geometric information; however, point cloud data has the defects of disorder, sparsity and uneven density; existing point cloud processing methods mainly include multi-view methods, voxel-based methods and point-based methods; multi-view methods learn features by projecting three-dimensional objects into multiple two-dimensional viewpoints, but there are problems of high computing resource consumption and information loss; point-based methods directly process point cloud data, but still have shortcomings in expressing local geometric structures and modeling global context; in addition, with the increase in the number of neural network layers, high-frequency geometric information is easily lost, affecting the performance of the model. How to efficiently capture local and global features in point cloud data while retaining high-frequency details has become a difficulty in current research. Therefore, we propose an efficient point cloud classification and accurate segmentation method based on high-frequency polyhedron representation. Summary of the Invention

[0003] To solve the above technical problems, the present invention is achieved through the following technical solutions:

[0004] The present invention is a method for efficient classification and accurate segmentation of point clouds based on high-frequency polyhedron representation, comprising the following steps:

[0005] Step S1: Construct local geometric features of high-frequency polyhedrons through neighborhood sampling and triangle patch normal vector calculation, and aggregate them into umbrella-shaped polyhedron representation;

[0006] Step S2: Use multi-frequency trigonometric functions to perform high-frequency encoding on point coordinates, enhance edge and corner information, and fuse them with polyhedral features to improve detail expression capabilities;

[0007] Step S3: Design local self-attention and self-localization attention mechanisms, combine spatial-semantic weights, and dynamically fuse global and local context features;

[0008] Step S4: Construct an encoding-decoding network architecture to achieve point cloud classification and semantic segmentation tasks respectively through progressive downsampling and feature interpolation;

[0009] Step S5: Propose dynamic neighborhood adjustment, high-frequency parameter adaptive optimization, and lightweight MLP design to balance computational efficiency and high-frequency detail expression capabilities, and ultimately output efficient features adapted to downstream tasks.

[0010] Furthermore, the step S1 includes the following steps:

[0011] Step S11: Input point cloud data is an unordered point set X = {x1, x2, ..., xN}∈RN×3, where each point xi contains three-dimensional coordinates (xi, yi, zi), xi is the i-th point in the point cloud, and for each target point xi, its two nearest neighbors are selected using the k-nearest neighbor algorithm to form a local triangle patch;

[0012] Step S12: For each triangle patch, calculate two edge vectors v1 and v2, and obtain the normal vector ni = (ai, bi, ci) through the cross product:

[0013] ni=v1×v2;

[0014] Calculate the surface position di of the patch, that is, the orthogonal projection distance of point xi to the plane, the formula is as follows:

[0015] di=aixi+biyi+cizi;

[0016] Concatenate the normal vector and the surface position into a triangle representation ti:

[0017] ti=Concat(ai,bi,ci,di);

[0018] Step S13: For each target point xi, expand the neighborhood to k nearest neighbor points, generate k triangular facets respectively, and use the multi-layer perceptron to normalize the coordinates of each facet. (with the centroid as reference) and triangles representing tij for feature extraction, and summing and aggregation along the third dimension, the formula is as follows:

[0019]

[0020] In the formula, k is the number of neighborhood points of each target point, is the normalized coordinate of the neighborhood point xj relative to the target point xi, τ is a lightweight multi-layer perceptron used for feature extraction, and ui is the umbrella polyhedron feature of the target point xi, which is obtained by aggregating the triangular patch features of the neighborhood points.

[0021] Furthermore, step S2 includes the following steps:

[0022] Step S21: Normalize the coordinates of each point Use multi-frequency trigonometric function for high frequency mapping, set the frequency parameter L=10, and the frequency range

[0023] For each coordinate component x, y, z, calculate the sine and cosine values of the corresponding frequency:

[0024] sini(x)=sin(freqi×x), cosi(x)=cos(freqi×x);

[0025] Concatenate the original coordinates with all frequency components to generate high-frequency encoding features:

[0026]

[0027] Where, High-frequency encoding features, including the original coordinates and the sine and cosine values of all frequency components, L is the number of frequency parameters, set to 10, freqi is the i-th frequency value, sin(x), cosi(x) is the sine and cosine functions of the i-th frequency applied to the coordinate component x;

[0028] Step S22: High frequency coding features Replace the normalized coordinates in step 1 and recalculate the umbrella polyhedron features. The formula is as follows:

[0029]

[0030] Where, It is the umbrella-shaped polyhedron feature after fusing high-frequency coding features.

[0031] Furthermore, step S3 includes the following steps:

[0032] Step S31: for each neighborhood point xj∈N(xi) of the target point xi, extract the feature difference Δfij=fj-fi and the coordinate difference Δxij=xj-xi;

[0033] The coordinate difference hf(Δxij) is enhanced by the high-frequency function and concatenated with the feature difference before inputting the aggregation function A:

[0034]

[0035] Where N(xi) is the set of neighborhood points of the target point xi, Δfij is the feature difference between the neighborhood point xj and the target point xi, Δxij is the coordinate difference between the neighborhood point xj and the target point xi, hf(Δxij) is the high-frequency encoding of the coordinate difference, and A is the aggregation function used to calculate the attention weight;

[0036] Calculate the attention weight wij and weighted aggregate neighborhood features:

[0037]

[0038] fLSA is the neighborhood feature aggregated by the local self-attention mechanism, wij is the attention weight of the neighborhood point xi, and gij is the median value of the attention weight calculation;

[0039] Step S32: Generate self-locating points xs∈RM×3 and weight the global point features by the learnable parameter matrix zs:

[0040]

[0041] Where xs is the self-locating point, which represents the key point of the global context, and zs is the learnable parameter matrix used to generate the self-locating point;

[0042] Calculate the spatial-semantic similarity weight ψs between the self-localized point and the original point cloud, and aggregate the global features fSPA through the cross-attention mechanism;

[0043] In step S33, the features generated by LSA and SPA are first processed through their respective transformation paths and fused using a weighted summation with a learnable parameter α:

[0044] fattn=α×fLSA+(1-α)×fSPA;

[0045] Where α is a learnable parameter used to balance the contribution of local features fLSA and global features fSPA.

[0046] Furthermore, the step S4 includes the following steps:

[0047] Step S41, feature encoder:

[0048] Input layer: Combines the original point cloud X with high-frequency polyhedron features Input shared MLP, mapped to the high-dimensional feature space F∈RN×C, F∈RN×C is the encoded high-dimensional feature, containing N points, and the feature dimension of each point is C;

[0049] Set abstraction: Select the seed point {xs} by sampling the farthest point and divide the local block around the seed;

[0050] For each block, K-nearest neighbor is used to extract neighborhood points, and local features are generated by combining high-frequency polyhedral features with the context fusion module;

[0051] Progressive downsampling: Repeating the set abstraction operation, gradually reducing the point cloud resolution while increasing the feature dimension;

[0052] Step S42, classification task: perform maximum pooling on the encoded global features, input them into the fully connected layer and output category probabilities;

[0053] Segmentation task: An encoder-decoder structure is used to gradually restore the point cloud resolution through feature interpolation, combined with jump connections to retain local details, and finally output semantic labels through point-by-point classification.

[0054] Furthermore, in step S5:

[0055] Dynamic domain adjustment: During training, the number of neighbors K is dynamically adjusted according to the point cloud density to balance the capture of details in sparse areas with the computational efficiency in dense areas.

[0056] High-frequency parameter adaptation: Optimize the trigonometric function frequency distribution through back-propagation to make the high-frequency encoding more adaptable to the geometric characteristics of the target dataset;

[0057] Lightweight MLP design: Using depthwise separable convolution and residual connection to reduce the computational overhead of MLPτ in step 1 while maintaining feature expression capability.

[0058] The present invention has the following beneficial effects:

[0059] 1. This paper significantly enhances the description of local geometric structures by constructing triangular neighborhoods and umbrella-shaped surface features, combined with high-frequency trigonometric function encoding. This method effectively improves the classification accuracy on the Model 1 Net40 and ScanObjectNN datasets, outperforming existing mainstream algorithms. It is particularly outstanding when processing complex shapes and real-world point clouds, providing more accurate 3D data understanding capabilities for practical applications such as autonomous driving and robot navigation.

[0060] 2. The present invention dynamically fuses spatial and semantic information through local self-attention and self-localization attention mechanisms, combined with high-frequency enhancement functions. Experiments on the S3DIS dataset show that its mIoU reaches 68.5% and OA is 89.6%, significantly outperforming benchmark methods such as PointNet++. This fusion strategy not only improves the model's ability to understand complex indoor scenes, but also reduces computational complexity, enabling it to achieve faster inference speed while maintaining high accuracy, making it suitable for application scenarios with high real-time requirements.

[0061] 3. The present invention has significant advantages in computing efficiency and is suitable for actual deployment. By optimizing the network structure and algorithm design, the model parameters are only 1.5M, the FLOPs is 1.8G, the inference time is as low as 66.7 milliseconds, and the throughput reaches 1920 instances / second. Compared with models such as PointMLP, it significantly reduces computing resource consumption while maintaining high classification and segmentation performance. This high efficiency is due to the lightweight design of high-frequency polyhedron representation and the optimization of the context fusion module, which enables the model to achieve real-time processing with limited hardware resources, providing a feasible technical solution for practical applications in smart cities, medical image analysis and other fields.

[0062] Of course, any product implementing the present invention does not necessarily need to achieve all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0064] Figure 1 Schematic diagram of the process of the method for efficient classification and accurate segmentation of point clouds based on high-frequency polyhedron representation of the present invention;

[0065] Figure 2 This is a network structure diagram of the point cloud efficient classification and precise segmentation method based on high-frequency polyhedron representation of the present invention. DETAILED DESCRIPTION

[0066] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making any creative efforts shall fall within the scope of protection of the present invention.

[0067] See also Figure 1-2 As shown, the present invention is a method for efficient classification and accurate segmentation of point clouds based on high-frequency polyhedron representation, comprising the following steps:

[0068] Step S1: Construct local geometric features of high-frequency polyhedrons through neighborhood sampling and triangle patch normal vector calculation, and aggregate them into umbrella-shaped polyhedron representation;

[0069] Step S2: Use multi-frequency trigonometric functions to perform high-frequency encoding on point coordinates, enhance edge and corner information, and fuse them with polyhedral features to improve detail expression capabilities;

[0070] Step S3: Design local self-attention and self-localization attention mechanisms, combine spatial-semantic weights, and dynamically fuse global and local context features;

[0071] Step S4: Construct an encoding-decoding network architecture to achieve point cloud classification and semantic segmentation tasks respectively through progressive downsampling and feature interpolation;

[0072] Step S5: Propose dynamic neighborhood adjustment, high-frequency parameter adaptive optimization, and lightweight MLP design to balance computational efficiency and high-frequency detail expression capabilities, and ultimately output efficient features adapted to downstream tasks.

[0073] Step S1 includes the following steps:

[0074] Step S11: Input point cloud data is an unordered point set X = {x1, x2, ..., xN}∈RN×3, where each point xi contains three-dimensional coordinates (xi, yi, zi), xi is the i-th point in the point cloud, and for each target point xi, its two nearest neighbors are selected using the k-nearest neighbor algorithm to form a local triangle patch;

[0075] Step S12: For each triangle patch, calculate two edge vectors v1 and v2, and obtain the normal vector ni = (ai, bi, ci) through the cross product:

[0076] ni=v1×v2;

[0077] Calculate the surface position di of the patch, that is, the orthogonal projection distance of point xi to the plane, the formula is as follows:

[0078] di=aixi+biyi+cizi;

[0079] Concatenate the normal vector and the surface position into a triangle representation ti:

[0080] ti=Concat(ai,bi,ci,di);

[0081] Step S13: For each target point xi, expand the neighborhood to k nearest neighbor points, generate k triangular facets respectively, and use the multi-layer perceptron to normalize the coordinates of each facet. (with the centroid as reference) and triangles representing tij for feature extraction, and summing and aggregation along the third dimension, the formula is as follows:

[0082]

[0083] In the formula, k is the number of neighborhood points of each target point, is the normalized coordinate of the neighborhood point xj relative to the target point xi, τ is a lightweight multi-layer perceptron used for feature extraction, and ui is the umbrella polyhedron feature of the target point xi, which is obtained by aggregating the triangular patch features of the neighborhood points.

[0084] Step S2 includes the following steps:

[0085] Step S21: Normalize the coordinates of each point Use multi-frequency trigonometric function for high frequency mapping, set the frequency parameter L=10, and the frequency range

[0086] For each coordinate component x, y, z, calculate the sine and cosine values of the corresponding frequency:

[0087] sini(x)=sin(freqi×x), cosi(x)=cos(freqi×x);

[0088] Concatenate the original coordinates with all frequency components to generate high-frequency encoding features:

[0089]

[0090] Where, High-frequency encoding features, including the original coordinates and the sine and cosine values of all frequency components, L is the number of frequency parameters, set to 10, freqi is the i-th frequency value, sin(x), cosi(x) is the sine and cosine functions of the i-th frequency applied to the coordinate component x;

[0091] Step S22: High frequency coding features Replace the normalized coordinates in step 1 and recalculate the umbrella polyhedron features. The formula is as follows:

[0092]

[0093] Where, It is the umbrella-shaped polyhedron feature after fusing high-frequency coding features.

[0094] Step S3 includes the following steps:

[0095] Step S31: for each neighborhood point xj∈N(xi) of the target point xi, extract the feature difference Δfij=fj-fi and the coordinate difference Δxij=xj-xi;

[0096] The coordinate difference hf(Δxij) is enhanced by the high-frequency function and concatenated with the feature difference before inputting the aggregation function A:

[0097]

[0098] Where N(xi) is the set of neighborhood points of the target point xi, Δfij is the feature difference between the neighborhood point xj and the target point xi, Δxij is the coordinate difference between the neighborhood point xj and the target point xi, hf(Δxij) is the high-frequency encoding of the coordinate difference, and A is the aggregation function used to calculate the attention weight;

[0099] Calculate the attention weight wij and weighted aggregate neighborhood features:

[0100]

[0101] fLSA is the neighborhood feature aggregated by the local self-attention mechanism, wij is the attention weight of the neighborhood point xi, and gij is the median value of the attention weight calculation;

[0102] Step S32: Generate self-locating points xs∈RM×3 and weight the global point features by the learnable parameter matrix zs:

[0103]

[0104] Where xs is the self-locating point, which represents the key point of the global context, and zs is the learnable parameter matrix used to generate the self-locating point;

[0105] Calculate the spatial-semantic similarity weight ψs between the self-localized point and the original point cloud, and aggregate the global features fSPA through the cross-attention mechanism;

[0106] In step S33, the features generated by LSA and SPA are first processed through their respective transformation paths and fused using a weighted summation with a learnable parameter α:

[0107] fattn=α×fLSA+(1-α)×fSPA;

[0108] Where α is a learnable parameter used to balance the contribution of local features fLSA and global features fSPA.

[0109] Step S4 includes the following steps:

[0110] Step S41, feature encoder:

[0111] Input layer: Combines the original point cloud X with high-frequency polyhedron features Input shared MLP, mapped to the high-dimensional feature space F∈RN×C, F∈RN×C is the encoded high-dimensional feature, containing N points, and the feature dimension of each point is C;

[0112] Set abstraction: Select the seed point {xs} by sampling the farthest point and divide the local block around the seed;

[0113] For each block, K-nearest neighbor is used to extract neighborhood points, and local features are generated by combining high-frequency polyhedral features with the context fusion module;

[0114] Progressive downsampling: Repeating the set abstraction operation, gradually reducing the point cloud resolution while increasing the feature dimension;

[0115] Step S42, classification task: perform maximum pooling on the encoded global features, input them into the fully connected layer and output category probabilities;

[0116] Segmentation task: An encoder-decoder structure is used to gradually restore the point cloud resolution through feature interpolation, combined with jump connections to retain local details, and finally output semantic labels through point-by-point classification.

[0117] In step S5:

[0118] Dynamic domain adjustment: During training, the number of neighbors K is dynamically adjusted according to the point cloud density to balance the capture of details in sparse areas with the computational efficiency in dense areas.

[0119] High-frequency parameter adaptation: Optimize the trigonometric function frequency distribution through back-propagation to make the high-frequency encoding more adaptable to the geometric characteristics of the target dataset;

[0120] Lightweight MLP design: Using depthwise separable convolution and residual connection to reduce the computational overhead of MLPτ in step 1 while maintaining feature expression capability.

[0121] A specific application of this embodiment is:

[0122] Example 1, experimental environment:

[0123] The experiments in this example were conducted on a machine equipped with an NVIDIA A6000 GPU and a 2-core Intel Xeon 2.50GHz CPU. The datasets used were ScanObjectNN and ModelNet40, which were used to verify the classification performance of real scenes and synthetic data, respectively.

[0124] ScanObjectNN: Contains 15,000 real-world point cloud objects from 15 categories, with occlusions and background interference;

[0125] ModelNet40: contains 12,311 CAD models from 40 categories, of which 9,843 are used for training and 2,468 are used for testing;

[0126] Method implementation steps:

[0127] 1. For each point in the input point cloud, query its two nearest neighbors to construct a triangular patch, calculate the patch normal vector and surface position, and form a local geometric feature representation;

[0128] Use high-frequency trigonometric functions (frequency range 20, 210, L = 10) to encode point coordinates to enhance the ability to capture high-frequency geometric details;

[0129] The high-frequency encoded coordinates are fused with the polyhedral features to generate the final point cloud representation;

[0130] 2. Local self-attention (LSA) and self-localization attention (SPA) are used to aggregate local and global features respectively;

[0131] Dynamically adjust the fusion ratio of local and global features through the learnable weight parameter α;

[0132] 3. Optimizer: AdamW, initial learning rate 0.001, weight decay 0.05;

[0133] Batch size: 32, training cycle: 250 rounds for ScanObjectNN and 600 rounds for Model l Net40;

[0134] Experimental results:

[0135] 1. ScanObjectNN dataset:

[0136] The overall accuracy (OA) reached 88.7%, and the average category accuracy (mAcc) reached 86.9%, an improvement of 10.8% over Poi ntNet++;

[0137] The accuracy rate exceeds 95% in categories such as "chair" and "toilet";

[0138] 2. Mode l Net40 dataset

[0139] OA is 93.7% and mAc is 91.5%, which is 4.5% higher than the baseline method;

[0140] Achieve 100% accuracy in categories such as "airplane" and "bed";

[0141] Computational efficiency:

[0142] Parameter count: 1.5M, FLOPs: 1.8G;

[0143] Inference speed: 66.7 milliseconds / sample, throughput 1920 instances / second, outperforming the comparison method.

[0144] Example 2

[0145] Experimental environment and dataset:

[0146] We use the S3D IS dataset, which contains RGB-D scans of six indoor scenes annotated with 13 semantic labels. The experiment is divided into Area-5 single-area testing and 6-fold cross-validation.

[0147] Method implementation steps

[0148] 1. Voxel downsampling (voxel size 0.04m), normalized point coordinates;

[0149] 2. The encoder uses a high-frequency polyhedron representation module, and the decoder combines skip connections and an upsampling loss function: cross entropy loss;

[0150] 3. Optimizer: AdamW, weight decay 10^-4, learning rate cosine decay;

[0151] Batch size: 32;

[0152] Experimental results:

[0153] Area-5 test: mIoU is 68.5%, OA is 89.6%;

[0154] 6-Fo ld cross-validation: mI oU was 72.4%, OA was 89.3%;

[0155] Compared with Poi ntNet++, it improves mIoU by 15.0%, proving the effectiveness of high-frequency polyhedron representation for complex scene segmentation;

[0156] Visual analysis:

[0157] High-frequency polyhedron features significantly enhance the ability to focus on edge and corner areas.

[0158] Example 3:

[0159] Comparison method:

[0160] Including mainstream models such as Poi ntNet, Poi ntNet++, Poi ntMLP;

[0161] Performance indicators:

[0162] Parameter count: 1.5M (only 11.4% of PointMLP);

[0163] Inference time: 66.7 milliseconds, 13.3 milliseconds faster than RepSurf-U;

[0164] Throughput: 1920 instances / second, suitable for real-time applications;

[0165] in conclusion:

[0166] This method significantly improves classification and segmentation accuracy while maintaining low computational resource consumption, making it suitable for practical scenarios with limited resources.

[0167] Throughout this specification, references to terms such as "one embodiment," "example," or "specific example" indicate that the specific features, structures, materials, or characteristics described in conjunction with that embodiment or example are included in at least one embodiment or example of the present invention. In this specification, schematic representations of these terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in any one or more embodiments or examples.

[0168] The preferred embodiments of the present invention disclosed above are intended only to help illustrate the present invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the present invention to the specific embodiments described. Obviously, many modifications and variations are possible based on the content of this specification. These embodiments are selected and described in detail in this specification to better explain the principles and practical applications of the present invention, thereby enabling those skilled in the art to better understand and utilize the present invention. The present invention is limited only by the claims and their full scope and equivalents.

Claims

1. An efficient point cloud classification and accurate segmentation method based on high-frequency polyhedron representation, characterized by: The following steps are involved: Step S1: Construct local geometric features of high-frequency polyhedrons through neighborhood sampling and triangle patch normal vector calculation, and aggregate them into umbrella-shaped polyhedron representation; Step S2: Use multi-frequency trigonometric functions to perform high-frequency encoding on point coordinates, enhance edge and corner information, and fuse them with polyhedral features to improve detail expression capabilities; Step S3: Design local self-attention and self-localization attention mechanisms, combine spatial-semantic weights, and dynamically fuse global and local context features; Step S4: Construct an encoding-decoding network architecture to achieve point cloud classification and semantic segmentation tasks respectively through progressive downsampling and feature interpolation; Step S5: Propose dynamic neighborhood adjustment, high-frequency parameter adaptive optimization, and lightweight MLP design to balance computational efficiency and high-frequency detail expression capabilities, and ultimately output efficient features adapted to downstream tasks.

2. The method for efficient classification and accurate segmentation of point clouds based on high-frequency polyhedron representation according to claim 1, characterized in that: The step S1 includes the following steps: Step S11: Input point cloud data is an unordered point set X = {x1, x2, ..., xN}∈RN×3, where each point xi contains three-dimensional coordinates (xi, yi, zi), xi is the i-th point in the point cloud, and for each target point xi, its two nearest neighbors are selected using the k-nearest neighbor algorithm to form a local triangle patch; Step S12: For each triangle patch, calculate two edge vectors v1 and v2, and obtain the normal vector ni = (ai, bi, ci) through the cross product: ni=v1×v2; Calculate the surface position di of the patch, that is, the orthogonal projection distance of point xi to the plane, the formula is as follows: di=aixi+biyi+cizi; Concatenate the normal vector and the surface position into a triangle representation ti: ti=Concat(ai,bi,ci,di); Step S13: For each target point xi, expand the neighborhood to k nearest neighbor points, generate k triangular facets respectively, and use the multi-layer perceptron to normalize the coordinates of each facet. (with the centroid as reference) and triangles representing tij for feature extraction, and summing and aggregation along the third dimension, the formula is as follows: In the formula, k is the number of neighborhood points of each target point, is the normalized coordinate of the neighborhood point xj relative to the target point xi, τ is a lightweight multi-layer perceptron used for feature extraction, and ui is the umbrella polyhedron feature of the target point xi, which is obtained by aggregating the triangular patch features of the neighborhood points.

3. The method for efficient classification and accurate segmentation of point clouds based on high-frequency polyhedron representation according to claim 1, characterized in that: The step S2 includes the following steps: Step S21: Normalize the coordinates of each point Use multi-frequency trigonometric function for high frequency mapping, set the frequency parameter L=10, and the frequency range For each coordinate component x, y, z, calculate the sine and cosine values of the corresponding frequency: sini(x)=sin(freqi×x), cosi(x)=cos(freqi×x); Concatenate the original coordinates with all frequency components to generate high-frequency encoding features: Where, High-frequency encoding features, including the original coordinates and the sine and cosine values of all frequency components, L is the number of frequency parameters, set to 10, freqi is the i-th frequency value, sin(x), cosi(x) is the sine and cosine functions of the i-th frequency applied to the coordinate component x; Step S22: High frequency coding features Replace the normalized coordinates in step 1 and recalculate the umbrella polyhedron features. The formula is as follows: Where, It is the umbrella-shaped polyhedron feature after fusing high-frequency coding features.

4. The method for efficient classification and accurate segmentation of point clouds based on high-frequency polyhedron representation according to claim 1, characterized in that: The step S3 includes the following steps: Step S31: for each neighborhood point xj∈N(xi) of the target point xi, extract the feature difference Δfij=fj-fi and the coordinate difference Δxij=xj-xi; The coordinate difference hf(Δxij) is enhanced by the high-frequency function and concatenated with the feature difference before inputting the aggregation function A: Where N(xi) is the set of neighborhood points of the target point xi, Δfij is the feature difference between the neighborhood point xj and the target point xi, Δxij is the coordinate difference between the neighborhood point xj and the target point xi, hf(Δxij) is the high-frequency encoding of the coordinate difference, and A is the aggregation function used to calculate the attention weight; Calculate the attention weight wij and weighted aggregate neighborhood features: fLSA is the neighborhood feature aggregated by the local self-attention mechanism, wij is the attention weight of the neighborhood point xi, and gij is the median value of the attention weight calculation; Step S32: Generate self-locating points xs∈RM×3 and weight the global point features by the learnable parameter matrix zs: Where xs is the self-locating point, which represents the key point of the global context, and zs is the learnable parameter matrix used to generate the self-locating point; Calculate the spatial-semantic similarity weight ψs between the self-localized point and the original point cloud, and aggregate the global features fSPA through the cross-attention mechanism; Step S33: Fusion is performed using a weighted summation method using a learnable parameter α: fattn=α×fLSA+(1-α)×fSPA; Where α is a learnable parameter used to balance the contribution of local features fLSA and global features fSPA.

5. The method for efficient classification and accurate segmentation of point clouds based on high-frequency polyhedron representation according to claim 1, characterized in that: The step S4 includes the following steps: Step S41, feature encoder: Input layer: Combines the original point cloud X with high-frequency polyhedron features Input shared MLP, mapped to the high-dimensional feature space F∈RN×C, F∈RN×C is the encoded high-dimensional feature, containing N points, and the feature dimension of each point is C; Set abstraction: Select the seed point {xs} by sampling the farthest point and divide the local block around the seed; For each block, K-nearest neighbor is used to extract neighborhood points, and local features are generated by combining high-frequency polyhedral features with the context fusion module; Progressive downsampling: Repeating the set abstraction operation, gradually reducing the point cloud resolution while increasing the feature dimension; Step S42, classification task: perform maximum pooling on the encoded global features, input them into the fully connected layer and output category probabilities; Segmentation task: An encoder-decoder structure is used to gradually restore the point cloud resolution through feature interpolation, combined with jump connections to retain local details, and finally output semantic labels through point-by-point classification.

6. The method for efficient classification and accurate segmentation of point clouds based on high-frequency polyhedron representation according to claim 1, characterized in that: In the step S5: Dynamic domain adjustment: During training, the number of neighbors K is dynamically adjusted according to the point cloud density to balance the capture of details in sparse areas with the computational efficiency in dense areas. High-frequency parameter adaptation: Optimize the trigonometric function frequency distribution through back-propagation to make the high-frequency encoding more adaptable to the geometric characteristics of the target dataset; Lightweight MLP design: Using depthwise separable convolution and residual connection to reduce the computational overhead of MLPτ in step 1 while maintaining feature expression capability.