Crop point cloud segmentation self-supervision method based on point-voxel double branches

By employing the point-voxel dual-branch Crop-PVCNet method, combined with a point-voxel cross-self-attention encoder and an improved SimSiam network, the problem of extracting detailed features in non-rigid crop point cloud segmentation is solved, achieving efficient and accurate crop point cloud segmentation and improving the efficiency of plant phenotypic analysis.

CN121329982APending Publication Date: 2026-01-13HENAN NORMAL UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202511486802.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-10-17
Publication Date
2026-01-13

AI Technical Summary

Technical Problem

Existing self-supervised learning methods struggle to effectively extract detailed features when dealing with non-rigid crops, and traditional point cloud segmentation methods are computationally expensive when dealing with sparsity and GPU memory limitations, leading to information loss or increased computation.

Method used

We employ the Crop-PVCNet method based on point-voxel dual branches, combining a point-voxel cross-self-attention encoder and an improved SimSiam network. Through data augmentation and feature fusion, we extract local details and global semantic information from crop point clouds, and utilize a cross-cube multi-head attention module to reduce computational complexity.

Benefits of technology

It achieves more accurate crop point cloud segmentation, significantly improves the quality of point cloud organ segmentation, reduces computational complexity, and provides an efficient plant phenotypic analysis tool.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121329982A_ABST
    Figure CN121329982A_ABST
Patent Text Reader

Abstract

The invention provides a crop point cloud segmentation self-supervision method based on point-voxel double branches, and the method comprises the following steps: enabling obtained crop point clouds to form a data set, dividing the data set into a training set, a test set and a verification set, carrying out the data enhancement of the crop point clouds in the training set, and obtaining an extended training set; a Crop-PVCNet segmentation network model is constructed, the Crop-PVCNet segmentation network model comprises an online branch and a target branch, and both the online branch and the target branch are provided with point-voxel cross self-attention encoders; a Crop-PVCNet segmentation network model is trained by using the extended training set, and a trained Crop-PVCNet network model is obtained; and inputting the crop point cloud in the test set into the trained Crop-PVCNet network model for testing to obtain a segmented crop point cloud file. According to the method, the structure and details of the crop point cloud can be accurately captured, accurate organ segmentation is realized, and an efficient labeling method is provided for plant phenotype analysis and research.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of artificial intelligence and computer vision, and particularly relates to a crop point cloud segmentation self-supervised method. BACKGROUND

[0002] Plant phenotype refers to the physiological and biochemical characteristics and morphological structure parameters exhibited by plants under the interaction of specific genotypes and environmental conditions, and is an important basis for plant growth state analysis in modern agricultural production. Real-time monitoring and analysis of phenotypic characteristics such as leaf shape, color and size are of great significance for early diagnosis of diseases and pests, crop yield estimation and excellent variety selection. Traditional phenotype collection methods usually rely on manual measurement, which is low in operation efficiency, high in cost, and easy to introduce subjective errors, and some measurement methods are also destructive. Although deep learning methods based on two-dimensional images have been widely used in plant organ segmentation tasks, they still have significant limitations in organ overlap, occlusion and spatial structure differentiation due to the lack of three-dimensional spatial information.

[0003] In recent years, with the development of laser scanning technology, point cloud segmentation technology has become an important tool in plant phenotype research. Some studies have proposed various traditional methods, most of which are mainly based on edge detection, region growing and clustering. These methods require a lot of human involvement and face challenges in handling large-scale 3D datasets. The development of deep learning has promoted plant point cloud segmentation to become a key research area, but the success of these models largely depends on a large number of finely annotated labels. One of the main methods for point cloud segmentation is the PointNet network, however, most existing deep learning methods rely on manually annotated labels, and the acquisition of these labels is both time-consuming and expensive. Given the high cost of annotating 3D point clouds, self-supervised learning has become an important research direction.

[0004] Self-supervised learning based methods leverage intrinsic signals in unlabelled data to learn representations, reducing the reliance on labelled data. Point clouds are sparse and irregular 3D representations, requiring aggregation of information from neighbouring points to extract local features. However, this requires computationally expensive operations such as K-Nearest Neighbour (KNN) search due to the lack of spatial consistency in storage. In contrast, voxel-based methods map points to a regular grid stored contiguously in memory according to spatial location. This design allows efficient memory access in a sequential manner during processing. However, the sparsity of point clouds and the limitation of GPU memory often restricts the grid resolution, losing details. Point-Voxel CNN for Efficient 3D Deep Learning (PVCNN) integrates both representations by leveraging point-based representation to reduce memory and using voxel convolution to improve data locality, achieving efficient feature extraction. However, PVCNN performs poorly in handling non-rigid objects due to the dynamic deformation nature of plant organs.

[0005] The invention patent with application number 202411900711.4 discloses a farmland feature semantic segmentation method based on laser radar and SegNet. The method uses a vehicle-mounted or unmanned aerial vehicle-mounted laser radar device to scan the farmland, obtains raw point cloud data, and pre-processes the raw point cloud data to obtain point cloud data of the target farmland. The point cloud data is voxelized to obtain a training data set. The segmentation model is trained based on the training data set, and the target farmland is segmented and predicted based on the segmentation model. According to the training result of the segmentation model and the segmentation result of the semantic segmentation prediction, the final semantic segmentation result of the target farmland is obtained, and the semantic segmentation result is visualized and displayed. The above invention realizes end-to-end farmland feature semantic segmentation and provides intuitive visual display. The effect of laser radar point cloud in complex farmland scene semantic segmentation is remarkable. However, when the above invention discretizes continuous three-dimensional point cloud into a fixed resolution voxel grid, it inevitably leads to the loss of subtle geometric features. Setting the resolution too low will cause serious information loss, while setting it too high will cause a sharp increase in computational complexity and memory overhead. This contradiction constitutes a key challenge in the field of three-dimensional point cloud processing that needs to be addressed. SUMMARY

[0006] In view of the technical problem that the existing self-supervised learning method cannot well extract detailed features when processing non-rigid crops, the present application proposes a crop point cloud segmentation self-supervised method based on point-voxel dual branches (Crop-PVCNet) to achieve more accurate and clear segmentation effect.

[0007] To achieve the above objectives, the technical solution of this invention is as follows: a self-supervised method for crop point cloud segmentation based on point-voxel bi-branching, comprising the following steps:

[0008] Step 1: Compile the obtained crop point clouds into a dataset and divide the dataset into a training set, a test set, and a validation set. Perform data augmentation on the crop point clouds in the training set to obtain an expanded training set.

[0009] Step 2: Construct the Crop-PVCNet segmentation network model, including online branches and target branches, with point-voxel cross self-attention encoders on both the online and target branches;

[0010] Step 3: Train the Crop-PVCNet segmentation network model using the expanded training set to obtain the trained Crop-PVCNet network model;

[0011] Step 4: Input the crop point cloud data from the test set into the trained Crop-PVCNet network model for testing, and obtain the segmented crop point cloud file.

[0012] Preferably, the leaf sheath method is used to label the points in the crop point cloud of the dataset, and different numerical labels are assigned to the stems and leaves of different plants, with the values ​​increasing sequentially; points marked as ground are removed from each point cloud file.

[0013] The data augmentation includes translation, scaling, rotation, jittering, and filtering of point clouds, with filtering including cropping and dicing.

[0014] Preferably, the online branch and the target branch are dual branches with shared weights. The input of the point-voxel cross-attention encoder is connected to the enhancement channel, and the output of the point-voxel cross-attention encoder is connected to the projection network. The output of the projection network of the online branch is connected to the prediction network, and the output of the projection network of the target branch is connected to the SimSiam network. The loss function L is obtained by comparing the similarity loss of the feature representations output by the online branch and the target branch.

[0015] Preferably, the enhancement channel enhances the input point cloud sample through rotation, coordinate jitter, translation, and point dropping operations. Process the data to generate an enhanced point cloud view. ;in, This indicates the number of points in a point cloud sample. Additional feature dimensions besides coordinates;

[0016] The point-voxel cross-attention encoder enhances the input point cloud view. Mapping to high-dimensional semantic features ; wherein D is the feature dimension of the high-dimensional semantic feature;

[0017] The projection network performs a nonlinear transformation on the features extracted by the point-voxel cross self-attention encoder and maps them to a high-dimensional feature space; the feature representation output by the projection network of the online branch is further input into the prediction network, which maps the feature representation into a prediction representation by batch normalization and LeakyReLU activation function ; in the target branch, the output of the projection network is kept stable by the SimSiam network through the stop gradient operation.

[0018] Preferably, the point-voxel cross self-attention encoder comprises three PVE layers connected in sequence, each PVE layer extracting complementary features through a parallel point branch and a voxel branch; the point branch is provided with a point self-attention mechanism, and the voxel branch is provided with voxelization, VTB module and inverse voxelization connected in sequence;

[0019] The high-dimensional semantic feature or the point cloud enhanced view is input into the PVE layer of the first layer, and the global geometric feature is obtained on the point branch through the point self-attention mechanism ; the local geometric feature is obtained on the voxel branch through voxelization, VTB module and inverse voxelization , and finally the global geometric feature and the local geometric feature are fused and output as the high-dimensional semantic feature as the output of the point-voxel cross self-attention encoder or the input of the next PVE layer, wherein , D is the feature dimension of the high-dimensional semantic feature, and D' is the feature dimension after projection.

[0020] Preferably, the point self-attention mechanism projects the input high-dimensional semantic feature or the point cloud enhanced view into a high-dimensional vector space through a multilayer perceptron to obtain a high-dimensional feature , and then inputs the high-dimensional feature into three linear transformation layers , and in sequence to generate a query matrix , a key matrix and a value matrix ; taking the query matrix and the key matrix K as inputs, the query matrix is calculated with K​T The matrix multiplication is performed and then normalized using the Softmax function to obtain the attention weight matrix. The attention weight matrix and the value matrix are then compared. Multiply to output the attention matrix. Attention matrix and high-dimensional features The data is then concatenated, and the global output features of the point branches are obtained through a multilayer perceptron. ;

[0021] The voxelization receives from the first High-dimensional semantic features of layer input Alternatively, a point cloud augmented view V is generated, and normalization centers the coordinates to the centroid and scales them to a [0, 1] unit cube, transforming the point cloud into a unified coordinate system; the normalized space is discretized through voxel mapping. The grid; each point is mapped to a corresponding grid index along the x, y, and z axes based on its normalized coordinates, with the index range being [0, ..., ...]. The average feature vector of all points in each occupied voxel is calculated using feature averaging, and then aggregated to generate voxel features. ;in, For target voxel resolution, It is the feature dimension;

[0022] The devoxometry uses trilinear interpolation to apply voxel features to the output of the VTB module. Processing: Based on the normalized coordinates of each point in the voxel grid, the voxel features are... Accurately map to the corresponding point features.

[0023] Preferably, the VTB module includes a first-layer normalization, a CCMHA module, a first residual connection, a second-layer normalization, a multilayer perceptron, and a second residual connection connected in sequence, wherein the first residual connection is connected to the input voxel features. The input of the second residual connection is connected to the output of the first residual connection;

[0024] Input voxel features After the first layer of normalization, the data is processed for layer normalization and then input into the CCMHA module; the output of the CCMHA module is compared with the input voxel features. Feature fusion is performed through the first residual connection to obtain features. ;feature After undergoing layer normalization in the second layer, the data is fed into a multilayer perceptron for nonlinear transformation. The output is then connected to the features via a second residual connection. Fusion, outputting enhanced voxel features .

[0025] Preferably, the CCMHA module is a multi-head self-attention mechanism based on three-dimensional cuboid partitioning. For each query voxel in the input, the CCMHA module constructs three mutually perpendicular cuboid regions in its neighborhood, which are perpendicular to the X, Y, and Z coordinate axes, respectively. Local attention is calculated in each direction through each cuboid region. In the multi-head self-attention mechanism, the total number of attention heads K is set to be a multiple of 3, and K / 3 attention heads are assigned to each direction to process the local features in the corresponding direction.

[0026] Preferably, the loss function L consists of two symmetric terms, each of which measures the similarity between the predicted output of the online branch for one enhanced view-level feature and the output of the target branch after applying the stopping gradient operation to another enhanced view-level feature.

[0027] The loss function L, which alternately predicts the target across different views, is represented as follows:

[0028]

[0029] in, This represents the operation function of a prediction network that operates only on the online branches. This indicates stopping the gradient operation, and D(·) represents the feature similarity measure, which is calculated using cosine similarity.

[0030] Preferably, the method for calculating the local attention of the cuboid region perpendicular to the X-axis is as follows: The features after layer normalization... Divide the x-axis into n non-overlapping cuboid units. The height of each cuboid unit is cw. For each partitioned cuboid unit Each is passed through three independent linear transformation layers. , , Generate the query matrix for the nth attention head respectively. Key matrix AND-value matrix Each attention head computes the query matrix within its corresponding cuboid cell. transpose of the key matrix Matrix multiplication and normalization using the Softmax function, and the sum of the matrix values. Multiplying them together yields the local attention vector features. The local attention vector features of all K cuboids in the same direction are merged to form a local self-attention representation in that direction. Perform the same operation on the Y-axis and Z-axis respectively to obtain the local self-attention representation. and The local self-attention representations in the X, Y, and Z axes are concatenated and then projected through a linear projection layer. Feature fusion is performed to output comprehensive multi-directional local self-attention features. , This represents the operation functions of the CCMHA module;

[0031] The feature fusion module is used to integrate the output features of the first PVE layer. and the output characteristics of the third PVE layer The fusion is performed to obtain the fusion features. The output of the point-voxel cross-attention encoder, and the fused features

[0032]

[0033]

[0034] in, This indicates a feature concatenation operation. These are the operation functions of a multilayer perceptron. This indicates a max pooling operation. It is a repetitive operation. It is a global context feature and has the same feature dimension as before the max pooling operation.

[0035] The attention weight matrix is ​​as follows:

[0036]

[0037] The global geometric features:

[0038]

[0039] in, Here, is the attention function, and Concat() represents the matrix concatenation operation along the feature dimension. Represents the operation functions of a multilayer perceptron;

[0040] The multilayer perceptron is composed of batch normalization and ReLU activation function;

[0041] After training, the weights of the point-voxel cross-self-attention encoder are retained; the crop point cloud from the test set is input into the augmentation channel, and the coordinates are normalized. The point-voxel cross-self-attention encoder is then used to extract point cloud features to obtain features that fuse local details and global context. ,feature Input a fully connected layer, apply the argmax function to the output of the fully connected layer to obtain the final semantic category label for each point; write the original coordinate information and predicted semantic category label information of the crop point cloud into a txt file and save it to obtain a 3D point cloud file after semantic segmentation.

[0042] Compared with existing technologies, the advantages of this invention are as follows: This invention proposes a feature encoder (PVCE) based on a dual-branch Transformer for crop point cloud segmentation. The feature encoder integrates point cloud and voxel representations, enabling simultaneous capture of local detail features and global semantic information at different scales. This effectively captures the complex spatial relationships and diversity in crop point clouds, thus significantly improving the quality of point cloud organ segmentation. Since traditional single attention modules struggle to balance global style and local detail, and voxel-based techniques suffer from information loss and high memory consumption, this invention introduces a cross-cube multi-head attention (CCMHA) module into the voxel branch of the network. This achieves selective enhancement of key crop organ features, effectively extracting fine features and geometric information while reducing computational complexity. Furthermore, this invention utilizes an improved SimSiam network to prevent feature loss through a stopping gradient (sg) operation, thereby reducing computational costs. The overall solution of this invention can accurately capture the structure and details of crop point clouds, achieving accurate organ segmentation and providing an efficient annotation method for plant phenotypic analysis research. Attached Figure Description

[0043] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0044] Figure 1 This is a flowchart of the present invention.

[0045] Figure 2 The above are examples of point cloud images after centralized enhancement of the data set in this invention, where (a) is the original image, (b) is translation, (c) is scaling, (d) is rotation, (e) is jitter, (f) is cropping, and (g) is slicing.

[0046] Figure 3 This is a network framework diagram of the Crop-PVCNet segmentation network model of the present invention.

[0047] Figure 4 The diagram shows the network structure of the PVCE module of the present invention; where (a) is the PVE module and (b) is the VTB module.

[0048] Figure 5 This is a data processing flowchart for the PVE module branching of the present invention.

[0049] Figure 6 This is a network framework diagram of the CCMHA module of the present invention.

[0050] Figure 7 This is a network framework diagram of the feature fusion module of the present invention.

[0051] Figure 8 This is a comparison chart of the segmentation results of this invention on the Pheno4D dataset and the segmentation results of other methods.

[0052] Figure 9 This is a comparison chart of the segmentation results of this invention on the Crops3D dataset with the segmentation results of other methods.

[0053] Figure 10 This is a comparison chart of the segmentation results of this invention on the ShapeNet dataset and the segmentation results of other methods. Detailed Implementation

[0054] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0055] like Figure 1 As shown, a self-supervised method for crop point cloud segmentation based on point-voxel bi-branch (Crop-PVCNet) includes the following steps:

[0056] Step 1: Compile the obtained crop point clouds into a dataset and divide the dataset into a training set, a test set, and a validation set. Perform data augmentation on the crop point clouds in the training set to obtain an expanded training set.

[0057] This invention utilizes the publicly available datasets Pheno4D and Crops3D. The Pheno4D dataset includes seven corn plants and seven tomato plants. A high-precision laser triangulation scanner (Perceptron Scan WorksV5) was used to scan the same plant daily. The corn was scanned continuously for 12 days, resulting in 84 point clouds (49 labeled, 35 unlabeled), with 7 labeled points per plant. The tomato was scanned continuously for 20 days, resulting in 140 point clouds (77 labeled, 63 unlabeled), forming complete growth cycle data. Each point is labeled as soil, stem, or leaf, and each leaf is assigned a unique ID, supporting cross-time-series tracking. Each labeled point cloud is labeled using two methods: leaf sheath method and tip method. This invention uses the leaf sheath method to label the corn dataset. Initially, ground points are labeled as 0, stems as 1, and leaves as 2, 3, 4, and 5. Points labeled as ground are removed from each point cloud file before preprocessing. This invention re-labels the points, assigning different numerical labels to the stems and leaves of different plants according to the order from M01 to M07, with the values ​​increasing sequentially, resulting in 35 different labels. The Crops3D dataset, a 3D crop point cloud dataset specifically designed for agricultural applications, includes 1230 point cloud samples from eight important crops: cabbage, cotton, corn, potato, rapeseed, rice, tomato, and wheat.

[0058] Specifically, preprocessing included: dividing the dataset into training, test, and validation sets in a 7:2:1 ratio. Information about the labeled corn dataset in Pheno4D is shown in Table 1. A corn plant's point cloud was downsampled to 5000 points, and the stem and leaves were each downsampled to 2500 points. If there were no leaves, only the stem was downsampled to 2500 points. Gaussian noise with a standard deviation of 0.005 was applied to the stem and leaves, generating 50 synthetic variants for each sample. The entire dataset contained 2450 variants (7 original samples × 50 perturbations × 7 plants), with the original samples being labeled point clouds. These data were divided into training (70%), validation (10%), and test (20%) sets. The validation set was strictly isolated from the training data to prevent information leakage. Partial information about the Crops3D dataset is shown in Table 2. The point cloud of a single plant was downsampled to 10,000 points using the farthest point sampling (FPS) method. The Crops3D dataset was partitioned using the same method as the corn dataset in Pheno4D. For data augmentation, an augmentation factor aug_factor=10 was introduced to expand the original number of samples in the training set by a factor of 10. See [link to documentation]. Figure 2 Data augmentation techniques, including translation, scaling, geometric rotation, and Gaussian filtering, are applied to the crop point cloud in the training set. Specifically, translation (… Figure 2(b) : Randomly move the entire point cloud along the X, Y, or Z axis, with a maximum movement of 10% of the original spatial extent in each dimension. Scaling ( Figure 2 (c) : Uniformly scales the point cloud to between 80% and 125% of its original size. Rotation ( Figure 2 (d): Randomly rotate the point cloud around three axes, with a rotation angle range of ±15°. Jitter ( Figure 2 (e): Randomly perturb the point coordinates along the X, Y, and Z axes, with a perturbation range of [0, 0.05]. Clipping ( Figure 2 f): Samples a cuboid sub-volume from the original point cloud, the cuboid volume occupying 60% to 100% of the original point cloud volume, with aspect ratios varying within the range of [0.25, 3.32]. Slicing ( Figure 2 (g) This method removes randomly sampled cuboids. The lengths of the cuboids in each dimension are scaled within the range [0.1, 0.4] of the corresponding dimensions of the original point cloud to obtain an expanded training set. The expanded training set obtained after data augmentation can simulate real-world disturbances (noise, occlusion), making the model more stable in complex agricultural environments, preventing the model from memorizing training data, learning more essential features to reduce overfitting, and improving metrics for tasks such as classification and segmentation (e.g., IoU, OA, AP), supporting high-precision agricultural applications.

[0059] Table 1. Information on the corn dataset in Pheno4D dataset.

[0060]

[0061] Table 2 Information about the Crops3D dataset

[0062]

[0063] Step 2: Construct the Crop-PVCNet segmentation network model.

[0064] For details, see Figure 3 The Crop-PVCNet segmentation model uses an improved SimSiam self-supervised architecture, abandoning the Convolutional Neural Network (CNN) encoder designed for images and introducing a point-voxel cross-Cuboid self-attention (PVCE) encoder specifically tailored for point clouds. The Crop-PVCNet segmentation network model is constructed as a weight-sharing, two-branch Siamese network structure, defined as an online branch and a target branch. Each branch contains three core modules: enhancement channels... Point-voxel cross-attention encoder (PVCE) and projection network For input point cloud samples ,in, This indicates the number of points in a point cloud sample. Additional feature dimensions besides coordinates (such as color and normal vector). The online branch and the target branch each use two independent enhancement channels. and Generate two enhanced views and The enhancement channel uses rotation, coordinate jitter, translation, and point dropping operations to change the spatial morphology of the point cloud without altering its dimensions, thus constructing semantically consistent but morphologically diverse point cloud sample pairs.

[0065] In the online branch, projection network Output feature representation Further input to the prediction network Features are mapped to predicted representations using batch normalization (BN) and the LeakyReLU activation function. In the objective branch, to avoid model collapse, a stop-gradient (sg) operation is introduced to stabilize its output. (Projection network) The purpose is to perform a nonlinear transformation on the features extracted by the PVCE encoder, mapping them to a high-dimensional feature space, thereby unifying the feature dimensions under different augmented views and improving the discriminative ability of the features. The similarity loss and loss function L are obtained by comparing the feature representations output by the two branches.

[0066] Specifically, the PVCE encoder will enhance the input point cloud view. Mapping to high-dimensional semantic features .in, This indicates the number of points in a point cloud sample. In addition to coordinates, there are additional feature dimensions (such as color and normal vector). D is the feature dimension of the high-dimensional semantic features, representing the length of the high-dimensional semantic features extracted by the PVCE encoder for each point. The PVCE encoder consists of three point-voxel encoder (PVE) layers, which are connected sequentially. Each PVE layer extracts complementary features through a dual-branch architecture.

[0067] For details, see Figure 4 (a) The PVE layer includes parallel point branches and voxel branches. (The second part is incomplete and requires further context.) Taking layers as an example, high-dimensional semantic features Or point cloud augmented view The input is fed into the PVE layer, where global geometric features are obtained through point self-attention mechanisms on the point branches. Local geometric features are obtained on the voxel branch through voxelization, VTB modules, and devoxelization. The global geometric features obtained at the end of each PVE layer and local geometric features Fusion is simply adding together to output semantic features. , ( ). The feature dimension is the feature dimension of the high-dimensional semantic features. is the feature dimension after projection.

[0068] For details, see Figure 5 The point branches in the PVE layer use a point self-attention mechanism module to model the global dependencies between points in the point cloud. This mechanism utilizes high-dimensional semantic features. Or point cloud augmented view As input, the high-dimensional semantic features of the input are first processed through a multilayer perceptron (MLP) consisting of batch normalization (BN) and ReLU activation function. Or point cloud augmented view Projecting onto a high-dimensional vector space yields a high-dimensional feature representation. Subsequently, high-dimensional features The inputs are fed into three linear transformation layers respectively. , and Generate query matrices sequentially. Key matrix Sum matrix Among them, the query matrix AND key matrix The channel dimension is set as a high-dimensional feature. The one-quarter of the channel count is used to learn diverse dependencies in point clouds across multiple low-dimensional spaces, thereby capturing global contextual information from multiple semantic perspectives. Ultimately, by fusing this diverse information, the diversity and discriminative power of feature representation are enhanced.

[0069] Point self-attention mechanism module for query matrix The key matrix K is taken as input, and the query matrix is ​​calculated. With K T The matrix multiplication is performed and normalized using the Softmax function to obtain the attention weight matrix. Finally, this weight matrix is ​​combined with the value matrix. Multiply to output the attention matrix. :

[0070]

[0071] Among them, attention matrix The attention matrix R and high-dimensional features are combined. The pieces are stitched together using a multilayer sensor. Obtain the global output features of the point branch:

[0072]

[0073] in, Concat represents the matrix concatenation operation along the feature dimension. This represents the operation functions of a multilayer perceptron.

[0074] Specifically, the voxel branch in the PVE layer first undergoes voxelization to receive data from the first voxel branch. High-dimensional semantic features of layer input Or point cloud enhancement view 𝑉 and target voxel resolution Output voxel features , It refers to the feature dimension. The voxelization process mainly includes three key steps: normalization, point-to-voxel mapping, and feature aggregation. Specifically: First, normalization centers the coordinates to the centroid and scales them to a [0, 1] unit cube, transforming the point cloud into a unified coordinate system. Then, voxel mapping discretizes the normalized space into... The grid; each point is mapped to a corresponding grid index along the x, y, and z axes based on its normalized coordinates, with the index range being [0, ..., ...]. Finally, feature aggregation is performed by calculating the average feature vector of all points in each occupied voxel using feature averaging, and then aggregating these features to generate voxel features. .

[0075] For details, see Figure 4 (b) In the PVE layer, a VTB module is introduced. The VTB module includes a first-layer normalization, a CCMHA module, a first residual connection, a second-layer normalization, a multilayer perceptron, and a second residual connection, which are connected in sequence. The first residual connection is also connected to the input voxel features. The input of the second residual is connected to the output of the first residual. The VTB module outputs voxel features after voxelization. As input. During feature processing, the input features... First, the data undergoes a layer normalization process, followed by input to the CCMHA module. The output of the CCMHA module is compared with the voxel features of the original input. Feature fusion is performed through the first residual connection to obtain features. .feature After further layer normalization in the second layer, the data is fed into a multilayer perceptron for nonlinear transformation. The output is then connected to the features from the previous stage via a second residual connection. The fusion process ultimately outputs enhanced voxel features. .

[0076] For details, see Figure 4 (b) The VTB module includes a CCMHA module, which is a multi-head self-attention mechanism based on three-dimensional cuboid partitioning, used to compute self-attention between voxels within a local region. The input to the CCMHA module is layer-normalized features. , This is the operation function for layer normalization.

[0077] For details, see Figure 6 For each query voxel in the input, the CCMHA module constructs three mutually perpendicular cuboid regions in its neighborhood, perpendicular to the X, Y, and Z coordinate axes, respectively. Each cuboid region is used to compute local attention in that direction. In the multi-head self-attention mechanism, the total number of attention heads K is set to a multiple of 3, and K / 3 attention heads are assigned to each direction to process the local features in the corresponding direction.

[0078] Taking the attention calculation of a cuboid perpendicular to the X-axis as an example, the CCMHA module first processes the input features... Divide the x-axis into n non-overlapping cuboid units, denoted as . The height of each cuboid is cw. For each partitioned cuboid element Each is achieved through three independent linear transformation layers. , , Generate query matrix Key matrix AND-value matrix Where 𝑘 represents the 𝑘-th attention head, and 𝑚 represents the 𝑚-th cuboid cell. Each attention head computes a query matrix within its corresponding cuboid cell. transpose of the key matrix Matrix multiplication and normalization using the Softmax function, and the sum of the matrix values. Multiplying them together yields the local attention vector features. Furthermore, the CCMHA module merges the local attention vector features of all K cuboids in the same direction to form a local self-attention representation in that direction, denoted as . Similarly, performing the same operation on the Y-axis and Z-axis directions respectively yields the following results. and Finally, the local self-attention outputs in the X, Y, and Z axes are concatenated and passed through a linear projection layer. Feature fusion is performed to output comprehensive multi-directional local self-attention features. , This represents the operation functions of the CCMHA module.

[0079] For details, see Figure 4 (a) The VTB module in the PVE layer outputs voxel features. To fuse the features from both the point and voxel branches, trilinear interpolation is used to integrate the voxel features. Perform devoxometry. Based on the normalized coordinates of each point in the voxel grid, convert the voxel features... Accurately map to the corresponding point features. For a point Calculate the normalized coordinates to obtain Calculate the index of the voxel unit to which it belongs based on its coordinates. .in, , , .in, This is the floor function. Furthermore, the relative position of the point within the cell is obtained. .in, , , Based on the above index and relative position The local features of a point are obtained by weighted fusion of the eigenvalues ​​at the eight vertices of the voxel cube. Simultaneously, based on the point-to-voxel mapping process described above, the position index of the point in the original point cloud sequence is recorded. By analyzing all points in the cloud Perform the above trilinear interpolation operation on each point sequentially, generating local features for each point. Index the local features of all points by their location. The points are sequentially aggregated to obtain a local geometric feature representation of the entire point cloud. ,in, The number of points in the point cloud. For feature dimensions.

[0080] For details, see Figure 7 and Figure 4 In (a), to simultaneously utilize the features extracted from both shallow and deep layers of the network, the PVE layer proposes a feature fusion module. This module integrates the output features of the first-layer PVE. and the output characteristics of the third PVE layer Fusion, resulting in fusion characteristics After processing by the feature fusion module, the output of the PVCE encoder is:

[0081]

[0082]

[0083] in, This indicates a feature concatenation operation. It is an operation of a multilayer perceptron. This indicates a max pooling operation. It is a repetitive operation. It is a global context feature, through Operation on output features A nonlinear transformation is performed, followed by max pooling to obtain the feature vector. Finally, a repeat operation is used to copy and expand the feature vector to obtain features with the same dimensions as before max pooling.

[0084] Step 3: Train the Crop-PVCNet segmentation network model using the expanded training set to obtain the trained Crop-PVCNet network model.

[0085] The loss function used during training mainly includes similarity loss.

[0086] Similarity loss measures the similarity between the output features of the online branch and the target branch, driving the model to learn augmentation-invariant feature representations. A symmetric similarity loss is constructed, consisting of two symmetric terms, each balancing the similarity between the predicted output of the online branch for a given augmented view-level feature and the output of the target branch after applying the stopping gradient (sg) operation to another augmented view-level feature. The model is forced to learn invariant features under two different augmentation transformations, thus driving the network to extract the common semantic information behind the two views, rather than relying on a specific augmentation transformation.

[0087] Specifically, the loss function consists of two terms: the first term calculates the online branch pair of view features. The prediction results and the target branch for view features The similarity between outputs; the second term calculates the online branch pair of view features. The prediction results and the target branch for view features Similarity between outputs. This symmetric design forces the model to learn consistent visual content from two enhanced perspectives. This is achieved through the representation of features at both view levels. and The symmetric computation of this loss function ensures that the model maintains consistency in feature representation across different data augmentation perspectives, thereby improving the model's generalization ability and robustness in point cloud segmentation tasks. The overall pre-trained loss function L alternately predicts the target across different views, and can be expressed as follows:

[0088]

[0089] in, This represents the operation function of the prediction network that operates only on the online branches, resulting in the predicted feature representation. This indicates stopping the gradient operation, which is applied to the feature representation of the target branch, blocking gradient backpropagation and ensuring training stability. D(·) represents the feature similarity measure, calculated using cosine similarity.

[0090]

[0091] Where ∥·∥2 represents the L2 norm.

[0092] The proposed PVCE serves as the encoder network, generating output features of dimension 1024. The feature fusion module uses an MLP with fully connected layers of output dimensions [512, 1024]. Furthermore, a projection network... The prediction network consists of a three-layer MLP [512, 256, 256] that does not contain an output dimension. A two-layer MLP with dimensions [512, 256] is used. All MLP layers include batch normalization (BN) and LeakyReLU activation functions, except for the output of the last layer. The feature dimensions in the PVE layer are [96, 192, 192], the resolution is [30, 15, 15], the voxel width is [1, 3, 3], and the number of attention heads is [3, 6, 6]. The Crop-PVCNet of this invention is trained on the Pheno4D and Crops3D datasets with the following configuration: learning rate set to 0.001, batch size of 16, and maximum number of iterations set to 250 epochs. To prevent overfitting, dropout regularization is introduced into the model with a dropout rate of 0.5, and the momentum coefficient of the optimizer is set to 0.9. The maximum capacity of the input point cloud is limited to 4096 points, and the computational cost is stable, ensuring the efficiency and stability of the training process. Cosine learning rate decay is used for optimization (scheduler='cos') to make the model converge more smoothly and stably.

[0093] Step 4: Input the crop point cloud data from the test set into the trained Crop-PVCNet network model for testing, and obtain the segmented crop point cloud file.

[0094] After training, only the encoder (PVCE) weights are retained, and the projection network is removed. and prediction networks Partially, using the trained weights, crop point cloud data from the test set is input into the network. The enhancement channel only normalizes the coordinates of the input point cloud. The PVCE encoder is used to extract point cloud features, resulting in features that fuse local details and global context. Then, the fused features output by the PVCE encoder. Next, a classification head (fully connected layer) is added. The argmax function is applied to the output of the classification head to obtain the final semantic category label for each point. The original coordinate information and the predicted semantic category label information are written to a txt file and saved, finally resulting in a semantically segmented 3D point cloud file.

[0095] To verify the performance of the model in this invention, a systematic evaluation of the Crop-PVCNet network segmentation model was conducted compared with existing methods to verify its effectiveness. Specifically, the point cloud segmentation quality was evaluated using the evaluation metrics IoU (Intersection Over Union) and mIoU (Mean Intersection Over Union). The method proposed in this invention achieves significant improvements in both of these quantitative evaluation metrics. Existing methods include: PointNet (Deep Learning on Point Sets for 3D Classification and Segmentation), PointNet++ (Deep Hierarchical Feature Learning on Point Sets in a Metric Space), DGCNN (Dynamic Graph Convolutional Neural Network), FF-Net (Feature-Fusion-Based Network for Semantic Segmentation of 3D Plant Point Cloud), Win-Former (Window-Based Transformer for Maize Plant Point Cloud Semantic Segmentation), and PSegNet (Simultaneous Semantic and Instance Segmentation for Point Clouds of Plants).

[0096] IoU evaluates model performance by calculating the similarity between the generated point cloud and the real point cloud using three metrics: True Positive Examples (TP), False Positive Examples (FP), and False Negative Examples (FN). Wherein, IoU is the intersection-union ratio; the larger the value, the closer the point cloud organ segmentation result is to the true result.

[0097] mIoU evaluates the model's performance by using the average IoU across all N crop organs:

[0098]

[0099] Where N represents the total number of categories in the plant point cloud, Indicates the first Organ categories (1≤ ≤N).

[0100] For details, see Figure 8 We analyzed the qualitative results of maize in the Pheno4D dataset and compared the model performance with existing methods. As shown in the figure, PointNet++ has limited ability to maintain the natural curvature of leaves on March 17th; the leaf edges are almost completely connected to the stem, lacking a precise segmentation boundary. In contrast, our method achieves highly accurate part-level segmentation. This consistency demonstrates the robustness of our method in capturing the dynamic morphological changes of individual plants over time. The segmentation evaluation results are shown in Table 3. Table 3 shows that our method achieves 98.51% mIoU, higher than PSegNet's performance of 11.2%. Furthermore, the IoU values ​​of our method are 98.68% and 98.45%, respectively, significantly higher than PointNet++'s values ​​of 64.38% and 96.65%, further confirming the advantage of our method in segmenting the details of non-rigid object parts.

[0101] Table 3. Evaluation results of objective indicators compared with other methods on maize.

[0102]

[0103] For details, see Figure 9Further qualitative analysis was conducted on the eight crops in the Crops3D dataset, comparing the model performance with existing methods. The segmentation evaluation results are shown in Table 4, where existing methods include CurveNet (Curvature-Based Multitask Learning Deep Networks), PointMLP (An EfficientMLP-like Point Cloud Learning Network), and PointCloudMamba (Point Cloud Learning via State Space Model). Table 4 shows that for maize and tomato, although PointCloudMamba achieved a higher mIoU, the component analysis revealed that it only performed slightly better in the soil portion, with an advantage of 0.6%. Conversely, our method consistently outperformed PointCloudMamba in all other plant organs. Overall, our method achieved higher overall accuracy and demonstrated more balanced segmentation performance across different plant organs, exhibiting superior robustness and generalization ability. Figure 9 As can be seen, the method achieves the best segmentation results on a variety of crops, and even surpasses the PointCloudMamba model in segmentation of stems, leaves, and fruits of complex crops such as corn and tomatoes. Although the PointCloudMamba model achieves complete segmentation of leaf regions, it still displays blurred boundaries between adjacent leaves and at stem-leaf junctions. The difference between the predicted segmentation results and the actual segmentation results of the method of this invention is minimal, demonstrating the effectiveness of the method in capturing fine-grained plant morphology and showcasing superior crop point cloud segmentation capabilities.

[0104] Table 4. Objective metrics evaluation results (%) of the Crops3D dataset compared with other methods

[0105]

[0106] Specifically, further analysis of the generalization ability on the Shapenet dataset was conducted, and the model performance was compared with existing methods. The segmentation evaluation results are shown in Table 5 and... Figure 10As shown, existing methods include RSCNN (Relation-Shape Convolutional Neural Network for Point Cloud Analysis), PCT (Point Cloud Transformer), LGAN (Learning Representations and Generative Models for 3D Point Clouds), HNS (Self-Contrastive Learning with Hard Negative Sampling for Self-supervised Point Cloud Learning), Jigsaw3D (Self-supervised Deep Learning on Point Clouds by Reconstructing Space), OcCo (Unsupervised Point Cloud Pre-training via Occlusion Completion), and STRL (Spatio-temporal Self-supervised Representation Learning for 3D Point Clouds). The method of this invention achieves an average IoU score of 85.8%, demonstrating that the pre-trained features of this model can be effectively transferred to dense prediction tasks. Furthermore, compared to supervised methods, the method of this invention outperforms DGCNN by 0.6% in instance mIoU. It also achieves comparable segmentation results compared to state-of-the-art RS-CNN and PCT methods. The method of this invention achieves the highest mIoU among self-supervised algorithms in most categories, and in some cases even surpasses the segmentation performance of supervised methods. These results highlight the powerful generalization ability of this method. Figure 10 As can be seen, the method of this invention can achieve segmentation highly consistent with the actual annotations, effectively distinguishing and separating different parts of an object. It is noteworthy that even in complex structural categories such as airplanes, automobiles, and motorcycles, where precise component definitions are required, Crop-PVCNet can successfully identify each component.

[0107] Table 5. Segmentation mIoU results (%) after fine-tuning on a subset of ShapeNet datasets.

[0108]

[0109] To further demonstrate the effectiveness of the present invention, ablation experiments were conducted to verify the effectiveness of the PVCE module and the CCMHA module and their effect on improving crop segmentation quality.

[0110] Specifically, a series of experiments compared the impact of voxel branching and point branching in the PVCE module on model performance. The experimental results are shown in Table 6. If only point branching is introduced, the overall accuracy (OA) is only 73.5%; if only voxel branching is introduced, the overall accuracy (OA) is only 73.1%. Introducing both branches simultaneously increases the OA accuracy to 76.2%, an increase of approximately 2.9%, indicating better segmentation accuracy. The experimental results demonstrate the effectiveness of the point-voxel dual-branching mechanism in the PVCE module and its significant improvement in point cloud segmentation accuracy.

[0111] Table 6. Impact of PVCE module voxel branching and point branching on model performance

[0112]

[0113] Specifically, a series of experiments compared the performance of the CCMHA module in the model, evaluating six different attention mechanisms, including sliding window, offset window, spatial separation, sequential axis, cross-cube, and cross cube. As shown in Table 7, the cross cube attention mechanism (CCMHA) achieved the highest accuracy of 76.2%, significantly outperforming other variants. In contrast, strategies such as sliding window and sequential axis have limited ability to model long-range contextual relationships in sparse and irregularly distributed point clouds, thus performing relatively poorly. Although these architectures can capture certain spatial patterns, they still struggle to effectively model complex organ-level structures in 3D scenes. These results demonstrate that the cross cube attention design excels at capturing multi-scale contextual information and global shape features of plants, thus significantly enhancing the accuracy of feature learning and semantic segmentation.

[0114] Table 7 Experimental Results of CCMHA Module

[0115]

[0116] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A self-supervised method for crop point cloud segmentation based on point-voxel bibranching, characterized in that, The steps are as follows: Step 1: Compile the obtained crop point clouds into a dataset and divide the dataset into a training set, a test set, and a validation set. Perform data augmentation on the crop point clouds in the training set to obtain an expanded training set. Step 2: Construct the Crop-PVCNet segmentation network model, including online branches and target branches, with point-voxel cross self-attention encoders on both the online and target branches; Step 3: Train the Crop-PVCNet segmentation network model using the expanded training set to obtain the trained Crop-PVCNet network model; Step 4: Input the crop point cloud data from the test set into the trained Crop-PVCNet network model for testing, and obtain the segmented crop point cloud file.

2. The self-supervised method for crop point cloud segmentation based on point-voxel bi-branching as described in claim 1, characterized in that, The leaf sheath method was used to label the points in the crop point cloud of the dataset. Different numerical labels were assigned to the stems and leaves of different plants, and the values ​​were sequentially increased. Points marked as ground were removed from each point cloud file. The data augmentation includes translation, scaling, rotation, jittering, and filtering of point clouds, with filtering including cropping and dicing.

3. The self-supervised method for crop point cloud segmentation based on point-voxel bibranching according to claim 1 or 2, characterized in that, The online branch and the target branch are dual branches with shared weights. The input of the point-voxel cross-attention encoder is connected to the enhancement channel, and the output of the point-voxel cross-attention encoder is connected to the projection network. The output of the projection network of the online branch is connected to the prediction network, and the output of the projection network of the target branch is connected to the SimSiam network. The loss function L is obtained by comparing the similarity loss of the feature representations output by the online branch and the target branch.

4. The self-supervised method for crop point cloud segmentation based on point-voxel bi-branching as described in claim 3, characterized in that, The enhancement channel enhances the input point cloud sample through rotation, coordinate jitter, translation, and point dropping operations. Process the data to generate an enhanced point cloud view. ;in, This indicates the number of points in a point cloud sample. Additional feature dimensions besides coordinates; The point-voxel cross-attention encoder enhances the input point cloud view. Mapping to high-dimensional semantic features Where D is the feature dimension of the high-dimensional semantic feature; The projection network performs a nonlinear transformation on the features extracted by the point-voxel cross-attention encoder, mapping them to a high-dimensional feature space; the feature representation output by the online branch projection network... The input is further fed into a prediction network, which then uses batch normalization and the LeakyReLU activation function to represent the features. Mapping to prediction representation In the target branch, the SimSiam network stabilizes the output of the projection network by stopping the gradient operation.

5. The self-supervised method for crop point cloud segmentation based on point-voxel bi-branching according to claim 4, characterized in that, The point-voxel cross self-attention encoder includes three PVE layers, which are connected in sequence. Each PVE layer extracts complementary features through parallel point branches and voxel branches. The point branch is equipped with a point self-attention mechanism, and the voxel branch is equipped with voxelization, VTB module and anti-voxelization connected in sequence. High-dimensional semantic features Or point cloud augmented view Enter the number The PVE layer of the layer obtains global geometric features through a point self-attention mechanism on the point branches. Local geometric features are obtained on the voxel branch through voxelization, VTB modules, and devoxelization. Finally, the obtained global geometric features and local geometric features Fusion to output high-dimensional semantic features As the output of the point-voxel cross-attention encoder or the input of the next PVE layer, where, , The feature dimension of high-dimensional semantic features is the feature dimension after projection.

6. The self-supervised method for crop point cloud segmentation based on point-voxel bi-branching according to claim 5, characterized in that, The point self-attention mechanism uses a multilayer perceptron to process the high-dimensional semantic features of the input. Or point cloud augmented view Projecting onto a high-dimensional vector space yields high-dimensional features. Then the high-dimensional features The inputs are fed into three linear transformation layers respectively. , and Generate query matrices sequentially. Key matrix Sum matrix ; using query matrix The key matrix K is taken as input, and the query matrix is ​​calculated. With K T The matrix multiplication is performed and then normalized using the Softmax function to obtain the attention weight matrix. The attention weight matrix and the value matrix are then compared. Multiply to output the attention matrix. Attention matrix and high-dimensional features The data is then concatenated, and the global output features of the point branches are obtained through a multilayer perceptron. ; The voxelization receives from the first High-dimensional semantic features of layer input Alternatively, the point cloud augmented view V can be normalized to center the coordinates to the centroid and scaled to a [0, 1] unit cube, thus transforming the point cloud into a unified coordinate system. The normalized space is discretized using voxel mapping. The grid; each point is mapped to a corresponding grid index along the x, y, and z axes based on its normalized coordinates, with the index range being [0, ..., ...]. The average feature vector of all points in each occupied voxel is calculated using feature averaging, and then aggregated to generate voxel features. ;in, For target voxel resolution, It is the feature dimension; The devoxometry uses trilinear interpolation to apply voxel features to the output of the VTB module. Processing: Based on the normalized coordinates of each point in the voxel grid, the voxel features are... Accurately map to the corresponding point features.

7. The self-supervised method for crop point cloud segmentation based on point-voxel bi-branching according to claim 5 or 6, characterized in that, The VTB module comprises a first-layer normalization, a CCMHA module, a first residual connection, a second-layer normalization, a multilayer perceptron, and a second residual connection, connected in sequence. The first residual connection is connected to the input voxel features. The input of the second residual connection is connected to the output of the first residual connection; Input voxel features After the first layer of normalization, the data is processed for layer normalization and then input into the CCMHA module; the output of the CCMHA module is compared with the input voxel features. Feature fusion is performed through the first residual connection to obtain features. ;feature After undergoing layer normalization in the second layer, the data is fed into a multilayer perceptron for nonlinear transformation. The output is then connected to the features via a second residual connection. Fusion, outputting enhanced voxel features .

8. The self-supervised method for crop point cloud segmentation based on point-voxel bi-branching according to claim 7, characterized in that, The CCMHA module is a multi-head self-attention mechanism based on three-dimensional cuboid partitioning. For each query voxel in the input, the CCMHA module constructs three mutually perpendicular cuboid regions in its neighborhood, which are perpendicular to the X, Y, and Z coordinate axes, respectively. Local attention is calculated in each cuboid region in that direction. In the multi-head self-attention mechanism, the total number of attention heads K is set to be a multiple of 3, and K / 3 attention heads are assigned to each direction to process the local features in the corresponding direction.

9. The self-supervised method for crop point cloud segmentation based on point-voxel bi-branching according to claim 8, characterized in that, The loss function L consists of two symmetric terms, each of which measures the similarity between the predicted output of the online branch for one enhanced view-level feature and the output of the target branch after applying the stopping gradient operation to another enhanced view-level feature. The loss function L, which alternately predicts the target across different views, is represented as follows: in, This represents the operation function of a prediction network that operates only on the online branches. This indicates stopping the gradient operation, and D(·) represents the feature similarity measure, which is calculated using cosine similarity.

10. The self-supervised method for crop point cloud segmentation based on point-voxel bi-branching according to claim 9, characterized in that, The method for calculating local attention in the cuboid region perpendicular to the X-axis is as follows: The features normalized by the layers are... Divide the x-axis into n non-overlapping cuboid units. The height of each cuboid unit is cw. For each partitioned cuboid unit Each is passed through three independent linear transformation layers. , , Generate the query matrix for the nth attention head respectively. Key matrix AND-value matrix Each attention head computes the query matrix within its corresponding cuboid cell. transpose of the key matrix Matrix multiplication and normalization using the Softmax function, and the sum of the matrix values. Multiplying them together yields the local attention vector features. The local attention vector features of all K cuboids in the same direction are merged to form a local self-attention representation in that direction. Perform the same operation on the Y-axis and Z-axis respectively to obtain the local self-attention representation. and The local self-attention representations in the X, Y, and Z axes are concatenated and then projected through a linear projection layer. Feature fusion is performed to output comprehensive multi-directional local self-attention features. , This represents the operation functions of the CCMHA module; The feature fusion module is used to integrate the output features of the first PVE layer. and the output characteristics of the third PVE layer The fusion is performed to obtain the fusion features. The output of the point-voxel cross-attention encoder, and the fused features in, This indicates a feature concatenation operation. These are the operation functions of a multilayer perceptron. This indicates a max pooling operation. It is a repetitive operation. It is a global context feature and has the same feature dimension as before the max pooling operation. The attention weight matrix is ​​as follows: The global geometric features: in, Here, is the attention function, and Concat() represents the matrix concatenation operation along the feature dimension. Represents the operation functions of a multilayer perceptron; The multilayer perceptron is composed of batch normalization and ReLU activation function; After training, the weights of the point-voxel cross-self-attention encoder are retained; the crop point cloud from the test set is input into the augmentation channel, and the coordinates are normalized. The point-voxel cross-self-attention encoder is then used to extract point cloud features to obtain features that fuse local details and global context. ,feature Input a fully connected layer, apply the argmax function to the output of the fully connected layer to obtain the final semantic category label for each point; write the original coordinate information and predicted semantic category label information of the crop point cloud into a txt file and save it to obtain a 3D point cloud file after semantic segmentation.

Citation Information

Patent Citations

  • Farmland feature semantic segmentation method based on laser radar and SegNet

    CN119360384A