Rape pod segmentation network model and system and pod counting method

By improving the deep learning model structure and combining three-dimensional reconstruction and clustering algorithms, an automated segmentation model of rapeseed fruit based on the improved three-dimensional point cloud segmentation network model was designed, which solved the problems of insufficient accuracy, low computational efficiency and poor noise robustness of rapeseed fruit segmentation and counting in the existing technology, and achieved high-precision and lightweight identification and statistics of rapeseed fruit, meeting the real-time processing needs of agricultural scenarios.

CN120182608AActive Publication Date: 2025-06-20SICHUAN AGRI UNIV
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510655388.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-21
Publication Date
2025-06-20
Estimated Expiration
2045-05-21

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient accuracy, low calculation efficiency and poor noise robustness in the segmentation and counting of rapeseed fruits, which is difficult to meet the real-time processing needs of agricultural scenarios.

Method used

By improving the deep learning model structure and combining three-dimensional reconstruction and clustering algorithms, a rapeseed fruit automatic segmentation model based on the improved three-dimensional point cloud segmentation network model was designed. The model includes an encoding layer and a decoding layer, and uses anti-bottleneck KAN convolution, GLFN feature modulation block and ContraNorm normalization module to achieve high-precision and lightweight horn fruit recognition and statistics.

Benefits of technology

It realizes high-precision segmentation and counting of rapeseed fruits, significantly improves the calculation efficiency and noise robustness, meets the real-time processing needs of agricultural scenarios, and provides reliable technical support for rapeseed phenotype research and breeding work.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120182608A_ABST
    Figure CN120182608A_ABST
Patent Text Reader

Abstract

The invention belongs to the crossing field of computer vision and agricultural information technology, and relates to a rape pod segmentation network model, a system and a pod counting method, the rape pod segmentation network model comprises a coding layer and a decoding layer, the coding layer comprises a grouping sampling module and a trunk feature extraction module; the decoding layer comprises an interpolation module and a local feature transformation module; the rape pod segmentation system comprises an image acquisition module, a three-dimensional reconstruction module, a data processing module, a rape pod segmentation network model and a rear pod counting module. The oilseed rape pod segmentation network model and the automatic counting method are superior to various advanced models in oilseed rape segmentation tasks, the precision is high, the model parameter is only 5.72 M, excellent balance is achieved between the precision and the parameter quantity, good practical prospects are shown, the model has wide application potential in plant phenotype research, and the application prospect is wide. The method is especially suitable for accurate segmentation of rape organs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the cross - field of computer vision and agricultural information technology, and relates to a rapeseed silique segmentation network model, system and silique counting method based on three - dimensional point cloud processing. This technical solution realizes the automatic and high - precision segmentation and counting of rapeseed siliques by improving the deep - learning model structure and combining three - dimensional reconstruction and clustering algorithms, and can be widely applied to plant phenotype analysis, crop yield prediction and precision agriculture management. Background Art

[0002] In the phenotypic research of crops such as rapeseed, the number of siliques is a key indicator for evaluating yield potential and growth status. It not only directly determines the grain yield, but also reflects the genetic characteristics of the variety and the response ability to environmental stress. Therefore, accurately obtaining the number of siliques is of great significance for rapeseed breeding screening, growth monitoring and agricultural management.

[0003] Currently, the counting of rapeseed siliques mainly relies on manual methods, that is, staff observe and record the number of siliques in the sampling area. Although this method is intuitive and easy to implement, it has problems such as low efficiency, high labor intensity, easy omission and misjudgment, and it is difficult to obtain the spatial structure information of siliques, which limits subsequent modeling and dynamic analysis.

[0004] With the development of computer vision technology, some studies have tried to realize automatic silique detection through two - dimensional image processing techniques (such as color and texture feature extraction of RGB images). However, two - dimensional methods have the following limitations: Overlap and occlusion problems: Two - dimensional images cannot effectively distinguish dense or overlapping siliques, resulting in false detection or missed detection; Lack of three - dimensional information: It is impossible to restore the true spatial distribution of siliques, affecting the accuracy of counting; Complex background interference: Factors such as field lighting changes and foliage occlusion will significantly reduce the robustness of the algorithm.

[0005] In recent years, three - dimensional point cloud technology has provided a new direction for automatic silique counting due to its powerful spatial representation ability. However, current three - dimensional point cloud technology in rapeseed silique segmentation and counting has problems such as insufficient model accuracy, low computational efficiency, and poor noise robustness. There is an urgent need for a solution that combines lightweight and high - precision to meet the real - time processing requirements of agricultural scenarios. Summary of the Invention

[0006] In view of the problems of insufficient accuracy, low calculation efficiency, and poor noise robustness in rapeseed pod segmentation and counting in the prior art, the present invention proposes an automated rapeseed pod segmentation model based on an improved three-dimensional point cloud segmentation network model, a rapeseed pod segmentation system and a counting method based on the segmentation model. By innovating the network structure, optimizing the three-dimensional reconstruction process, and introducing an efficient clustering algorithm, the present invention realizes high-precision and lightweight pod recognition and statistics.

[0007] Through long-term exploration and attempts, as well as multiple experiments and efforts, and continuous reform and innovation, in order to solve the above technical problems, the technical solution provided by the present invention is to provide a rapeseed pod segmentation network model, including an encoding layer and a decoding layer, wherein: The encoding layer includes a grouping sampling module and a backbone feature extraction module; The decoding layer includes an interpolation module and a local feature transformation module; The backbone feature extraction module realizes multi-scale feature extraction through the following structure: a. Inverse bottleneck KAN convolution, which includes two parallel branches. The left branch sequentially performs SiLU activation and K×K convolution, and the right branch sequentially performs 1×1 convolution to expand the number of channels, Gaussian kernel radial basis function transformation, and 1×1 convolution to restore the number of channels; b. GLFN feature modulation block, which consists of a global-local spatial attention module and a partial convolutional feed-forward network, and is used to fuse local details and global context information; c. ContraNorm normalization module based on contrast learning, which suppresses noise and alleviates dimensional collapse by constraining the feature distribution.

[0008] Preferably, the grouping sampling module uses farthest point sampling (FPS) to generate centroid points, and searches for neighborhood points through the Ball Query algorithm. If the number of neighborhood points is insufficient, the center point is repeated to make up the deficiency.

[0009] Compared with the prior art, the beneficial effects of the present invention are: The rapeseed silique segmentation network model of the present invention realizes efficient rapeseed silique segmentation through a unique encoding layer and decoding layer structure. The grouped sampling module and backbone feature extraction module in the encoding layer work together to extract feature information of siliques from different scales. The inverse bottleneck KAN convolution uses two parallel branch structures, combines SiLU activation, K×K convolution, and 1×1 convolution to expand and restore the number of channels, enhancing the feature expression ability. At the same time, the Gaussian kernel radial basis function transformation further improves the accuracy of feature extraction. The GLFN feature modulation block integrates the global-local spatial attention module and the partial convolutional feed-forward network, effectively integrating local details and global context information, enabling the model to more accurately identify silique features. The ContraNorm normalization module, based on contrastive learning, constrains the feature distribution, effectively suppressing noise interference, alleviating the problem of dimensional collapse, and improving the stability and segmentation accuracy of the model. Overall, while maintaining a low number of parameters, this network model significantly improves the accuracy and efficiency of rapeseed silique segmentation, provides a reliable basis for subsequent silique counting and analysis, and promotes the automation process of rapeseed phenotype research and breeding work.

[0010] On the basis of the above technical solution, the present invention can also be improved as follows: Further: The Gaussian kernel radial basis function of the inverse bottleneck KAN convolution is defined as: Where represents the radial distance, is the standard deviation, which is used to control the width of the Gaussian function.

[0011] Preferably, σ is the standard deviation, and the initial value is 0.2.

[0012] Preferably, batch normalization (BatchNorm) and GELU activation are added after the 3×3 convolutional layer of the left branch.

[0013] Preferably, the output of the right branch is added to the left branch through a residual connection to accelerate the convergence of the model.

[0014] On the basis of the above technical solution, the present invention can also be improved as follows: Further: In the GLFN feature modulation block: The global-local spatial attention module evenly splits the input features into two parts, processes them through the global spatial attention module and the local spatial attention module respectively, and finally combines and outputs them; The partial convolutional feed-forward network encodes local context information through 3×3 convolution, and after splicing with the original features, fuses them through 1×1 convolution.

[0015] On the basis of the above technical solution, the present invention can also be improved as follows: Further: The calculation formula of the ContraNorm normalization module is: Where, is the input feature matrix, with dimensions D×N, where D is the feature dimension and N is the number of point clouds; is the transpose of the input feature matrix; s is the step parameter, used to control the amplitude of feature update; is the update amplitude control parameter, LN is the standard layer normalization; is the output feature matrix, used to suppress noise and alleviate dimensional collapse.

[0016] The present invention also provides a rapeseed pod segmentation system, including the following modules: a. An image acquisition module, used to obtain continuous images of rapeseed plants through a rotating platform and multi-view shooting; b. A three-dimensional reconstruction module, which generates a dense three-dimensional point cloud of rapeseed plants based on the structure from motion algorithm and neural radiance field technology; c. A data processing module, used to preprocess, enhance, and label the point cloud to construct a segmentation data set; d. The rapeseed pod segmentation network model described above, used to perform semantic segmentation on the preprocessed point cloud; e. A pod counting module, which performs clustering statistics on the segmented pod point cloud based on the DBSCAN clustering algorithm.

[0017] Compared with the prior art, the beneficial effects of the present invention are: The system of the present invention not only effectively improves the efficiency and accuracy of rapeseed pod counting, but also significantly reduces the intensity and error of manual operations, provides strong technical support for rapeseed phenotype research and breeding work, and has good practicability and promotion prospects.

[0018] On the basis of the above technical solution, the present invention can also be improved as follows: Further: The three-dimensional reconstruction module uses the Nerfacto model for dense point cloud reconstruction, specifically including: Generating sparse point clouds and camera poses through the structure from motion algorithm, inputting them into the Nerfacto model to optimize the implicit scene representation, and sampling from the density field to generate dense point clouds.

[0019] On the basis of the above technical solutions, the present invention can also be improved as follows: Further: The point cloud preprocessing of the data processing module includes: Removing noise points through radius statistical filtering and compressing the single-plant point cloud to 20,000 to 30,000 points through uniform downsampling; The enhancement includes random translation, flipping, cropping, rotation, and adding Gaussian noise.

[0020] The present invention also provides an automatic rapeseed pod counting method, including the following steps: S1. Collect multi-view images of rapeseed plants and reconstruct three-dimensional point clouds through the structure from motion algorithm and neural radiance field technology; S2. Preprocess, enhance, and annotate the point clouds to construct a training set and a test set; S3. Use the described segmentation network model to perform semantic segmentation on the point clouds to distinguish pod and stem point clouds; S4. Apply the DBSCAN clustering algorithm to the segmented pod point clouds to count the number of pods; Among them, the parameters of the DBSCAN algorithm are optimized through grid search, and the optimal parameter combination is ε = 0.10, MinPts = 7.

[0021] Compared with the prior art, the beneficial effects of adopting the above further technical solutions are as follows: The automatic rapeseed pod counting method of the present invention realizes the full-process automation from image acquisition to pod counting, effectively improving the counting efficiency and accuracy. By collecting multi-view images and combining the structure from motion algorithm and neural radiance field technology for three-dimensional reconstruction, an accurate three-dimensional point cloud of rapeseed plants can be generated, providing a high-quality data basis for subsequent processing. The preprocessing, enhancement, and annotation operations on the point clouds ensure the accuracy and diversity of the data, helping to improve the performance of the segmentation model. Using the segmentation network model for semantic segmentation can efficiently distinguish pod and stem point clouds, providing a guarantee for accurate counting. Finally, applying the DBSCAN clustering algorithm and optimizing the parameters realizes the accurate statistics of the number of pods, meeting the requirements of fast and accurate counting in the field, and providing strong support for rapeseed phenotype analysis and breeding screening.

[0022] On the basis of the above technical solutions, the present invention can also be improved as follows: Further: The performance of the DBSCAN clustering is evaluated by the following indicators: Average recognition accuracy, mean absolute percentage error, and coefficient of determination, where the mean absolute percentage error is defined as: is the total number of samples, is the true silique number, is the predicted number.

[0023] Furthermore, the coefficient of determination is: The mPr is calculated by the sliding window method as: The present invention also provides a computer-readable storage medium storing a computer program, and when the program is executed by a processor, it implements the automatic silique counting method for rapeseed as described above. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required to be used in the embodiments. It should be understood that the following drawings only show some embodiments of the present invention, and thus should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0025] Figure 1 is the structural diagram of the rapeseed silique segmentation network model based on the improved PointNet++.

[0026] Figure 2 is the network structure diagram of the backbone feature extraction network KGL-PointNet of the present invention.

[0027] Figure 3 is the structural diagram of the anti-bottleneck KAN convolution of the present invention.

[0028] Figure 4 is the global and local attention module diagram of the present invention.

[0029] Figure 5 is the rapeseed silique segmentation and counting flow chart of the present invention. Figure 5 In, A is the SfM sparse point cloud reconstruction, B is the Nerfacto dense point cloud reconstruction, C is the point cloud after preprocessing, D is the point cloud after data augmentation, E is the labeled rapeseed point cloud label, F is the semantic segmentation prediction point cloud, G is the semantic segmentation prediction silique point cloud, and H is the DBSCAN silique clustering result.

[0030] Figure 6 is the rapeseed data augmentation flow chart of the present invention.

[0031] Figure 7 is the visualization diagram of the DBSCAN clustering parameter optimization of the present invention. Figure 7In it, A shows the clustering results of the point cloud with 220 true silique numbers under different DBSCAN parameters, and B shows the clustering results of the point cloud with 173 true silique numbers under different DBSCAN parameters. Detailed implementation mode

[0032] The following is described in conjunction with the accompanying drawings and a specific embodiment.

[0033] To make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present invention. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the present invention to be protected, but merely represents the selected embodiments of the present invention.

[0034] It should be noted that similar reference numerals and letters represent similar items in the following drawings. Therefore, once an item is defined in one drawing, it may not be further defined and explained in subsequent drawings.

[0035] Embodiment 1 The rapeseed silique segmentation network model of this embodiment is designed based on an improved encoding-decoding structure. As Figure 1 shown, it includes the following core modules: Encoding layer: It consists of a grouped sampling module and a backbone feature extraction module (KGL-PointNet) for multi-scale feature extraction; Decoding layer: It consists of an interpolation module and a local feature transformation module for restoring the point cloud resolution and enhancing local semantic information.

[0036] Grouped sampling module: Generate centroid points based on farthest point sampling (FPS), select neighboring points within a preset radius through the Ball Query algorithm to construct local neighborhoods, and form a set of regions covering the spatial structure of the point cloud; if the number of neighboring points is insufficient, repeat the center point to make up for it to ensure the integrity of the local area. This mechanism provides a multi-scale spatial association basis for subsequent hierarchical feature learning.

[0037] Backbone feature extraction module (KGL-PointNet): KGL-PointNet, as the backbone feature extraction network designed and proposed by the present invention, realizes three improvements: First, an anti-bottleneck KAN convolution is proposed (the structure is shown in Figure 3Replace some of the original convolutional layers, achieving more efficient geometric feature extraction with fewer parameters compared to the standard KAN convolution.

[0038] Secondly, design the GLFN feature modulation block, which fuses global-local spatial attention (GLS Attention) and the partial convolution network (PCFN) to jointly optimize local details and global context representations; finally, introduce the ContraNorm contrast normalization module, which uses contrast learning to constrain the feature distribution, suppress noise, and alleviate the problems of dimensional collapse and over-smoothing during training.

[0039] In the KGL-PointNet network, the input point cloud per batch is (D + 3, K, N), where the feature dimension is D, each sampling point selects K neighboring points, and the total number of sampling points is N. After passing through the inverse bottleneck KAN convolution, the dimension of the input point cloud is compressed, and the output is (D, K, N). The GLFN feature modulation block consists of GLS attention and PCFN. Through normalization of the intermediate connection process, the point cloud data is feature-modulated by GLFN to obtain a shallow output. Subsequently, through two layers of 1 ×1 convolution, the dimension of the point cloud is expanded, and the output result is (D', K, N). Finally, the output result (D', N) is obtained through max-pooling.

[0040] The KAN convolution can be expressed as follows: The convolution kernel consists of a set of unary non-linear functions. Assume that the input point cloud per batch is , where D is the feature dimension, K is the number of neighboring points, and N is the number of sampling points. At this time, the KAN convolution with a kernel size of k can be defined as: , where, is a unary non-linear learnable function, and the specific form of this function is: , where, and are trainable parameters used to control the overall scale of the function, is a spline function, and by default, it is a B-spline basis function.

[0041] This embodiment proposes an inverse bottleneck KAN convolution. First, use the Gaussian kernel radial basis function to replace the B-spline basis function. The Gaussian kernel radial basis function is defined as follows: where, represents the radial distance, is the standard deviation, which is used to control the width of the Gaussian function.

[0042] Secondly, refer to Figure 3, in this embodiment, an anti-bottleneck structure is designed, which includes two parallel branches: the left branch first applies the SiLU activation to the input features, and then performs a K×K convolution with the K value set to 3; the right branch first uses a 1×1 convolution to expand the number of channels by 4 times, then applies the Gaussian kernel radial basis function, and finally uses a 1×1 convolution to restore the original number of channels. This structure similar to a single-layer encoding-decoding enhances the feature expression ability while reducing the number of parameters.

[0043] See Figure 2 , in this embodiment, a GLFN feature modulation block is designed to strengthen the multi-scale feature extraction ability of the network. The GLFN feature modulation block consists of GLS attention and PCFN: The GLS attention module includes global spatial attention ( attention) and local spatial attention ( attention), which can effectively enhance features at different scales. A tensor with features is evenly split into and and , and after being processed by 1×1 convolution, they are respectively input into and attention modules, and finally merged and output through another layer of 1×1 convolution: where is the output tensor; The specific structures of attention and attention are as shown in Figure 4 ; a. Attention module; this module focuses on the long-range relationships of points in the point cloud in the spatial dimension, complementing the local spatial attention. The specific calculation process is as follows: where is the global spatial attention operation, represents 1×1 convolution, is matrix multiplication, and MLP includes two fully connected layers, a ReLU activation function, and a normalization layer.

[0044] b. Attention module; This module is good at enhancing the feature representation of the local region of interest in the point cloud, especially improving the ability to depict the details of small objects. Its calculation process is: where consists of three layers of 1×1 convolution and one layer of 3×3 depth convolution. For pointwise multiplication, this design efficiently integrates local spatial information with fewer parameters.

[0045] PCFN is an efficient local convolutional feed-forward network designed to optimize features, reduce noise, and fuse local and global information to improve overall performance. See Figure 2 , first PCFN receives the input features , and realizes the interaction between channels through 1×1 convolution and the GELU activation function. Subsequently, the output hidden features are divided into two parts , where is encoded with local context information through 3×3 convolution and GELU activation. Finally, the processed and are concatenated, and the features are further fused through another 1×1 convolution, and the number of hidden channels is restored to the original dimension. This process can be expressed as: .

[0046] This embodiment introduces the contrast learning-based normalization method ContraNorm and applies it to the field of 3D point clouds, effectively preventing dimensional collapse and over-smoothing, and enhancing the model's representation ability for complex geometric structures to a certain extent; , In the formula represents the input feature matrix, is the transpose of the input feature matrix, is the step parameter, is the update amplitude control parameter. Softmax represents the softmax normalization of the similarity of the feature representation matrix to ensure a reasonable distribution of the similarity between feature vectors. The subsequent LN operation performs standard layer normalization to ensure the numerical stability of the features.

[0047] In the decoding stage, interpolation is used to gradually restore the sparse high-level semantic features to the resolution of the original point cloud, restoring the spatial distribution information. The local feature transformation introduces an additional MLP to perform non-linear transformation on the fused features, enhancing the local semantic expression ability, thereby improving the classification accuracy of each point.

[0048] Overall, the role of the decoding layer is to accurately restore the local details while retaining the global semantics, and complete the semantic segmentation at the point level.

[0049] On the Ubuntu 22.04 operating system (NVIDIA RTX 4090 graphics card), the rapeseed silique segmentation model based on the improved PointNet++ was trained using the training set. The total number of iterations of the model was set to 300, the initial learning rate was 0.01, the learning rate decay factor was 0.7, each batch of training data contained 4,096 points, the batch size was 8, and the AdamW optimizer was used. After training, the test set was input into the model to obtain the prediction results, and the widely recognized evaluation metrics for point cloud semantic segmentation were used to evaluate the model performance during the training and testing phases; In the experiment, we used accuracy (Acc) and intersection over union (IoU) to evaluate the segmentation effect of this method on the rapeseed dataset. Specifically, let the row index of the confusion matrix be the true class and the column index be the predicted class, then is the number of points with true class predicted as class , is the number of points with true class predicted as class , is the number of points with true class predicted as class , C represents the total number of classes, then the overall accuracy (OAcc), class accuracy (Acc), and mean class accuracy (mAcc) are defined as follows: In addition, the calculation formulas for intersection over union (IoU) and mean intersection over union (mIoU) are as follows: To verify the advantages of the rapeseed silique segmentation model based on the improved PointNet++, several common deep learning models were compared in this embodiment, including PointNet, PointNext (including two variants, L and XL), and PointVector.

[0050] Table 1 Model comparison results: To evaluate the lightweight of the model simultaneously, the parameter quantity (Params) index was introduced in this embodiment. As can be seen from Table 1, the rapeseed silique segmentation model based on the improved PointNet++ outperformed all benchmark models comprehensively, demonstrating excellent segmentation performance. Specifically, this model achieved 90.19% mIoU, 95.45% mAcc, and 96.25% OAcc, which were 6.32, 7.26, and 2.17 percentage points higher than the second-ranked PointNext-XL, respectively.

[0051] Although the number of parameters of the segmentation model is slightly higher than that of the baseline model PointNet++, the significantly improved segmentation accuracy of the segmentation model demonstrates a breakthrough performance advantage. More importantly, the segmentation model not only surpasses all SOTA models in terms of segmentation accuracy but also maintains the lowest number of parameters relative to SOTA models, proving that the segmentation model has the ability to achieve high-precision segmentation under lightweight conditions in complex plant scenarios. This excellent balance between model complexity and segmentation accuracy highlights the outstanding computational efficiency and practical application potential of the segmentation model.

[0052] Table 2 Results of ablation experiments: The three core modules of inverse bottleneck KAN convolution, GLFN, and ContraNorm were separately and combinatorially integrated into the baseline PointNet++ model, and the independent effects and combined efficacy of each module were systematically evaluated. As shown in Table 2, each module can independently contribute to performance improvement, and their combinatorial integration ultimately obtained the optimal performance indicators.

[0053] Table 3 Quantization results of inverse bottleneck KAN convolution: To further evaluate the performance advantage of the inverse bottleneck KAN convolution, a comparative experiment of three convolution structures was conducted on the first layer of the PointNet++ feature extraction network: the basic KAN convolution, FastKAN convolution, and the inverse bottleneck KAN convolution proposed in this embodiment were used to replace the original traditional convolution respectively. As can be seen from Table 3, although the baseline model has the fewest number of parameters, its segmentation accuracy is limited. While the KAN convolution, FastKAN convolution, and inverse bottleneck KAN convolution all perform well in the rapeseed pod and stem segmentation task, the first two have a higher number of parameters while improving the segmentation accuracy because simply replacing the basis function still introduces a large number of parameters. In contrast, the inverse bottleneck KAN convolution not only significantly reduces the number of model parameters but also has a more significant improvement effect on segmentation accuracy.

[0054] Example 2 The rapeseed pod segmentation and counting system and counting method of this embodiment are based on the improved segmentation network model described in Example 1, combined with three-dimensional reconstruction and clustering algorithms, to achieve full-process automation from image acquisition to pod counting. The system process is as Figure 5 shown, including the following core modules: Image acquisition module and image acquisition: Obtain rapeseed plant images by multi-view shooting; Three-dimensional reconstruction module and three-dimensional reconstruction: Generate dense point clouds based on SfM and Nerfacto; Data processing module and data processing: Point cloud preprocessing and enhancement; Segmentation model module and segmentation model: The rapeseed pod segmentation network model described in Embodiment 1; Pod counting module and pod counting: DBSCAN clustering and post-processing.

[0055] The specific process is as follows: S1. Image acquisition and picture input Collect rapeseed video data during the pod stage, and extract frames from the video to obtain a corresponding rapeseed picture set.

[0056] Select 50 rapeseed plants during the pod stage for video acquisition. Fix each plant sample on an electric rotating platform (rotation speed: 1 r / min), use and fix a smartphone for continuous shooting. The distance between the phone and the rapeseed potted plant is 5 meters, the acquisition time for each single plant is 2 minutes, and the turntable adopts a uniform rotation mode to ensure the imaging stability of the plant.

[0057] Transfer the videos of the 50 rapeseed plants during the pod stage from the phone to a computer folder, and use the FFmpeg tool to extract frames from each video at a rate of 1 frame per second. Extract 120 pictures (format:.jpg) for each rapeseed plant, and finally generate 50 sets of picture data, which are all saved in the local folder.

[0058] S2. 3D reconstruction Based on monocular vision technology, perform sparse reconstruction on the rapeseed picture set through the Structure from Motion (SfM) algorithm, and use Neural Radiance Field (NeRF) to further complete dense reconstruction.

[0059] Sparse reconstruction: Input each group of rapeseed pictures (a total of 50 groups, 120 pictures in each group) into the COLMAP software in turn, and complete sparse reconstruction through the SfM algorithm. The specific process is divided into the following two steps: a. Feature extraction and matching: COLMAP extracts the SIFT feature points of each rapeseed picture, and establishes the corresponding relationship between images from different perspectives through feature matching; b. Sparse reconstruction: Based on the matching results, use the incremental SfM algorithm to estimate the camera pose and sparse 3D point cloud; Traditional dense reconstruction based on monocular vision technology mainly relies on multi-view stereo vision (MVS) technology. Although MVS technology has the advantages of low equipment cost and high reconstruction quality, it has an obvious bottleneck in computational efficiency - dense reconstruction of a single point cloud often requires several hours of processing time. Recently, NeRF technology has achieved a technological breakthrough in point cloud reconstruction through implicit neural representation. Existing studies have shown that in agricultural scenarios, the reconstruction accuracy of NeRF reaches the level of MVS technology. As an improved model of NeRF, Nerfacto not only significantly shortens the training time, but also further improves the geometric fidelity. Based on these technical advantages, this embodiment uses Nerfacto to perform dense reconstruction of rapeseed plants. The specific implementation process includes the following two key steps: a. NeRF model training: The sparse point cloud, camera pose and original multi-view images generated by COLMAP are input into the Nerfacto model, and the implicit scene representation is optimized through differentiable volume rendering to model the geometry and appearance of rapeseed; b. Dense point cloud extraction: After training convergence, the dense point cloud of rapeseed is generated by sampling from the density field of NeRF, and low-confidence points are removed through threshold filtering to retain the high-precision surface structure.

[0060] S3. Data processing The original point cloud of rapeseed is preprocessed and enhanced, and the processed point cloud is labeled to obtain the point cloud segmentation dataset of rapeseed in the silique stage, which is divided into training set and test set.

[0061] Point cloud preprocessing: The dense point cloud contains rapeseed plants, culture dishes, turntables, and background noise points. To accurately segment rapeseed siliques, CloudCompare software was first used to manually remove interfering components such as culture dishes and turntables, and then radius-based statistical filtering was used to remove abnormal noise points. Then, uniform downsampling was used to compress the point cloud data of each rapeseed plant to 20,000-30,000 points while retaining key morphological features. This operation significantly reduced the data size while ensuring a balance between model accuracy and computational efficiency.

[0062] Point cloud enhancement; see Figure 6 In order to improve the scale and diversity of the NeRF rapeseed point cloud dataset, the following enhancement methods were used in turn: random translation (±0.1m), random flipping (50% probability), random cropping (retaining 70%), random rotation (0-360°), Gaussian noise (1mm standard deviation), etc. Through this enhancement strategy, the rapeseed dataset was expanded from 50 samples to 220, effectively simulating key features such as plant displacement, multi-angle observation, leaf occlusion and equipment noise in real environments, significantly improving the generalization ability and robustness of the segmentation model.

[0063] The rapeseed point cloud is labeled through the CloudCompare software. The label value of the stem point cloud is set to 0, and the label value of the silique point cloud is set to 1. The labeled point cloud file is constructed into a rapeseed point cloud segmentation dataset. The dataset is randomly divided into a training set and a test set at a ratio of 8:2 to ensure the uniformity of data distribution and the objectivity of model evaluation.

[0064] S4. Semantic segmentation As described in Example 1.

[0065] S5. Silique counting Apply DBSCAN clustering to the rapeseed point cloud predicted by the trained network model to obtain the silique quantity result.

[0066] Perform threshold filtering on the point cloud test result generated by the segmentation model, and only retain the points with a predicted label of 1 as the rapeseed silique point cloud. Use the DBSCAN algorithm to cluster it to count the number of rapeseed siliques, and evaluate the silique prediction effect through widely recognized clustering accuracy metrics.

[0067] The DBSCAN algorithm is a density-based clustering algorithm. Its core idea is to aggregate the points in the high-density region into clusters through density reachability and can automatically identify noise points. Its main process includes the following steps: a. Parameter setting; is the neighborhood radius, which defines the scale of density, and MinPts is the density threshold, that is, the minimum number of neighbors required for a point to become a core point.

[0068] b. Neighborhood query. For each unvisited point in the dataset : Find all the points whose distance does not exceed , and denote the set as .

[0069] c. Core point determination; If , then mark as a core point, otherwise, first mark it as a border point or a noise point, depending on whether it will be included in a certain cluster later.

[0070] d. Cluster expansion; For each core point , if it has not been included in any cluster, create a new cluster , and add to , then traverse all the neighbor points of . If is also a core point and has not been visited, then add its neighborhood All the points in are also added to the list to be expanded, and is added to the cluster .

[0071] e. Mark noise; After each round of traversal, all points not included in any cluster are marked as noise points.

[0072] To determine the optimal DBSCAN parameter combination for rapeseed silique clustering, the grid search method was used to systematically evaluate the influence of two key parameters on the clustering performance. The parameter search ranges were set as ∈[0.05, 0.20] and MinPts ∈[4, 15], with step sizes of 0.05 and 3 respectively. As Figure 7 shown, when the parameter combination ( = 0.10, MinPts = 7) was adopted, the model achieved the optimal clustering performance in the regions with significantly different silique distribution densities. This study found that: a. A smaller value can effectively distinguish adjacent silique clusters with close spacing, but may misjudge some siliques as noise due to uneven point cloud density; b. A moderate MinPts value can effectively suppress the over-segmentation phenomenon caused by point cloud noise or residual segmentation artifacts, and can also avoid losing effective silique point cloud data due to too high a threshold.

[0073] To evaluate the performance of the clustering stage, three indicators were introduced in this embodiment: mean recognition precision (mPr), mean recognition precision (mPr), mean absolute percentage error (MAPE), and coefficient of determination . Among them, is the total number of samples, is the true number of siliques of the th sample, is the mean value predicted as the true number, and is defined as follows: .

[0074] Table 4 DBSCAN clustering results: As can be seen from Table 4, the clustering accuracies are that the mean recognition accuracy is 97.42%, MAPE = 2.47%, and R² = 0.989. These indicators jointly verify the accuracy of the clustering task from three dimensions of matching accuracy, error range, and model goodness of fit, effectively identifying and distinguishing individual siliques of rapeseed plants, and achieving a relatively high recognition accuracy.

[0075] The rapeseed silique segmentation network model and automatic counting method proposed by the present invention based on the improved PointNet++ are superior to a variety of advanced models in the rapeseed segmentation task, achieving an mIoU of 90.19%, an mAcc of 95.45%, and an OAcc of 96.25%. Moreover, the model parameters are only 5.72M, achieving an excellent balance between accuracy and the number of parameters, showing a good practical prospect. The model has broad application potential in plant phenotype research and is particularly suitable for the precise segmentation of rapeseed organs.

[0076] Embodiment 3 In this embodiment, the computer-readable storage medium is a USB flash drive, which uses the FAT32 file system format, has a storage capacity of 128GB, and both the read and write speeds reach more than 100MB / s, ensuring the rapid storage and loading of the program.

[0077] The computer program stored in the USB flash drive is a complete software system for implementing the rapeseed silique counting method described in Embodiment 2.

[0078] The above are only the preferred embodiments of the present invention. It should be noted that the above preferred embodiments should not be regarded as limitations of the present invention. The protection scope of the present invention should be determined by the scope defined by the claims. For those of ordinary skill in the art, without departing from the spirit and scope of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as within the protection scope of the present invention.

Claims

1. A rapeseed silique segmentation network model, characterized in that: It includes encoding layer and decoding layer, where: The coding layer includes a group sampling module and a backbone feature extraction module; The decoding layer includes an interpolation module and a local feature transformation module; The backbone feature extraction module realizes multi-scale feature extraction through the following structure: a. Anti-bottleneck KAN convolution, which contains two parallel branches. The left branch performs SiLU activation and K×K convolution in sequence, and the right branch performs 1×1 convolution to expand the number of channels, Gaussian kernel radial basis function transformation and 1×1 convolution to restore the number of channels; b. GLFN feature modulation block, which consists of a global-local spatial attention module and a partial convolutional feed-forward network to fuse local details and global context information; c. The ContraNorm normalization module based on contrastive learning suppresses noise and alleviates dimensionality collapse by constraining feature distribution.

2. The rapeseed silique segmentation network model according to claim 1, characterized in that: The Gaussian kernel radial basis function of the anti-bottleneck KAN convolution is defined as: in, represents the radial distance, is the standard deviation, which is used to control the width of the Gaussian function.

3. The rapeseed silique segmentation network model according to claim 1, characterized in that: In the GLFN feature modulation block: The global-local spatial attention module evenly splits the input features into two parts, processes them through the global spatial attention module and the local spatial attention module respectively, and finally merges the outputs; The partial convolutional feed-forward network encodes local context information through 3×3 convolution, and concatenates it with the original features and fuses it through 1×1 convolution.

4. The rapeseed silique segmentation network model according to claim 1, characterized in that: The calculation formula of the ContraNorm normalization module is: in, is the input feature matrix with a dimension of D×N, where D is the feature dimension and N is the number of point clouds; is the transpose of the input feature matrix; s is the step size parameter, which is used to control the amplitude of feature update; To update the amplitude control parameters, LN is standard layer normalization; It is the output feature matrix used to suppress noise and alleviate dimensionality collapse.

5. A rapeseed silique segmentation system, characterized in that: Includes the following modules: a. An image acquisition module for acquiring continuous images of rapeseed plants through a rotating platform and multi-viewing angle shooting; b. 3D reconstruction module, which generates dense 3D point clouds of rapeseed plants based on the structure-from-motion algorithm and neural radiation field technology; c. Data processing module, used to preprocess, enhance and label point clouds and construct segmentation data sets; d. The rapeseed silique segmentation network model according to any one of claims 1 to 4, used for semantic segmentation of the preprocessed point cloud; e. Silique counting module, which performs clustering statistics on the segmented silique point cloud based on the DBSCAN clustering algorithm.

6. The rapeseed silique segmentation system according to claim 5, characterized in that: The 3D reconstruction module uses the Nerfacto model to reconstruct dense point clouds, specifically including: The sparse point cloud and camera pose are generated by the motion recovery structure algorithm, which is input into the Nerfacto model to optimize the implicit scene representation, and the dense point cloud is generated by sampling from the density field.

7. The rapeseed silique segmentation system according to claim 5, characterized in that: The point cloud preprocessing of the data processing module includes: Noise points are removed through radius statistical filtering, and the point cloud of a single plant is compressed to 20,000 to 30,000 points through uniform downsampling; The enhancements include random translation, flipping, cropping, rotation and adding Gaussian noise.

8. A method for counting rapeseed siliques, characterized in that: The following steps are involved: S1. Collect multi-view images of rapeseed plants and reconstruct 3D point clouds using structure-from-motion algorithm and neural radiation field technology; S2. Preprocess, enhance and annotate the point cloud and construct training and test sets; S3. Using the segmentation network model according to any one of claims 1 to 4 to perform semantic segmentation on the point cloud, distinguishing between siliques and stem point clouds; S4. Apply the DBSCAN clustering algorithm to the segmented silique point cloud to count the number of siliques; Among them, the parameters of the DBSCAN algorithm are optimized through grid search, and the optimal parameter combination is ε=0.10 and MinPts=7.

9. The method for counting rapeseed siliques according to claim 8, characterized in that: The performance of the DBSCAN clustering is evaluated by the following metrics: Average recognition accuracy, mean absolute percentage error, and coefficient of determination, where the mean absolute percentage error is defined as: is the total number of samples, is the actual number of siliques, For the predicted quantity.

10. A computer-readable storage medium, characterized in that: A computer program is stored, and when the program is executed by a processor, the rapeseed silique counting method according to claim 8 or 9 is implemented.

Citation Information

Patent Citations

  • Diaphragm flaw detection method based on reverse bottleneck structure deep convolutional network

    CN110415238A

  • Hyperspectral image segmentation method based on non-local feature fusion

    CN113743450A

  • Rape phenotype parameter extraction method and system based on three-dimensional point cloud

    CN116704497A

  • Multi-scale multi-perception real-time image segmentation method and system, terminal and medium

    CN117576118A

  • Polyp image segmentation method and system based on parallel double-encoder feature fusion

    CN118781351A