Rice germplasm yield prediction method and system based on panicle type classification and growth stage alignment

CN122435368BActive Publication Date: 2026-09-18HUAZHONG AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610902237.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-06-23
Publication Date
2026-09-18
Estimated Expiration
2046-06-23

AI Technical Summary

Technical Problem

[0005]此外,虽然深度学习算法已广泛应用于稻穗检测,但真实的育种大田环境十分复杂,存在光照不均、背景复杂及器官遮挡严重等问题

Benefits of technology

[0018] (1) Overcoming spatiotemporal noise and significantly improving yield estimation accuracy. This invention innovatively proposes a yield estimation strategy based on growth period alignment and panicle type classification. A multi-label classification model is used to lock in the "maturity period" as a unified yield estimation benchmark, eliminating temporal noise caused by phenological asynchrony. Simultaneously, independent yield models are established for "compact" and "spreading" rice varieties, eliminating feature conflicts caused by spatial morphology. After implementing this strategy, the mean absolute percentage error (MAPE) of the predicted yield of compact and spreading rice varieties decreased to 10.51% and 9.83%, respectively. Ablation experiments show that without growth period alignment, the error would increase by approximately 7.87%, and without panicle type classification, the explanatory power of the model would decrease by more than 60%.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122435368B_ABST
    Figure CN122435368B_ABST
Patent Text Reader

Abstract

This invention discloses a method and system for predicting rice germplasm yield by aligning panicle type classification with growth stage. First, time-series images of paddy rice plots are acquired. The ST-YOLO single-panicle detection model is used to identify the rice growth stage and extract single-panicle images. These single-panicle images are then input into the PP-LCNet multi-label classification model to simultaneously identify panicle type and growth stage. Maturity data are filtered and processed by panicle type, achieving growth stage alignment and panicle type classification, eliminating spatiotemporal noise introduced by phenological asynchrony and morphological differences. Geometric features of the detection box are extracted at the single-panicle scale, and features such as quantity, color, texture, spectrum, and phenology are extracted at the population scale using the Swin Transformer segmentation model. Machine learning algorithms are then used to establish yield prediction models for different panicle types. This invention effectively solves the complex spatiotemporal noise interference problem in large-scale breeding experiments, achieving high-throughput, accurate, and lossless rice yield prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of smart agricultural information technology applications, artificial intelligence and deep learning, agricultural automation, and agricultural bioinformatics, specifically to a method and system for predicting rice germplasm yield by aligning panicle type classification with growth period. Background Technology

[0002] Rice is the world's most important food crop, and its genetic improvement and the breeding of superior varieties are the main drivers for increasing yields. In modern breeding experiments, researchers increasingly rely on large-scale screening of germplasm resources. Therefore, conducting high-throughput, accurate, and non-destructive yield estimation is the foundation for guiding precise field management and the screening of high-quality germplasm.

[0003] Traditional agronomic yield estimation methods mainly rely on manual field surveys and destructive sampling, which are labor-intensive, inefficient, and susceptible to human error, failing to meet the needs of large-scale monitoring. In recent years, non-destructive yield estimation methods based on remote sensing technology have become mainstream. However, most existing remote sensing yield estimation methods rely on empirical regression by extracting spectral features of crop canopy, which faces two major limitations in large-scale breeding trials: first, spatially, the canopy is treated as a homogeneous whole, failing to distinguish specific yield-contributing organs, thus introducing noise from leaves and the background; second, temporally, the asynchronous phenological development between different breeding materials is ignored, failing to align growth stages, thus introducing temporal noise.

[0004] To address these shortcomings, research needs to shift its focus from general canopy analysis to precise panicle characteristic analysis, concentrating on organs directly related to yield. However, using panicles for direct yield estimation faces significant challenges. Spatially, panicle morphology exhibits diversity; the geometry and texture of compact and spreading panicles are completely opposite. Without classification, ambiguous characteristics can easily arise, affecting the accuracy of yield estimation. Temporally, rice phenotypic characteristics change drastically with physiological processes, and significant phenological differences exist among field breeding populations. Without alignment of growth stages, temporal noise will inevitably be introduced.

[0005] Furthermore, although deep learning algorithms have been widely applied to rice panicle detection, the real breeding field environment is extremely complex, with problems such as uneven lighting, complex backgrounds, and severe organ occlusion. Most existing mainstream convolutional neural networks (CNNs) are limited by local receptive fields, making it difficult to fully capture global contextual information, and exhibit poor robustness under extreme lighting conditions such as severe underexposure or strong shadows. At the same time, current agricultural phenotyping systems are often fragmented, limited to single tasks, and lack an integrated technical framework capable of simultaneously handling differences in panicle type and phenological characteristics. Summary of the Invention

[0006] To address the aforementioned issues, this invention proposes a method and system for predicting rice germplasm yield by aligning panicle type classification with growth period. By simultaneously identifying rice growth period and panicle type, aligning maturity period first, and then independently modeling panicle type, the method effectively eliminates spatiotemporal noise interference caused by asynchronous phenology and morphological differences, significantly improving the accuracy and robustness of large-scale rice germplasm yield prediction.

[0007] The present invention discloses a method for predicting rice germplasm yield by aligning panicle type classification with growth period, comprising the following steps: Step S1: Obtain time-series images of the rice experimental field, and obtain the cell images of each sampling plot after preprocessing; Step S2: Use the single ear detection model to detect rice ears in the image of the plot. If rice ears are detected, output the single ear image and the bounding box information of the single ear. Step S3: Input the single ear image into the multi-label single ear classification model, and simultaneously identify and output the growth period label and ear type label of the single ear; wherein, the growth period label includes at least the heading and flowering period and the maturity period, and the ear type label includes at least the compact type and the spreading type. Step S4: Based on the classification results of step S3, implement the growth period alignment and ear type classification strategy, filter out the data in the maturity stage, and divide them into two independent data subsets: compact type and spreading type. Step S5: Extract multi-scale features from the filtered data, including single-ear scale features and population scale features. Step S6: Using the multi-scale features extracted in step S5 as input, construct independent yield prediction models for compact and spread rice varieties respectively, to predict rice yield.

[0008] Furthermore, in step S2, the single ear detection model is the ST-YOLO single ear detection model, which is constructed by introducing the Swin Transformer module into the backbone network of the YOLO network; the ST-YOLO single ear detection model adopts a P3-P5 multi-scale output structure and retains the SPPF module and C2PSA module, and the detection head adopts continuous upsampling and feature splicing to achieve multi-scale feature fusion.

[0009] Furthermore, in step S3, the multi-label single-ear classification model uses the PP-LCNet network as the backbone network to simultaneously identify and output the growth period label and ear type label of a single ear; the ear type and growth period categories output by the model are locally mutually exclusive labels, and the confidence comparison method is used to compare the confidence of the pairwise mutually exclusive labels, and the label with the higher confidence is taken as the final classification result.

[0010] Further, in step S5, the single-ear scale feature extraction refers to: extracting the geometric features of the single-ear frame from the single-ear boundary frame information described in step S2. The geometric features of the single-ear frame include one or more of the following: single-ear frame length, single-ear frame width, single-ear frame area, single-ear frame perimeter, and single-ear frame length-to-width ratio.

[0011] Further, in step S5, the group-scale feature extraction includes: The Swin Transformer semantic segmentation model based on the Transformer architecture was used to perform pixel-level segmentation of rice ears in images of mature plots to generate rice ear mask images. Based on the rice ear mask image, population features are extracted, including one or more of the following: rice ear quantity features, morphological features, spectral features, color features, and texture features. The phenological characteristics are obtained by combining the classification results of step S3. The phenological characteristics include the number of days for heading and flowering and the number of days for maturity.

[0012] Furthermore, in step S6, the machine learning algorithm employs one of random forest, support vector regression, extreme gradient boosting, or backpropagation neural network.

[0013] Based on the same inventive concept, this invention also designs a system for predicting rice germplasm yield by aligning panicle type classification with growth period, comprising: The data acquisition module is used to acquire time-series images of rice experimental fields, and after preprocessing, obtains the cell images of each sampling plot. The single ear detection module is used to detect rice ears in the image of the plot using a single ear detection model. If a rice ear is detected, the single ear image and the bounding box information of the single ear are output. A multi-label classification module is used to input the single ear image into a multi-label single ear classification model, and simultaneously identify and output the growth period label and ear type label of the single ear; wherein, the growth period label includes at least the heading and flowering period and the maturity period, and the ear type label includes at least the compact type and the spreading type; The data filtering module is used to perform a growth period alignment and ear type classification strategy based on the classification results of the multi-label classification module, filter out the data in the maturity stage, and divide them into two independent data subsets: compact and spreading. The feature extraction module is used to extract multi-scale features from the filtered data, including single-ear scale features and population scale features. The yield prediction module is used to construct independent yield prediction models for compact and spread rice varieties, respectively, using the multi-scale features extracted by the feature extraction module as input, so as to predict the yield of rice.

[0014] Furthermore, the single-ear detection module adopts the ST-YOLO single-ear detection model; the multi-label classification module uses the PP-LCNet network as the backbone network, and uses the confidence comparison method to compare the confidence of mutually exclusive labels, and takes the label with the higher confidence as the final classification result.

[0015] Furthermore, the feature extraction module includes: The single-ear scale feature extraction unit is used to extract the geometric features of the single-ear frame from the single-ear bounding box information output by the single-ear detection module; The group-scale feature extraction unit is used to perform pixel-level segmentation of rice ears in mature plot images using the Swin Transformer semantic segmentation model, generate rice ear mask images, extract group features based on the mask images, and obtain phenological features by combining the classification results.

[0016] Based on the same inventive concept, the present invention also designs a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements a method for predicting rice germplasm yield by aligning panicle type classification with growth period.

[0017] Compared with the prior art, the present invention has significant positive technical effects, which are specifically manifested in the following aspects.

[0018] (1) Overcoming spatiotemporal noise and significantly improving yield estimation accuracy. This invention innovatively proposes a yield estimation strategy based on growth period alignment and panicle type classification. A multi-label classification model is used to lock in the "maturity period" as a unified yield estimation benchmark, eliminating temporal noise caused by phenological asynchrony. Simultaneously, independent yield models are established for "compact" and "spreading" rice varieties, eliminating feature conflicts caused by spatial morphology. After implementing this strategy, the mean absolute percentage error (MAPE) of the predicted yield of compact and spreading rice varieties decreased to 10.51% and 9.83%, respectively. Ablation experiments show that without growth period alignment, the error would increase by approximately 7.87%, and without panicle type classification, the explanatory power of the model would decrease by more than 60%.

[0019] (2) Significantly enhanced robustness of rice panicle detection in complex field environments. The ST-YOLO detection model proposed in this invention breaks the limitation of the local receptive field of traditional convolutional kernels by introducing the Swing Transformer module. Under extreme field lighting conditions such as severe underexposure, strong shadows, and high-density overlap, the model can still accurately locate rice panicles by capturing global contextual information, achieving an average accuracy (mAP@50) of 0.918 and an F1 score of 0.880, significantly reducing the false negative and false positive rates, and laying a high-quality data foundation for downstream feature extraction.

[0020] (3) It achieves complementary advantages and precise expression of feature dimensions. Starting from two spatial scales, single panicle and population, this invention uses the Swin Transformer network for high-quality semantic segmentation, integrating the geometric attributes of single panicles with multi-dimensional features of population, such as phenology, quantity, morphology, spectrum, color, and texture. This multi-scale fusion mechanism can accurately capture the core factors that determine rice yield. Feature contribution rate analysis shows that, in addition to the common dominant factor of population quantity, the two panicle types show significant differences in their dependence on multi-scale features: compact rice is highly dependent on the geometric features at the single panicle level (contribution rate of 24.98%); while for spreading rice, due to the ease with which mutual occlusion occurs, the contribution rate of single panicle geometric features decreases (16.25%), and its model instead relies on the texture features (10.05%) and morphological features (8.85%) at the population level. The macroscopic spatial arrangement pattern effectively compensates for the information loss caused by single panicle occlusion.

[0021] (4) Improved the high-throughput and non-destructive level of breeding screening. This invention integrates ST-YOLO detection, PP-LCNet multi-label classification, and multi-scale feature modeling to construct a complete automated analysis framework, effectively replacing the traditional destructive manual seed selection process. Quantitative verification results show that the optimal yield prediction model (BP neural network) constructed by this framework achieved prediction determination coefficients (R²) of 0.71 and 0.75 for compact and spread-type rice yields, respectively, in the current season's validation set, with mean absolute percentage error (MAPE) reduced to 10.51% and 9.83%, respectively. When dealing with independent test data across years, the MAPE can still be stably controlled at 18.19% and 16.33%. The above quantitative indicators show that this invention can efficiently achieve digital screening of high-yield germplasm resources in large-scale rice breeding populations with a stable and controllable error level. Attached Figure Description

[0022] Figure 1 The overall flowchart of a method and system for predicting rice germplasm yield by aligning panicle type classification with growth period provided by the present invention.

[0023] Figure 2 This is a schematic diagram of the technical process for identifying the rice growth period and panicle type in an embodiment of the present invention.

[0024] Figure 3 This is a schematic diagram of the network architecture of the ST-YOLO single-ear detection model used in this embodiment of the invention.

[0025] Figure 4 This is a schematic diagram of the network architecture of the PP-LCNet multi-label single-ear classification model used in the embodiments of the present invention.

[0026] Figure 5This is a schematic diagram of the technical process for multi-scale feature extraction of rice in an embodiment of the present invention.

[0027] Figure 6 This is a schematic diagram of the technical process for estimating rice yield and analyzing key characteristics in an embodiment of the present invention.

[0028] Figure 7 This is a schematic diagram of the contribution rate of the SHAP feature group for different panicle type yield estimation models in this embodiment of the invention.

[0029] Figure 8 This is a schematic diagram illustrating the performance evaluation of rice yield estimation models for different panicle types in embodiments of the present invention.

[0030] Figure 9 This is the Swin Transformer rice ear segmentation network structure in an embodiment of the present invention.

[0031] Figure 10 This is a dataset for identifying the rice growth period and panicle type in an embodiment of the present invention.

[0032] Figure 11 This is a rice ear segmentation dataset in an embodiment of the present invention. Detailed Implementation

[0033] The parameter tuning method proposed in this invention will be further described below with reference to specific embodiments. Those skilled in the art should understand that the following embodiments are only used to explain the technical principles of this invention and are not intended to limit the scope of protection of this invention.

[0034] Example 1 This embodiment provides a method for predicting rice germplasm yield by aligning panicle type classification with growth period. For example... Figure 1 As shown, it includes the following steps: Step 1, Data Acquisition and Preprocessing: A drone platform equipped with a high-definition image sensor was used to acquire high-definition time-series images of the experimental rice plots in the field, and then preprocessed. Specifically, a drone supporting RTK centimeter-level positioning was used with a high-definition full-frame visible light camera (e.g., 45 megapixels) for all-weather data acquisition. During flight operations, the drone observed vertically downwards (90°) at a specific altitude (e.g., 15m) to acquire images at the extremely high ground sampling distance (GSD). To avoid high-contrast shadows caused by direct sunlight, it was preferred to conduct the operation during periods of uniform illumination. After acquiring the images, image processing software was used to convert them to a standard format, stitch them together to generate orthophotos with geographic coordinates, and then crop the orthophotos according to the field plot vector boundaries (ROI) to obtain the time-series images of each experimental plot.

[0035] Step 2, growth stage identification and single ear detection, such as Figure 2As shown: An improved ST-YOLO single-ear detection model is used to detect rice panicles in plot images, distinguishing between the vegetative and reproductive growth stages. The cropped plot images are input into the ST-YOLO single-ear detection model. If no panicles are detected, the rice is determined to be in the vegetative growth stage; if panicles are detected, it is determined to be in the reproductive growth stage, and the detection bounding box for each panicle is output. This task uses a YOLO format dataset, and the rice panicles in the plot images are manually labeled. The labeling results are shown below. Figure 10 As shown, this embodiment uses Labelme software for manual annotation. After data augmentation operations such as rotation, flipping, and Gaussian blur, and with the addition of the public dataset DRPD, the dataset currently contains 10787 images. The training set, validation set, and test set account for 70%, 20%, and 10% respectively. The training set has 7551 images, the validation set has 2158 images, and the test set has 1078 images. This ST-YOLO model is based on the YOLOv11n architecture, and the feature extraction system is constructed by replacing the backbone network with a layered Swin Transformer module. The backbone network is divided into 4 stages. The images are first segmented by the Patch Partition layer and mapped by the Linear Embedding layer before entering the Swin Transformer Block using an offset window mechanism. In the specific engineering implementation, the window size is set to 7×7, the number of Blocks in each stage is [2, 2, 6, 2], and the number of attention heads is [3, 6, 12, 24]. The backbone network retains the hierarchical output characteristics, extracting feature maps from Stage 2, Stage 3, and Stage 4 outputs at downsampling ratios of 8x, 16x, and 32x, respectively, and inputting them into the neck network. The network structure employs a P3-P5 multi-scale output structure, retaining SPPF and C2PSA modules to improve feature extraction robustness. Continuous upsampling and feature concatenation are used in the detection head to achieve multi-scale feature fusion. Detailed network structure is as follows: Figure 3 As shown. After detection, the rice ears in the cell image are cropped according to the bounding box coordinates to obtain a single ear image.

[0036] Step 3, Simultaneous Classification of Ear Type and Growth Stage: The single ear images cropped in Step 2 are input into the PP-LCNet multi-label classification model to identify the specific ear type and growth stage. This dataset was created by manually outlining the rice ear frame using Labelme software, cropping the image content within the rectangle, thus obtaining images of individual rice ears. Each ear was then labeled with multiple labels. Ear type was manually classified based on the angle between the primary branch and the ear neck, into compact and spreading types. Growth stage was classified based on the ear's state and color, into heading and flowering stage and maturity stage. This dataset contains 48,923 images, divided into training, validation, and test sets in a 70:15:15 ratio. The detailed network structure of this classification model is as follows... Figure 4 As shown, this model uses PP-LCNet_x1_0 as the backbone network. Its core architecture consists of stacked BaseBlocks composed of depthwise separable convolutions, and an SE attention module is integrated deep within the network to enhance channel feature learning capabilities. After features are extracted from a single ear image by the backbone network, they are transformed into feature vectors through a global average pooling layer, and then input into a fully connected layer with 4 nodes. The Sigmoid activation function is used to map the output to the [0, 1] interval, corresponding to the probability scores of the four categories: (Compact) (Spreading type) (Heading and flowering period) and (Maturity stage). Since the panicle type and growth stage of the same single panicle are logically locally mutually exclusive, the model uses a "confidence comparison method" for the final decision: within the panicle type branch, comparison... and The size of the spikelets is used, and the one with the higher confidence level is taken as the spikelet type label for that sample; within the fertility period branches, comparisons are made. and The size of the sample is used, and the sample with the higher confidence level is selected as the reproductive period label. During training, a binary cross-entropy loss function is used for optimization, enabling a single model to simultaneously predict multiple attributes, thus improving the system's computational efficiency when processing large-scale germplasm resources.

[0037] Step 4, Maturity Period Image Screening and Growth Stage Dynamic Alignment Strategy: Based on the classification results of Step 3, a growth stage dynamic alignment and panicle type classification strategy is implemented to construct a unified yield prediction benchmark dataset. The growth stage dynamic alignment strategy proposed in this study effectively eliminates characteristic noise introduced by phenological asynchrony among varieties by converting the observation benchmark of different germplasm resources in a large-scale breeding population from absolute time coordinates (such as the number of days after sowing, DAS) to a unified relative physiological development stage (maturity period). During implementation, the system uses the ST-YOLO model to identify and extract single panicle images, and then uses the PP-LCNet model to capture the evolution of the physiological state of the rice panicle, thereby dynamically capturing the nodes when each plot enters the "maturity period," and extracting only the data from this stage for yield prediction, achieving strict alignment along the physiological dimension. The objective basis for classification mainly stems from the differences in the optical, morphological, and geometric characteristics of rice. In the growth stage, the criteria are the color evolution and gravitational posture changes of the panicle. During the heading and flowering stage, the panicle is yellowish-green and upright, while at maturity, the husk turns significantly golden yellow and, due to increased dry matter accumulation, droops under the influence of gravity. In the panicle type dimension, the criteria are the spatial topology and occlusion relationships of the panicle. Compact panicles have tightly packed branches and stalks, resulting in continuous boundaries and less occlusion in images, while spreading panicles have diffuse boundaries in image projection and are prone to overlapping and occlusion with adjacent leaves. Based on this logic, the system strictly removes non-mature stage data and separates the retained mature stage images according to panicle type, constructing independent data subsets for compact and spreading panicles respectively, thereby eliminating spatial feature ambiguities caused by morphological differences.

[0038] Step 5, Multi-scale Feature Extraction: Technical approach as follows Figure 5 As shown, multidimensional feature pools for yield prediction are extracted at both the single-ear and population scales.

[0039] Single ear scale feature extraction: Using the single ear detection bounding box output by the ST-YOLO model in step 2, calculate and extract geometric features such as single ear frame length, single ear frame width, single ear frame perimeter, single ear frame area, and single ear frame length-to-width ratio, which are used to quantify the individual development status and single ear load capacity.

[0040] Group-scale feature extraction: using the Swin Transformer semantic segmentation model based on the Transformer architecture ( Figure 9 The task involves pixel-level segmentation of mature plot images to generate rice ear mask images. This is achieved by manually extracting the rice ear regions using Adobe Photoshop, setting the grayscale value of the target rice ear region (foreground) to 1, and setting the grayscale value of the remaining background regions to 0, thus constructing a binary mask label. The labeling results are shown below. Figure 11As shown, the processed dataset is stored in the standard PASCAL VOC format, and both the original images and their corresponding label images are uniformly cropped to a standard size of 512×512 pixels. The model for this task adopts a symmetric encoder-decoder architecture: the encoder consists of four hierarchical stages, using a patch merging layer for downsampling to construct multi-scale feature maps. At each stage, window-based self-attention and moving window self-attention mechanisms are used to establish long-range dependencies between pixels. The window size is set to 7×7 to balance local details and global contextual information. The decoder fuses features from the corresponding levels of the encoder through linear upsampling and skip connections to compensate for spatial information loss during downsampling. The function of this architecture is to separate complex rice spike organs from the leaf and soil background through pixel-level classification capabilities, providing a foundation for subsequent feature pool construction. Based on the obtained mask images, the following multidimensional features are extracted: 1) Quantitative features: number of rice spikes, projected area of ​​rice spikes, etc. 2) Morphological features: area, perimeter, and fractal dimension of the bounding rectangle of the population, etc. 3) Color and spectral characteristics: First to third order statistical moments (mean, standard deviation, uniformity, etc.) of color extracted from the visible light band, and visible light vegetation indices (such as ExG, NDI, etc.). 4) Texture characteristics: Second order angular moments, contrast, correlation, dissimilarity, homogeneity, etc., extracted from the gray-level co-occurrence matrix (GLCM), reflecting canopy surface roughness and spatial arrangement patterns. 5) Phenological characteristics: Combining the classification results from step 3, calculating the number of days after sowing (DAS) for key growth nodes (such as from heading and flowering to maturity).

[0041] Step 6, Production Prediction Model Construction and Validation: Technical route as follows Figure 6 As shown, machine learning algorithms were used to establish rice yield prediction models for different panicle types. The single panicle and population multi-scale features extracted in step 5 were cleaned, missing values ​​were imputed, and standardized to eliminate dimensional differences. The extracted feature sets were used as input, and the measured rice yield of the corresponding plot was used as the output label. For compact and scattered data subsets, backpropagation (BP) neural networks, random forests (RF), support vector regression (SVR), or extreme gradient boosting (XGBoost) algorithms were applied to train the yield estimation models (training results are shown in the figure). Figure 8 (As shown). Preferably, a BP neural network algorithm is used to determine the optimal hyperparameters through grid search and cross-validation. After training, the predicted yields of compact rice and spreading rice are output. This step can also quantitatively analyze the contribution weights of feature groups at each scale (such as geometric features, quantitative features, texture features, etc.) to the yield of different panicle types using the SHAP value algorithm, in order to analyze the intrinsic agronomic mechanism of yield formation. The SHAP analysis results are shown below. Figure 7 As shown.

[0042] Example 2 This embodiment provides a large-scale rice germplasm resource yield prediction system based on a panicle type classification and dynamic alignment algorithm of growth period, used to execute the method described in Embodiment 1. The system includes the following modules: The data acquisition module is used to acquire time-series images of the rice experimental fields, and after preprocessing, obtains the images of each sampling plot. Specifically, a drone supporting RTK centimeter-level positioning is used to carry a high-definition full-frame visible light camera for all-weather data acquisition. During flight operations, the drone observes vertically downwards, and after acquiring images, image processing software is used to convert the images into a standard format, stitch them together to generate orthophotos with geographic coordinates, and then crop the orthophotos according to the field plot vector boundaries to obtain the time-series images of each experimental plot.

[0043] The single ear detection module is used to detect rice ears in the image of the plot using a single ear detection model. If a rice ear is detected, the single ear image and the bounding box information of the single ear are output.

[0044] Preferably, the single ear detection module adopts the ST-YOLO single ear detection model, which is constructed by introducing the Swin Transformer module into the backbone network of the YOLO network to enhance the ability to capture global context information.

[0045] More preferably, this model is based on the YOLOv11n architecture. The feature extraction system is constructed by replacing the backbone network with a hierarchical Swin Transformer module. The backbone network is divided into four stages, with the number of blocks in each stage being 2, 2, 6, and 2 respectively, and the number of attention heads being 3, 6, 12, and 24 respectively. Furthermore, the ST-YOLO single-ear detection model adopts a P3-P5 multi-scale output structure and retains the SPPF and C2PSA modules. The detection head uses continuous upsampling and feature concatenation to achieve multi-scale feature fusion. After detection, the rice ears in the plot image are cropped according to the bounding box coordinates to obtain single-ear images.

[0046] A multi-label classification module is used to input the single-ear image into a multi-label single-ear classification model, and simultaneously identify and output the growth period label and ear type label of the single ear. The growth period label includes at least the heading and flowering stage and the maturity stage, and the ear type label includes at least compact and spreading types.

[0047] Preferably, the multi-label single-ear classification model uses PP-LCNet_x1_0 as the backbone network. Its core architecture consists of stacked BaseBlocks composed of depthwise separable convolutions, and an SE attention module is integrated in the deep layers of the network to enhance the channel feature learning ability. Specifically, after the single-ear image extracts features through the backbone network, it is transformed into a feature vector through a global average pooling layer, and then input to a fully connected layer with 4 nodes. The Sigmoid activation function is used to map the output to the 0,1 interval, corresponding to the probability scores of four categories: compact, spreading, heading and flowering stage, and maturity stage.

[0048] Since the ear type and growth period of the same single ear are logically mutually exclusive, preferably, the multi-label single ear classification model adopts the "confidence comparison method" for the final decision: in the ear type branch, the probability values ​​of compact type and spreading type are compared, and the one with higher confidence is taken as the ear type label of the sample; in the growth period branch, the probability values ​​of heading and flowering period and maturity period are compared, and the one with higher confidence is taken as the growth period label of the sample.

[0049] The data filtering module is used to implement growth period alignment and panicle type classification strategies based on the classification results of the multi-label classification module, filtering out data in the mature stage and dividing it into two independent data subsets: compact and spreading. Specifically, this module uses the ST-YOLO model to identify and extract single panicle images, and then uses the PP-LCNet model to capture the evolution of the physiological state of rice panicles, dynamically capturing the nodes of each plot entering the "mature stage," and extracting only the data of this stage for subsequent processing. In terms of panicle type, it is classified according to the spatial topology and occlusion relationship of rice panicles: compact rice panicles have tightly packed branches and stalks, with continuous boundaries and less occlusion in the image; spreading rice panicles have a structure that spreads outwards, with diffuse boundaries in the image projection and are very easy to overlap and occlude with adjacent leaves. The system strictly removes non-mature stage data and separates the retained mature stage images according to panicle type, constructing independent data subsets of compact and spreading panicles respectively.

[0050] The feature extraction module is used to extract multi-scale features from the filtered data, including single-ear scale features and population scale features.

[0051] Preferably, the feature extraction module includes a single-ear-scale feature extraction unit and a population-scale feature extraction unit.

[0052] The single-ear scale feature extraction unit is used to extract the geometric features of the single-ear frame from the single-ear boundary box information output by the single-ear detection module. More preferably, the single-ear frame geometric features include one or more of the following: single-ear frame length, single-ear frame width, single-ear frame area, single-ear frame perimeter, and single-ear frame aspect ratio, used to quantify the individual development status and single-ear load capacity.

[0053] The group-scale feature extraction unit is used to perform pixel-level segmentation of rice ears in mature plot images using the SwinTransformer semantic segmentation model based on the Transformer architecture, generating rice ear mask images. More preferably, the model adopts a symmetrical encoder-decoder architecture: the encoder contains four hierarchical stages, downsampling through the PatchMerging layer to construct multi-scale feature maps, and establishing long-range dependencies between pixels using window-based self-attention and moving window self-attention mechanisms in each stage. The window size is set to 7×7 to balance local details and global contextual information; the decoder fuses features from the corresponding levels of the encoder through linear upsampling and skip connections to compensate for spatial information loss during downsampling. Based on the obtained masked image, the following population features are extracted: quantitative features, including the number of rice ears and the projected area of ​​the rice ears; morphological features, including the area, perimeter, and fractal dimension of the population's bounding rectangle; color and spectral features, including first to third order statistical moments of color extracted based on the visible light band and the visible light vegetation index; and texture features, including second-order angular moments, contrast, correlation, dissimilarity, and homogeneity extracted based on the gray-level co-occurrence matrix, used to reflect the surface roughness and spatial arrangement pattern of the canopy. Furthermore, the population-scale feature extraction unit also combines the classification results of the multi-label classification module to calculate the number of days after sowing at key growth nodes as phenological features, whereby key growth nodes include the number of days from heading and flowering to maturity.

[0054] The yield prediction module uses the multi-scale features extracted by the feature extraction module as input to construct independent yield prediction models for compact and spread-type rice, respectively, to predict rice yield. Specifically, this module cleans, imputes missing values, and standardizes the extracted single panicle and population multi-scale features to eliminate dimensional differences. Using the processed feature set as input and the measured rice yield of the corresponding plot as the output label, machine learning algorithms are applied to train the yield estimation model for the compact and spread-type data subsets, respectively.

[0055] Preferably, the machine learning algorithm employs one of random forest, support vector regression, extreme gradient boosting, or backpropagation neural networks.

[0056] More preferably, a backpropagation neural network algorithm is used to determine the optimal hyperparameters through grid search and cross-validation. After training, the predicted yields of compact rice and spread rice are output.

[0057] Furthermore, this module can also quantitatively analyze the contribution weight of each scale feature group to the yield of different panicle types through the SHAP value algorithm, so as to analyze the intrinsic agronomic mechanism of yield formation.

[0058] Example 3 Based on the same inventive concept, the present invention also designs a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of a method and system for predicting rice germplasm yield by aligning panicle type classification with growth period.

[0059] Since the computer-readable storage medium described in Embodiment 3 of this invention is the same computer-readable medium used in the rice germplasm yield prediction method and system for aligning panicle type classification and growth period in Embodiment 1 of this invention, those skilled in the art can understand the specific structure and variations of this computer-readable storage medium based on the method described in Embodiment 1 of this invention, and therefore will not be repeated here. All computer-readable media used in any method of this invention are within the scope of protection of this invention.

[0060] The specific embodiments described in this application are merely illustrative of the spirit of the invention. Those skilled in the art can make various modifications or additions to the described embodiments or use similar methods to substitute them, without departing from the spirit of the invention or exceeding the scope defined by the appended claims.

Claims

1. A method for predicting rice germplasm yield by aligning panicle type classification with growth period, characterized in that, Includes the following steps: Step S1: Obtain time-series images of the rice experimental field, and obtain the cell images of each sampling plot after preprocessing; Step S2: Use the single ear detection model to detect rice ears in the image of the plot. If rice ears are detected, output the single ear image and the bounding box information of the single ear. Step S3: Input the single ear image into the multi-label single ear classification model, and simultaneously identify and output the growth period label and ear type label of the single ear; wherein, the growth period label includes at least the heading and flowering period and the maturity period, and the ear type label includes at least the compact type and the spreading type; wherein: The multi-label single-ear classification model uses the PP-LCNet network as the backbone network to simultaneously identify and output the growth period label and ear type label of a single ear. The ear type and growth period categories output by the model are locally mutually exclusive labels. The confidence comparison method is used to compare the confidence of the pairwise mutually exclusive labels, and the label with the higher confidence is taken as the final classification result. Step S4: Based on the classification results of step S3, implement the growth period alignment and ear type classification strategy, filter out the data in the maturity stage, and divide them into two independent data subsets: compact type and spreading type. Step S5: Perform multi-scale feature extraction on the filtered data. The multi-scale features include single-ear scale features and population scale features. Population scale feature extraction includes: The Swin Transformer semantic segmentation model based on the Transformer architecture was used to perform pixel-level segmentation of rice ears in images of mature plots to generate rice ear mask images. Based on the rice panicle mask image, population features are extracted, including rice panicle quantity and morphological features; The phenological characteristics are obtained by combining the classification results of step S3, including the number of days for heading and flowering and the number of days for maturity. Step S6: Using the multi-scale features extracted in step S5 as input, construct independent yield prediction models for compact and spread-type rice respectively, and output the rice yield prediction results.

2. The method for predicting rice germplasm yield by aligning panicle type classification with growth period according to claim 1, characterized in that: In step S2, the single ear detection model is the ST-YOLO single ear detection model, which is constructed by introducing the Swin Transformer module into the backbone network of the YOLO network. The ST-YOLO single ear detection model adopts a P3-P5 multi-scale output structure and retains the SPPF module and C2PSA module. The detection head part adopts continuous upsampling and feature splicing to achieve multi-scale feature fusion.

3. The method for predicting rice germplasm yield by aligning panicle type classification with growth period according to claim 1, characterized in that, In step S5, the single-ear scale feature extraction refers to: extracting the geometric features of the single-ear frame from the boundary box information of the single ear in step S2. The geometric features of the single ear frame include one or more of the following: single ear frame length, single ear frame width, single ear frame area, single ear frame perimeter, and single ear frame length-to-width ratio.

4. The method for predicting rice germplasm yield by aligning panicle type classification with growth period according to claim 1, characterized in that: The yield prediction model in step S6 employs a machine learning algorithm, which may be one of random forest, support vector regression, extreme gradient boosting, or backpropagation neural network.

5. A system for implementing the rice germplasm yield prediction method according to any one of claims 1-4, characterized in that, include: The data acquisition module is used to acquire time-series images of rice experimental fields, and after preprocessing, obtains the cell images of each sampling plot. The single ear detection module is used to detect rice ears in the image of the plot using a single ear detection model. If a rice ear is detected, the single ear image and the bounding box information of the single ear are output. A multi-label classification module is used to input the single ear image into a multi-label single ear classification model, and simultaneously identify and output the growth period label and ear type label of the single ear; wherein, the growth period label includes at least the heading and flowering period and the maturity period, and the ear type label includes at least the compact type and the spreading type; The data filtering module is used to perform a growth period alignment and ear type classification strategy based on the classification results of the multi-label classification module, filter out the data in the maturity stage, and divide them into two independent data subsets: compact and spreading. The feature extraction module is used to extract multi-scale features from the filtered data, including single-ear scale features and population scale features. The yield prediction module is used to construct independent yield prediction models for compact and spread rice varieties, respectively, using the multi-scale features extracted by the feature extraction module as input, so as to predict the yield of rice.

6. The system according to claim 5, characterized in that: The single-ear detection module adopts the ST-YOLO single-ear detection model; the multi-label classification module uses the PP-LCNet network as the backbone network and uses the confidence comparison method to compare the confidence of mutually exclusive labels and take the label with the higher confidence as the final classification result.

7. The system according to claim 5, characterized in that, The feature extraction module includes: The single-ear scale feature extraction unit is used to extract the geometric features of the single-ear frame from the single-ear bounding box information output by the single-ear detection module; The group-scale feature extraction unit is used to perform pixel-level segmentation of rice ears in mature plot images using the Swin Transformer semantic segmentation model, generate rice ear mask images, extract group features based on the mask images, and obtain phenological features by combining the classification results.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the rice germplasm yield prediction method that aligns panicle type classification with growth period as described in any one of claims 1-4.