Seed lodging-resistant breeding method based on big data
By constructing a multi-source heterogeneous database and performing spatiotemporal alignment and fusion, and utilizing tensor networks and hypergraph attention networks, combined with reinforcement learning and mechanical surrogate models, the problems of data association feature loss and phenotypic identification lag in seed lodging resistance breeding were solved, achieving efficient and accurate breeding decisions and seed lodging resistance assessment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-16
- Publication Date
- 2026-03-20
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies lack multi-source heterogeneous data fusion methods in seed lodging resistance breeding, resulting in the loss of correlation features between data. Traditional phenotypic identification relies on destructive measurement lag, leading to low breeding efficiency and difficulty in quickly identifying high-yield and high lodging resistance germplasm.
A multi-source heterogeneous lodging-resistant breeding database was constructed, multimodal data spatiotemporal alignment and fusion were performed, latent correlation features were extracted using tensor network technology, a spatiotemporal hypergraph attention network was constructed for breeding decision-making, and non-destructive evaluation was carried out by combining reinforcement learning and cross-modal mechanical surrogate models to optimize the breeding path.
Effectively extract latent association features between genotype and environment to achieve accurate prediction of lodging resistance traits and breeding decisions, improve breeding efficiency, reduce resource waste, and ensure the lodging resistance of seeds in extreme environments.
Smart Images

Figure CN121709038A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the interdisciplinary field of smart agriculture and bioinformatics, specifically to a seed lodging resistance breeding method based on big data. Background Technology
[0002] With the rapid development of high-throughput sequencing technology and crop phenomics, modern agricultural breeding has gradually transformed from traditional experience-based breeding to big data-driven intelligent breeding. Seed lodging resistance is one of the key factors determining the final yield of crops and the efficiency of mechanized harvesting. Especially in the planting of major food crops such as corn and rice, lodging not only leads to significant yield losses but also causes a decline in quality. Therefore, how to mine the key genetic laws of lodging resistance from massive breeding data and achieve accurate prediction at the seed stage is of great strategic significance for shortening the breeding cycle and ensuring food security.
[0003] Currently, in the fields of seed lodging resistance breeding and data management, existing technical solutions mainly focus on the functional verification of single genes or the digital management of basic data. For example, existing technologies disclose a maize gene that controls plant growth and organ size. This approach mainly uses bioengineering to increase the activity of specific polypeptides to stimulate root growth, thereby improving root morphology and enhancing the upright stability of the plant. This method focuses on the modification of microscopic molecular biological mechanisms. In addition, there are also technical solutions that disclose data-driven management methods and systems for individual seeds, focusing on establishing a database to identify and track the genetic and phenotypic information of individual seeds, emphasizing data recording and query functions.
[0004] However, the aforementioned existing technologies still have significant shortcomings in addressing the complex needs of lodging resistance breeding. First, lodging resistance is the result of a high degree of interaction between genotype, environment, and field management measures. Existing technologies lack effective means of fusing multi-source heterogeneous data, making it difficult to align microscopic one-dimensional genotype data with macroscopic three-dimensional phenotypic data and dynamic time-series environmental data in a unified spatiotemporal dimension, resulting in the loss of deep correlation features between data. Second, traditional lodging resistance phenotypic identification heavily relies on destructive measurements or manual observations after natural disasters, lacking a proxy mechanism based on non-destructive image back-inference of mechanical indicators, leading to delayed identification and the inability to conduct large-scale operations. Finally, faced with a massive number of potential parental combinations, existing breeding path planning often relies on linear statistics or field trial and error, lacking an intelligent decision-making mechanism that can simulate virtual evolution and automatically prune high-risk paths, resulting in low breeding efficiency and difficulty in quickly identifying superior germplasm with both high yield and high lodging resistance characteristics. Summary of the Invention
[0005] The purpose of this application is to provide a big data-based method for seed lodging resistance breeding, comprising: constructing a multi-source heterogeneous lodging resistance breeding database, wherein the lodging resistance breeding database includes genome sequencing data of the target crop population, high-throughput field phenotypic data throughout the entire growth cycle, and synchronous micro-meteorological environment data; performing spatiotemporal alignment and fusion processing of multimodal data, mapping one-dimensional genotype features, three-dimensional phenotypic spatial features, and time-series environmental features to a unified high-dimensional tensor space, and using tensor network technology to perform Tucker decomposition or tensor orthogonal rank-one decomposition on the constructed high-dimensional tensor to extract latent association features of genotype-environment interaction; constructing and training a spatiotemporal hypergraph attention network, inputting the latent association features into the spatiotemporal hypergraph attention network, updating feature weights through a node aggregation mechanism, and outputting a spatiotemporal hypergraph feature matrix containing lodging resistance genetic potential; calculating the comprehensive lodging resistance index of each breeding material based on the spatiotemporal hypergraph feature matrix, and performing parent selection and offspring optimization based on the comprehensive lodging resistance index to complete the seed lodging resistance breeding decision-making process.
[0006] By adopting the above technical solution, the latent association features between genotype and environment can be effectively extracted, solving the problem of feature distortion caused by the mismatch of dimensions in multi-source data.
[0007] Optionally, the processing steps of the high-throughput phenotypic data in the field include a semantic segmentation process for phenotypic micro-features: using an improved multi-scale convolutional coding-decoding semantic segmentation network to perform pixel-level semantic segmentation on the collected crop canopy orthophoto and side-view projection images; extracting key morphological micro-features for lodging resistance from the segmented image mask, the morphological micro-features including basal stem thickness, third internode length, stem fullness, and root soil heave volume; quantizing the morphological micro-features into numerical vectors, which are then loaded into the high-dimensional tensor space as input components of the phenotypic modality.
[0008] By adopting the above technical solution, the problems of low accuracy in traditional phenotypic data extraction and inability to quantify microscopic features are solved, ensuring that the data input to the model contains biologically significant fine structural information.
[0009] Optionally, a cross-modal mechanical surrogate model technique is employed: Field-measured destructive mechanical data are acquired, including crop stem puncture strength, critical breaking force, and root pull-out strength; a nonlinear mapping model based on a multilayer perceptron is constructed, using the morphological micro-feature numerical vectors as input and the destructive mechanical data as supervisory labels for model training; during the breeding prediction stage, the trained nonlinear mapping model is used to convert the non-destructive image features of the test sample into virtual mechanical strength indices, which participate in the lodging resistance performance evaluation.
[0010] By adopting the above technical solution, the problem that traditional mechanical measurement relies on destructive experiments and is difficult to apply on a large scale has been solved.
[0011] Optionally, the micrometeorological environment data includes the instantaneous maximum wind speed, prevailing wind direction, short-term heavy rainfall, and soil volumetric moisture content at different soil depths in the field; in the spatiotemporal alignment and fusion processing, the micrometeorological environment data is used to construct a dynamic stress simulation field, and the intensity of combined wind and rain stress encountered by crops at different growth stages is used as a dynamic weighting factor to weight and correct the expression effect value of genotype characteristics.
[0012] By adopting the above technical solutions, breeding evaluation is no longer based on static environmental assumptions, but fully considers dynamic adaptability under extreme weather conditions.
[0013] Optionally, the breeding decision-making process includes a path simulation step based on reinforcement learning: constructing a virtual breeding environment that includes historical meteorological scenarios and genetic laws; defining the agent's action space as parental mating operations and offspring screening strategies, and defining the state space as the gene frequency distribution of the current generation population and the predicted lodging resistance trait; designing a multi-objective reward function, which is composed of a weighted positive feedback of yield gain per unit area and a negative feedback of crop lodging rate, to guide the agent to autonomously explore the optimal breeding path in the virtual breeding environment.
[0014] By adopting the above technical solution, a multi-objective reward function is designed, which is composed of positive feedback of yield gain per unit area and negative feedback of crop lodging rate. This function guides the agent to explore autonomously in a virtual environment, thereby evolving the optimal breeding path without consuming actual field resources.
[0015] Optionally, the reinforcement learning process includes an automatic pruning mechanism: during the agent's exploration process, the probability of lodging risk of the current hybridization combination path is calculated in real time; when the predicted probability of lodging risk of a certain genetic combination path exceeds a preset safety threshold, the pruning operation is automatically triggered to block the subsequent simulation calculation of the genetic combination path, and the combination is marked as a high-risk elimination combination and fed back to the parent selection database.
[0016] By adopting the above technical solution, the amount of invalid computation is significantly reduced, and resources are avoided from being wasted on hybridization combinations that are destined to fail.
[0017] Optionally, it also includes constructing a lodging resistance-enhanced genomic selection model: using the spatiotemporal hypergraph feature matrix as a fixed effect factor and fusing it into a hybrid linear model; calculating the genomic estimated breeding value after fusion features, correcting the prediction bias of lodging resistance traits caused by the traditional best linear unbiased prediction model relying solely on pedigree relationships, and outputting a corrected lodging resistance breeding value ranking list.
[0018] By adopting the above technical solution, the prediction bias of lodging resistance traits caused by the traditional best linear unbiased prediction model relying solely on pedigree relationships is corrected, and a more accurate ranking list of lodging resistance breeding values is output.
[0019] Optionally, it also includes data preprocessing, specifically including: filling in phenotypic data with spatiotemporal gaps using Kriging interpolation or long short-term memory recurrent neural networks; and performing standard score standardization on multi-source heterogeneous data to map physical quantities of different dimensions to a distribution range with a mean of zero and a standard deviation of one, thereby eliminating the impact of differences in data magnitude on the accuracy of tensor decomposition.
[0020] Optionally, the parameter range for model training is set as follows: the initial learning rate of the spatiotemporal hypergraph attention network and reinforcement learning model is set between one-thousandth and one-hundredth; the number of iterations for model training is set between 1,000 and 5,000, and an early stopping strategy is adopted to prevent overfitting; the rank parameter of tensor decomposition is dynamically adjusted according to the data sparsity and kept within the range of ten percent to thirty percent of the feature dimension.
[0021] Optionally, it can be applied to field big data fusion breeding scenarios for maize or rice: for maize crops, the focus is on optimizing the weighted analysis of the number of aerial root layers and the bending resistance of the internodes at the base of the stem; for rice crops, the focus is on optimizing the synergistic evaluation of the bending moment of the panicle neck and the canopy ventilation; and the breeding decision-making process outputs a recommended planting scheme for lodging-resistant new varieties adapted to specific ecological zones. Attached Figure Description
[0022] Figure 1 This is a schematic diagram of the seed lodging resistance breeding method based on big data in this application. Detailed Implementation
[0023] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present application.
[0024] like Figure 1 As shown in the figure, this application provides a seed lodging resistance breeding method based on big data, including the following steps.
[0025] S01: Construct a multi-source heterogeneous lodging resistance breeding database, which includes genome sequencing data of the target crop population, high-throughput field phenotypic data throughout the entire growth cycle, and synchronous microclimate environment data.
[0026] Specifically, a multi-source heterogeneous database covering the entire life cycle of the target crop is constructed. The input data includes genome sequencing data, field high-throughput phenotypic data, and micro-meteorological environment data, specifically including genome single nucleotide polymorphism chip data or whole genome resequencing data. The data storage format is a common variant retrieval format. Field phenotypic data is acquired through multispectral cameras on UAVs and ground-based lidar, including time-series point clouds of parameters such as canopy coverage, plant height, and stem angle. Environmental data comes from minute-level wind speed, rainfall, and soil tension collected by field weather stations.
[0027] S02: Perform spatiotemporal alignment and fusion processing of multimodal data, mapping one-dimensional genotype features, three-dimensional phenotypic spatial features, and time-series environmental features to a unified high-dimensional tensor space. Use tensor network technology to perform Tucker decomposition or tensor orthogonal rank-one decomposition on the constructed high-dimensional tensor to extract latent association features of genotype-environment interaction.
[0028] Specifically, the spatiotemporal alignment and fusion processing of multimodal data first aligns data from different sampling frequencies based on GPS coordinates and a unified timestamp index. For example, static genotype data is used as a constant dimension, while high-frequency environmental data is averaged through a sliding window and adapted to key crop growth stages, such as jointing, tasseling, and grain filling. A third-order or higher tensor is constructed, with its dimensions representing individual samples, time nodes, and multimodal feature channels, respectively. Tensor network technology is then used to decompose the constructed high-dimensional tensor. Specifically, the Tucker decomposition algorithm can be used. This algorithm iteratively optimizes the original high-dimensional sparse tensor through alternating least squares to decompose it into a core tensor and multiple factor matrices, compressing and preserving the latent interactions between different modalities, such as the latent interactions between strong wind environments and specific single nucleotide polymorphism sites and the high elastic modulus of stems. During the operation, the rank parameter is set to a range of 10% to 30% of the original dimension to remove data noise and prevent overfitting.
[0029] S03: Construct and train a spatiotemporal hypergraph attention network, input latent association features into the spatiotemporal hypergraph attention network, update feature weights through node aggregation mechanism, and output a spatiotemporal hypergraph feature matrix containing lodging resistance genetic potential.
[0030] Specifically, the input layer of the spatiotemporal hypergraph attention network receives the latent association feature vectors after the above decomposition. The network architecture includes a hyperedge construction module, which dynamically constructs hyperedges connecting different samples and environmental factors based on feature similarity using the K-nearest neighbor algorithm. The attention mechanism layer calculates the importance weight of each node within the hyperedge, and node aggregation and updating of features are achieved through multi-layer graph convolution operations. The training process uses an adaptive moment estimation optimizer with weight decay, and the loss function is set as the cross-entropy loss between the predicted lodging resistance level and the measured lodging level in the field. The number of iterations is set to 1,000 to 5,000 until the validation set error converges. The final output is a spatiotemporal hypergraph feature matrix containing the genetic potential for lodging resistance.
[0031] S04: Calculate the comprehensive lodging resistance index of each breeding material based on the spatiotemporal hypergraph feature matrix, and select parent lines and offspring based on the comprehensive lodging resistance index to complete the seed lodging resistance breeding decision-making process.
[0032] Understandably, the spatiotemporal hypergraph feature matrix is transformed into a normalized lodging resistance index, ranging from zero to one. This lodging resistance index not only reflects the crop's own genetic resistance but also includes its probability of adapting to extreme weather conditions. Breeders rank tens of thousands of breeding materials based on this lodging resistance index, eliminating materials with an index below a preset threshold, such as 0.75, thereby achieving precise targeting of lodging resistance traits at the molecular design breeding stage.
[0033] To address the practical problems of delayed non-destructive prediction of lodging phenotypes in the field and the difficulty of capturing early lodging micro-features in traditional manual surveys, the processing steps of high-throughput field phenotypic data include a semantic segmentation process for phenotypic micro-features. This semantic segmentation process can keenly capture key morphological structural parameters that determine lodging resistance before lodging occurs. In the embodiments of this application, the processing steps of high-throughput field phenotypic data include a semantic segmentation process for phenotypic micro-features: pixel-level semantic segmentation of the collected crop canopy orthophoto and side-view projection images is performed using an improved multi-scale convolutional coding-decoding semantic segmentation network; key morphological micro-features for lodging resistance are extracted from the segmented image mask, including basal stem thickness, third internode length, stem fullness, and root soil heave volume; the morphological micro-features are quantized into numerical vectors and loaded into a high-dimensional tensor space as input components of the phenotypic modality.
[0034] Specifically, the semantic segmentation process includes: acquiring high-resolution image data, the data sources of which include orthophotos taken by low-altitude UAVs at a height of 10 to 30 meters above the canopy, with a resolution better than 1 centimeter per pixel, and side-view projection images of crop stem bases taken by ground-based phenotyping robots. The image formats are lossless labeled image files or high-compression Joint Image Experts Group (JEAG) formats, covering the visible and near-infrared bands; performing image preprocessing operations, including enhancing the contrast between stems and background using histogram equalization, and normalizing the image size to the fixed resolution required by the model input using bilinear interpolation, such as 512 by 512 pixels; and performing pixel-level semantic segmentation using an improved multi-scale convolutional encoder-decoder semantic segmentation network. The semantic segmentation network architecture adopts a symmetrical encoder-decoder structure. The encoder consists of multiple convolutional layers and max-pooling layers to extract deep abstract features of the image and reduce spatial dimensions. The decoder gradually restores spatial resolution through upsampling operations and fuses shallow texture features transmitted by the encoder through skip connections, thereby solving the problem of blurred edges of small stems. The output layer of the semantic segmentation network uses a normalized exponent. The function generates a probability map, classifying each pixel as stem, leaf, soil, or background. Key morphological micro-features for lodging resistance are extracted from the segmented binary mask. Specific algorithm logic includes: for stem thickness at the base, calculating the average pixel width of the stem base region in the side-view mask and converting it to the actual physical diameter in millimeters using camera calibration parameters and shooting distance; for the length of the third internode, identifying the stem centerline using a skeleton extraction algorithm, locating the third leaf node, and calculating its Euclidean distance to the second node; for stem fullness, calculating the stem thickness using the transmittance difference in the near-infrared band. The average gray value of the region is used as a proxy indicator of tissue compactness. For the root soil heave volume, the texture changes and height anomalies of the soil around the roots in the orthophoto are analyzed. It is necessary to combine the digital surface model to estimate the small soil displacement caused by insufficient root anchoring force. The extracted morphological micro-features are numerically processed to construct a feature vector, where each component represents stem thickness, internode length, fullness and heave volume, respectively. The standard fractions of each component are standardized to make the mean zero and the variance one to eliminate the difference in dimensions. Finally, the vector is used as the input of the phenotypic mode and loaded into the high-dimensional tensor space.
[0035] To address the bottleneck of current destructive mechanical measurements in breeding processes, such as stem breakage, which cannot be fully implemented across large breeding populations and cannot be tested on valuable individual plants, this application employs a cross-modal mechanical surrogate model technology to achieve non-destructive testing that reveals force from images. Specifically, this includes: acquiring field-measured destructive mechanical data, including crop stem puncture strength, critical breaking force, and root pull-out strength; constructing a nonlinear mapping model based on a multilayer perceptron, using morphological micro-feature numerical vectors as input and destructive mechanical data as supervisory labels for model training; and in the breeding prediction stage, using the trained nonlinear mapping model to convert the non-destructive image features of the sample to be tested into virtual mechanical strength indices, which participate in the evaluation of lodging resistance.
[0036] Specifically, a field-measured mechanical dataset was established. A small, representative sample, such as 5% of the total population, was selected for destructive experiments. A portable plant stem strength meter was used to record the force-displacement curves during the probe's penetration of the stem at a sampling frequency of 50 Hz. The puncture strength, critical breaking force, and root vertical pull-out force of the crop stem were extracted as real physical labels. The unit of crop stem puncture strength was Newtons per square millimeter, and the unit of force was Newton. Simultaneously, the estimated values of cellulose and lignin content of the corresponding samples were obtained through hyperspectral imaging. A nonlinear mapping model based on a multilayer perceptron was constructed. This nonlinear mapping model establishes a black-box mapping relationship between image micro-features and physical mechanical indicators. The input layer of the nonlinear mapping model receives morphological micro-feature vectors, and the hidden layer contains three to five fully connected layers. The number of neurons in each layer is set to 64 to 256. The activation function uses a modified linear unit to introduce nonlinearity. To prevent gradient vanishing, the output layer corresponds to the predicted three mechanical indices. Model training and parameter tuning are implemented, using mean squared error as the loss function to measure the difference between predicted and measured mechanical values. The backpropagation algorithm is used to update network weights, and a random inactivation regularization strategy is introduced, with a dropout rate set to 0.2 to 0.5. Randomly blocking neuron connections enhances the model's generalization ability and prevents overfitting to specific varieties. In the breeding prediction stage, the lossless image data of the large-scale breeding population to be tested is input into the trained nonlinear mapping model. The nonlinear mapping model outputs the virtual mechanical strength index of each individual plant in real time. This virtual mechanical strength index not only fills the gaps in the destructive data, but also reveals the nonlinear compensation mechanism between stem diameter, wall thickness and strength through feature fusion. For example, some thin-stemmed varieties have high bending resistance due to extremely high cell wall density, thus significantly improving the biophysical interpretability of lodging resistance assessment.
[0037] To address the problem that traditional breeding focuses only on static genotypes while neglecting dynamic environmental stresses, resulting in high yields but poor lodging resistance or limited adaptability due to lodging resistance, this application introduces dynamic stress simulation field technology. By quantifying environmental stress and correcting genetic assessments, micrometeorological environmental data include instantaneous maximum wind speed, prevailing wind direction, short-term heavy rainfall, and soil volumetric moisture content at different soil depths. In the spatiotemporal alignment and fusion processing, a dynamic stress simulation field is constructed using micrometeorological environmental data. The combined wind and rain stress intensity encountered by crops at different growth stages is used as a dynamic weighting factor to weight and correct the expression effect value of genotype characteristics.
[0038] Specifically, high spatiotemporal resolution micro-meteorological environmental data are collected. Using ultrasonic anemometers, tipping bucket rain gauges, and frequency domain reflectometry soil moisture sensors deployed at field IoT nodes, real-time data are acquired to obtain instantaneous maximum wind speed (gust speed), sampling interval of one second, prevailing wind direction (range 0-360 degrees), short-duration heavy rainfall (unit: millimeters per hour), and volumetric water content of the 0-20 cm soil layer. A dynamic stress simulation field is constructed. Based on fluid mechanics principles, the crop canopy is simplified as a wind-receiving body. The physical process of calculating wind load is specifically defined as half the product of air density, wind-receiving area, the square of wind speed, and drag coefficient. This is combined with the canopy interception weight gain effect caused by rainfall and the root anchoring stiffness reduction effect caused by increased soil moisture content (i.e., soil liquefaction risk), generating a comprehensive stress intensity curve that varies over time. The dynamic weighting factor is calculated by dividing the entire crop growth cycle into seedling, jointing, heading, and grain-filling maturity stages. Based on historical meteorological big data, the probability of lodging disasters at each stage is statistically analyzed. Combined with the current real-time stress intensity, the environmental pressure is mapped to a dynamic weight between zero and one using an S-shaped growth curve function. When the maximum wind speed exceeds the threshold, such as a gale of level 8, accompanied by heavy rainfall, the weight approaches one. The genotype feature expression effect is weighted and corrected. During the spatiotemporal alignment and fusion process, the dynamic weighting factor and genotype feature values are subjected to a matrix element-to-element Hadamard product operation. This Hadamard product operation amplifies the signal intensity of gene loci that can still maintain upright growth under extreme weather conditions and suppresses the weight of pseudo-high resistance genes that only perform well under stable weather conditions, thereby ensuring that the selected varieties have real field lodging resistance and stable yield capabilities.
[0039] To address the bottlenecks of massive hybridization path combinations, long and costly field trial-and-error cycles, and inefficient virtual simulation pruning of breeding paths, the breeding decision-making process in this application includes a path simulation step based on reinforcement learning: constructing a virtual breeding environment that includes historical weather scenarios and genetic laws; defining the agent's action space as parental mating operations and offspring selection strategies, and defining the state space as the gene frequency distribution of the current generation population and the predicted lodging resistance trait; designing a multi-objective reward function, which is composed of a weighted positive feedback of yield gain per unit area and a negative feedback of crop lodging rate, to guide the agent to autonomously explore the optimal breeding path in the virtual breeding environment.
[0040] Specifically, a virtual breeding environment incorporating historical meteorological scenarios and genetic patterns is constructed. This virtual breeding environment integrates a crop growth simulation system and a genome prediction model. Based on the input parental genotypes and set future climate scenarios, such as simulating a once-in-fifty-year windy year, it can predict the growth, development, and lodging performance of offspring populations. The action space of the agent is defined as parental selection operations and offspring screening strategies, where the agent represents the breeding decision-making system. The action space is defined as the gene frequency distribution of the current generation population and the predicted lodging resistance trait, allowing for the selection of parents from the germplasm resource bank for hybridization, self-pollination, or backcrossing. The state space consists of the allele frequency distribution of the current generation population, the predicted average yield, and the lodging rate. A multi-objective reward function is designed, which is composed of a weighted positive feedback of yield gain per unit area and a negative feedback of crop lodging rate. This guides the agent to autonomously explore the optimal breeding path in the virtual breeding environment. The mathematical... The logic is to multiply the yield gain relative to the control group by a first weighting coefficient, and subtract the model-predicted lodging ratio multiplied by a second weighting coefficient. This multi-objective reward function explicitly penalizes combinations with high lodging risk and rewards high-yield and stable-yield combinations. The training process, based on a proximal strategy optimization algorithm, involves the agent performing millions of simulated hybridization attempts in a virtual breeding environment. Through trial and error learning, the parameters of the strategy network are continuously updated, enabling it to identify which parental pairing patterns, such as specific complementarities between short-stalked, multi-resistant and tall-stalked, large-eared varieties, can maximize cumulative rewards. The optimal breeding path is output. The trained agent no longer performs random trials but directly generates the shortest genetic improvement path from existing germplasm to the ideal new variety, clearly indicating the required number of hybridization generations and the key molecular markers to be screened in each generation. This compresses the traditional eight-to-ten-year breeding cycle into preliminary screening within hours in computer simulation, significantly reducing the resource waste of blind field testing.
[0041] To address the technical bottlenecks of existing breeding simulation technologies, such as severe waste of computational resources, lagging identification of high-risk genetic pathways, and slow convergence speed due to invalid iterations when dealing with massive parental hybridization combinations, the reinforcement learning process of this application includes an automatic pruning mechanism: during the agent's exploration process, the lodging risk probability of the current hybridization combination path is calculated in real time; when the predicted lodging risk probability of a certain genetic combination path exceeds a preset safety threshold, a pruning operation is automatically triggered to block the subsequent simulation calculation of that genetic combination path, and the combination is marked as a high-risk elimination combination and fed back to the parental selection database.
[0042] Understandably, this automatic pruning mechanism, by establishing a real-time dynamic risk assessment model, eliminates genetic pathways that are destined to fail to meet lodging resistance requirements in the early stages of virtual evolution. The operation process of the automatic pruning mechanism includes, but is not limited to: calculating the lodging risk probability of the current hybridization combination path in real time during the agent's exploration process; when the predicted lodging risk probability of a certain genetic combination path exceeds a preset safety threshold, automatically triggering the pruning operation, blocking the subsequent simulation calculation of the genetic combination path, marking the combination as a high-risk elimination combination, and feeding it back to the parent selection database.Specifically, a risk probability calculation module based on Bayesian inference is constructed; real-time path safety monitoring and threshold interpretation are implemented; a circuit breaker pruning operation is triggered and a negative sample feedback library is generated; the input data of the risk probability calculation module comes from the genotype matrix of intermediate hybrid offspring generated by the reinforcement learning agent in the virtual breeding environment. This genotype matrix contains single nucleotide polymorphism site information of tens of thousands of simulated individuals, and the data scale is usually a sparse matrix with millions of rows and hundreds of thousands of columns. At the same time, the simulated breeding environment stress parameters corresponding to the current virtual generation are also input, such as the wind load pressure value during the simulated typhoon, in Newtons per square meter, and the soil liquefaction coefficient, which is zero to... The system is dimensionless; it does not wait for the simulation of the entire growth cycle to end, but instead, when the virtual crop develops to a critical disaster-causing stage, such as the tasseling stage of corn or the grain-filling stage of rice, it calls a lightweight proxy model to quickly estimate the lodging tendency score of the current population. This lodging tendency score is calculated based on the ratio of the predicted value of the crop's center of gravity height, the mechanical strength of the stem base, and the environmental wind load. If more than 30% of the offspring population produced by a certain hybrid combination path has a lodging tendency score higher than the critical lodging threshold, which is derived from the 95th percentile of historical field measurement data, then the path is determined to be a high-risk genetic trajectory; real-time monitoring of path safety is implemented. In the threshold determination process, a dynamic safety threshold is set. This threshold is not fixed but tightens exponentially as the simulation generations progress. A certain degree of risk is allowed in the initial exploration phase to preserve genetic diversity, while risk is strictly limited in the later optimization phase. Each time the system performs a parental selection action, it calculates the risk probability implicit in the current action value function. If this risk probability exceeds the threshold, an interrupt command is immediately activated. Triggering a circuit breaker pruning operation not only stops all subsequent computational resource allocation for the current branch—such as stopping high-computational-power-consuming steps like yield simulation, quality analysis, and disease resistance prediction for the offspring of that combination—but also instantly releases approximately 40% to 100% of the computational resources. Sixty percent of the computing power was used to explore other potential pathways. At the same time, the blocked parental combination was marked as a lodging-sensitive pair and its feature vector was stored in the negative sample experience replay pool. The output results showed significantly improved search efficiency and excellent lodging-resistant germplasm locking ability. The reinforcement learning model after pruning optimization can shorten the time to converge to the Pareto optimal frontier by more than 75% when dealing with 100,000 potential hybrid combinations. Moreover, the deviation between the measured and predicted lodging rates of the final recommended parental combination in subsequent field verification was controlled within 5%, effectively avoiding recommending varieties with high yields but high lodging risk to the expensive field testing stage.
[0043] To address the problem that traditional genomic selection models, particularly the optimal linear unbiased prediction algorithm, over-rely on pedigree relationships and lack the ability to capture non-additive effects and complex gene-environment interactions, resulting in low accuracy in predicting complex quantitative traits like lodging resistance, this application proposes a lodging resistance-enhanced genomic selection model that integrates spatiotemporal hypergraph features. This model utilizes the spatiotemporal hypergraph feature matrix as a fixed-effect factor, integrating it into a hybrid linear model. The genomically estimated breeding values after feature integration are calculated, correcting the prediction bias of lodging resistance caused by the traditional optimal linear unbiased prediction model's reliance solely on pedigree relationships. The corrected lodging resistance breeding value ranking list is then output.
[0044] Specifically, the construction process of the lodging resistance-enhanced genome selection model includes, but is not limited to: extracting and reducing the dimensionality of the spatiotemporal hypergraph feature matrix; constructing a mixed linear model containing fixed and random effects; using the spatiotemporal hypergraph feature matrix as a fixed effect factor and fusing it into the mixed linear model; calculating the genome-estimated breeding value after feature fusion; correcting the prediction bias of lodging resistance traits caused by the traditional best linear unbiased prediction model relying solely on pedigree relationships; and outputting a corrected ranking list of lodging resistance breeding values. The input to the feature extraction step directly comes from the output layer of the spatiotemporal hypergraph attention network in the aforementioned embodiment. This spatiotemporal hypergraph feature matrix is a high-dimensional dense vector, for example, with dimensions of 1 x 1024, which highly condenses the lodging resistance biophysical response pattern of crop genotypes under specific spatiotemporal environmental stresses. To solve the curse of dimensionality, principal component analysis is used to reduce the dimensionality of this matrix, retaining the top principal components with a cumulative variance contribution rate of over 95%, and using these principal components as lodging resistance prior factors. When constructing the mixed linear model, the limitation of the traditional whole-genome best linear unbiased prediction model, which only uses the molecular marker matrix as a random effect, is broken. A hybrid equation was established, incorporating conventional environmental fixed effects, additive genetic effects based on kinship matrices, and regression terms from the spatiotemporal hypergraph feature factor matrix. This explicitly embeds nonlinear lodging resistance features mined by deep learning into a statistical genetics framework. During the solution process, the variance components were iteratively estimated using restricted maximum likelihood estimation. By introducing regression terms containing deep environmental interaction information, more reasonable prediction values can be assigned to materials with distant kinship but similar lodging resistance mechanisms, such as those with well-developed aerial roots or high lignin content. This corrects the biases in traditional models caused by kinship. The bias of underestimating breeding potential due to distant kinship is eliminated; the output is a double-corrected ranking list of lodging resistance breeding values. In a validation population test containing 500 maize inbred lines, the prediction accuracy of this mixed linear model, i.e., the Pearson correlation coefficient between predicted values and field measurements, improved from 0.45 in the traditional model to 0.78. In particular, the accuracy improvement in predicting performance in years of extreme wind disasters exceeded 60%, enabling breeders to use this quantitative indicator to screen for high-yield potential high-quality germplasm at the seed stage with extremely high confidence.
[0045] To address the issues of spatiotemporal gaps in data collection under open field environments, sensor drift, and model training divergence caused by extreme inconsistencies in the dimensions of multimodal data, this embodiment proposes an engineered data preprocessing strategy to ensure high-quality consistency and completeness of the data input to the tensor network. This data preprocessing includes, but is not limited to: filling in spatiotemporal gaps in phenotypic data using Kriging interpolation or long short-term memory recurrent neural networks; and performing standard score standardization on multi-source heterogeneous data to map physical quantities of different dimensions to a distribution range with a mean of zero and a standard deviation of one, thereby eliminating the impact of differences in data magnitude on the accuracy of tensor decomposition.
[0046] Specifically, Kriging interpolation or long short-term memory recurrent neural networks are used to fill in phenotypic data with spatiotemporal gaps; standard score standardization is performed on multi-source heterogeneous data to map physical quantities of different dimensions to a distribution range with a mean of zero and a standard deviation of one, thereby eliminating the impact of differences in data magnitude on the accuracy of tensor decomposition. Specifically, the project implements geostatistical-based spatial interpolation repair; performs dynamic imputation based on time series deep learning; and conducts standardization and normalization mapping of multi-dimensional data. Spatial interpolation repair primarily addresses the discontinuous sampling problem of environmental data such as soil moisture and nutrients. The input data is sparsely distributed in the field, such as IoT sensor readings with one node per ten acres. Using Kriging interpolation, spatial autocorrelation is analyzed by calculating the semi-variogram, expanding the discrete monitoring point data into a continuous raster map covering the entire field. The raster resolution is refined to one meter by one meter, ensuring that the root environment of each crop has a corresponding estimated value. Time series imputation addresses data gaps caused by drone battery replacements, inclement weather, or sensor malfunctions. A long short-term memory recurrent neural network model is used, inputting meteorological and growth data sequences before and after the missing points, with a time window of seven to fourteen days. Network gating mechanisms are used to capture time dependencies, predicting and filling in the missing moments. For wind speed, light intensity, or crop height data, the filling error is controlled within a root mean square error of less than 5%. Faced with the significant differences in discrete gene values, phenotypic data such as plant height and stem diameter, and environmental data such as rainfall and wind speed, standard score standardization is performed. The calculation logic is to subtract the mean from the original value and divide by the standard deviation, mapping all physical quantities to a standard normal distribution interval with a mean of zero and a standard deviation of one. Simultaneously, for some long-tailed data, such as extreme rainfall, Box-Cox transform is performed to correct skewness. The cleaned panoramic dataset generated in the intermediate steps possesses characteristics of no missing data, identical distribution, and high signal-to-noise ratio, serving as direct input for subsequent tensor decomposition. Field cross-validation of the output results shows that after this preprocessing process, the model's convergence speed is improved by 2.5 times, and the number of abnormal prediction points caused by data noise is reduced by more than 90%, effectively ensuring the robustness and stability of the breeding decision-making system in variable field environments.
[0047] To ensure optimal convergence of the spatiotemporal hypergraph attention network and reinforcement learning model on complex biological breeding data and avoid getting trapped in local minima or experiencing catastrophic forgetting, this embodiment specifies in detail the key hyperparameter ranges and dynamic adjustment strategies for model training. The parameter ranges for model training are set as follows: the initial learning rate of the spatiotemporal hypergraph attention network and reinforcement learning model is set between 0.1% and 0.1%; the number of iterations for model training is set between 1,000 and 5,000, and an early stopping strategy is adopted to prevent overfitting; the rank parameter of tensor decomposition is dynamically adjusted according to the data sparsity and kept within the range of 10% to 30% of the feature dimension.
[0048] Specifically, the training parameter setting and optimization process includes, but is not limited to: setting the initial learning rate of the spatiotemporal hypergraph attention network and reinforcement learning model between 0.1% and 0.1%; setting the number of model training iterations between 1000 and 5000, and employing an early stopping strategy to prevent overfitting; dynamically adjusting the rank parameter of the tensor decomposition according to the data sparsity, maintaining it within the range of 10% to 30% of the feature dimension; for the spatiotemporal hypergraph attention network, an adaptive moment estimation optimizer with weight decay is used, with the initial learning rate set between 0.1% and 0.1%. This range is a golden interval determined based on numerous pre-experiments, ensuring a rapid decrease in the loss function in the initial stage while avoiding excessive step size that skips the optimal solution. As the training rounds increase, a cosine annealing strategy is used to gradually decay the learning rate to a minimum value for fine-tuning deep within the solution space; the maximum number of training iterations is set to 5000, but an early stopping strategy is introduced to monitor the anti-fall prediction on the validation set. If the accuracy fails to improve or even declines after fifty consecutive rounds, training is immediately terminated and the system rolls back to the optimal weight state. Regarding the rank parameter in the tensor decomposition process, this embodiment does not use a fixed rank but dynamically adjusts it based on the sparsity of the input tensor. Specifically, by analyzing the singular value decay curve after singular value decomposition, a threshold of 85% to 90% cumulative energy is selected as the rank selection criterion, typically maintained within 10% to 30% of the original feature dimension. Furthermore, a training process visualization monitoring platform is used to monitor the gradient distribution and loss curve in real time, ensuring that the gradient vector magnitude remains stable within a reasonable range. The output is a set of finely refined model weight parameter files. This model exhibits extremely high generalization ability on independent external test sets. For new varieties that did not participate in training, the accuracy of its lodging resistance prediction remains above 85%, demonstrating the effectiveness and universality of this parameter system in engineering practice.
[0049] Considering the significant differences between maize and rice in morphology, lodging mechanisms, and planting environments, this embodiment develops targeted application optimization schemes to achieve precise breeding decisions tailored to specific crops. These schemes are applied to field big data fusion breeding scenarios for maize or rice: For maize, the focus is on optimizing the weighted analysis of the number of aerial root layers and the bending resistance of the internodes at the base of the stem; for rice, the focus is on optimizing the synergistic evaluation of the bending moment of the panicle neck and the canopy's air permeability; and the breeding decision-making process outputs recommended planting schemes for lodging-resistant new varieties adapted to specific ecological zones.
[0050] Specifically, for maize, lodging is mainly divided into root-up lodging and stem-breakage lodging. The input data focuses on enhancing the side-view images of the stem base obtained by low-altitude UAV flights. The algorithm pays special attention to the number of aerial root layers, the angle of penetration into the soil, and the root width. At the same time, combined with the puncture mechanical proxy index of the third internode at the base of the stem, the weight of these two types of features is increased by 2 times in the model to construct a root-stem dual-locking evaluation system and output a new maize variety with strong grip and tough stems suitable for planting in windy areas. For rice, lodging mostly occurs during the heading and grain-filling stage, often due to excessive bending moment caused by excessive ear weight. The input data focuses on collecting the wind permeability index of the canopy. Through leaf area index inversion and the bending stiffness of the ear neck, a fluid-structure interaction model is used to simulate the domino effect of the canopy under monsoon rain conditions. The system prioritizes selecting rice varieties with high stem elastic modulus, compact plant type, and good air permeability. Based on the decision-making results, it outputs planting recommendation maps adapted to specific ecological zones. For example, for the Huang-Huai-Hai summer maize region, the system recommends varieties with a root lodging resistance index greater than 0.9 and suggests supporting deep soil loosening and appropriate deep sowing measures. For the Yangtze River mid-lower reaches rice region, it recommends short-stalked, thick-stemmed varieties with a low center of gravity and provides optimal planting density suggestions, such as adjusting from 18,000 holes per mu to 15,000 holes per mu to reduce wind resistance. This application scheme has shown significant results in actual field demonstrations. Compared to the control group, the lodging rate of maize varieties selected using this method decreased by 40% in years of severe wind disasters, and the grain filling rate of rice varieties increased by 15%, fully verifying the practical value and economic benefits of the technical solution in real agricultural production scenarios.
[0051] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0052] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces, or indirect coupling or communication connection between apparatuses or units, and may be electrical, mechanical, or other forms.
[0053] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0054] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0055] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks.
[0056] The above are all preferred embodiments of this application and are not intended to limit the scope of protection of this application. Any feature disclosed in this specification (including the abstract and drawings) may be replaced by other equivalent or similar features unless specifically stated otherwise. That is, unless specifically stated otherwise, each feature is only one example of a series of equivalent or similar features.
Claims
1. A seed lodging resistance breeding method based on big data, characterized in that, include: A multi-source heterogeneous lodging resistance breeding database was constructed, which includes genome sequencing data of the target crop population, high-throughput field phenotypic data throughout the entire growth cycle, and synchronous microclimate environment data. The spatiotemporal alignment and fusion processing of multimodal data is performed, mapping one-dimensional genotype features, three-dimensional phenotypic spatial features and time series environmental features to a unified high-dimensional tensor space. Tensor network technology is used to perform Tucker decomposition or tensor orthogonal rank-one decomposition on the constructed high-dimensional tensor to extract latent association features of genotype-environment interaction. A spatiotemporal hypergraph attention network is constructed and trained. The latent association features are input into the spatiotemporal hypergraph attention network. The feature weights are updated through a node aggregation mechanism, and the spatiotemporal hypergraph feature matrix containing the lodging resistance genetic potential is output. Based on the spatiotemporal hypergraph feature matrix, the comprehensive lodging resistance index of each breeding material is calculated. Parental selection and offspring optimization are carried out based on the comprehensive lodging resistance index to complete the seed lodging resistance breeding decision-making process.
2. The seed lodging resistance breeding method based on big data according to claim 1, characterized in that, The processing steps for the high-throughput phenotypic data from the field include semantic segmentation of phenotypic micro-features: An improved multi-scale convolutional coding-decoding semantic segmentation network was used to perform pixel-level semantic segmentation on the acquired crop canopy orthophotos and side-view projections. Key morphological micro-features for lodging resistance are extracted from the segmented image mask. These morphological micro-features include basal stem thickness, third internode length, stem fullness, and root soil uplift volume. The morphological micro-features are quantized into numerical vectors and loaded into the high-dimensional tensor space as input components of the phenotypic modality.
3. The seed lodging resistance breeding method based on big data according to claim 2, characterized in that, Employing cross-modal mechanics surrogate modeling techniques: Obtain field-measured destructive mechanical data, including crop stem puncture strength, critical breaking force, and root pull-out force; A nonlinear mapping model based on a multilayer perceptron is constructed, and the model is trained using the numerical vector of the morphological micro-features as input and the destructive mechanical data as supervision labels. In the breeding prediction stage, the trained nonlinear mapping model is used to transform the non-destructive image features of the test sample into virtual mechanical strength indicators, which are then used to evaluate the lodging resistance performance.
4. The seed lodging resistance breeding method based on big data according to claim 1, characterized in that, The micrometeorological environment data includes the instantaneous maximum wind speed, prevailing wind direction, short-term heavy rainfall, and soil volumetric moisture content at different soil depths in the field. In the spatiotemporal alignment and fusion process, a dynamic stress simulation field is constructed using the micrometeorological environment data. The combined wind and rain stress intensity encountered by crops at different growth stages is used as a dynamic weighting factor to weight and correct the expression effect value of genotype characteristics.
5. The seed lodging resistance breeding method based on big data according to claim 1, characterized in that, The breeding decision-making process includes path simulation steps based on reinforcement learning: Construct a virtual breeding environment that incorporates historical meteorological scenarios and genetic inheritance patterns; The action space of the agent is defined as the parental mating operation and the offspring selection strategy, and the state space is defined as the gene frequency distribution of the current generation population and the predicted lodging resistance trait. A multi-objective reward function is designed, which is composed of a weighted positive feedback of yield gain per unit area and a negative feedback of crop lodging rate, to guide the agent to autonomously explore the optimal breeding path in the virtual breeding environment.
6. The seed lodging resistance breeding method based on big data according to claim 5, characterized in that, The reinforcement learning process includes an automatic pruning mechanism: During the agent's exploration process, the probability of lodging risk for the current hybridization combination path is calculated in real time; When the predicted lodging risk probability of a certain genetic combination path exceeds a preset safety threshold, a pruning operation is automatically triggered to block the subsequent simulation calculation of that genetic combination path, and the combination is marked as a high-risk elimination combination and fed back to the parent selection database.
7. The seed lodging resistance breeding method based on big data according to claim 6, characterized in that, This also includes constructing lodging-resistant enhanced genomic selection models: The spatiotemporal hypergraph feature matrix is used as a fixed effect factor and fused into the hybrid linear model; Calculate the genomically estimated breeding value after fusion features, correct the prediction bias of lodging resistance traits caused by the traditional best linear unbiased prediction model relying solely on pedigree relationships, and output a sorted list of corrected lodging resistance breeding values.
8. The seed lodging resistance breeding method based on big data according to any one of claims 1 to 7, characterized in that, It also includes data preprocessing, specifically including: Kriging interpolation or long short-term memory recurrent neural networks are used to fill in phenotypic data with spatiotemporal gaps. Standardized fraction processing is performed on multi-source heterogeneous data to map physical quantities of different dimensions to a distribution range with a mean of zero and a standard deviation of one, thereby eliminating the impact of differences in data magnitude on the accuracy of tensor decomposition.
9. The seed lodging resistance breeding method based on big data according to any one of claims 1 to 7, characterized in that, The parameter range for model training is set as follows: The initial learning rate of the spatiotemporal hypergraph attention network and reinforcement learning model is set between one-thousandth and one-hundredth. The number of iterations for model training is set between 1,000 and 5,000, and an early stopping strategy is used to prevent overfitting. The rank parameter of tensor decomposition is dynamically adjusted according to the sparsity of the data, and is kept within the range of 10% to 30% of the feature dimension.
10. The seed lodging resistance breeding method based on big data according to any one of claims 1 to 7, characterized in that, Applications of big data fusion breeding in the field for corn or rice: For maize crops, the focus is on optimizing the weighted analysis of the number of adventitious root layers and the bending resistance of the internodes at the base of the stem; For rice crops, the focus is on optimizing the synergistic evaluation of the bending moment of the panicle neck and the canopy air permeability; The breeding decision-making process outputs recommended planting schemes for lodging-resistant new varieties adapted to specific ecological zones.