High-photosynthetic-efficiency sugarcane breeding method
By constructing a sugarcane F1 segregating population, collecting multi-dimensional data, and building a multi-task Bayesian fusion prediction model, the problem of long breeding cycles for high light efficiency sugarcane was solved, enabling early high-throughput screening and efficient breeding, thus improving breeding efficiency and accuracy.
Patent Information
- Application Number
- CN202511639893.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-11
- Publication Date
- 2026-02-06
AI Technical Summary
The breeding cycle for high light efficiency in sugarcane is relatively long, making it impossible to achieve high-throughput screening in the early stages, which restricts the breeding efficiency.
F1 segregating populations were constructed by screening high-photometric-efficiency parents and low-photometric-efficiency controls. Multidimensional data were collected throughout the entire growth period and integrated into a four-dimensional heterogeneous tensor. Features were extracted using PARAFAC2 sparse regularization. A multi-task Bayesian fusion prediction model was constructed by combining spatiotemporal attention mechanism and FvCB photosynthetic mechanism to conduct early screening of the new F1 population. Low-potential lines were eliminated through mid-term field validation and late-term genetic stability validation.
This approach enables early identification and efficient breeding of sugarcane with high light efficiency traits, reduces reliance on field phenotypic measurements, decreases the workload of ineffective experiments, improves breeding efficiency, and ensures the accuracy of predicted values and compliance with the laws of photosynthetic physics.
Smart Images

Figure CN121483384A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of seedling breeding technology, specifically a method for breeding sugarcane with high light efficiency. Background Technology
[0002] Photosynthesis is the core physiological process by which crops convert light energy into chemical energy. Its efficiency directly determines biomass accumulation and the formation of economic traits (such as yield and sugar content). High photosynthetic efficiency breeding has become a core direction for crop genetic improvement. Sugarcane, as the source of 80% of the world's sugar, relies on photosynthetic products for 60%-70% of its biomass. With the increasing demand for sugar and the constraints of arable land, improving its photosynthetic efficiency is the key to the sustainable development of the sugar industry.
[0003] Current methods for breeding sugarcane with high photosynthetic efficiency mainly focus on field phenotypic measurements. This involves screening parents with high photosynthetic efficiency phenotypes for hybridization, constructing segregating populations, and then repeatedly measuring photosynthetic physiological indicators throughout the population's growth period to select superior lines. However, this method relies on field phenotypic measurements throughout the entire sugarcane growth cycle, which lasts 10-12 months. Furthermore, photosynthetic phenotype measurements need to be performed in situ at specific growth stages using portable photosynthesis meters and other equipment. Each measurement is time-consuming and covers a limited sample size, resulting in a long breeding cycle and hindering early high-throughput screening, thus limiting breeding efficiency. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a method for breeding sugarcane with high light efficiency, which solves the problem that the current breeding cycle for sugarcane with high light efficiency is long, making it impossible to achieve early high-throughput screening and thus restricting breeding efficiency.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a method for breeding sugarcane with high light efficiency, comprising the following steps:
[0006] Screening treatment: High light efficiency parents and low light efficiency controls for sugarcane were screened, and all seedlings were standardized to construct an F1 segregating population;
[0007] Data collection: Genomic data, transcriptomic data, metabolomic data, dynamic photosynthetic phenotype data, and spatiotemporal environmental data were collected throughout the entire growth period for the F1 segregating population, high photosynthetic efficiency parents, and low photosynthetic efficiency controls to form multi-dimensional data.
[0008] Preprocessing: Standardize and handle outliers of multi-dimensional data, and integrate them into a four-dimensional heterogeneous tensor;
[0009] Extraction and Enhancement: Latent features are extracted from the four-dimensional heterogeneous tensor using a PARAFAC2-based sparse regularization method, and then spatiotemporal features are enhanced through a spatiotemporal attention mechanism to obtain a fused feature matrix;
[0010] Combined construction: By combining the FvCB photosynthetic mechanism parameters and the fusion feature matrix, a multi-task Bayesian fusion prediction model is constructed, and the predicted values are output for model validation.
[0011] Screening and validation: Early high-throughput screening of the new F1 segregating population was carried out based on a multi-task Bayesian fusion prediction model, followed by mid-term field dynamic phenotypic validation and late-term genetic stability validation to obtain sugarcane lines with high light efficiency.
[0012] By adopting the above technical solution, a segregating F1 population was constructed by screening high-photometric-efficiency parents and low-photometric-efficiency controls. Multi-dimensional data covering genetic, physiological, and environmental information were collected throughout the entire growth period. After preprocessing and integration into a four-dimensional heterogeneous tensor, features were extracted using PARAFAC2 sparse regularization and key features were enhanced using spatiotemporal attention mechanism. A predictive model was constructed based on the FvCB photosynthetic mechanism. Early screening of the new F1 population was carried out based on the model, eliminating low-potential lines without waiting for the entire growth period. Mid-term field validation and late-term genetic confirmation were then performed, thereby reducing the dependence on field phenotypic measurements throughout the entire growth period, reducing the workload of ineffective experiments, and improving the efficiency of early screening. This enabled early identification and efficient breeding of high-photometric-efficiency traits in sugarcane, solving the problem that the current breeding cycle for high-photometric-efficiency sugarcane is long and cannot achieve early high-throughput screening, thus restricting the breeding efficiency.
[0013] Preferably, the construction of the F1 segregating population specifically includes the following steps:
[0014] Sugarcane that meets the following criteria during the jointing stage: stable net photosynthetic rate ≥25μmol / m²・s, sucrose content ≥16%, and resistance to smut disease index ≤1, are selected as high photosynthetic efficiency parents.
[0015] Sugarcane that meets the criteria of net photosynthetic rate ≤18μmol / m²・s at the jointing stage and has completed whole-genome sequencing is selected as a low photosynthetic efficiency control.
[0016] Artificial pollination was carried out using the female and male parents from the high-light-efficiency parent line as a hybrid combination. After harvesting F1 generation seeds, tissue culture was used for rapid propagation on MS medium to obtain F1 generation seedlings. The MS medium contained 29.5-30.5 g / L sucrose and 6.8-7.2 g / L agar. The temperature for rapid tissue culture propagation was 24-26℃, the photoperiod was 15-17 h / d, and the light intensity was 2900-3100 lux.
[0017] High light-efficiency parent, low light-efficiency control and F1 generation seedlings were all treated into single-bud stems, which were then disinfected by soaking in a 480-520 times dilution of 48-52% carbendazim wettable powder for 28-32 minutes to obtain seedlings. The single-bud stems were required to have a stem diameter ≥2.5cm and full, undamaged buds.
[0018] The seedlings were hardened off for 7 days, and the F1 segregating population was obtained after transplanting. The 7-day hardening-off was carried out in a greenhouse under natural light with a humidity of 70%-80%.
[0019] Preferably, the genomic data includes whole-genome sequence data based on PacBio HiFi sequencing and genotyping data from an Illumina 100K SNP chip; the transcriptome data includes TPM values of photosynthetic pathway-related genes at four growth stages: seedling, tillering, jointing, and maturity; the photosynthetic pathway-related genes include genes related to carbon fixation and starch-sucrose metabolism pathways; the metabolome data includes the concentrations of photosynthetic metabolites and stress-resistance metabolites; the dynamic photosynthetic phenotypic data includes net photosynthetic rate, stomatal conductance, intercellular CO2 concentration, maximum photochemical efficiency, and actual photochemical efficiency; and the spatiotemporal environmental data includes light intensity, ambient temperature, atmospheric CO2 concentration, and canopy light distribution data collected by a drone.
[0020] Preferably, the standardization includes mean filling of genomic data, ComBat batch correction of transcriptome data, and Z-score conversion of metabolome data. The outlier handling includes 3σ rule removal of transcriptome data and box plot replacement of photosynthetic phenotype data, wherein the box plot replacement involves replacing outliers with the median of the index.
[0021] Preferably, obtaining the fused feature matrix specifically includes the following steps:
[0022] The alternating direction multiplier method is used to minimize the joint objective function of tensor reconstruction error and L1 regularization. The PARAFAC2 decomposition of the four-dimensional heterogeneous tensor is then performed to obtain low-rank latent features.
[0023] The dimension and regularization parameters of the low-rank latent features are determined by 5-fold cross-validation, the sample shared factor matrix and the reproductive period factor matrix are extracted, and the initial spatiotemporal feature matrix is obtained by outer product operation.
[0024] The initial spatiotemporal feature matrix is concatenated with the standardized spatiotemporal environment data to form a spatiotemporal feature embedding matrix.
[0025] An attention module is constructed using a spatiotemporal attention mechanism. The attention module is trained using the Adam optimizer and mean squared error loss function. The spatiotemporal attention weights are calculated and the spatiotemporal feature embedding matrix is weighted to obtain the fused feature matrix.
[0026] Preferably, the output model validation prediction values specifically include the following steps:
[0027] Based on measured net photosynthetic rate and intercellular CO2 concentration data, the Levenberg-Marquardt algorithm was used to fit the FvCB photosynthetic mechanism model, estimate the maximum carboxylation rate of Rubisco, electron transport rate, triose phosphate utilization rate and dark respiration rate, and form a mechanism parameter matrix.
[0028] Using the fusion feature matrix as input features and the mechanism parameter matrix as physiological constraints, a multi-task Bayesian fusion prediction model was constructed. With net photosynthetic rate and maximum photochemical efficiency as dual prediction tasks, the Hamilton Monte Carlo algorithm was used to sample parameters and obtain the sampling results.
[0029] The sampling results are subjected to convergence diagnosis. The mean of the converged sampling results is taken as the parameters of the multi-task Bayesian fusion prediction model. The net photosynthetic rate and maximum photochemical efficiency of the F1 segregated population are output as predicted values for model validation.
[0030] Preferably, the fitting satisfies a coefficient of determination ≥ 0.90, and the prior distribution of the parameters of the multi-task Bayesian fusion prediction model is that the regression coefficients follow a normal distribution, and the error variance and coefficient variance parameters follow an inverse gamma distribution.
[0031] Preferably, the early high-throughput screening specifically includes the following steps:
[0032] For the new F1 segregating population, genomic SNP data and transcriptome data were collected during the seedling stage. A simplified four-dimensional heterogeneous tensor was constructed according to the preprocessing steps, and input into a multi-task Bayesian fusion prediction model to predict the net photosynthetic rate, maximum photochemical efficiency and Rubisco maximum carboxylation rate of the new F1 segregating population during the jointing stage, and new predicted values were obtained.
[0033] We screened for new predicted values that met the following criteria: net photosynthetic rate at the jointing stage ≥ 26 μmol / m²·s, maximum photochemical efficiency ≥ 0.84, and maximum Rubisco carboxylation rate ≥ 80 μmol / m²·s.
[0034] The expression levels of photosynthetic core genes in the selected lines were verified by qPCR. Lines with relative expression levels below the mean were removed to obtain qualified lines. The photosynthetic core genes include the RBCS gene, the PEPC gene, and the SPS gene.
[0035] Preferably, the mid-term field dynamic phenotypic verification specifically includes the following steps:
[0036] After planting qualified lines, the net photosynthetic rate and maximum photochemical efficiency were measured weekly from the tillering stage to the jointing stage. The actual maximum carboxylation rate of Rubisco was fitted to obtain the actual value. Lines with deviations of less than 5% between the actual value and the new predicted value were screened.
[0037] The sucrose content was measured at the maturity stage of the deviation lines, and the sucrose content was calculated as sucrose content = 1.0625 × sucrose content - 7.7065. Lines with sucrose content ≥ 16% were retained to obtain the mid-term lines.
[0038] Preferably, the later genetic stability verification specifically includes the following steps:
[0039] Whole-genome resequencing was performed on the mid-term lines with a sequencing depth ≥20× to verify the homozygosity of photosynthesis-related QTLs. Mid-term lines with homozygosity ≥90% were retained to obtain homozygous lines. The photosynthesis-related QTLs include QTLs that control the maximum carboxylation rate of Rubisco.
[0040] After planting homozygous lines, variety comparison tests were conducted to determine the net photosynthetic rate, maximum photochemical efficiency, sucrose content, and yield per acre throughout the entire growth period, thus obtaining the experimental lines.
[0041] High-efficiency sugarcane lines were selected based on their net photosynthetic rate at the jointing stage being ≥27 μmol / m²・s, maximum photochemical efficiency being ≥0.85, sucrose content being ≥16.5%, and yield per mu being ≥10 tons.
[0042] This invention provides a method for breeding sugarcane with high light efficiency. It has the following beneficial effects:
[0043] 1. This invention constructs an F1 segregating population by screening high-photometric-efficiency parents and low-photometric-efficiency controls, and then collects multi-dimensional data covering genetic, physiological, and environmental information throughout the entire growth period. After preprocessing and integrating into a four-dimensional heterogeneous tensor, features are extracted using PARAFAC2 sparse regularization and key features are enhanced using spatiotemporal attention mechanism. Combined with the FvCB photosynthetic mechanism, a predictive model is constructed. Based on the model, early screening of the new F1 population can be carried out, eliminating low-potential lines without waiting for the entire growth period. After mid-term field verification and late-term genetic confirmation, the dependence on field phenotypic determination throughout the entire growth period is reduced, the workload of ineffective experiments is reduced, and the efficiency of early screening is improved. This realizes the early identification and efficient breeding of sugarcane with high photosynthetic efficiency traits, while reducing the field planting area and management costs.
[0044] 2. This invention extracts potential correlation features from genomic, transcriptomic, metabolomic, and spatiotemporal environmental data using PARAFAC2 sparse regularization, enhances the feature contribution of key scenarios such as high light environment during the jointing stage using spatiotemporal attention mechanism, and constructs a multi-task Bayesian fusion prediction model by combining the FvCB photosynthetic mechanism model. This improves the determination coefficient of the prediction of net photosynthetic rate and maximum photochemical efficiency across the reproductive stage. By constraining the physiological boundary of the mechanism parameters, unreasonable results of predicting net photosynthetic rate exceeding the physiological limit of sugarcane are avoided, ensuring that the predicted values are both accurate and in line with the laws of photosynthetic physics.
[0045] 3. This invention uses a multi-task Bayesian fusion prediction model to synergistically predict net photosynthetic rate, maximum photochemical efficiency, and sucrose-related metabolite concentrations. It leverages the genetic correlation between traits to achieve information complementarity. Mid-term verification retains lines with sucrose content ≥16%. In the later stage, the yield per mu (a Chinese unit of area, approximately 0.067 hectares) is ≥10 tons and sucrose content ≥16.5% as indicators. The obtained high photosynthetic efficiency lines not only have a net photosynthetic rate ≥27 μmol / m²・s at the jointing stage, but also simultaneously meet the requirements of sugarcane production for both quality and yield. Thus, through multi-task modeling and full-cycle trait monitoring, the synergistic improvement of high photosynthetic efficiency, high sucrose content, and high yield is achieved. Attached Figure Description
[0046] Figure 1 This is a flowchart of a high-light-efficiency sugarcane breeding method proposed in this invention. Detailed Implementation
[0047] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0048] Please see the appendix Figure 1 This invention provides a method for breeding sugarcane with high light efficiency, comprising the following steps:
[0049] Screening treatment: High light efficiency parents and low light efficiency controls for sugarcane were screened, and all seedlings were standardized to construct an F1 segregating population;
[0050] Furthermore, the construction of the F1 segregating population specifically includes the following steps:
[0051] Sugarcane that meets the following criteria during the jointing stage: stable net photosynthetic rate ≥25μmol / m²・s, sucrose content ≥16%, and resistance to smut disease index ≤1, are selected as high photosynthetic efficiency parents.
[0052] Sugarcane that meets the criteria of net photosynthetic rate ≤18μmol / m²・s at the jointing stage and has completed whole-genome sequencing is selected as a low photosynthetic efficiency control.
[0053] Artificial pollination was carried out using the female and male parents from the high-light-efficiency parent line as a hybrid combination. After harvesting F1 generation seeds, tissue culture was used for rapid propagation on MS medium to obtain F1 generation seedlings. The MS medium contained 29.5-30.5 g / L sucrose and 6.8-7.2 g / L agar. The temperature for rapid tissue culture propagation was 24-26℃, the photoperiod was 15-17 h / d, and the light intensity was 2900-3100 lux.
[0054] High light-efficiency parent, low light-efficiency control and F1 generation seedlings were all treated into single-bud stems, which were then disinfected by soaking in a 480-520 times dilution of 48-52% carbendazim wettable powder for 28-32 minutes to obtain seedlings. The single-bud stems were required to have a stem diameter ≥2.5cm and full, undamaged buds.
[0055] The seedlings were hardened off for 7 days, and the F1 segregating population was obtained after transplanting. The 7-day hardening-off was carried out in a greenhouse under natural light with a humidity of 70%-80%.
[0056] Specifically, the purpose of constructing the F1 segregating population is to obtain a segregating population with rich genetic diversity through genetic recombination between high-light-efficiency parents and low-light-efficiency controls, so as to provide a material basis for subsequent multi-dimensional data collection and genetic analysis of high-light-efficiency traits.
[0057] Generally, the selection of high photosynthetic efficiency parents requires a comprehensive consideration of photosynthetic efficiency, quality, and disease resistance indicators. A stable net photosynthetic rate of ≥25 μmol / m²·s at the jointing stage can ensure the initial frequency of high photosynthetic efficiency alleles in the population. A sucrose content of ≥16% ensures the genetic basis for excellent quality traits, and a disease resistance index of ≤1 grade helps to reduce the interference of disease on photosynthetic efficiency. Low photosynthetic efficiency controls need to meet the requirement of a net photosynthetic rate of ≤18 μmol / m²·s at the jointing stage to form phenotypic differences. Their whole-genome sequencing characteristics can provide convenience for the analysis of basal genetic information, facilitating subsequent marker development and QTL mapping.
[0058] In one possible implementation, the hybridization combination uses the female and male parents from the high-light-efficiency parent line for artificial pollination. This allows for control of the hybridization process, reduces interference from foreign pollen, and ensures the genetic purity of the F1 generation seeds. The harvested F1 generation seeds are then rapidly propagated through tissue culture on MS medium. 29.5-30.5 g / L sucrose provides a carbon source for the seedlings, while 6.8-7.2 g / L agar maintains the medium's solidification. A temperature of 24-26℃, a light duration of 15-17 h / d, and a light intensity of 2900-3100 lux simulate the suitable environment for sugarcane seedling growth, promoting the uniform and robust growth of the F1 generation seedlings.
[0059] In some embodiments, when high-light-efficiency parents, low-light-efficiency controls, and F1 generation seedlings are treated to produce single-bud seed stems, the standard of stem diameter ≥2.5cm and plump, undamaged buds ensures that the seed stems store sufficient nutrients and improves the germination rate. Disinfection by soaking in a 48-52% carbendazim wettable powder solution at a dilution of 480-520 times for 28-32 minutes can kill pathogens on the surface of the seed stems and reduce the risk of seedling diseases.
[0060] Specifically, acclimatizing seedlings for 7 days using natural light in a greenhouse while maintaining 70%-80% humidity allows them to gradually adapt to external environmental conditions, reducing post-transplant stress and improving survival rates. The resulting F1 segregating population exhibits a clear genetic background and significant segregation of light-related traits, providing reliable experimental materials for subsequent data collection throughout the entire growth period and model construction.
[0061] Data collection: Genomic data, transcriptomic data, metabolomic data, dynamic photosynthetic phenotype data, and spatiotemporal environmental data were collected throughout the entire growth period for the F1 segregating population, high photosynthetic efficiency parents, and low photosynthetic efficiency controls to form multi-dimensional data.
[0062] Furthermore, the genomic data includes whole-genome sequence data based on PacBio HiFi sequencing and genotyping data from the Illumina 100K SNP chip; the transcriptome data includes TPM values of photosynthetic pathway-related genes at four growth stages: seedling stage, tillering stage, jointing stage, and maturity stage; the photosynthetic pathway-related genes include genes related to carbon fixation pathway and starch-sucrose metabolism pathway; the metabolome data includes the concentrations of photosynthetic metabolites and stress-resistance metabolites; the dynamic photosynthetic phenotypic data includes net photosynthetic rate, stomatal conductance, intercellular CO2 concentration, maximum photochemical efficiency, and actual photochemical efficiency; and the spatiotemporal environmental data includes light intensity, ambient temperature, atmospheric CO2 concentration, and canopy light distribution data collected by UAV.
[0063] Specifically, the data collection aims to obtain multi-dimensional information on the F1 segregating population and the high-light-efficiency parent with low-light-efficiency control, to analyze the genetic basis and environmental response mechanism of the high-light-efficiency trait, and to provide comprehensive materials for subsequent data integration and model construction.
[0064] Generally, data collection throughout the entire growth period needs to cover key stages of crop growth to ensure the capture of dynamic changes in traits. Genomic data is obtained by PacBio HiFi sequencing to obtain whole genome sequences, which can provide high-accuracy long read information and facilitate the detection of large fragment variations. Combined with genotyping data from the Illumina 100KSNP chip, high-density marker genotyping can be achieved, laying the foundation for genetic map construction and QTL localization.
[0065] Transcriptome data focused on four growth stages: seedling, tillering, jointing, and maturity, with a particular emphasis on collecting TPM values for genes related to photosynthetic pathways. In some examples, the expression levels of carbon fixation pathway genes such as RBCS and PEPC, and starch and sucrose metabolism pathway genes such as SPS, were directly correlated with photosynthetic efficiency and nutrient accumulation, reflecting the photosynthetic regulatory network at different growth stages.
[0066] Metabolome data covers photosynthetic metabolites such as triose phosphate sucrose, and stress-resistant metabolites such as proline. The concentration changes of these substances are closely related to photosynthetic efficiency and environmental adaptability. Dynamic photosynthetic phenotype data includes net photosynthetic rate, stomatal conductance, etc., which are measured by a portable photosynthesis instrument at key nodes of each growth stage and can reflect the photosynthetic physiological state of plants in real time.
[0067] Spatio-temporal environmental data records light intensity, environmental temperature, and atmospheric CO2 concentration through environmental sensors. Combining with the canopy light distribution data collected by drones at regular intervals, it can analyze the impact of environmental factors on photosynthetic traits. The multi-dimensional data collected in this way can comprehensively associate genotypes, phenotypes, and environments, providing reliable inputs for the construction of subsequent four-dimensional heterogeneous tensors.
[0068] Preprocessing: Standardize and handle outliers for multi-dimensional data, and integrate them into a four-dimensional heterogeneous tensor;
[0069] Furthermore, the standardization includes mean filling of genomic data, ComBat batch correction of transcriptome data, and Z-score transformation of metabolome data. The outlier handling includes removing outliers from transcriptome data using the 3σ rule, and replacing outliers in photosynthetic phenotype data using the boxplot method. The boxplot method replacement is to replace outliers with the median of this metric.
[0070] Specifically, the purpose of preprocessing is to eliminate the heterogeneity and noise of multi-dimensional data and provide standardized inputs for subsequent feature extraction and model construction. Generally, the standardization needs to adapt methods for different data types: genomic data uses mean filling to supplement missing sites; transcriptome uses ComBat to correct batch effects; metabolome uses Z-score transformation, and the formula is , where is the metabolite concentration of a certain sample, is the mean of this metabolite in the population, is its standard deviation, so as to unify the dimension.
[0071] Outlier handling is divided into two categories: for transcriptome, the 3σ rule is used. When the outliers are removed, is the gene TPM value, is the mean of this gene TPM, is the standard deviation; for photosynthetic phenotypes, the boxplot method is used, and outliers (<Q1 - 1.5×IQR or >Q3 + 1.5×IQR, IQR = Q3 - Q1) are replaced with the median of this metric. Q1 and Q3 are quartiles.
[0072] Integrate the processed genomic, transcriptomic, metabolomic, photosynthetic phenotype, and spatio-temporal environmental data into a four-dimensional heterogeneous tensor. The dimensions correspond to samples, growth stages, feature types, and environmental factors, laying a foundation for subsequent PARAFAC2 feature extraction.
[0073] Extraction and Enhancement: Latent features are extracted from the four-dimensional heterogeneous tensor using a PARAFAC2-based sparse regularization method, and then spatiotemporal features are enhanced through a spatiotemporal attention mechanism to obtain a fused feature matrix;
[0074] Furthermore, obtaining the fused feature matrix specifically includes the following steps:
[0075] The alternating direction multiplier method is used to minimize the joint objective function of tensor reconstruction error and L1 regularization. The PARAFAC2 decomposition of the four-dimensional heterogeneous tensor is then performed to obtain low-rank latent features.
[0076] The dimension and regularization parameters of the low-rank latent features are determined by 5-fold cross-validation, the sample shared factor matrix and the reproductive period factor matrix are extracted, and the initial spatiotemporal feature matrix is obtained by outer product operation.
[0077] The initial spatiotemporal feature matrix is concatenated with the standardized spatiotemporal environment data to form a spatiotemporal feature embedding matrix.
[0078] An attention module is constructed using a spatiotemporal attention mechanism. The attention module is trained using the Adam optimizer and mean squared error loss function. The spatiotemporal attention weights are calculated and the spatiotemporal feature embedding matrix is weighted to obtain the fused feature matrix.
[0079] Specifically, the core features for mining high-efficiency traits from high-dimensional heterogeneous tensors are extracted and enhanced, and the contribution of key spatiotemporal information is strengthened to provide accurate input for subsequent prediction models and avoid redundant interference from high-dimensional data.
[0080] Generally, four-dimensional heterogeneous tensors have high dimensionality and complex features, requiring dimensionality reduction through PARAFAC2 decomposition. Specifically, the alternating direction multiplier method is used to minimize the joint objective function of tensor reconstruction error and L1 regularization, as shown in the following formula: ,middle, This is a preprocessed four-dimensional heterogeneous tensor (dimensions correspond to samples, features, omics types, and reproductive periods). This is a sample shared factor matrix, where each row represents a latent feature of one sample. A diagonal matrix specific to omics, characterizing the contributions of different omics disciplines. This is a matrix of omics feature factors. This is a factor matrix representing the reproductive period. For regularization parameters, The Frobenius norm measures the reconstruction error. Using the L1 norm, feature sparsity is achieved. This formula can remove redundant features while reducing dimensionality, resulting in low-rank latent features.
[0081] The dimensionality of low-rank latent features was determined using 5-fold cross-validation. The potential dimension candidate range is 20-100. The candidate range is 0.001-0.01. After selecting the optimal dimension and parameters, the extraction is performed. and An outer product operation is performed to obtain an initial spatiotemporal feature matrix, with each row corresponding to a feature combination of one sample and reproductive period. Alternatively, the initial spatiotemporal feature matrix is concatenated with standardized spatiotemporal environmental data (light, temperature, etc.) to form a spatiotemporal feature embedding matrix, integrating genetic and environmental information.
[0082] A three-head attention module is constructed using a spatiotemporal attention mechanism, first mapping the embedding matrix to the query. ,key ,value The attention weight formula is: ,in, Scaling factor ( (To avoid gradient vanishing, the Softmax function normalizes the weights, taking into account the feature dimension.) Calculate the correlation between features. The module is trained using the Adam optimizer (learning rate 1e-4) and mean squared error loss function. Spatiotemporal attention weights are calculated and... The weighted summation yields a fusion feature matrix, which enhances the feature contribution of key scenarios such as the growth stage and high light intensity, providing high-quality input for subsequent multi-task Bayesian models.
[0083] Combined construction: By combining the FvCB photosynthetic mechanism parameters and the fusion feature matrix, a multi-task Bayesian fusion prediction model is constructed, and the predicted values are output for model validation.
[0084] Furthermore, the output model validation predictions specifically include the following steps:
[0085] Based on measured net photosynthetic rate and intercellular CO2 concentration data, the Levenberg-Marquardt algorithm was used to fit the FvCB photosynthetic mechanism model, estimate the maximum carboxylation rate of Rubisco, electron transport rate, triose phosphate utilization rate and dark respiration rate, and form a mechanism parameter matrix.
[0086] Using the fusion feature matrix as input features and the mechanism parameter matrix as physiological constraints, a multi-task Bayesian fusion prediction model was constructed. With net photosynthetic rate and maximum photochemical efficiency as dual prediction tasks, the Hamilton Monte Carlo algorithm was used to sample parameters and obtain the sampling results.
[0087] The sampling results are subjected to convergence diagnosis. The mean of the converged sampling results is taken as the parameters of the multi-task Bayesian fusion prediction model. The net photosynthetic rate and maximum photochemical efficiency of the F1 segregated population are output as predicted values for model validation.
[0088] Furthermore, the fitting satisfies a determination coefficient ≥ 0.90, and the prior distribution of the parameters of the multi-task Bayesian fusion prediction model is that the regression coefficients follow a normal distribution, and the error variance and coefficient variance parameters follow an inverse gamma distribution.
[0089] Specifically, by combining the construction of a data-driven model based on photosynthetic physiological mechanisms, the reliability and physiological rationality of high light efficiency trait predictions can be improved, avoiding predictions that violate the laws of photosynthesis from pure data models, and providing accurate model support for subsequent screening and verification.
[0090] Generally, the prediction of photosynthetic traits needs to be anchored to physiological mechanisms. The FvCB photosynthetic mechanism model can characterize the three rate-limiting steps of sugarcane photosynthesis, and its formula is as follows: ,in, To measure the net photosynthetic rate, The maximum carboxylation rate of Rubisco is 40-100 μmol / m²·s. This represents the intercellular CO2 concentration (measured value). The CO2 compensation point is 40-60 μmol / mol. Rubisco's Michaelis constant for CO2 is 200-400 μmol / mol. This represents the intercellular O2 concentration, defaulted to 210 mmol / mol. Rubisco's Michaelis constant for O2 is 200-300 mmol / mol. The electron transport rate is 80-200 μmol / m²·s. The utilization rate of triose phosphate is 20-60 μmol / m²·s. The dark respiration rate is 1-3 μmol / m²·s. Specifically, the Levenberg-Marquardt algorithm was used to fit the model, based on actual measurements. and Given the input, estimate , , , A mechanism parameter matrix is formed, and the fitting must satisfy the coefficient of determination ≥ 0.90 to ensure that the mechanism parameters are consistent with the actual physiological process.
[0091] As an alternative, the multi-task Bayesian fusion prediction model uses the extracted and enhanced fusion feature matrix as input features and the mechanistic parameter matrix as physiological constraints to predict net photosynthetic rate and maximum photochemical efficiency in a dual-task manner. Its likelihood function formula is as follows: ,in, To fuse the feature matrix, For the mechanism parameter matrix, , These are the regression coefficients of fusion characteristics on net photosynthetic rate and maximum photochemical efficiency, respectively. , These are the regression coefficients of the mechanistic parameters on the two traits, respectively. , The error term (following a normal distribution) ).
[0092] In some embodiments, the prior distribution of the model parameters is set as: regression coefficients , , , Follows a normal distribution Error variance , With coefficient variance Follows an inverse gamma distribution This introduces uninformed priors to reduce the impact of subjective assumptions on prediction.
[0093] Hamiltonian Monte Carlo algorithm was used for parameter sampling, with three sampling chains, each sampling 15,000 times. The first 5,000 samples were discarded during the combustion period to eliminate the influence of initial values. Convergence diagnostics were performed on the sampling results, requiring all parameters to have an Rhat < 1.05 and an effective sample size ESS > 1000 to ensure the reliability of the sampling results. The mean of the converged sampling results was taken as the model parameters, and the net photosynthetic rate and maximum photochemical efficiency of the F1 segregating population were output as predicted values for model validation. These predicted values can be compared with measured values to verify the model accuracy.
[0094] Screening and validation: Early high-throughput screening of the new F1 segregating population was carried out based on a multi-task Bayesian fusion prediction model, followed by mid-term field dynamic phenotypic validation and late-term genetic stability validation to obtain sugarcane lines with high light efficiency.
[0095] Furthermore, the early high-throughput screening specifically includes the following steps:
[0096] For the new F1 segregating population, genomic SNP data and transcriptome data were collected during the seedling stage. A simplified four-dimensional heterogeneous tensor was constructed according to the preprocessing steps, and input into a multi-task Bayesian fusion prediction model to predict the net photosynthetic rate, maximum photochemical efficiency and Rubisco maximum carboxylation rate of the new F1 segregating population during the jointing stage, and new predicted values were obtained.
[0097] We screened for new predicted values that met the following criteria: net photosynthetic rate at the jointing stage ≥ 26 μmol / m²·s, maximum photochemical efficiency ≥ 0.84, and maximum Rubisco carboxylation rate ≥ 80 μmol / m²·s.
[0098] The expression levels of photosynthetic core genes in the selected lines were verified by qPCR. Lines with relative expression levels below the mean were removed to obtain qualified lines. The photosynthetic core genes include the RBCS gene, the PEPC gene, and the SPS gene.
[0099] Furthermore, the mid-term field dynamic phenotypic validation specifically includes the following steps:
[0100] After planting qualified lines, the net photosynthetic rate and maximum photochemical efficiency were measured weekly from the tillering stage to the jointing stage. The actual maximum carboxylation rate of Rubisco was fitted to obtain the actual value. Lines with deviations of less than 5% between the actual value and the new predicted value were screened.
[0101] The sucrose content was measured at the maturity stage of the deviation lines, and the sucrose content was calculated as sucrose content = 1.0625 × sucrose content - 7.7065. Lines with sucrose content ≥ 16% were retained to obtain the mid-term lines.
[0102] Furthermore, the subsequent genetic stability verification specifically includes the following steps:
[0103] Whole-genome resequencing was performed on the mid-term lines with a sequencing depth ≥20× to verify the homozygosity of photosynthesis-related QTLs. Mid-term lines with homozygosity ≥90% were retained to obtain homozygous lines. The photosynthesis-related QTLs include QTLs that control the maximum carboxylation rate of Rubisco.
[0104] After planting homozygous lines, variety comparison tests were conducted to determine the net photosynthetic rate, maximum photochemical efficiency, sucrose content, and yield per acre throughout the entire growth period, thus obtaining the experimental lines.
[0105] High-efficiency sugarcane lines were selected based on their net photosynthetic rate at the jointing stage being ≥27 μmol / m²・s, maximum photochemical efficiency being ≥0.85, sucrose content being ≥16.5%, and yield per mu being ≥10 tons.
[0106] Specifically, the screening and validation process involves a multi-stage process: early high-throughput screening, mid-term field phenotypic validation, and late-term genetic stability validation. This process precisely selects genetically stable, high-photosynthetic efficiency, and high-quality lines from the new F1 segregating population, ensuring that the breeding results meet actual production needs. At the same time, it connects with the multi-task Bayesian fusion prediction model trained earlier to achieve a closed loop between model prediction and field validation.
[0107] Generally, early high-throughput screening relies on seedling data for rapid initial screening to reduce the workload and cost of subsequent field trials. Specifically, for the new F1 segregating population (the breeding population from the same parental origin as the previously mentioned F1 segregating population), seedling genomic SNP and transcriptome data are collected. Following the preprocessing steps described above (standardization, outlier removal), a simplified four-dimensional heterogeneous tensor is constructed. This tensor retains only genomic SNP features and transcriptome photosynthetic pathway gene features closely associated with high photosynthetic efficiency traits, avoiding redundant information interference. The simplified tensor is input into a multi-task Bayesian fusion prediction model, outputting the net photosynthetic rate, maximum photochemical efficiency, and maximum Rubisco carboxylation rate at the jointing stage of the new F1 segregating population, which are the new predicted values. As an alternative, the screening criteria are set as a net photosynthetic rate ≥26 μmol / m²·s, a maximum photochemical efficiency ≥0.84, and a maximum Rubisco carboxylation rate ≥80 μmol / m²·s at the jointing stage. This criterion is determined based on the model validation results described above and balances screening efficiency and accuracy. In some embodiments, qPCR was used to verify the expression levels of photosynthetic core genes in the screened lines. The core genes include the RBCS gene (which regulates Rubisco synthesis), the PEPC gene (which participates in carbon fixation), and the SPS gene (which affects sucrose synthesis). The average expression levels of these genes in the screened lines were calculated, and lines with relative expression levels lower than the average were removed to obtain qualified lines. False positive lines with abnormal gene expression were further excluded.
[0108] During the mid-term field dynamic phenotypic validation phase, qualified lines are generally planted in at least two ecological zones, with three replicates per line and a plot area of 10 m², to verify the phenotypic stability of the lines under different environments. Specifically, during the tillering to jointing stage, the critical photosynthetic period of sugarcane, the net photosynthetic rate and maximum photochemical efficiency are measured weekly. The actual maximum carboxylation rate of Rubisco is fitted using the aforementioned FvCB model to obtain the actual value; the deviation rate between the actual value and the new predicted value is calculated using the formula: ,in, The actual maximum carboxylation rate of Rubisco fitted in the field. To obtain a new predicted value for the maximum carboxylation rate of Rubisco from the early screening output, this formula was used to quantify the consistency between the phenotype and the prediction, and to screen out deviation lines with a deviation rate of <5%. At maturity, the sucrose content of the juice was measured (using an Abbe refractometer), and the sucrose content was calculated using the formula: sucrose content = 1.0625 × sucrose content - 7.7065, where sucrose content is the mass fraction of soluble solids in the juice, and 1.0625 and 7.7065 are empirical coefficients for calculating sucrose content. Lines with a sucrose content ≥16% were retained to obtain mid-stage lines, ensuring that high photosynthetic efficiency lines also possess excellent quality.
[0109] Later-stage genetic stability verification requires confirmation of line stability from both molecular and field phenotypic perspectives. Specifically, whole-genome resequencing is performed on mid-stage lines with a sequencing depth ≥20× to accurately detect the genotypes of photosynthesis-related QTLs—with a focus on validating QTLs controlling the maximum carboxylation rate of Rubisco (such as qP10). The homozygosity of this QTL locus is calculated, and mid-stage lines with a homozygosity ≥90% are retained to obtain homozygous lines, reducing phenotypic segregation caused by heterozygous genotypes. In some embodiments, homozygous lines were planted in at least three ecological zones for comparative trials, with three replicates for each line. Net photosynthetic rate and maximum photochemical efficiency were measured monthly throughout the growth period, and sucrose content and yield per acre were measured at maturity to obtain experimental lines. Finally, experimental lines that met the following criteria were selected: net photosynthetic rate ≥27 μmol / m²·s at the jointing stage, maximum photochemical efficiency ≥0.85, sucrose content ≥16.5%, and yield ≥10 tons per acre. These were identified as high photosynthetic efficiency sugarcane lines. These lines possess both high photosynthetic efficiency and good yield and quality, while also having a stable genetic background, which can meet the needs of subsequent production applications.
[0110] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A high light efficiency sugarcane breeding method, characterized by, The method comprises the following steps: Screening treatment: screening high photosynthetic efficiency parents and low photosynthetic efficiency controls of sugarcane, and standardizing all seedlings to construct an F1 separation population; Data collection: collecting genomic data, transcriptomic data, metabolomic data, dynamic photosynthetic phenotype data and spatiotemporal environment data of the whole growth period of the F1 separation population and the high photosynthetic efficiency parents and the low photosynthetic efficiency controls to form multidimensional data; Preprocessing: standardizing and processing outliers of the multidimensional data to integrate into a four-dimensional heterogeneous tensor; Extraction and strengthening: extracting potential features from the four-dimensional heterogeneous tensor by using a sparse regularization method based on PARAFAC2, and then strengthening spatiotemporal features by using a spatiotemporal attention mechanism to obtain a fusion feature matrix; Combination construction: combining FvCB photosynthetic mechanism parameters and the fusion feature matrix to construct a multi-task Bayesian fusion prediction model to output prediction values for model verification; Screening verification: screening new F1 separation populations at an early stage by using the multi-task Bayesian fusion prediction model, and then verifying the dynamic phenotypes in the middle stage and the genetic stability in the later stage to obtain high photosynthetic efficiency sugarcane strains.
2. The method for breeding high light efficiency sugarcane according to claim 1, characterized in that: The construction of the F1 separation population specifically comprises the following steps: Screening sugarcane that meets the conditions of net photosynthetic rate stability ≥ 25 μmol / m²・s at the jointing stage, sucrose content ≥ 16%, and sugarcane smut disease index ≤ 1 grade as high photosynthetic efficiency parents; Screening sugarcane that meets the conditions of net photosynthetic rate ≤ 18 μmol / m²・s at the jointing stage and complete whole genome sequencing as low photosynthetic efficiency controls; Artificial pollination is performed by using female parents and male parents in the high photosynthetic efficiency parents as hybrid combinations, F1 generation seeds are harvested, tissue culture and rapid propagation are performed by using MS medium, F1 generation seedlings are obtained, the MS medium contains 29.5-30.5 g / L sucrose and 6.8-7.2 g / L agar, and the temperature of the tissue culture and rapid propagation is 24-26℃, the light duration is 15-17 h / d, and the light intensity is 2900-3100 lux; The high photosynthetic efficiency parents, the low photosynthetic efficiency controls and the F1 generation seedlings are all treated as single-bud seed stems, are soaked in 480-520 times of 48-52% carbendazim wettable powder for 28-32 min for disinfection to obtain seedling bodies, and the single-bud seed stems meet the conditions of stem diameter ≥ 2.5 cm and full and undamaged bud eyes; The seedling bodies are subjected to 7d hardening, and the F1 separation population is obtained after transplanting, and the 7d hardening is greenhouse natural light acclimation and humidity of 70%-80%.
3. The method for breeding high light efficiency sugarcane according to claim 1, characterized in that: The genomic data includes whole genome sequence data based on PacBio HiFi sequencing and typing data of Illumina 100K SNP chips, the transcriptomic data includes TPM values of photosynthetic pathway related genes in four growth periods of seedling stage, tillering stage, jointing stage and mature stage, the photosynthetic pathway related genes include carbon fixation pathway and starch sucrose metabolism pathway related genes, the metabolomic data includes concentrations of photosynthetic metabolites and stress resistance metabolites, the dynamic photosynthetic phenotype data includes net photosynthetic rate, stomatal conductance, intercellular CO2 concentration, maximum photochemical efficiency and actual photochemical efficiency, and the spatiotemporal environmental data includes light intensity, environmental temperature, atmospheric CO2 concentration and crown light distribution data collected by a drone.
4. The method for breeding high light efficiency sugarcane according to claim 1, characterized in that: The standardization includes mean filling of genomic data, ComBat batch correction of transcriptomic data and Z-score conversion of metabolomic data, and the outlier processing includes 3σ rule rejection of transcriptomic data and boxplot method replacement of photosynthetic phenotype data, wherein the boxplot method replacement is to replace the outliers with the median of the index.
5. The method for breeding high light efficiency sugarcane as claimed in claim 1, wherein: The obtaining of the fusion feature matrix specifically includes the following steps: performing PARAFAC2 decomposition on the four-dimensional heterogeneous tensor by minimizing a joint objective function of tensor reconstruction error and L1 regularization, to obtain low-rank latent features; determining the dimension and regularization parameter of the low-rank latent features through 5-fold cross-validation, extracting a sample sharing factor matrix and a growth period factor matrix, and performing outer product operation to obtain an initial spatiotemporal feature matrix; splicing the initial spatiotemporal feature matrix and the standardized spatiotemporal environmental data to form a spatiotemporal feature embedding matrix; constructing an attention module through a spatiotemporal attention mechanism, training the attention module using an Adam optimizer and a mean square error loss function, calculating spatiotemporal attention weights and weighting the spatiotemporal feature embedding matrix to obtain a fusion feature matrix.
6. The high light efficiency sugarcane breeding method according to claim 1, characterized by the fact that: The output model verification prediction value specifically includes the following steps: based on the measured net photosynthetic rate and intercellular CO2 concentration data, fitting the FvCB photosynthetic mechanism model using the Levenberg-Marquardt algorithm to estimate the maximum Rubisco carboxylation rate, electron transport rate, triose phosphate utilization rate and dark respiration rate, and forming a mechanism parameter matrix; taking the fusion feature matrix as input features and the mechanism parameter matrix as physiological constraints, constructing a multi-task Bayesian fusion prediction model with net photosynthetic rate and maximum photochemical efficiency as double prediction tasks, and performing parameter sampling using the Hamilton Monte Carlo algorithm to obtain sampling results; performing convergence diagnosis on the sampling results, taking the mean value of the converged sampling results as the multi-task Bayesian fusion prediction model parameters, and outputting the net photosynthetic rate and maximum photochemical efficiency of the F1 isolated population as the model verification prediction value.
7. The method for breeding high light efficiency sugarcane according to claim 6, characterized in that: The fitting satisfies a determination coefficient ≥ 0.90, and the parameter prior distribution of the multi-task Bayesian fusion prediction model is that the regression coefficients obey a normal distribution and the error variance and coefficient variance parameters obey an inverse gamma distribution.
8. The high light efficiency sugarcane breeding method according to claim 1, characterized by the fact that: The early high-throughput screening specifically includes the following steps: For the new F1 separation population, the genome SNP data and the transcriptome data at the seedling stage are collected, a simplified four-dimensional heterogeneous tensor is constructed according to the pretreatment steps, and a multi-task Bayesian fusion prediction model is input to predict the net photosynthetic rate, the maximum photochemical efficiency and the maximum Rubisco carboxylation rate of the new F1 separation population at the jointing stage, and new predicted values are obtained. Strains that meet the conditions of net photosynthetic rate ≥ 26 μmol / m²・s, maximum photochemical efficiency ≥ 0.84 and maximum Rubisco carboxylation rate ≥ 80 μmol / m²・s at the jointing stage are screened out from the new predicted values. The expression amounts of photosynthetic core genes of the screened strains are verified by qPCR, and strains with relative expression amounts lower than the average are eliminated to obtain qualified strains, and the photosynthetic core genes include RBCS genes, PEPC genes and SPS genes.
9. The high light efficiency sugarcane breeding method according to claim 1, characterized by the fact that: The mid-term field dynamic phenotype verification specifically includes the following steps: After the qualified strains are planted, the net photosynthetic rate and the maximum photochemical efficiency are measured once a week at the tillering stage to the jointing stage, the actual maximum Rubisco carboxylation rate is fitted to obtain actual values, and deviation strains with a deviation of <5% between the actual values and the new predicted values are screened out. The sucrose degree is measured at the mature stage of the deviation strains, and the sucrose content is calculated according to sucrose content = 1.0625 × degree of brix - 7.7065, and strains with sucrose content ≥ 16% are retained to obtain mid-term strains.
10. The method for breeding high light efficiency sugarcane according to claim 1, characterized in that: The late genetic stability verification specifically includes the following steps: The mid-term strains are subjected to whole genome resequencing with a sequencing depth of ≥ 20x, and the homozygosity of photosynthetic related QTLs is verified, and mid-term strains with homozygosity ≥ 90% are retained to obtain homozygous strains; the photosynthetic related QTLs include QTLs controlling the maximum Rubisco carboxylation rate; The homozygous strains are planted for a yield comparison test, and the net photosynthetic rate, the maximum photochemical efficiency, the sucrose content and the yield per mu are measured during the whole growth period to obtain test strains; Test strains that meet the conditions of net photosynthetic rate ≥ 27 μmol / m²・s, maximum photochemical efficiency ≥ 0.85, sucrose content ≥ 16.5% and yield per mu ≥ 10 tons at the jointing stage are screened out, which are high photosynthetic efficiency sugarcane strains.