Data fusion analysis method and system for maize whole genome selection breeding
By constructing a multi-source data fusion network and a data-genetic response relationship model, the problem of insufficient data fusion in maize breeding was solved, achieving efficient data integration and dynamic analysis, and improving the accuracy and efficiency of breeding.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING FENGJIE YIJIA AGRICULTURAL TECHNOLOGY CO LTD
- Filing Date
- 2025-07-23
- Publication Date
- 2026-04-24
AI Technical Summary
In existing maize breeding methods, there is a lack of effective integration mechanisms for phenotypic data, genotypic data, and environmental data. This results in the underutilization of the correlation and interaction between data, affecting breeding efficiency and accuracy. Furthermore, existing models struggle to capture the nonlinear relationships and spatiotemporal dynamic characteristics of multi-source data.
A multi-source data fusion network was constructed, noise filtering and format unification were performed, a data-genetic response relationship model was established, marker-trait association analysis and generation-related effect value assessment were conducted, the core selection range was identified, and a statistical prediction model was constructed for dynamic analysis of genome-wide selection breeding.
It significantly improved data utilization and analysis accuracy, enhanced the accuracy and efficiency of breeding selection, optimized selection strategies, strengthened the stability and adaptability of breeding, and realized the intelligence and controllability of the breeding process.
Smart Images

Figure CN121034431B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of whole-genome selection breeding technology for maize, specifically to a data fusion analysis method and system for whole-genome selection breeding of maize. Background Technology
[0002] As a crucial global food crop and feed source, maize's breeding efficiency directly impacts agricultural productivity and food security. Traditional maize breeding methods primarily rely on phenotypic selection and hybridization techniques, but these methods are time-consuming, inefficient, and ill-suited to addressing complex environmental changes and genetic diversity challenges. With the development of genomics technologies, genome-wide selection (GS) has become an important tool for improving breeding efficiency. However, existing GS methods still suffer from insufficient data integration and limited model prediction accuracy, hindering further improvements in breeding efficiency.
[0003] In maize breeding, phenotypic data, genotypic data, and environmental data are the three core data sources. Phenotypic data reflects the actual traits of maize plants, such as plant height, ear height, and yield; genotypic data, obtained through techniques such as single nucleotide polymorphism (SNP) markers, reveals genetic variation information; and environmental data includes external factors such as temperature, precipitation, and soil properties. Currently, these data are usually analyzed independently, lacking effective fusion mechanisms, resulting in the underutilization of the correlations and interactions between data. Furthermore, the impact of dynamic changes during the breeding process (such as generational succession and environmental fluctuations) on selection efficiency has not been systematically quantified, further limiting the accuracy and timeliness of breeding.
[0004] In existing technologies, data fusion often employs simple statistical models or linear regression methods, which struggle to capture the nonlinear relationships and spatiotemporal dynamics of multi-source data. For example, the association between marker effects and phenotypic variation may exhibit different patterns due to environmental changes, but existing methods have failed to effectively integrate these complex relationships. Furthermore, noisy data and inconsistent formats increase the difficulty of data analysis, reducing model reliability and prediction accuracy. Therefore, a method and system are needed that can efficiently integrate multi-source data, dynamically analyze genetic responses, and optimize breeding selection to overcome these technical bottlenecks.
[0005] To address these issues, this invention proposes a data fusion analysis method and system for whole-genome selection breeding of maize. By constructing a multi-source data fusion network, dynamic modeling, and real-time optimization, the accuracy and efficiency of breeding are significantly improved. Summary of the Invention
[0006] The purpose of this invention is to provide a data fusion analysis method and system for whole-genome selection breeding of maize, so as to solve the problems mentioned in the background art.
[0007] To achieve the above objectives, the present invention provides the following technical solution: a data fusion analysis method for whole-genome selection breeding of maize, the method comprising:
[0008] Step T1: Collect multi-dimensional data from the breeding population to obtain phenotypic, genotypic, and environmental data; determine key fusion nodes based on the correlation characteristics of phenotypic, genotypic, and environmental data, and construct a multi-source data fusion network;
[0009] Step T2: Perform noise filtering and format unification on the multi-source data fusion network, and integrate breeding parameters in real time to obtain a dynamic breeding dataset; perform spatiotemporal correlation processing on the dynamic breeding dataset based on label effects, phenotypic variability, and environmental interactions to generate feature correlation data;
[0010] Step T3: Establish a data-genetic response relationship model based on the characteristic association data and the background information of the breeding population; perform marker-trait association analysis based on the data-genetic response relationship model, evaluate the generation-related effect size, and generate dynamic genetic association data;
[0011] Step T4: Perform selection efficiency association processing based on the genetic association dynamic data and dynamic breeding dataset to generate selection efficiency dynamic characteristic data; identify breeding influence areas based on the selection efficiency dynamic characteristic data to generate core selection range data; and couple genetic effect patterns based on marker effects, phenotypic stability, and environmental adaptability based on the core selection range data to generate genetic action coupling data.
[0012] Step T5: Construct a statistical prediction model based on the genetic interaction coupling data, and use the statistical prediction model to perform dynamic analysis of whole-genome selection breeding to generate selection analysis data; optimize breeding parameters in real time based on the selection analysis data to obtain whole-genome selection breeding control feedback data.
[0013] Preferably, step T1 includes the following steps:
[0014] Step T11: Conduct field phenotypic measurements and laboratory tests on the breeding population to obtain phenotypic observation data such as plant height, ear height, and yield;
[0015] Step T12: Perform genome resequencing and SNP microarray scanning on the breeding population to obtain single nucleotide polymorphism marker data and genotype frequency data;
[0016] Step T13: Conduct meteorological station records and soil sampling analysis in the breeding area to obtain environmental background data such as temperature, precipitation, and soil organic matter content;
[0017] Step T14: Merge phenotypic observation data, genotype frequency data, and environmental background data into a breeding dataset;
[0018] Step T15: Perform significance analysis on the association strength of phenotype-genotype-environment based on the breeding dataset to obtain key association point data;
[0019] Step T16: Determine the node distribution of the multi-source data fusion network based on the key correlation point data, and construct a multi-source data fusion network that includes a phenotypic database, a genotypic database, and an environmental database.
[0020] Preferably, step T16 includes the following steps:
[0021] Step T161: Perform mutation analysis on the covariation relationship between phenotype and genotype based on the breeding dataset, identify the locations where phenotypic variation and genotype frequency changes are significantly correlated, and obtain phenotypic mutation point data;
[0022] Step T162: Perform sensitivity analysis on the environment-genotype interaction effect based on the breeding data set, assess the degree of influence of different environmental factors on genotype expression, and obtain data on environmentally sensitive areas;
[0023] Step T163: Determine the spatial distribution of fusion nodes based on preset breeding target parameters, trait mutation point data, and environmentally sensitive area data to obtain node layout data;
[0024] Step T164: Determine the data interface type based on the node layout data, and construct a multi-source data fusion network including phenotypic data interface, genotypic data interface and environmental data interface.
[0025] Preferably, step T2 includes the following steps:
[0026] Step T21: Perform outlier removal, noise filtering and format unification on the multi-source data fusion network, and integrate real-time parameters from the breeding process to obtain a dynamic breeding dataset.
[0027] Step T22: Segment the marker effect values in the dynamic breeding dataset based on the generation sequence to establish the effect-generation relationship curve;
[0028] Step T23: Calculate the frequency characteristics and amplitude distribution of effect fluctuations based on the effect-generation relationship curve, and perform correlation analysis between effect fluctuations and generations to obtain effect-generation evolutionary characteristic data;
[0029] Step T24: Establish the relationship curve between phenotypic variability and generations based on the dynamic breeding dataset, and calculate the variation rate of different generations to obtain the spatiotemporal evolutionary characteristic data of variation.
[0030] Step T25: Record the real-time changes of environmental interaction values in the dynamic breeding dataset, and perform correlation analysis between environmental interaction values and generations to obtain spatiotemporal evolutionary characteristic data of interactions;
[0031] Step T26: Based on the effect-generation evolutionary characteristic data, variation spatiotemporal evolutionary characteristic data, and interaction spatiotemporal evolutionary characteristic data, conduct a comprehensive impact analysis on selection efficiency of data combination, and identify the critical value and fit interval of key data combination, thereby generating spatiotemporal evolutionary characteristic data of data combination.
[0032] Step T27: Establish a dynamic evolution model of the three-dimensional data space based on the spatiotemporal evolution characteristic data, and extract feature changes to obtain feature-related data.
[0033] Preferably, step T21 includes the following steps:
[0034] The multi-source data fusion network is subjected to outlier removal based on Z-score standardization and format processing with unified unit dimensions. Phenotypic measurements, genotypic detection values, and environmental records from different generations during the breeding process are integrated to obtain a dynamic breeding dataset. Outlier removal includes outlier identification based on quartile range, and format unification includes converting phenotypic data into dimensionless indicators, genotypic data into binary encoded values, and environmental data into relative index values.
[0035] Preferably, step T27 includes the following steps:
[0036] Step T271: Construct a three-dimensional data space based on spatiotemporal evolution characteristic data, with the labeling effect, phenotypic variability, and environmental interaction values as coordinate axes, thereby obtaining the data space coordinate data;
[0037] Step T272: Establish the feature change trajectory based on the data spatial coordinate data, and construct the feature motion trajectory curve through the spatial mapping of time-series sampling points, thereby obtaining the feature trajectory data;
[0038] Step T273: Perform generation-based hierarchical processing on the feature trajectory data to identify the characteristic change characteristics of different generation segments, including change trends, change rates, and interrelationships, thereby obtaining hierarchical characteristic data;
[0039] Step T274: Construct a dynamic model of the three-dimensional data space based on the hierarchical characteristic data, and establish a continuous expression of feature changes through linear interpolation and polynomial fitting to obtain dynamic evolution model data;
[0040] Step T275: Extract the characteristics of the dynamic evolution model data based on gradient features, curvature features, and rate features according to feature changes, thereby obtaining model characteristic data;
[0041] Step T276: Based on the model characteristic data, perform data correlation analysis based on the relationship and interaction mechanism between data to obtain data correlation data;
[0042] Step T277: Extract indicators that represent the dynamic change patterns of features based on model characteristic data and data association data, thereby generating feature association data.
[0043] Preferably, step T3 includes the following steps:
[0044] Step T31: Based on the feature association data and breeding population background information, establish a data-genetic characteristic correspondence table that includes the mapping relationship between data changes and genetic responses in different generations, thereby obtaining data response data;
[0045] Step T32: Perform principal component analysis and feature screening on the data response data, and establish a statistical model of data-genetic response to obtain response model data;
[0046] Step T33: Perform linear regression training based on the response model data to construct a linear mapping relationship between the data and the genetic response, thereby obtaining the data-genetic response relationship model;
[0047] Step T34: Perform genetic association analysis based on the data-genetic response relationship model, including marker interpretation rate, trait heritability, and environmental contribution rate, to obtain association characteristic data;
[0048] Step T35: Calculate generation-related effect values based on association characteristic data to obtain effect value data. The effect value calculation includes cumulative analysis of marker effects and product analysis of environmental interactions.
[0049] Step T36: Identify key characteristics and mutation points in the genetic association process of the associated dynamic data to obtain the genetic association dynamic data.
[0050] Preferably, step T4 includes the following steps:
[0051] Step T41: Analyze the impact of different data combinations on selection efficiency based on genetic association dynamic data and dynamic breeding datasets to obtain efficiency impact data;
[0052] Step T42: Efficiency performance evaluation of the efficiency impact data based on selection accuracy, genetic progress rate, and stability indicators to obtain performance evaluation data;
[0053] Step T43: Analyze the evolutionary characteristics of selection efficiency over generations on the performance evaluation data to obtain dynamic characteristic data of selection efficiency;
[0054] Step T44: Based on the selection efficiency dynamic characteristic data, simulate the breeding impact area and establish a multi-field coupling analysis model including the marker effect field, phenotypic variation field and environmental interaction field to obtain the impact area data;
[0055] Step T45: Perform boundary identification and spatial partitioning on the affected area data to obtain the core selection range data;
[0056] Step T46: Based on the core selection range data, perform genetic effect pattern coupling based on marker effect, phenotypic stability and environmental adaptation, identify the evolution law of effect field and critical state characteristics, and generate genetic action coupling data.
[0057] Preferably, step T5 includes the following steps:
[0058] Step T51: Perform feature screening on the genetic interaction coupling data and construct a training sample set for the statistical prediction model to obtain training sample data, which includes input features and selected labels;
[0059] Step T52: Construct a statistical prediction model structure based on a linear mixture model based on the training sample data, and estimate the model parameters to obtain the selected prediction model;
[0060] Step T53: Perform model prediction accuracy calibration based on cross-validation method on the selected prediction model to obtain prediction model data;
[0061] Step T54: Utilize the predictive model data to perform dynamic analysis of whole-genome selection breeding, thereby obtaining selection analysis data, which includes marker preference sets, phenotypic predicted values, and environmental adaptation schemes;
[0062] Step T55: Identify the key stages and selection mutation characteristics in the breeding process from the selection analysis data to obtain selection evolution data. The key stages include the basic population construction stage, the early generation selection stage, the high generation stabilization stage, and the multi-environment verification stage.
[0063] Step T56: Optimize breeding parameters at each stage in real time based on the selection evolution data to obtain whole-genome selection breeding control feedback data.
[0064] Preferably, the present invention further includes a data fusion analysis system for maize whole-genome selection breeding, used to perform the above-described data fusion analysis method for maize whole-genome selection breeding, wherein the data fusion analysis system for maize whole-genome selection breeding includes:
[0065] The data acquisition module is used to collect multi-dimensional data from the breeding population, including phenotypic, genotypic, and environmental data; it identifies key fusion nodes based on the correlation characteristics of phenotypic, genotypic, and environmental data, and constructs a multi-source data fusion network.
[0066] The data processing module is used to perform noise filtering and format unification processing on the multi-source data fusion network, and to integrate breeding parameters in real time to obtain a dynamic breeding dataset; the dynamic breeding dataset is then subjected to spatiotemporal correlation processing based on label effects, phenotypic variability and environmental interactions to generate feature correlation data.
[0067] The association modeling module is used to establish a data-genetic response relationship model based on feature association data and breeding population background information; perform marker-trait association analysis based on the data-genetic response relationship model; evaluate generation-related effect values; and generate dynamic genetic association data.
[0068] The efficiency evaluation module is used to perform selection efficiency association processing based on genetic association dynamic data and dynamic breeding datasets to generate selection efficiency dynamic characteristic data; to identify breeding influence areas based on selection efficiency dynamic characteristic data to generate core selection range data; and to couple genetic effect patterns based on marker effects, phenotypic stability, and environmental adaptability based on core selection range data to generate genetic action coupling data.
[0069] The selection analysis module is used to construct a statistical prediction model based on genetic interaction coupling data, and to perform dynamic analysis of whole-genome selection breeding using the statistical prediction model to generate selection analysis data; based on the selection analysis data, breeding parameters are optimized in real time to obtain whole-genome selection breeding control feedback data.
[0070] Compared with the prior art, the beneficial effects of the present invention are:
[0071] This invention achieves efficient integration of phenotypic, genotypic, and environmental data through multi-dimensional data collection and fusion, significantly improving data utilization and analytical accuracy. It overcomes the limitations of isolated data analysis in traditional methods, fully exploring the correlations and interactions among multi-source data, and providing a more comprehensive scientific basis for breeding decisions.
[0072] This invention employs noise filtering and format standardization to effectively address data anomalies and heterogeneity issues. Through Z-score standardization and unit dimension unification, data quality is significantly improved, laying a reliable foundation for subsequent modeling and analysis. The construction of the dynamic breeding dataset further enables real-time integration of breeding parameters, making the analysis process more flexible and precise.
[0073] By establishing a data-genetic response relationship model, this invention can quantify the dynamic associations between marker effects, phenotypic variation, and environmental interactions, providing scientific theoretical support for breeding selection. The combination of marker-trait association analysis and generational effect value assessment clearly reveals the spatiotemporal evolutionary patterns of genetic variation, significantly improving the accuracy and predictability of selection.
[0074] The selection efficiency correlation processing and core selection range identification proposed in this invention can dynamically assess the impact of different data combinations on breeding efficiency and optimize selection strategies. Based on the coupling of genetic effect patterns of marker effects, phenotypic stability, and environmental adaptability, it further enhances the stability and adaptability of breeding, providing a powerful tool for variety breeding in complex environments.
[0075] The construction and dynamic analysis capabilities of statistical prediction models make the whole-genome selection breeding process more intelligent and controllable. Through cross-validation and model calibration, prediction accuracy is significantly improved, and real-time optimization of breeding parameters further enhances the efficiency and success rate of selection.
[0076] This invention is not only applicable to maize breeding, but can also provide technical reference for genome-wide selection in other crops, exhibiting broad application prospects and promotional value. Its systematic and modular design makes the method easy to implement and expand, providing crucial support for the advancement of modern agricultural breeding technology. Attached Figure Description
[0077] Figure 1 This is a schematic diagram illustrating the working principle of the data fusion analysis method for whole-genome selection breeding of maize described in this invention.
[0078] Figure 2 Design diagram for constructing a multi-source data fusion network;
[0079] Figure 3 Design diagram for modeling data-genetic response relationships;
[0080] Figure 4 To select a design diagram for dynamic efficiency evaluation;
[0081] Figure 5 A design diagram for dynamic analysis of genome-wide selection breeding. Detailed Implementation
[0082] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0083] Please see Figures 1-5 The data fusion analysis method for whole-genome selection breeding of maize involved in this invention has the following specific implementation steps:
[0084] Step T1: Collect multi-dimensional data from the breeding population to obtain phenotypic, genotypic, and environmental data; determine key fusion nodes based on the correlation characteristics of phenotypic, genotypic, and environmental data, and construct a multi-source data fusion network.
[0085] Step T2: Perform noise filtering and format unification processing on the multi-source data fusion network, and integrate breeding parameters in real time to obtain a dynamic breeding dataset; perform spatiotemporal correlation processing on the dynamic breeding dataset based on label effect, phenotypic variability and environmental interaction to generate feature correlation data.
[0086] Step T3: Establish a data-genetic response relationship model based on the characteristic association data and the background information of the breeding population; perform marker-trait association analysis based on the data-genetic response relationship model, and evaluate the generation-related effect value to generate dynamic genetic association data.
[0087] Step T4: Perform selection efficiency association processing based on the dynamic genetic association data and the dynamic breeding dataset to generate selection efficiency dynamic characteristic data; identify the breeding influence area based on the selection efficiency dynamic characteristic data to generate core selection range data; and couple genetic effect patterns based on marker effect, phenotypic stability, and environmental adaptability based on the core selection range data to generate genetic action coupling data.
[0088] Step T5: Construct a statistical prediction model based on the genetic interaction coupling data, and use the statistical prediction model to perform dynamic analysis of whole-genome selection breeding to generate selection analysis data; optimize breeding parameters in real time based on the selection analysis data to obtain whole-genome selection breeding control feedback data.
[0089] Example 1:
[0090] During step T1, multi-dimensional data collection of the breeding population is required. For phenotypic data acquisition, phenotypic measurements of the breeding population must be conducted in a field setting, combined with laboratory testing methods. Specifically, field phenotypic measurements include plant height, an important indicator, which is measured vertically from a ground reference point to the top of the plant's growing point to accurately reflect its longitudinal growth. Ear height is measured as the vertical distance from the ground to the node where the ear attaches; this data is crucial for assessing maize's lodging resistance. Yield data is obtained by harvesting, threshing, and weighing individual plants or plots of each breeding material after maize maturity, yield observation data. In addition to these, other phenotypic observation data related to maize growth, development, and yield formation are also included. Through these comprehensive measurements and tests, rich and accurate phenotypic observation data are ensured.
[0091] In terms of genotype data acquisition, genome resequencing and SNP microarray scanning were performed on the breeding population. Genome resequencing utilizes high-throughput sequencing technology to comprehensively sequence the maize genome, thereby obtaining the sequence information of the entire genome. By comparing with a reference genome, a large number of genetic variations can be discovered, including single nucleotide polymorphisms (SNPs) and insertions / deletions (InDels). SNP microarray scanning is a highly efficient genotyping technique. It designs probes based on known SNP sites, hybridizes with sample DNA, and determines the genotype of an individual by detecting the hybridization signal. The combination of these two techniques allows for the acquisition of single nucleotide polymorphism marker data and genotype frequency data, providing a foundation for subsequent genetic analysis.
[0092] Environmental data collection is equally crucial, requiring meteorological station recordings and soil sampling analysis of the breeding area. Meteorological station records provide real-time temperature data, including daily maximum, minimum, and average temperatures, which are essential for understanding the thermal conditions during maize growth. Precipitation data, including rainfall amount, duration, and intensity, directly impacts water supply to maize. Additionally, records of meteorological elements such as sunshine duration and wind speed are necessary. For soil sampling analysis, soil samples must be collected within the breeding area according to specific sampling rules. Soil organic matter content must be measured, and chemical analysis methods used to determine the organic matter content in the soil, a key indicator for assessing soil fertility. Simultaneously, analysis of soil pH, nutrient content (such as nitrogen, phosphorus, and potassium), soil texture, and other environmental background data is required to comprehensively understand the soil environmental conditions of the breeding area.
[0093] After collecting phenotypic observation data, genotype frequency data, and environmental background data, these data were merged to form a breeding dataset. This dataset integrates various aspects of information involved in maize breeding, providing a unified data foundation for subsequent analysis.
[0094] A significance analysis of the phenotypic-genotypic-environment association strength was performed based on the breeding dataset. This process required the use of statistical methods to calculate the correlation between phenotypic and genotypic data, as well as the degree of influence of environmental data on phenotypic and genotypic expression. By setting appropriate significance thresholds, points with significant associations were selected, thus obtaining key association point data. These key association points reflect the important connections between phenotypic, genotypic, and environment, and are crucial for the subsequent construction of fusion networks.
[0095] After identifying key association points, a multi-source data fusion network was constructed. Mutational analysis was performed on the covariation relationships between phenotype and genotype based on the breeding dataset. By analyzing the phenotypic differences between individuals with different genotypes and the patterns of phenotypic frequency changes with genotype frequency, the association locations between phenotypic variation and significant changes in genotype frequency were identified, yielding phenotypic mutation point data. These mutation points may correspond to gene loci associated with important agronomic traits.
[0096] Sensitivity analysis was conducted on the environment-genotype interaction effect. The influence of different environmental factors (such as temperature, precipitation, and soil nutrients) on genotype expression was assessed to determine which environmental factors play a key role in the expression of specific genotypes, thus obtaining data on environmentally sensitive regions. These regions reflect the differences in genotype expression under different environmental conditions and are of great significance for understanding the environmental adaptability of maize.
[0097] Based on preset breeding target parameters, trait mutation point data, and environmentally sensitive area data, the spatial distribution of fusion nodes is determined, resulting in node layout data. Breeding target parameters, such as high yield, lodging resistance, and disease resistance, combined with trait mutation points and environmentally sensitive areas, can determine where to place fusion nodes to better capture key information related to the breeding target.
[0098] Based on the node deployment data, the data interface types are determined, and a multi-source data fusion network is constructed, including phenotypic data interfaces, genotypic data interfaces, and environmental data interfaces. Different types of data interfaces are used to enable data interaction and fusion between phenotypic, genotypic, and environmental databases, ensuring effective transmission and integration of multi-source data within the network, and providing a solid infrastructure for subsequent data processing and analysis.
[0099] Example 2:
[0100] During step T2, noise filtering and format standardization are performed on the data in the multi-source data fusion network to ensure data quality and consistency. Specifically, the noise filtering operation is based on outlier removal methods, using Z-score normalization to process the data across all dimensions. Z-score normalization standardizes the data by calculating the difference between each data point and the dataset mean, and dividing by the standard deviation, ensuring the data follows a standard normal distribution. Based on this, outlier identification is performed using the interquartile range (IQR). First, the first quartile (Q1) and third quartile (Q3) of the dataset are calculated, and the IQR value is determined as the difference between Q3 and Q1. Then, data points below Q1 - 1.5IQR or above Q3 + 1.5IQR are considered outliers and removed, thus eliminating noise interference in the data.
[0101] Standardized format processing involves unifying the units of measurement for different types of data. For phenotypic data, such as plant height, ear height, and yield, since their original units differ (e.g., length in centimeters, yield in grams), they need to be converted to dimensionless indicators. Specifically, a normalization method can be used, dividing each phenotypic data point by the maximum or mean value of that data type, mapping its value range to the [0,1] interval, thus eliminating the influence of unit differences. Genotype data mainly consists of single nucleotide polymorphism (SNP) marker data and genotype frequency data, which are converted to binary encoded values. For example, the three genotypes of SNP sites (AA, Aa, aa) can be encoded as 0, 1, 2, or other binary encoding methods to facilitate computer processing and analysis. Environmental data, such as temperature, precipitation, and soil organic matter content, are converted to relative index values. These can be compared with historical averages or specific benchmark values to calculate the relative deviation or index value, achieving standardized processing of environmental data.
[0102] After noise filtering and format standardization, it is necessary to integrate real-time parameters from the breeding process to obtain a dynamic breeding dataset. These real-time parameters include phenotypic measurements from different generations, such as plant height, ear height, and yield data for each generation of maize; genotype detection values, i.e., SNP marker data and genotype frequency data for each generation of individuals; and environmental records, including environmental data such as temperature, precipitation, and soil nutrients during each generation of breeding. By integrating these real-time data from different generations, a dynamic dataset that changes with each generation is formed, reflecting the dynamic changes of various factors during maize breeding.
[0103] After obtaining the dynamic breeding dataset, the marker effect values are segmented based on generation sequences. Marker effect values reflect the degree of influence of each genetic marker on the target trait. The dataset is segmented by generation; for example, the marker effect values of different generations are divided into several time periods. Then, an effect-generation relationship curve is constructed for the marker effect values within each time period. This curve, with generations on the x-axis and marker effect values on the y-axis, visually demonstrates the changing trend of marker effects over generations.
[0104] The frequency characteristics and amplitude distribution of effect fluctuations are calculated based on the effect-generation relationship curve. Frequency characteristic analysis primarily involves statistically analyzing the number and period of effect value fluctuations across different generations to understand the frequency of marker effect changes. Amplitude distribution calculation determines the range of effect value fluctuations and analyzes the severity of these changes. Simultaneously, correlation analysis between effect fluctuations and generations is performed. By calculating correlation coefficients and other statistical measures, it is determined whether there is a significant correlation between marker effect fluctuations and generations, thus obtaining effect-generation evolutionary characteristic data that reflects the evolutionary pattern of the marker effect during generational transmission.
[0105] Simultaneously, a relationship curve between phenotypic variability and generations was established based on the dynamic breeding dataset. Phenotypic variability measures the degree of variation of the same trait across different generations. By calculating indicators such as variance, standard deviation, or coefficient of variation of representative phenotypic data for each generation, a curve was plotted with generations on the x-axis and phenotypic variability on the y-axis. Then, the variation rate for different generation segments was calculated. For example, the entire breeding process was divided into early, middle, and late generations, and the phenotypic variation rate within each stage was calculated separately, thereby obtaining spatiotemporal evolutionary characteristic data of variation. This data reflects the evolutionary characteristics of phenotypic variation in time and space (different generations).
[0106] The environmental interaction values in the dynamic breeding dataset are recorded in real time. Environmental interaction values represent the degree of influence of environmental factors on genotype expression. As breeding generations progress, environmental conditions may change, necessitating real-time recording of environmental interaction values for each generation. Then, a correlation analysis is performed between environmental interaction values and generations to analyze the changing trends of environmental interaction values across generations and their correlation with other factors, thereby obtaining spatiotemporal evolutionary characteristic data of the interaction. This data reflects the changing patterns of the environment-genotype interaction effect in the spatiotemporal dimensions.
[0107] After acquiring effect-generational evolutionary characteristic data, variation spatiotemporal evolutionary characteristic data, and interaction spatiotemporal evolutionary characteristic data, these data are comprehensively analyzed to assess the overall impact of data combinations on selection efficiency. Statistical methods and data analysis techniques are used to identify the critical values and fit intervals of key data combinations. The critical value refers to the threshold at which a data combination reaches a certain state, and the fit interval is the range within which a data combination can effectively exert its influence. This generates spatiotemporal evolutionary characteristic data of data combinations, which describes the impact of the spatiotemporal evolution of different data combinations on selection efficiency.
[0108] A dynamic evolution model of a three-dimensional data space is established based on spatiotemporal evolution characteristic data, and feature changes are extracted. First, a three-dimensional data space is constructed with marker effect, phenotypic variability, and environment interaction values as coordinate axes. Data from each generation is mapped into this space to obtain data space coordinate data. Then, feature change trajectories are established based on the data space coordinate data. Feature motion trajectory curves are constructed through spatial mapping of time-series sampling points, thereby obtaining feature trajectory data. These trajectory curves intuitively demonstrate the motion trajectory of each feature in three-dimensional space.
[0109] The characteristic trajectory data is processed using a generation-based stratification method, dividing the entire breeding process into different levels according to generations. This identifies the characteristic changes in different generation segments, including trends, rates of change, and interrelationships, thus obtaining stratified characteristic data. Based on this stratified characteristic data, a dynamic model in a three-dimensional data space is constructed. Continuous expressions of characteristic changes are established using methods such as linear interpolation and polynomial fitting, thereby obtaining dynamic evolutionary model data.
[0110] Feature extraction is performed on the dynamic evolution model data, including gradient features, curvature features, and rate features based on feature changes. Gradient features reflect the direction and intensity of feature changes, curvature features describe the degree of curvature of the feature trajectory, and rate features indicate the speed of feature changes. By extracting these features, model feature data is obtained. Based on the model feature data, data correlation analysis is performed to analyze the relationships and interaction mechanisms between data, thereby obtaining data association data. Finally, indicators characterizing the dynamic changes of features are extracted from the model feature data and data association data, thereby generating feature association data, providing key feature inputs for subsequent analysis and modeling.
[0111] Example 3:
[0112] In step T3, a data-genetic characteristic correspondence table needs to be established based on the feature association data and the breeding population background information. The feature association data is generated by performing spatiotemporal correlation processing on the dynamic breeding dataset in step T2, and includes the dynamic changes of features such as marker effects, phenotypic variability, and environmental interactions during generational transmission. The breeding population background information includes the population's genetic composition, pedigree, and breeding history, which are crucial for understanding the foundation of genetic responses.
[0113] When establishing a data-genetic trait mapping table, it is necessary to map changes in data across different generations to genetic responses. For example, in a given generation, changes in marker effects may correspond to genetic variations in a specific trait, while changes in environment interaction values may affect the expression level of that trait. By analyzing in detail the relationship between data characteristics and genetic responses in each generation, a one-to-one mapping table is formed, thus obtaining data-response data. This data-response data clearly records the association between data changes and genetic responses, providing a direct basis for subsequent modeling and analysis.
[0114] After obtaining the response data, principal component analysis (PCA) and feature selection are required. PCA is a dimensionality reduction technique that transforms multiple variables into a few principal components through linear transformation. These principal components retain as much information as possible from the original data. In PCA, the covariance matrix of the response data is first calculated, then its eigenvalues and eigenvectors are solved. The number of principal components is determined based on the magnitude of the eigenvalues. Typically, principal components with a cumulative variance contribution rate reaching a certain threshold (e.g., above 85%) are selected as variables for subsequent analysis.
[0115] Feature selection, building upon principal component analysis, further identifies features that significantly influence the genetic response. Statistical tests, such as t-tests and F-tests, can be used to calculate the significance level between each feature and the genetic response. A significance threshold (e.g., p < 0.05) is set, retaining significant features and discarding insignificant ones, thereby reducing data dimensionality and improving model efficiency and accuracy. Through principal component analysis and feature selection, response model data is obtained, containing key features that significantly impact the genetic response.
[0116] Linear regression training is performed based on response model data to construct a linear mapping relationship between the data and the genetic response. Linear regression is a commonly used statistical method to establish a linear relationship between dependent and independent variables. During training, features from the response model data are used as independent variables, and the genetic response as the dependent variable. Optimization methods such as least squares are used to solve for the regression coefficients, minimizing the error between the model's predicted and actual values.
[0117] During training, careful data partitioning is crucial. Data is typically divided into training and test sets. The training set is used for parameter estimation, while the test set is used to evaluate model performance. By continuously adjusting the model's parameters and structure, we ensure good fitting and generalization abilities, thus obtaining a data-genetic response relationship model. This model can quantitatively describe the linear relationship between data features and genetic responses, providing mathematical model support for subsequent genetic association analysis.
[0118] Genetic association analysis is performed based on data-genetic response models, including the calculation of marker-explained variance, trait heritability, and environmental contribution. Marker-explained variance refers to the proportion of trait variation explained by all markers in the model, reflecting the degree of influence of markers on the trait. To calculate marker-explained variance, the total variance of the model is first calculated, and then the variance explained by the marker effect is calculated; the ratio of the two is the marker-explained variance.
[0119] Heritability of a trait refers to the proportion of genetic variation in the total variation of a trait, reflecting the degree to which a trait is influenced by genetic factors. Calculating heritability requires estimating both genetic and environmental variances; the ratio of these two is the heritability. The environmental contribution rate refers to the proportion of environmental factors contributing to trait variation, determined by calculating the proportion of environmental variance to the total variance. These analyses yield association characteristic data, which comprehensively describes various characteristics of genetic associations and provides a basis for assessing generation-related effect sizes.
[0120] Effect values were calculated based on association characteristic data across generations. The calculation included cumulative analysis of marker effects and product analysis of environmental interactions. Cumulative analysis of marker effects summed the effect values of each marker in each generation to obtain the total marker effect value for that generation, reflecting the combined influence of multiple markers on the trait. Product analysis of environmental interactions considered the interaction effect between environmental factors and genotype; by calculating the product of the environmental factor and genotype effects, the influence of environmental interactions on the trait was obtained.
[0121] The calculation process needs to consider the genetic transmission patterns between generations, such as dominant effects and epistatic effects. By calculating the effect value in detail for each generation, effect value data is obtained, which reflects the changing pattern of the effect value during the generational transmission process.
[0122] Key characteristics and mutation points in the genetic association process are identified from the dynamic data of the association studies. This dynamic data, generated during the genetic association analysis, includes data on changes in effect size, marker explained rate, and other data over generations. Identification of key characteristics primarily involves analyzing the trends and periodicity of these data, such as whether effect size increases or decreases with generations, and whether periodic fluctuations exist.
[0123] Mutation point identification involves detecting significant mutations in the data. For example, a sudden and significant increase or decrease in the marker interpretation rate in a particular generation may indicate an important genetic event, such as gene recombination or mutation, that occurred in that generation. By identifying key traits and mutation points, dynamic data on genetic associations are obtained. This data provides crucial genetic information for whole-genome selection breeding of maize, guiding subsequent breeding and selection efforts. Example 4:
[0124] In step T4, it is necessary to analyze the impact of different data combinations on selection efficiency based on genetic association dynamic data and dynamic breeding datasets. Genetic association dynamic data includes marker-trait association analysis and generational effect value assessment results, such as the cumulative value of marker effects across generations and the multiplicative effect of environmental interactions. The dynamic breeding dataset integrates real-time data on phenotype, genotype, and environment, such as plant height phenotype data, corresponding SNP marker genotype frequencies, and temperature and precipitation data for a given generation. At this point, the two types of data need to be matched along the generational dimension. For example, the marker effect data of the 5th generation can be extracted and combined with the phenotypic variability and environmental interaction values of that generation to form a data combination. By cross-referencing selection efficiency indicators (such as the deviation of the trait mean of the candidate population) under different combinations, it is analyzed whether the combination of high marker effect and high phenotypic variability has a synergistic effect on selection efficiency, or whether the selection efficiency of the data combination shows a specific trend when the environmental interaction value is high. This yields efficiency impact data, which records the correlation patterns between different data combinations and selection efficiency.
[0125] Efficiency performance is evaluated based on selection accuracy, genetic progress rate, and stability indices. Selection accuracy can be measured by the correlation between predicted trait values and measured values in the candidate population. For example, in a breeding cycle, the yield value of candidate individuals is predicted using a model, and then the Pearson correlation coefficient is calculated with the measured yield after harvest. Genetic progress rate requires comparing the change in the mean trait value of the population before and after selection, such as comparing the difference in mean plant height between the base population and the population after three generations of selection. Stability indices can be reflected by the coefficient of variation of trait performance in multi-environment trials, such as the coefficient of variation of yield of a variety at five different experimental sites. During the evaluation, the three indices corresponding to each data combination need to be standardized to avoid the influence of dimensional differences on the evaluation results, thereby obtaining performance evaluation data that quantitatively reflects the efficiency performance of different data combinations.
[0126] The evolutionary characteristics of selection efficiency across generations were analyzed using performance evaluation data. An evolutionary curve was plotted with generations on the x-axis and indicators such as selection accuracy and genetic progress rate on the y-axis. For example, selection accuracy was observed to increase rapidly in early generations (generations 1-3), stabilize in mid-generations (generations 4-6), and slightly decrease in late generations (generations 7 and beyond) due to reduced genetic variation. Simultaneously, the evolutionary rate and inflection points of each indicator were analyzed. For instance, the genetic progress rate showed a slowdown in its growth rate in the 4th generation, and this inflection point was investigated to determine whether it was related to a decrease in population genetic diversity. This yielded dynamic characteristics data of selection efficiency, revealing the dynamic changes in selection efficiency throughout the breeding process.
[0127] Based on the dynamic characteristics of selection efficiency, simulation calculations of the breeding impact area were performed, and a multi-field coupling analysis model was established. Here, "multi-field" includes the marker effect field, the phenotypic variation field, and the environmental interaction field. Taking the marker effect field as an example, SNP markers across the entire genome can be arranged according to chromosomal location, and a heatmap can be plotted using the marker effect value as the intensity, with color intensity representing the magnitude of the effect. The phenotypic variation field uses the coefficient of variation of traits such as plant height and yield as indicators to spatially display the variation distribution of different traits. The environmental interaction field spatially interpolates the interaction effect values of environmental factors such as temperature and precipitation with genotypes to generate a continuous interaction effect surface.
[0128] During simulation calculations, the three data fields need to be spatially superimposed. For example, in a certain chromosome region, if the marker effect field shows a high effect value, the phenotypic variation field shows a low coefficient of variation, and the environmental interaction field shows a moderate interaction effect, then this region is likely to have a significant impact on selection efficiency. By setting thresholds (such as an absolute value of the marker effect greater than 0.5 and a coefficient of variation less than 0.3), regions with overlapping fields are screened out, thereby obtaining the data of the influencing region. This data locates the spatiotemporal range that plays a key role in breeding selection.
[0129] Boundary identification and spatial partitioning are performed on the data of the affected areas. Boundary identification can be achieved using density clustering algorithms, such as DBSCAN, which determines the regional boundaries based on the density distribution of multi-field data, classifying high-density areas as core affected areas and low-density areas as peripheral affected areas. Spatial partitioning needs to consider breeding objectives. For example, for high-yield objectives, areas with high yield-related marker effects are classified as high-yield core areas; for stress resistance objectives, areas with high environmental interaction effects are classified as stress-resistant core areas. After partitioning, key features need to be labeled for each region. For example, the high-yield core area includes the 5-10 Mb range of chromosome 1, with a mean marker effect of 0.8 and a phenotypic coefficient of variation of 0.25, thus obtaining core selection range data, which provides clear spatial targets for breeding selection.
[0130] Genetic effect model coupling based on marker effects, phenotypic stability, and environmental adaptability is performed using data from a core selection area. Taking a specific core selection area as an example, the marker effect within this region is predominantly additive. Phenotypic stability is calculated using multi-year, multi-location experimental data (e.g., the coefficient of variation for plant height in different years is 0.15), while environmental adaptability is determined by analyzing the interaction effect values between the marker and temperature factors in this region (e.g., the interaction effect value is 0.6 under high-temperature conditions). During coupling, an effect model matrix needs to be constructed, with rows representing different environmental conditions (high temperature, low temperature, drought, etc.), columns representing different levels of phenotypic stability (high, medium, low), and matrix elements representing the combined marker effect values under the corresponding conditions. This allows for the identification of the evolutionary patterns of the effect field and the characteristics of critical states.
[0131] For example, it was found that when the temperature exceeds 35℃ and the phenotypic stability level is low, the composite value of the marker effect suddenly decreases. This critical state characteristic suggests that special attention should be paid to this region when selecting in high-temperature environments. Through comprehensive analysis, coupled genetic action data were generated, which integrates multi-dimensional genetic effect information and provides systematic theoretical support for breeding decisions.
[0132] Example 5:
[0133] In step T5, feature screening of the genetic interaction coupling data is required, and a training sample set for the statistical prediction model needs to be constructed. The genetic interaction coupling data includes the coupling relationship between marker effects, phenotypic stability, and environmental adaptability within the core selection range. For example, in a maize breeding population, the combined marker effect value of the 20-25Mb region of chromosome 3 under high temperature conditions is 0.7, and the corresponding phenotypic stability index (plant height coefficient of variation) is 0.12. During feature screening, screening criteria need to be set according to the breeding objectives (such as high yield and high temperature resistance). For example, features with an absolute marker effect value greater than 0.5 and an environmental adaptability index (high temperature interaction effect) greater than 0.6 should be retained, while redundant or low-correlation features, such as environmental factor features with a correlation less than 0.3 with the target trait, should be removed.
[0134] After screening, the feature data is combined with corresponding selection labels (such as "selected" or "rejected") to form training samples. Taking high-temperature resistant breeding as an example, the training samples may include: a genotype with a marker effect value of 0.8, a phenotypic stability coefficient of variation of 0.1, and an environmental adaptability index of 0.7 under high-temperature conditions, with the selection label "selected"; and another genotype with corresponding feature values of 0.3, 0.25, and 0.4, with the selection label "rejected". By collecting such samples from multiple generations and under multiple environments, training sample data containing input features and selection labels is constructed, providing a sufficient dataset for model training.
[0135] A statistical prediction model based on a linear mixture model is constructed using training sample data, and model parameters are estimated. Linear mixture models can simultaneously handle fixed effects (such as marker effects and environmental factors) and random effects (such as differences in genetic background among individuals), making them suitable for complex breeding data. In setting the model structure, selected characteristics (such as high-temperature interaction effects and phenotypic coefficients of variation) are used as fixed effects, and additive effects within the population are used as random effects, establishing a model framework such as "selection probability = fixed effect coefficient × eigenvalue + random effect value".
[0136] During parameter estimation, the restricted maximum likelihood (REML) method is used to solve for the fixed effects coefficients and variance components in the model. For example, through iterative calculation, the coefficient of the high-temperature interaction effect is determined to be 0.6, the coefficient of phenotypic variation is -0.4, and the variance of the random effects is 0.2, so that the model's predicted values match the selection labels of the training samples as closely as possible, thus obtaining the selection prediction model. This model can predict the probability of an individual in breeding selection based on the input feature data.
[0137] The selected prediction model is calibrated for prediction accuracy using a cross-validation method. Cross-validation typically employs 10-fold cross-validation, where the training sample set is divided into 10 equal parts. Each time, 9 parts are used as the training set to fit the model, and the remaining part is used as the test set to evaluate prediction accuracy. This process is repeated 10 times, and the average value is taken. Evaluation metrics include accuracy (the proportion of correctly predicted samples out of the total sample size), precision (the proportion of samples predicted as "selected" that were actually selected), and recall (the proportion of actually selected samples that were predicted as "selected").
[0138] For example, in a certain round of cross-validation, the model's prediction accuracy for 200 samples in the test set is 85%, precision is 80%, and recall is 88%. If the accuracy does not meet the preset standard (such as accuracy ≥ 90%), the model parameters are adjusted or the feature selection is repeated until the model accuracy meets the requirements, thereby obtaining the prediction model data.
[0139] Dynamic analysis of genome-wide selection breeding is performed using predictive model data to obtain selection analysis data. Dynamic analysis requires incorporating real-time data from the breeding process. For example, in the 6th generation of breeding, the input includes the genotype marker effect values (e.g., SNP1 effect value 0.5, SNP2 effect value -0.3), phenotypic predicted values (yield values predicted by the model), and environmental adaptation schemes (e.g., the individual's adaptability score in high-temperature environments). Based on this data, the model outputs a marker optimization set, selecting marker combinations that contribute significantly to the target trait (e.g., SNP1, SNP5, SNP7), generating phenotypic predicted values (e.g., a predicted yield of 850 kg / mu for an individual), and providing environmental adaptation schemes (e.g., the individual is more suitable for planting in areas with an average annual temperature above 25℃), thus forming the selection analysis data.
[0140] The selection analysis data is used to identify key stages and selection mutation characteristics in the breeding process. Key stages include the basic population construction stage, the early generation selection stage, the high generation stabilization stage, and the multi-environment validation stage. For example, in the basic population construction stage, attention is paid to population genetic diversity indicators (such as allele richness); in the early generation selection stage, the transmission stability of marker effects is analyzed; in the high generation stabilization stage, the convergence of phenotypic variation coefficients is assessed; and in the multi-environment validation stage, the significance of environmental interaction effects is examined.
[0141] Selection mutation identification involves detecting significant changes during the selection process. For example, if the frequency of a disease resistance marker suddenly increases from 0.3 to 0.7 after the fourth generation of selection, it may indicate the introduction of a new disease resistance gene resource in that generation. By identifying these key stages and mutation characteristics, selection evolution data are obtained, which records the dynamic changes in the breeding process.
[0142] Based on evolutionary selection data, breeding parameters at each stage are optimized in real time to obtain whole-genome selection breeding control feedback data. For example, in the early generation selection stage, if the transmission of marker effects is found to be unstable (e.g., the effect value of a certain marker fluctuates by more than 0.4 between generations), the selection intensity is adjusted, reducing the proportion of phenotype-based selection from 50% to 30% and increasing the weight of genotype selection; in the multi-environment validation stage, if the environmental interaction effect value of a certain variety in arid environments is -0.5 (poor performance), its promotion area is adjusted to exclude drought-prone areas.
[0143] Optimization parameters include selection intensity, selection method (such as BLUP, GBLUP), number of experimental sites, and phenotypic measurement frequency. These are adjusted through real-time feedback to make the breeding process more precise and efficient. For example, increasing the selection intensity from 0.2 to 0.3, or adding two high-temperature stress experimental sites, ultimately generates whole-genome selection breeding control feedback data, providing direct operational guidance for breeding decisions. Through these specific steps, the implementation process of predictive model construction and breeding parameter optimization in step T5 is completed, achieving closed-loop control from data modeling to breeding practice, and improving the efficiency and accuracy of whole-genome selection breeding.
[0144] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus.
[0145] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. A data fusion analysis method for whole-genome selection breeding of maize, characterized in that, Includes the following steps: Step T1: Collect multi-dimensional data from the breeding population to obtain phenotypic, genotypic, and environmental data; determine key fusion nodes based on the correlation characteristics of phenotypic, genotypic, and environmental data, and construct a multi-source data fusion network; Step T2: Perform noise filtering and format unification on the multi-source data fusion network, and integrate the breeding parameters of different generations in real time to obtain a dynamic breeding dataset; perform spatiotemporal correlation processing on the dynamic breeding dataset based on label effect, phenotypic variability and environmental interaction to generate feature correlation data; Step T3: Establish a data-genetic response relationship model based on the characteristic association data and the background information of the breeding population; perform marker-trait association analysis based on the data-genetic response relationship model, and evaluate the generation-related effect value to generate dynamic genetic association data; Step T4: Perform selection efficiency association processing based on the genetic association dynamic data and dynamic breeding dataset to generate selection efficiency dynamic characteristic data; identify breeding influence areas based on the selection efficiency dynamic characteristic data to generate core selection range data; and couple genetic effect patterns based on marker effects, phenotypic stability, and environmental adaptability based on the core selection range data to generate genetic action coupling data. Step T5: Construct a statistical prediction model based on the genetic interaction coupling data, and use the statistical prediction model to perform dynamic analysis of whole-genome selection breeding, generating selection analysis data; Breeding parameters are optimized in real time based on selection analysis data to obtain whole-genome selection breeding control feedback data.
2. The data fusion analysis method for whole-genome selection breeding of maize according to claim 1, characterized in that, Step T1 includes the following steps: Step T11: Conduct field phenotypic measurements and laboratory tests on the breeding population to obtain phenotypic observation data, including plant height, ear height, and yield; Step T12: Perform genome resequencing and SNP microarray scanning on the breeding population to obtain single nucleotide polymorphism marker data and genotype frequency data; Step T13: Conduct meteorological station records and soil sampling analysis in the breeding area to obtain environmental background data on temperature, precipitation, and soil organic matter content; Step T14: Merge phenotypic observation data, genotype frequency data, and environmental background data into a breeding dataset; Step T15: Perform significance analysis on the association strength of phenotype-genotype-environment based on the breeding dataset to obtain key association point data; Step T16: Determine the node distribution of the multi-source data fusion network based on the key correlation point data, and construct a multi-source data fusion network that includes a phenotypic database, a genotypic database, and an environmental database.
3. The data fusion analysis method for whole-genome selection breeding of maize according to claim 2, characterized in that, Step T16 includes the following steps: Step T161: Perform mutation analysis on the covariation relationship between phenotype and genotype based on the breeding dataset, identify the locations where phenotypic variation and genotype frequency changes are significantly correlated, and obtain phenotypic mutation point data; Step T162: Perform sensitivity analysis on the environment-genotype interaction effect based on the breeding data set, assess the degree of influence of different environmental factors on genotype expression, and obtain data on environmentally sensitive areas; Step T163: Determine the spatial distribution of fusion nodes based on preset breeding target parameters, trait mutation point data, and environmentally sensitive area data to obtain node layout data; Step T164: Determine the data interface type based on the node layout data, and construct a multi-source data fusion network including phenotypic data interface, genotypic data interface and environmental data interface.
4. The data fusion analysis method for whole-genome selection breeding of maize according to claim 3, characterized in that, Step T2 includes the following steps: Step T21: Perform outlier removal, noise filtering and format unification on the multi-source data fusion network, and integrate the breeding parameters of different generations in real time to obtain a dynamic breeding dataset. Step T22: Segment the marker effect values in the dynamic breeding dataset based on the generation sequence, that is, divide the marker effect values of different generations into several time periods, thereby establishing the effect-generation relationship curve; Step T23: Calculate the frequency characteristics and amplitude distribution of effect fluctuations based on the effect-generation relationship curve, and perform correlation analysis between effect fluctuations and generations to obtain effect-generation evolutionary characteristic data; Step T24: Establish the relationship curve between phenotypic variability and generations based on the dynamic breeding dataset, and calculate the variation rate of different generations to obtain the spatiotemporal evolutionary characteristic data of variation. Step T25: Record the real-time changes of environmental interaction values in the dynamic breeding dataset, and perform correlation analysis between environmental interaction values and generations to obtain spatiotemporal evolutionary characteristic data of interactions; Step T26: Based on the effect-generation evolutionary characteristic data, variation spatiotemporal evolutionary characteristic data, and interaction spatiotemporal evolutionary characteristic data, conduct a comprehensive impact analysis on selection efficiency of data combination, and identify the critical value and fit interval of key data combination, thereby generating spatiotemporal evolutionary characteristic data of data combination. Step T27: Establish a dynamic evolution model of the three-dimensional data space based on the spatiotemporal evolution characteristic data, and extract feature changes to obtain feature-related data.
5. The data fusion analysis method for whole-genome selection breeding of maize according to claim 4, characterized in that, Step T21 includes the following steps: The multi-source data fusion network is subjected to outlier removal based on Z-score standardization and format processing with unified unit dimensions. Phenotypic data, genotypic data and environmental data from different generations in the breeding process are integrated to obtain a dynamic breeding dataset. Outlier removal includes outlier identification based on quartile range, and format unification includes converting phenotypic data into dimensionless indicators, genotypic data into binary encoded values, and environmental data into relative index values.
6. The data fusion analysis method for whole-genome selection breeding of maize according to claim 5, characterized in that, Step T27 includes the following steps: Step T271: Construct a three-dimensional data space based on spatiotemporal evolution characteristic data, with the labeling effect, phenotypic variability, and environmental interaction values as coordinate axes, thereby obtaining the data space coordinate data; Step T272: Establish the feature change trajectory based on the data spatial coordinate data, and construct the feature motion trajectory curve through the spatial mapping of time-series sampling points to obtain the feature trajectory data; Step T273: Perform generation-based hierarchical processing on the feature trajectory data to identify the characteristic change characteristics of different generation segments, including change trends, change rates, and interrelationships, thereby obtaining hierarchical characteristic data; Step T274: Construct a dynamic model of the three-dimensional data space based on the hierarchical characteristic data, and establish a continuous expression of feature changes through linear interpolation and polynomial fitting to obtain dynamic evolution model data; Step T275: Extract the characteristics of the dynamic evolution model data based on gradient features, curvature features, and rate features according to feature changes, thereby obtaining model characteristic data; Step T276: Based on the model characteristic data, perform data correlation analysis based on the relationship and interaction mechanism between data to obtain data correlation data; Step T277: Extract indicators that represent the dynamic change patterns of features based on model characteristic data and data association data, thereby generating feature association data.
7. The data fusion analysis method for whole-genome selection breeding of maize according to claim 6, characterized in that, Step T3 includes the following steps: Step T31: Based on the feature association data and breeding population background information, establish a data-genetic characteristic correspondence table that includes the mapping relationship between data changes and genetic responses in different generations, thereby obtaining the response data; Step T32: Perform principal component analysis and feature screening on the response data, and establish a statistical model of data-genetic response to obtain response model data; Step T33: Perform linear regression training based on the response model data to construct a linear mapping relationship between the data and the genetic response, thereby obtaining the data-genetic response relationship model; Step T34: Perform genetic association analysis based on the data-genetic response relationship model, including marker interpretation rate, trait heritability, and environmental contribution rate, to obtain association characteristic data; Step T35: Calculate generation-related effect values based on association characteristic data to obtain effect value data. The effect value calculation includes cumulative analysis of marker effects and product analysis of environmental interactions. Step T36: Generate dynamic association data during the genetic association analysis process, including data on the changes in effect size, marker explanation rate, trait heritability, and environmental contribution rate over generations; identify key characteristics and mutation points in the genetic association process from the dynamic association data to obtain the dynamic genetic association data.
8. The data fusion analysis method for whole-genome selection breeding of maize according to claim 7, characterized in that, Step T4 includes the following steps: Step T41: Analyze the impact of different data combinations on selection efficiency based on genetic association dynamic data and dynamic breeding datasets to obtain efficiency impact data; Step T42: Efficiency performance evaluation of the efficiency impact data based on selection accuracy, genetic progress rate, and stability indicators to obtain performance evaluation data; Step T43: Analyze the evolutionary characteristics of selection efficiency over generations on the performance evaluation data to obtain dynamic characteristic data of selection efficiency; Step T44: Based on the selected efficiency dynamic characteristic data, simulate the breeding impact area and establish a multi-field coupling analysis model including the marker effect field, phenotypic variation field and environmental interaction field to obtain the impact area data; Step T45: Perform boundary identification and spatial partitioning on the affected area data to obtain the core selection range data; Step T46: Based on the core selection range data, perform genetic effect pattern coupling based on marker effect, phenotypic stability and environmental adaptation, identify the evolution law of effect field and critical state characteristics, and generate genetic action coupling data.
9. The data fusion analysis method for whole-genome selection breeding of maize according to claim 8, characterized in that, Step T5 includes the following steps: Step T51: Perform feature filtering on the genetic interaction coupling data and construct a training sample set for the prediction model to obtain training sample data, which includes input features and selected labels; Step T52: Construct a prediction model structure based on a linear mixture model based on the training sample data, and estimate the model parameters to obtain the prediction model; Step T53: Perform model prediction accuracy calibration based on cross-validation method on the prediction model to obtain the prediction model; Step T54: Utilize the predictive model to perform dynamic analysis of whole-genome selection breeding, thereby obtaining selection analysis data, which includes marker preference sets, phenotypic data, and environmental adaptation schemes; Step T55: Identify the key stages and selection mutation characteristics in the breeding process from the selection analysis data to obtain selection evolution data. The key stages include the basic population construction stage, the early generation selection stage, the high generation stabilization stage, and the multi-environment verification stage. Step T56: Optimize breeding parameters at each stage in real time based on the selection evolution data to obtain whole-genome selection breeding control feedback data.
10. A data fusion analysis system for maize whole-genome selection breeding, used to implement the data fusion analysis method for maize whole-genome selection breeding as described in any one of claims 1-9, characterized in that, For performing the data fusion analysis method for maize whole-genome selection breeding as described in claim 1, the maize whole-genome selection breeding data fusion analysis system comprises: The data acquisition module is used to collect multi-dimensional data from the breeding population, including phenotypic, genotypic, and environmental data; it identifies key fusion nodes based on the correlation characteristics of phenotypic, genotypic, and environmental data, and constructs a multi-source data fusion network. The data processing module is used to perform noise filtering and format unification processing on the multi-source data fusion network, and to integrate the breeding parameters of different generations in real time to obtain a dynamic breeding dataset; the dynamic breeding dataset is then subjected to spatiotemporal correlation processing based on label effects, phenotypic variability and environmental interactions to generate feature correlation data. The association modeling module is used to establish a data-genetic response relationship model based on feature association data and breeding population background information; perform marker-trait association analysis based on the data-genetic response relationship model; evaluate generation-related effect values; and generate dynamic genetic association data. The efficiency evaluation module is used to perform selection efficiency association processing based on genetic association dynamic data and dynamic breeding datasets to generate selection efficiency dynamic characteristic data; to identify breeding influence areas based on selection efficiency dynamic characteristic data to generate core selection range data; and to couple genetic effect patterns based on marker effects, phenotypic stability, and environmental adaptability based on core selection range data to generate genetic action coupling data. The selection analysis module is used to construct a predictive model based on genetic interaction coupling data, and to perform dynamic analysis of whole-genome selection breeding using the predictive model, generating selection analysis data; and to optimize breeding parameters in real time based on the selection analysis data, obtaining whole-genome selection breeding control feedback data.
Citation Information
Patent Citations
Titer phenotype data recombination method based on crop breeding platform
CN118538290A
Statistical validation of candidate genes
US20100145624A1