A wheat molecular design breeding auxiliary method, system, device and medium

By employing a dual-axis parallel feature selection and an independent subprocess sandbox isolation mechanism, combined with virtual DOM deferred rendering, the problems of limited functionality and high-concurrency computing bottlenecks in wheat breeding databases were solved, enabling high-precision phenotypic prediction and comprehensive breeding decision support.

CN122435975APending Publication Date: 2026-07-21CHINA AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
CHINA AGRI UNIV
Filing Date
2026-04-23
Publication Date
2026-07-21

AI Technical Summary

Technical Problem

Existing wheat breeding databases are limited in function, lack integrated analysis tools, have insufficient correlation between genotype and phenotype, and cause browser rendering blockage due to high-concurrency computing, making it difficult to meet high availability requirements.

Method used

By employing a dual-axis parallel feature screening and evolutionary signal quantification strategy, combined with an independent subprocess sandbox isolation mechanism and virtual DOM deferred rendering technology, we can achieve accurate haplotype visualization, dynamic assessment of genetic background, and real-time biaxial phenotypic prediction of multiple traits.

Benefits of technology

It significantly improved the model's accuracy and biological significance, broke through the blind spot of single phenotypic evaluation, overcame the bottleneck of high-concurrency computing power, and achieved efficient breeding decision support.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122435975A_ABST
    Figure CN122435975A_ABST
Patent Text Reader

Abstract

The application discloses a kind of wheat molecular design breeding auxiliary methods, system, equipment and medium, comprising: obtaining the genotype VCF data file of measured wheat and pre-processing;Provide whole genome haplotype distribution atlas display based on reference population database, extract the genotype information of target site in VCF data file, dynamically map digital genotype into haplotype string;Provide genetic similarity threshold search based on pre-stored database to internal sample of reference population, based on effective site in VCF data file, dynamically calculate the N×N two genetic similarity matrix between internal samples of measured sample;When phenotype prediction, output wheat target agronomic trait prediction value and adaptability index prediction value carry out data fusion and statistical quantity calculation;Haplotype atlas, mapping string, genetic similarity matrix and phenotype prediction value are visually displayed.It realizes haplotype accurate visualization, dynamic evaluation of genetic background and multi-target double-axis instant phenotype prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the fields of bioinformatics and crop genetics and breeding, specifically relating to a method, system, device and medium for assisting wheat molecular design breeding based on Web architecture and parallel computing. Background Technology

[0002] Wheat is one of my country's most important food crops, and increasing wheat yield is a crucial strategic goal for ensuring food security. Traditional wheat breeding methods rely solely on phenotypic selection, failing to directly track gene delivery, leading to problems such as high breeding blindness, low efficiency, and long cycles. With the development of high-throughput sequencing technology, genomic selection (GS) has become an important tool for accelerating breeding. However, existing public wheat databases (such as Wheat@URGI and the IWGSC data portal) primarily serve as repositories for genomic variation, expression data, and reference sequences, lacking integrated analytical tools and breeding decision support functions. Existing technologies have the following shortcomings:

[0003] (1) Existing databases focus on basic data queries and lack integrated analysis tools that directly link genotypes and phenotypes;

[0004] (2) Most existing genome prediction models use whole-genome random markers, which fail to decode the core genetic signals in adaptive differentiation and breeding selection trajectories. This results in a lack of biological basis for feature marker screening, which limits the accuracy and generalization ability of genome design prediction for complex traits.

[0005] (3) There is a lack of a one-stop platform that can simultaneously realize haplotype visualization, genetic background analysis and phenotypic prediction, making it difficult for breeders to quickly obtain genetic information that can be used for decision-making;

[0006] (4) Due to the extremely high dimensionality of genotype data (usually involving more than 100,000 mutation sites) and the intensive computation of machine learning model inference, the existing conventional Web architecture is prone to browser rendering blockage and server main process memory overflow when performing large-scale haplotype matrix rendering and real-time phenotypic prediction, making it difficult to meet the high availability requirements of industrial applications.

[0007] Therefore, developing a molecular design breeding auxiliary system that can integrate genetic resources, overcome the bottlenecks of massive data rendering and high-concurrency computing, and provide visualization analysis and accurate prediction is of great practical significance. Summary of the Invention

[0008] The technical problem to be solved by this invention is to overcome the shortcomings of existing wheat breeding databases, such as limited functionality and disconnect between basic genomics discoveries and actual data-driven breeding. This invention provides a method, system, equipment, and medium for assisting wheat molecular design breeding that integrates multidimensional genetic resources, breaks through the bottlenecks of massive data rendering and high-concurrency computing, and achieves accurate haplotype visualization, dynamic assessment of genetic background, and real-time biaxial phenotypic prediction of multiple traits.

[0009] In a first aspect, the first embodiment of the present invention provides a method for assisting molecular design breeding of wheat, comprising:

[0010] Obtain the VCF data file of the wheat genotype to be tested uploaded by the user, and preprocess the VCF data file;

[0011] It provides a genome-wide haplotype distribution map based on a pre-constructed reference population database. The reference population database contains reference population genotype data, GWAS associated locus information, haplotype block division information, and superior haplotype identifiers. Using the associated QTL loci pre-stored in the database as target loci, the genotype information of the target loci in the VCF data file of the wheat genotype to be tested is extracted. By comparing the physical location with the variant alleles, the digital genotype is dynamically mapped to a haplotype string. The mapped haplotype string is then compared with the superior haplotype patterns defined in the reference population database, and a regular expression matching algorithm is used to automatically identify and mark superior haplotypes.

[0012] Genetic similarity threshold retrieval based on a pre-stored database is provided for samples within the reference population. Based on the effective sites in the VCF data file of the wheat genotype to be tested, the N×N pairwise genetic similarity matrix between samples within the test sample is dynamically calculated through vectorization operations.

[0013] When performing phenotypic prediction, in response to user-triggered prediction configuration commands, the core features of the target agronomic traits and fitness index are extracted independently based on a dual-axis parallel prediction strategy. Missing genotypes are automatically imputed to reference alleles and converted into mutually independent numerical feature matrices. After feature name normalization and alignment, the numerical feature matrices are input into the corresponding pre-trained genome selection model using an independent subprocess sandbox isolation mechanism. The output predicted values ​​of wheat target agronomic traits and fitness index are then fused and statistically calculated.

[0014] The whole genome haplotype distribution map, haplotype strings, genetic similarity matrix and phenotypic prediction values ​​are visualized and can be exported locally in spreadsheet format.

[0015] Furthermore, the pre-trained genome selection model includes a target agronomic trait prediction model and a fitness index prediction model. It is constructed using a biaxial parallel feature screening and model training pipeline, specifically including: based on genotype data and agronomic trait phenotypic data from a reference population database, performing genome-wide association analysis using a mixed linear model, estimating strong linkage disequilibrium using D' confidence intervals to divide haplotype blocks, extracting several candidate SNPs with the smallest P-values ​​within each block and retaining the SNP with the highest phenotypic variation explanation rate as the block representative marker; filtering the block representative markers according to their P-values ​​for extreme significance, and extracting a preset number of extremely significant sites to construct a core feature set of agronomic traits.

[0016] Using the EigenGWAS method, association analysis was performed on the first few principal components as phenotypes. The overlapping regions that showed significant differences in multiple principal components were extracted as selected evolutionary sites. The absolute values ​​of the allele frequency differences between different ecological subpopulations were calculated as site weights. The whole genome weighted score was calculated based on the homozygous genotype matching status of the sample variation sites. After normalization, the fitness index of each sample was generated and used as the prediction target and corresponding core feature set of the fitness index prediction model.

[0017] The core feature sets of agronomic traits and the core feature sets of fitness indices are respectively input into the corresponding target agronomic trait prediction models and fitness index prediction models. Cross-validation combined with grid search technology is used to independently perform step-by-step grid search optimization, and heuristic fine-tuning strategies of reducing the learning rate and increasing the iterative tree are used to generate and save the final pre-trained genome selection models.

[0018] Furthermore, the step-by-step grid search optimization specifically includes: model complexity optimization, random sampling rate optimization, and learning rate and regularization coefficient optimization.

[0019] Furthermore, the optimization of model complexity includes: on the basis of a fixed learning rate, performing a grid search to obtain the optimal parameters in the first round for the maximum depth of the tree, the number of trees, and the minimum weight of the leaf nodes;

[0020] The optimization of the random sampling rate includes: inheriting the best parameters from the first round, and performing a grid search on the sample sampling rate and the tree feature sampling rate to obtain the best parameters for the second round;

[0021] The optimization of the learning rate and regularization coefficient includes: inheriting the best parameters from the first and second rounds, performing a grid search on the learning rate, L1 regularization coefficient, and L2 regularization coefficient to obtain the best parameters from the third round.

[0022] Furthermore, the dynamic calculation of the N×N pairwise genetic similarity matrix among samples under test through vectorization operations specifically includes:

[0023] Missing sites are automatically filtered out through vectorization operations. The number of non-missing sites with completely identical genotypes between any two test samples is counted to obtain the statistical count. The statistical count is then divided by the total number of non-missing sites in both samples to obtain the percentage of genetic similarity between the samples.

[0024] Furthermore, the independent subprocess sandbox isolation mechanism specifically includes: upon receiving a dual-axis prediction instruction covering the target phenotype and fitness index, generating corresponding numerical feature matrices and saving them to a temporary cache directory, dynamically generating independent prediction execution scripts for each prediction target; and using the operating system's underlying process creation module to invoke the model inference environment to execute the scripts and load the serialized model files for inference.

[0025] The system reads the prediction results and statistics written to a temporary file, performs data fusion on the multi-axis prediction results based on the sample number, and encapsulates the data into a JSON data stream before returning it to the front-end for rendering. Afterward, the temporary script and cache directory are automatically destroyed. Furthermore, the visualization display for the variant site data table enables the virtual DOM deferred rendering mode of the front-end rendering plugin, and disables automatic column width calculation and global front-end sorting in the initial configuration.

[0026] Secondly, another embodiment of the present invention provides a wheat molecular design breeding assistance system for implementing the method described in the first embodiment of the present invention, comprising:

[0027] The data acquisition and preprocessing module is used to acquire the VCF data file of the wheat genotype to be tested uploaded by the user, and to preprocess the VCF data file of the wheat genotype to be tested.

[0028] The haplotype parsing and display module is used to provide a whole-genome haplotype distribution map display based on a pre-constructed reference population database; extract the target site information from the VCF data file of the wheat genotype to be tested, dynamically map the digital genotype to a haplotype string by comparing the physical location with the variant allele, and automatically identify and mark superior haplotypes using a regular expression matching algorithm;

[0029] The genetic background analysis module is used to provide genetic similarity threshold retrieval based on a pre-stored database for samples within the reference population, and to dynamically calculate the N×N pairwise genetic similarity matrix between samples within the test sample through vectorization operations based on the effective sites in the VCF data file of the wheat genotype to be tested.

[0030] The phenotypic real-time prediction module, in response to user-triggered prediction configuration commands, independently extracts the core features of the target agronomic traits and fitness index based on a dual-axis parallel prediction strategy. It automatically imputes missing genotypes with reference alleles and converts them into mutually independent numerical feature matrices. After feature name normalization and alignment, it uses an independent subprocess sandbox isolation mechanism to input the numerical feature matrices into the corresponding pre-trained genome selection model, and performs data fusion and statistical calculation on the output wheat target agronomic trait prediction values ​​and fitness index prediction values.

[0031] The results display and export module is used to visualize the whole genome haplotype distribution map, haplotype strings, genetic similarity matrix and phenotypic prediction values, and supports local export in spreadsheet format.

[0032] Thirdly, another embodiment of the present invention provides an electronic device comprising: a processor, an input device, an output device, and a memory, wherein the processor, the input device, the output device, and the memory are interconnected, the memory is used to store a computer program, the computer program includes program instructions, and the processor is configured to invoke the program instructions to execute the method described in the first embodiment of the present invention.

[0033] Fourthly, an embodiment of the present invention provides a computer-readable storage medium storing a computer program, the computer program including program instructions, which, when executed by a processor, cause the processor to perform the method described in the first embodiment of the present invention.

[0034] The beneficial effects of this invention are:

[0035] (1) A pioneering dual-axis parallel feature screening and evolutionary signal quantification strategy significantly improves model accuracy and biological significance: This invention abandons the blindness and crude feature pooling drawbacks of traditional whole-genome random markers. On the one hand, for agronomic traits, core loci are accurately locked through strong linkage disequilibrium (LD) block dimensionality reduction and highly significant filtering; on the other hand, based on the population ecological subpopulation structure, selected evolutionary loci are extracted, and a quantified fitness index is innovatively constructed based on "allele frequency difference and homozygous matching status" as the prediction target. Combined with step-by-step grid cross-optimization, model overfitting is effectively prevented. Experiments show that the model achieves a prediction determination coefficient (R²) of 0.85 for plant height (PH) and a prediction correlation coefficient of 0.72 for normal sowing grain width (GW_NS), realizing high-precision phenotypic prediction.

[0036] (2) Pioneering a dual-axis parallel prediction and data fusion architecture, breaking the blind spot of single-phenotype evaluation: Traditional prediction models often only target a single trait, resulting in the selected materials potentially having poor overall adaptability. This system innovatively designs a dual-axis parallel prediction strategy of "target phenotype + adaptability index". The system's underlying layer can automatically perform dual-stream concurrent feature extraction, missing value imputation, and sandbox inference for different prediction targets, and finally perform accurate data fusion based on sample numbers. This mechanism enables breeders to simultaneously control the comprehensive environmental adaptability potential of materials while focusing on specific yield or quality traits, significantly improving the overall decision-making quality of parent selection.

[0037] (3) Innovative underlying sandbox isolation and multi-process computing architecture to overcome the high-concurrency computing power bottleneck of the Web client: In response to the problem of server main process memory overflow and GIL lock blocking that is easily caused when massive genotype data is input into large machine learning models, this invention innovatively develops an independent child process sandbox isolation mechanism. Through dual-axis dynamic cache generation, environment isolation inference and process instant destruction lifecycle management, the molecular breeding auxiliary platform is guaranteed to have extremely high stability and data security under high concurrency requests.

[0038] (4) Extreme front-end performance optimization based on virtual DOM mechanism, achieving second-level display of data matrices of hundreds of thousands: In response to the freezing phenomenon when browsers render ultra-large-scale genotype matrices, this system innovatively integrates DataTables deferred rendering technology, supplemented by a conditional style dynamic binding algorithm, to achieve millisecond-level display and smooth scrolling of large-scale data matrices. This greatly reduces the time cost for breeders to compare superior and inferior haplotypes in large-scale populations.

[0039] (5) Constructing a dual-track analysis closed loop of "static benchmark + dynamic tracking" to empower accurate decision-making: This invention breaks through the limitation of existing public databases that only perform passive storage. For haplotypes, it not only provides a static haplotype distribution map based on the whole genome variation of the reference population as a decision benchmark, but also uses a regular matching algorithm to target and highlight superior alleles in dynamically uploaded test samples; for genetic background, it provides efficient threshold retrieval within the reference population and supports real-time N×N similarity vectorization calculation between unknown test materials. By quantifying bloodline differences, it can effectively avoid bottlenecks caused by narrow genetic bases in breeding and accelerate the breeding process of breakthrough new varieties. Attached Figure Description

[0040] Figure 1 A flowchart of a wheat molecular design breeding-assisted method provided in the first embodiment of the present invention;

[0041] Figure 2 A detailed flowchart of a wheat molecular design breeding-assisted method provided in an embodiment of the present invention;

[0042] Figure 3 An architectural diagram of a wheat molecular design breeding auxiliary system provided in an embodiment of the present invention;

[0043] Figure 4 A structural block diagram of a wheat molecular design breeding auxiliary system provided in an embodiment of the present invention;

[0044] Figure 5 This is a schematic diagram of the haplotype accurate tracking and visualization analysis interface in an embodiment of the present invention;

[0045] Figure 6 This is a schematic diagram of the interface for dynamic calculation of genetic background and display of N×N similarity matrix in an embodiment of the present invention;

[0046] Figure 7 This is a schematic diagram of the prediction result display interface in an embodiment of the present invention. Detailed Implementation

[0047] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, and to make the above-mentioned objectives, features and advantages of the embodiments of the present invention more apparent and understandable, the technical solutions in the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0048] In the description of this invention, unless otherwise specified and limited, it should be noted that the term "connection" should be interpreted broadly. For example, it can be a mechanical connection or an electrical connection, or it can be a connection between two internal components. It can be a direct connection or an indirect connection through an intermediate medium. Those skilled in the art can understand the specific meaning of the above term according to the specific circumstances.

[0049] Unless otherwise specified, the data processing, model building, and system development methods described in the following embodiments are conventional methods in the art, performed according to the techniques or conditions described in the literature or in accordance with relevant software libraries / product manuals. Unless otherwise specified, the datasets, algorithm packages, and software used in the following embodiments are all available from public or commercial sources.

[0050] Example 1

[0051] See Figure 1 2. The first embodiment of the present invention provides a method for assisting wheat molecular design breeding, comprising the following steps:

[0052] Obtain the VCF data file of the wheat genotype to be tested uploaded by the user, and perform parsing and effective site preprocessing;

[0053] It provides a genome-wide haplotype distribution map based on a pre-constructed reference population database. The reference population database contains reference population genotype data, GWAS associated locus information, haplotype block division information, and superior haplotype identifiers. Using the associated QTL loci pre-stored in the database as target loci, the genotype information of the target loci in the VCF data file of the wheat genotype to be tested is extracted. By comparing the physical location with the variant alleles, the digital genotype is dynamically mapped to a haplotype string. The mapped haplotype string is then compared with the superior haplotype patterns defined in the reference population database. A regular expression matching algorithm is used to automatically identify and highlight superior haplotypes.

[0054] For samples within the reference population, a genetic similarity threshold retrieval based on a pre-stored database is provided. Based on the effective sites in the wheat genotype data file to be tested, an N×N pairwise genetic similarity matrix between samples to be tested is dynamically calculated through vectorization operations.

[0055] In performing phenotypic prediction, the system innovatively employs a dual-axis parallel prediction strategy. Responding to user-triggered prediction configuration commands, the system independently loads the model feature lists for the target agronomic trait and the fitness index, performing two isolated feature site extractions on the uploaded VCF file. After automatically imputing missing genotypes (. / .) to reference alleles (0 / 0), two independent pure numerical feature matrices are generated. After feature name normalization and alignment, the system uses an independent subprocess sandbox isolation mechanism to invoke independent underlying environments and input the above matrices into the corresponding pre-trained XGBoost genome selection model. Based on the operating system's underlying layer, mutually isolated subprocess runtime environments are created. In each independent subprocess, the corresponding pre-trained genome selection model is loaded for inference calculations, while the main process asynchronously retrieves the results and releases subprocess resources. After inference is complete, the system reads the output results, performs data fusion (Merge) between the target phenotypic prediction value and the fitness index prediction value based on the sample's unique ID, and calculates statistics.

[0056] The whole genome haplotype distribution map, haplotype strings, genetic similarity matrix and phenotypic prediction values ​​are visualized through a front-end page based on virtual DOM deferred rendering, and can be exported locally in Excel / CSV spreadsheet format.

[0057] In the section providing a genome-wide haplotype distribution map based on a pre-constructed reference population database, the logic for constructing and displaying the underlying benchmark map specifically includes: identifying and extracting significantly associated QTL regions based on the genome-wide association analysis (GWAS) results of the reference population on 14 key agronomic traits; for each QTL, haplotypes are classified according to their corresponding significant SNP genotypes, and combined with the statistical differences in the corresponding phenotypic data, defining superior haplotypes (HAPs) and inferior haplotypes (haps) that enhance the target trait; the haplotype genotyping data of 483 reference populations at all key QTL loci are structured and pre-stored in the underlying database to construct a genome-wide haplotype matrix; during front-end display, the haplotype matrix is ​​read globally, and the superior haplotypes are highlighted and visualized using a conditional rendering mechanism, providing a static map benchmark for the dynamic mapping and comparison of haplotypes of subsequent test samples.

[0058] The wheat genotype data file to be tested is a standard VCF format file uploaded by the user. To ensure the absolute accuracy of the physical coordinates of subsequent feature mapping and the strict alignment of model features, the system mandates that the VCF file uploaded by the user must be generated based on a reference genome (Chinese Spring IWGSC RefSeq v1.0 version) consistent with the underlying model training benchmark, and must only contain variation data files of biallelic alleles. The system parses the acquired VCF file. When executing the phenotypic prediction branch, in response to the user's prediction configuration instructions, it automatically imputes missing genotypes (. / .) with reference alleles (0 / 0) and converts them into mutually independent numerical feature matrices. In the haplotype parsing stage, the system accurately extracts target site information, compares the chromosomal physical location with the variant alleles, and dynamically maps the numerical genotype to a haplotype string. A case-sensitive regular expression matching algorithm is used to automatically identify and highlight superior haplotypes (HAP) and inferior haplotypes (hap).

[0059] In the genetic background analysis phase, the system provides a similarity threshold retrieval function for pre-built reference populations; for user-uploaded test samples, the system performs high-performance real-time N×N matrix similarity calculations. Specifically, through NumPy vectorized matrix operations, missing sites (-1 or NA) are automatically filtered out, and the number of genotype-completely identical and non-missing sites between any two samples is counted. This number is then divided by the total number of non-missing sites in both samples to obtain the percentage of genetic similarity between the samples. Highly similar related samples are extracted from the top-ranked samples, accurately quantifying the genetic lineage between unknown materials and effectively avoiding the selection of parents with a narrow genetic background.

[0060] The pre-trained genome selection model described above employs a multi-stage optimized feature selection and parameter tuning pipeline, specifically including:

[0061] (1) Block division and representative extraction: Resequencing data of 483 wheat germplasms and phenotypic data of key agronomic traits (such as plant height, grain width, etc.) were obtained. Genome-wide association analysis (GWAS) was performed using a mixed linear model (MLM) with a significance threshold of -log10(P) = 7.0.

[0062] Based on GWAS results, PLINK software (v1.9) was used to segment whole-genome haplotype regions. Specifically, following the method of Gabriel et al., strong linkage disequilibrium (LD) was estimated by calculating D' confidence intervals, with a lower bound of 0.70 and an upper bound of 0.98 for the 90% confidence interval. Within each independent block, the top three candidate SNPs (single nucleotide polymorphisms) with the smallest P-values ​​were first extracted, and then the SNP with the highest phenotypic variation explained (PVE) was selected as the representative marker for that block. Subsequently, the representative markers of the blocks were filtered for extreme significance based on P-values, and the top 10,000 extremely significant sites were selected to construct the core feature set of agronomic traits.

[0063] (2) Construction of Adaptive Phenotypes and Feature Sets: Based on the population structure, different ecological subgroups were divided. Using the EigenGWAS method, association analysis was performed using the first 10 principal components (PCs) as phenotypes. Significantly overlapping intervals in 3 or more PCs were extracted, resulting in a total of 522 selected evolutionary sites to construct the core feature set of the adaptation index. The quantitative calculation logic is as follows: For any selected evolutionary site i, the absolute value of the difference in allele frequency between ecological subgroup A and ecological subgroup B is recorded as the site weight. The score S of the sample at this site. i The determination rule is as follows: if the sample genotype is homozygous mutant and the frequency of the mutant allele in subgroup A is higher than that in subgroup B, or if the genotype is homozygous reference and the frequency of the reference allele in subgroup A is lower than that in subgroup B, then... ,

[0064] otherwise The whole-genome weighted total score of the sample is Normalized fitness index This serves as the direct prediction target for the adaptive prediction model.

[0065] (3) Dual-axis parallel grid search and model training: The 483 samples were randomly divided into a training set and an independent test set at an 8:2 ratio. The 10,000 agronomic features and 522 adaptive features were extracted and input into the corresponding XGBoost models for parallel training. For both sets of core features, 5-fold cross-validation combined with GridSearchCV was used to perform stepwise hyperparameter optimization of the XGBoost model, sequentially optimizing model complexity, random sampling rate, learning rate, and regularization coefficient to prevent overfitting.

[0066] Step-by-step hyperparameter optimization is divided into the following three tuning stages:

[0067] 1. Model complexity optimization: With a fixed learning rate (0.1), a grid search is performed to obtain the optimal parameters in the first round for the maximum depth of the tree (max_depth is [3, 5, 7]), the number of trees (n_estimators is [100, 200]), and the minimum weight of the leaf nodes (min_child_weight is [1, 3, 5]).

[0068] 2. Model randomness optimization: Inheriting the best parameters from the previous round, a grid search is performed on the sample sampling rate (subsample takes values ​​of [0.7, 0.8, 0.9]) and the tree feature sampling rate (colsample_bytree takes values ​​of [0.5, 0.7, 0.9]) to obtain the second round of best parameters.

[0069] 3. Learning Rate and Regularization Optimization: Inheriting the best parameters from the first two rounds, a fine search is performed on the learning rate (learning_rate takes values ​​of [0.05, 0.1]) and L1 / L2 regularization coefficients (reg_alpha takes values ​​of [0, 0.1, 1] and reg_lambda takes values ​​of [1, 2, 5]) to obtain the best parameters for the third round.

[0070] (4) Heuristic fine-tuning and model serialization:

[0071] After obtaining the complete set of optimal hyperparameters, the final heuristic optimization strategy is implemented: the optimal learning rate obtained from grid search is reduced to 0.1 times its original value, while the number of decision trees is increased to 3 times its original value to improve the model's generalization ability and prediction smoothness. The optimized XGBoost model is retrained using the full training set, and its coefficient of determination (R²), root mean square error (RMSE), and Pearson correlation coefficient are calculated on the independent test set. After training, the dual-axis model weights and corresponding feature lists are serialized and saved using joblib for immediate access by the web platform.

[0072] This method pioneers a dual-axis parallel feature screening and evolutionary signal quantification strategy. On one hand, it accurately identifies core agronomic trait loci through LD block dimensionality reduction and highly significant filtering; on the other hand, it innovatively constructs a quantified fitness index based on the absolute value of allele frequency differences and homozygous matching status. Combined with step-by-step grid cross-validation optimized by parameter pruning, this not only significantly shortens model training time but also effectively prevents overfitting. Experiments show that the model achieves a prediction coefficient of determination R² of 0.85 for plant height (PH) and a correlation coefficient of 0.72 for normal sowing grain width (GW_NS), placing its prediction accuracy at an industry-leading level.

[0073] Employing deferred rendering technology based on a virtual DOM to process massive genotype matrices (with over 100,000 cells), it achieves second-level browsing and smooth retrieval of large-scale data. It supports one-click export of haplotype matrices, dynamic similarity matrices, and comprehensive decision reports based on biaxial phenotypic predictions in Excel / CSV format. This not only significantly reduces the time cost of analyzing and visualizing massive genetic data but also, through a dual-axis linkage mechanism of "yield and quality + environmental adaptability," enables breeders to simultaneously control the comprehensive adaptability potential of materials while focusing on specific agronomic traits, greatly improving the overall decision-making quality and success rate of molecular breeding.

[0074] Example 2

[0075] like Figure 3 , 4 As shown, another embodiment of the present invention provides a wheat molecular design breeding auxiliary system for implementing the method described in the first embodiment above.

[0076] The system adopts a classic B / S architecture. The backend is built on the Flask micro-framework in Python, responsible for route distribution and high-intensity matrix calculations; the frontend uses HTML5 + Bootstrap 5 + jQuery to build a responsive user interface. Data interaction uses AJAX asynchronous communication technology based on HTTP to achieve refresh-free loading. The system has a built-in global session repeater, supporting one-click dynamic switching between Chinese and English (zh / en) template variables.

[0077] The system includes the following functional modules:

[0078] Data Acquisition and Preprocessing Module: This module acquires and preprocesses the VCF data files of wheat genotypes uploaded by users. As the data entry point for dynamic analysis, this module receives standard VCF format files. To ensure the validity of the analysis results, this module requires that the received input data be: biallelic data files generated based on alignment with the underlying benchmark reference genome (Chinese Spring IWGSC RefSeq v1.0 version) and with multiple allelic variations removed. This module integrates a genotype missing value handling mechanism, providing a standardized data source for subsequent model inference.

[0079] The haplotype parsing and display module provides a genome-wide haplotype distribution map based on a pre-constructed reference population database. This database contains reference population genotype data, GWAS-associated loci information, haplotype block division information, and superior haplotype identifiers. Using the pre-stored associated QTL loci in the reference population database as target loci, the module extracts the genotype information of the target loci from the wheat genotype VCF data file to be tested. By comparing physical location with variant alleles, the digital genotype is dynamically mapped to a haplotype string. The mapped haplotype string is then compared with the superior haplotype patterns defined in the reference population database. A regular expression matching algorithm is used to automatically identify and label superior haplotypes. In practical implementation, for static benchmarks, the system globally renders the haplotype distribution matrices of 483 reference populations and globally visualizes the 274 pre-identified key QTLs. For dynamic parsing, the parsing engine compares chromosome numbers (CHROMs) and physical locations (POSs) row by row. If a target QTL is matched, regular expressions and conditional style bindings are used to render cells containing superior haplotype identifiers (HAPs) with a prominent highlight background color (such as soft pink), achieving... Figure 5 The intuitive visual tracking shown.

[0080] The genetic background analysis module provides genetic similarity threshold retrieval based on a pre-stored database for samples within the reference population. Based on valid loci in the VCF data file of the wheat genotype to be tested, it dynamically calculates the N×N pairwise genetic similarity matrix between samples using vectorized operations. This module comprises two functions: first, it allows users to dynamically retrieve genetic relationships within the reference population by setting similarity thresholds; second, it automatically filters out missing loci (-1 or NA) using NumPy vectorized matrix operations, dynamically calculates the N×N pairwise similarity matrix between unknown materials (the formula is: number of genotypes that are completely identical and not missing / total number of non-missing loci), and the backend extracts the data subset with the highest correlation and sends it back to the frontend to generate... Figure 6 The multi-tabbed high-similarity sample list view shown provides real-time background screening for parent selection.

[0081] The phenotypic real-time prediction module, in response to user-triggered prediction configuration commands during phenotypic prediction, independently extracts the core features of the target agronomic traits and fitness index based on a dual-axis parallel prediction strategy. It automatically imputes missing genotypes with reference alleles and converts them into independent numerical feature matrices. After feature name normalization and alignment, an independent subprocess sandbox isolation mechanism is used to input the numerical feature matrices into the corresponding pre-trained genome selection model. The output predicted values ​​of wheat target agronomic traits and fitness index are then fused and statistically calculated. To avoid memory overflow and thread blocking (GIL) in the web main process due to loading a large machine learning model, and to prevent feature pollution during multi-model concurrency, the backend innovatively employs an independent subprocess sandbox isolation mechanism and a data fusion pipeline. Specifically, after receiving the prediction command, the main program loads the feature dictionaries of the target trait and fitness index, performs two rigorous targeted parsings on the VCF file, generates two independent numerical feature matrices, and saves them to a temporary cache directory. Subsequently, the system dynamically generates dedicated prediction execution scripts, which concurrently invoke isolated model inference environments (such as Conda environments) through the underlying subprocess module to perform computations independently. After loading the corresponding XGBoost model for inference and outputting results in the isolated sandbox, the main process uses the Pandas library to merge the multi-stream prediction results into a complete decision wide table based on sample identifiers, calculates statistics such as mean and extreme values, and finally encapsulates it into a JSON data stream and returns it to the front-end for rendering (e.g., ...). Figure 7 (As shown), and automatically destroys temporary scripts and caches. This architecture greatly ensures extremely high stability and data security during multi-target concurrent prediction.

[0082] The results display and export module visualizes the whole-genome haplotype distribution map, haplotype strings, genetic similarity matrix, and phenotypic predictions through a front-end page based on virtual DOM deferred rendering, and supports local export in Excel / CSV spreadsheet format. This module integrates the DataTables plugin for massive genotype matrices, forcibly enabling deferRender (virtual DOM deferred rendering) mode, and disabling automatic column width calculation (autoWidth: false) and global front-end sorting (ordering: false) in the initial configuration. This ensures that the system can still achieve millisecond-level display and smooth scrolling even when dealing with hundreds of samples and a large number of variant sites.

[0083] The execution process of each of the above modules can be carried out in accordance with the steps of the wheat molecular design breeding auxiliary method provided in the first embodiment, and will not be described in detail in this embodiment.

[0084] The wheat molecular design breeding assistance system provided in this embodiment of the invention is based on the same inventive concept and has the same beneficial effects as the wheat molecular design breeding assistance method provided in the first embodiment above, and will not be described again here.

[0085] Another embodiment of the present invention provides an electronic device, which includes a processor, an input device, an output device, and a memory. The processor, the input device, the output device, and the memory are interconnected. The memory is used to store a computer program, which includes program instructions. The processor is configured to call the program instructions to execute the method described in the first embodiment above.

[0086] It should be understood that, in the embodiments of the present invention, the processor may be a Central Processing Unit (CPU), but it may also be other general-purpose processors, graphics processing units (GPUs), digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.

[0087] Input devices may include touchpads, microphones, etc., and output devices may include displays (LCDs, etc.), speakers, etc.

[0088] The memory may include read-only memory and random access memory, and provides instructions and data to the processor. A portion of the memory may also include non-volatile random access memory. For example, the memory may also store information about the device type.

[0089] In specific implementations, the processor, input device, and output device described in the embodiments of the present invention can execute the implementation methods described in the method embodiments of the present invention, or they can execute the implementation methods described in the system embodiments of the present invention, which will not be repeated here.

[0090] The present invention also provides an embodiment of a computer-readable storage medium storing a computer program, the computer program including program instructions, which, when executed by a processor, cause the processor to perform the methods described in the above embodiments.

[0091] The computer-readable storage medium can be an internal storage unit of the terminal described in the foregoing embodiments, such as the terminal's hard drive or memory. The computer-readable storage medium can also be an external storage device of the terminal, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card equipped on the terminal. Furthermore, the computer-readable storage medium can include both internal storage units and external storage devices of the terminal. The computer-readable storage medium is used to store the computer program and other programs and data required by the terminal. The computer-readable storage medium can also be used to temporarily store data that has been output or will be output.

[0092] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0093] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the terminals and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0094] In the several embodiments provided in this application, it should be understood that the disclosed terminals and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices or units, or may be electrical, mechanical or other forms of connection.

[0095] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention, and they should all be covered within the scope of the claims and specification of the present invention.

Claims

1. A method for assisting molecular design breeding of wheat, characterized in that, include: Obtain the VCF data file of the wheat genotype to be tested uploaded by the user, and preprocess the VCF data file; It provides a genome-wide haplotype distribution map based on a pre-constructed reference population database. The reference population database contains reference population genotype data, GWAS associated loci information, haplotype block division information, and superior haplotype identifiers. Using the associated QTL loci pre-stored in the reference population database as target loci, the genotype information of the target loci in the VCF data file of the wheat genotype to be tested is extracted. By comparing the physical location with the variant alleles, the digital genotype is dynamically mapped to a haplotype string. The mapped haplotype string is then compared with the superior haplotype patterns defined in the reference population database. A regular expression matching algorithm is used to automatically identify and mark superior haplotypes. Genetic similarity threshold retrieval based on a pre-stored database is provided for samples within the reference population. Based on the effective sites in the VCF data file of the wheat genotype to be tested, the N×N pairwise genetic similarity matrix between samples within the test sample is dynamically calculated through vectorization operations. When performing phenotypic prediction, in response to the prediction configuration command triggered by the user, the core features of the target agronomic traits and fitness index are extracted independently based on the dual-axis parallel prediction strategy. Missing genotypes are automatically imputed to reference alleles and converted into mutually independent numerical feature matrices. After feature name normalization and alignment, the numerical feature matrices are input into the corresponding pre-trained genome selection model using an independent subprocess sandbox isolation mechanism. The output wheat target agronomic trait prediction values ​​and fitness index prediction values ​​are fused and statistically calculated. The whole genome haplotype distribution map, haplotype strings, genetic similarity matrix and phenotypic prediction values ​​are visualized and can be exported locally in spreadsheet format.

2. The wheat molecular design breeding-assisted method according to claim 1, characterized in that, The pre-trained genome selection model includes a target agronomic trait prediction model and a fitness index prediction model, and is constructed using a biaxial parallel feature selection and model training pipeline, specifically including: Based on genotype data and agronomic phenotypic data from a reference population database, genome-wide association analysis was performed using a mixed linear model. Strong linkage disequilibrium was estimated using D' confidence intervals to divide haplotype blocks. Within each block, several candidate SNPs with the smallest P-values ​​were extracted, and the SNP with the highest phenotypic variation explanatory power was retained as the block representative marker. The block representative markers were filtered for extreme significance based on P-values, and a predetermined number of extremely significant sites were selected to construct a core feature set of agronomic traits. Using the EigenGWAS method, association analysis was performed on the first few principal components as phenotypes. The overlapping regions that showed significant differences in multiple principal components were extracted as selected evolutionary sites. The absolute values ​​of the allele frequency differences between different ecological subpopulations were calculated as site weights. The whole genome weighted score was calculated based on the homozygous genotype matching status of the sample variation sites. After normalization, the fitness index of each sample was generated and used as the prediction target and corresponding core feature set of the fitness index prediction model. The core feature sets of agronomic traits and the core feature sets of fitness indices are respectively input into the corresponding target agronomic trait prediction models and fitness index prediction models. Cross-validation combined with grid search technology is used to independently perform step-by-step grid search optimization, and heuristic fine-tuning strategies of reducing the learning rate and increasing the iterative tree are used to generate and save the final pre-trained genome selection models.

3. The wheat molecular design breeding-assisted method according to claim 2, characterized in that, The step-by-step grid search optimization specifically includes: model complexity optimization, random sampling rate optimization, and learning rate and regularization coefficient optimization.

4. The wheat molecular design breeding-assisted method according to claim 3, characterized in that, The optimization of model complexity includes: on a fixed learning rate, performing a grid search to obtain the first round of optimal parameters for the maximum depth of the tree, the number of trees, and the minimum weight of the leaf nodes; The optimization of the random sampling rate includes: inheriting the best parameters from the first round, and performing a grid search on the sample sampling rate and the tree feature sampling rate to obtain the best parameters for the second round; The optimization of the learning rate and regularization coefficient includes: inheriting the best parameters from the first and second rounds, performing a grid search on the learning rate, L1 regularization coefficient, and L2 regularization coefficient to obtain the best parameters from the third round.

5. The wheat molecular design breeding-assisted method according to claim 1, characterized in that, The dynamic calculation of the N×N pairwise genetic similarity matrix among samples under test through vectorization operations specifically includes: Missing sites are automatically filtered out through vectorization operations. The number of non-missing sites with completely identical genotypes between any two test samples is counted to obtain the statistical count. The statistical count is divided by the total number of non-missing sites in both samples to obtain the percentage of genetic similarity between the samples.

6. The wheat molecular design breeding-assisted method according to claim 1, characterized in that, The independent subprocess sandbox isolation mechanism specifically includes: Upon receiving a dual-axis prediction configuration instruction covering the target phenotype and fitness index, the corresponding numerical feature matrices are generated and saved to a temporary cache directory. Independent prediction execution scripts are dynamically generated for each prediction target. The operating system creates modules through underlying processes, which then invoke isolated model inference environments to execute the scripts and load serialized model files for inference. Read the prediction results and statistics written to the temporary file, perform data fusion on the multi-axis prediction results based on the sample number, encapsulate them into a JSON data stream and return it to the front end for rendering, and then automatically destroy the temporary script and cache directory.

7. The wheat molecular design breeding-assisted method according to claim 1, characterized in that, The visualization is for the table of variant site data. The virtual DOM deferred rendering mode of the front-end rendering plugin is enabled, and the automatic column width calculation and global front-end sorting functions are disabled in the initialization configuration.

8. A wheat molecular design breeding auxiliary system, characterized in that, For implementing the method as described in any one of claims 1-7, comprising: The data acquisition and preprocessing module is used to acquire the VCF data file of the wheat genotype to be tested uploaded by the user, and to preprocess the VCF data file of the wheat genotype to be tested. The haplotype parsing and display module is used to provide a whole-genome haplotype distribution map display based on a pre-constructed reference population database. The reference population database pre-stores reference population genotype data, GWAS associated locus information, haplotype block division information, and superior haplotype identifiers. Using the associated QTL loci pre-stored in the database as target loci, the genotype information of the target loci in the VCF data file of the wheat genotype to be tested is extracted. By comparing the physical location with the variant alleles, the digital genotype is dynamically mapped to a haplotype string. The mapped haplotype string is then compared with the superior haplotype patterns defined in the reference population database, and a regular expression matching algorithm is used to automatically identify and mark superior haplotypes. The genetic background analysis module is used to provide genetic similarity threshold retrieval based on a pre-stored database for samples within the reference population, and to dynamically calculate the N×N pairwise genetic similarity matrix between samples within the test sample through vectorization operations based on the effective sites in the VCF data file of the wheat genotype to be tested. The phenotypic real-time prediction module, in response to user-triggered prediction configuration commands, independently extracts the core features of the target agronomic traits and fitness index based on a dual-axis parallel prediction strategy. It automatically imputes missing genotypes with reference alleles and converts them into mutually independent numerical feature matrices. After feature name normalization and alignment, it creates mutually isolated subprocess runtime environments based on the operating system. In each independent subprocess, it loads the corresponding pre-trained genome selection model for inference calculation. The main process asynchronously obtains the results and releases the subprocess resources. It then performs data fusion and statistical calculation on the output predicted values ​​of wheat target agronomic traits and fitness index. The results display and export module is used to visualize the whole genome haplotype distribution map, haplotype strings, genetic similarity matrix and phenotypic prediction values, and supports local export in spreadsheet format.

9. An electronic device, comprising: The processor, input device, output device, and memory are interconnected, the memory being used to store a computer program, the computer program including program instructions, characterized in that the processor is configured to invoke the program instructions to perform the method as described in any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program including program instructions that, when executed by a processor, cause the processor to perform the method as described in any one of claims 1-7.