Crop breeding decision analysis system

The crop breeding decision analysis system, which integrates data input, processing, analysis, and display modules, solves the problems of complex operation and low data utilization efficiency in existing technologies, and simplifies the breeding process and enhances decision support capabilities.

CN121961273APending Publication Date: 2026-05-01INSTITUTE OF CROP SCIENCE CHINESE ACADEMY OF AGRICULTURAL SCIENCES
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INSTITUTE OF CROP SCIENCE CHINESE ACADEMY OF AGRICULTURAL SCIENCES
Filing Date
2025-12-18
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing technologies in maize breeding suffer from several drawbacks: they require specialized bioinformatics knowledge and programming skills for operation; the analysis process is fragmented; the functional platforms are limited; they cannot systematically support the entire breeding process; SNP data is not closely linked to breeding decisions; and data utilization efficiency is low.

Method used

This invention provides a crop breeding decision analysis system that integrates data input, processing, analysis, and display modules, including material genetic background analysis, breeding utilization guidance, whole genome selection, and knowledge question answering modules. Through technologies such as sliding window algorithm, preprocessing, and visualization, it simplifies the breeding process and improves data analysis efficiency.

Benefits of technology

It reduces the reliance of breeders on professional skills, simplifies the entire process of analysis, improves the efficiency of data analysis and the integrated support capability for breeding decisions, and promotes the application of SNP data in breeding practice.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121961273A_ABST
    Figure CN121961273A_ABST
Patent Text Reader

Abstract

The invention provides a crop breeding decision analysis system. The crop breeding decision analysis system comprises a data input module, a data processing module, a data analysis module and a result display module, the data input module is used for receiving user material data uploaded by a user; the data processing module is connected with the data input module and is used for preprocessing the user material data; the data analysis module is connected with the data input module and the data processing module, and is used for analyzing the preprocessed user material data based on the sub-module selected by the user and / or the parameters set by the user to obtain an analysis result; and the result display module is connected with the data analysis module and is used for generating a visual chart and / or a data table based on the analysis result and displaying the visual chart and / or the data table to the user. According to the method, the dependence of breeding personnel on professional bioinformatics knowledge and programming skills can be reduced, key analysis steps in the whole breeding process are simplified, the data analysis efficiency and the integrated support capability of breeding decision are improved, and the application of SNP data in breeding practice is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of agricultural information technology, and in particular to a crop breeding decision analysis system. Background Technology

[0002] In modern maize breeding, marker-assisted selection has become a key method to improve breeding efficiency. Among them, single nucleotide polymorphism (SNP) markers are widely used in breeding practices due to their advantages such as large number, wide coverage, and low detection cost.

[0003] Currently, the main technologies for acquiring SNP data include: (1) high-throughput sequencing technologies, such as whole-genome resequencing, which can obtain comprehensive variation information, but are costly and not suitable for large-scale population applications; (2) simplified genome sequencing technologies, such as genotyping-by-sequencing (GBS) and specific-locus amplified fragment sequencing (SLAF-seq), which use enzyme digestion to enrich partial genomic regions for sequencing, reducing costs, but have limited coverage, missing data, and affect the accuracy of analysis; (3) SNP chip technology, which uses pre-set probes for targeted typing, has the characteristics of moderate cost, high data quality, and reliable loci, and is suitable for genotyping of large-scale populations. In terms of data analysis, existing methods mainly rely on local bioinformatics tools (such as ADMIXTURE, CoreHunter) or programming scripts (R, Python, etc.) for processing. There are also some online platforms with more specialized functions, such as systems that focus on genome selection or multi-omics data query, but their functional integration is generally not high.

[0004] While existing technologies provide basic support for genotyping and some analyses, their application throughout the entire breeding process still has significant limitations: First, operation relies on specialized bioinformatics knowledge and programming skills, posing a barrier to entry for most breeders; second, the analysis stages are fragmented, with different tools required for steps such as material evaluation, population division, parent identification, fragment tracing, and core germplasm screening, resulting in cumbersome and inefficient processes; third, existing platforms have relatively limited functionality and lack integrated support capabilities covering key breeding decisions, making it difficult to systematically meet actual breeding needs; furthermore, the connection between SNP data and breeding decisions is not close enough, and the overall data utilization efficiency is low, limiting its further promotion and application in breeding practice. Summary of the Invention

[0005] This invention provides a crop breeding decision analysis system to solve the above-mentioned problems in the prior art.

[0006] This invention provides a crop breeding decision analysis system, comprising the following modules: The system includes a data input module, a data processing module, a data analysis module, and a results display module. The data input module is used to receive user material data uploaded by the user; the user material data includes single nucleotide polymorphism (SNP) chip data; The data processing module is connected to the data input module and is used to preprocess the user material data; The data analysis module connects the data input module and the data processing module, and is used to analyze the preprocessed user material data based on the sub-modules selected by the user and / or the parameters set by the user, and obtain the analysis results. The results display module is connected to the data analysis module and is used to generate visual charts and / or data tables based on the analysis results and display them to the user. The data analysis module includes multiple sub-modules, including at least one of the following: a material genetic background analysis sub-module, a breeding utilization guidance sub-module, and a whole-genome selection sub-module. The material genetic background analysis submodule is used to analyze the genetic background of the user's material data; The breeding utilization guidance submodule is used to make breeding decisions based on the user material data; The whole-genome selection submodule is used to perform phenotypic prediction on the user material data.

[0007] According to the crop breeding decision analysis system provided by the present invention, the material genetic background analysis submodule is specifically used for one or more of the following: Based on a reference population divided into multiple heterosis groups, determine the type of heterosis group to which the user material data belongs; Based on a reference population divided into multiple heterosis groups, the genetic composition of the user material data was analyzed; Using a sliding window algorithm and a pre-set matching threshold, the source of the backbone system of the user material data of the improved system is identified based on the user material data of the improved system and the user material data of the backbone system uploaded by the user. The genomic fragments from the user-uploaded backbone lines in the user-uploaded improved lines are identified by using a sliding window algorithm, a pre-set matching degree threshold, and a minimum recombinant fragment length threshold. When a user uploads multiple sets of user material data for improved lines, the conserved genomic fragments of the multiple sets of user material data for improved lines are determined based on the conservation threshold. The group of the user material data is determined to be either a paternal-biased group or a maternal-biased group.

[0008] According to the crop breeding decision analysis system provided by the present invention, the result display module is specifically used for: Based on the genetic distance between the user material data and the reference population, a phylogenetic tree merging the user material data and the reference population is constructed. Based on the genetic distance between the user material data and the reference population, a list of a preset number of individuals in the reference population is generated, prioritizing those with the closest genetic distance. Based on the genetic components of the user material data, a stacked histogram of the population structure is generated.

[0009] According to the crop breeding decision analysis system provided by the present invention, the breeding utilization guidance submodule is specifically used for one or more of the following: Based on the genotypes of known variety patterns, determine the potential variety patterns of the user material data; Determine the simulated offspring genotypes of the user material data; The application potential of the simulated offspring genotypes is evaluated based on the known superior hybrid genotypes.

[0010] According to the crop breeding decision analysis system provided by the present invention, the breeding utilization guidance submodule is specifically used for: Based on user material data of the target group uploaded by users, calculate the genetic distance matrix between any two individuals in the target group; Based on the genetic distance matrix between any two individuals and the core group ratio set by the user, the core group in the target group is selected.

[0011] According to a crop breeding decision analysis system provided by the present invention, the whole-genome selection submodule is specifically used for one or more of the following: Based on the user-uploaded user material data and the corresponding phenotypic data, the prediction model selected by the user is run to obtain the genomic breeding value of each sample in the user material data.

[0012] According to a crop breeding decision analysis system provided by the present invention, the whole-genome selection submodule is further used for: The hyperparameters of the prediction model are tuned using a grid search method. The prediction model is evaluated using cross-validation.

[0013] According to a crop breeding decision analysis system provided by the present invention, the crop breeding decision analysis system further includes: Knowledge Q&A module; The knowledge-based question-answering module includes a large-scale language model and a crop breeding knowledge base, specifically used for: Transform the user's input question into a question vector; Based on the question vector, a matching process is performed in the crop breeding knowledge base to obtain the multiple text fragments with the highest similarity. The multiple text fragments are input into the large language model to obtain the answer to the input question and provide feedback to the user.

[0014] According to the crop breeding decision analysis system provided by the present invention, the crop breeding knowledge base is constructed in the following manner: We will disassemble, clean, and slice professional articles related to crop breeding collected from multiple sources; The crop breeding knowledge base is constructed based on the embedded model and the sliced ​​text data.

[0015] According to the crop breeding decision analysis system provided by the present invention, the data processing module is specifically used for: The user material data is filtered based on the heterozygosity of each sample in the user material data, and the heterozygosity and / or missing rate of each marker in the user material data; Based on the major allele imputation method, missing markers in the screened user material data are imputed.

[0016] The crop breeding decision analysis system provided by this invention integrates multiple functional sub-modules into a unified platform and combines data preprocessing and visualization functions. This reduces the reliance of breeders on professional bioinformatics knowledge and programming skills, simplifies key analysis steps in the entire breeding process, improves data analysis efficiency and integrated support capabilities for breeding decisions, and promotes the effective application of SNP data in breeding practice. Attached Figure Description

[0017] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0018] Figure 1 This is a schematic diagram of the structure of the crop breeding decision analysis system provided by the present invention.

[0019] Figure 2 This is a schematic diagram of the overall analysis process of MBDP provided by the present invention.

[0020] Figure 3 This is a box plot showing the accuracy distribution of genome prediction based on flowering period phenotypes provided by this invention.

[0021] Figure 4 This is a comparison chart of the accuracy of the model provided in this embodiment of the invention with other publicly available models. Detailed Implementation

[0022] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0023] Figure 1 This is a schematic diagram of the structure of the crop breeding decision analysis system provided by the present invention, as shown below. Figure 1 As shown, the present invention provides a crop breeding decision analysis system, comprising the following modules: The system includes a data input module 100, a data processing module 110, a data analysis module 120, and a result display module 130. The data input module 100 is used to receive user material data uploaded by the user; the user material data includes single nucleotide polymorphism (SNP) chip data; Data processing module 110 is connected to data input module 100 and is used to preprocess user material data; The data analysis module 120 is connected to the data input module 100 and the data processing module 110. It is used to analyze the preprocessed user material data based on the sub-modules selected by the user and / or the parameters set by the user, and to obtain the analysis results. The results display module 130 is connected to the data analysis module 120 and is used to generate visual charts and / or data tables based on the analysis results and display them to the user.

[0024] Specifically, the crop breeding decision analysis system provided in this embodiment of the invention includes at least a data input module, a data processing module, a data analysis module, and a result display module.

[0025] Users can upload their material data through the data input module. It should be noted that, in different embodiments provided by this invention, users can also use the data input module to select the sub-module to be used in the data analysis module, upload relevant information required by that sub-module, and set relevant parameters for that sub-module.

[0026] In some implementations, the crop breeding decision analysis system can pre-set the required format for the data input module and provide format conversion tools for various data types to facilitate data format conversion for users.

[0027] After the crop breeding decision analysis system obtains the user material data uploaded by users, it can preprocess the received user material data through the data processing module. For example, it can filter out data that does not meet the requirements and supplement missing values ​​in the user material data, thereby providing standardized and high-quality data for the subsequent data analysis module.

[0028] In some embodiments, the data processing module 110 is specifically used for: The user material data is filtered based on the heterozygosity of each sample in the user material data, as well as the heterozygosity and / or missing rate of each marker in the user material data. Based on the major allele imputation method, missing markers in the screened user material data are imputed.

[0029] Specifically, the data processing module can perform the following steps to preprocess the user's material data: First, the data processing module can calculate the heterozygosity of each sample (i.e., each individual) in the user's material data across the entire genome. If a sample has a high heterozygosity, it can be removed. The data processing module can also calculate the heterozygosity of each SNP marker (i.e., each SNP locus). If a marker has a high heterozygosity or a high deletion rate, it can be removed. Understandably, the data processing module can filter the user's material data using pre-set thresholds.

[0030] After screening the user material data, some missing SNP markers may still remain. Without processing, this could affect the completeness and accuracy of subsequent analyses. Therefore, the data processing module can further fill in the missing markers in the screened user material data using the major allele imputation method. This method calculates the most common allele (i.e., the major allele) for each SNP locus across all retained samples and uses this allele to fill in the missing values ​​at the corresponding loci. This allows the data processing module to effectively restore the integrity of the data while preserving the original genetic structure as much as possible, enabling subsequent analytical algorithms to run smoothly.

[0031] After the data processing module completes the preprocessing of the user material data, the crop breeding decision analysis system sends the preprocessed user material data to the data analysis module.

[0032] It should be noted that the data analysis module may include multiple functional sub-modules. These sub-modules include at least one of the following: a material genetic background analysis sub-module, a breeding utilization guidance sub-module, and a genome-wide selection sub-module. Each sub-module can be responsible for addressing a specific analytical need in breeding decision-making.

[0033] The Material Genetic Background Analysis Submodule is used to analyze the genetic background of the user's material data. For example, it can perform heterosis group classification and population genetic background analysis on SNP chip data. The Material Genetic Background Analysis Submodule can help users understand the classification and population relationships of each sample in the SNP chip data, which is the basis for hybridization and breeding strategy formulation.

[0034] The breeding utilization guidance submodule is used to make breeding decisions based on the user's material data. For example, it can determine the source of the backbone lines of the improved lines uploaded by the user, providing a basis for understanding the genetic basis of superior traits. In the case of a large population, this submodule can help screen a smaller core population subset that can retain most of the genetic variation from the perspective of genetic diversity, based on the population's SNP chip data, thereby providing support for field trials and resource conservation optimization.

[0035] The whole-genome selection submodule is used to perform phenotypic prediction on the user material data, for example, determining the genomic breeding value of each sample in the SNP chip data. This submodule can integrate phenotypic data and use statistical models or machine learning methods to quantitatively predict the breeding potential of each sample in the user material data, thereby providing data support for users to select samples.

[0036] The data analysis module can perform targeted analysis on the preprocessed data based on the specific sub-modules selected by the user in the data input module and the parameters set by the user, and obtain the corresponding analysis results. Thus, the crop breeding decision analysis system can flexibly meet the specific analysis needs of users in different breeding scenarios.

[0037] Finally, the crop breeding decision analysis system can automatically create visual charts to display to users based on the analysis results generated by the analysis module; or create data tables to display to users; or create both visual charts and data tables to display to users. This helps users understand and utilize the analysis results intuitively and efficiently.

[0038] Compared to existing technologies, the crop breeding decision analysis system provided by this invention effectively solves the problems of reliance on professional skills, fragmented analysis processes, and limited functional platforms in traditional analysis methods by combining data input, processing, analysis, and visualization into a coherent workflow. According to the crop breeding decision analysis system provided by this invention, users can complete the entire process from raw data to visualized results without switching between different tools or writing complex code, significantly lowering the technical barrier to entry and improving the overall efficiency and integrated support capabilities of breeding data analysis.

[0039] In some implementations, the materials genetic background analysis submodule can be used for: Based on a reference population divided into multiple heterosis groups, the heterosis group type to which the user material data belongs is determined.

[0040] Specifically, the materials genetic background analysis submodule can first calculate the genetic distance between the user-uploaded material data and the reference population. This calculation process can be expressed by the following formula: In the formula, This represents the total number of SNP markers, i.e., the number of SNP sites on the genome; It is an index of SNP tags. and These represent two individuals (e.g., an individual in the sample and a reference group) at the [missing information] th [missing information]. Genotype values ​​on SNP markers; IBS value represents the average genotype difference between two individuals across all SNP markers. The closer the IBS value is to 0, the more similar the two individuals are (high genotype similarity) and the closer their genetic distance; the larger the IBS value, the less similar the two individuals are (large genotype difference) and the farther their genetic distance.

[0041] Based on the obtained genetic distance data, the material genetic background analysis submodule can determine the heterotic group affiliation of the user-uploaded material data by comparing the genetic distance between each individual in the reference population and the user-uploaded material data. By comparing the genetic distance between the test material and a reference population of known taxa, the most similar heterotic group type can be found for the user-uploaded material data.

[0042] In some implementations, the materials genetic background analysis submodule can be used for: The genetic composition of the user material data was analyzed based on a reference population divided into multiple heterosis groups.

[0043] Specifically, the materials genetic background analysis submodule can also calculate the proportion of the genetic contribution of the user's material data to the reference population based on the population structure information of the reference population obtained in advance.

[0044] For example, the reference population comprises multiple subpopulations. The P-matrix of the reference population (K=7, representing 7 subpopulations) is pre-obtained using ADMIXTURE software. Then, the genetic contribution of the user material data to each subpopulation is calculated. This analysis can determine the proportional relationship of the genomic composition of the user material data originating from different reference subpopulations within the reference population, thus providing more detailed information about its genetic background.

[0045] In some implementations, the materials genetic background analysis submodule can be used for: Using a sliding window algorithm and a pre-set matching threshold, the source of the backbone of the improved user material data is identified based on the user material data of the improved system and the user material data of the backbone system uploaded by the user.

[0046] Specifically, the material genetic background analysis submodule can use a sliding window algorithm and a pre-set matching threshold to identify the source of the backbone of the user material data of the improved lines based on the user-uploaded user material data of the improved lines and the user material data of the backbone lines.

[0047] In the sliding window algorithm, the material genetic background analysis submodule can slide along the genome step by step with a step size of 1 SNP based on a fixed-size window, for example, a window size of 20 SNPs, and perform a systematic comparison of the genotypes in each window.

[0048] The formula for the sliding window method is as follows: In the formula, Display window They were determined to be of the same origin; , Let A and B represent the SNP genotype vectors, respectively. This indicates an indicator function that returns 1 if its internal condition is true, and 0 otherwise. This is the window size, which defaults to 20. This is the consistency threshold, which defaults to 19.

[0049] Within each sliding window, the materials genetic background analysis submodule compares the user material data of the improved line with the user material data of each backbone line locus by locus. By using a pre-set matching threshold, such as requiring at least 19 SNP loci within the window to have completely identical genotypes, it determines whether the window segment originates from a specific backbone line. When the matching degree between the user material data of the improved line and the user material data of a backbone line within a certain window segment reaches or exceeds the set threshold, the materials genetic background analysis submodule can determine that the genomic segment of the improved line originates from that backbone line.

[0050] In some implementations, the materials genetic background analysis submodule can be used for: By using a sliding window algorithm, a pre-set matching degree threshold, and a minimum recombinant fragment length threshold, genomic fragments from user-uploaded backbone lines in user-uploaded improved lines are identified.

[0051] Specifically, for user material data of improved lines and user material data of backbone lines uploaded by users, a whole genome window scan can be performed as described in the previous embodiments to determine whether the genome segments of improved lines originate from backbone lines.

[0052] The Materials Genetic Background Analysis submodule can merge consecutive windows that are determined to originate from the same lineage and are located adjacently.

[0053] After merging consecutive windows, the materials genetic background analysis submodule can filter the merged fragments based on a set minimum recombination fragment length threshold, for example, a default value of 3 Mb. The materials genetic background analysis submodule can then identify genomic fragments whose length reaches or exceeds this minimum recombination fragment length threshold as originating from the backbone lineage.

[0054] In some implementations, the materials genetic background analysis submodule can be used for: When a user uploads multiple sets of improved strain user material data, the conserved genomic fragments of the multiple sets of improved strain user material data are determined based on the conservation threshold.

[0055] Specifically, when a user uploads multiple sets of improved user material data, the material genetic background analysis submodule can further compare multiple sets of user material data originating from the same backbone line, based on the identification of the backbone line origin of each improved line.

[0056] By setting a conservation threshold, the material genetic background analysis submodule can identify conserved genomic segments in multiple sets of user material data that originate from a certain backbone line. These segments may contain key genetic components related to important agricultural traits.

[0057] In some implementations, the materials genetic background analysis submodule can be used for: The group of the user material data is determined to be either a paternal-biased group or a maternal-biased group.

[0058] Specifically, the Material Genetic Background Analysis submodule can be used to analyze the parental bias characteristics of user-provided materials. This submodule can determine whether the genetic background of the entire offspring population is more paternal or maternal by calculating the genetic relationships between the user-provided parents and the offspring population.

[0059] For example, user material data includes SNP microarray data for both parents and SNP microarray data for multiple double haploid (DH) lines. The material genetic background analysis submodule can calculate the genetic distance between each DH line and its parents to perform parental bias analysis. The genetic distance results are obtained by comparing the genotypic matching degree between each DH line and its father and mother at SNP loci.

[0060] Based on the calculated genetic distance results, the material genetic background analysis submodule can further analyze the genetic distance relationship between each DH line and its parents. By comparing the differences in distance between each DH line and its paternal and maternal parents, it can be determined which parent its genetic background is more closely related to.

[0061] After completing the individual-level analysis, the genetic background analysis submodule can determine the overall parental bias of the entire DH population by statistically analyzing the proportions of individuals biased towards the father and mother. This analysis can reflect the genetic structural characteristics of the DH population at the population level.

[0062] According to the crop breeding decision analysis system provided by the present invention, the result display module 130 is specifically used for: Based on the genetic distance between the user material data and the reference population, a phylogenetic tree merging the user material data and the reference population was constructed. Based on the genetic distance between the user's material data and the reference population, a list of a preset number of individuals in the reference population is generated, prioritizing those with the closest genetic distance. Generate a stacked histogram of population structure based on the genetic components of the user's material data.

[0063] Specifically, the results display module can construct a phylogenetic tree merging the user material data and the reference population based on the genetic distance between them. This step places the user material data and the reference population with known backgrounds in the same structure, visually displaying their kinship and clustering in the form of a tree diagram.

[0064] Based on the obtained genetic distance data, the results display module can further generate a list of a preset number of individuals in the reference population, prioritizing those with the closest genetic distance, according to the genetic distance between the user's material data and the reference population. This list directly identifies several known reference individuals in the reference population that are genetically closest to the user's material data, providing a basis for determining the population affiliation of the user's material data.

[0065] Furthermore, the results display module can generate a stacked bar chart of population structure based on the genetic composition of the user material data, i.e., the proportion of the user material data's genetic contribution to the reference population. This chart can clearly show the specific proportion of each reference subpopulation in the genomic composition of the user material data from the reference population in a stacked bar format, thus demonstrating its complex genetic background composition.

[0066] According to the crop breeding decision analysis system provided by the present invention, the breeding utilization guidance submodule can be used for: Based on the genotypes of known variety patterns, potential variety patterns in user material data are determined.

[0067] Specifically, the breeding guidance submodule can use genotype data from user-uploaded user materials and genotype data from known varietal patterns as the basis for analysis. These known varietal patterns typically include validated superior parent combinations and their genotype information.

[0068] The breeding guidance submodule can calculate the genetic distance between user material data and the parents of each known varietal model. By comparing the genotypic matching degree between user material data and each known parent at SNP loci, a quantitative genetic similarity index is obtained.

[0069] Based on the calculated genetic distance, the breeding guidance submodule can select parental materials that are highly similar to the user's material data from known varietal patterns according to a preset similarity threshold. For example, when the genetic distance between the user's material data and a known parent is below a certain threshold, the two can be considered to have high genetic similarity.

[0070] After identifying highly similar parental materials, the breeding utilization guidance submodule can replace the similar parent with user material data and pair it with another parent in the original variety combination to recommend new potential variety patterns. For example, when user material data M is found to be highly similar to parent A in the known variety pattern A×B, the breeding utilization guidance submodule can use M×B as a potential new variety pattern.

[0071] According to the crop breeding decision analysis system provided by the present invention, the breeding utilization guidance submodule can be used for: Determine the genotype of the simulated offspring of the user material data; Based on the known superior hybrid genotypes, the application potential of simulated offspring genotypes is evaluated.

[0072] Specifically, the breeding utilization guidance submodule can first determine the simulated progeny genotypes of the user's material data.

[0073] In some implementations, when a user provides a set of parental material data, the breeding utilization guidance submodule can simulate the genotypic composition of the F1 offspring that may be produced after the two parents are crossed, based on genetic principles.

[0074] Based on the obtained simulated progeny genotypes, the breeding utilization guidance submodule can further compare and analyze the simulated progeny genotypes with the known superior hybrid genotypes, thereby evaluating the application potential of the simulated progeny genotypes.

[0075] In some implementations, this assessment process can be achieved by calculating the genetic distance between the simulated offspring genotype and the genotype of a known superior hybrid. A quantified genetic similarity index is obtained by comparing the degree of genotypic matching at SNP loci.

[0076] Based on the calculated genetic distance results, the breeding guidance submodule can further analyze the similarity between simulated offspring and superior hybrids. Understandably, the closer the genetic distance between the simulated offspring and superior hybrids, the better the performance potential of the offspring produced by that combination.

[0077] It should be noted that the breeding utilization guidance submodule can combine the genotypes based on known variety patterns to determine the potential variety patterns of user material data, as well as the simulated progeny genotypes of user material data; based on the genotypes of known superior hybrids, it can evaluate the application potential of simulated progeny genotypes, thereby providing one-stop breeding guidance.

[0078] The following is an example of a specific application scenario to illustrate this, and the process of this example is as follows: 1. Built into the system (Z58) Chang7-2:ZD958, PH6WC PH4CV:XY335) are two known varietal models and hybrid genotypes.

[0079] 2. Users upload their own materials, and the genetic distance between them and all their parents (Z58, Chang7-2, PH6WC, PH4CV) is calculated. A set of materials with high similarity is then identified based on a threshold. The logical expression is as follows: In the formula: Indicates the genotype of the material uploaded by the user; The set of parental genotypes provided by the system; Parental materials in systems that represent high similarity to user materials; Indicates the genetic distance threshold; Represents the genetic distance function; This represents a set of parent materials that are highly similar to the user's materials.

[0080] 3. Based on the set of parents with high similarity to the user material obtained in the previous step, replace the user material with another parent to form a recommended set of variety patterns. The logical expression is as follows: In the formula: This refers to the variety model genotype database provided by the system; This represents a set of parent materials that are highly similar to the user's materials. Parental materials in systems that represent high similarity to user materials; express His original wife; Hybrids representing varietal patterns; This represents a set of recommendation patterns.

[0081] 4. Based on the obtained potential variety pattern combinations, simulate the genotypes of the F1 hybrids, calculate the genetic distance between the hybrids and the original variety patterns, and obtain the final recommended pattern set. The logical expression is as follows: In the formula: Represents a set of recommendation patterns; Represents a hybridization simulation function; This indicates the genetic distance between the simulated genotype and the actual genotype of the offspring. This represents the final set of recommended patterns.

[0082] According to the crop breeding decision analysis system provided by the present invention, the breeding utilization guidance submodule is specifically used for: Based on user material data of the target group uploaded by users, calculate the genetic distance matrix between any two individuals in the target group; Based on the genetic distance matrix between any two individuals and the core group ratio set by the user, the core group in the target group is selected.

[0083] Specifically, the breeding utilization guidance submodule can calculate the genetic distance matrix between any two individuals within the target population based on user-uploaded data of the target population. Through this calculation process, the genotypic differences of any two individuals within the target population across all SNP markers can be compared, constructing a symmetric matrix that comprehensively reflects the genetic diversity and inter-individual kinship within the target population, providing a quantitative data foundation for subsequent breeding utilization guidance.

[0084] After obtaining the genetic distance matrix, the breeding guidance submodule can call a specialized optimization algorithm, such as the Parallel Tempering Algorithm, which takes the genetic distance matrix as input and the core population ratio set by the user (e.g., setting the core population ratio to be 30% of the target population), to automatically select a core subset from the original population that can best represent the genetic diversity of the original population.

[0085] It should be noted that the embodiments of the present invention do not limit which optimization algorithm the breeding utilization guidance submodule calls. The parallel annealing algorithm is just an example given in the embodiments of the present invention. The optimization algorithm called by the breeding utilization guidance submodule can be selected according to the actual situation.

[0086] According to the crop breeding decision analysis system provided by the present invention, the whole-genome selection submodule is specifically used for one or more of the following: Based on the user-uploaded material data and the corresponding phenotypic data, the prediction model selected by the user is run to obtain the genomic breeding value of each sample in the user material data.

[0087] Specifically, the whole-genome selection submodule can establish the association between genotype and phenotype based on the user-uploaded user material data and the corresponding phenotypic data, train statistical models and / or machine learning models, and thus predict the genomic breeding value of each sample in the user material data through the trained models.

[0088] In this embodiment of the invention, the whole-genome selection submodule can train a variety of different prediction models, such as statistical models like rrBLUP and BayesC, and machine learning models like Support Vector Regression (SVR) and LightGBM. Users can choose from a variety of different prediction models to meet the prediction needs of different types of traits.

[0089] The whole-genome selection submodule can run the prediction model selected by the user and return the genomic breeding value for each sample in the user's material data.

[0090] In some implementations, the whole-genome selection submodule can also return the prediction accuracy assessment results corresponding to the genomic breeding values ​​of each sample in the user's material data, so as to facilitate the user's evaluation of the prediction results.

[0091] According to the crop breeding decision analysis system provided by the present invention, the whole-genome selection submodule is further used for: The hyperparameters of the prediction model are tuned using a grid search method. The prediction model is evaluated using cross-validation.

[0092] Specifically, during the training of the prediction model, the whole-genome selection submodule can fine-tune the hyperparameters of the prediction model through grid search. Through this process, the whole-genome selection submodule can automatically traverse preset hyperparameter combinations, train a model for each combination, and find the optimal hyperparameter configuration through performance evaluation, thereby improving the model's prediction accuracy.

[0093] After hyperparameter tuning, the whole-genome selection submodule can evaluate the prediction model using cross-validation. This process involves dividing the dataset into multiple subsets, using a portion as the validation set and the remainder as the training set in turn, training and validating the model multiple times to obtain a robust assessment of the model's generalization ability.

[0094] By using this combination, the whole-genome selection submodule can ensure that the final prediction model has both good parameter configuration and reliable generalization ability.

[0095] According to the crop breeding decision analysis system provided by the present invention, the crop breeding decision analysis system further includes: Knowledge Q&A module; The knowledge-based question-and-answer module includes a large language model and a crop breeding knowledge base, specifically used for: Transform the user's input question into a question vector; Based on the question vector, a matching process is performed in the crop breeding knowledge base to obtain the multiple text fragments with the highest similarity. Multiple text fragments are input into a large language model to obtain the answer to the input question and provide feedback to the user.

[0096] Specifically, the crop breeding decision analysis system provided by the present invention may also include a knowledge question-and-answer module, which provides users with professional and accurate breeding-related knowledge services by combining a large language model with a crop breeding knowledge base with professional domain knowledge.

[0097] The knowledge-based question-and-answer module can include two core components: a large-scale language model and a crop breeding knowledge base.

[0098] In some implementations, large-scale language models can employ advanced natural language processing models such as Qwen3:14b, while crop breeding knowledge bases consist of a large number of organized breeding professional knowledge documents and data.

[0099] First, the knowledge-based question-answering module can transform the user's input question into a question vector. This transformation process converts natural language questions into a numerical vector representation that can be understood by machines.

[0100] Subsequently, the knowledge question answering module can perform similarity matching in the crop breeding knowledge base based on the generated question vectors.

[0101] In some implementations, the similarity matching rule is a hybrid retrieval. Specifically, firstly, keyword retrieval is performed. Based on the generated question vector, keywords of the user's question are fuzzily matched in the crop breeding knowledge base, and multiple candidate results with the highest matching degree are returned, for example, the top 10 keyword candidate results with the highest matching degree are returned. Secondly, semantic retrieval is performed. By calculating the similarity between the question vector and the text fragment vectors in the knowledge base, multiple candidate results with the highest relevance to the user's question are returned, for example, the top 10 text fragment candidate results with the highest similarity are returned. Finally, the two retrieval results are merged and duplicates are removed. The relevance between the duplicate-removed candidate results and the question vector is calculated. The multiple candidate results with the highest scores are selected according to the score ranking, for example, the top 5 candidate results with the highest scores are used as the final retrieval results.

[0102] It should be noted that the similarity matching rules can also use keyword retrieval or semantic retrieval alone. In this embodiment of the invention, the specific form of the similarity matching rules is not limited.

[0103] Finally, the knowledge-based question-answering module can input multiple relevant text fragments retrieved into a large language model. Based on these professional knowledge fragments and its own language comprehension capabilities, the large language model can generate accurate and logically clear answers and provide them to the user, thus ensuring the professionalism and accuracy of the responses.

[0104] As shown in Table 1, the knowledge-based question-and-answer module combined with the crop breeding knowledge base provides more professional and accurate answers than ordinary large-scale language models.

[0105] Table 1. Examples of question-and-answer formats for different models of the same problem.

[0106] According to the crop breeding decision analysis system provided by the present invention, the crop breeding knowledge base is constructed in the following manner: We will disassemble, clean, and slice professional articles related to crop breeding collected from multiple sources; A crop breeding knowledge base is constructed based on the embedded model and sliced ​​text data.

[0107] Specifically, in the process of constructing a crop breeding knowledge base, professional articles related to crop breeding can be collected from multiple sources. These sources can include Chinese literature, English literature, and patent literature, etc., and high-quality professional content can be obtained by setting corresponding search strategies and screening criteria.

[0108] After obtaining the original documents, the collected full texts can be broken down and processed.

[0109] In some implementations, specialized text recognition tools can be used to convert documents in formats such as PDF into a processable text format, such as Markdown, in order to prepare for subsequent text processing.

[0110] After the full text is decomposed, text cleaning and slicing can be performed. Text cleaning removes irrelevant formatting symbols and noisy data. The cleaned text is then sliced ​​into segments of a preset size, such as approximately 350 tokens, to maintain information integrity and retrieval accuracy.

[0111] Based on the processed text data, a crop breeding knowledge base can be further constructed.

[0112] In some implementations, an open-source embedding model can be used to vectorize the sliced ​​text, converting the text into numerical vectors, and then using a dedicated vector database framework to store and manage this vector data, forming a crop breeding knowledge base that can be efficiently retrieved. The following examples from specific application scenarios further illustrate the crop breeding decision analysis system provided by this invention.

[0113] The specific implementation of this embodiment is an online analysis platform called Maize Breeding Decision Platform (MBDP), and its technical solution is described below in conjunction with the accompanying drawings.

[0114] System Architecture: The platform adopts a front-end and back-end separation architecture. The back-end is based on a Linux + Nginx + MySQL + Python (LNMP) technology stack. The Django framework in Python handles business logic, and the Django REST framework is used to build APIs for communication with the front-end. A MySQL database is used for data storage. The front-end uses the Vue.js framework to build the user interface, ElementPlus as the component library, and Axios handles HTTP communication with the back-end.

[0115] Overall analysis process: Figure 2 This is a schematic diagram of the overall analysis process of MBDP provided by the present invention, as shown below. Figure 2 As shown, the overall process for users to use this platform is as follows: Step 1: Data Preparation and Upload. Users organize the 20K SNP chip data into the format required by the platform. The platform is compatible with various data inputs and provides a genotype filling tool (based on LB-Impute) for converting 3K chips to 20K chips, as well as a conversion tool from VCF format to the platform format.

[0116] Step 2: Data Preprocessing. After users upload their data, the system automatically performs data quality control, including removing samples with a heterozygosity >5% and discarding markers with a heterozygosity >5% or a deletion rate >20%. Subsequently, the major allele imputation method is used to impute missing markers to ensure data integrity.

[0117] Step 3: Functional Module Selection and Analysis. Users select the appropriate analysis module and set the parameters according to their breeding needs.

[0118] Step 4: Viewing and Downloading Results. After completing the analysis, the platform displays the results in the form of visual charts and data tables, which users can view or download online.

[0119] At the same time, users can ask questions on the platform and receive answers, which can be used in conjunction with the platform's analysis results to guide the breeding decision-making process.

[0120] The following are Figure 2 The various modules within will be described in detail.

[0121] I. Data Preprocessing The core data is in 20K chip format.

[0122] (1) Data padding: If a 3K chip is used, it can be padded (LB-Impute) to the 20K level based on the parent 20K.

[0123] (2) Format conversion: If the 20K sites obtained by resequencing are in VCF format, they can be converted to the chip format supported by this platform.

[0124] (3) Preprocessing: After the user uploads the data, the system automatically performs data quality control, including: removing samples with a heterozygosity of >5% and removing markers with a heterozygosity of >5% and a missing rate of >20%.

[0125] II. Analysis of the genetic background of the materials Function: Analyzes the genetic background of user materials from multiple dimensions.

[0126] (1) Hybrid group: Determine the type of maize hybrid group to which the user's material belongs.

[0127] Implementation method: 1. The system has a built-in reference population that can be clearly divided into 7 heterosis groups (BSSS, Flint, Iodent, Lancaster, Lvda, Reid, TSPT). 2. Users upload 20K chip data of the material to be analyzed. 3. Combine user data and reference population data, and calculate genetic distance using the IBS method. The formula is as follows: In the formula, This represents the total number of SNP markers, i.e., the number of SNP sites on the genome; It is an index of SNP tags. and These represent two individuals (e.g., an individual in the sample and a reference group) at the [missing information] th [missing information]. Genotype values ​​on SNP markers; IBS value represents the average genotype difference between two individuals across all SNP markers. The closer the IBS value is to 0, the more similar the two individuals are (high genotype similarity) and the closer their genetic distance; the larger the IBS value, the less similar the two individuals are (large genotype difference) and the farther their genetic distance.

[0128] 4. Based on the calculated genetic distance matrix, the UPGMA clustering function is constructed using the R package ggtree. Simultaneously, information on each user's material and the five genetically closest materials in the system is output to help users determine the heterosis group of their materials. (2) Population structure: Analyze the genetic components of user materials.

[0129] Implementation method: 1. The system has a built-in reference population (same as the heterosis group), and the P matrix is ​​calculated in advance using ADMIXTURE software when K is 7.

[0130] 2. Users upload 20K chip data of the material to be analyzed.

[0131] 3. Use ADMIXTURE software to calculate the proportion of the genetic contribution of user samples to each subpopulation and generate a stacked bar chart of the population structure.

[0132] (3) Backbone line genome segment analysis: Determine which genome segments in the user's improved line originate from the backbone line.

[0133] Implementation method: 1. Users upload a backbone genotype and an improved genotype (or multiple improved lines).

[0134] 2. The sliding window method is used for analysis, and the formula is as follows: In the formula, Display window They were determined to be of the same origin; , Let A and B represent the SNP genotype vectors, respectively. This indicates an indicator function that returns 1 if its internal condition is true, and 0 otherwise. This is the window size, which defaults to 20. This is the consistency threshold, which defaults to 19.

[0135] Logic: If the genotype matching degree between the improved line and the backbone line reaches a threshold (e.g., matching 19 or more SNPs) within the window, then the window segment is considered to originate from the backbone line.

[0136] 3. Merge consecutive windows that meet the requirements and set a minimum recombinant fragment length threshold (default is 3 Mb) to finally identify the genomic fragments from the backbone line in the improved line (Note: If multiple improved lines are uploaded at the beginning, a conservation threshold can be set to identify conserved genomic segments).

[0137] (4) Parental inference: Determine the true kinship between the user's parents and offspring materials.

[0138] Implementation method: 1. Users upload the genotypes of their parents and the genotypes of their offspring to be tested.

[0139] 2. Using the sliding window method (same as above) and setting a threshold, determine which offspring are the true offspring of the parents.

[0140] (5) Parent bias: Determine whether the user's DH series material is biased towards the parent or the mother.

[0141] Implementation method: 1. Users upload the genotypes of their parents and multiple DH lineages.

[0142] 2. Calculate the genetic distance between each DH and its parents (same as IBS above) to determine the genetic distance between each DH and its parents, and thus determine whether the entire DH population is more paternal or maternal.

[0143] III. Breeding Utilization Guidance Function: Provides decision-making services for multiple breeding scenarios. (1) Variety pattern recommendation: Based on the materials provided by the user and the genotype of the variety pattern, infer the potential variety pattern.

[0144] Implementation method: 1. Users upload the genotypes of their own materials as well as the genotypes of known variety patterns.

[0145] 2. Calculate the genetic distance between the user material and each known varietal model parent, and find highly similar materials based on the threshold.

[0146] 3. Subsequently, by using the user's material replacement and its highly similar material, and another parent of the combination recommended as a potential variety pattern.

[0147] Example: Suppose a user uploads a set of known varietal patterns A×B. By calculating the genetic distance between M and A and B, it is found that the genetic distance (IBS) between M and A is only 0.2, that is, their similarity is as high as 80%. Therefore, it is considered that M can replace A and B as potential varietal patterns, i.e., M×B.

[0148] (2) Hybrid simulation and evaluation: Simulate the genotype of the offspring of the parent materials provided by the user and compare it with the genotype of the superior hybrid to evaluate the application potential of the hybrid.

[0149] Implementation method: 1. Users upload a set of their parents' genotypes (who they believe are a good combination), as well as the genotypes of known superior hybrids.

[0150] 2. Simulate the genotypes of the offspring F1 based on the genotypes of the parents, and then calculate the genetic distance (IBS) between the offspring and the hybrid. If the genetic distance is close, it means that the similarity is high, indicating that the offspring of this combination may have better performance.

[0151] Example: Suppose a user's parents are A and B, and they upload a superior hybrid H. First, simulate the genotypes of the offspring of A×B, and then calculate the genetic distance between the offspring and H to evaluate the application potential of the hybrid.

[0152] (3) One-stop breeding guidance: This module combines the functions of (1) and (2). Implementation method: 1. Built into the system (Z58) Chang7-2:ZD958, PH6WC PH4CV:XY335) are two known varietal models and hybrid genotypes.

[0153] 2. Users upload their own materials, and the genetic distance between them and all their parents (Z58, Chang7-2, PH6WC, PH4CV) is calculated. A set of materials with high similarity is then identified based on a threshold. The logical expression is as follows: In the formula: Indicates the genotype of the material uploaded by the user; The set of parental genotypes provided by the system; Parental materials in systems that represent high similarity to user materials; Indicates the genetic distance threshold; Represents the genetic distance function; This represents a set of parent materials that are highly similar to the user's materials.

[0154] 3. Based on the set of parents with high similarity to the user material obtained in the previous step, replace the user material with another parent to form a recommended set of variety patterns. The logical expression is as follows: In the formula: This refers to the variety model genotype database provided by the system; This represents a set of parent materials that are highly similar to the user's materials. Parental materials in systems that represent high similarity to user materials; express His original wife; Hybrids representing varietal patterns; This represents a set of recommendation patterns.

[0155] 4. Based on the obtained potential variety pattern combinations, simulate the genotypes of the F1 hybrids, calculate the genetic distance between the hybrids and the original variety patterns, and obtain the final recommended pattern set. The logical expression is as follows: In the formula: Represents a set of recommendation patterns; Represents a hybridization simulation function; This indicates the genetic distance between the simulated genotype and the actual genotype of the offspring. This represents the final set of recommended patterns.

[0156] (4) Core DH selection: Users submit DH group data of the same batch, select how much material to retain, and then perform redundancy removal.

[0157] Implementation method: 1. Users upload a batch of DH line genotype data, then calculate the IBS between each pair to construct a genetic distance matrix.

[0158] 2. Using the CoreHunter3 tool, select the core DH lineage based on the percentage of the retained population set by the user (e.g., 30%) using the parallel annealing algorithm.

[0159] IV. Genome-wide selection Function: Used for phenotypic prediction.

[0160] It provides linear and machine learning models, such as Ridge Regression BLUP (rrBLUP) model, BayesC model, Support Vector Regression (SVR) model, and Light Gradient Boosting Machine (LightGBM) model.

[0161] Example / Effect: Figure 3 This is a box plot showing the accuracy distribution of genome prediction based on flowering period phenotype provided by this invention, such as... Figure 3 As shown, the test data used consisted of 20K microarray data from 240 natural maize populations and phenotypic data from two days to anthesis (DTA) flowering periods over two years. Genotypic data were used to predict phenotypes, and the Pearson correlation coefficient between predicted and actual values ​​was calculated as an accuracy indicator. Each box plot represents the distribution of prediction accuracy under a given environment.

[0162] V. Maize Breeding Knowledge Base Technology: Retrieval-Augmented Generation (RAG) is an artificial intelligence technique that combines information retrieval technology with language generation models. This technique retrieves relevant information from external knowledge bases and uses it as prompts to feed into Large Language Models (LLMs), thereby enhancing the model's ability to handle knowledge-intensive tasks such as question answering, text summarization, and content generation.

[0163] Purpose and Significance: By collecting published high-quality maize breeding-related patents, Chinese literature, and English literature, a professional-grade knowledge base for maize breeding will be constructed. Based on the natural language processing capabilities of a large model, it will provide users with accurate retrieval and answers to breeding-related knowledge. In addition, users can ask questions about the knowledge base at any time when they encounter problems or have needs in the previous analysis modules.

[0164] The knowledge base construction process includes the following steps: 1. Literature and patent collection: (1) Collection of Chinese literature: Using “maize breeding” as the keyword, we searched for Chinese journal articles on CNKI and only retained articles from journals included in the “Peking University Core Journals”.

[0165] (2) English literature collection: Using “maize breeding” as the keyword, we searched for SCI journals in Web of Science, with the time range set from 2015 to 2025. We also manually selected journals to retain papers from high-impact journals.

[0166] (3) Patent collection: Search the State Intellectual Property Office using keywords such as “maize variety”, “maize gene” or “maize locus”.

[0167] 2. Full text breakdown The MinerU tool was used to perform text recognition on the entire text and output it in Markdown format.

[0168] 3. Text Slicing The entire text is cleaned and sliced, with each slice being 350 tokens in size. Two adjacent slices overlap by 70 tokens.

[0169] 4. Construction of a maize breeding knowledge base The above-mentioned sliced ​​text is used to construct maize breeding knowledge based on the open-source embedding model "nomic-embed-text" and the open-source Milvus vector database framework.

[0170] Knowledge Q&A Process: Large model: Qwen3:14b.

[0171] Process: The user inputs a question, which is then converted into a vector. This vector is matched against a knowledge base, and the five most similar fragments are output. These fragments are then fed into a large model, which uses prompts to refine the text into a precise and logically clear answer, which is then returned to the user.

[0172] Figure 4 This is a comparison chart of the accuracy of the model provided in this embodiment of the invention with other publicly available models, such as... Figure 4 As shown, the breeding knowledge base in the maize breeding decision platform MBDP provided in this embodiment corresponds to qwen3-14b-RAG, and its accuracy rate in answering search questions reaches 87%, which is significantly better than Kimi2-1024b, Qwen3-256b and DeepSeek-V3-671b.

[0173] Compared with the prior art, the technical solution of this embodiment has the following significant technical advantages: 1. Significantly lowers the technical barrier and improves the utilization rate of genomic data: This platform encapsulates complex bioinformatics analysis processes into simple web-based operations, allowing breeders to complete the entire process from data upload to result interpretation without programming. This greatly promotes the popularization and application of SNP chip technology in grassroots breeding units, freeing breeders from tedious data analysis and allowing them to focus on breeding decisions.

[0174] 2. High degree of functional integration, providing one-stop breeding decision support: The platform integrates multiple key modules such as phylogenetic analysis, parent identification, core population screening, and genetic testing (GS), covering the main decision-making stages in the breeding process. Breeders can complete a systematic evaluation of breeding materials on a single platform, solving the problems of limited functionality and fragmented operations of existing tools, and improving work efficiency.

[0175] 3. Improve breeding efficiency and shorten the breeding cycle: Through the online GS module, breeders can quickly predict the genetic potential of a large number of candidate materials, enabling early and efficient selection. The core DH population screening function can effectively reduce the scale and cost of field trials. The parent inference function can avoid breeding errors caused by material mixing. These functions work together to significantly accelerate the breeding process.

[0176] 4. Intelligent and Highly Accurate: The platform incorporates advanced machine learning algorithms such as LightGBM and improves the predictive accuracy of the GS model through automatic hyperparameter tuning. Phylogenetic analysis based on a built-in, rigorously defined reference population ensures the reliability of heterosis group identification, providing breeders with a more scientific basis for developing hybridization strategies.

[0177] The platform architecture and core concept of this embodiment have good scalability, and the following alternative or extended solutions also exist: 1. Expanded Applicability to Multiple Species: While this platform currently focuses on maize, its underlying architecture and analysis modules (such as genetic distance calculation and GS models) are universal. By replacing the species-specific reference panel and genomic information, the platform can be quickly expanded to other major crops such as rice, wheat, and soybeans, serving a wider range of breeding needs.

[0178] 2. Integrating More Advanced AI Algorithms: The GS module can further integrate deep learning models (such as Convolutional Neural Networks, CNNs) to explore non-additive effects and uncover more complex genotype-phenotype relationships, potentially leading to higher prediction accuracy. Furthermore, the integration of Large Language Models (LLMs) into the platform can be explored to provide intelligent interpretation of results and breeding suggestions, enabling interactive breeding decision-making.

[0179] 3. Compatibility with more data types: In addition to SNP microarray data, the platform can be upgraded in the future to be compatible with higher-density microarray data or low-depth sequencing data. When the data density is high enough, alternative analysis algorithms can also be added, such as introducing a more accurate IBD (Identity By Descent) algorithm to replace the sliding window algorithm in improved line source tracing analysis.

[0180] 4. Add new functional modules: To address other pain points in the breeding process, new functional modules can be developed, such as optimized design of breeding population size, comprehensive analysis of multi-year and multi-location experimental data, and marker-assisted selection modules for specific functional genes / QTLs, making the platform more complete.

[0181] It should be noted that the system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.

[0182] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a read-only memory (ROM), a random access memory (RAM), a magnetic disk, an optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.

[0183] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A crop breeding decision analysis system, characterized in that, include: The system includes a data input module, a data processing module, a data analysis module, and a results display module. The data input module is used to receive user material data uploaded by the user; the user material data includes single nucleotide polymorphism (SNP) chip data; The data processing module is connected to the data input module and is used to preprocess the user material data; The data analysis module connects the data input module and the data processing module, and is used to analyze the preprocessed user material data based on the sub-modules selected by the user and / or the parameters set by the user, and obtain the analysis results. The results display module is connected to the data analysis module and is used to generate visual charts and / or data tables based on the analysis results and display them to the user. The data analysis module includes multiple sub-modules, including at least one of the following: a material genetic background analysis sub-module, a breeding utilization guidance sub-module, and a whole-genome selection sub-module. The material genetic background analysis submodule is used to analyze the genetic background of the user's material data; The breeding utilization guidance submodule is used to make breeding decisions based on the user material data; The whole-genome selection submodule is used to perform phenotypic prediction on the user material data.

2. The crop breeding decision analysis system according to claim 1, characterized in that, The material genetic background analysis submodule is specifically used for one or more of the following: Based on a reference population divided into multiple heterosis groups, determine the type of heterosis group to which the user material data belongs; Based on a reference population divided into multiple heterosis groups, the genetic composition of the user material data was analyzed; Using a sliding window algorithm and a pre-set matching threshold, the source of the backbone system of the user material data of the improved system is identified based on the user material data of the improved system and the user material data of the backbone system uploaded by the user. The genomic fragments from the user-uploaded backbone lines in the user-uploaded improved lines are identified by using a sliding window algorithm, a pre-set matching degree threshold, and a minimum recombinant fragment length threshold. When a user uploads multiple sets of user material data for improved lines, the conserved genomic fragments of the multiple sets of user material data for improved lines are determined based on the conservation threshold. The group of the user material data is determined to be either a paternal-biased group or a maternal-biased group.

3. The crop breeding decision analysis system according to claim 2, characterized in that, The results display module is specifically used for: Based on the genetic distance between the user material data and the reference population, a phylogenetic tree merging the user material data and the reference population is constructed. Based on the genetic distance between the user material data and the reference population, a list of a preset number of individuals in the reference population is generated, prioritizing those with the closest genetic distance. Based on the genetic components of the user material data, a stacked histogram of the population structure is generated.

4. The crop breeding decision analysis system according to claim 1, characterized in that, The breeding utilization guidance submodule is specifically used for one or more of the following: Based on the genotypes of known variety patterns, determine the potential variety patterns of the user material data; Determine the simulated offspring genotypes of the user material data; The application potential of the simulated offspring genotypes is evaluated based on the known superior hybrid genotypes.

5. The crop breeding decision analysis system according to claim 1, characterized in that, The breeding utilization guidance submodule is specifically used for: Based on user material data of the target group uploaded by users, calculate the genetic distance matrix between any two individuals in the target group; Based on the genetic distance matrix between any two individuals and the core group ratio set by the user, the core group in the target group is selected.

6. The crop breeding decision analysis system according to claim 1, characterized in that, The whole-genome selection submodule is specifically used for one or more of the following: Based on the user-uploaded user material data and the corresponding phenotypic data, the prediction model selected by the user is run to obtain the genomic breeding value of each sample in the user material data.

7. The crop breeding decision analysis system according to claim 6, characterized in that, The whole-genome selection submodule is also used for: The hyperparameters of the prediction model are tuned using a grid search method. The prediction model is evaluated using cross-validation.

8. The crop breeding decision analysis system according to any one of claims 1 to 7, characterized in that, The crop breeding decision analysis system also includes: Knowledge Q&A module; The knowledge-based question-answering module includes a large-scale language model and a crop breeding knowledge base, specifically used for: Transform the user's input question into a question vector; Based on the question vector, a matching process is performed in the crop breeding knowledge base to obtain the multiple text fragments with the highest similarity. The multiple text fragments are input into the large language model to obtain the answer to the input question and provide feedback to the user.

9. The crop breeding decision analysis system according to claim 8, characterized in that, The crop breeding knowledge base is constructed in the following manner: We will disassemble, clean, and slice professional articles related to crop breeding collected from multiple sources; The crop breeding knowledge base is constructed based on the embedded model and the sliced ​​text data.

10. The crop breeding decision analysis system according to any one of claims 1 to 7, characterized in that, The data processing module is specifically used for: The user material data is filtered based on the heterozygosity of each sample in the user material data, and the heterozygosity and / or missing rate of each marker in the user material data; Based on the major allele imputation method, missing markers in the screened user material data are imputed.