Rice simulation breeding platform capable of predicting breeding
By developing a rice simulation breeding platform with built-in multiple genome selection algorithms and rich data management functions, the problem of insufficient support for rice on the existing breeding platform is solved, and efficient, accurate and diverse applications of breeding work are achieved.
Patent Information
- Application Number
- CN202510223803.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-27
AI Technical Summary
The existing breeding platforms lack sufficient support for rice, which has problems such as strong enclosure, limited algorithm types, incomplete data management, unfriendly operational functions, poor visualization of result reports, incomplete rice data resources and limited data analysis capacity, which seriously limits the efficiency and accuracy of breeding work.
A predictable breeding rice simulation breeding platform has been developed, including Home, Phenotype, GS and Help interfaces, with up to 8 genome selection algorithms built-in, supporting multi-environment data fusion and rich variation data management, providing task management systems and efficient result report generation.
The platform is highly open and has a variety of algorithms built into it, which greatly broadens the breeding application scenarios, is simple and efficient in data management, is easy and friendly in operation, has excellent visualization effect on the result report, and is built-in task management system, which supports rapid breeding and genome selection of various economic animals, improving the efficiency and accuracy of breeding.
Smart Images

Figure CN120215939A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of genomic breeding, and specifically to a rice simulation breeding platform capable of predictable breeding selection. Background Art
[0002] In the history of selective breeding, there has been an exploration from empirical breeding to breeding theories and methods, including the selection theory, pure line theory, backcross breeding, recurrent breeding, mutagenic breeding, single-seed descent, ideal plant type; and then to molecular marker-assisted selection breeding, where various molecular markers have been explored, such as amplified fragment length polymorphism marker-assisted selection (AFLP), microsatellite marker-assisted selection (SSR), and single nucleotide polymorphism marker-assisted selection (SNP); with the development of sequencing technology, the sequencing throughput has become higher and the cost has become lower, and coupled with the continuous improvement of computer computing power, this has created technical conditions for the development of new breeding technologies, giving rise to the wave of genomic selection (GS) breeding. Genomic selection breeding can effectively solve the limitations of difficult-to-measure traits, high luck factor, long time consumption, high technical difficulty, etc., and accelerate the breeding pace. Genomic selection breeding is a breeding method of marker-assisted selection using high-density molecular genetic markers covering the entire genome. With the concept of "intelligent breeding" proposed, intelligent breeding has attracted more and more attention from breeders and is gradually moving towards application. Intelligent breeding is a new opportunity brought about by the development of science and technology. It is expected that in the next 10 - 20 years, intelligent breeding will become the core driving force for the development of the seed industry. From traditional breeding to molecular breeding and then to intelligent breeding, the "scientific" component content in breeding is increasing, while the "artistic" component content is decreasing. The "scientific" component mainly comes from the construction and upgrading of the intelligent breeding platform, establishing a crop breeding data extraction, mining, storage, analysis, and sharing database based on germplasm resource information, integrating multi-omics data for joint analysis, breaking through underlying support technologies such as the acquisition, analysis, and mining of biological big data, establishing an "data - technology - algorithm - decision" integrated intelligent breeding strategy, improving breeding accuracy, shortening the breeding cycle, and establishing an efficient crop intelligent breeding system.
[0003] Currently, there is no dedicated breeding platform for rice, and other existing breeding platforms such as CropGS-HUB and AI-breeder are also unavailable. These breeding platforms have the following deficiencies: 1. They are highly closed and lack openness to the outside world, making them unavailable for other users; 2. The types of built-in algorithms are few, restricting the flexibility of breeding selection; 3. Data management is not comprehensive enough to effectively integrate multi-source data; 4. The operation functions of the platform are not user-friendly, resulting in a poor user experience; 5. The visualization effect of the result report is poor, lacking in operability and aesthetics, and is not convenient for data interpretation and decision-making; 6. The rice data resources are not yet perfect, and the data coverage is limited; 7. The platform has limited capacity for data analysis and only supports the processing of small-scale data sets. The above deficiencies seriously limit the efficiency and accuracy of breeding work. Summary of the Invention
[0004] The purpose of the present invention is to provide a rice simulation breeding platform for predictable breeding to solve the technical problems mentioned in the above background technology.
[0005] To achieve the above purpose, the present invention provides the following technical solutions: A rice simulation breeding platform for predictable breeding, including the following steps:
[0006] S1. Develop the Home interface: Display data classification, basic overview, source distribution, and information of cooperation units to provide an overview for users.
[0007] S2. Develop the Phenotype interface: A phenotypic data management system and an integrated multi-environment phenotypic fusion algorithm platform.
[0008] S3. Develop the GS interface: A breeding platform that provides introductions, training platforms, and prediction platforms for 8 genomic selection algorithms, including the self-developed algorithm LGBMY and other algorithms GBLUP, LGBM, XGB, HGB, SVR, CATB.
[0009] S4. Develop the Help interface: Provide an operation manual for the Rice3kGS platform to help users quickly familiarize themselves with the platform functions and operation processes and achieve efficient use.
[0010] Preferably, the platform adopts an independently designed interface, including Home, Phenotype, GS, and Help interfaces. The layout of each interface is intuitive and simple, the operation is convenient, the task management is efficient, and the result report is clear and easy to understand.
[0011] Preferably, the platform is rich in phenotypic data resources provided by multiple cooperation units. The platform can quickly retrieve all sample data of target traits by constructing a database and supports statistical analysis and visual display.
[0012] Preferably, the database includes rich mutation data, covering a variety of mutation types, including SNP mutations within about 1 Mb, conservative SNP locus mutations of about 404 Kb, and large fragment mutations such as Deletion, Insertion, Inversion, and Duplication mutations. In addition, it also includes various types, namely STR mutations and gene deletion / insertion gene-PAV mutations, providing comprehensive mutation data support for breeding analysis.
[0013] Preferably, the platform integrates multi-environment algorithms, which can quickly and conveniently fuse and normalize multi-environment data, supporting three algorithm selections: BLUP, BLUE, and AVERAGE.
[0014] Preferably, the platform integrates 8 genome-wide selection breeding algorithms. Users can select one or more algorithms to train phenotypic and genotypic data. During the training process, 10-fold cross-validation is adopted, and the model is evaluated through four indicators: PCC, RMSE, MSE, and MAE.
[0015] Preferably, the platform integrates the self-developed LGBMY algorithm. By optimizing the parameters of the LGBM algorithm and selecting the best parameters to improve the model accuracy, it can improve the precision of breeding selection.
[0016] Preferably, the number of individuals for which the platform in S3 predicts the offspring to be used for germplasm conservation or breeding is default set to the top 100 samples, or the user can also select the top n number of germplasm conservation samples by themselves.
[0017] Preferably, the platform in S2 and S3 provides task process viewing, result viewing, and downloading.
[0018] Preferably, the application scope of the platform includes plants, animals, and their mutation types. The plants include rice, corn, wheat, and cotton, and the animals include pigs, cows, sheep, chickens, ducks, and geese.
[0019] Compared with the prior art, the beneficial effects of the present invention are:
[0020] 1. The simulation breeding platform of the present invention has strong openness. After the user applies for an account and is approved, all functions can be used, which greatly facilitates the breeding work. It has up to 8 algorithms built-in, covering traditional models and various machine learning algorithms, greatly expanding the breeding application scenarios and meeting different breeding needs.
[0021] 2. The data management of the simulation breeding platform of the present invention is simple and efficient. It comprehensively integrates the rice phenotypic database and supports fuzzy retrieval, enabling users to quickly find the required phenotypic data. At the same time, the platform also provides data statistics and visualization functions, greatly improving the efficiency of data processing and making the breeding work more efficient and accurate.
[0022] 3. The simulation breeding platform of the present invention is easy and friendly to operate, significantly improving the user experience. Whether it is data import, model training or result report generation, it can be easily completed. The generated result report has excellent visualization effects, is convenient to operate and beautiful, facilitating data interpretation and scientific decision-making.
[0023] 4. The simulation breeding platform of the present invention is built-in with a task management system, which can pause, restart and other operations on analysis tasks, facilitating efficient management of task progress. This function makes the breeding work more flexible and controllable, helping to improve breeding efficiency and quality.
[0024] 5. The simulation breeding platform of the present invention also supports selecting and training multiple GS model algorithms through phenotypic and genotypic data, screening out the optimal model for accurate prediction of offspring data. This function helps to achieve rapid breeding, shorten the new variety cultivation cycle, and provide a rapid selection platform for breeders.
[0025] 6. The simulation breeding platform of the present invention is applicable not only to major plants such as rice, but also to various economic animals such as pigs, cows and sheep. It can process various types of variant data, including SNPs, InDels generated by resequencing or chips, etc., improving the accuracy and universality of prediction while meeting different experimental requirements, and contributing to promoting the innovation and development of breeding work. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a schematic diagram of the operation process of the rice simulation breeding platform that can be predictively selected and bred according to the present invention;
[0027] Figure 2 It is a schematic diagram of the result fused by the AVERGE algorithm based on the thousand-grain weight phenotypic data for three years according to the present invention;
[0028] Figure 3 It is a schematic diagram of the result fused by the BLUP algorithm based on the thousand-grain weight phenotypic data for three years according to the present invention;
[0029] Figure 4 It is a schematic diagram of the result fused by the BLUE algorithm based on the thousand-grain weight phenotypic data for three years according to the present invention;
[0030] Figure 5 It is a schematic diagram of the result of the fuzzy retrieval related to grain phenotypes according to the present invention;
[0031] Figure 6 It is a schematic diagram of the results of GS models trained respectively based on multiple algorithms according to the present invention;
[0032] Figure 7 It is a schematic diagram of the results of the breeding values of candidate samples predicted respectively based on multiple algorithms according to the present invention. Detailed implementation manners
[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0034] Please refer to Figures 1-7 , the present invention provides a technical solution:
[0035] A rice simulation breeding platform for predictable breeding, including the following steps:
[0036] S1. Develop a Home interface: display data classification, basic overview, source distribution, and information of cooperation units, providing an overview for users;
[0037] S2. Develop a Phenotype interface: a phenotype data management system and an integrated multi-environment phenotype fusion algorithm platform. The phenotype data management system provides phenotype data from multiple cooperation units and uniquely constructs the most comprehensive phenotype database of rice currently. The platform supports phenotype data retrieval, visualization, and phenotype distribution fusion based on multi-environment algorithms. After submitting data and parameters, users manage tasks through the task management system, including pausing / resuming, viewing visualization analysis reports, downloading reports, and analysis results.
[0038] S3. Develop a GS interface: a breeding platform that provides introductions, training platforms, and prediction platforms for 8 genomic selection algorithms, including the self-developed algorithm LGBMY and other algorithms such as GBLUP, LGBM, XGB, HGB, SVR, and CATB. In the GS training submission column, users add data and parameters according to their needs, select one or more algorithms for model training, and manage tasks through the task management system, supporting pausing / resuming, viewing visualization analysis reports, and manually downloading reports and results. After completing the training, users predict the phenotypes of the offspring through the GS prediction submission column and perform relevant operations in the task management system.
[0039] S4. Develop a Help interface: provide an operation manual for the Rice3kGS platform to help users quickly familiarize themselves with the platform functions and operation processes and achieve efficient use.
[0040] The rice simulation breeding platform for predictable breeding is named Rice3kGS. It uses a self-designed interface, including Home, Phenotype, GS, and Help interfaces. The layout of each interface is intuitive and simple, the operation is convenient, the task management is efficient, and the result report is clear and easy to understand. Moreover, the platform has rich phenotypic data resources provided by multiple collaborating units. By constructing a database, it can quickly retrieve all sample data of target traits, support statistical analysis and visual display. And the database includes rich variation data, covering various variation types, including SNP variations within about 1 Mb, conserved SNP locus variations of about 404 Kb, and large fragment variations such as Deletion, Insertion, Inversion, and Duplication variations. In addition, it also includes various types, namely STR variations and gene deletion / insertion gene-PAV variations, providing comprehensive variation data support for breeding analysis;
[0041] By integrating multi-environment algorithms, this rice simulation breeding platform can quickly and conveniently fuse and normalize multi-environment data, supporting three algorithm selections: BLUP, BLUE, and AVERAGE. And the platform integrates 8 genome-wide selection breeding algorithms. Users can choose one or more algorithms to train phenotypic and genotypic data. During the training process, 10-fold cross-validation is adopted, and the model is evaluated through four indicators: PCC, RMSE, MSE, and MAE. Moreover, the platform integrates the self-developed LGBMY algorithm. By optimizing the parameters of the LGBM algorithm and selecting the best parameters to improve the model accuracy, it can improve the precision of breeding selection;
[0042] In S3, the number of samples for which the platform predicts offspring that can be used for germplasm conservation or breeding individuals is default set to the top 100 samples, or users can also select the top n samples for germplasm conservation by themselves. And in S2 and S3, the platform provides task progress viewing, result viewing, and downloading. And the application scope of this platform includes plants, animals, and their variation types. Plants include rice, corn, wheat, and cotton, and animals include pigs, cows, sheep, chickens, ducks, and geese.
[0043] Example 1
[0044] Refer to Figure 2 As can be seen, using the data of 2412 samples in 3 years such as 2015, 2016, and 2017, the AVERAGE algorithm will be used for fusion and normalization. The operation process of the platform using this method is as follows: First, submit the data and corresponding parameters. After completion of submission, the task enters the task management interface. After running, a "Finish" sign indicating completion appears. Click the button to view the sign, and the interface generates a result report. By viewing the content in the report, the fusion result and visual result of each sample can be obtained.
[0045] Example 2
[0046] Refer to Figure 3 It can be seen that using the data of the 2412 samples in 2015, 2016, 2017 and other 3 years, the BLUP algorithm is used to fuse and normalize. The operation process of the platform using this method is as follows: First, submit the data and the corresponding parameters. After the submission is completed, the task enters the task management interface. After running, a "Finish" flag indicating completion appears. Click the button to view the flag, and the interface generates a result report. By viewing the content in the report, the fusion results and visualization results of each sample can be obtained.
[0047] Example 3
[0048] Refer to Figure 4 It can be seen that using the data of the 2412 samples in 2015, 2016, 2017 and other 3 years, the BLUE algorithm is used to fuse and normalize. The operation process of the platform using this method is as follows: First, submit the data and the corresponding parameters. After the submission is completed, the task enters the task management interface. After running, a "Finish" flag indicating completion appears. Click the button to view the flag, and the interface generates a result report. By viewing the content in the report, the fusion results and visualization results of each sample can be obtained.
[0049] Example 4
[0050] Refer to Figure 5 It can be seen that in the platform phenotype management system, all phenotypes related to "grain" are retrieved, and then the phenotypes of the target traits are selected from them for statistics and visualization. The process of the platform phenotype management system retrieving phenotypes related to "grain" is as follows: First, enter the character "grain" in the search box, and all phenotypes related to the character "grain" will automatically appear in the search box; click the "search" button, and all traits with the character "grain" and the statistical data will be displayed in the result box. Then click the "Total" number in the target row of "Thousand_grain_weight", and a new page will jump to, and the content is the detailed information of this trait, including the statistical results and visualization results.
[0051] Example 5
[0052] Refer to Figure 6It can be seen that on the GS interface of the platform, the interfaces and results of GS models trained by using the SNP data of 3k middle indica rice and the Lesion_height phenotype data and selecting various algorithms are presented. In the process of training the GS model on the platform, first, input data and corresponding parameters are submitted on the interface. The input uses the SNP data of 3k indica rice and the Lesion_height phenotype data file. After supplementing other references, click Submit to complete the submission. The task enters the task management interface. After running, a "Finish" sign appears. Click the button to view the sign, and the interface generates a result report. View the content in the report to obtain the integrated results of each sample and the visualization report. In the report, a specified model can be selected, and the model evaluation results trained by the corresponding algorithm are given, as well as the results that can be downloaded.
[0053] Example Six
[0054] Refer to Figure 7 It can be seen that on the genomic selection GS interface of the platform, using the SNP data of 3K indica rice and the Lesion_height phenotype data, various algorithms are selected to train the GS model respectively, demonstrating the specific process of training the GS model on the platform: First, input the data file and related parameters on the submission interface, including the SNP data of 3K indica rice and the Lesion_height phenotype data file, and supplement other reference information. After completing the input, click Submit. The task enters the task management interface. After running, a "Finish" sign is displayed. Click the button to view the sign, and a result report can be generated. In the report, the comprehensive results of each sample and the visualization analysis can be viewed. In addition, the report allows selecting a specific model to view the model evaluation results of the corresponding algorithm, and also supports downloading the relevant results.
[0055] In summary, the present invention not only solves the pain points of existing websites but also introduces new algorithms, thus having the following advantages:
[0056] (1) The platform has strong openness. After a user applies for an account and is approved, all functions can be used.
[0057] (2) The platform has up to 8 built-in algorithms, covering traditional models and various machine learning algorithms, greatly broadening the breeding application scenarios.
[0058] (3) The data management on the platform is simple and efficient. It comprehensively integrates the rice phenotype database and supports fuzzy retrieval, facilitating users to quickly find the required phenotype data. The platform also provides data statistics and visualization functions, greatly improving the efficiency of data processing.
[0059] (4) The platform is easy to operate and user-friendly, significantly improving the user experience.
[0060] (5) The visualization of the generated result report is excellent, with convenient operation, good aesthetics, and is conducive to data interpretation and scientific decision-making;
[0061] (6) It has a built-in task management system, which can pause, restart and other operations on the analysis tasks, facilitating the efficient management of task progress;
[0062] (7) The simulation breeding platform of the present invention also supports selecting and training multiple GS model algorithms through phenotypic and genotypic data, and screening out the optimal model for accurate prediction of offspring data. This function helps to achieve rapid breeding, shorten the breeding cycle of new varieties, and provide a rapid selection platform for breeders;
[0063] (8) The present invention also has wide applicability and can flexibly adapt to GS model training and prediction under various material backgrounds and multi-environment fusion algorithm calculations. It not only supports genomic selection of major plants such as rice, corn, wheat, cotton, etc., but also applies to various economic animals such as pigs, cows, sheep, chickens, ducks, geese, etc. This system can process various types of variant data, including SNPs, InDels, SVs, and CNV data generated by resequencing or chips, improving the accuracy and universality of prediction while meeting different experimental requirements, effectively reducing cluster resource consumption, reducing labor costs and time costs, improving the accuracy of breeding, and promoting the breeding process.
[0064] The platform of the present invention is not only applicable to high-throughput sequencing of high-density markers, namely SNPs, InDels, SVs, but also can be used for genomic breeding requirements of chips and reduced-representation genome sequencing, providing a rapid selection platform for breeders to cultivate their own varieties.
[0065] Meanwhile, the content not described in detail in this specification belongs to the prior art well-known to those skilled in the art.
[0066] Although the present invention has been described in detail with reference to the foregoing embodiments, for those skilled in the art, they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A rice breeding simulation platform capable of predicting breeding, characterized in that: The following steps are involved: S1. Develop the Home interface: display data classification, basic profile, source distribution and partner information to provide users with an overview; S2. Develop Phenotype interface: phenotypic data management system and integrated multi-environment phenotypic fusion algorithm platform; S3. Develop GS interface: breeding platform, providing introduction, training platform and prediction platform of 8 genomic selection algorithms, including self-developed algorithm LGBMY and other algorithms GBLUP, LGBM, XGB, HGB, SVR and CATB; S4. Develop Help interface: Provide Rice3kGS platform operation manual to help users quickly familiarize themselves with platform functions and operation procedures and achieve efficient use.
2. The rice breeding simulation platform capable of predicting breeding according to claim 1, characterized in that: The platform uses self-designed interfaces, including Home, Phenotype, GS and Help interfaces. The layout of each interface is intuitive and concise, the operation is convenient, the task management is efficient, and the result report is clear and easy to understand.
3. The rice simulation breeding platform capable of predicting breeding according to claim 1, characterized in that: The platform is rich in phenotypic data resources provided by multiple cooperative units. By building a database, the platform can quickly retrieve all sample data of target traits and support statistical analysis and visual display.
4. The rice simulation breeding platform capable of predicting breeding according to claim 1, characterized in that: The database includes rich variation data, covering various variation types, including SNP variations within a range of about 1 Mb, conservative SNP site variations of about 404 Kb, and large-fragment variations such as Deletion, Insertion, Inversion and Duplication variations. In addition, it also includes multiple types, namely STR variations and gene deletion / insertion gene-PAV variations, providing comprehensive variation data support for breeding analysis.
5. The rice breeding simulation platform capable of predicting breeding according to claim 1, characterized in that: The platform integrates multi-environment algorithms, can quickly and conveniently integrate and normalize multi-environment data, and supports three algorithm options: BLUP, BLUE and AVERAGE.
6. The rice simulation breeding platform capable of predicting breeding according to claim 1, characterized in that: The platform integrates 8 whole-genome selection breeding algorithms. Users can select one or more algorithms to train phenotypic and genotypic data. A 10-fold cross-validation is used during the training process, and the model is evaluated using four indicators: PCC, RMSE, MSE, and MAE.
7. The rice breeding simulation platform capable of predicting breeding according to claim 1, characterized in that: The platform integrates the self-developed LGBMY algorithm, optimizes the parameters of the LGBM algorithm, selects the best parameters to improve the model accuracy, and can improve the accuracy of breeding.
8. The rice breeding simulation platform capable of predicting breeding according to claim 1, characterized in that: In the S3, the platform predicts that the number of offspring that can be used for seed preservation or breeding is selected by default from the top 100 samples, and the top n seed preservation samples can also be selected by the platform.
9. The rice breeding simulation platform capable of predicting breeding according to claim 8, characterized in that: The platforms in S2 and S3 provide task progress viewing, result viewing and downloading.
10. The rice breeding simulation platform capable of predicting breeding according to claim 1, characterized in that: The application scope of the platform includes plants, animals and their variant types. The plants include rice, corn, wheat, and cotton. The animals include pigs, cattle, sheep, chickens, ducks, and geese.