Big data-based biological breeding management method and system
Through big data and cloud computing technology, the whole genome correlation analysis and closed-loop verification are carried out, which solves the problems of low efficiency and inflexible adjustment in traditional breeding methods, and improves the accuracy and flexibility of biological breeding management.
Patent Information
- Application Number
- CN202510928523.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-07
- Publication Date
- 2025-08-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional breeding methods are inefficient, limited by the lack of sample size and genetic knowledge, and it is difficult to explore the complex relationship between genotypes and phenotypes. The breeding plan is fixed and difficult to adjust, and lacks flexibility, resulting in low accuracy and adjustment flexibility in biological breeding management.
By acquiring multi-source heterogeneous data, genotype sequencing and historical phenotype integration, integrated germplasm data are generated and distributed storage, combined with crop image characteristics and environmental response parameters, genome-phenotype prediction models are performed, and closed-loop verification is performed to dynamically optimize breeding schemes.
It has achieved in-depth exploration of the relationship between complex genotypes and phenotypes, improved the accuracy and efficiency of breeding, supported real-time adjustment of breeding strategies, and ensured the effectiveness and flexibility of breeding programs.
Smart Images

Figure CN120492544A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of biological breeding technology, and in particular to a biological breeding management method and system based on big data. Background Art
[0002] Early breeding methods relied primarily on artificial selection and crossbreeding, aiming to improve desirable traits in crops or animals through natural selection or hybridization. However, these methods were inefficient and limited by sample size and limited genetic knowledge. Breakthroughs in genomics have revolutionized breeding, enabling researchers to analyze plant or animal genomes to identify genes associated with desirable traits, thereby improving the speed and accuracy of breeding. By integrating vast amounts of information from genomics, phenotyping, climatology, and other data sources, researchers can make more refined breeding decisions. Processing this data with artificial intelligence and machine learning algorithms can uncover more complex genetic and environmental relationships, significantly improving the efficiency and adaptability of breeding. However, traditional genotype-phenotype association analysis often relies on simple statistical methods, making it difficult to explore the deeper relationships between complex genotypes and phenotypes. Furthermore, traditional breeding plans are often fixed and difficult to adjust, lacking flexibility to adapt to changing environmental conditions and crop growth conditions. This, in turn, results in limited precision and flexibility in biological breeding management. Summary of the Invention
[0003] Based on this, it is necessary to provide a biological breeding management method and system based on big data to solve at least one of the above technical problems.
[0004] To achieve the above objectives, a biological breeding management method based on big data is provided, the method comprising the following steps: Step S1: Obtaining original germplasm data; integrating multi-source heterogeneous data on the original germplasm data, wherein the data integration includes genotype sequencing integration and historical phenotypic integration, generating integrated germplasm data and distributively storing it in a database; Step S2: Acquire germplasm resource images and field meteorological data and perform data preprocessing to generate standardized germplasm resource comprehensive data and synchronize them to the database; Step S3: Perform genome-wide association analysis based on the integrated germplasm data to analyze local gene sequence characteristics and phenotypic deep characteristics; perform rule-based association between local gene sequence characteristics and phenotypic deep characteristics to obtain breeding gene association rules; and screen the integrated germplasm data based on the standardized germplasm resource comprehensive data to generate a genotype-phenotype prediction model; Step S4: Based on the standardized germplasm resource comprehensive data, the genotype-phenotype prediction model is closed-loop verified to generate dynamic optimization breeding program data, the database is marked with breeding molecules according to the dynamic optimization breeding program data, and the marked breeding molecules are observed in the field offspring performance to generate field offspring observation data; based on the preset promotion standards, the field offspring observation data is screened for variety promotion to execute the full process management of biological breeding.
[0005] By integrating genotype sequencing data and historical phenotypic data, the present invention can fully exploit information from various data sources and achieve comprehensive germplasm resource assessment. Distributed storage in a database ensures efficient data storage and access. Data preprocessing generates standardized comprehensive germplasm resource data, helping to eliminate noise and inconsistencies, making subsequent analysis more accurate and reliable. Crop image feature sets and environmental response parameter sets can reflect the growth status and environmental adaptability of crops. Genome-wide association analysis helps identify gene regions associated with specific phenotypes and reveals potential genetic markers, providing a precise genetic basis for subsequent breeding targets. Through in-depth analysis of genotype and phenotypic data, the generated breeding gene association rules can help determine which genes or gene regions have a significant impact on the phenotypic characteristics of crops, providing theoretical support for precision breeding. Based on dynamically optimized breeding plan data, breeding strategies can be adjusted in real time, thereby improving breeding efficiency and accuracy. Closed-loop verification processing ensures the sustainability and verifiability of breeding plan results. Prediction of breeding molecular markers can identify superior genotypes earlier, reducing the time and resource waste in traditional breeding and improving the precision of crop breeding. Therefore, the present invention integrates genotype, phenotype and environmental data through big data and cloud computing technology, thereby improving the accuracy and adjustment flexibility of biological breeding management.
[0006] Preferably, step S1 includes the following steps: Step S11: Screening germplasm phenotypic data according to breeding objectives to obtain original germplasm data; Step S12: parsing the genotype sequencing information of the original germplasm, and performing cross-platform sequencing alignment on the genotype sequencing information to generate genotype data in a unified format; Step S13: fusing historical phenotype information on the unified format genotype data to generate genotype-phenotype fusion data; Step S14: performing time series consistency calibration on the genotype-phenotype fusion data to generate time series unified integrated data; extracting multi-source features of the time series unified integrated data and performing structure mapping to generate structurally consistent integrated germplasm data; Step S15: Distribute and store the integrated germplasm data with consistent structure into the database.
[0007] Through cross-platform sequencing alignment (step S12), the present invention ensures the consistency and comparability of genotype data across different platforms. This provides stable foundational data for subsequent cross-platform analysis, making genotype data integration more reliable and enabling comprehensive evaluation of germplasm resources. By integrating genotype sequencing data with historical phenotypic information, more comprehensive information can be provided for breeding. This integration not only helps identify potential genotype-phenotype associations but also provides more precise guidance for crop improvement. Time series consistency calibration in step S14 ensures that phenotypic data collected at different time points are consistent across time and space. This process ensures the consistency of data across multiple time points, providing a reliable data foundation for long-term breeding research. By extracting multi-source features from the time-series unified integrated data and performing structural mapping, step S14 further enhances the multi-dimensional data fusion effect. The resulting structurally consistent integrated germplasm data has higher data quality, facilitating subsequent in-depth analysis and modeling. The distributed database storage mechanism in step S15, combining advanced storage architecture with flexible parameter settings, enables efficient and stable storage and management of integrated germplasm data in the cloud. The database allows researchers to easily access and share data, promoting collaboration and sharing. The distributed nature of database storage and flexible parameter settings ensure that data storage not only meets current needs but also adapts to future data expansion. As research deepens, storage capacity and data processing capabilities can be seamlessly expanded to ensure long-term data stability.
[0008] Preferably, step S15 includes the following steps: Step S151: indexing the integrated germplasm data with consistent structure to generate a germplasm data index table; Step S152: Divide the integrated germplasm data and the germplasm data index table into multiple data blocks, perform data segmentation and distribution, and generate distributed data blocks; Step S153: performing redundant backup of the distributed data blocks to generate redundant backup data blocks; Step S154: store the redundant backup data blocks in the distributed database.
[0009] The present invention accelerates data retrieval and access by indexing the integrated germplasm data (step S151) and generating a germplasm data index table. Indexing makes data queries more efficient, reducing retrieval time and computational complexity. By dividing the integrated germplasm data and the germplasm data index table into multiple data blocks (step S152) and assigning them to different storage locations, storage pressure is reduced and data distribution is evenly distributed. This helps improve the scalability and load balancing of the storage system. Redundant backup (step S153) prevents data loss or corruption and improves data fault tolerance. Even if a data block fails, the redundant backup ensures data integrity and recoverability. Storing the redundant backup data blocks in a distributed database (step S154) leverages the high availability and fault tolerance of the distributed system, ensuring real-time synchronization and efficient storage of data across multiple nodes, thereby improving system stability and service continuity. Through distributed data block storage and management, the system can handle larger data sets while supporting parallel computing and analysis tasks. This provides a solid foundation for subsequent data mining, analysis, and application.
[0010] Preferably, step S2 includes the following steps: Step S21: using drones, mobile data collection apps, and IoT devices to acquire germplasm resource images and field meteorological data; Step S22: integrating germplasm resource images and field meteorological data into original germplasm growth data; Step S23: performing image preprocessing on the germplasm resource images in the original germplasm growth data to generate a crop image feature set, wherein the image preprocessing includes image filtering, image binarization, and image feature point extraction; performing data preprocessing on the field meteorological data in the original germplasm growth data to generate an environmental response parameter set, wherein the data preprocessing includes data cleaning, data denoising, missing value filling, and data standardization; Step S24: Integrate the crop image feature set and the environmental response parameter set into standardized germplasm resource comprehensive data; perform data quality detection on the standardized germplasm resource comprehensive data, and synchronously store the detected standardized germplasm resource comprehensive data in the database.
[0011] The present invention uses drones and a mobile data collection app to capture germplasm resource images and field meteorological data (step S21), ensuring extensive and comprehensive data collection. The drone's efficient coverage and the app's convenience enable real-time acquisition of dynamic field information, providing high-quality raw data for subsequent analysis. The germplasm resource images and field meteorological data are integrated into raw germplasm growth data (step S22), providing comprehensive input for subsequent data processing. This data fusion helps establish more accurate crop growth models and, based on this, further extract the complex relationship between the environment and crop growth. The germplasm resource images are preprocessed to generate a crop image feature set (step S23), including image filtering, binarization, and feature point extraction. These processing steps remove image noise, improve image clarity and analysis accuracy, and extract key features for subsequent analysis. Preprocessing of field meteorological data (data cleaning, denoising, missing value filling, and standardization) also helps improve data quality and avoid errors caused by incomplete or noisy data. By integrating the crop image feature set and the environmental response parameter set (step S24), standardized germplasm resource comprehensive data is generated, ensuring data consistency and comparability. Standardized data can effectively avoid bias caused by different collection sources and methods, ensuring data reliability during analysis. Data quality testing of standardized germplasm resource comprehensive data further ensures data accuracy and integrity, preventing data errors from affecting subsequent analysis. Synchronous storage in a database facilitates long-term data management and traceability, ensuring data availability and security.
[0012] Preferably, performing genome-wide association analysis based on integrated germplasm data in step S3 includes: Extract single nucleotide polymorphism sites from integrated germplasm data to obtain standardized SNP feature data; Conduct stratified screening of trait stability on integrated germplasm data to generate stability trait subset data; Perform site-by-site linear regression modeling on the standardized SNP feature data and trait subset data to generate initial association signal matrix data; Perform multiple hypothesis testing correction on the initial correlation signal matrix data to generate significance screening result data; Perform Bayesian sparse representation-driven association path reconstruction on the significant screening result data to generate optimized gene-trait association network data; The optimized gene-trait association network data is subjected to model structure nested encoding to obtain genotype-phenotype association model data.
[0013] The present invention can effectively reduce noise by extracting standardized SNP feature data and performing trait stability stratification screening, ensuring that the analysis is based only on stable and representative trait data, thereby improving the accuracy of genotype-phenotype association analysis. By performing multiple hypothesis testing corrections on the initial association signal matrix data, the false positive rate can be effectively controlled, ensuring that only truly significant association signals are screened out, and improving the reliability of the analysis results. The Bayesian sparse representation-driven association path reconstruction method can further refine the gene-trait association model, making the association network more accurate and compact, avoiding overfitting, and improving the generalization ability of the model. The genotype-phenotype association model data is generated by optimizing the nested coding of the model structure, so that the final association model is more adaptable to different data sets and actual application requirements, enhancing the flexibility and applicability of the model. This process efficiently integrates multi-source data (genotype data and phenotypic data), providing more comprehensive and accurate basic data for subsequent analysis and decision-making, and ensuring the consistency and integrity of the data.
[0014] Preferably, the data quality test of the standardized germplasm resource comprehensive data includes: Conduct consistency verification on the comprehensive data of standardized germplasm resources and generate consistency verification reports; Conduct data distribution analysis on standardized germplasm resource comprehensive data and generate data distribution maps and analysis reports; Conduct duplication detection on standardized germplasm resource comprehensive data and generate duplication data detection report; Conduct integrity verification on standardized germplasm resource comprehensive data and generate integrity verification reports; The consistency verification report, data distribution map and analysis report, duplicate data detection report and integrity verification report are integrated as the data quality detection results; based on the data quality detection results, the standardized germplasm resource comprehensive data after detection is synchronously stored in the database.
[0015] The present invention ensures the coordination between different data sources and the consistency of the data itself by performing consistency verification on the comprehensive data of standardized germplasm resources (consistency verification report in the step). This reduces errors caused by data inconsistencies and improves the accuracy of subsequent analysis. Through data distribution analysis (generating data distribution graphs and analysis reports), the distribution of data can be intuitively presented, helping to identify data deviations, outliers or skewed distributions. This visualization effect provides data scientists with a clearer image, making the analysis results more intuitive and enabling the detection of potential problems in the early stages. Repeatability detection (generating duplicate data detection reports) can ensure that duplicate data in the dataset is discovered and processed in a timely manner, preventing the reliability and validity of the analysis results from being affected by data duplication. This provides a guarantee for ensuring the accuracy of the final model or decision. Through integrity verification (generating integrity verification reports), missing data or incomplete data items can be discovered and corrected in a timely manner to ensure the integrity of the dataset. This verification step can reduce analysis deviations or result errors caused by missing data and ensure the availability of data in subsequent processing. Integrating all test reports (including consistency verification reports, data distribution graphs and analysis reports, duplicate data detection reports, and integrity verification reports) into a complete data quality test result comprehensively demonstrates the data quality status and provides a scientific basis for subsequent data application and decision-making. This report not only provides a comprehensive analysis of the data but also serves as a basis for data quality control, supporting continuous data optimization and quality improvement. Based on the data quality test results, only standardized germplasm resource comprehensive data that has passed quality inspection can be synchronously stored in the database, ensuring that the data in the database meets certain quality standards. This mechanism ensures high-quality data storage and avoids subsequent processing difficulties caused by storage errors or quality issues.
[0016] Preferably, the analysis of the local features of the gene sequence and the deep features of the phenotype in step S3 includes: Extract candidate gene segments from genotype-phenotype association model data to obtain local gene sequence data; perform sequence k-mer encoding conversion on the local gene sequence data to generate gene sequence structured feature data; Extract multi-scale deep features of genotype-phenotype association model data to obtain phenotypic deep feature data; reconstruct feature space mapping between gene sequence structured feature data and phenotypic deep feature data to generate homologous nested feature collaborative data; Correlation density clustering is performed on the homologous nested feature collaboration data to generate local gene-phenotype significant coupling feature data; the local gene-phenotype significant coupling feature data is visualized and thermally encoded to obtain the local gene sequence features and phenotypic depth features of the genotype-phenotype association model data.
[0017] The present invention converts gene sequence data into structured feature data through k-mer encoding conversion, which can better reveal the potential patterns and laws in the gene sequence and provide efficient and accurate genotype information for subsequent association analysis. By extracting multi-scale phenotypic depth features, the detailed information in the phenotypic image can be captured more comprehensively, providing richer multi-dimensional data support for the analysis of phenotypic characteristics, and improving the depth and accuracy of phenotypic data. By performing feature space mapping reconstruction on gene sequence structured feature data and phenotypic depth feature data, different types of data sources can be effectively integrated to generate homologous nested feature collaborative data, thereby improving the correlation and matching degree between genotype and phenotype. Through correlation density clustering, local gene-phenotype significant coupling features can be identified and extracted, further improving the correlation strength and scientificity between genotype and phenotype, and helping to accurately identify key genotype-phenotype association relationships. Through visual thermal coding, complex genotype-phenotype association data is converted into a form that is easy to understand and analyze, making the analysis results more intuitive and easy to interpret, facilitating subsequent decision-making and further breeding research. Overall, through the sophisticated feature extraction, fusion and optimization process, the accuracy of the association analysis between genotypes and phenotypes has been effectively improved, providing reliable data support for subsequent precision breeding, crop improvement and genetic research.
[0018] Preferably, in step S3, rule-association of the local features of the gene sequence with the deep features of the phenotype includes: Perform feature item encoding on local features of gene sequences to generate encoded gene feature item data; perform pattern discretization on phenotypic depth features to generate discrete phenotypic image feature data; Align the coded gene feature item data with the discrete phenotype image feature data to generate gene-phenotype feature pair data; perform frequent item set mining on the gene-phenotype feature pair data to generate frequent gene-phenotype combination data; Association rules are generated for frequent gene-phenotype combination data to obtain breeding gene association rules.
[0019] The present invention can convert raw data into more representative and structured feature data by encoding local features of gene sequences and discretizing patterns of phenotypic deep features, thereby improving the effect and efficiency of subsequent analysis. By aligning the feature items of the encoded gene feature item data with the discrete phenotypic image feature data, the relationship between genotype and phenotype can be effectively matched, data noise can be reduced, and the precision and accuracy of association analysis can be ensured. By mining frequent item sets of genotype-phenotype feature pairs, potential frequent genotype-phenotype combinations can be revealed, providing more accurate candidate genes and phenotype combinations for breeding research, reducing unnecessary calculations and assumptions, and improving analysis efficiency. By generating breeding gene association rules, the potential laws between genes and phenotypes can be excavated, providing a scientific basis for precision breeding. These association rules help predict the interaction between genotype and phenotype, thereby optimizing breeding strategies and decisions. Overall, the deep relationship between genotype and phenotype can be effectively established through the process of rule association, providing systematic and operational decision support for crop improvement and breeding, and improving breeding efficiency and success rate.
[0020] Preferably, step S4 includes the following steps: Step S41: performing a closed-loop genetic effect verification on the genotype-phenotype prediction model based on the standardized germplasm resource comprehensive data to obtain a genetic effect verification result; Step S42: Pairing the integrated germplasm data and the standardized germplasm resource comprehensive data into breeding schemes based on the genetic effect verification results to generate dynamically optimized breeding scheme data; Step S43: marking the database for breeding molecules according to the dynamically optimized breeding scheme data, and observing the performance of offspring in the field for the marked breeding molecules to generate field offspring observation data; Step S44: Perform variety promotion screening on the field offspring observation data based on the preset promotion criteria to execute the full process management of biological breeding.
[0021] The present invention ensures the authenticity and validity of every genetic effect during the breeding process by performing closed-loop validation of the genetic effects of the genotype-phenotype prediction model based on comprehensive standardized germplasm resource data (step S41). This validation process provides a scientific basis, helping researchers identify truly beneficial genotype-phenotype relationships, thereby reducing the impact of inaccurate predictions and invalid genetic effects and improving breeding success rates. Breeding strategies are then paired with germplasm data based on the genetic effect validation results (step S42). By generating dynamically optimized breeding strategy data, breeding strategies can be adjusted in real time to ensure that the breeding process consistently targets optimal breeding goals. The dynamic optimization strategy can be continuously updated based on actual validation results, supporting more precise and flexible breeding plans. Based on the dynamically optimized breeding strategy data, breeding molecular markers in the database are evaluated, and field progeny performance observations of these markers are performed (step S43), enabling accurate prediction of the performance of different genotypes in actual field environments. This combination of genetic markers and field performance observations can significantly improve breeding efficiency, avoid premature elimination of varieties with good potential, and provide reliable genetic guidance for future generation breeding. Variety promotion screening (step S44) based on pre-set promotion criteria based on field progeny observation data ensures that only varieties meeting the specified criteria are promoted, reducing the time and resources wasted on inefficient or unsuitable varieties. This screening mechanism effectively ensures efficient resource utilization during the breeding process and accelerates the selection of superior varieties. By implementing full-process management of biological breeding, not only can the systematic and scientific nature of breeding be strengthened, but the efficiency of the entire breeding cycle can also be improved. Each stage is precisely verified and optimized, reducing manual intervention and unnecessary trial and error, making the breeding process more automated, precise, and efficient.
[0022] In this specification, a biological breeding management system based on big data is provided, which is used to implement the above-mentioned biological breeding management method based on big data. The biological breeding management system based on big data includes: The genotype and phenotype data management module is used to obtain original germplasm data; integrate multi-source heterogeneous data of the original germplasm data, where data integration includes genotype sequencing integration and historical phenotypic integration, generate integrated germplasm data and store it in a distributed database; A germplasm resource management module is used to obtain germplasm resource images and field meteorological data and perform data preprocessing, generate standardized germplasm resource comprehensive data and synchronize it to the database; The gene analysis module is used to perform genome-wide association analysis on integrated germplasm data based on standardized germplasm resource comprehensive data, analyze local gene sequence characteristics and phenotypic deep characteristics; perform rule association between local gene sequence characteristics and phenotypic deep characteristics to obtain breeding gene association rules. Based on standardized germplasm resource comprehensive data, integrated germplasm data is screened to generate a genotype-phenotype prediction model; The breeding prediction and promotion module is used to perform closed-loop verification processing on the genotype-phenotype prediction model based on standardized comprehensive data of germplasm resources, generate dynamic optimization breeding plan data, mark breeding molecules in the database according to the dynamic optimization breeding plan data, and observe the performance of the marked breeding molecules in the field to generate field offspring observation data; based on the preset promotion standards, the field offspring observation data is screened for variety promotion to execute the full process management of biological breeding.
[0023] The beneficial effect of the present invention is that the original germplasm data is integrated with genotype sequencing and historical phenotypes through the germplasm management module to generate integrated germplasm data, and a database is used for distributed storage, which can efficiently manage massive multi-source heterogeneous data, improve data access speed and reliability, and provide high-quality data support for subsequent analysis. The gene phenotype management module can efficiently pre-process the original growth data and generate standardized germplasm resource comprehensive data, including crop image feature sets and environmental response parameter sets. This process ensures the consistency and high quality of the data, and synchronizes the standardized data to the database for subsequent analysis and real-time updates. The gene analysis module generates genotype-phenotype association model data through whole-genome association analysis, analyzes the local characteristics of the gene sequence and the deep characteristics of the phenotype, and performs rule association to obtain breeding gene association rules. This process provides a data-based scientific basis for subsequent breeding decisions, helps to discover potential breeding traits and gene markers, and improves breeding efficiency. The breeding prediction module performs closed-loop verification on the genotype-phenotype association model data based on the breeding gene association rules to generate dynamically optimized breeding program data. This module can also construct breeding molecular markers in the database according to the optimization scheme, and predict the performance of the marked breeding molecules in the field offspring. This process helps to accurately predict the breeding effect, provide decision support for biological breeding management, and ensure the efficiency and scientificity of the variety improvement process. The overall system not only improves the accuracy and efficiency of breeding, but also optimizes the planting management strategy through efficient data management, scientific gene analysis, accurate breeding prediction and optimization scheme. The system provides strong support for intelligent management in the field of biological breeding and promotes the rapid development of precision agriculture. Therefore, the present invention integrates genotype, phenotypic and environmental data through big data and cloud computing technology, thereby improving the accuracy and adjustment flexibility of biological breeding management. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1This is a flowchart of the steps of a biological breeding management method based on big data; Figure 2 for Figure 1 Detailed implementation steps of step S2 in FIG. Figure 3 for Figure 1 Detailed implementation steps of step S4 in FIG. The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION
[0025] The following is a clear and complete description of the technical method of the present invention in conjunction with the accompanying drawings. It is obvious that the embodiments described are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts are within the scope of protection of the present invention.
[0026] In addition, the accompanying drawings are merely schematic illustrations of the present invention and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor and / or microcontroller approaches.
[0027] It should be understood that although the terms "first," "second," and the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used solely to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. The term "and / or" as used herein includes any and all combinations of one or more of the listed associated items.
[0028] To achieve this, please refer to Figures 1 to 3 , a biological breeding management method based on big data, the method comprising the following steps: Step S1: Obtaining original germplasm data; integrating multi-source heterogeneous data on the original germplasm data, wherein the data integration includes genotype sequencing integration and historical phenotypic integration, generating integrated germplasm data and distributively storing it in a database; Step S2: Acquire germplasm resource images and field meteorological data and perform data preprocessing to generate standardized germplasm resource comprehensive data and synchronize them to the database; Step S3: Perform genome-wide association analysis based on the integrated germplasm data to analyze local gene sequence characteristics and phenotypic deep characteristics; perform rule-based association between local gene sequence characteristics and phenotypic deep characteristics to obtain breeding gene association rules; and screen the integrated germplasm data based on the standardized germplasm resource comprehensive data to generate a genotype-phenotype prediction model; Step S4: Based on the standardized germplasm resource comprehensive data, the genotype-phenotype prediction model is closed-loop verified to generate dynamic optimization breeding program data, the database is marked with breeding molecules according to the dynamic optimization breeding program data, and the marked breeding molecules are observed in the field offspring performance to generate field offspring observation data; based on the preset promotion standards, the field offspring observation data is screened for variety promotion to execute the full process management of biological breeding.
[0029] In an embodiment of the present invention, raw germplasm data is collected, including genotype sequencing data (such as SNP data), historical phenotypic data (such as plant growth indicators, etc.), and field meteorological data (such as temperature, humidity, precipitation, light, etc.). In addition, germplasm resource image data, such as plant appearance images and leaf morphology images, need to be collected. The above raw data are integrated, and the integration process includes: sorting the raw genotype sequencing data to ensure data consistency and labeling the genetic information. The phenotypic data of many years are sorted and standardized into a unified format for subsequent use. The integrated germplasm data is distributed and stored for subsequent rapid access and query. The germplasm resource image data is subjected to pre-processing steps such as image denoising, segmentation, and calibration to extract the morphological characteristics of the plant, such as leaf size and color distribution. The field meteorological data is time-series processed to consider the impact of multi-dimensional factors such as temperature and humidity on plant growth, and the meteorological data is standardized (for example, unit unification, filling in missing data values, eliminating outliers, etc.). Based on image processing and meteorological data, standardized phenotypic data is generated, including phenotypic characteristics, morphological features, and growth conditions at various growth stages. Genome-wide association studies (GWAS) are conducted on the integrated germplasm data and standardized germplasm resource data to explore the relationship between specific genotypes and phenotypes. Through GWAS analysis, genotype-phenotype association models are constructed to reveal which loci significantly influence specific phenotypes. In genotype-phenotype association models, the association between local features of the gene sequence and deeper phenotypic features is analyzed. For example, whether a gene affects plant growth, stress resistance, and other phenotypes under specific environments is determined. The associations between genotypes and phenotypes are converted into breeding gene association rules. For example, certain genotypes can lead to greater drought resistance or higher yield in plants under specific climatic conditions. These rules provide a basis for subsequent breeding efforts. Based on the derived breeding gene association rules, the genotype-phenotype association model data is closed-loop validated. This process verifies that the effect of each genotype on the phenotype is consistent with expectations, ensuring the effectiveness of the breeding plan. Based on the results of the closed-loop validation, breeding molecular markers are generated for the germplasm data in the database. These markers can indicate which gene loci have a significant effect on the control of target traits. Breeding molecular markers are used to predict the performance of offspring in the field and simulate the performance of different genotypes under different environmental conditions. Through multiple simulations, it is predicted whether the performance of offspring meets the breeding goals, such as performance in disease resistance, drought resistance, etc. Based on the predicted data of offspring in the field and the preset promotion criteria, candidate varieties are screened and the best breeding materials are selected to enter the next stage. These screening criteria include the performance of target traits, genetic stability, environmental adaptability, etc. Through all the above steps, precision breeding is implemented, from the acquisition of germplasm resources to the generation of breeding molecular markers to the prediction of field performance, and finally the best varieties that are most suitable for the environment are screened.The entire breeding process is based on an intelligent optimization system that automatically adjusts breeding strategies based on real-time data feedback to ensure the maximum achievement of breeding goals. By monitoring changes in breeding data in real time, we can promptly adjust the selection of germplasm resources and optimize breeding plans, ensuring efficient execution of the entire breeding process, ultimately achieving the cultivation of high-yield, stress-resistant, and superior varieties.
[0030] Preferably, step S1 includes the following steps: Step S11: Screening germplasm phenotypic data according to breeding objectives to obtain original germplasm data; Step S12: parsing the genotype sequencing information of the original germplasm, and performing cross-platform sequencing alignment on the genotype sequencing information to generate genotype data in a unified format; Step S13: fusing historical phenotype information on the unified format genotype data to generate genotype-phenotype fusion data; Step S14: performing time series consistency calibration on the genotype-phenotype fusion data to generate time series unified integrated data; extracting multi-source features of the time series unified integrated data and performing structure mapping to generate structurally consistent integrated germplasm data; Step S15: Distribute and store the integrated germplasm data with consistent structure into the database.
[0031] In this embodiment of the present invention, germplasm resources are initially screened based on breeding objectives (e.g., improving drought resistance, disease resistance, and yield). Phenotypic data (e.g., plant height, yield, and resistance) is collected from these resources to identify potential germplasm resources. This phenotypic data includes historical planting information, field performance, and environmental conditions. The collected data comes from a variety of sources, including field trials, germplasm image data, and environmental sensor data. The screened germplasm resources serve as the foundation for subsequent genotyping analysis and fusion. Genotypic sequencing information within the raw germplasm data is parsed. This data includes genotypic SNPs (single nucleotide polymorphisms) and gene variant sites. Data generated by different sequencing platforms (e.g., Illumina and PacBio) may differ in format, necessitating cross-platform sequencing alignment of genotypic data from different sources. Sequencing results are converted to a unified format using standardized alignment methods (e.g., using GATK or the BWA alignment tool) to ensure data comparability and consistency. Through this alignment and alignment, a unified genotypic dataset is generated. This dataset contains genotypic information for all germplasm resources and standardizes differences between platforms. Genotypic and phenotypic data are fused together using existing historical phenotypic data. Historical phenotypic data includes performance data on different germplasm resources during cultivation, such as yield, resistance, and quality. Through data matching and association analysis, genotypic data and corresponding phenotypic data are correlated to generate genotype-phenotype fusion data. For example, by correlating changes in specific loci in the genotype with the expression of related traits in the phenotype, a relationship model between genotype and phenotype is constructed. During the germplasm breeding process, phenotypic data may be generated at different time points. For example, the expression of certain traits varies at different growth stages, necessitating time-based consistency calibration. Time series analysis techniques are used to calibrate the genotype-phenotype data across time periods to ensure that data at different time points share the same reference standard. This process can be achieved through interpolation and smoothing algorithms to ensure temporal consistency of the data. After time series consistency calibration, unified and integrated temporal data are generated. This data contains both phenotypic and genotypic characteristics at each time point, enabling more accurate subsequent breeding analysis. Extract multi-source features from unified, integrated time-series data, such as genotypes, environmental factors (such as climate and soil), and phenotypic traits. These features comprehensively describe the performance of germplasm resources and their relationship to the environment. Through multi-dimensional analysis, the performance characteristics of germplasm resources under different environmental conditions are obtained. Data mapping techniques (such as multidimensional space mapping and feature dimensionality reduction) are used to structurally map the extracted multi-source features. The goal is to unify the structure of different types of data (phenotype, genotype, environment, etc.) and generate a consistent data model.After structural mapping, the generated integrated germplasm data contains genotype data, phenotypic data, environmental factors, and other information in a unified format, facilitating subsequent data analysis, breeding optimization, and decision support. Cloud storage or distributed file systems ensure data security, scalability, and efficient access. Data storage must be synchronized in real time to ensure that subsequent queries and analyses utilize the latest data. A database structure tailored to the needs of germplasm resource management should be adopted to support rapid data retrieval, analysis, and decision support.
[0032] Preferably, step S15 includes the following steps: Step S151: indexing the integrated germplasm data with consistent structure to generate a germplasm data index table; Step S152: Divide the integrated germplasm data and the germplasm data index table into multiple data blocks, perform data segmentation and distribution, and generate distributed data blocks; Step S153: performing redundant backup of the distributed data blocks to generate redundant backup data blocks; Step S154: store the redundant backup data blocks in the distributed database.
[0033] In an embodiment of the present invention, integrated germplasm data is indexed. Indexing involves optimizing the sorting and tagging of key fields within the data (such as germplasm resource ID, genotype, and phenotypic characteristics) to facilitate rapid retrieval. By establishing a data index table, data can be quickly queried, for example, searching for corresponding genotype or phenotypic data based on a germplasm ID. The generated germplasm data index table will contain index information for data fields, such as the ID of each germplasm resource, related phenotypic information, and genotype information, ensuring that subsequent queries and analyses can accelerate data access through indexing. The integrated germplasm data and the germplasm data index table are segmented and divided according to specific rules. This division can be based on data volume, data type, or application scenario requirements. Segmentation can be based on germplasm ID range, genotype characteristic interval, phenotypic characteristic category, and so on. Each data block contains a specific range of germplasm resource information for distributed storage. Based on the system load and computing requirements, the divided data blocks are distributed to different distributed storage nodes to achieve load balancing and improve storage and processing efficiency. Each data block contains a portion of the integrated germplasm data and its index information. Data blocks are stored in a distributed manner across different nodes to ensure parallel and efficient data access. Redundant backups are implemented for distributed data blocks. Redundant backups are a data protection strategy used to prevent data loss or corruption. Redundant backups ensure that data can be restored from backups even if a storage node fails, ensuring high data availability and security. Replica backups (e.g., creating multiple copies of each data block) can be used to ensure data is backed up on multiple nodes. The degree of redundancy (e.g., the number of replicas) is typically determined based on the requirements of the storage system to ensure reliable data recovery. These redundant backup data blocks have the same structure and content as the original data blocks but are stored on different nodes. Creating redundant backups ensures that data is not lost during storage and processing. Redundant backup data blocks are stored in a distributed database. Distributed databases such as Hadoop, Cassandra, and other NoSQL databases can be used for large-scale data storage and management and are particularly well-suited for processing large-scale, high-frequency data such as germplasm data. Appropriate storage strategies and algorithms should be employed based on factors such as data access frequency, storage capacity, and computing requirements. Typically, redundant backup data blocks are distributed across multiple physical nodes, leveraging the distributed nature of databases to ensure high availability and scalability of data storage. Using a distributed database not only ensures data security and stability, but also improves query efficiency, data fault tolerance, and optimizes data processing capabilities through a distributed architecture.
[0034] Preferably, step S2 includes: Step S21: using drones, mobile data collection apps, and IoT devices to acquire germplasm resource images and field meteorological data; Step S22: integrating germplasm resource images and field meteorological data into original germplasm growth data; Step S23: performing image preprocessing on the germplasm resource images in the original germplasm growth data to generate a crop image feature set, wherein the image preprocessing includes image filtering, image binarization, and image feature point extraction; performing data preprocessing on the field meteorological data in the original germplasm growth data to generate an environmental response parameter set, wherein the data preprocessing includes data cleaning, data denoising, missing value filling, and data standardization; Step S24: Integrate the crop image feature set and the environmental response parameter set into standardized germplasm resource comprehensive data; perform data quality detection on the standardized germplasm resource comprehensive data, and synchronously store the detected standardized germplasm resource comprehensive data in the database.
[0035] In an embodiment of the present invention, a drone flies over a pre-defined field area to capture high-resolution images of germplasm resources. These images provide information on crop growth status, morphological characteristics, and environmental impacts. Ensure that the images captured by the drone have sufficient resolution and detail to support subsequent image processing and analysis, especially when extracting image feature points. A mobile data acquisition app, combined with sensors, captures field meteorological data, including temperature, humidity, wind speed, and precipitation. This data has a significant impact on crop growth and phenotypic performance. Field meteorological data is synchronized and recorded in real time via the app and uploaded to a database for subsequent use. The crop images are aligned with the corresponding field meteorological data to ensure temporal and spatial synchronization. This can be achieved through geolocation information or timestamp matching. The integrated germplasm images and field meteorological data together form the original germplasm growth dataset, which serves as the basis for subsequent analysis. Filtering algorithms (such as Gaussian filtering and median filtering) are used to remove noise from the images to reduce errors during image analysis. Image binarization converts the images into black and white, highlighting the key structures or features of the crops and facilitating subsequent feature point extraction. Image processing techniques, such as corner detection, edge detection, or other feature extraction algorithms, are used to extract key crop feature points from images. These feature points help describe crop growth morphology and changes. Outliers and erroneous data are removed from meteorological data to ensure data quality. De-noising is performed on meteorological data to eliminate inaccuracies caused by sensor errors or external interference. Appropriate methods (such as interpolation and mean filling) are used to fill in missing meteorological data to ensure dataset completeness. Meteorological data is standardized to unify data of different dimensions (such as temperature and humidity) to a common standard, facilitating subsequent analysis and model training. The preprocessed crop image feature set and the environmental response parameter set are fused to generate a standardized germplasm resource comprehensive dataset. This dataset not only includes crop morphological characteristics but also takes into account the influence of environmental factors. Data quality testing tools are used to evaluate the standardized germplasm resource comprehensive data for completeness, consistency, and validity. Statistical analysis, machine learning models, and other methods are used to identify and correct outliers and erroneous data to ensure high data quality. The quality-tested standardized germplasm resource comprehensive data is synchronously stored in the aforementioned database. These data can provide important input for subsequent breeding analysis, model training, and decision support systems.
[0036] Preferably, performing genome-wide association analysis based on integrated germplasm data in step S3 includes: Extract single nucleotide polymorphism sites from integrated germplasm data to obtain standardized SNP feature data; Conduct stratified screening of trait stability on integrated germplasm data to generate stability trait subset data; Perform site-by-site linear regression modeling on the standardized SNP feature data and trait subset data to generate initial association signal matrix data; Perform multiple hypothesis testing correction on the initial correlation signal matrix data to generate significance screening result data; Perform Bayesian sparse representation-driven association path reconstruction on the significant screening result data to generate optimized gene-trait association network data; The optimized gene-trait association network data is subjected to model structure nested encoding to obtain genotype-phenotype association model data.
[0037] In an embodiment of the present invention, single nucleotide polymorphism (SNP) site information is extracted from integrated germplasm data, typically derived from genotype sequencing data. The extracted SNP data is preprocessed, including removing low-quality sites, addressing missing values, and removing duplicates, to ensure data quality. The SNP data is then normalized to ensure a consistent numerical range for each site. Normalization methods can include Z-score normalization or Min-Max normalization to ensure data comparison is performed on the same scale. The processed SNP sites are coded, typically using 0, 1, or 2 to represent different genotypes (for example, 0 represents homozygous, 1 represents heterozygous, and 2 represents homozygous for the other allele). The resulting normalized SNP signature data can be used as input for subsequent analyses and is typically represented as a matrix with rows representing samples and columns representing SNP sites. Stability analysis is performed on various traits in the phenotypic data (such as crop yield, drought tolerance, and pest and disease resistance). Statistical methods (such as analysis of variance and homeostasis index analysis) are used to assess the stability of each trait under different environments. Based on screening criteria (e.g., a stability index greater than a certain threshold), traits that exhibit stability are selected to form a stability trait subset data set. Trait data with high stability are extracted from the original phenotypic data to generate a stability trait subset data set. This data set contains information on traits that exhibit stability under various environmental conditions. A site-by-site linear regression analysis is performed on each SNP locus and each trait in the stability trait subset data set. The model form is: ;in, represents the phenotypic value of a stable trait, Indicates the genotype data of a certain SNP site, is the regression coefficient, is the intercept term, The error term is σ. Through site-by-site regression, the regression coefficients between each SNP and each trait are obtained, thereby calculating the initial association signal matrix. This initial association signal matrix contains the regression coefficients between each SNP and each trait, with rows representing SNP loci and columns representing phenotypic traits. Because genome-wide association studies involve a large number of hypothesis tests (site-by-site regression), multiple hypothesis testing correction is necessary. Common correction methods include adjusting the significance level based on the number of tests (e.g., setting the significance threshold to 0.05 divided by the total number of tests). The Benjamini-Hochberg method is used to control the false discovery rate to avoid false positive results due to multiple testing. The corrected results provide a significance indicator for the association between each SNP and the trait. Based on the adjusted P-value or false discovery rate, significant SNP-trait associations are screened to generate a significance screening result dataset containing association information between the significant SNP loci and the traits. The SNP-trait association data from the significance screening results are modeled using Bayesian sparse representation methods. Bayesian sparse representation can optimize the model using sparsity constraints to extract the most predictive gene-trait association pathways. The model estimates the importance of each gene-trait pathway using prior distributions, likelihood functions, and posterior distributions, and then reconstructs the pathways. Using Bayesian sparse representation, the most influential association pathways are screened and the gene-trait association network is optimized. This process identifies core genes and their pathways with traits, further removing redundant or irrelevant associations. The resulting optimized gene-trait association network reflects the complex relationships between genes and traits, providing more accurate genotype-phenotype association data. Nested encoding is then performed on the resulting optimized gene-trait association network. This process hierarchically nests gene nodes and trait nodes within the network, encoding gene-trait relationships at different levels. Nested coding can use methods such as graph convolutional neural networks (GCN) to encode association networks, thereby extracting more fine-grained gene-trait relationship patterns. The generated genotype-phenotype association model data contains comprehensive association information between each genotype and phenotype, which is suitable for subsequent breeding decisions and crop improvement.
[0038] Preferably, the data quality test of the standardized germplasm resource comprehensive data includes: Conduct consistency verification on the comprehensive data of standardized germplasm resources and generate consistency verification reports; Conduct data distribution analysis on standardized germplasm resource comprehensive data and generate data distribution maps and analysis reports; Conduct duplication detection on standardized germplasm resource comprehensive data and generate duplication data detection report; Conduct integrity verification on standardized germplasm resource comprehensive data and generate integrity verification reports; The consistency verification report, data distribution map and analysis report, duplicate data detection report and integrity verification report are integrated as the data quality detection results; based on the data quality detection results, the standardized germplasm resource comprehensive data after detection is synchronously stored in the database.
[0039] In an embodiment of the present invention, consistency check rules are compiled to check whether the values of the same fields in different data sources are consistent. For example, for the same batch of data, the phenotypic values of each sample are checked to ensure they are within the expected range, and fields such as date and timestamps are verified for consistency. The data is then compared with historical data, literature, or existing datasets to check consistency, and any outliers are corrected. The results of the consistency check are summarized in a consistency check report, which lists the verification results for each data item and identifies the specific inconsistencies and their resolution. Statistical analysis tools (such as histograms, box plots, and density curves) are used to generate distribution maps of the standardized germplasm resource comprehensive data. These charts can display information such as the central tendency, dispersion, and skewness of the data. The generated distribution maps are analyzed to identify any abnormal distribution patterns or whether the data conform to an expected distribution pattern (such as a normal distribution). The report includes a description of the distribution map and an analysis of the overall trend of the dataset, specifically any potential deviations from the normal pattern and the possible causes of these deviations. Algorithms (such as hashing algorithms and data deduplication algorithms) are used to check the duplication of the standardized germplasm resource comprehensive data. Duplicate data is identified by checking fields such as sample number and sample characteristics. For detected duplicate data, choose to retain one item, merge duplicates, or remove duplicates as needed. The report will list all duplicate data items, the reasons for the duplication, and the measures taken to address it, as well as the status of the data after repair. Each field in the dataset will be examined to identify missing values. Missing values can be addressed through methods such as filling, interpolation, or deletion. Missing phenotypic data will be supplemented using appropriate methods (such as mean filling and model-based prediction filling) to ensure data integrity. The report will list all missing fields and data items and provide detailed information on the treatment methods and the results of the data repair. The consistency check report, data distribution chart and analysis report, duplicate data detection report, and integrity check report will be combined to generate a comprehensive data quality test report. The report should include the specific methods used for each data quality check, any issues found, the remediation measures taken, and the final data status. The report can be presented in charts, text descriptions, and data tables to facilitate subsequent review and processing. Ensure that the standardized germplasm resource comprehensive data that has undergone quality inspection and correction is stored synchronously in the designated database. Data should be stored efficiently to ensure access and query speed, while also ensuring data security and backup. Use storage formats suitable for big data processing, such as CSV, Parquet, or database table formats, and encrypt storage as needed to protect data privacy.
[0040] Preferably, the analysis of the local features of the gene sequence and the deep features of the phenotype in step S3 includes: Extract candidate gene segments from genotype-phenotype association model data to obtain local gene sequence data; perform sequence k-mer encoding conversion on the local gene sequence data to generate gene sequence structured feature data; Extract multi-scale deep features of genotype-phenotype association model data to obtain phenotypic deep feature data; reconstruct feature space mapping between gene sequence structured feature data and phenotypic deep feature data to generate homologous nested feature collaborative data; Correlation density clustering is performed on the homologous nested feature collaboration data to generate local gene-phenotype significant coupling feature data; the local gene-phenotype significant coupling feature data is visualized and thermally encoded to obtain the local gene sequence features and phenotypic depth features of the genotype-phenotype association model data.
[0041] In this embodiment of the present invention, candidate gene segments associated with the target trait are identified from genotype-phenotype association model data based on genome-wide association analysis results (e.g., SNPs significantly associated with the trait). These gene segments generally include functional regions such as promoters and coding regions. Based on the correlation signals between genotype and phenotypic data, candidate gene regions closely associated with phenotypic traits are selected as research targets. Specific sequence information for the candidate gene segments is extracted. Gene sequences are typically stored in FASTA format, containing all nucleotide information for that segment. The extracted gene sequences are then converted to k-mer encoding. k-mer encoding is a method that segments gene sequences into subsequences (k-mers) of length k. For example, if k = 3, the gene sequence "AGCT" will be segmented into 3-mers: "AGC" and "GCT." Using k-mer encoding, gene sequence data can be converted into fixed-length numerical features, facilitating subsequent machine learning and pattern recognition processing. k-mer encoding can be performed using a sliding window method, and different k values can be set to capture sequence patterns of varying lengths. Frequency statistics are then performed on all k-mer sequences to generate structured feature data for the gene sequence. The frequency of each k-mer is used as a feature, ultimately generating a feature vector representing the structural information of the gene sequence. Multi-scale feature extraction is performed on phenotypic image data (such as crop growth images or other visual phenotypic data). Deep learning methods such as convolutional neural networks (CNNs) are used to extract multi-level features from the images. Existing image processing tools (such as deep convolutional neural networks like ResNet and Inception) can be used to process the phenotypic images and extract deep features at different scales. Deep features typically include image texture, morphology, color, and edge information, effectively describing the detailed characteristics of crop phenotypes. Multi-scale deep features are integrated to generate a deep feature dataset for the phenotypic images. Each feature can represent specific information in the image (such as shape, texture, or surface characteristics). The structured feature data of the gene sequence and the deep feature data of the phenotypic data are mapped into a unified feature space. During this feature space mapping process, joint embedding techniques can be used to associate gene sequence and phenotypic characteristics by learning a common feature space. Feature space mapping can employ techniques such as principal component analysis (PCA), multidimensional scaling (MDS), or neural network-based autoencoders to identify correlations between two types of feature data. In the mapped feature space, features of gene sequences and phenotypic images can be nested into the same vector space, generating homologous nested feature synergy data. Homologous nested feature synergy data combines genotypic and phenotypic information, enabling subsequent association analysis to consider the in-depth features of both genotypes and phenotypes.Cluster analysis of homologous nested feature synergy data uses density-based clustering methods, such as DBSCAN (Density-Based Spatial Clustering of Applications with Noise), to identify coupling patterns between genotypes and phenotypes. Clustering algorithms identify significantly correlated genotype-phenotype feature combinations. Clustering results reveal which genotypes and phenotypic features exhibit significant coupling relationships. Cluster analysis generates a dataset of local genotype-phenotype significant coupling features. This dataset includes pairs of features that are significantly coupled between gene sequence features and phenotypic image features, helping to reveal the complex relationships between specific genes and phenotypes. Visualization techniques such as heatmaps are used to encode these local genotype-phenotype significant coupling features to reveal association patterns between genotypes and phenotypes. Heatmaps display the strength of correlation between each pair of genotype-phenotype features. Heatmap colors indicate correlation or similarity between features, helping researchers quickly understand the relationship between genes and phenotypes. The resulting genotype-phenotype association model data will contain significant coupling relationships between local features of gene sequences (such as k-mer features) and deep features of phenotypic images (such as texture, shape, etc.).
[0042] Preferably, in step S3, rule-association of the local features of the gene sequence with the deep features of the phenotype includes: Perform feature item encoding on local features of gene sequences to generate encoded gene feature item data; perform pattern discretization on phenotypic depth features to generate discrete phenotypic image feature data; Align the coded gene feature item data with the discrete phenotype image feature data to generate gene-phenotype feature pair data; perform frequent item set mining on the gene-phenotype feature pair data to generate frequent gene-phenotype combination data; Association rules are generated for frequent gene-phenotype combination data to obtain breeding gene association rules.
[0043] In embodiments of the present invention, local features of gene sequences (e.g., k-mer features) are converted into discrete feature items. Common encoding methods include frequency-based feature extraction and position-based feature extraction. Each local feature of a gene sequence can be mapped into a set of discrete feature values using specific encoding rules (such as hash encoding or integer encoding). These feature values serve as input for subsequent analysis. The encoded gene feature item data is typically a numerical feature set, such as a specific k-mer pattern or mutation site within a gene segment. Feature items are encoded for all candidate gene segments to generate a complete gene feature item dataset. This dataset contains all encoded features of the gene sequence and is suitable for association analysis with phenotypic data. The deep feature data of the phenotypic image (e.g., texture, shape, color, and edges) is discretized. Discretization converts continuous image deep features into discrete patterns or categories for comparison with the gene data. Common discretization methods include clustering-based discretization (e.g., K-means clustering) and threshold-based discretization (which categorizes feature values into high, medium, and low categories). Discretization can convert complex image features into easily processable categorical data, facilitating subsequent association analysis. After pattern discretization, a phenotypic dataset containing discrete features is generated. This dataset contains various discrete features of the phenotypic image (for example, certain shape or color features of the phenotype are classified into different categories) and can be aligned and analyzed with the genotypic feature data. The encoded genetic feature item data and the discretized phenotypic image feature data are aligned. The goal of alignment is to match genetic data with corresponding phenotypic data on a sample or individual basis to analyze their associations. During the alignment process, individual identifiers (such as sample ID and crop variety) can be used to ensure that each genetic feature is correctly paired with the corresponding phenotypic image feature. The aligned data forms a dataset of genotype-phenotype feature pairs, with each feature pair consisting of a genotype and a phenotypic feature. Frequent itemset mining is performed on the genotype-phenotype feature pair data. Frequent itemset mining is a method based on association rule learning that aims to identify frequently occurring patterns of combinations between genetic and phenotypic features. Common frequent itemset mining algorithms include the Apriori algorithm and the FP-growth algorithm. These algorithms can help identify frequently occurring gene-phenotype combinations as potential breeding targets. The mining process automatically counts which gene-phenotype combinations occur most frequently within the sample set. Frequent item set mining helps uncover potential associations between genes and phenotypes. The mining results form a frequent gene-phenotype combination dataset, listing all frequently occurring gene-phenotype combinations. These combinations serve as the basis for generating breeding gene association rules. Based on this frequent gene-phenotype combination data, association rule learning algorithms (such as the Apriori algorithm and association rule expansion algorithms) are used to generate specific association rules.Association rules are usually in the form of "gene feature A → phenotypic feature B", indicating that there is a significant association between a specific feature of gene A and a specific expression of phenotype B. Association rules can be evaluated for strength and credibility by calculating metrics such as support, confidence, and lift. Support indicates the frequency of the rule in the dataset, confidence indicates the probability of phenotype B occurring when gene A appears, and lift indicates the actual value of the rule. Through the association rule generation process, a set of breeding gene association rules is ultimately obtained. Each rule reflects the potential association between genotypic characteristics and phenotypic expressions. These rules can guide the formulation of breeding strategies, such as selecting crop varieties with specific genotypic characteristics and predicting their performance under specific environments.
[0044] Of particular importance is that in step S3, the generation of a genotype-phenotype prediction model by screening the integrated germplasm data based on the standardized germplasm resource comprehensive data includes: Extract single nucleotide polymorphism sites from integrated germplasm data to obtain standardized SNP feature data; Conduct stratified screening of trait stability on integrated germplasm data to generate stability trait subset data; Select the model algorithm and cross-validation times for the standardized SNP feature data and trait subset data to generate a genotype-phenotype prediction model; A comparative analysis was conducted on the genotype-phenotype prediction models, and the genotype-phenotype prediction models were compared in terms of prediction accuracy, mean square error, and root mean square error to generate a genotype-phenotype prediction model.
[0045] In this embodiment, tools such as GATK and PLINK are used to extract SNP (Single Nucleotide Polymorphism) sites from integrated germplasm data. The extracted SNP data are filtered, for example, by removing sites with a minimum allele frequency (MAF) less than 0.05 and sites or samples with a missingness rate greater than 10%. The retained SNPs are standardized (e.g., 0, 1, and 2 represent homozygous and heterozygous allelic states). A standardized SNP feature data matrix is generated, with each row representing a germplasm sample and each column representing a standardized SNP site. Phenotypic trait data across multiple environments and years is extracted from the integrated germplasm resources. Trait stability indices under different environments are assessed using methods such as BLUP (Best Linear Unbiased Prediction) and analysis of variance (ANOVA). Trait subsets with high trait stability are screened based on stability thresholds (e.g., Shukla stability index and Finlay–Wilkinson regression). A data table of the stability trait subsets is output, containing stability evaluation indicators and the selected highly stable phenotypic traits. Based on standardized SNP signature data and a subset of stability traits, the system performs training and test set partitioning (e.g., 5-fold cross-validation). Modeling is performed using a variety of mainstream genotype-phenotype association prediction algorithms, such as GBLUP (Genomic Best Linear Unbiased Prediction), Ridge Regression BLUP (RR-BLUP), LASSO regression, random forest, and support vector regression (SVR). Cross-validation is performed over a set number of times (e.g., 5-fold and 10-fold cross-validation), and average prediction performance is calculated. Multiple candidate genotype-phenotype prediction models and their cross-validation performance metrics are generated and compared. The prediction results of each candidate model are then compared using the following performance metrics: predictive accuracy (e.g., Pearson correlation coefficient), mean squared error (MSE), and root mean squared error (RMSE). Multi-dimensional visualization of these performance metrics is performed (e.g., radar charts and boxplots). The optimal model is selected based on its overall performance. The final genotype-phenotype prediction model parameters and structure are generated and described, which can be used to predict traits in unknown germplasm materials.
[0046] Preferably, step S4 includes: Step S41: performing a closed-loop genetic effect verification on the genotype-phenotype prediction model based on the standardized germplasm resource comprehensive data to obtain a genetic effect verification result; Step S42: Pairing the integrated germplasm data and the standardized germplasm resource comprehensive data into breeding schemes based on the genetic effect verification results to generate dynamically optimized breeding scheme data; Step S43: marking the database for breeding molecules according to the dynamically optimized breeding scheme data, and observing the performance of offspring in the field for the marked breeding molecules to generate field offspring observation data; Step S44: Perform variety promotion screening on the field offspring observation data based on the preset promotion criteria to execute the full process management of biological breeding.
[0047] In this embodiment of the present invention, relevant breeding gene association rules are extracted by utilizing data from a genotype-phenotype prediction model. These rules describe the relationship between genotype and phenotype and reveal the underlying patterns of genetic effects. A genetic effect validation model is designed to verify the genetic effects of these gene association rules by comparing genotype and phenotype data. This validation model can analyze the data using statistical methods (such as linear regression and generalized linear models) to examine the contribution of genotype to phenotype. Retrospective analysis of genotype-phenotype correlations is performed to ensure the model's predictive accuracy, identify any potential errors or deviations, and make necessary corrections. This process generates a genetic effect validation result for the genotype-phenotype relationship, ensuring the accuracy and reliability of the genetic effect model. Based on the genetic effect validation results from step S41, the validity of the genotype-phenotype association model is evaluated to determine which genetic effects are important for the current breeding goals. Based on the genetic effect validation results and in combination with actual breeding needs (such as crop disease resistance and yield improvement), the integrated germplasm data and standardized germplasm resource comprehensive data are paired. The goal of pairing is to find the genotype and phenotype combination that best meets the breeding objectives. Based on these matching genotype-phenotype combinations, a dynamically optimized breeding plan is generated. This plan, based on the results of genetic effect verification, can be continuously optimized over multiple breeding cycles to improve breeding efficiency. Key molecular markers are selected from a database based on the key genotype data in the dynamically optimized breeding plan. Molecular biology techniques (such as SNP markers and microsatellite markers) are used to identify molecular markers associated with the target traits. Using this molecular marker information, statistical prediction models (such as genetic prediction models and machine learning models) are used to conduct preliminary field progeny screening based on the performance of the tagged breeding molecules. This process can be integrated with historical planting data, meteorological data, and other information. Field trial data is observed on retained breeding materials to generate performance predictions for field progeny. This data provides a basis for progeny selection and breeding decisions. Clear criteria for variety advancement are established based on the target market, crop variety requirements, and actual needs. For example, these criteria may include disease resistance, yield, and quality. Field progeny observation data is compared with the variety advancement criteria to select breeding progeny that meet the advancement criteria. Statistical methods (such as analysis of variance and cluster analysis) are used to determine which varieties have the best performance. Based on the screening results, subsequent breeding processes are initiated, including screening the next generation, adjusting breeding strategies, and planning the next round of experiments. This comprehensive process management ensures the achievement of breeding goals and enhances the controllability and accuracy of the breeding process. Data on the promotion of selected varieties, breeding progress, and results are updated in real time in the breeding database, ensuring transparency of the breeding process and traceability of data.
[0048] It is particularly important that step S43 further includes the following steps: Step S431: performing molecular marker site identification on the dynamically optimized breeding scheme data to generate candidate breeding molecular marker data; Step S432: performing accuracy assessment and environmental adaptability screening on candidate breeding molecular marker data to generate high-confidence molecular marker data; Step S433: Standardizing the format and filing the high-confidence molecular marker data, generating injectable molecular marker coding data, and writing the data into the database; Step S434: screening the molecular marker coding data injected into the database for the marker plants, and screening representative marker plants based on historical phenotypic performance; Step S435: Observe the offspring performance in the field of the breeding materials retained after screening to generate field offspring observation data.
[0049] In an embodiment of the present invention, a genome-wide SNP database or candidate gene locus database is used to compare key trait genes involved in a dynamic optimization breeding plan (such as drought resistance, disease resistance, and high yield). SNP or InDel sites highly associated with the target trait are identified through variant enrichment analysis and LD (linkage disequilibrium) analysis. The identified key sites are extracted and packaged as structured data, including site number, chromosome location, variant type, allele frequency, etc., and output as candidate breeding molecular marker data, which serves as the basis for subsequent screening. Based on existing annotated germplasm sample data, the accuracy of each site in predicting the target trait (such as AUC, F1-score) is evaluated. Sites with high noise or low repetition rate are removed, and the stability performance of candidate sites in different ecological environments is evaluated using multi-environment test data (MET). GxE interaction effect modeling or Bayesian stability index screening methods are used to retain environmentally robust molecular markers, and a molecular marker set that has been dual-screened for accuracy and adaptability is output as high-confidence molecular marker data. Marker data is standardized into a standard VCF (Variant Call Format) or JSON structured format, including metadata (such as sequencing method, screening source, and application scenario). Each marker data entry is formatted using standardized fields (such as marker_id, position, effect, and target_trait). A unique identifier (UID) is assigned to each high-confidence marker, and an indexing system is established. A marker traceability path is constructed to support tracing back to the original breeding plan and verification process. Standardized molecular marker encoding data is uploaded and written to a database for subsequent marker application and traceability. Molecular marker comparisons are performed on plants in the germplasm resource library to screen for individuals carrying the target molecular marker. Filtering modes include full marker matching, partial marker matching, and optimal matching. Historical field phenotypic data of plants carrying the corresponding markers is retrieved, including growth period, disease incidence, and yield. Marker plants with excellent phenotypic performance and genetic stability are selected, while samples with high volatility or significant performance differences are eliminated. The output of labeled plant sample data includes marker information, phenotypic performance, selection source, and intended application scenario, forming a field prediction plant sample dataset. Performance prediction models (such as GBDT, neural networks, or crop simulation models) are fitted based on the genotypes, marker information, historical phenotypes, and target environmental variables of the field prediction samples. Prediction metrics include, but are not limited to, yield per plant, survival rate, disease resistance score, and reproductive cycle. Based on the prediction results, varieties are prioritized, breeding mating combinations, field trial arrangements, and resource allocation plans are developed. Management recommendations, crop cultivation advice, and breeding cycle plans are generated. Breeding trials are executed within the biobreeding management system, collecting real-time field feedback. The database is updated with marker performance records, optimizing subsequent breeding recommendation algorithms and completing dynamic iterations of genetic resources.
[0050] More importantly, step S435 further includes the following steps: Step S4351: performing environmental factor matching on the field prediction plant sample data, integrating field variables such as climate, soil, and irrigation, and generating multivariate field simulation input data; Step S4352: performing high-throughput field simulation modeling on the multivariate field simulation input data to generate simulated environment growth prediction data; Step S4353: performing target trait growth response calculation on the simulated environment growth prediction data, and generating target trait dynamic response data based on the trait growth curve model; Step S4354: performing time series normalization on the dynamic response data of the target trait, strengthening the periodic expression characteristics, and generating standardized phenotypic time series data; Step S4355: integrating and deducing the standardized phenotypic time series data, combining the molecular marker action mechanism, and generating integrated field progeny performance prediction data; Step S4356: Perform breeding target consistency evaluation on the integrated field offspring performance prediction data to generate field target adaptation score data; sort the breeding priority of the field prediction plant sample data according to the field target adaptation score data to generate field offspring prediction data.
[0051] In this embodiment, historical and real-time field variables, including climate data such as temperature, humidity, precipitation, light intensity, and wind speed, are acquired from the target experimental plots. Soil properties (such as pH, organic matter content, soil type, and water retention capacity) and irrigation strategy data are also acquired. These field variables are then integrated with the genetic characteristics and phenotypic performance of predicted plant samples to form a multivariate input structure centered around "sample-environmental factors." This outputs standardized, multidimensional, and structured multivariate field simulation input data, paving the way for subsequent modeling. Process-based crop models (such as DSSAT and APSIM) or data-driven models (such as LSTM and deep regression networks) can be employed. Modeling is performed on multiple samples in parallel, supporting high-throughput batch processing. Based on this integrated input data, the growth history of different plant samples under different environmental conditions (such as changes in plant height, flowering period, and fruit set) is simulated to generate predicted growth data for the simulated environment, recording the dynamic values of key agronomic traits at each growth stage. Target traits requiring key analysis, such as yield, stress resistance, and maturity, are identified. Based on simulated data, a growth curve model, such as a Sigmoid function, Gompertz model, or Piecewise linear model, is established for the target trait. The response trajectory of the trait over time and environmental changes is examined, and the dynamic response data of the target trait for each sample is output to reflect its growth adaptability and potential. The trait expression value at each time point is normalized (e.g., maximum normalization method, Z-score standardization), ensuring comparability between samples of different proportions and time lengths. Methods such as wavelet transform and Fourier transform are introduced to extract periodic expression patterns, strengthen the expressiveness of the trait during key periods, and obtain standardized phenotypic time series data, laying a temporal foundation for the next step of integrated deduction and prediction. The regulatory sites of molecular markers are matched to the target trait time series; a three-layer mapping model of "molecular markers-expression dynamics-phenotypic effects" is established. Graph convolutional networks (GCNs) or structural equation models (SEMs) are introduced to model the molecular marker pathways. The final phenotypic integration of different marker combinations in field environments is simulated to obtain the final field progeny performance prediction data for each predicted field plant, taking into account both gene regulation and environmental response. Weights and expected value ranges are set for each target trait (e.g., a 40% weight for yield, a 20% weight for maturity, etc.). A target fit score is calculated for each sample, using methods such as weighted Euclidean distance, TOPSIS, or a fuzzy composite score. The field target fit score for each predicted plant is output, indicating its degree of fit with the ideal breeding target. The predicted field plant samples are sorted in descending order, with high-priority samples marked as candidates for breeding mating or retention. The breeding management system is simultaneously updated to generate individual-level breeding operation recommendations, promoting precise management and iterative evolution.
[0052] The present invention is therefore intended to be illustrative and non-restrictive in all respects, with the scope of the invention being defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the application documents are intended to be embraced therein.
[0053] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is to be construed in the widest possible manner consistent with the principles and novel features disclosed herein.
Claims
1. A biological breeding management method based on big data, characterized in that: The following steps are involved: Step S1: Obtaining original germplasm data; integrating multi-source heterogeneous data on the original germplasm data, wherein the data integration includes genotype sequencing integration and historical phenotypic integration, generating integrated germplasm data and distributively storing it in a database; Step S2: Acquire germplasm resource images and field meteorological data and perform data preprocessing to generate standardized germplasm resource comprehensive data and synchronize them to the database; Step S3: Perform genome-wide association analysis based on the integrated germplasm data to analyze local gene sequence characteristics and phenotypic deep characteristics; perform rule-based association between local gene sequence characteristics and phenotypic deep characteristics to obtain breeding gene association rules; and screen the integrated germplasm data based on the standardized germplasm resource comprehensive data to generate a genotype-phenotype prediction model; Step S4: Based on the standardized germplasm resource comprehensive data, the integrated germplasm data is generated into a genotype-phenotype prediction model for closed-loop verification processing, and dynamic optimization breeding program data is generated. Breeding molecules are marked on the germplasm data in the database according to the dynamic optimization breeding program data, and the marked breeding molecules are observed in the field offspring performance to generate field offspring observation data; based on the preset promotion standards, the field offspring observation data is screened for variety promotion to execute the full process management of biological breeding.
2. The biological breeding management method based on big data according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: Screening germplasm phenotypic data according to breeding objectives to obtain original germplasm data; Step S12: parsing the genotype sequencing information of the original germplasm, and performing cross-platform sequencing alignment on the genotype sequencing information to generate genotype data in a unified format; Step S13: fusing historical phenotype information on the unified format genotype data to generate genotype-phenotype fusion data; Step S14: performing time series consistency calibration on the genotype-phenotype fusion data to generate time series unified integrated data; extracting multi-source features of the time series unified integrated data and performing structure mapping to generate structurally consistent integrated germplasm data; Step S15: Distribute and store the integrated germplasm data with consistent structure into the database.
3. The biological breeding management method based on big data according to claim 2, characterized in that: Step S15 includes the following steps: Step S151: indexing the integrated germplasm data with consistent structure to generate a germplasm data index table; Step S152: Divide the integrated germplasm data and the germplasm data index table into multiple data blocks, perform data segmentation and distribution, and generate distributed data blocks; Step S153: performing redundant backup of the distributed data blocks to generate redundant backup data blocks; Step S154: store the redundant backup data blocks in the distributed database.
4. The biological breeding management method based on big data according to claim 1, characterized in that: Step S2 includes the following steps: Step S21: using drones, mobile data collection apps, and IoT devices to acquire germplasm resource images and field meteorological data; Step S22: integrating germplasm resource images and field meteorological data into original germplasm growth data; Step S23: performing image preprocessing on the germplasm resource images in the original germplasm growth data to generate a crop image feature set, wherein the image preprocessing includes image filtering, image binarization, and image feature point extraction; performing data preprocessing on the field meteorological data in the original germplasm growth data to generate an environmental response parameter set, wherein the data preprocessing includes data cleaning, data denoising, missing value filling, and data standardization; Step S24: Integrate the crop image feature set and the environmental response parameter set into standardized germplasm resource comprehensive data; perform data quality detection on the standardized germplasm resource comprehensive data, and synchronously store the detected standardized germplasm resource comprehensive data in the database.
5. The biological breeding management method based on big data according to claim 1, characterized in that: The genome-wide association analysis based on the integrated germplasm data in step S3 includes: Extract single nucleotide polymorphism sites from integrated germplasm data to obtain standardized SNP feature data; Conduct stratified screening of trait stability on integrated germplasm data to generate stability trait subset data; Perform site-by-site linear regression modeling on the standardized SNP feature data and trait subset data to generate initial association signal matrix data; Perform multiple hypothesis testing correction on the initial correlation signal matrix data to generate significance screening result data; Perform Bayesian sparse representation-driven association path reconstruction on the significant screening result data to generate optimized gene-trait association network data; The optimized gene-trait association network data is subjected to model structure nested encoding to obtain genotype-phenotype association model data.
6. The biological breeding management method based on big data according to claim 4, characterized in that: The data quality detection of the standardized germplasm resource comprehensive data includes: Conduct consistency verification on the comprehensive data of standardized germplasm resources and generate consistency verification reports; Conduct data distribution analysis on standardized germplasm resource comprehensive data and generate data distribution maps and analysis reports; Conduct duplication detection on standardized germplasm resource comprehensive data and generate duplication data detection report; Conduct integrity verification on standardized germplasm resource comprehensive data and generate integrity verification reports; The consistency verification report, data distribution map and analysis report, duplicate data detection report and integrity verification report are integrated as the data quality detection results; based on the data quality detection results, the standardized germplasm resource comprehensive data after detection is synchronously stored in the database.
7. The biological breeding management method based on big data according to claim 1, characterized in that: The analysis of local gene sequence features and phenotypic deep features in step S3 includes: Extract candidate gene segments from genotype-phenotype association model data to obtain local gene sequence data; perform sequence k-mer encoding conversion on the local gene sequence data to generate gene sequence structured feature data; Extract multi-scale deep features of genotype-phenotype association model data to obtain phenotypic deep feature data; reconstruct feature space mapping between gene sequence structured feature data and phenotypic deep feature data to generate homologous nested feature collaborative data; Correlation density clustering is performed on the homologous nested feature collaboration data to generate local gene-phenotype significant coupling feature data; the local gene-phenotype significant coupling feature data is visualized and thermally encoded to obtain the local gene sequence features and phenotypic depth features of the genotype-phenotype association model data.
8. The biological breeding management method based on big data according to claim 1, characterized in that: In step S3, rule association between local features of gene sequences and deep features of phenotypes includes: Perform feature item encoding on local features of gene sequences to generate encoded gene feature item data; perform pattern discretization on phenotypic depth features to generate discrete phenotypic image feature data; Align the coded gene feature item data with the discrete phenotype image feature data to generate gene-phenotype feature pair data; perform frequent item set mining on the gene-phenotype feature pair data to generate frequent gene-phenotype combination data; Association rules are generated for frequent gene-phenotype combination data to obtain breeding gene association rules.
9. The biological breeding management method based on big data according to claim 1, characterized in that: Step S4 includes the following steps: Step S41: performing a closed-loop genetic effect verification on the genotype-phenotype prediction model based on the standardized germplasm resource comprehensive data to obtain a genetic effect verification result; Step S42: Pairing the integrated germplasm data and the standardized germplasm resource comprehensive data into breeding schemes based on the genetic effect verification results to generate dynamically optimized breeding scheme data; Step S43: marking the database for breeding molecules according to the dynamically optimized breeding scheme data, and observing the performance of offspring in the field for the marked breeding molecules to generate field offspring observation data; Step S44: Perform variety promotion screening on the field offspring observation data based on the preset promotion criteria to execute the full process management of biological breeding.
10. A biological breeding management system based on big data, characterized in that: For executing the biological breeding management method based on big data as claimed in claim 1, the biological breeding management system based on big data comprises: The genotype and phenotype data management module is used to obtain original germplasm data; integrate multi-source heterogeneous data of the original germplasm data, where data integration includes genotype sequencing integration and historical phenotypic integration, generate integrated germplasm data and store it in a distributed database; A germplasm resource management module is used to obtain germplasm resource images and field meteorological data and perform data preprocessing, generate standardized germplasm resource comprehensive data and synchronize it to the database; The gene analysis module is used to perform genome-wide association analysis on integrated germplasm data based on standardized germplasm resource comprehensive data, analyze local gene sequence characteristics and phenotypic deep characteristics; perform rule-based association between local gene sequence characteristics and phenotypic deep characteristics to obtain breeding gene association rules; and screen integrated germplasm data based on standardized germplasm resource comprehensive data to generate a genotype-phenotype prediction model; The breeding prediction and promotion module is used to perform closed-loop verification processing on the genotype-phenotype prediction model based on the comprehensive data of standardized germplasm resources, generate dynamic optimization breeding plan data, mark the breeding molecules in the database according to the dynamic optimization breeding plan data, and observe the field offspring performance of the marked breeding molecules to generate field offspring observation data; based on the preset promotion standards, the field offspring observation data is screened for variety promotion to execute the full process management of biological breeding.
Citation Information
Patent Citations
Breeding processing method and device and computer readable storage medium
CN114912040A
Unmanned agricultural precise planting intelligent management and control method
CN118095645A
Directional breeding method for improving wheat quality
CN120032710A
Latent Representations of Phylogeny to Predict Organism Phenotype
US20190130999A1
Method for efficiently optimizing a phenotype with a combination of a generative and a predictive model
WO2021217138A1
Cited By
Semen euryales breeding database construction system based on big data acquisition
CN120808904A
Intelligent navigation breeding QTN data analysis method based on large model
CN121415876A
Intelligent navigation breeding qtn data analysis method based on large model
CN121415876B