Multi-source data collaborative pear precise breeding planning large model construction method and system
By constructing a large-scale model for multi-source data-driven precision pear breeding planning, the problems of long breeding cycles and blind parent selection in traditional pear breeding have been solved. This has resulted in shorter breeding cycles, improved accuracy in parent selection, and increased utilization of heterosis, ensuring stable gene transmission and environmental adaptability.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ANHUI AGRICULTURAL UNIVERSITY
- Filing Date
- 2026-01-27
- Publication Date
- 2026-07-07
Smart Images

Figure CN121999882B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of breeding planning technology, specifically involving a method and system for constructing a large model for multi-source data collaborative precision breeding planning of pears. Background Technology
[0002] Pear is an important economic fruit crop in my country, and its variety improvement plays a crucial role in industrial development. However, traditional breeding relies on empirical parent selection. Due to the scattered storage and inconsistent standards of multi-source data (such as phenotypic, genomic, and environmental data), cross-dimensional correlation analysis cannot be achieved, resulting in a breeding cycle of 15-20 years. Furthermore, parent selection is highly unpredictable, making it difficult to fully realize hybrid vigor.
[0003] With the application of molecular biology and artificial intelligence technologies, existing methods have attempted to transition towards design breeding, but many are limited to single data dimensions or single phenotype optimization, neglecting the synergistic effects of multiple sources. For example, some technologies sacrifice data integrity to emphasize privacy protection, or lack constraint mechanisms in large-scale model applications, leading to planning bias. Furthermore, intent recognition relies on simple part-of-speech statistics, failing to meet the core needs of breeding scenarios. Consequently, it is difficult to uniformly guarantee balanced synergy among multiple phenotypes, stable gene transmission, and environmental adaptability, thus restricting breeding efficiency and variety adaptability.
[0004] Therefore, there is an urgent need for a multi-source data collaboration method based on constraint reasoning to fundamentally solve the systemic inefficiency caused by the lack of a framework. Summary of the Invention
[0005] To address the problems existing in the prior art, the purpose of this invention is to provide a large-scale model construction method and system for multi-source data collaborative precision pear breeding planning. This significantly shortens the breeding cycle, improves the accuracy of parent selection and the utilization rate of heterosis, promotes the transformation of pear breeding from experience-driven to constraint-driven, multi-source collaborative precision design, ensures stable gene transmission, and enhances the adaptability of varieties to the target production area environment.
[0006] The technical solution of this invention is:
[0007] A method for constructing a large-scale model for multi-source data-driven collaborative precision breeding planning of pears includes the following steps:
[0008] Obtain the target variety type, core improved traits, yield target, quality indicators, and stress resistance requirements;
[0009] Input the target variety type, core improved traits, yield target, quality indicators and stress resistance requirements into the breeding planning reasoning model to obtain a set of joint constraints that include gene feasible domains and multi-phenotype synergistic quantitative targets;
[0010] Based on the aforementioned set of joint constraints, germplasm resources are invoked within the framework of cross-domain feature alignment and feature-level desensitization to perform phenotypic-genotypic collaborative screening, dynamic parental matching, and multi-generational genetic evolution extrapolation, thereby generating breeding planning simulation data.
[0011] A multi-dimensional evaluation system was constructed based on the breeding planning simulation data, including gene stability, multi-phenotype synergistic consistency, environmental adaptability, and implementation feasibility, and the system was comprehensively ranked to obtain the optimal breeding planning scheme.
[0012] The method for constructing the breeding planning reasoning model includes the following steps:
[0013] Based on a multi-source collaborative dataset composed of multi-phenotype quantitative data, whole genome data, environmental data, and domain knowledge data, a five-dimensional association matrix of "variety ID-multi-phenotype-whole genome-environment-knowledge" is constructed.
[0014] Based on the aforementioned five-dimensional association matrix, a constrained knowledge representation structure is constructed. A constrained reasoning engine is introduced, and constrained-driven training is performed by combining professional corpus covering all phenotypes in the pear domain. Furthermore, a multimodal knowledge index and dynamic association retrieval mechanism are integrated to form a breeding planning reasoning model.
[0015] Preferably, the multi-source collaborative dataset includes multi-phenotype quantitative data, whole-genome data, environmental data, and domain knowledge data, wherein each phenotype in the multi-phenotype quantitative data is completely equal in terms of data structure, calling interface, and optimization status.
[0016] Preferably, the multi-phenotypic quantitative data includes sclereid characteristic data, fruit quality data, stress resistance data, and growth characteristic data, with all phenotypic data collected under consistent standards and of equal importance. Specifically, the sclereid characteristic data includes the percentage area of sclereids, equivalent particle size, and distribution density; the fruit quality data includes sugar content, fruit firmness, and fruit shape index; the stress resistance data includes drought resistance, disease resistance, and cold resistance; and the growth characteristic data includes maturity date, plant height, and the proportion of fruiting branches. The whole-genome data includes SNP sites and InDel markers obtained from resequencing, as well as candidate functional gene information related to the phenotype. The environmental data includes average annual temperature, precipitation, humidity, and soil organic matter content.
[0017] Preferably, the construction of the five-dimensional association matrix includes: storing each phenotype field and genotype field in parallel with the variety ID as the primary key, and establishing a bidirectional index for each phenotype field and genotype variation features. The bidirectional index is used to calculate the association strength of gene-phenotype cooperating units, and the same index generation rules and the same calling priority are used for each phenotype.
[0018] Preferably, the constraint-type knowledge representation structure is a rule-based structure used to constrain the boundaries of feasible relationships, and includes at least two or more of the following constraint relationship types: regulation constraint, collaboration constraint, inhibition constraint, environmental impact constraint, and production area adaptation constraint.
[0019] Preferably, the constraint-driven training includes: constructing a phenotypic coverage consistency loss function, a gene stable transmission loss function, and an environment adaptation constraint loss function, which together serve as training objectives; and the consistency loss function, gene stable transmission loss function, and environment adaptation constraint loss function are combined into a joint loss function through a weighted summation method, and optimized using a gradient descent algorithm, so as to realize that the breeding planning reasoning model simultaneously satisfies the triple constraints of gene stable transmission, multi-phenotypic balanced coordination, and environment adaptation, and the conflict between loss functions is dynamically adjusted through a constraint-type knowledge representation structure during the training process.
[0020] Preferably, the generation of the joint constraint set includes: decomposing the user breeding instructions into a target phenotype set, a target production area set, and a genetic restriction set, and generating a target threshold or target range for each target phenotype.
[0021] Preferably, the feature-level desensitization framework includes: retaining the phenotypic quantification values and key variant site information required for gene-phenotype co-screening, and masking the variety origin unit, original collection geographical coordinates, non-related genomic fragments, and identity information.
[0022] Preferably, the parental dynamic matching includes using a genetic algorithm to search for parental combinations with fitness constraints based on multi-phenotype balanced synergistic gain and stable gene transmission; the multi-generational genetic evolution extrapolation includes: under the conditions of hybridization, backcrossing and molecular marker-assisted selection, jointly updating the transmission probability of key gene loci and the probability of achieving each phenotype in each generation, and simultaneously outputting the gene homozygosity trend, genetic stability score and multi-phenotype synergistic consistency score for each generation.
[0023] Preferably, a large-scale model construction system for multi-source data collaborative pear precision breeding planning is used to implement any of the methods described above, including:
[0024] The instruction acquisition module is used to acquire the target variety type, core improved traits, yield target, quality indicators, and stress resistance requirements.
[0025] The joint constraint generation module is used to input the target variety type, core improved traits, yield target, quality index and stress resistance requirements into the breeding planning reasoning model to obtain a set of joint constraint conditions that includes gene feasible domain and multi-phenotype synergistic quantitative targets.
[0026] The collaborative screening and genetic deduction module, based on the joint constraint set, calls germplasm resources under the framework of cross-domain feature alignment and feature-level desensitization, performs phenotypic-genotypic collaborative screening, parental dynamic matching and multi-generation genetic evolution deduction, and generates breeding planning simulation data;
[0027] The evaluation and output module is used to construct a multi-dimensional evaluation system and perform comprehensive ranking on the breeding planning simulation data to obtain the optimal breeding planning scheme.
[0028] Compared with existing technologies, the large-scale model construction method and system for multi-source data collaborative pear precision breeding planning of the present invention has the following beneficial effects:
[0029] This invention addresses the lack of cross-dimensional correlation analysis caused by scattered data and inconsistent standards in traditional breeding by constructing a multi-source collaborative dataset and a five-dimensional association matrix. It fundamentally overcomes the planning bias defects caused by existing methods being limited to single data dimensions or phenotypic optimization and lacking constraint mechanisms by utilizing a constraint-based knowledge representation structure and a constraint-driven training-based constraint consistency inference model. By mapping user instructions to a joint constraint set of gene feasible domains and multi-phenotypic collaborative quantification objectives, and performing collaborative screening and genetic deduction within a cross-domain feature alignment framework, it ensures unified protection of balanced multi-phenotypic collaboration, stable gene transmission, and environmental adaptability. Finally, by combining multi-dimensional evaluation and model incremental update mechanisms, it significantly shortens the breeding cycle, improves the accuracy of parent selection and the utilization rate of heterosis, and promotes the transformation of pear breeding from experience-driven to constraint-driven, multi-source collaborative precision design, while ensuring stable gene transmission and enhancing the adaptability of varieties to the target production area environment. Attached Figure Description
[0030] Figure 1 This is a flowchart of multi-source data acquisition and standardization processing provided in an embodiment of the present invention;
[0031] Figure 2 A flowchart illustrating the generation process of the constraint-based knowledge representation structure and breeding planning reasoning model provided in this embodiment of the invention;
[0032] Figure 3 A flowchart for analyzing breeding requirements and generating a set of joint constraint conditions provided in an embodiment of the present invention;
[0033] Figure 4 A flowchart of collaborative screening and genetic deduction provided in this embodiment of the invention;
[0034] Figure 5 A flowchart illustrating the multidimensional evaluation and optimal solution output provided in this embodiment of the invention;
[0035] Figure 6 A flowchart illustrating the method architecture provided in this embodiment of the invention. Detailed Implementation
[0036] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.
[0037] Based on the embodiments of this invention, all other embodiments obtained by those skilled in the art without inventive effort are within the scope of protection of this invention.
[0038] Furthermore, the technical solutions of the various embodiments of the present invention can be combined with each other, but this must be based on the ability of those skilled in the art to implement them. When the combination of technical solutions is contradictory or cannot be implemented, it should be considered that such combination of technical solutions does not exist and is not within the scope of protection claimed by the present invention.
[0039] See Figures 1 to 6 As shown, in order to significantly shorten the breeding cycle, improve the accuracy of parent selection and the utilization rate of heterosis, promote the transformation of pear breeding from experience-driven to constraint-driven, multi-source collaborative precision design, ensure stable gene transmission, and enhance the adaptability of varieties to the target production area environment, this embodiment provides a large-scale model construction method for multi-source data collaborative precision pear breeding planning based on constraint reasoning, including the following steps:
[0040] Step 1:
[0041] like Figure 3 As shown, this step involves analyzing breeding needs and generating a set of joint constraints. It aims to transform user natural language instructions into a quantitative and complete set of joint constraints, avoiding bias from a single phenotype and providing clear guidance for subsequent breeding planning.
[0042] Specifically, this involves: acquiring user-input breeding instructions; parsing the instruction functional requirement vector through a breeding planning reasoning model; performing constraint mapping based on a constraint-based knowledge representation structure; generating a joint constraint condition set containing gene feasible domains and multi-phenotype collaborative quantification objectives; providing equally important target thresholds / intervals for each phenotype; and automatically completing phenotypic constraints not explicitly mentioned. The specific implementation process is as follows:
[0043] The breeding requirements input stage serves as the starting point of the process. Users can flexibly submit instructions via text, supporting a mix of natural language descriptions and industry-specific terminology. The input content comprehensively covers various aspects, such as... Figure 3The clearly defined core dimensions include: target variety type (such as early-maturing fresh-eating pears, late-maturing storage-resistant pears, processing pears, etc.), core improvement traits (such as targeted improvement directions such as reducing the proportion of stone cells, increasing fruit sugar content, and enhancing drought resistance), yield targets (such as quantitative expectations such as yield per acre during peak production and single fruit weight), quality indicators (such as specific standards such as fruit firmness and stone cell ratio), and stress resistance requirements (such as drought resistance level, black spot disease resistance, and low temperature tolerance threshold). Supplementary information such as target planting area, cultivation mode (open field cultivation / facility cultivation), and market application scenarios can also be provided to ensure comprehensive demand delivery.
[0044] Furthermore, requirements analysis is conducted, strictly following the logic of extracting key parameters through text semantic analysis and prioritizing requirements. First, the input instructions are systematically preprocessed using a breeding planning reasoning model: core terms, target thresholds, and scenario information are extracted through natural language processing technology, automatically filtering out redundant content such as interjections, colloquialisms, and repetitive expressions. Then, key parameters are categorized and organized, mapping the extracted information to three main categories of structured data: phenotypic requirements, environmental requirements, and genetic requirements, generating a vector of instruction functional requirements. Throughout the requirement prioritization process, the core principle of complete equality among all phenotypes is upheld, without assigning primary or secondary importance; only conflicting requirements are initially marked to avoid bias towards a single phenotype due to priority imbalances.
[0045] Constraint identification is conducted, closely linked to the results of demand analysis. It primarily focuses on four types of constraints: biological, resource, environmental, and policy constraints. A comprehensive identification is carried out using a constraint-based knowledge representation structure and a five-dimensional association matrix. Biological constraints focus on the genetic and phenotypic levels, including requirements for the stable transmission of key genes strongly associated with the target phenotype (such as functional genes controlling high sugar content), requirements for avoiding unfavorable genes (such as pathogenic genes causing fruit deformities), and constraints on synergistic / inhibitory relationships between phenotypes (such as strong synergistic constraints between fruit firmness and storage and transportation tolerance, and inhibitory constraints between plant height and the proportion of fruiting branches). Resource constraints mainly refer to limitations on the availability of germplasm resources; that is, the parental germplasm selected must come from a resource bank that has been entered into the database and completed standardized characterization to ensure that subsequent screening can be implemented. Environmental constraints are derived from environmental data of the target production area (average annual temperature, precipitation, soil organic matter content, etc.). For example, there are drought resistance constraints in arid production areas of North China and disease resistance constraints in rainy production areas of South China; policy constraints are set according to industry rules such as pear variety approval standards and germplasm resource utilization norms, such as compliance requirements for variety safety and minimum standards for phenotypic indicators, to ensure that the constraints are in line with industry policy guidance.
[0046] A joint constraint set is constructed, and the final integration is completed according to the process of constraint conflict detection, constraint weight allocation, and constraint standardization. In the constraint conflict detection stage, the system performs compatibility checks on the identified constraints by calling the synergistic / inhibition relation matrix in the constraint-type knowledge representation structure. For example, when a user simultaneously points out a strong inhibitory relationship between "early maturity" and "extra-large fruit," the system will provide conflict resolution suggestions based on historical data and genetic patterns in the five-dimensional association matrix to ensure that there are no logical contradictions between constraints. In the constraint weight allocation stage, the principle of equal weight for each phenotype is strictly adhered to, with all phenotypes having the same constraint weight ratio and no priority differences, only considering environmental adaptability and genetic stability. The system sets minimum weights for core fundamental constraints to ensure that the core breeding objectives do not deviate. During the constraint standardization phase, all constraints are transformed into quantifiable and calculable forms, generating clear thresholds or ranges for each target phenotype. For phenotypes not explicitly mentioned by the user, the system queries the industry's conventional reasonable range for that phenotype based on a five-dimensional association matrix and automatically completes the neutral constraints to ensure the completeness of multi-phenotype objectives. At the same time, the system clarifies the gene feasible domain based on the genome-wide association index, ultimately forming a set of joint constraints that includes "gene feasible domain + multi-phenotype collaborative objectives + environmental adaptation requirements + resource and policy boundaries," ensuring the comprehensiveness, compatibility, and operability of the constraints.
[0047] Step Two:
[0048] like Figure 4 As shown, this step is collaborative screening and genetic extrapolation, which is the core link in breeding planning. Through cross-domain feature alignment, collaborative screening, parent matching and multi-generation extrapolation, breeding planning simulation data that meets joint constraints is generated.
[0049] Specifically, based on the joint constraint set, germplasm resources are invoked under the framework of cross-domain feature alignment and feature-level desensitization. Classical correlation analysis is used to achieve phenotypic-genotypic collaborative screening and generate a candidate set of germplasm resources. The parent combination set is obtained by optimizing the dynamic matching of parents through genetic algorithms. Multi-generation genetic evolution is performed by combining Bayesian networks to output the gene transmission stability and multi-phenotype achievement probability of each generation, and generate breeding planning simulation data.
[0050] In the cross-domain feature alignment process, a pear breeding-specific feature dictionary is first constructed to clarify the standard definitions and quantification rules of all phenotypic indicators, genotype loci, and environmental factors (without differences in phenotypic priority), and to unify the feature descriptions and units of different institutions. Then, a domain-adaptive feature mapping algorithm is used to map the local heterogeneous features of each institution to a unified feature space, control the mapping error to a low range, and achieve cross-institutional data compatibility and interoperability.
[0051] In the feature-level desensitization process, sensitive information (variety breeding unit, original collection address, non-related genomic fragments, breeder information) is masked by a rule engine, while core breeding data such as phenotypic quantification values, key genomic sites, and environmental adaptation parameters are retained. No privacy protection noise is introduced, ensuring that the desensitized data can still be used for phenotypic-genotypic collaborative analysis.
[0052] Phenotypic-genotypic collaborative screening employs canonical correlation analysis, performed in three dimensions: phenotypic matching quantifies the similarity between germplasm phenotypes and the joint constraint condition set to screen for germplasm with high similarity; genotypic matching constructs a genotypic similarity matrix and calculates the matching degree based on the allele frequency of key genes to screen for germplasm with high matching degree; phenotypic-genotypic association matching quantifies the association strength to screen for strongly associated germplasm. The comprehensive screening of these three factors generates a candidate set of germplasm resources.
[0053] The dynamic matching of parent lines is based on a genetic algorithm. It sets reasonable population size, crossover probability, mutation probability and number of iterations, and aims at both balanced synergistic gain of multiple phenotypes and stable gene transmission as the dual fitness objectives. It calculates the gene complementarity and genetic distance between parents (controlled within a reasonable range), combines the environmental adaptability score of the target production area, and sorts them according to preset weights. It verifies the genetic diversity of the top-ranked candidate parent combinations to avoid genetic bottlenecks, and finally generates breeding parent matching data.
[0054] A multi-generational genetic evolution extrapolation model for constructing a multi-phenotype balanced and coordinated genetic recombination model was developed, with all phenotypic recombination probabilities uniformly set within a reasonable range. The specific steps are as follows: trait-driven genetic encoding reconstruction was performed on parental matching data to generate a set of parental genetic representation units; target trait-guided path mapping was performed on the joint constraint set to generate data on the evolutionary trajectory of the intended trait; intergenerational recombination probability modeling was performed using a Bayesian network model based on the two types of data; Monte Carlo simulation was conducted on the matrix for 3-5 generations to generate a path-level pedigree evolution map; trait achievement rate and gene transmission efficiency of each path were analyzed through map inversion to generate a multi-path constraint satisfaction matrix. This matrix characterizes the degree to which each path satisfies constraints on stable gene transmission, multi-phenotype balanced coordination, and environmental adaptation; paths that do not meet the constraint thresholds were eliminated based on the constraint satisfaction matrix, and a set of candidate paths that satisfy the constraints was output.
[0055] Step 3:
[0056] like Figure 5 As shown, step 5 is the multi-dimensional evaluation and optimal solution output. This step uses multi-dimensional quantitative evaluation to select the optimal breeding solution that combines theoretical advantages and practical operability.
[0057] Specifically, this involves constructing a multi-dimensional evaluation system for breeding program simulation data. This scheme establishes a four-dimensional evaluation system based on "gene stability, multi-phenotype synergistic consistency, environmental adaptability, and implementation feasibility." Furthermore, it employs an AHP-TOPSIS fusion evaluation method (AHP: Analytic Hierarchy Process, TOPSIS: Ranking Method for Approximating Ideal Solutions; the fusion of these two methods balances subjective weighting with objective data ranking, improving evaluation accuracy) to comprehensively rank the breeding program simulation data, eliminating phenotypic biases or paths that violate gene stability constraints, and outputting the optimal breeding program scheme. Detailed steps are as follows:
[0058] First, calculate the indicator values of the breeding plan simulation data item by item according to the four-dimensional evaluation system: the gene stability indicator focuses on the transmission efficiency of key gene loci, the homozygosity rate of target genes, and genetic diversity, all of which are set with quantitative standards; the multi-phenotype synergistic consistency indicator requires that the probability of all target phenotypes meeting the standard and the synergistic score between phenotypes reach the preset threshold; the environmental adaptability indicator is calculated based on the ecological niche matching degree and a scoring standard is set; the implementation feasibility indicator focuses on the breeding cycle, the selection pressure of each generation, and the complexity of field management, and controls them within a reasonable range.
[0059] The weights of the four-dimensional indicators are determined using the analytic hierarchy process (AHP). A decision matrix is then constructed, and after standardizing the values of each indicator, positive and negative ideal solutions are determined (positive ideal solutions represent the industry-optimal values for each indicator, and negative ideal solutions represent the industry-worst values). The similarity between each planned path and the positive and negative ideal solutions is calculated to determine the relative closeness. Combined with constraint consistency verification, paths with phenotypic bias or violations of gene stability constraints are eliminated. Finally, the path with a comprehensive score reaching a preset threshold is selected as the optimal solution. The solution includes detailed information such as parental combinations, hybridization methods, selection markers for each generation, field planting conditions, trait detection nodes, and expected results, facilitating implementation by breeders.
[0060] The construction of the breeding planning reasoning model is mainly completed in the following two steps:
[0061] Step 1:
[0062] 1. Obtain multi-source collaborative datasets. Multi-source collaborative datasets include multi-phenotype quantitative data, whole genome data, environmental data, and domain knowledge data. Each phenotype is completely equal in terms of data structure, calling interface, and optimization status. The multi-phenotypic quantitative data includes sclereid characteristics, fruit quality, stress resistance, and growth characteristics. All phenotypes were collected using consistent standards and were considered equally. Sclereid characteristics data were obtained using phloroglucinol staining combined with ImageJ image analysis (ImageJ is a Java-based public image processing software used for image analysis and processing), calculating the area of sclereids, equivalent particle size, and distribution density. Fruit quality data included sugar content, fruit firmness, and fruit shape index. Stress resistance data included drought resistance, disease resistance, and cold resistance. Growth characteristics data included maturity date, plant height, and proportion of fruiting branches. Whole-genome data included single nucleotide polymorphism (SNP) sites, insertion-deletion (InDel) markers, and functional genes associated with each phenotype obtained from deep whole-genome resequencing (30×) of pear. Environmental data included annual average temperature, precipitation, humidity, and soil organic matter content.
[0063] 2. The multi-source collaborative dataset is standardized for consistency. The improved KNN algorithm (k=5) is used to fill missing values. Interference is eliminated through robust standardization and outlier consistency correction. The whole genome variation features and multi-phenotype indicators are constructed into computable gene-phenotype collaborative units. Finally, a five-dimensional association matrix of "variety ID-multi-phenotype-whole genome-environment-knowledge" is constructed.
[0064] Specifically, standardization processes include:
[0065] Multiphenotypic data is analyzed using the 3σ principle to remove outliers. This method, based on the properties of normal distribution, is suitable for multiphenotypic quantitative data that follows or approximately follows a normal distribution (such as the proportion of stone cells in the area, fruit sugar content, etc.). The core logic is to define the normal range of values using the mean and standard deviation of the data, and then accurately remove extreme outliers. The specific operation process is as follows: First, calculate the mean and standard deviation of all sample data for a certain phenotypic indicator. Determine the normal data range as the interval from the mean minus 3 times the standard deviation to the mean plus 3 times the standard deviation. Then, remove samples outside the interval and identify them as outliers. Perform the above operation sequentially for all phenotypic indicators to ensure the reasonableness of the data distribution. After removing outliers, a modified KNN algorithm (k-nearest neighbor algorithm combined with phenotypic association strength weighting) was used to fill missing values. Robust standardization (RobustScaler) was then applied to eliminate dimensional differences; this standardization method is insensitive to outliers and achieves dimensional uniformity through the median and interquartile range. Whole-genome data were screened for high-quality variant sites using GATK software (a genome analysis toolkit), and functionally annotated using Annovar software (a variant annotation tool) to construct a bidirectional phenotype-genotype association index. Environmental data were weighted using the entropy weighting method and then normalized to the [0,1] interval using min-max normalization. The normalization formula is:
[0066]
[0067] in, The minimum value of the data. This sets the maximum data value to ensure that data from all dimensions can be directly analyzed collaboratively.
[0068] like Figure 1 As shown, step 1 is multi-source data collection and standardization processing. This step serves as the foundation of the entire breeding program and aims to construct a high-quality, structured five-dimensional correlation matrix to provide data support for subsequent constraint reasoning.
[0069] Furthermore, the collection of multi-source collaborative datasets covers four core data categories and ensures that each phenotypic data is of equal status, including: multi-phenotypic quantitative data, whole genome data, environmental data, and domain knowledge data.
[0070] Multiphenotypic quantitative data collection must adhere to a unified accuracy standard: When collecting data on stone cell characteristics, undamaged tissue from the middle of pear pulp is selected to prepare continuous sections. After staining with phloroglucinol and mounting, the sections are magnified and photographed under an optical microscope. Using ImageJ software, an adaptive threshold segmentation algorithm is employed to accurately extract the stone cell outlines. The size and pixel ratio are calibrated using a section scale, and the area, equivalent particle size, and distribution density of stone cells are calculated. Each sample is measured three times, and the average value is taken to control the error within a reasonable range. In terms of fruit quality data, sugar content is measured using a handheld saccharimeter at three or more sites on the equator of the fruit, mixed juice. Fruit firmness is measured using a texture analyzer at symmetrical sites, and the average value is taken. Fruit shape index is calculated by manually measuring the longitudinal and transverse diameters. Stress resistance data is obtained through simulated adverse stress (such as drought resistance through drought treatment and disease resistance through artificial inoculation with pathogens) or natural field identification. Growth characteristic data is obtained through fixed-point and timed field observations and records, covering key nodes throughout the entire growth period.
[0071] Whole-genome data were deep resequencing using a high-throughput sequencing platform. After quality control, alignment, sorting and deduplication processes, the raw data were used for variant detection and screening using professional software to retain high-quality SNP sites and InDel markers. After functional annotation, a bidirectional association index of "phenotype-genotype" was constructed. The index generation rules were consistent for all phenotypes to ensure consistent priority in access.
[0072] Environmental data are derived from authoritative meteorological data centers (average daily temperature, precipitation, and humidity over the past 10 years) and field soil measurements (soil organic matter content and pH value).
[0073] Domain knowledge data is collected from pear breeding monographs, approval standards, core journal articles and practical experience records, and structured knowledge items (including phenotypic quantification standards, genetic laws, environmental adaptation requirements, etc.) are extracted using text mining tools.
[0074] In the data standardization stage, outliers in multi-phenotypic data are first removed using outlier removal rules, and then the improved KNN algorithm is used to fill in missing values. Compared with the traditional KNN, the algorithm is weighted by quantifying the phenotypic association strength between the sample to be filled and its neighboring samples to ensure that the missing value filling is more in line with the phenotypic association rules. After filling, robust standardization is performed to effectively reduce the residual impact of outliers.
[0075] After screening and annotation, the whole genome data is used to construct a bidirectional "phenotype-genotype" association index, which matches the corresponding functional genes and SNP sites for each phenotype to ensure index query efficiency. The environmental data uses the entropy weight method to calculate the weight of each factor (the smaller the entropy value, the greater the weight), and then normalizes it to a uniform interval to ensure that the weight allocation of each environmental factor is objective.
[0076] Finally, using "Variety ID" as the primary key, we store each phenotype field, genotype field, environment field, and knowledge entry in parallel to construct a five-dimensional association matrix. This matrix is stored in a professional database with dual backups. At the same time, we generate a gene-phenotype synergistic unit association strength matrix (quantifying the association coefficient between any phenotype and genotype), providing a foundation for the subsequent construction of a constrained knowledge structure.
[0077] Step Two:
[0078] like Figure 2 As shown, step 2 involves generating a constrained knowledge representation structure and a breeding planning reasoning model. This step aims to construct a reasoning engine subject to triple constraints, providing core technical support for breeding intention analysis and planning generation. Its main components include: constructing a constrained knowledge representation structure based on a five-dimensional association matrix; extracting four core entities: phenotype, genotype, environmental factors, and rules; defining four types of constraint relationships: "regulation," "synergy," "inhibition," and "adaptation"; introducing a large model as a constrained reasoning engine; and conducting constraint-driven training using a professional corpus that evenly covers all phenotypes in the pear field. This ensures the model satisfies the triple constraints of stable gene transmission, balanced synergy among multiple phenotypes, and environmental adaptation. Finally, it integrates a multimodal knowledge index and a dynamic association retrieval mechanism to generate a breeding planning reasoning model. The specific implementation process is as follows:
[0079] The data input layer uses a five-dimensional correlation matrix formed by "variety ID-multiphenotype-whole genome-environment-knowledge" as its core input, covering phenotypic quantitative data, whole genome data (single nucleotide polymorphism (SNP) sites, insertion-deletion (InDel) markers, functional genes), environmental data (annual mean temperature, precipitation, etc.) and domain knowledge data (breeding standards, genetic laws), providing complete and standardized basic data support for subsequent stages.
[0080] The construction of a constrained knowledge representation structure (knowledge graph) follows a process of "entity definition → relation extraction → attribute annotation." Based on a five-dimensional association matrix, four core entities are extracted: phenotypic entities (multi-phenotypic indicators such as the proportion of stone cells and sugar content), genotype entities (SNP sites, InDel markers, and functional genes associated with the phenotype), environmental factor entities (annual mean temperature, precipitation, soil organic matter content, etc.), and rule entities (breeding quantitative standards, genetic transmission laws, etc.). Each entity is assigned a unique identifier to ensure traceability. Simultaneously, four core constraint relationships are extracted and defined: regulatory relationships (positive or negative regulatory effects of genes on phenotypes), synergistic relationships (positive associations between phenotypes), inhibitory relationships (negative constraints between phenotypes), and adaptation relationships (matching relationships between environmental factors and phenotypes / genotypes), clarifying the feasible boundaries between entities. Furthermore, quantitative attributes are added to entities and relationships. Entity attributes include phenotypic collection standards and genotype variation frequencies, while relationship attributes cover the association strength of synergistic / inhibitory relationships and the efficiency of regulatory relationships, constructing a complete entity-relationship-attribute system.
[0081] Constraint Database Establishment: A structured constraint database is constructed based on the classification logic of "explicit constraints → implicit constraints → dynamic constraints" to ensure that constraint relationships are callable and updatable. Explicit constraints refer to clearly defined and directly enforceable rigid constraints that can be applied without derivation, including breeding rule constraints (such as phenotypic quantitative standards and genetic restrictions for variety approval), regulatory relationship constraints (explicit regulatory rules of genes on phenotypes), and rigid environmental adaptation constraints (environmental thresholds for target production areas). Implicit constraints are potential constraints that need to be derived based on data association analysis and are obtained through quantitative model mining, covering implicit constraints of phenotypic synergy (phenotypic associations that are not explicitly defined but actually exist) and implicit constraints of genotype interaction (potential synergy / inhibition between functional genes). Dynamic constraints are flexible constraints that are dynamically adjusted according to breeding scenarios and data updates, including dynamic environmental constraints (adjustments to adaptation thresholds caused by interannual changes in environmental data of target production areas), dynamic genetic constraints (constraint updates caused by changes in gene homozygosity rates during multi-generational evolution), and dynamic retrieval constraints (dynamic adjustments based on the strength of entity associations in a five-dimensional association matrix).
[0082] Inference Model Training: A large model is introduced as the constrained inference engine. The basic model selected is ChatGLM3-6B (a lightweight, high-performance open-source dialogue model). A dedicated corpus for multi-phenotype collaboration in pear breeding is constructed, with the corpus evenly distributed according to the number of phenotypes (without any single phenotype bias). It covers phenotype collaboration terminology, breeding technology literature fragments, and multi-phenotype collaboration instruction samples. The training objective focuses on the three constraints of stable gene transmission, balanced multi-phenotype collaboration, and environmental adaptation. Correspondingly, three types of loss functions are constructed: phenotype coverage consistency loss function, gene stable transmission loss function, and environmental adaptation constraint loss function, to ensure that the model meets the preset accuracy requirements in all three dimensions. During training, the constraint database is called in real time, and explicit / implicit / dynamic constraints are embedded into the model's inference logic to avoid the model outputting results that violate the constraints.
[0083] The training mechanisms and coordination methods of the three loss functions are as follows:
[0084] The training mechanism of the three-loss function is a joint optimization process that uses gradient descent to collaboratively update model parameters, ensuring the unified satisfaction of the three constraints. The specific mechanism is as follows:
[0085] Joint loss function construction: During training, the three loss functions are combined into a total loss function (joint loss function) through weighted summation.
[0086] ,
[0087] Among them, weight , , The training objective is to minimize the total loss, based on domain prior knowledge (e.g., multi-phenotypic collaborative weights may be slightly higher to emphasize balance). For the joint loss function, To cover the consistency loss function, For the loss function of stable gene transmission, The loss function is constrained by environmental adaptation. This design allows the model to simultaneously consider multiple phenotypic coverage, genetic stability, and environmental adaptation during backpropagation, avoiding bias towards any one dimension.
[0088] The coordination mechanism of the three loss functions:
[0089] Parallel computation and gradient fusion: During training on each batch of data, the gradients of the three loss functions are computed in parallel. The optimizer (such as Adam) updates the model parameters based on gradient weights, enabling the model to automatically balance different constraints during inference. For example, when user instructions emphasize disease resistance, gene stability loss and environmental adaptation loss work together to ensure the stable transmission of disease-resistant genes without violating the humidity constraints of the production area.
[0090] Constraint Prioritization: During training, if there are conflicts among the three loss functions (e.g., increasing sugar content may inhibit premature maturation), the model will dynamically adjust the gradient direction through a constraint-based knowledge representation structure (e.g., synergistic / inhibition relationships), prioritizing the satisfaction of rigid constraints (e.g., gene stability). This mechanism stems from the integration of a "constraint database," ensuring that training conforms to biological principles.
[0091] Validation and tuning: After training, the model is evaluated on the validation set to assess the satisfaction rate of the triple constraints (such as requiring a phenotypic prediction accuracy of ≥93%). Only when all three constraints are met can the model be deployed, forming a closed-loop optimization.
[0092] Role in the overall scheme: This training mechanism enables the breeding planning reasoning model to become a "constrained reasoning engine," which can accurately map user instructions into a set of joint constraints, providing a reliable foundation for subsequent parent matching and genetic deduction.
[0093] The feature fusion layer transforms multi-source data into a unified feature space, enabling efficient fusion of constraints and knowledge. Specifically, it constructs a multimodal knowledge index library, converting phenotypic data, genotypic data, environmental data, and knowledge data into vector representations, establishing a one-to-one mapping with entities in the knowledge graph and constraints in the constraint database. Simultaneously, it builds a dynamic association retrieval mechanism, prioritizing the retrieval of entity knowledge and constraints with high association strength based on the entity association strength of the five-dimensional association matrix, generating dynamic knowledge call paths, and ensuring the targeted and efficient nature of feature fusion.
[0094] The constraint reasoning layer integrates the three types of constraints from the constraint database, the entity-relationship system of the knowledge graph, and the multimodal retrieval mechanism, enabling the model to have constraint-driven reasoning capabilities. Its reasoning logic is as follows: after receiving the instruction, it first calls the explicit constraints from the constraint database for preliminary screening, then supplements the derivation with implicit constraints, and finally adjusts the reasoning direction by combining dynamic constraints to ensure accurate parsing of multi-phenotype collaborative needs, avoid single-phenotype bias, and simultaneously satisfy the triple constraints of stable gene transmission, balanced multi-phenotype collaboration, and environmental adaptation.
[0095] Finally, by integrating the constrained knowledge representation structure with the multimodal retrieval mechanism, a breeding planning reasoning model is generated. The model needs to be validated: the parsing accuracy of multi-phenotype collaborative instructions, the satisfaction rate of gene stability transmission constraints, the satisfaction rate of multi-phenotype balanced collaborative constraints, and the satisfaction rate of environmental adaptation constraints all need to meet the preset standards.
[0096] This invention also includes a breeding planning reasoning model update step:
[0097] For dynamic model evolution: This step establishes a closed-loop optimization mechanism of "feedback-sample-training-iteration" to ensure that the breeding planning inference model continuously adapts to new data and actual breeding needs, and maintains constraint consistency.
[0098] Specifically, this involves: collecting field full-phenotype validation data, genome validation data, and user feedback data; constructing an incremental sample set according to the format of "user instructions - multi-source knowledge - breeding plan - full-phenotype validation data - genome validation data - feedback"; and using knowledge distillation combined with LoRA fine-tuning to incrementally train the breeding plan reasoning model, thereby completing the update and deployment of the breeding plan reasoning model and maintaining consistent satisfaction of the triple constraints.
[0099] Feedback data is collected through multiple channels: user feedback data includes satisfaction ratings and adjustment suggestions for breeding programs; field full-phenotype identification data includes quantitative values of all phenotypes measured in each generation and their compliance status; genome validation data includes measured values of target gene homozygosity and key SNP site transmission efficiency, ensuring the authenticity and relevance of the data.
[0100] The incremental sample set is constructed in a fixed format, each sample contains core fields, and the sample size reaches the preset standard to ensure balanced coverage of the entire table of data without any single table bias.
[0101] Incremental training employs a strategy combining knowledge distillation and low-rank adaptation fine-tuning: knowledge distillation uses a high-performance large model as the teacher model and the fine-tuned ChatGLM3-6B as the student model to achieve knowledge transfer; low-rank adaptation fine-tuning only updates the low-rank adaptation matrix of the model and freezes the main parameters of the basic model, significantly reducing computational power consumption, and a single round of training can be completed on a regular GPU server.
[0102] Training parameters were set to reasonable learning rates, batch sizes, and number of iterations, and an appropriate optimizer was used. Model validation employed cross-validation and independent test set validation (with a test set sample size meeting preset standards). Validation standards included: all phenotypic prediction accuracies, genomic locus transmission prediction accuracies, and triple constraint satisfaction rates all meeting preset requirements. Once these standards were met, the model was updated and deployed, with a fixed update cycle set to enable dynamic model evolution and continuously improve the accuracy and adaptability of breeding plans.
[0103] This invention provides a method for constructing a large-scale model for multi-source data collaborative precision pear breeding planning based on constrained reasoning. Using a five-dimensional association matrix as the data core, it innovatively constructs a constrained knowledge representation structure and a constrained large-scale model reasoning mechanism, completely avoiding planning biases dominated by a single phenotype or gene, and achieving a unified approach of equal collaboration among multiple phenotypes and stable gene transmission. Through constraint mapping, user requirements are transformed into a quantified set of joint constraint conditions, ensuring both the integrity of phenotypic goals and clarifying gene transmission rules. A feature-level desensitization framework ensures data security while preserving the integrity of core breeding data, abandoning redundant designs of federated learning and differential privacy. Dynamic parental matching and multi-generational extrapolation are integrated with genetic algorithms and Bayesian networks, combined with AHP-TOPSIS fusion evaluation, significantly improving the scientific rigor and feasibility of breeding programs. Knowledge distillation combined with a LoRA fine-tuning model update strategy reduces computational consumption while maintaining constraint consistency, promoting the transformation of pear breeding from "single-trait optimization" to "multi-phenotype collaborative evolution."
[0104] The above inventions are merely a few specific embodiments of the present invention. However, the embodiments of the present invention are not limited thereto, and any variations that can be conceived by those skilled in the art should fall within the protection scope of the present invention.
[0105] Furthermore, to achieve the above steps, this invention provides a multi-source data collaborative pear precision breeding planning system based on constrained reasoning, comprising:
[0106] The multi-source data acquisition and standardization module is configured to: acquire multi-phenotype quantitative data, whole genome data, environmental data, and domain knowledge data; perform improved KNN (k-nearest neighbor algorithm combined with phenotypic association strength weighting) filling, robust standardization, and outlier consistency correction processing on the premise that all phenotypes are completely equal; construct a five-dimensional association matrix of "variety ID-multi-phenotype-whole genome-environment-knowledge"; and generate gene-phenotype co-operating units and their associated indexes, ensuring that the generation rules and calling priorities of each phenotype index are consistent.
[0107] The constrained knowledge representation and reasoning model construction module is configured to: extract four core entities—phenotype, genotype, environmental factors, and rules—based on a five-dimensional association matrix; define four types of constraint relationships—"regulation," "synergy," "inhibition," and "adaptation"—to construct a constrained knowledge representation structure; introduce a large model as a constrained reasoning engine, and conduct constraint-driven training by combining professional corpora that evenly cover all phenotypes in the pear field, so that it meets the triple constraints of stable gene transmission, balanced synergy of multiple phenotypes, and environmental adaptation; and integrate multimodal knowledge indexing and dynamic association retrieval mechanisms to generate a breeding planning reasoning model (the large model does not directly output breeding decision results).
[0108] The instruction acquisition module is used to receive breeding instructions input by the user via text, including the target variety type, core improved traits, yield target, quality indicators, and stress resistance requirements.
[0109] The joint constraint generation module is configured to: input user instructions (target variety type, core improved traits, yield target, quality indicators, and stress resistance requirements) into the breeding planning reasoning model and generate an instruction functional requirement vector; perform constraint mapping based on a constraint-based knowledge expression structure, decompose the instructions into a target phenotype set, a target production area set, and a genetic restriction set; generate equally important target thresholds or intervals for each phenotype, automatically complete phenotype-neutral constraints or minimum constraints not explicitly mentioned, and output a set of joint constraint conditions including gene feasible domains, multi-phenotype synergistic targets, and environmental adaptation requirements.
[0110] The collaborative screening and genetic inference module is configured to: invoke germplasm resources within a cross-domain feature alignment and feature-level desensitization framework; perform collaborative screening using canonical correlation analysis to match phenotypes, genotypes, and phenotype-genotype associations, generating a candidate set of germplasm resources; optimize dynamic parental matching using genetic algorithms with dual constraints of balanced synergistic gain across multiple phenotypes and stable gene transmission, outputting a set of parental combinations after genetic diversity verification; construct a balanced synergistic genetic recombination model using a Bayesian network, perform 3-5 generation genetic evolution inference, and simultaneously output gene transmission stability and multi-phenotype achievement probabilities for each generation, generating breeding planning simulation data.
[0111] The evaluation and output module is configured to: perform multi-dimensional quantitative evaluation of breeding program simulation data based on a four-dimensional evaluation system of "gene stability - multi-phenotype synergistic consistency - environmental adaptability - implementation feasibility"; determine the weight of indicators using the analytic hierarchy process (AHP), calculate the comprehensive score using the TOPSIS (Topology of Ideal Solutions) method, and eliminate phenotypic biased or genetically unstable paths through constraint consistency checks; and output the optimal breeding program scheme, which includes parental combinations, hybridization methods, selection markers for each generation, field planting conditions, trait detection nodes, and expected results.
[0112] The model update module is configured to: collect field full-phenotype validation data, genome validation data, and user feedback data; construct an incremental sample set in the format of "user instructions - multi-source knowledge - breeding plan - full-phenotype validation data - genome validation data - feedback"; update the breeding plan inference model using knowledge distillation combined with a low-rank fitting fine-tuning strategy, using Llama 2-13B as the teacher model to achieve knowledge transfer, updating only the model's low-rank fitting matrix, and freezing the main parameters of the basic model; after multi-dimensional validation (prediction accuracy of all phenotypes ≥93%, prediction accuracy of genomic locus transfer ≥91%), complete the model update deployment, maintaining its consistent satisfaction of the triple constraints.
[0113] In this embodiment, the system constructs a complete technical chain from data integration, constraint modeling, intent parsing, plan generation, evaluation and screening to model evolution through the coordinated linkage of six major modules. The functions of each module are progressive and mutually supportive, ensuring the consistency of constraints, equality of multiple phenotypes and gene stability of breeding plans.
[0114] The multi-source data acquisition and standardization module, as the core of the system, undertakes the function of comprehensive coverage and standardized integration of multi-dimensional data. Multi-phenotypic quantitative data are collected through professional testing (staining + image analysis for stone cell characteristics), field observation (growth characteristics, stress resistance), and instrumental measurements (fruit quality), ensuring consistent standards and equal status for each phenotypic data. Whole-genome data is acquired through a high-throughput sequencing platform, and after quality control, comparison, and screening, high-quality single nucleotide polymorphism (SNP) sites, insertion-deletion (InDel) markers, and functional genes are retained to construct a bidirectional "phenotype-genotype" association index. Environmental data integrates authoritative meteorological data (temperature, humidity, and precipitation over the past 10 years) and field soil measurement data (organic matter content and pH value), with weights allocated using the entropy weighting method. Domain knowledge data is collected from breeding monographs, approval standards, core literature, and practical experience, and structured knowledge entries are extracted through text mining. In the standardization stage, multi-phenotype data are filled with missing values using an improved KNN algorithm and robust standardization is used to eliminate dimensional differences. After functional annotation, whole-genome data are used to generate gene-phenotype cooperating units. Environmental data are normalized to the [0,1] interval. Finally, a five-dimensional association matrix is constructed with "variety ID" as the primary key. The data is stored in a professional database with dual backups to ensure data security and support rapid association and retrieval of multi-source data.
[0115] The constrained knowledge representation and reasoning model construction module is the core reasoning support of the system, focusing on constrained modeling and large-scale model adaptation. The constrained knowledge representation structure is based on a five-dimensional association matrix. Core entities cover all phenotypic indicators, genotype loci, environmental factors, and breeding rules. Four types of constraint relationships clearly define the feasible boundaries between entities: "regulation" relationships limit the logic of gene effects on phenotypes; "synergy" and "inhibition" relationships quantify the strength of associations between phenotypes; and "adaptation" relationships define the matching rules between environment and variety. Furthermore, no phenotypic dominance priority is set for any of these relationships. Large-scale model constraint-driven training is based on ChatGLM3-6B (a lightweight, high-performance open-source dialogue model, developed jointly by the Tsinghua University team and Zhipu AI, representing the 3rd generation of a dialogue generation language model with 6 billion parameters). A dedicated corpus evenly covers terminology, technical literature, and collaborative instruction samples related to each phenotype. The training process incorporates three loss functions: phenotypic coverage consistency, stable gene transmission, and environmental adaptation, ensuring that the model strictly adheres to the constraint boundaries during reasoning. The multimodal knowledge index transforms various types of data into vector representations, and the dynamic association retrieval mechanism prioritizes calling highly relevant knowledge based on the strength of entity association, ultimately generating a breeding planning reasoning model.
[0116] The instruction parsing and joint constraint generation module serves as the linchpin connecting user needs and system functions, transforming natural language instructions into quantifiable constraints. This module allows users to submit instructions in natural language or industry terminology. First, it extracts core terms, target thresholds, and scenario information using word segmentation tools, removing redundant content to generate structured parsing results. Then, based on a constraint-based knowledge representation structure, it performs constraint mapping, breaking down the instructions into three sets: a target phenotype set clarifying the phenotypes and thresholds to be improved; a target production area set defining the environmental adaptability range; and a genetic restriction set clarifying the requirements for stable gene transmission. For phenotypes not explicitly mentioned by the user, the system automatically completes neutral constraints based on a five-dimensional correlation matrix (maintaining industry-standard levels) to ensure the completeness of multi-phenotype targets. The final generated set of joint constraints sets equally important achievement standards for each phenotype and clarifies the transmission or avoidance requirements for key genes, thus avoiding planning biases dominated by a single phenotype or gene from the outset.
[0117] The collaborative screening and genetic deduction module is the core execution unit of the system, realizing the entire process planning from germplasm screening to path simulation. Cross-domain feature alignment unifies the data representation and units of different institutions by constructing a dedicated feature dictionary, and uses a domain-adaptive mapping algorithm to map heterogeneous features to a unified space with a mapping error ≤3%. Feature-level desensitization only masks sensitive information such as variety breeding units and collection coordinates, while retaining core data such as phenotypic quantification values and key genotype loci, without introducing privacy noise. Phenotypic-genotype collaborative screening is executed in three dimensions: phenotypic matching quantifies the similarity between germplasm and constraints, genotype matching calculates the allele frequency matching degree of key genes, and phenotypic-genotype association matching screens strongly associated germplasm. The three are combined to generate a candidate set. Parental dynamic matching is based on a genetic algorithm, setting reasonable population size and iteration parameters, and using multi-phenotypic synergistic gain and stable gene transmission as dual fitness constraints. It calculates gene complementarity and genetic distance between parents, and combines environmental adaptability scores for weighted sorting. After genetic diversity verification (Shannon index ≥ 0.35), the parental combination is output. A multi-generational genetic evolution extrapolation model is constructed to achieve balanced and coordinated genetic recombination of multiple phenotypes. The probability of intergenerational recombination is modeled using Bayesian networks. After 3-5 generations of simulation extrapolation, the gene transmission stability score and the probability of achieving multiple phenotypes in each generation are output, generating breeding planning simulation data.
[0118] The evaluation and output module selects the optimal breeding scheme through multi-dimensional evaluation and constraint verification. Each indicator in the four-dimensional evaluation system has clearly defined standards: the gene stability indicator focuses on the efficiency of key gene transmission, the homozygosity of target genes, and genetic diversity; the multi-phenotype synergy consistency indicator requires that the probability of all target phenotypes meeting the standard be ≥0.92 and the synergy score between phenotypes be ≥0.85; the environmental adaptability indicator is calculated based on niche matching degree, with a score ≥0.83; and the feasibility indicator controls the breeding cycle to ≤5.5 years and the selection pressure to ≤0.3. During the evaluation process, the weights of the indicators are first determined using the analytic hierarchy process (gene stability 27%, multi-phenotype synergy 40%, environmental adaptability 22%, and feasibility 11%), then the comprehensive score is calculated using the TOPSIS method. Constraint consistency verification is performed simultaneously to eliminate paths with phenotypic bias or violations of gene stability constraints. Finally, schemes with a comprehensive score ≥85 are selected for output. These schemes include complete breeding process details, making them easy for breeders to implement directly.
[0119] The model update module constructs a closed-loop optimization mechanism of "feedback-sample-training-iteration" to ensure the long-term adaptability of the system. Feedback data is collected from multiple channels, covering user satisfaction ratings, full phenotypic data from field measurements, and genome validation data, ensuring data authenticity and coverage of all phenotypes. The incremental sample set is constructed according to a fixed format, with ≥1000 samples per update, ensuring balanced coverage of all phenotypes. Model fine-tuning employs knowledge distillation combined with LoRA (low-rank adaptation) strategy, using a high-performance large model as a teacher model to achieve knowledge transfer, updating only the low-rank adaptation matrix (rank 8), freezing the main parameters of the basic model, and significantly reducing computational consumption. After fine-tuning, multi-dimensional validation ensures that the prediction accuracy of all phenotypes is ≥93% and the prediction accuracy of genomic locus transfer is ≥91%. After achieving these targets, an update and deployment is completed every 3 months to continuously maintain the model's consistent satisfaction of the triple constraints, improving the accuracy and adaptability of breeding plans.
[0120] The six modules mentioned above support and coordinate with each other to form a complete technical closed loop: the multi-source data acquisition and standardization module provides a high-quality data foundation; the constraint-based knowledge representation module constructs a constraint reasoning framework; the instruction parsing module transforms and quantifies constraint objectives; the collaborative screening module generates targeted breeding paths; the evaluation module selects the optimal solution; and the model update module ensures long-term adaptability. The entire chain is centered on "constraint reasoning" and ensured by "equal synergy among multiple phenotypes" and "stable gene transmission," completely avoiding planning biases dominated by a single phenotype or gene. This drives the transformation of pear breeding from "single-trait optimization" to "multi-phenotype co-evolution," significantly improving the accuracy, success rate, and environmental adaptability of variety breeding.
[0121] In summary, the present invention has the following technical advantages:
[0122] The large-scale model construction method and system for multi-source data collaborative pear precision breeding planning provided in this invention constructs a constraint-type knowledge expression structure based on a five-dimensional correlation matrix. Through constraint-driven training, the large model is transformed into an inference engine subject to triple constraints of gene-phenotype-environment, completely avoiding planning bias dominated by a single phenotype or a single gene, and achieving the unity of equal collaboration of multiple phenotypes and stable gene transmission.
[0123] By leveraging improved KNN filling and robust normalization consistency processing techniques, gene-phenotype collaborative units and bidirectional indexes are constructed to ensure equal status for each phenotype at the data level, resolving the phenotypic weight imbalance problem in traditional breeding. A constraint mapping mechanism transforms user requirements into a joint constraint set, automatically completing phenotypic constraints and avoiding misinterpretation of intent. Simultaneously, a feature-level anonymization framework preserves the integrity of core breeding data while ensuring data security, eliminating the redundant design and accuracy loss associated with federated learning and differential privacy.
[0124] Dynamic parental matching employs dual fitness constraints of balanced multi-phenotype coordination and stable gene transmission. Multi-generational genetic evolution is used to simultaneously track gene transmission and phenotypic achievement. A multi-dimensional evaluation system rigorously eliminates planning paths that violate constraints, significantly improving the long-term reliability of breeding programs. A model update strategy combining knowledge distillation and low-rank adaptation fine-tuning is used to reduce computational consumption while maintaining consistent model compliance with constraints, continuously adapting to new data and actual breeding needs.
[0125] The overall technical solution breaks through the limitations of existing technologies that prioritize privacy protection, optimize single traits, and focus on genotype variation deconstruction. It promotes the transformation of pear breeding from "experience-driven" and "single trait optimization" to "constraint-driven" and "multi-phenotype co-evolution," significantly shortening the breeding cycle and improving the accuracy, success rate, and environmental adaptability of variety breeding.
[0126] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A method for constructing a large-scale model for multi-source data-driven collaborative precision breeding planning of pears, characterized in that: Includes the following steps: Obtain the target variety type, core improved traits, yield target, quality indicators, and stress resistance requirements; Input the target variety type, core improved traits, yield target, quality indicators and stress resistance requirements into the breeding planning reasoning model to obtain a set of joint constraints that include gene feasible domains and multi-phenotype synergistic quantitative targets; Based on the aforementioned set of joint constraints, germplasm resources are invoked within the framework of cross-domain feature alignment and feature-level desensitization to perform phenotypic-genotypic collaborative screening, dynamic parental matching, and multi-generational genetic evolution extrapolation, thereby generating breeding planning simulation data. The feature-level desensitization framework includes: retaining phenotypic quantification values and key variant site information required for gene-phenotype co-screening, while masking the variety's origin unit, original collection geographical coordinates, unrelated genomic fragments, and identification information; the parental dynamic matching includes using a genetic algorithm to search for parental combinations with multi-phenotype balanced synergistic gain and stable gene transmission as fitness constraints; the multi-generational genetic evolution extrapolation includes: under hybridization, backcrossing, and marker-assisted selection conditions, jointly updating the key gene locus transmission probability and the probability of each phenotype achieving the target for each generation, and simultaneously outputting the gene homozygosity trend, genetic stability score, and multi-phenotype synergistic consistency score for each generation; A multi-dimensional evaluation system was constructed based on the breeding planning simulation data, including gene stability, multi-phenotype synergistic consistency, environmental adaptability, and implementation feasibility, and the system was comprehensively ranked to obtain the optimal breeding planning scheme. The method for constructing the breeding planning reasoning model includes the following steps: Based on a multi-source collaborative dataset composed of multi-phenotypic quantitative data, whole-genome data, environmental data, and domain knowledge data, a five-dimensional association matrix of "variety ID-multi-phenotypic-whole-genome-environment-knowledge" is constructed. The multi-phenotypic quantitative data includes sclereid characteristic data, fruit quality data, stress resistance data, and growth characteristic data. All phenotypic data are collected according to consistent standards and are of equal importance. Among them, sclereid characteristic data includes sclereid area, equivalent particle size, and distribution density; fruit quality data includes sugar content, fruit firmness, and fruit shape index; stress resistance data includes drought resistance, disease resistance, and cold resistance; and growth characteristic data includes maturity period, plant height, and proportion of fruiting branches. The whole-genome data includes SNP loci and InDel markers obtained from resequencing, as well as candidate functional gene information related to the phenotype. The environmental data includes annual average temperature, precipitation, humidity, and soil organic matter content. Based on the aforementioned five-dimensional association matrix, a constrained knowledge representation structure is constructed. A constrained reasoning engine is introduced, and constrained-driven training is performed by combining professional corpora that evenly cover various phenotypes in the pear field. Furthermore, a multimodal knowledge index and a dynamic association retrieval mechanism are integrated to form a breeding planning reasoning model.
2. The method for constructing a large-scale model for multi-source data collaborative pear precision breeding planning according to claim 1, characterized in that, The multi-source collaborative dataset includes multi-phenotype quantitative data, whole genome data, environmental data, and domain knowledge data. Among them, each phenotype in the multi-phenotype quantitative data is completely equal in terms of data structure, calling interface, and optimization status.
3. The method for constructing a large-scale model for multi-source data collaborative pear precision breeding planning according to claim 1, characterized in that, The construction of the five-dimensional association matrix includes: storing each phenotype field and genotype field in parallel with the variety ID as the primary key, and establishing a bidirectional index for each phenotype field and genotype variation features. The bidirectional index is used to calculate the association strength of gene-phenotype co-operating units, and the same index generation rules and the same calling priority are used for each phenotype.
4. The method for constructing a large-scale model for multi-source data collaborative pear precision breeding planning according to claim 1, characterized in that, The constrained knowledge representation structure is a rule-based structure used to constrain the boundaries of feasible relationships, and includes at least two or more of the following constraint relationship types: regulation constraint, collaboration constraint, inhibition constraint, environmental impact constraint, and production area adaptation constraint.
5. The method for constructing a large-scale model for multi-source data collaborative precision breeding planning of pears according to claim 1, characterized in that, The constraint-driven training includes: constructing a phenotypic coverage consistency loss function, a gene stable transmission loss function, and an environment adaptation constraint loss function, which together serve as training objectives; and combining the consistency loss function, gene stable transmission loss function, and environment adaptation constraint loss function into a joint loss function through a weighted summation method, and optimizing it using a gradient descent algorithm, so as to achieve the breeding planning reasoning model simultaneously satisfying the triple constraints of gene stable transmission, multi-phenotypic balanced coordination, and environment adaptation, and dynamically adjusting the conflicts between loss functions through a constraint-type knowledge representation structure during training.
6. The method for constructing a large-scale model for multi-source data collaborative pear precision breeding planning according to claim 1, characterized in that, The generation of the joint constraint set includes: decomposing the user breeding instructions into a target phenotype set, a target production area set, and a genetic restriction set, and generating a target threshold or target range for each target phenotype.
7. A large-scale model construction system for multi-source data collaborative pear precision breeding planning, used to implement the method described in any one of claims 1-6, characterized in that, include: The instruction acquisition module is used to acquire the target variety type, core improved traits, yield target, quality indicators, and stress resistance requirements. The joint constraint generation module is used to input the target variety type, core improved traits, yield target, quality index and stress resistance requirements into the breeding planning reasoning model to obtain a set of joint constraint conditions that includes gene feasible domain and multi-phenotype synergistic quantitative targets. The collaborative screening and genetic deduction module, based on the joint constraint set, calls germplasm resources under the framework of cross-domain feature alignment and feature-level desensitization, performs phenotypic-genotypic collaborative screening, parental dynamic matching and multi-generation genetic evolution deduction, and generates breeding planning simulation data; The evaluation and output module is used to construct a multi-dimensional evaluation system and perform comprehensive ranking on the breeding planning simulation data to obtain the optimal breeding planning scheme.
Citation Information
Patent Citations
CN120430659A
CN121369215A