An intelligent breeding planning and decision-making method and system based on a large model

By constructing a structured breeding knowledge graph and a large model for joint phenotype-genotype analysis, the problems of low efficiency and inaccurate decision-making in traditional breeding methods are solved, and efficient and scientific decision-making in intelligent breeding planning is achieved.

CN120430659BActive Publication Date: 2025-09-30CHANGSHA BAIAOYUN DATA TECH CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510928527.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-07
Publication Date
2025-09-30
Estimated Expiration
2045-07-07

AI Technical Summary

Technical Problem

Traditional breeding methods rely on experience and manual analysis, are inefficient, and are unable to cope with the complexity and diversity of large-scale agricultural production. They also fail to make accurate decisions based on specific user needs and ignore the joint analysis of phenotypes and genotypes.

Method used

Build a structured breeding knowledge graph, integrate multi-source breeding data, conduct phenotype-genotype joint analysis through large models, simulate multi-generation pedigree paths based on user breeding intentions, evaluate the feasibility of breeding plans, and generate intelligent breeding decision results.

Benefits of technology

It improves the intelligence and operability of breeding planning, enhances breeding efficiency and success rate, ensures the scientific nature and executability of decision-making, and reduces waste of resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120430659B_ABST
    Figure CN120430659B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of breeding planning, and in particular to an intelligent breeding planning and decision-making method and system based on a large model. The method comprises the following steps: obtaining a multi-source breeding data set; constructing a structured breeding knowledge graph based on the multi-source breeding data set; associating the structured breeding knowledge graph with breeding data according to a preset large model to generate a breeding-specific basic model; obtaining breeding instructions input by a user; performing user semantic recognition on the breeding instructions input by the user to generate breeding semantic recognition data; inputting the breeding semantic recognition data into the breeding-specific basic model to perform breeding intention analysis to generate user breeding intention data; and confirming the data information that needs to be called based on the user breeding intention data, analyzing and screening to generate germplasm resource screening data and breeding planning / breeding decision-making schemes. The present invention improves the intelligence and operability of breeding planning by integrating multi-source data, intelligent semantic recognition, joint genetic analysis and executability evaluation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of breeding planning, and in particular to an intelligent breeding planning and decision-making method and system based on a large model. Background Art

[0002] Traditional breeding relies on experience and manual analysis, resulting in low efficiency and difficulty coping with the complexity and diversity of large-scale agricultural production. With the development of computer technology, data processing capabilities have gradually increased, and breakthroughs in genomics have made precision breeding possible. In the 1990s, advances in genomics and molecular biology promoted the application of marker-assisted breeding (MAS), enabling more refined breeding decisions. With the rapid development of artificial intelligence and big data, machine learning and deep learning techniques have been introduced to the breeding field, enabling the processing of increasingly complex and massive data sets and improving the accuracy of breeding predictions. In particular, the application of large models (such as deep neural networks and generative adversarial networks) in agriculture has greatly enhanced the intelligence and automation of breeding programs. However, most existing technologies are still based on general breeding models, failing to make precise decisions tailored to specific user needs. Furthermore, traditional breeding often relies on single-phenotype or genotype analysis, neglecting the combined analysis of both. This results in low intelligence and low operability in breeding planning. Summary of the Invention

[0003] Based on this, it is necessary to provide an intelligent breeding planning and decision-making method and system based on a large model to solve at least one of the above technical problems.

[0004] To achieve the above objectives, a large-scale model-based intelligent breeding planning and decision-making method is provided, the method comprising the following steps:

[0005] Step S1: Acquire a multi-source breeding data set; construct a structured breeding knowledge graph and vector repository based on the multi-source breeding data set; associate breeding data with the structured breeding knowledge graph according to a preset large model to generate a breeding-specific basic model;

[0006] Step S2: Obtaining breeding instructions input by the user; performing user semantic recognition on the breeding instructions input by the user to generate breeding semantic recognition data; performing matching retrieval on the breeding semantic recognition data and the structured breeding knowledge graph and the vector repository, and inputting the retrieval results and the user-input breeding instructions into a breeding-specific basic model to perform breeding intention analysis and generate user breeding intention data;

[0007] Step S3: Based on the user's breeding intention data, germplasm resources are called through a preset differentially privacy-protected federated learning framework and phenotype-genotype joint analysis is performed to generate germplasm resource screening data, where the phenotype-genotype joint analysis includes phenotype data resource matching, genotype data resource matching, and phenotype-genotype data resource matching; parents are dynamically matched on the germplasm resource screening data to generate breeding parent matching data; based on the breeding parent matching data, multi-generation pedigree path simulation is performed on the user's breeding intention data to generate breeding planning simulation data;

[0008] Step S4: Evaluate the executability of the breeding planning simulation data, and output the breeding planning simulation data by path sorting according to the executability evaluation results to obtain the intelligent breeding decision results.

[0009] The present invention realizes systematic and semantic management of knowledge by constructing a structured breeding knowledge graph, integrating multi-dimensional information such as phenotype, genotype, ecological environment, and germplasm resources, and provides a basis for intelligent reasoning and in-depth analysis. A breeding-specific basic model is introduced to accurately obtain the user's breeding goals and constraints through semantic understanding and intention recognition of user input instructions, thereby improving the system's responsiveness to complex needs. Through the joint analysis mechanism of genotype and phenotype, the scientificity and accuracy of germplasm resource screening are improved, providing high-quality data support for subsequent parent configuration and path simulation. Multi-generation breeding path simulation is carried out in combination with the user's breeding intention, and the optimal parent combination is dynamically matched to ensure multi-dimensional optimization of the planned path in terms of genetic potential, adaptability, etc. Through the feasibility evaluation and intelligent sorting of the simulated path, it is ensured that the breeding path finally output not only has theoretical advantages, but also has practical execution value, helping to implement scientific decision-making. This method constructs a complete intelligent breeding closed loop from data integration, intention recognition, resource screening, path simulation to decision output, effectively reducing the dependence on manual experience and improving the efficiency and success rate of modern breeding. Therefore, the present invention improves the intelligence and operability of breeding planning by integrating multi-source data, intelligent semantic recognition, joint genetic analysis and executable evaluation.

[0010] Preferably, step S1 includes the following steps:

[0011] Step S11: Acquire a multi-source breeding dataset, wherein the multi-source breeding dataset includes approved variety data / genotype / phenotype association data, environmental response parameters, and a literature knowledge base;

[0012] Step S12: performing data preprocessing on the multi-source breeding dataset to generate a standard multi-source breeding dataset, wherein the data preprocessing includes data cleaning, data denoising, missing value filling and data standardization;

[0013] Step S13: extracting breeding features from a standard multi-source breeding dataset to construct a structured breeding knowledge graph;

[0014] Step S14: embedding multi-dimensional entities and relationships in the structured breeding knowledge graph according to the preset large model to generate graph vectorized representation data; performing data association and storage on the graph vectorized representation data based on the graph neural network and RAG engine to generate semantically enhanced breeding association data and a vector repository;

[0015] Step S15: Based on the preset domain adaptation training corpus, the semantically enhanced breeding-related data is subjected to structural compression and model mapping to generate a breeding-specific basic model.

[0016] The present invention solves the problems of heterogeneity, noise interference and data missing in the original breeding data by cleaning, denoising, filling missing values ​​and standardizing multi-source breeding data, generating a high-quality, unified format standard multi-source breeding data set, and laying a solid foundation for subsequent analysis. A structured breeding knowledge map is constructed through feature extraction, and the core information such as varieties, traits, genes, environment, and technical paths are systematically integrated to achieve visualization and structured expression of complex breeding knowledge, thereby improving the reusability and reasonability of knowledge. The structured map is embedded with multi-dimensional entities and relationships to construct a map vector representation of data, and semantically enhanced breeding-related data is further generated through semantic association mining, so that the system can understand complex relationships at the semantic level and improve the semantic accuracy of subsequent model reasoning. The semantically enhanced data is structurally compressed and mapped to construct a breeding-specific basic model with both knowledge coverage and computational efficiency. The model can efficiently support tasks such as intention recognition, parent selection and path simulation under user breeding instructions, and has good versatility and portability. With the help of graph vectorization and semantic enhancement mechanisms, the model has good concept generalization and semantic association capabilities, can accurately understand user breeding needs, and realize intelligent processing from knowledge reasoning to solution simulation.

[0017] Preferably, step S13 includes the following steps:

[0018] Step S131: extracting breeding characteristics from a standard multi-source breeding data set to obtain initial breeding characteristic data;

[0019] Step S132: semantically classifying the initial breeding feature data to generate semantically classified feature data; performing ontology mapping on the semantically classified feature data to generate feature ontology mapping data;

[0020] Step S133: identifying entities in the feature ontology mapping data, and constructing relationships between the identified entities to generate candidate knowledge triple data;

[0021] Step S134: perform knowledge fusion and redundancy elimination on the candidate knowledge triple data to generate a refined breeding knowledge triple set; construct a graph based on the refined breeding knowledge triple set, use the breeding entities in the refined breeding knowledge triple set as nodes, and the breeding relationships as edges to generate a structured breeding knowledge graph.

[0022] The present invention ensures the comprehensiveness and accuracy of the basic data for breeding knowledge construction and improves the structural integrity of the knowledge graph by deeply mining the key features (such as genotype, phenotype, trait expression, cultivation method, etc.) in standard multi-source breeding data. The feature data is classified through a semantic classification mechanism, and ontology mapping is performed in combination with the breeding field ontology standard, effectively solving the problems of ambiguity and heterogeneous expression of breeding terms, and achieving semantic consistency and standardized expression of concepts that can be machine-recognized. The core entities in the feature ontology mapping data are identified and the semantic relationships between the entities are constructed to form preliminary candidate knowledge triples, which provide basic semantic structure support for graph construction and help the system to perform complex knowledge reasoning such as causality and association. By integrating candidate knowledge triples with redundancy resolution strategies, duplicate, conflicting and invalid knowledge fragments are eliminated, and a refined triple set with high expressiveness and explanatory power is retained to ensure that the graph structure is clear and the semantics are clear, which facilitates subsequent efficient indexing and semantic reasoning. Through a graph-based structure generation process using entities as nodes and relationships as edges, a structured breeding knowledge graph is ultimately formed. This not only visually represents the complex breeding knowledge system but also supports diverse intelligent analysis tasks within the graph space, such as graph mining, path search, and deep relationship exploration. This graph, serving as the foundation for semantic understanding and knowledge reasoning within large-scale breeding models, can be widely applied in subsequent intelligent breeding scenarios, including breeding intent analysis, parent recommendation, and path simulation, significantly improving model interpretability and output decision accuracy.

[0023] Preferably, step S2 includes the following steps:

[0024] Step S21: obtaining breeding instructions input by the user;

[0025] Step S22: performing instruction text translation on the breeding instruction input by the user to generate user input text data; performing instruction parsing on the user input text data to generate initial breeding instruction parsed data;

[0026] Step S23: extracting the syntax of the initial breeding instruction parsed data, and using the syntax to perform semantic modeling on the initial breeding instruction parsed data to generate breeding semantic recognition data;

[0027] Step S24: Match and search the breeding semantic recognition data with the structured breeding knowledge graph and vector repository, analyze the part-of-speech ratio of the breeding semantic recognition data, thereby confirming the breeding intention of the breeding semantic recognition data and generating breeding intention mapping data; input the breeding intention mapping data and the user-input breeding instructions into the breeding-specific basic model for breeding intention model inference to generate user breeding intention data.

[0028] This system accepts breeding instructions expressed in natural language or industry terminology, fully meeting the user habits of professional breeders and enhancing the system's practicality and user-friendliness. By translating and parsing user instruction text, it can adapt to complex contexts such as dialects, professional abbreviations, and polysemous expressions, effectively improving the machine understandability of the original input text and ensuring the accuracy of subsequent processing. By analyzing the instruction grammatical structure and integrating it with semantic modeling, it can identify key elements within the command, such as core actions, objects, and constraints, generating structured breeding semantic recognition data, significantly improving the system's processing accuracy for complex instructions. By analyzing the proportion of part-of-speech structures (e.g., verbs, nouns, and adjectives) within the semantic recognition data, it identifies user focus and breeding task objectives at the linguistic level, constructing semantically accurate and logically complete breeding intent mapping data. This intent mapping data is then fed into a dedicated breeding-specific basic model for inference. This model integrates knowledge graphs and semantic vector representations to intelligently determine the user's breeding objectives (e.g., target traits, target germplasm, and technical approaches), generating accurate breeding intent data.

[0029] Preferably, in step S24, analyzing the part-of-speech ratio of the breeding semantic identification data to confirm the breeding intention of the breeding semantic identification data includes:

[0030] Perform part-of-speech tagging on breeding semantic recognition data to generate breeding semantic part-of-speech tagging data;

[0031] Statistical analysis was performed on breeding semantic part-of-speech tagging data to calculate the proportion of each part-of-speech type and construct a high-dimensional part-of-speech vector feature matrix. Part-of-speech categories include nouns, verbs, adjectives, proper nouns, question words, and conjunctions.

[0032] Set the following logical rules to perform numerical judgment and intent recognition on the high-dimensional part-of-speech vector feature matrix:

[0033] When the high-dimensional part-of-speech vector feature matrix has a noun ratio of <40%, a verb ratio of >30%, and an adjective ratio of >15%, it is identified as a representational breeding intention;

[0034] When the combined proportion of proper nouns and technical verbs in the high-dimensional part-of-speech vector feature matrix is ​​greater than 50% and contains the keywords "improve" and "resistance", it is identified as an improvement breeding intention;

[0035] When the total proportion of question words and conditional clause structures in the high-dimensional part-of-speech vector feature matrix is ​​greater than 20% and contains hypothetical structures, it is identified as exploratory breeding intention;

[0036] The data of representation-based breeding intentions, improvement-based breeding intentions and exploratory breeding intentions are integrated to generate breeding intention mapping data.

[0037] The present invention accurately tags breeding semantic recognition data with parts of speech and constructs a high-dimensional part-of-speech feature vector matrix including six types of language components: nouns, verbs, adjectives, proper nouns, interrogative words, and conjunctions. This allows natural language information to be converted into structured data with statistical characteristics, improving the accuracy and stability of language semantic processing. Based on the constructed high-dimensional part-of-speech vector feature matrix, the system can automatically identify the categories of user breeding intentions in combination with preset logical rules, including representational breeding intentions (such as focusing on phenotypic descriptions), improvement breeding intentions (such as targeting trait optimization), and exploratory breeding intentions (such as exploring potential paths based on hypothetical conditions), effectively supporting the upstream and downstream process connection of complex and personalized breeding tasks. By setting clear numerical thresholds and keyword judgment rules, such as identifying nouns <40%, verbs >30%, and adjectives >15% as representational, a traceable conversion path from language structure features to semantic intentions is achieved, giving the breeding model reasoning process a clear logical basis and explainability. Special introduction of the judgment of keywords such as "improve" and "resistance" as well as the grammatical recognition of conditional clauses and hypothetical structures can accurately determine whether the user has the motivation for improvement or the need to ask exploratory questions, and enhance the system's adaptability to breeding industry terminology and scientific research instructions.

[0038] Preferably, in step S3, calling germplasm resources based on the user's breeding intention data through a preset differential privacy-protected federated learning framework and performing phenotype-genotype joint analysis includes:

[0039] The user's breeding intention data is protected for data privacy through a preset differentially private federated learning framework to generate user breeding intention protection data. Resource matching is performed based on the user's breeding intention protection data to obtain germplasm resources, where resource matching includes phenotypic data resource matching, genotypic data resource matching, and phenotype-genotype data resource matching. Multi-channel phenotypic representation modeling is performed on the germplasm resources to generate a multidimensional trait vector set.

[0040] Perform high-dimensional genomic mosaic mapping on the multidimensional trait vector set to generate a nested map of genotype-phenotype associations; perform structure-preserving variation deconstruction on the nested map of genotype-phenotype associations to generate characteristic mutation location data;

[0041] Perform target trait allele inference on characteristic mutation mapping data to generate a set of functional genetic loci; quantify the pan-genetic group value of the functional genetic loci set to generate a set of multidimensional optimization factors for germplasm resources;

[0042] The multi-dimensional optimization factor set of germplasm resources is intelligently decoupled with intention weighting to generate germplasm resource screening data.

[0043] The present invention uses the user's breeding intention data as a guiding factor to carry out germplasm resource matching oriented to trait requirements, effectively avoiding the drawbacks of traditional germplasm screening that relies on manual rules and lacks intelligent guidance, and improving the relevance and adaptability of the initial resource library. By performing multi-channel acquisition and structural modeling of the phenotypic characteristics of the target germplasm using images, spectra, texts, and sensors, a multidimensional trait vector set with high information richness and discrimination is generated, providing a high-quality phenotypic expression basis for subsequent gene association analysis. Based on the trait vector set, the system constructs a nested phenotype-genotype association map, which not only accurately reflects the coupling relationship between traits and genes, but also has structural progressiveness and association depth, and can support multi-level genetic traceability analysis. The nested map is subjected to structure-preserving variation deconstruction processing, which can identify and locate variation characteristics in multi-source structures, generate characteristic mutation location data with high biological significance, and provide precise targets for subsequent trait factor modeling. By functionally reasoning about alleles in mutation regions and assessing their contributions to the pan-genetic group, we can screen and prioritize functional genetic loci, thereby constructing a multidimensional set of optimal factors with explanatory power for biological effects and establishing a new paradigm for measurable genetic value. Ultimately, we embed the user intent model into the intelligent decoupling process of the optimal factor set, comprehensively considering target traits, genetic value, and usage scenarios, and outputting semantically reinforced and intent-aligned germplasm resource screening data, achieving a closed-loop decision-making process from data analysis to intelligent recommendation.

[0044] Preferably, performing dynamic parent matching on germplasm resource screening data in step S3 includes:

[0045] Extract parental pairing trait characteristics from germplasm resource screening data, evaluate the correlation between parents, and generate a parental similarity map; perform dynamic gene combination optimization on the parental similarity map to generate a parental genotype matching matrix;

[0046] Perform environmental adaptability analysis on the parental genotype matching matrix to generate parental environmental adaptability data;

[0047] Perform multi-dimensional priority sorting on the parental genotype matching matrix according to the parental environmental adaptability data to generate priority matching parent candidate data;

[0048] The genetic diversity of the priority matching parent candidate data is verified to generate breeding parent matching data.

[0049] The present invention extracts and calculates semantic associations of key pairing traits for screening germplasm to generate a parent similarity map with trait correlation and functional comparability, providing structured and quantitative evaluation support for parent screening, significantly improving the scientificity and accuracy of parent matching. A multi-objective genetic optimization algorithm is introduced on the basis of the parent similarity map to construct a genotype matching matrix, taking into account allele complementarity, linkage relationship deconstruction and fusion of superior traits, thereby improving the breeding potential of offspring combinations. Combined with the performance data and adaptability model of the parents in the target ecological zone, the genotype matching matrix is ​​evaluated for environmental adaptability, and parent combinations that can stably express excellent traits under multiple ecological backgrounds are screened out, thereby enhancing the generalization ability and practical feasibility of the pairing results. Based on comprehensive modeling of dimensions such as parental genotype fusion degree, adaptability score, and target trait proportion, a multi-dimensional priority sorting mechanism is constructed to perform gradient optimization of candidate parents to ensure that high-value parent combinations are given priority in the subsequent breeding process. After the initial screening of parental combinations, further genetic background analysis and diversity testing are carried out to avoid genetic bottlenecks and risks of inbreeding, ensure that the combinations have good genetic breadth and breeding innovation potential, and meet long-term variety optimization and improvement needs.

[0050] Preferably, in step S3, performing multi-generation pedigree path simulation on the user's breeding intention data based on the breeding parent matching data includes:

[0051] Reconstruct trait-driven genetic coding of breeding parent matching data to generate parent genetic representation unit sets;

[0052] Map the target trait-oriented path of the user's breeding intention data to generate the intended trait evolution trajectory data;

[0053] Based on the parental genetic representation unit set and the intended trait evolution trajectory data, intergenerational recombination probability modeling is performed to generate a genetic recombination transfer matrix;

[0054] Perform multi-generational simulation of genetic transmission evolution on the genetic recombination transfer matrix to generate a path-level pedigree evolution map;

[0055] Through the path-level pedigree evolution map, the trait achievement rate inversion analysis of the genetic recombination transfer matrix is ​​performed to generate a multi-path breeding achievement rate matrix;

[0056] The target driving probability of the multi-path breeding achievement rate matrix is ​​calculated by clustering to obtain the breeding planning simulation data.

[0057] The present invention reconstructs the genetic coding driven by the breeding parent matching data to generate the parent genetic characterization unit set, which can accurately reflect the genetic characteristics and key traits of each parent and provide a solid genetic basis for subsequent intergenerational recombination. According to the user's breeding intention data, a guide path mapping of the target trait is constructed, and the intended trait evolution trajectory data is generated, so as to clearly outline the evolution trend of the target trait in future generations and enhance the accuracy and operability of the breeding goals. By modeling the intergenerational recombination probability based on the parent genetic characterization unit set and the intended trait evolution trajectory data, a genetic recombination transfer matrix is ​​generated, which can efficiently simulate the transfer and variation of each genotype and trait during multi-generation transmission, and provide a powerful prediction tool for the trait optimization of offspring. By performing multi-generational simulation of genetic transmission evolution on the genetic recombination transfer matrix, a path-level pedigree evolution map is generated, which can intuitively display the trait expression and genetic drift in the multi-generation genetic evolution process, helping breeding experts make scientific adjustments and decisions in multi-generation plans. By using path-level pedigree evolution maps and inversely analyzing trait achievement rates on the genetic recombination transfer matrix, the probability of trait achievement can be calculated for each genetic path, providing a clear assessment basis for breeding planning and enhancing the scientific nature of decision-making. By clustering the target-driven probabilities of multi-path breeding achievement rate matrices, the achievement scenarios under different breeding paths can be accurately simulated, providing an optimized decision-making basis for breeding planning simulation data and enabling breeding programs to achieve their intended goals to the greatest extent possible in actual operation.

[0058] Preferably, the step S4 of evaluating the feasibility of the breeding planning simulation data includes:

[0059] Import the breeding planning simulation data into the preset biological breeding management system for field extraction and field format verification to ensure that the integrity of each data item is ≥98%, and generate a structured planning simulation data set. The extracted fields include the probability of achieving the target trait, expected number of generations, genetic diversity index, stability coefficient, and environmental adaptation value;

[0060] The target trait achievement probability field in the structured planning simulation data set is screened using a probability threshold, with the minimum threshold set to 0.75 and the maximum threshold set to 1.0, to screen out eligible individual paths. If the target trait achievement probability is lower than 0.75, the path is marked as having no achievement potential, and a potential achievement screening result graph is generated.

[0061] The genetic diversity of the pathway data that passed the target trait screening was evaluated and scored using the Shannon diversity index calculation formula, with the minimum Shannon index set at 0.30. If the Shannon index of an individual pathway was lower than this value, it was marked as having a genetic bottleneck, and a genetic diversity score map was generated.

[0062] Perform genetic stability analysis on qualified genetic diversity paths and calculate the standard deviation σ. The σ range is set to 0.00 to 0.20. If the σ value of a path exceeds 0.20, it is judged to have large genetic fluctuations and is not recommended for execution. Finally, a stability judgment matrix diagram is generated;

[0063] Environmental adaptability analysis was performed on the paths that passed the stability screening. The environmental adaptability score of each path was compared with the target planting area environment using the environmental index matching calculation formula, with a matching score threshold set at ≥0.70. If the path score was lower than 0.70, it was marked as having low adaptability, and an environmental adaptability screening image was generated.

[0064] All paths that pass the screening are scored using a multi-indicator weighted approach, with the weights for each indicator set as follows: 30% for trait achievement rate, 25% for diversity, 20% for stability, and 25% for environmental adaptability. Weighted scores are assigned based on a normalized scoring formula, ranging from 0 to 1.0, with a recommended execution threshold set at ≥0.75. This ultimately generates a weighted scoring result for the breeding planning path.

[0065] The weighted scoring results of the breeding planning path are visualized as a score distribution graph, and the output is the feasibility assessment result.

[0066] This invention imports breeding planning simulation data into a biobreeding management system, performs field extraction and format verification, and ensures that the integrity of each data item reaches ≥98%. The resulting structured planning simulation dataset ensures high data quality and consistency, providing a reliable foundation for subsequent analysis and decision-making. By setting a screening threshold (0.75 to 1.0) for the probability of achieving target traits, breeding pathways with high potential can be effectively identified. Pathways below this threshold are marked to ensure that individual pathways with the greatest potential for achieving breeding goals are selected, avoiding wasting resources on pathways lacking the potential to meet targets. The Shannon Diversity Index is used to assess the genetic diversity of pathways, ensuring that the selected pathways possess sufficient genetic diversity, avoiding genetic bottlenecks, and thus improving the adaptability and long-term stability of the breeding program. A genetic diversity score chart provides intuitive data support for breeding experts. By calculating the standard deviation (σ) and setting it within a range of 0.00 to 0.20, pathways with large genetic fluctuations can be identified and unstable pathways can be eliminated, thereby ensuring the genetic stability of the selected pathways. The resulting stability judgment matrix provides a scientific basis for breeding decisions and avoids the implementation of unstable pathways. Environmental adaptability analysis uses an environmental index matching formula to compare the adaptability of each breeding path with the target planting area's environment, ensuring that the selected path performs well in the actual environment. Pathways with low adaptability are eliminated, reducing the risk of failure due to environmental inadaptability. A comprehensive evaluation based on weighted scores for multiple indicators, including trait achievement rate, diversity, stability, and environmental adaptability, ensures that the selected breeding path possesses comprehensive advantages. This weighted scoring intuitively demonstrates the comprehensive potential of each path, helping breeders make optimal decisions.

[0067] In this specification, a large-scale model-based intelligent breeding planning and decision-making system is provided for executing the above-mentioned large-scale model-based intelligent breeding planning and decision-making method. The large-scale model-based intelligent breeding planning and decision-making system includes:

[0068] A breeding map construction module is used to obtain multi-source breeding data sets; build a structured breeding knowledge map and vector repository based on the multi-source breeding data sets; associate breeding data with the structured breeding knowledge map according to a preset large model to generate a breeding-specific basic model;

[0069] The breeding instruction recognition module is used to obtain the breeding instructions input by the user; perform user semantic recognition on the breeding instructions input by the user to generate breeding semantic recognition data; match and search the breeding semantic recognition data with the structured breeding knowledge graph and vector repository, and input the search results and the user-input breeding instructions into the breeding-specific basic model for breeding intention analysis to generate user breeding intention data;

[0070] The breeding planning module is used to call germplasm resources based on user breeding intention data through a preset differentially privacy-protected federated learning framework and perform phenotype-genotype joint analysis to generate germplasm resource screening data. The phenotype-genotype joint analysis includes phenotypic data resource matching, genotype data resource matching, and phenotype-genotype data resource matching; dynamically match parents on germplasm resource screening data to generate breeding parent matching data; and simulate multi-generation pedigree paths on user breeding intention data based on breeding parent matching data to generate breeding planning simulation data.

[0071] The decision output module is used to evaluate the executability of the breeding planning simulation data, and output the breeding planning simulation data in a path sorting manner according to the executability evaluation results to obtain the intelligent breeding decision results.

[0072] The beneficial effect of the present invention is that it constructs a structured breeding knowledge graph through the integration of multi-source breeding data sets, generates a dedicated basic model, and provides comprehensive data support for subsequent decision-making. The breeding instruction recognition module intelligently identifies and analyzes breeding intentions based on the breeding instructions input by the user to ensure that user needs are accurately captured. The breeding planning module accurately matches breeding resources through germplasm resource screening, phenotype-genotype joint analysis and dynamic matching of parents, and performs multi-generation pedigree path simulation to provide scientific path planning for breeding decisions. Finally, the decision output module performs an executable assessment on the breeding planning simulation data, intelligently sorts and generates the optimal breeding path, thereby greatly improving the scientific nature and executability of breeding decisions. The comprehensive application of this system effectively improves the achievement rate of breeding goals, reduces resource waste, promotes the intelligence and automation of the breeding process, and significantly improves breeding efficiency and success rate. Therefore, the present invention improves the intelligence and operability of breeding planning by integrating multi-source data, intelligent semantic recognition, joint genetic analysis and executable assessment. BRIEF DESCRIPTION OF THE DRAWINGS

[0073] Figure 1 A schematic diagram of the steps of an intelligent breeding planning and decision-making method based on a large model;

[0074] Figure 2 for Figure 1 Detailed implementation steps of step S1 in FIG.

[0075] Figure 3 for Figure 1 Detailed implementation steps of step S2 in FIG.

[0076] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0077] The following is a clear and complete description of the technical method of the present invention in conjunction with the accompanying drawings. It is obvious that the embodiments described are part of the embodiments of the present invention, but not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without making any creative efforts are within the scope of protection of the present invention.

[0078] In addition, the accompanying drawings are merely schematic illustrations of the present invention and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the block diagrams shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor and / or microcontroller approaches.

[0079] It should be understood that although the terms "first," "second," and the like may be used herein to describe various elements, these elements should not be limited by these terms. These terms are used solely to distinguish one element from another. For example, a first element may be referred to as a second element, and similarly, a second element may be referred to as a first element, without departing from the scope of the exemplary embodiments. The term "and / or" as used herein includes any and all combinations of one or more of the listed associated items.

[0080] To achieve this, please refer to Figures 1 to 3 , a large-scale model-based intelligent breeding planning and decision-making method and system, the method comprising the following steps:

[0081] Step S1: Acquire a multi-source breeding data set; construct a structured breeding knowledge graph and vector repository based on the multi-source breeding data set; associate breeding data with the structured breeding knowledge graph according to a preset large model to generate a breeding-specific basic model;

[0082] Step S2: Obtaining breeding instructions input by the user; performing user semantic recognition on the breeding instructions input by the user to generate breeding semantic recognition data; performing matching retrieval on the breeding semantic recognition data and the structured breeding knowledge graph and the vector repository, and inputting the retrieval results and the user-input breeding instructions into a breeding-specific basic model to perform breeding intention analysis and generate user breeding intention data;

[0083] Step S3: Based on the user's breeding intention data, germplasm resources are called through a preset differentially privacy-protected federated learning framework and phenotype-genotype joint analysis is performed to generate germplasm resource screening data, where the phenotype-genotype joint analysis includes phenotype data resource matching, genotype data resource matching, and phenotype-genotype data resource matching; parents are dynamically matched on the germplasm resource screening data to generate breeding parent matching data; based on the breeding parent matching data, multi-generation pedigree path simulation is performed on the user's breeding intention data to generate breeding planning simulation data;

[0084] Step S4: Evaluate the executability of the breeding planning simulation data, and output the breeding planning simulation data by path sorting according to the executability evaluation results to obtain the intelligent breeding decision results.

[0085] The present invention realizes systematic and semantic management of knowledge by constructing a structured breeding knowledge graph, integrating multi-dimensional information such as phenotype, genotype, ecological environment, and germplasm resources, and provides a basis for intelligent reasoning and in-depth analysis. A breeding-specific basic model is introduced to accurately obtain the user's breeding goals and constraints through semantic understanding and intention recognition of user input instructions, thereby improving the system's responsiveness to complex needs. Through the joint analysis mechanism of genotype and phenotype, the scientificity and accuracy of germplasm resource screening are improved, providing high-quality data support for subsequent parent configuration and path simulation. Multi-generation breeding path simulation is carried out in combination with the user's breeding intention, and the optimal parent combination is dynamically matched to ensure multi-dimensional optimization of the planned path in terms of genetic potential, adaptability, etc. Through the feasibility evaluation and intelligent sorting of the simulated path, it is ensured that the breeding path finally output not only has theoretical advantages, but also has practical execution value, helping to implement scientific decision-making. This method constructs a complete intelligent breeding closed loop from data integration, intention recognition, resource screening, path simulation to decision output, effectively reducing the dependence on manual experience and improving the efficiency and success rate of modern breeding. Therefore, the present invention improves the intelligence and operability of breeding planning by integrating multi-source data, intelligent semantic recognition, joint genetic analysis and executable evaluation.

[0086] In the embodiment of the present invention, reference Figure 1 FIG. 1 is a flow chart of a large-scale model-based intelligent breeding planning and decision-making method according to the present invention. In this example, the large-scale model-based intelligent breeding planning and decision-making method includes the following steps:

[0087] Step S1: Acquire a multi-source breeding data set; construct a structured breeding knowledge graph and vector repository based on the multi-source breeding data set; associate breeding data with the structured breeding knowledge graph according to a preset large model to generate a breeding-specific basic model;

[0088] Step S2: Obtaining breeding instructions input by the user; performing user semantic recognition on the breeding instructions input by the user to generate breeding semantic recognition data; performing matching retrieval on the breeding semantic recognition data and the structured breeding knowledge graph and the vector repository, and inputting the retrieval results and the user-input breeding instructions into a breeding-specific basic model to perform breeding intention analysis and generate user breeding intention data;

[0089] Step S3: Based on the user's breeding intention data, germplasm resources are called through a preset differentially privacy-protected federated learning framework and phenotype-genotype joint analysis is performed to generate germplasm resource screening data, where the phenotype-genotype joint analysis includes phenotype data resource matching, genotype data resource matching, and phenotype-genotype data resource matching; parents are dynamically matched on the germplasm resource screening data to generate breeding parent matching data; based on the breeding parent matching data, multi-generation pedigree path simulation is performed on the user's breeding intention data to generate breeding planning simulation data;

[0090] Step S4: Evaluate the executability of the breeding planning simulation data, and output the breeding planning simulation data by path sorting according to the executability evaluation results to obtain the intelligent breeding decision results.

[0091] In this embodiment of the present invention, breeding datasets are obtained from multiple sources, including farm data, laboratory genotype data, climate data, and agricultural literature. This data ensures coverage of genotypes, phenotypes, environmental conditions, and management practices for different crops. The raw data is cleaned and standardized, including deduplication, missing value processing, and normalization. A structured breeding knowledge graph is constructed based on the acquired multi-source data. Graph nodes include genotypes, phenotypes, environmental factors, and breeding methods, while edges represent relationships between them. Graph databases (such as Neo4j) can be used to store and manage breeding-related information. The data is annotated to identify associations between different data types (such as the association between genotype and phenotype), structured, and mapped into a knowledge graph. Each entity in the knowledge graph (such as genotype, phenotype, etc.) is converted into a vector representation using natural language processing (NLP) or embedding models (such as Word2Vec and BERT). Each entity is converted into a low-dimensional vector to facilitate subsequent retrieval and analysis. These vectors are stored in a vector database (such as FAISS and Milvus) for efficient similarity calculation and retrieval. Based on the structured breeding knowledge graph and vector repository, a specialized breeding foundation model is trained. Deep learning or other machine learning methods (such as convolutional neural networks and graph neural networks) can be used to enable the model to understand the relationships between breeding expertise and data. The foundational model is trained using a pre-defined model training framework, enabling it to provide accurate recommendations during the breeding decision-making process. User breeding instructions are obtained through natural language processing interfaces (such as voice recognition and text input), allowing users to describe their requirements for breeding goals, variety selection, environmental requirements, and so on. Deep learning-based semantic analysis models (such as BERT and Transformer) perform semantic recognition on the user-entered breeding instructions, extracting key information such as breeding goals, target crops, variety requirements, and climate adaptability. The identified semantic information is structured and converted into a data format compatible with the knowledge graph. The user's breeding semantic recognition data is matched with data in the knowledge graph and vector repository, and similarity searches are performed to identify the most relevant breeding data (such as suitable genotypes and phenotypes). Combined with the user's breeding instructions, breeding intention data is generated. This data contains the user's desired breeding direction and goals, further guiding the subsequent breeding process. Based on the user's breeding intent, germplasm data matching the target phenotypic characteristics is screened. Germplasm data matching the target genotype is screened, taking into account the genetic diversity and stability of the genotype. Phenotypic and genotypic data are jointly matched to determine the most appropriate germplasm combination. Based on phenotype-genotype matching, multi-objective optimization algorithms (such as genetic algorithms and particle swarm optimization) are used to screen the optimal germplasm combination. Dynamic parent matching is performed based on the selected germplasm.By simulating parent pairings and their offspring performance, the most suitable parent combination is determined. Intelligent algorithms such as genetic algorithms or simulated annealing are used to dynamically match parents to ensure the achievement of breeding goals. Based on parent matching data, breeding algorithms simulate multi-generational pedigree pathways, analyze genetic transmission and selective breeding results, and generate breeding plan simulation data. By simulating different breeding pathways, the effectiveness of each pathway is evaluated, and the pathway most likely to achieve the desired breeding goals is selected. The feasibility of different breeding pathways is evaluated based on simulated data and external data sources. Factors such as climate adaptability, genetic stability, and production capacity are evaluated to select the most practical breeding pathway. The feasibility of breeding plans is scored using mathematical modeling or machine learning algorithms (such as decision trees and support vector machines). Based on the feasibility assessment results, the breeding pathways are ranked using ranking algorithms (such as multi-criteria decision-making methods like AHP and TOPSIS) to select the optimal pathway. Based on the evaluation results and the pathway ranking, the system outputs the final intelligent breeding decision, providing the user with the most suitable breeding plan and implementation steps.

[0092] As an example of the present invention, refer to Figure 2 As shown, in this example, step S1 includes:

[0093] Step S11: Acquire a multi-source breeding dataset, wherein the multi-source breeding dataset includes approved variety data / genotype / phenotype association data, environmental response parameters, and a literature knowledge base;

[0094] Step S12: performing data preprocessing on the multi-source breeding dataset to generate a standard multi-source breeding dataset, wherein the data preprocessing includes data cleaning, data denoising, missing value filling and data standardization;

[0095] Step S13: extracting breeding features from a standard multi-source breeding dataset to construct a structured breeding knowledge graph;

[0096] Step S14: embedding multi-dimensional entities and relationships in the structured breeding knowledge graph according to the preset large model to generate graph vectorized representation data; performing data association and storage on the graph vectorized representation data based on the graph neural network and RAG engine to generate semantically enhanced breeding association data and a vector repository;

[0097] Step S15: Based on the preset domain adaptation training corpus, the semantically enhanced breeding-related data is subjected to structural compression and model mapping to generate a breeding-specific basic model.

[0098] In this embodiment of the present invention, data on approved varieties, including variety name, approval year, and regional adaptability, is collected from agricultural research institutions, agricultural research publications, and other sources. This data can include basic information about the variety, as well as indicators such as yield, resistance, and adaptability. Genotypic data relevant to breeding can be collected from genomic research databases, gene chip experiments, and other sources. This includes information on gene mutations and genetic markers. Phenotypic data on crops, including growth rate, fruit morphology, and disease resistance, is collected from experimental fields or farms. Genotypic and phenotypic data are combined to conduct phenotype-genotype association analysis to identify genotypic traits that influence phenotypic expression. Agricultural production environment data, including climate, soil, temperature, humidity, and precipitation, can be obtained from meteorological stations, remote sensing data, and agricultural meteorological models. The responses of different varieties to different environments are recorded, including characteristics such as drought resistance and pest and disease resistance. Scientific research papers, breeding books, and agricultural technology reports in related fields are collected. These can be obtained through agricultural research databases (such as CNKI) and international literature databases (such as PubMed and Google Scholar). Use text mining techniques (such as natural language processing) to extract information from the literature on breeding methods, variety characteristics, and genotype-phenotype relationships. Clean up duplicate and inconsistent data to ensure data uniqueness and integrity. Detect and remove or correct outliers to ensure data accuracy. Identify and remove noise from the data. Use filtering algorithms (such as low-pass filtering and median filtering) to remove irrelevant or erroneous data points. Use interpolation methods (such as mean imputation, regression imputation, and KNN imputation) to fill in missing phenotypic or genotypic data. For areas with a high number of missing values, use data augmentation techniques to generate approximate data to reduce the impact of missing data on model training. Standardize numerical data (such as Z-score normalization and Min-Max normalization) to ensure that different features are trained on the same scale. Extract features such as gene markers, mutation types, and gene expression levels from genotypic data. Extract crop growth, yield, and resistance indicators from phenotypic data. Extract environmental factors that affect crop growth, such as climate, soil type, humidity, and temperature, from environmental data. Extracted breeding characteristics are used as nodes in the graph to define relationships between entities (such as the association between genotype and phenotype, the adaptability of varieties to the environment, etc.). A breeding knowledge graph is constructed and managed using a graph database (such as Neo4j) to facilitate subsequent query and analysis. Graph embedding techniques (such as Node2Vec and GraphSAGE) are used to vectorize the entities and relationships in the knowledge graph. Each node and edge is converted into a high-dimensional vector, capturing its structural and semantic information in the graph. The vectors of each entity and relationship in the graph are stored in a vector database (such as FAISS and Milvus) for efficient vector retrieval and analysis. The relationships between different genotypes, phenotypes, and environmental factors are analyzed by calculating the similarity of nodes and edges in the graph.Graph neural networks are used to further learn the nodes and edges in the graph, improving the model's understanding of breeding data. The RAG (Retrieval-Augmented Generation) engine enhances the model's generation capabilities, retrieving and integrating external data to improve the relevance and accuracy of breeding knowledge. Graph vectors are compressed and reduced using methods such as autoencoders and principal component analysis (PCA) to improve computational efficiency and storage space utilization. Domain adaptation techniques (such as transfer learning) are used to train the compressed model using a pre-defined domain-adapted training corpus (e.g., breeding data for a specific crop), making it more adaptable to specific domain needs. A breeding-specific base model is trained based on the compressed data and mapped model. This model incorporates domain expertise to provide users with precise breeding decision support. Model parameters are adjusted through cross-validation and performance evaluation to ensure accuracy and efficiency in real-world applications.

[0099] Preferably, step S13 includes the following steps:

[0100] Step S131: extracting breeding characteristics from a standard multi-source breeding data set to obtain initial breeding characteristic data;

[0101] Step S132: semantically classifying the initial breeding feature data to generate semantically classified feature data; performing ontology mapping on the semantically classified feature data to generate feature ontology mapping data;

[0102] Step S133: identifying entities in the feature ontology mapping data, and constructing relationships between the identified entities to generate candidate knowledge triple data;

[0103] Step S134: perform knowledge fusion and redundancy elimination on the candidate knowledge triple data to generate a refined breeding knowledge triple set; construct a graph based on the refined breeding knowledge triple set, use the breeding entities in the refined breeding knowledge triple set as nodes, and the breeding relationships as edges to generate a structured breeding knowledge graph.

[0104] In this embodiment of the present invention, core breeding indicators, including but not limited to genotypic characteristics (e.g., SNP markers, gene expression levels), phenotypic characteristics (e.g., yield, growth cycle, disease resistance), and environmental interaction characteristics (e.g., temperature and humidity effects, soil nutrient adaptability), are extracted from a standard multi-source breeding dataset using an expert rule base and statistical learning methods (e.g., principal component analysis (PCA) and random forest feature importance scoring). The raw feature data is uniformly converted into a standard structure, such as a key-value pair or vector representation, to generate initial breeding feature data. Natural language processing techniques (e.g., word2vec and BERT embedding) are used to semantically cluster feature names and their descriptions. Features with similar meanings are grouped into a unified semantic category, e.g., "growth rate" and "seedling growth rate" are grouped into the "growth rate category." The output is semantically categorized feature data. Domain ontologies (e.g., plant ontology PO and crop feature ontology TO) are used to map semantically categorized features to standard concepts. Entity alignment and knowledge reasoning algorithms (e.g., OWL reasoning) are combined to match corresponding ontology entities, generating feature ontology mapping data. Standardized entities, such as "maize," "drought resistance," and "ZmDREB2A gene," are extracted from feature ontology mapping data. Relationships between entities (e.g., "possess," "regulate," and "exhibit") are identified. Semantic relationships between entities are constructed using rule bases, co-occurrence analysis, and graph completion techniques. Preliminary knowledge triples are formed, such as (maize, possess, drought resistance), (ZmDREB2A, regulate, drought resistance), and (drought resistance, exhibit, high growth rate), generating candidate knowledge triples. Candidate knowledge triples are merged and deduplicated, integrating semantically similar triples from different data sources. Entity ambiguity and redundancy are addressed using logical reasoning and semantic alignment (e.g., unifying "maize variety A" and "Zea Mays A"). The output is a refined set of breeding knowledge triples. A breeding knowledge graph is constructed, using entities in the triples as nodes and relationships as edges. The graph can be implemented in graph database formats such as RDF and Neo4j to support subsequent querying and reasoning. Finally, a structured breeding knowledge graph is generated, providing a data basis for subsequent graph embedding and intelligent breeding decision-making.

[0105] It is particularly important that the knowledge fusion and redundancy elimination of the candidate knowledge triple data in step S134 also includes:

[0106] Perform semantic synonymy normalization on the candidate knowledge triple data to obtain triple data with unified terminology;

[0107] Extract the relational paradigm of the triple data with unified terminology and fuse them to generate semantically classified triple data;

[0108] Perform entity disambiguation mapping on semantically classified triple data to generate entity unified triple data;

[0109] Perform semantic confidence filtering on entity unified triple data to generate highly reliable triple data;

[0110] Perform ontological logical expansion on high-credibility triple data to generate closed semantic triple data;

[0111] The closed semantic triple data is subjected to heterogeneous structural consistency processing to generate structurally consistent triple data; the structurally consistent triple data is subjected to triple density compression processing to generate a refined breeding knowledge triple set.

[0112] In an embodiment of the present invention, term extraction is performed on the head entity, relationship predicate, and tail entity in the candidate knowledge triple data. A domain ontology dictionary or term mapping table is used to match terms with synonymous relationships (such as "stress resistance" and "stress tolerance", "yield" and "yield per mu"). Standardized terms are uniformly used to replace the original terms to ensure that the term expressions in all triples are consistent, and triple data with unified terms are generated as the basis for subsequent fusion. Semantically similar relationship types are extracted from the unified term triples and classified into the same relationship paradigm (such as "increase yield" and "increase output" are classified as "increase yield"). Triples belonging to the same paradigm are merged, and the relationship predicate representation is unified to generate semantically classified triple data with clear relationship classification. Fuzzy or ambiguous entities appearing in the triples are identified (such as "indica rice" represents "indica rice variety" or "indica rice type"). Disambiguation is performed based on contextual semantics, entity type labels, entity word vectors, and other methods. Different expressions of the same entity are mapped to standard entity IDs to generate unified entity triple data, ensuring entity uniqueness. The semantic confidence of each triple is calculated based on factors such as triple source, contextual consistency, entity strength, and relationship strength. A minimum confidence threshold (e.g., 0.8) is set for filtering, retaining high-confidence triples and removing unstable or noisy triples to generate highly reliable triple data. Agricultural breeding-related ontology logic rules (such as subclass inheritance and attribute transfer) are imported, and closed reasoning is performed on triples based on ontology reasoning rules. For example, from "rice has disease resistance" and "disease resistance increases yield," "rice can increase yield" is inferred. Logically closed triple content is added to generate closed semantic triple data. Structural inconsistencies caused by different data sources, such as nested relationships and hierarchical entity expressions, are identified and standardized (e.g., flattening nested structures into primary relationships and unifying entity hierarchies). The original semantics are preserved during the structural transformation to prevent information loss and generate structurally consistent triple data. Analyze the triple density under each relational paradigm, identify high-redundancy areas, and compress the number of triples by merging equivalent paths and removing duplicate relations without affecting knowledge coverage. Retain the semantic core, compress redundant paths and equivalent relation chains, and generate a high-quality, low-redundancy refined breeding knowledge triple set for constructing an efficient agricultural knowledge graph.

[0113] As an example of the present invention, refer to Figure 3 As shown, in this example, step S2 includes:

[0114] Step S21: obtaining breeding instructions input by the user;

[0115] Step S22: performing instruction text translation on the breeding instruction input by the user to generate user input text data; performing instruction parsing on the user input text data to generate initial breeding instruction parsed data;

[0116] Step S23: extracting the syntax of the initial breeding instruction parsed data, and using the syntax to perform semantic modeling on the initial breeding instruction parsed data to generate breeding semantic recognition data;

[0117] Step S24: Match and search the breeding semantic recognition data with the structured breeding knowledge graph and vector repository, analyze the part-of-speech ratio of the breeding semantic recognition data, thereby confirming the breeding intention of the breeding semantic recognition data and generating breeding intention mapping data; input the breeding intention mapping data and the user-input breeding instructions into the breeding-specific basic model for breeding intention model inference to generate user breeding intention data.

[0118] In embodiments of the present invention, a concise and intuitive user interface (such as a web interface or app) allows users to submit breeding instructions through keyboard input, voice input, or a selection box. The system supports multiple input formats, including natural language text, standardized templates, and voice recognition. The system receives user input instructions through the interface and initializes the input data structure. The instruction content is converted into a computer-processable format (such as JSON or XML). Natural language processing technology is used to translate and convert the user-entered instruction text. This step can handle instruction text in different languages ​​and dialects. For example, if the user enters instructions in Chinese, the system first translates them into standard breeding instruction syntax. If the user enters instructions via voice, voice recognition technology is used to convert the speech into text. Standardized user-entered text data is generated to ensure that the text content is unambiguous in subsequent processing. Irrelevant content (such as colloquialisms and modal particles) in the user-entered text is cleaned up. Input text of different formats and lengths is standardized to ensure uniformity and ease of parsing. Dependency parsing technology is used to parse the user-entered text. By analyzing the relationships between words in a sentence, the subject, predicate, object, attributive, and adverbial modifiers are identified. For example, the sentence "Select drought-resistant varieties" is parsed as follows: the verb "select" directs the instruction, and the target is "drought-resistant varieties." Key entities, such as variety, climate conditions, and yield targets, are extracted from the sentence. Syntactic analysis is used to identify the user's intent. For example, a user might express the intent to "select a variety" or "increase yield." The parsed results are converted into structured instruction data for easier processing. For example, {"action": "select", "target": "drought-resistant varieties", "condition": "drought-resistant"}. Key grammatical elements, such as verbs, nouns, and adjectives, are extracted from the instruction parsed data. For example, in the instruction "Select drought-resistant varieties," the verb is "select," the noun is "variety," and the adjective is "drought-resistant." Relationships between the various components of the instruction are identified, such as the relationship between the verb and its target, and the relationship between the condition and the target. Based on the syntactic analysis results, deep learning models (such as BERT and GPT) are used for semantic understanding and modeling, converting the instruction content into a semantic vector representation. Semantic embedding technology is used to model each command, converting it into a vector so the system can understand its underlying meaning. Deep learning models are used to process the commands at a semantic level, identifying the breeding intent expressed by the user. For example, the breeding intent in "select drought-resistant varieties" may refer to variety selection, targeting drought-resistant varieties. Based on the results of semantic modeling, breeding semantic recognition data is generated.The data structure can include: {"action": "select", "target": "variety", "attributes": {"drought resistance": "strong"}}. This concise and clear data format facilitates subsequent reasoning and decision-making. The breeding semantic recognition data is matched and retrieved from the graph vector repository. Using embedding vector technology, the repository searches for breeding knowledge with semantically similar meanings to the input instruction. The system identifies knowledge graph elements relevant to the breeding instruction by calculating text similarity and part-of-speech ratios. For example, after identifying "variety with strong drought resistance" in the instruction, the system can retrieve relevant drought-resistant varieties and their characteristics. The proportion of different parts of speech (such as verbs, nouns, and adjectives) in the instruction text can be analyzed to better understand the user's breeding intent. For example, in the sentence "select variety with strong drought resistance," "select" is a verb with a large proportion, indicating the primary intention is selection, while "strong drought resistance" is an adjective, representing a condition. Based on the part-of-speech analysis and matching results, the user's breeding intent, such as variety selection or drought resistance enhancement, is confirmed. The user's breeding instructions are mapped with relevant breeding knowledge and models in the database to generate breeding intention mapping data. For example: {"intended_action": "Select", "target": "Variety", "condition": "Highly drought-resistant"}. The generated breeding intention mapping data, along with the original breeding instructions, is input into a dedicated breeding basic model. This basic model can utilize a deep learning-based neural network model with reasoning and decision-making capabilities. The model analyzes the input data and infers the action path that best meets the user's breeding intention. For example, it determines whether the user needs to further screen varieties, select a planting area, or conduct an environmental adaptability analysis. Based on the input data and inference results, the basic model generates detailed user breeding intention data, including specific breeding decisions and recommendations. For example: {"action": "Select", "target": "Highly drought-resistant varieties", "next_step": "Evaluate variety adaptability"}.

[0119] Preferably, in step S24, analyzing the part-of-speech ratio of the breeding semantic identification data to confirm the breeding intention of the breeding semantic identification data includes:

[0120] Perform part-of-speech tagging on breeding semantic recognition data to generate breeding semantic part-of-speech tagging data;

[0121] Statistical analysis was performed on breeding semantic part-of-speech tagging data to calculate the proportion of each part-of-speech type and construct a high-dimensional part-of-speech vector feature matrix. Part-of-speech categories include nouns, verbs, adjectives, proper nouns, question words, and conjunctions.

[0122] Set the following logical rules to perform numerical judgment and intent recognition on the high-dimensional part-of-speech vector feature matrix:

[0123] When the high-dimensional part-of-speech vector feature matrix has a noun ratio of <40%, a verb ratio of >30%, and an adjective ratio of >15%, it is identified as a representational breeding intention;

[0124] When the combined proportion of proper nouns and technical verbs in the high-dimensional part-of-speech vector feature matrix is ​​greater than 50% and contains the keywords "improve" and "resistance", it is identified as an improvement breeding intention;

[0125] When the total proportion of question words and conditional clause structures in the high-dimensional part-of-speech vector feature matrix is ​​greater than 20% and contains hypothetical structures, it is identified as exploratory breeding intention;

[0126] The data of representation-based breeding intentions, improvement-based breeding intentions and exploratory breeding intentions are integrated to generate breeding intention mapping data.

[0127] In this embodiment, semantic data is tagged word-by-word using natural language processing tools (such as spaCy, HanLP, LTP, and the BERT-CRF model). Parts of speech include, but are not limited to, nouns, verbs, adjectives, proper nouns, interrogative words (Wh-words, such as "whether" and "which"), and conjunctions (such as "and" and "or"). This generates structured breeding semantic part-of-speech tagged data, which is then categorized and counted to calculate the frequency of each part of speech. After normalizing the statistical results (e.g., percentage), a high-dimensional part-of-speech vector feature matrix is ​​generated: {"noun": 35.2, "verb": 38.6, "adjective": 17.3, "proper noun": 21.5, "interrogative word": 8.2, "conjunction": 4.0}. The breeding intent type is determined based on the proportion of parts of speech in the feature matrix, combined with logical rules. The rules are as follows: Condition: Noun proportion <40%, verb proportion >30%, and adjective proportion >15%. This type of instruction emphasizes the description of the target trait and is biased towards tasks such as trait identification, expression, and data labeling. Condition: Proper nouns + technical verbs (such as "transgenic," "mutagenized," and "backcross") account for >50%, including keywords such as "improve," "enhance," "resistance," and "adaptability." This type of intent focuses on optimizing breeding goals and improving performance and is commonly used in variety improvement tasks. Condition: The combined proportion of question words + conditional clause structures is >20% and contains hypothetical structures (such as "if...whether..."). This type of intent emphasizes exploring the unknown and is suitable for experimental hypothesis verification, strategy selection, or intelligent question and answering. The recognition results are integrated into structured labels and output as breeding intent mapping data, which serves as input for subsequent model inference.

[0128] Preferably, in step S3, calling germplasm resources based on the user's breeding intention data through a preset differential privacy-protected federated learning framework and performing phenotype-genotype joint analysis includes:

[0129] The user's breeding intention data is protected for data privacy through a preset differentially private federated learning framework to generate user breeding intention protection data. Resource matching is performed based on the user's breeding intention protection data to obtain germplasm resources, where resource matching includes phenotypic data resource matching, genotypic data resource matching, and phenotype-genotype data resource matching. Multi-channel phenotypic representation modeling is performed on the germplasm resources to generate a multidimensional trait vector set.

[0130] Perform high-dimensional genomic mosaic mapping on the multidimensional trait vector set to generate a nested map of genotype-phenotype associations; perform structure-preserving variation deconstruction on the nested map of genotype-phenotype associations to generate characteristic mutation location data;

[0131] Perform target trait allele inference on characteristic mutation mapping data to generate a set of functional genetic loci; quantify the pan-genetic group value of the functional genetic loci set to generate a set of multidimensional optimization factors for germplasm resources;

[0132] The multi-dimensional optimization factor set of germplasm resources is intelligently decoupled with intention weighting to generate germplasm resource screening data.

[0133] In this embodiment of the present invention, differential privacy technology is incorporated into the transmission and processing of user breeding intention data to protect user privacy by introducing noise (such as Laplace noise or Gaussian noise). The quantification of this noise is based on the user's privacy protection requirements and specific risk assessment. Thanks to differential privacy protection, the user breeding intention data generated by the system does not contain any information that directly identifies the user or specific breeding requirements, ensuring that sensitive information is not leaked during data processing and analysis. Using a federated learning framework, data is processed and trained locally, avoiding the transmission of user data to a central server, further enhancing privacy protection. Each local node (such as an agricultural research institute or farm) trains a model based on differentially private user data, calculates updated model parameters, and synchronizes them to a central server, ultimately integrating the global model. Using the target trait information in the user's breeding intention data, phenotypic data resources relevant to the user's requirements are matched. For example, a breeding instruction targeting "high drought resistance" can match phenotypic data related to drought resistance. Combined with genotypic information, genotypic data matching is performed to meet the user's breeding requirements. By correlating genotypes with phenotypes, genotypic information that meets target breeding requirements is obtained. Phenotypic and genotypic information is further combined for joint matching to identify germplasm resources that meet specific trait and genotypic requirements. Using existing breeding databases and linked data, the matching of genotypes and phenotypes is calculated and confirmed. Phenotypic data, genotypic data, and combined phenotype-genotype data are centrally stored to establish a germplasm resource database. This database should support rapid query and filtering to facilitate subsequent multi-channel modeling and analysis. Multiple dimensions of trait information, such as growth rate, disease resistance, and adaptability, are extracted from different phenotypic data. Multi-channel modeling is performed on the phenotypic data based on the characteristics of different phenotypic traits. For example, one channel can be used for growth traits, another for drought resistance, and another for yield. Deep learning techniques (such as autoencoders and convolutional neural networks) are used to convert the phenotypic characteristics of each channel into a multidimensional trait vector set. The phenotypic characteristics of each variety are represented as a vector, with the dimensions of the vector representing the characteristics of the different traits. Data from all channels are fused and the multidimensional trait vector set is standardized to ensure data consistency and comparability. The relationship between genotype and phenotype is constructed through association analysis between genotype data and the multidimensional trait vector set. Regression analysis, correlation analysis, or machine learning models (such as random forests and support vector machines) are used to identify the association between genotype and phenotype. Genotype information is combined with phenotypic data to generate a high-dimensional mosaic map. Each node represents a genotype or phenotype, and each edge represents the relationship between the two. A nested map of genotype-phenotype associations is generated, where each node in the map has multidimensional features and the edge weight represents the degree of association between the genotype and phenotype.Graphs are optimized using methods such as graph convolutional neural networks (GCNs) to maximize the connectivity of each node and edge in the graph. Based on the structural information in the graph, variation deconstruction analysis is performed. Variant detection techniques are used to identify the location of mutations in the genotype and their impact on the phenotype. High-throughput genomic data and variant detection algorithms (such as SNP detection and GWAS analysis) are used to determine the location of mutations in the genotype associated with a specific phenotype. All relevant signature mutations are annotated in the graph to generate signature mutation mapping data. Each mutation and its associated trait information are annotated on the graph for subsequent analysis. Based on signature mutation mapping data, allele information associated with the target trait is inferred. For example, alleles associated with the trait "drought resistance" can be inferred. Functional loci associated with target traits (such as drought resistance and disease resistance) are identified from the variation data. The relationship between functional genetic loci and target traits is clearly mapped to form a functional genetic locus set. These loci will serve as key indicators for breeding selection. Based on the functional genetic locus set, the genetic value of each germplasm resource is quantitatively assessed. Use mathematical models or machine learning methods (such as genetic algorithms and deep neural networks) to evaluate the optimization potential of each germplasm resource. Evaluate multiple traits of each germplasm resource to generate multi-dimensional optimization factors. Integrate factors such as genetic value, trait characteristics, and environmental adaptability to construct a multi-dimensional optimization factor set. Each factor represents a optimization criterion to facilitate breeding decisions. Design a weighted model based on the user's breeding intentions and the optimization factors of the germplasm resources. Use intelligent algorithms (such as weighted decision trees and neural networks) to weight each factor. Adjust the weight of the optimization factors according to different breeding intentions, perform intelligent decoupling, and generate the germplasm resource screening results that best meet the user's needs. Generate the final germplasm resource screening data based on the weighted optimization factor set. This data can be used by breeding experts or systems for further analysis and decision-making.

[0134] Of particular importance, the structure-preserving variation deconstruction of the nested genotype-phenotype association map also includes:

[0135] Perform node module layering on the genotype-phenotype association nested map, dividing the gene subnodes in the map into first-level modules according to chromosome numbers and into second-level modules according to functional domain relationships. The number of nodes in each module should be between 20 and 200, maintaining the topological continuity of the original map, and obtaining modular nested map data.

[0136] The modular nested graph data was subjected to structure-preserving rarefaction, retaining the genotype core path and its nearest neighbor edges, removing redundant edge connections with node degrees below 2, ensuring that the graph redundancy rate was less than 15%, and generating a structure-preserving graph subset.

[0137] Confidence-driven mutation hotspot detection was performed on the retained structural map subset, with the minimum confidence threshold of mutation-associated edges set to 0.85. Potential mutation paths were screened, candidate mutation signal regions were identified, and a hotspot mutation region map was generated.

[0138] Perform graph structure deconstruction analysis on candidate mutation signal regions and use the structure-preserving variation deconstruction algorithm to calculate the topological stability index of the mutation node. The topological stability index is set in the range of 0.2–1.0. Mutations below 0.4 are marked as structurally unstable mutations, generating a structural deconstruction evaluation matrix.

[0139] We extract gene mutation feature vectors for structurally unstable mutation nodes. The feature vectors are 128-dimensional and contain multidimensional descriptors such as base substitution probability distribution, functional region displacement parameters, and local flux change rates. We use UMAP for dimensionality reduction and visualization to locate key mutation sites and generate a feature mutation vector map.

[0140] The characteristic mutation vector map was clustered and located using the DBSCAN clustering algorithm. ε was set to 0.5, the minimum number of samples within a cluster was set to 10, and the center of gravity analysis was performed on each mutation cluster to generate the final characteristic mutation location data.

[0141] In this embodiment of the present invention, the input is nested graph data containing the relationship between genotype and phenotype, with nodes including genes, phenotypes, and regulatory factors. Nodes are grouped according to the chromosome numbers of their gene subnodes to form first-level modules. Each first-level module corresponds to a segment of chromosome numbers. Within the first-level modules, the relationships between gene functional domains (such as transcription factors, signaling pathways, and metabolic genes) are further refined into second-level modules. The number of nodes in each module is controlled between 20 and 200 to ensure appropriate module granularity and facilitate subsequent analysis. During the partitioning process, main paths and high-weight connections across modules are retained to ensure that the global topology of the graph is not disrupted, thus forming modular nested graph data. High-confidence paths between genes and phenotypes are identified as core paths. The nearest neighbor connections of nodes on the core paths are retained to maintain the representativeness of the local topology. Edge connections with a connectivity degree less than 2 (such as isolated paths and weakly connected edges) are deleted. A rarefaction operation is used to control the redundancy rate of the graph to less than 15%, forming a subset of the graph that retains the structure. The minimum confidence threshold for edges linking gene variants to phenotypes is set to 0.85. A depth-first search is performed on edges that meet the criteria to screen for potential mutation transmission pathways. High-confidence mutation pathways identified are aggregated and identified as candidate mutation signal regions, and a hotspot mutation region map corresponding to the candidate regions is generated. A structure-preserving mutation deconstruction algorithm is applied to the hotspot regions. The topological stability index (TSI) of the mutation nodes within the region is calculated, with a range of 0.2–1.0. Nodes with a TSI below 0.4 are marked as structurally unstable mutation nodes. All mutation nodes and their TSI indicators are recorded and output as a structural deconstruction assessment matrix. A 128-dimensional feature vector is extracted for each structurally unstable mutation node, primarily containing the base substitution probability distribution (e.g., A>G, C>T), functional region displacement parameters (e.g., promoter migration distance), and local flux change rate (e.g., neighborhood signal variation intensity). The Unified Mapping (UMAP) algorithm is used to map the 128-dimensional vector to a two-dimensional plane for visualization and analysis. Clustered regions are identified on the UMAP map, and mutation hotspots are preliminarily located. The DBSCAN algorithm was applied to the characteristic mutation vector map with parameters set to ε = 0.5 and a minimum number of samples within a cluster of 10. Multiple mutation clusters were identified and outliers were filtered out. The center point of each cluster was calculated and mapped back to the original map node for biological interpretation. The center point location data for each mutation cluster was generated, which became the final characteristic mutation location data.

[0142] Preferably, performing dynamic parent matching on germplasm resource screening data in step S3 includes:

[0143] Extract parental pairing trait characteristics from germplasm resource screening data, evaluate the correlation between parents, and generate a parental similarity map; perform dynamic gene combination optimization on the parental similarity map to generate a parental genotype matching matrix;

[0144] Perform environmental adaptability analysis on the parental genotype matching matrix to generate parental environmental adaptability data;

[0145] Perform multi-dimensional priority sorting on the parental genotype matching matrix according to the parental environmental adaptability data to generate priority matching parent candidate data;

[0146] The genetic diversity of the priority matching parent candidate data is verified to generate breeding parent matching data.

[0147] In the embodiment of the present invention, the core traits required for pairing are extracted from the germplasm screening data obtained by screening, such as: phenotype: plant height, stress resistance, growth period, yield, grain quality; genotype: QTL locus distribution, SNP density, dominant and recessive gene locus expression, etc. The pairing trait similarity function is used to perform bidirectional pairing scoring, for example: , where For the Local similarity function of individual traits; To automatically adjust trait weights based on breeding objectives, a parental similarity map is generated in the form of a graph or similarity matrix. Genotype optimization is performed on the similarity map: a genetic algorithm (GA) or reinforcement learning (RL) model is used to simulate the mating paths of parental combinations and optimize the allelic recombination potential. The genomic complementarity function (GCF) is introduced: GCF(A,B) = number of effective recombination sites ÷ total number of effective alleles. Regions of linkage disequilibrium and negatively interacting gene pairs (epistasis conflict) are considered and screened out to obtain a parental genotype matching matrix containing the genetic complementarity potential score for each parental combination. Parental adaptability is assessed based on the original ecological data of the germplasm and climate and soil data of the target breeding area (which can be connected to a remote sensing agricultural database): historical environmental response models are matched; time series environmental response data are modeled using LSTM or Transformer architectures; and environmental response scores are calculated based on the parents' past field test performance. A multi-objective prioritization function is constructed, for example: ;in , , , Adjustable weights are assigned based on user breeding intent; the total priority score for each parent pair is used for ranking; and visual adjustments based on trait strength are supported to generate a dataset of candidate parent matches. Each parent combination comes with priority ranking information and predicted trait distribution. Genetic diversity analysis is performed on candidate parent combinations: using population genetics metrics such as Nei's genetic distance (Nei's D), population structure coefficient (Fst), principal component clustering (PCA), and t-SNE projections to filter out close-kin and highly homologous combinations. Prioritizing combinations with a broad genetic base is strengthened, generating a final breeding parent matching dataset. Each pair is accompanied by genetic distance, trait complementarity map, matching priority index, and a visual genetic diversity score chart.

[0148] Preferably, the step S4 of evaluating the feasibility of the breeding planning simulation data includes:

[0149] Import the breeding planning simulation data into the preset biological breeding management system for field extraction and field format verification to ensure that the integrity of each data item is ≥98%, and generate a structured planning simulation data set. The extracted fields include the probability of achieving the target trait, expected number of generations, genetic diversity index, stability coefficient, and environmental adaptation value;

[0150] The target trait achievement probability field in the structured planning simulation data set is screened using a probability threshold, with the minimum threshold set to 0.75 and the maximum threshold set to 1.0, to screen out eligible individual paths. If the target trait achievement probability is lower than 0.75, the path is marked as having no achievement potential, and a potential achievement screening result graph is generated.

[0151] The genetic diversity of the pathway data that passed the target trait screening was evaluated and scored using the Shannon diversity index calculation formula, with the minimum Shannon index set at 0.30. If the Shannon index of an individual pathway was lower than this value, it was marked as having a genetic bottleneck, and a genetic diversity score map was generated.

[0152] Perform genetic stability analysis on qualified genetic diversity paths and calculate the standard deviation σ. The σ range is set to 0.00 to 0.20. If the σ value of a path exceeds 0.20, it is judged to have large genetic fluctuations and is not recommended for execution. Finally, a stability judgment matrix diagram is generated;

[0153] Environmental adaptability analysis was performed on the paths that passed the stability screening. The environmental adaptability score of each path was compared with the target planting area environment using the environmental index matching calculation formula, with a matching score threshold set at ≥0.70. If the path score was lower than 0.70, it was marked as having low adaptability, and an environmental adaptability screening image was generated.

[0154] All paths that pass the screening are scored using a multi-indicator weighted approach, with the weights for each indicator set as follows: 30% for trait achievement rate, 25% for diversity, 20% for stability, and 25% for environmental adaptability. Weighted scores are assigned based on a normalized scoring formula, ranging from 0 to 1.0, with a recommended execution threshold set at ≥0.75. This ultimately generates a weighted scoring result for the breeding planning path.

[0155] The weighted scoring results of the breeding planning path are visualized as a score distribution graph, and the output is the feasibility assessment result.

[0156] The present invention is therefore intended to be illustrative and non-restrictive in all respects, with the scope of the invention being defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the application documents are intended to be embraced therein.

[0157] The foregoing description is intended only to provide specific embodiments of the present invention, which will enable those skilled in the art to understand and implement the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not intended to be limited to the embodiments shown herein, but is to be construed in the widest possible manner consistent with the principles and novel features disclosed herein.

Claims

1. An intelligent breeding planning and decision-making method based on a large model, characterized in that: The following steps are involved: Step S1: Acquire a multi-source breeding dataset; construct a structured breeding knowledge graph and vector repository based on the multi-source breeding dataset; According to the preset large model, the structured breeding knowledge graph is associated with breeding data to generate a breeding-specific basic model; Step S2: Obtaining breeding instructions input by the user; performing user semantic recognition on the breeding instructions input by the user to generate breeding semantic recognition data; performing matching retrieval on the breeding semantic recognition data and the structured breeding knowledge graph and the vector repository, and inputting the retrieval results and the user-input breeding instructions into a breeding-specific basic model to perform breeding intention analysis and generate user breeding intention data; Step S3: Based on the user's breeding intention data, germplasm resources are called through a preset differentially private federated learning framework and phenotype-genotype joint analysis is performed to generate germplasm resource screening data, wherein the phenotype-genotype joint analysis includes phenotype data resource matching, genotype data resource matching, and phenotype-genotype data resource matching; parents are dynamically matched on the germplasm resource screening data to generate breeding parent matching data; based on the breeding parent matching data, multi-generation pedigree path simulation is performed on the user's breeding intention data to generate breeding planning simulation data; wherein, the multi-generation pedigree path simulation of the user's breeding intention data based on the breeding parent matching data in step S3 includes: Reconstruct trait-driven genetic coding of breeding parent matching data to generate parent genetic representation unit sets; Map the target trait-oriented path of the user's breeding intention data to generate the intended trait evolution trajectory data; Based on the parental genetic representation unit set and the intended trait evolution trajectory data, intergenerational recombination probability modeling is performed to generate a genetic recombination transfer matrix; Perform multi-generational simulation of genetic transmission evolution on the genetic recombination transfer matrix to generate a path-level pedigree evolution map; Through the path-level pedigree evolution map, the trait achievement rate inversion analysis of the genetic recombination transfer matrix is ​​performed to generate a multi-path breeding achievement rate matrix; The target driving probability of the multi-path breeding achievement rate matrix is ​​calculated by clustering to obtain the breeding planning simulation data; Step S4: Evaluate the executability of the breeding planning simulation data, and output the breeding planning simulation data by path sorting according to the executability evaluation results to obtain the intelligent breeding decision results.

2. The intelligent breeding planning and decision-making method based on a large model according to claim 1, characterized in that: Step S1 includes the following steps: Step S11: Acquire a multi-source breeding dataset, wherein the multi-source breeding dataset includes approved variety data, genotype-phenotype association data, environmental response parameters, and a literature knowledge base; Step S12: performing data preprocessing on the multi-source breeding dataset to generate a standard multi-source breeding dataset, wherein the data preprocessing includes data cleaning, data denoising, missing value filling and data standardization; Step S13: extracting breeding features from a standard multi-source breeding dataset to construct a structured breeding knowledge graph; Step S14: embedding multi-dimensional entities and relationships in the structured breeding knowledge graph according to the preset large model to generate graph vectorized representation data; performing data association and storage on the graph vectorized representation data based on the graph neural network and RAG engine to generate semantically enhanced breeding association data and a vector repository; Step S15: Based on the preset domain adaptation training corpus, the semantically enhanced breeding-related data is subjected to structural compression and model mapping to generate a breeding-specific basic model.

3. The intelligent breeding planning and decision-making method based on a large model according to claim 2, characterized in that: Step S13 includes the following steps: Step S131: extracting breeding characteristics from a standard multi-source breeding data set to obtain initial breeding characteristic data; Step S132: semantically classifying the initial breeding feature data to generate semantically classified feature data; performing ontology mapping on the semantically classified feature data to generate feature ontology mapping data; Step S133: identifying entities in the feature ontology mapping data, and constructing relationships between the identified entities to generate candidate knowledge triple data; Step S134: perform knowledge fusion and redundancy elimination on the candidate knowledge triple data to generate a refined breeding knowledge triple set; construct a graph based on the refined breeding knowledge triple set, use the breeding entities in the refined breeding knowledge triple set as nodes, and the breeding relationships as edges to generate a structured breeding knowledge graph.

4. The intelligent breeding planning and decision-making method based on a large model according to claim 1, characterized in that: Step S2 includes the following steps: Step S21: obtaining breeding instructions input by the user; Step S22: performing instruction text translation on the breeding instruction input by the user to generate user input text data; performing instruction parsing on the user input text data to generate initial breeding instruction parsed data; Step S23: extracting the syntax of the initial breeding instruction parsed data, and using the syntax to perform semantic modeling on the initial breeding instruction parsed data to generate breeding semantic recognition data; Step S24: Match and search the breeding semantic recognition data with the structured breeding knowledge graph and vector repository, analyze the part-of-speech ratio of the breeding semantic recognition data, thereby confirming the breeding intention of the breeding semantic recognition data and generating breeding intention mapping data; input the breeding intention mapping data and the user-input breeding instructions into the breeding-specific basic model for breeding intention model inference to generate user breeding intention data.

5. The large-scale model-based intelligent breeding planning and decision-making method according to claim 4, characterized in that: In step S24, the proportion of parts of speech of the breeding semantic identification data is analyzed to confirm the breeding intention of the breeding semantic identification data, including: Perform part-of-speech tagging on breeding semantic recognition data to generate breeding semantic part-of-speech tagging data; Statistical analysis was performed on breeding semantic part-of-speech tagging data to calculate the proportion of each part-of-speech type and construct a high-dimensional part-of-speech vector feature matrix. Part-of-speech categories include nouns, verbs, adjectives, proper nouns, question words, and conjunctions. Set the following logical rules to perform numerical judgment and intent recognition on the high-dimensional part-of-speech vector feature matrix: When the high-dimensional part-of-speech vector feature matrix has a noun ratio of <40%, a verb ratio of >30%, and an adjective ratio of >15%, it is identified as a representational breeding intention; When the combined proportion of proper nouns and technical verbs in the high-dimensional part-of-speech vector feature matrix is ​​greater than 50% and contains the keywords "improve" and "resistance", it is identified as an improvement breeding intention; When the total proportion of question words and conditional clause structures in the high-dimensional part-of-speech vector feature matrix is ​​greater than 20% and contains hypothetical structures, it is identified as exploratory breeding intention; The data of representation-based breeding intentions, improvement-based breeding intentions and exploratory breeding intentions are integrated to generate breeding intention mapping data.

6. The large-scale model-based intelligent breeding planning and decision-making method according to claim 1, characterized in that: In step S3, based on the user's breeding intention data, germplasm resources are called through a preset differential privacy-protected federated learning framework and phenotype-genotype joint analysis is performed, including: The user's breeding intention data is protected for data privacy through a preset differentially private federated learning framework to generate user breeding intention protection data. Resource matching is performed based on the user's breeding intention protection data to obtain germplasm resources, where resource matching includes phenotypic data resource matching, genotypic data resource matching, and phenotype-genotype data resource matching. Multi-channel phenotypic representation modeling is performed on the germplasm resources to generate a multidimensional trait vector set. Perform high-dimensional genomic mosaic mapping on the multidimensional trait vector set to generate a nested map of genotype-phenotype associations; perform structure-preserving variation deconstruction on the nested map of genotype-phenotype associations to generate characteristic mutation location data; Perform target trait allele inference on characteristic mutation mapping data to generate a set of functional genetic loci; quantify the pan-genetic group value of the functional genetic loci set to generate a set of multidimensional optimization factors for germplasm resources; The multi-dimensional optimization factor set of germplasm resources is intelligently decoupled with intention weighting to generate germplasm resource screening data.

7. The intelligent breeding planning and decision-making method based on a large model according to claim 1, characterized in that: Performing dynamic parent matching on germplasm resource screening data in step S3 includes: Extract parental pairing trait characteristics from germplasm resource screening data, evaluate the correlation between parents, and generate a parental similarity map; perform dynamic gene combination optimization on the parental similarity map to generate a parental genotype matching matrix; Perform environmental adaptability analysis on the parental genotype matching matrix to generate parental environmental adaptability data; Perform multi-dimensional priority sorting on the parental genotype matching matrix according to the parental environmental adaptability data to generate priority matching parent candidate data; The genetic diversity of the priority matching parent candidate data is verified to generate breeding parent matching data.

8. The large-scale model-based intelligent breeding planning and decision-making method according to claim 1, characterized in that: The feasibility of evaluating the breeding planning simulation data in step S4 includes: Import the breeding planning simulation data into the preset biological breeding management system for field extraction and field format verification to ensure that the integrity of each data item is ≥98%, and generate a structured planning simulation data set. The extracted fields include the probability of achieving the target trait, expected number of generations, genetic diversity index, stability coefficient, and environmental adaptation value; The target trait achievement probability field in the structured planning simulation data set is screened using a probability threshold, with the minimum threshold set to 0.75 and the maximum threshold set to 1.0, to screen out eligible individual paths. If the target trait achievement probability is lower than 0.75, the path is marked as having no achievement potential, and a potential achievement screening result graph is generated. The genetic diversity of the pathway data that passed the target trait screening was evaluated and scored using the Shannon diversity index calculation formula, with the minimum Shannon index set at 0.

30. If the Shannon index of an individual pathway was lower than this value, it was marked as having a genetic bottleneck, and a genetic diversity score map was generated. Perform genetic stability analysis on qualified genetic diversity paths and calculate the standard deviation σ. The σ range is set to 0.00 to 0.

20. If the σ value of a path exceeds 0.20, it is judged to have large genetic fluctuations and is not recommended for execution. Finally, a stability judgment matrix diagram is generated; Environmental adaptability analysis was performed on the paths that passed the stability screening. The environmental adaptability score of each path was compared with the target planting area environment using the environmental index matching calculation formula, with a matching score threshold set at ≥0.

70. If the path score was lower than 0.70, it was marked as having low adaptability, and an environmental adaptability screening image was generated. All paths that pass the screening are scored using a multi-indicator weighted approach, with the weights for each indicator set as follows: 30% for trait achievement rate, 25% for diversity, 20% for stability, and 25% for environmental adaptability. Weighted scores are assigned based on a normalized scoring formula, ranging from 0 to 1.0, with a recommended execution threshold set at ≥0.

75. This ultimately generates a weighted scoring result for the breeding planning path. The weighted scoring results of the breeding planning path are visualized as a score distribution graph, and the output is the feasibility assessment result.

9. An intelligent breeding planning and decision-making system based on a large model, characterized by: For executing the large-model-based intelligent breeding planning and decision-making method according to claim 1, the large-model-based intelligent breeding planning and decision-making system comprises: A breeding map construction module is used to obtain multi-source breeding data sets; build a structured breeding knowledge map and vector repository based on the multi-source breeding data sets; associate breeding data with the structured breeding knowledge map according to a preset large model to generate a breeding-specific basic model; The breeding instruction recognition module is used to obtain the breeding instructions input by the user; perform user semantic recognition on the breeding instructions input by the user to generate breeding semantic recognition data; match and search the breeding semantic recognition data with the structured breeding knowledge graph and vector repository, and input the search results and the user-input breeding instructions into the breeding-specific basic model for breeding intention analysis to generate user breeding intention data; The breeding planning module is used to call germplasm resources based on user breeding intention data through a preset differentially privacy-protected federated learning framework and perform phenotype-genotype joint analysis to generate germplasm resource screening data, where the phenotype-genotype joint analysis includes phenotypic data resource matching, genotype data resource matching, and phenotype-genotype data resource matching; perform dynamic parent matching on germplasm resource screening data to generate breeding parent matching data; perform multi-generation pedigree path simulation on user breeding intention data based on breeding parent matching data to generate breeding planning simulation data; where the multi-generation pedigree path simulation on user breeding intention data based on breeding parent matching data includes: Reconstruct trait-driven genetic coding of breeding parent matching data to generate parent genetic representation unit sets; Map the target trait-oriented path of the user's breeding intention data to generate the intended trait evolution trajectory data; Based on the parental genetic representation unit set and the intended trait evolution trajectory data, intergenerational recombination probability modeling is performed to generate a genetic recombination transfer matrix; Perform multi-generational simulation of genetic transmission evolution on the genetic recombination transfer matrix to generate a path-level pedigree evolution map; Through the path-level pedigree evolution map, the trait achievement rate inversion analysis of the genetic recombination transfer matrix is ​​performed to generate a multi-path breeding achievement rate matrix; The target driving probability of the multi-path breeding achievement rate matrix is ​​calculated by clustering to obtain the breeding planning simulation data; The decision output module is used to evaluate the executability of the breeding planning simulation data, and output the breeding planning simulation data in a path sorting manner according to the executability evaluation results to obtain the intelligent breeding decision results.