Digital planting management method and system for good seed breeding

CN122548331APending Publication Date: 2026-08-11GAOZHOU CITY BREED BREEDING FARM
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-30
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0005]本申请实施例通过提供面向良种繁育的数字化种植管理方法及系统,解决了现有良种繁育中品种数量多、单个品种杂株标注样本稀缺,导致纯度鉴定模型难以快速适配新品种的技术问题

Benefits of technology

本申请实施例通过提供面向良种繁育的数字化种植管理方法及系统,首先,以繁育品种为节点搭建包含遗传关系、扩繁世代、形态相似度的作物系谱知识图谱,能够整合已有的品种系谱与形态信息,为新品种的迁移源域选取提供全面的关联依据。其次,结合遗传关系和形态相似度确定近缘品种,相比仅依靠遗传关系选取迁移源域的方法,能够选择出形态更接近待管理新品种的源域品种,有效提升迁移适配后模型的精度。最后,结合当前繁育世代提取对应的标准形态特征,将其作为约束引入模型微调过程,能够适配不同繁育世代的纯度要求,仅依靠少量标注样本即可快速得到精度满足要求的纯度鉴定模型。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122548331A_ABST
    Figure CN122548331A_ABST
Patent Text Reader

Abstract

This application discloses a digital planting management method and system for improved variety breeding, relating to the field of seed and seedling cultivation technology. The method includes: constructing a crop pedigree knowledge graph and obtaining the variety identifier and current breeding generation of the variety to be managed; locating the variety node corresponding to the variety to be managed in the crop pedigree knowledge graph; performing a graph traversal with a limited number of hops along genetic and feature edges to extract a set of closely related varieties; simultaneously extracting associated standard morphological feature descriptions; selecting varieties with labeled purity training data as the migration source domain based on the phylogenetic distance between each closely related variety in the set and the variety to be managed; and performing migration adaptation using the corresponding pre-trained purity identification model as the base network to obtain a purity identification model. This solves the technical problem in existing improved variety breeding where the large number of varieties and the scarcity of labeled hybrid samples for individual varieties make it difficult for purity identification models to quickly adapt to new varieties.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of seed and seedling cultivation technology, specifically to digital planting management methods and systems for the propagation of superior varieties. Background Technology

[0002] With the development of agricultural digital technology, intelligent management of the seed industry has become an inevitable trend. Among them, seed purity control in the process of breeding improved varieties is the core link to ensure the quality of seed sources.

[0003] However, traditional purity control relies heavily on manual field weed removal and the experience of breeders, which is not only inefficient but also prone to missed detections and misjudgments. In recent years, purity identification models based on deep learning have been gradually applied to hybrid identification. However, deep learning models rely on a large amount of labeled training data, and for newly bred varieties, it is difficult to quickly obtain sufficient labeled purity training data, making it difficult to meet the accuracy requirements for direct training of the model.

[0004] Furthermore, existing transfer learning methods select source domain models based solely on the genetic relationships between varieties, neglecting the correlation of morphological characteristics among varieties. The accuracy of the transferred models still has room for improvement. At the same time, they do not incorporate generational information from crop pedigrees to introduce standardized morphological constraints, making it impossible to adapt to the purity requirements of different breeding generations. Summary of the Invention

[0005] This application provides a digital planting management method and system for improved seed breeding, which solves the technical problem that the existing improved seed breeding has a large number of varieties and a scarcity of labeled samples of hybrid plants of a single variety, making it difficult for the purity identification model to be quickly adapted to new varieties.

[0006] The technical solution to the above-mentioned technical problems in this application is as follows:

[0007] Firstly, this application provides a digitalized planting management method for breeding improved varieties, the method comprising: Using breeding varieties as nodes, parent-offspring relationships as genetic edges, propagation associations from original varieties to improved varieties as generational edges, and morphological similarity among varieties as characteristic edges, a crop pedigree knowledge graph is constructed. Obtain the variety identifier and current breeding generation of the variety to be managed, and use the variety identifier as the query key to locate the variety node corresponding to the variety to be managed in the crop pedigree knowledge graph; Centered on the variety node, a graph traversal with a limited number of jumps is performed along the genetic edge and the feature edge to extract a set of closely related varieties. At the same time, the variety node corresponding to the current breeding generation is retrieved along the generation edge to extract the associated standard morphological feature description. Based on the phylogenetic distance between each closely related variety in the set of closely related varieties and the variety to be managed, varieties with labeled purity training data are selected as migration source domains according to the closeness priority strategy. Using the pre-trained purity identification model corresponding to the migration source domain as the base network, and employing the standard morphological feature description for migration adaptation, a purity identification model for the variety to be managed is obtained.

[0008] Secondly, this application provides a digital planting management system for breeding improved varieties, including: The knowledge graph construction module is used to construct a crop pedigree knowledge graph with breeding varieties as nodes, parent-offspring relationships as genetic edges, propagation associations from original varieties to improved varieties as generational edges, and morphological similarity between varieties as feature edges. The variety node query module is used to obtain the variety identifier and current breeding generation of the variety to be managed, and use the variety identifier as the query key to locate the variety node corresponding to the variety to be managed in the crop pedigree knowledge graph; The associated feature extraction module is used to perform a graph traversal with a limited number of jumps along the genetic edge and the feature edge, centered on the variety node, to extract a set of closely related varieties, and at the same time, to search for the variety node corresponding to the current breeding generation along the generation edge and extract the associated standard morphological feature description. The migration source domain selection module is used to select varieties with labeled purity training data as migration source domains according to the phylogenetic distance between each closely related variety in the set of closely related varieties and the variety to be managed, and according to the closeness priority strategy. The identification model acquisition module is used to obtain a purity identification model for the variety to be managed by using the pre-trained purity identification model corresponding to the migration source domain as the base network and performing migration adaptation using the standard morphological feature description.

[0009] This application provides one or more technical solutions, which have at least the following technical effects or advantages: This application provides a digital planting management method and system for improved variety breeding. First, it constructs a crop pedigree knowledge graph, including genetic relationships, propagation generations, and morphological similarity, using the breeding variety as a node. This integrates existing pedigree and morphological information, providing comprehensive correlation evidence for selecting migration source regions for new varieties. Second, it identifies closely related varieties by combining genetic relationships and morphological similarity. Compared to methods that rely solely on genetic relationships to select migration source regions, this method can select source region varieties with morphology more similar to the new variety to be managed, effectively improving the accuracy of the model after migration adaptation. Finally, it extracts corresponding standard morphological features based on the current breeding generation and introduces them as constraints into the model fine-tuning process. This allows for adaptation to the purity requirements of different breeding generations, and a purity identification model with satisfactory accuracy can be quickly obtained using only a small number of labeled samples.

[0010] Through the above technical solutions, this application solves the problem of scarce labeled samples of hybrid plants in newly bred varieties and the difficulty in quickly adapting purity identification models, and realizes digital and efficient management of improved seed breeding. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 This is a flowchart illustrating the digital planting management method for improved seed breeding provided in the embodiments of this application; Figure 2 This is a schematic diagram of the structure of a digital planting management system for breeding improved varieties provided in the embodiments of this application.

[0013] The components represented by each number in the attached diagram are explained below: Module 11: Knowledge Graph Construction; Module 12: Variety Node Query; Module 13: Association Feature Extraction; Module 14: Migration Source Domain Selection; Module 15: Identification Model Acquisition. Detailed Implementation

[0014] This application provides a digital planting management method and system for breeding improved varieties, which addresses the technical problem in existing improved variety breeding where there are many varieties and a scarcity of labeled samples of hybrid plants of individual varieties, making it difficult for purity identification models to be quickly adapted to new varieties.

[0015] Example 1, as Figure 1 As shown in the embodiments of this application, a digital planting management method for breeding improved varieties is provided, including: S10: Using breeding varieties as nodes, parent-offspring relationships as genetic edges, propagation associations from original varieties to improved varieties as generational edges, and morphological similarity between varieties as characteristic edges, a crop pedigree knowledge graph is constructed. In this embodiment of the application, a crop pedigree knowledge graph is constructed. The parent-offspring relationship can be directly obtained from the existing variety registration data. The propagation association of each generation from the original seed to the improved seed can be extracted from the breeding archives. The morphological similarity between varieties can be obtained by comparing the morphological characteristics of each registered variety and calculating based on the semantic similarity of the text.

[0016] Specifically, step S10 in the method includes: Obtain the variety name, variety approval number, parent combination information, and generational information from the variety approval information; Each variety name is associated with a variety entity as a node, and each variety node is assigned variety attributes, which include at least variety type, morphological characteristics, approval year, breeding unit, and generation level. Based on the parental combination information, a genetic edge is established between two variety nodes that have a parent-offspring relationship, with the direction of the genetic edge pointing from the parent node to the offspring node. Based on the generation level information, generation edges are established between adjacent generation variety nodes that have a propagation association. The generation edges connect the original seed node with the original seed node, and the original seed node with the improved seed node. Calculate the morphological similarity between any two variety nodes, and establish feature edges between two variety nodes whose morphological similarity exceeds a preset similarity threshold; The variety nodes, genetic edges, generation edges, and feature edges are stored in a graph database to form a crop pedigree knowledge graph.

[0017] In this embodiment of the application, firstly, variety approval information is obtained, from which variety name, variety approval number, parent combination information and generation level information are extracted. Among them, the variety entity corresponding to the variety name is used as a node, and each node is configured with basic attributes including variety type, morphological characteristics, approval year, breeding unit and generation level.

[0018] Secondly, based on the parent combination information, the parent-offspring relationship is clarified, and directional genetic edges are added. Then, based on the generation level information, generation edges are established between nodes of adjacent propagation generations. Specifically, the original seed node is connected to the original seed node obtained from its propagation, and the original seed node is then connected to the improved seed node obtained from its propagation, thus fully presenting the propagation path of the variety from basic seed source to production seed.

[0019] Next, the morphological feature text of all nodes is extracted, and a pre-trained language model is used to convert the morphological features into text embedding vectors. The cosine similarity between the embedding vectors of two nodes is calculated as the morphological similarity. When the similarity exceeds the preset similarity threshold, an undirected feature edge is added between the two nodes. Finally, all nodes and edges are stored in the graph database to obtain a crop pedigree knowledge graph that integrates genetic, generational, and morphological information.

[0020] Specifically, the morphological similarity between any two variety nodes is calculated, and feature edges are established between two variety nodes whose morphological similarity exceeds a preset similarity threshold, including: Obtain the morphological feature description field associated with the two variety nodes, input the morphological feature description field into the pre-trained feature encoding model, and output the initial morphological similarity by the feature encoding model; Obtain breeding log text, extract morphological comparison description statements related to the two variety nodes from the breeding log text, input the morphological comparison description statements into a pre-trained semantic scoring model, and output a morphological similarity correction score from the semantic scoring model. The initial morphological similarity score and the morphological similarity correction score are weighted and fused together, and the fused value is used as the morphological similarity between the two variety nodes. Obtain the current breeding generation of the variety to be managed, and determine the corresponding generation purity requirement based on the current breeding generation; A preset similarity threshold is dynamically determined based on the generation purity requirement, wherein the higher the generation purity requirement, the higher the preset similarity threshold. The morphological similarity is compared with the preset similarity threshold, and a feature edge is established between two variety nodes whose morphological similarity exceeds the preset similarity threshold.

[0021] In this embodiment of the application, firstly, the morphological feature description field associated with two variety nodes is obtained, and the morphological feature description field is converted into a fixed-dimensional feature embedding vector, which is then input into a pre-trained feature encoding model. The initial morphological similarity between the two varieties is then calculated, wherein the pre-trained feature encoding model is based on the initial feature encoding model.

[0022] Secondly, from the archived breeding log texts, content mentioning the two varieties is retrieved, descriptive statements involving morphological comparisons between the two are extracted, and input into a pre-trained semantic scoring model to obtain a morphological similarity correction score. This score can reflect the actual judgment of the morphological differences between the two varieties in the experience of breeding experts.

[0023] Next, the initial morphological similarity output from the feature encoding model and the semantic scoring model are combined with the corrected score according to preset weights to obtain the final morphological similarity. The weight coefficients are preset according to the actual application scenario, or can be adaptively adjusted according to historical labeled data.

[0024] Next, the purity requirements are determined based on the current breeding generation of the variety to be managed. The higher the generation level, such as the original seed generation, the higher the required purity, and the preset similarity threshold is also increased accordingly. Conversely, the threshold is appropriately lowered. Finally, the calculated morphological similarity is compared with the dynamically adjusted preset threshold. Feature edges are only established between two variety nodes whose similarity exceeds the threshold to ensure the reasonableness of the feature edge association. For example, if the preset similarity threshold is set in the range of 0.85 to 0.95, the threshold corresponding to the original seed generation is taken at the upper limit of the range, and the threshold corresponding to the improved seed generation is taken at the lower limit of the range, which can adapt to the association requirements of different purity requirements.

[0025] Furthermore, the training process of the feature encoding model includes: Collect morphological feature description fields from variety approval information and construct a morphological description corpus; An initial feature encoding model is selected, which is a language representation model pre-trained by a masked language modeling task; The morphological description corpus is input into the initial feature encoding model, and the initial feature encoding model is pre-trained a second time using the masked language modeling task to obtain the pre-trained initial feature encoding model. Construct a similarity-labeled dataset, which includes multiple pairs of morphological feature description fields for varieties and their corresponding similarity labels; The pre-trained initial feature encoding model is then fine-tuned using the similarity-labeled dataset to obtain the trained feature encoding model.

[0026] In this embodiment of the application, firstly, publicly available crop variety approval announcements are collected, and morphological feature description fields corresponding to each variety are extracted in batches, including multi-dimensional descriptions such as plant height, leaf shape, flower color, ear type, and grain shape, and then compiled and summarized to form a domain-specific morphological description corpus.

[0027] Secondly, a language representation model pre-trained on a general corpus-masked language modeling task is selected as the initial feature encoding model. For example, the existing Sinongda language model, which is pre-trained on a general corpus, already has basic text semantic understanding capabilities. By organizing a domain-specific morphological description corpus, the model can further learn the professional expression logic of crop morphological description and encode the semantic information of variety morphological features more accurately.

[0028] Next, the prepared morphological description corpus is input into the initial feature encoding model. The initial feature encoding model is then pre-trained a second time using the masked language modeling task, allowing the initial feature encoding model to learn professional semantic expressions in the field of crop morphological description, thus obtaining a pre-trained initial feature encoding model adapted to this task.

[0029] Next, the similarity of multiple sets of morphological descriptions is manually labeled. Each label contains a pair of complete morphological feature descriptions of the varieties and a similarity score in the range of 0 to 1 as the similarity label of the pair. A labeled similarity label dataset is constructed. Then, the constructed similarity label dataset is used to perform supervised fine-tuning on the initial feature encoding model after secondary pre-training, and the morphological similarity results are output. Finally, the trained feature encoding model is obtained.

[0030] For example, the feature encoding model is trained based on the initial feature encoding model. The specific steps are as follows: the collected morphological description corpus is divided into training set and validation set in an 8:2 ratio. The initial learning rate is set to 0.001, the batch size is set to 32, and the training rounds are set to 100 rounds. Secondary pre-training is performed on the GPU. Then, the manually labeled similarity annotation dataset is also divided into training subset and validation subset in the same ratio. The mean squared error is used as the loss function. Supervised fine-tuning is performed in the same hardware environment. Training is stopped early when the validation set loss does not decrease for 5 consecutive rounds, and the final usable feature encoding model is obtained.

[0031] Furthermore, the training process of the semantic scoring model includes: Collect morphological comparison descriptions from breeding log texts and construct a morphological comparison corpus; An initial semantic scoring model is selected, and a Sigmoid output layer is added on top of the initial semantic scoring model, wherein the initial semantic scoring model is a language representation model pre-trained by a masked language modeling task; The morphological comparison corpus is input into the initial semantic scoring model, and the initial semantic scoring model is pre-trained a second time using the masked language modeling task to obtain the pre-trained initial semantic scoring model. Construct a rating labeling dataset, which contains multiple morphological comparison description statements and corresponding similarity rating labels; The pre-trained initial semantic scoring model is fine-tuned using the aforementioned scoring annotation dataset to obtain the trained semantic scoring model.

[0032] In this embodiment of the application, firstly, publicly available and archived crop breeding logs are collected, and descriptive statements that mention the morphological comparison of two or more varieties are extracted in batches, such as "This variety is slightly taller than XX variety, and its ear type is similar to XX variety" containing direct comparison information. These statements are then organized and summarized to obtain a morphological comparison corpus.

[0033] Secondly, a language representation model pre-trained on a general corpus mask language modeling task is selected as the initial semantic scoring model, such as the existing Sinnon language model. A sigmoid output layer with an output dimension of 1 is added to the top of the model to constrain the output results to the similarity score range of 0 to 1.

[0034] Next, the prepared morphological comparison corpus is input into the initial semantic scoring model. The initial semantic scoring model is then pre-trained a second time using the masked language modeling task, allowing the initial semantic scoring model to learn the professional expression logic of variety morphological comparison in the breeding log, thus obtaining a pre-trained initial semantic scoring model adapted to this task.

[0035] Next, the descriptive statements in the morphological comparison corpus are manually annotated, and similarity score labels ranging from 0 to 1 are given according to the degree of similarity of the varieties reflected in the statements. A labeled score label dataset is constructed, and then the pre-trained initial semantic scoring model is fine-tuned in a supervised manner using this dataset, and finally the trained semantic scoring model is obtained.

[0036] For example, the semantic scoring model is trained based on the initial semantic scoring model. The specific steps are as follows: the collected morphological comparison corpus is divided into training set and validation set in a ratio of 7:3. The initial learning rate is set to 0.0005, the batch size is set to 16, and the training epochs are set to 80. Secondary pre-training is performed on the GPU. Then, the manually annotated scoring dataset is divided into training subset and validation subset in the same ratio. The mean squared error is used as the loss function, and supervised fine-tuning is performed in the same hardware environment. Training is stopped early when the validation set loss does not decrease for 8 consecutive epochs, and the trained semantic scoring model can be used directly.

[0037] S20: Obtain the variety identifier and current breeding generation of the variety to be managed, and use the variety identifier as the query key to locate the variety node corresponding to the variety to be managed in the crop pedigree knowledge graph; In this embodiment, the node entity corresponding to the variety to be managed is located by matching the variety identifier of the variety to be managed with the number information of the corresponding variety node in the crop pedigree knowledge graph. At the same time, the generation level attribute of the node is read in combination with the current breeding generation information to confirm that the node attribute information matches the actual breeding scenario. If there are multiple nodes with the same name but different generations, the target node of the corresponding generation is selected according to the breeding generation information to ensure the accuracy of node positioning.

[0038] Furthermore, since the variety nodes of each generation are stored independently, different generations, even for the same variety, correspond to different nodes. Therefore, by combining the variety identifier with the current breeding generation, the unique target node can be accurately located without any location confusion.

[0039] S30: Using the variety node as the center, perform a graph traversal with a limited number of jumps along the genetic edge and the feature edge to extract the set of closely related varieties. At the same time, search along the generation edge for the variety node corresponding to the current breeding generation and extract the associated standard morphological feature description. In this embodiment, starting from the target variety node, closely related varieties with direct or indirect parent-child genetic relationships are obtained by traversing along the genetic edge. Varieties with highly similar morphological characteristics are obtained by traversing along the feature edge. A fixed number of jumps is set in the traversal process, and only all related varieties within the jump range are extracted to obtain a set of closely related varieties, so as to avoid irrelevant varieties from interfering with subsequent analysis. At the same time, the standard morphological characteristic description of the seed source node of the previous generation generation is obtained by searching upward along the generation edge, which is used for subsequent morphological comparison and verification of plants in the field.

[0040] Specifically, step S30 in the method includes: Using the variety node obtained from the location as the starting node, initialize the set of closely related varieties as an empty set, and preset the maximum number of traversal jumps; Starting from the starting node, a breadth-first traversal is performed along the genetic edges and the feature edges. For each variety node visited during the traversal, the number of genetic edges and the number of feature edges traversed from the starting node to the variety node are recorded, and the variety node is added to the set of closely related varieties until the maximum number of traversal jumps is reached. For each closely related variety node in the set of closely related varieties, the number of genetic edges traversed to reach the closely related variety node is multiplied by a first preset weight, and then the number of feature edges traversed is multiplied by a second preset weight to obtain the kinship distance between the closely related variety node and the starting node. Search along the generation edge for the variety node corresponding to the current breeding generation, and extract the standard plant height range, standard leaf shape description, standard ear shape description and standard grain shape description that the variety to be managed should present under the current breeding generation from the variety attributes associated with the variety node, to form a standard morphological feature description.

[0041] In this embodiment, the target variety node after positioning is used as the starting point of the traversal, and the maximum number of traversal jumps is preset. For example, the maximum number of traversal jumps is set to 2 to 3 jumps, which can cover closely related and similar varieties with a high degree of correlation with the target variety, without introducing too many unrelated varieties with a distant relationship, and the set of closely related varieties is initialized to an empty set.

[0042] Secondly, starting from the starting node, a breadth-first traversal is performed along the established genetic edges and feature edges. During the traversal, each time a new variety node is visited, the total number of genetic edges and feature edges traversed from the starting node to that node is recorded. Then, the variety node is added to the set of closely related varieties. The traversal stops when the number of traversal steps reaches the preset maximum number of traversal jumps.

[0043] Furthermore, after traversing to obtain the set of closely related varieties, for each closely related variety node in the set, the kinship distance between each closely related variety node and the starting node is calculated according to the preset weight. Specifically, the number of genetic edges passed through is multiplied by the preset first preset weight, and the number of feature edges passed through is multiplied by the preset second preset weight. The result is the final kinship distance between the closely related node and the target node. The smaller the value, the closer the kinship relationship.

[0044] For example, the final kinship distance = number of genetic edges passed × 0.6 plus number of feature edges passed × 0.4. The first preset weight 0.6 and the second preset weight 0.4 can be adjusted according to the breeding objectives of different crops. If the focus is on genetic association, the first weight is increased; if the focus is on morphological consistency verification, the second weight is increased, to adapt to different breeding management needs.

[0045] Finally, after organizing the set of closely related varieties and calculating the kinship distance, the upstream seed source node corresponding to the current breeding generation is retrieved along the generation edge. From the attribute fields associated with the corresponding node, the various morphological standards that the varieties to be managed under the current breeding generation should meet are extracted, including standard plant height range, standard leaf type description, standard ear type description and standard grain type description, and integrated to obtain the standard morphological feature description for subsequent comparison.

[0046] S40: Based on the phylogenetic distance between each closely related variety in the set of closely related varieties and the variety to be managed, select the variety with labeled purity training data as the migration source domain according to the closeness priority strategy; In this embodiment, after calculating the phylogenetic distance of all closely related varieties, they are sorted in ascending order of phylogenetic distance. Varieties with labeled purity training data that are closer in phylogenetic distance are selected as the transfer source domain. If there is no labeled purity training data for the closely related varieties at the top, the search continues until a set of labeled varieties that meets the data requirements for transfer learning is found, which is then used as the final transfer source domain. This ensures the correlation between the source domain and the target variety to be managed, and improves the effectiveness of transfer learning.

[0047] If no labeled purity training data is available for all closely related varieties, the number of traversal jumps is increased to re-extract the set of closely related varieties and select again. If no suitable varieties are found, a general purity detection model is used to complete the subsequent detection.

[0048] Specifically, step S40 in the method includes: Candidate varieties with labeled purity training data are selected from the set of closely related varieties. The labeled purity training data includes hybrid plant images from the seed breeding field of the candidate varieties and corresponding pixel-level purity labels. The candidate varieties are sorted in ascending order of their kinship distance; The candidate variety with the smallest phylogenetic distance was selected as the migration source region.

[0049] In this embodiment of the application, firstly, all candidate varieties with labeled purity training data are selected from the set of closely related varieties. Only candidate varieties with labeled data can serve as the transfer source domain to provide knowledge transfer support for the purity detection model of the target variety. If the number of candidate varieties is zero after screening, the number of traversal jumps is increased according to the aforementioned rules to obtain the set of closely related varieties again for screening.

[0050] Secondly, all the selected candidate varieties are sorted from smallest to largest according to the previously calculated kinship distance. The smaller the kinship distance, the higher the correlation with the target variety to be managed, and the better the transfer learning effect. Therefore, the candidate variety with the smallest kinship distance after sorting can be directly selected as the transfer source domain. In scenarios with large data requirements, a preset number of candidate varieties with high sorting can also be selected to form the transfer source domain to meet the data requirements of different model scales.

[0051] S50: Using the pre-trained purity identification model corresponding to the migration source domain as the base network, and using the standard morphological feature description for migration adaptation, a purity identification model for the variety to be managed is obtained.

[0052] In this embodiment, a pre-trained purity identification model that has completed basic training in the migration source domain is used as the base network. This model has learned the visual feature differences between hybrid plants and normal plants in the migration source domain and has basic plant purity identification capabilities.

[0053] Subsequently, the extracted standard morphological characteristics of the varieties to be managed are input into the pre-trained purity identification model. By learning the standard morphological characteristics of the varieties to be managed, the model's feature extraction preferences are adjusted to complete the model's transfer adaptation. Finally, a purity identification model adapted to the current varieties to be managed is obtained. Even with insufficient labeled data, the accuracy of purity identification can still be guaranteed, reducing the cost of manual labeling.

[0054] Specifically, step S50 in the method includes: Obtain the pre-trained purity assessment model corresponding to the source domain of the transfer, wherein the pre-trained purity assessment model is an instance segmentation network that has converged and trained on the purity training data of the source domain of the transfer. A small number of hybrid plant images were collected from the seed breeding field of the variety to be managed, and corresponding pixel-level purity labels were labeled to form a fine-tuning training dataset. The standard morphological feature description is transformed into a morphological feature constraint vector, and the deviation between the hybrid morphological features predicted by the model and the morphological feature constraint vector is calculated as the morphological feature constraint loss term. The original instance segmentation loss function is added to the morphological feature constraint loss term to obtain the total loss function; The parameters of the feature extraction layer of the pre-trained purity identification model are frozen. The parameters of the classification output layer and mask output layer of the pre-trained purity identification model are fine-tuned using the fine-tuning training dataset. The parameters of the classification output layer and mask output layer are updated by backpropagation based on the total loss function until the preset convergence condition is met. After fine-tuning and convergence, a purity identification model for the variety to be managed is obtained.

[0055] In this embodiment, firstly, the purity identification model that has been trained and converged in the migration source domain is extracted as a pre-trained purity identification model to be fine-tuned. This model is an instance segmentation network trained based on the purity training data labeled in the migration source domain, and it is already able to identify hybrids and normal plants of the corresponding source domain varieties.

[0056] Secondly, a small number of images of hybrid plants in the breeding fields of the varieties currently under management are collected to complete pixel-level purity annotation, thereby constructing a small-scale fine-tuning training dataset. This dataset can meet the fine-tuning requirements without large-scale annotation, significantly reducing annotation costs.

[0057] Then, the previously extracted standard morphological feature descriptions of the varieties to be managed are encoded into morphological feature constraint vectors. During the model fine-tuning process, the deviation between the hybrid morphological features predicted by the model and the morphological feature constraint vectors is calculated. This deviation is added to the total loss as an additional morphological feature constraint loss term, allowing the model to learn the standard morphological requirements of the varieties to be managed and constraining the prediction direction of the model.

[0058] Next, all the parameters of the feature extraction layer of the pre-trained purity identification model are frozen, and only the parameters of the classification output layer and the mask output layer are opened for updating. The parameters are fine-tuned using a small number of labeled fine-tuning training datasets. The parameters are updated through backpropagation using the total loss function until the model's accuracy on the validation set reaches the preset convergence condition, and then training is stopped. Finally, a purity identification model adapted to the current varieties to be managed is obtained.

[0059] For example, the purity identification model is trained by the following steps: Select the pre-trained instance segmentation network of the transfer source domain, freeze all parameters of its backbone feature extraction network, and retain only the parameters of the classification head and mask generation head for training. The standard morphological feature description of the variety to be managed is encoded by word embedding to obtain a morphological feature constraint vector with dimensions [1,256].

[0060] During each forward propagation, the feature map before the model's classification head is extracted, and the cosine similarity loss is calculated with the morphological feature constraint vector as the morphological feature constraint loss term. The total loss is the sum of the cross-entropy loss, mask loss, and cosine similarity loss of the original instance segmentation. The AdamW optimizer is used, with an initial learning rate of 1e-4 and a batch size of 8. Validation is performed every 5 training rounds. Training stops when the accuracy of the validation set improves by no more than 0.001 for 3 consecutive rounds. Finally, a target model that can be directly used for purity identification is obtained.

[0061] The cosine similarity loss is calculated using the formula L=1-cos(fe,vs), where fe is the average feature of the predicted region extracted by the model, vs is the morphological feature constraint vector obtained by encoding, and cos is the cosine similarity calculation function. The more similar the predicted region features are to the standard morphological features, the smaller the loss value and the stronger the constraint effect on the model.

[0062] Furthermore, digitalized planting management methods for breeding improved varieties also include: Acquire field image data of the seed breeding field of the variety to be managed at least two growth stages, namely the seedling stage, flowering stage, and maturity stage, to form a multi-growth-stage field image sequence; The multi-growth-stage field image sequence is input into the purity identification model for the variety to be managed. The purity identification model for the variety to be managed identifies hybrid plants in the field images of each growth stage and outputs the location coordinates of the hybrid plants in the seed production field and the morphological characteristics of the hybrid plants. A heatmap of hybrid plant distribution is generated based on the location coordinates of the hybrid plants. Based on the heatmap of hybrid plant distribution and the morphological characteristics of the hybrid plants, a strategy for removing hybrids and preserving purity is generated for the current breeding generation.

[0063] In this embodiment, firstly, when the varieties to be managed enter the key growth stages such as seedling stage, flowering stage, and maturity stage, aerial photography by drones is used to collect overall field images of the seed production field at the corresponding growth stages. Images from at least two different growth stages are selected and integrated to form a multi-growth-stage field image sequence, covering the morphological manifestations of hybrid plants at different growth stages, thus avoiding the problem of missed detection in single growth stage identification.

[0064] Secondly, each field image at multiple growth stages is sequentially input into the trained purity identification model. The model will identify the hybrid plants in the image one by one, output the coordinates of each hybrid plant in the field, and output the morphological characteristics description of the identified hybrid plants.

[0065] Furthermore, based on the location coordinates of all identified hybrid plants, a visual heat map of hybrid plant distribution is generated by mapping it to the global field coordinate system of the seed breeding field. This map can intuitively present the concentrated distribution areas of hybrid plants in the seed breeding field. Combined with the morphological characteristics of all hybrid plants, the main sources of hybrid plants are analyzed. Finally, based on the distribution density and type of hybrid plants, a strategy for removing hybrid plants and maintaining purity corresponding to this breeding generation is generated. This identifies the field areas that need to be focused on removing hybrid plants, as well as the methods for removing hybrid plants of different types, thus achieving precise management of hybrid plant removal.

[0066] In summary, compared with existing technologies, this application constructs a crop pedigree knowledge graph centered on genetic edges, generation edges, and characteristic edges. It uses graph traversal to retrieve closely related varieties and extract standard morphological feature descriptions. It selects migration source domains according to a close-relative priority strategy and transfers the pre-trained purity identification model of the source domain to the varieties to be managed. This enables the rapid construction of new variety purity identification models and the intelligent generation of multi-growth-stage impurity removal and purity preservation operation strategies. It effectively reduces the construction cost of new variety purity identification models and improves the intelligent level and cross-variety generalization ability of improved seed breeding purity management.

[0067] In summary, the embodiments of this application have at least the following technical effects: This application provides a digital planting management method for improved variety breeding. First, a crop pedigree knowledge graph is constructed using the breeding variety as a node, incorporating genetic relationships, propagation generations, and morphological similarity. This integrates existing pedigree and morphological information, providing comprehensive correlation data for selecting migration source regions for new varieties. Second, closely related varieties are identified by combining genetic relationships and morphological similarity. Compared to methods relying solely on genetic relationships to select migration source regions, this method can select source region varieties morphologically closer to the new variety to be managed, effectively improving the accuracy of the model after migration adaptation. Finally, standard morphological features corresponding to the current breeding generation are extracted and introduced as constraints into the model fine-tuning process. This allows for adaptation to the purity requirements of different breeding generations, and a purity identification model with satisfactory accuracy can be quickly obtained using only a small number of labeled samples.

[0068] Through the above technical solutions, this application solves the problem of scarce labeled samples of hybrid plants in newly bred varieties and the difficulty in quickly adapting purity identification models, and realizes digital and efficient management of improved seed breeding.

[0069] Example 2, as Figure 2 As shown, based on the same inventive concept as the digital planting management method for improved seed breeding provided in Embodiment 1, this application also provides a digital planting management system for improved seed breeding, including: The knowledge graph construction module 11 is used to construct a crop pedigree knowledge graph with breeding varieties as nodes, parent-offspring relationships as genetic edges, propagation associations from original varieties to improved varieties as generation edges, and morphological similarity between varieties as feature edges. The variety node query module 12 is used to obtain the variety identifier and current breeding generation of the variety to be managed, and to locate the variety node corresponding to the variety to be managed in the crop pedigree knowledge graph using the variety identifier as the query key. The associated feature extraction module 13 is used to perform a graph traversal with a limited number of jumps along the genetic edge and the feature edge, centered on the variety node, to extract a set of closely related varieties, and at the same time to search for the variety node corresponding to the current breeding generation along the generation edge and extract the associated standard morphological feature description. The migration source domain selection module 14 is used to select varieties with labeled purity training data as migration source domains according to the phylogenetic distance between each closely related variety in the set of closely related varieties and the variety to be managed, and according to the closeness priority strategy. The identification model acquisition module 15 is used to obtain a purity identification model for the variety to be managed by using the pre-trained purity identification model corresponding to the migration source domain as the base network and performing migration adaptation using the standard morphological feature description.

[0070] In one embodiment, the knowledge graph construction module 11 is specifically used for: Obtain the variety name, variety approval number, parent combination information, and generational information from the variety approval information; Each variety name is associated with a variety entity as a node, and each variety node is assigned variety attributes, which include at least variety type, morphological characteristics, approval year, breeding unit, and generation level. Based on the parental combination information, a genetic edge is established between two variety nodes that have a parent-offspring relationship, with the direction of the genetic edge pointing from the parent node to the offspring node. Based on the generation level information, generation edges are established between adjacent generation variety nodes that have a propagation association. The generation edges connect the original seed node with the original seed node, and the original seed node with the improved seed node. Calculate the morphological similarity between any two variety nodes, and establish feature edges between two variety nodes whose morphological similarity exceeds a preset similarity threshold; The variety nodes, genetic edges, generation edges, and feature edges are stored in a graph database to form a crop pedigree knowledge graph.

[0071] Furthermore, the morphological similarity between any two variety nodes is calculated, and feature edges are established between two variety nodes whose morphological similarity exceeds a preset similarity threshold, including: Obtain the morphological feature description field associated with the two variety nodes, input the morphological feature description field into the pre-trained feature encoding model, and output the initial morphological similarity by the feature encoding model; Obtain breeding log text, extract morphological comparison description statements related to the two variety nodes from the breeding log text, input the morphological comparison description statements into a pre-trained semantic scoring model, and output a morphological similarity correction score from the semantic scoring model. The initial morphological similarity score and the morphological similarity correction score are weighted and fused together, and the fused value is used as the morphological similarity between the two variety nodes. Obtain the current breeding generation of the variety to be managed, and determine the corresponding generation purity requirement based on the current breeding generation; A preset similarity threshold is dynamically determined based on the generation purity requirement, wherein the higher the generation purity requirement, the higher the preset similarity threshold. The morphological similarity is compared with the preset similarity threshold, and a feature edge is established between two variety nodes whose morphological similarity exceeds the preset similarity threshold.

[0072] Furthermore, the training process of the feature encoding model includes: Collect morphological feature description fields from variety approval information and construct a morphological description corpus; An initial feature encoding model is selected, which is a language representation model pre-trained by a masked language modeling task; The morphological description corpus is input into the initial feature encoding model, and the initial feature encoding model is pre-trained a second time using the masked language modeling task to obtain the pre-trained initial feature encoding model. Construct a similarity-labeled dataset, which includes multiple pairs of morphological feature description fields for varieties and their corresponding similarity labels; The pre-trained initial feature encoding model is then fine-tuned using the similarity-labeled dataset to obtain the trained feature encoding model.

[0073] Furthermore, the training process of the semantic scoring model includes: Collect morphological comparison descriptions from breeding log texts and construct a morphological comparison corpus; An initial semantic scoring model is selected, and a Sigmoid output layer is added on top of the initial semantic scoring model, wherein the initial semantic scoring model is a language representation model pre-trained by a masked language modeling task; The morphological comparison corpus is input into the initial semantic scoring model, and the initial semantic scoring model is pre-trained a second time using the masked language modeling task to obtain the pre-trained initial semantic scoring model. Construct a rating labeling dataset, which contains multiple morphological comparison description statements and corresponding similarity rating labels; The pre-trained initial semantic scoring model is fine-tuned using the aforementioned scoring annotation dataset to obtain the trained semantic scoring model.

[0074] In one embodiment, the associated feature extraction module 13 is specifically used for: Using the variety node obtained from the location as the starting node, initialize the set of closely related varieties as an empty set, and preset the maximum number of traversal jumps; Starting from the starting node, a breadth-first traversal is performed along the genetic edges and the feature edges. For each variety node visited during the traversal, the number of genetic edges and the number of feature edges traversed from the starting node to the variety node are recorded, and the variety node is added to the set of closely related varieties until the maximum number of traversal jumps is reached. For each closely related variety node in the set of closely related varieties, the number of genetic edges traversed to reach the closely related variety node is multiplied by a first preset weight, and then the number of feature edges traversed is multiplied by a second preset weight to obtain the kinship distance between the closely related variety node and the starting node. Search along the generation edge for the variety node corresponding to the current breeding generation, and extract the standard plant height range, standard leaf shape description, standard ear shape description and standard grain shape description that the variety to be managed should present under the current breeding generation from the variety attributes associated with the variety node, to form a standard morphological feature description.

[0075] In one embodiment of the application, the migration source domain selection module 14 is specifically used for: Candidate varieties with labeled purity training data are selected from the set of closely related varieties. The labeled purity training data includes hybrid plant images from the seed breeding field of the candidate varieties and corresponding pixel-level purity labels. The candidate varieties are sorted in ascending order of their kinship distance; The candidate variety with the smallest phylogenetic distance was selected as the migration source region.

[0076] In one embodiment of the application, the identification model acquisition module 15 is specifically used for: Obtain the pre-trained purity assessment model corresponding to the source domain of the transfer, wherein the pre-trained purity assessment model is an instance segmentation network that has converged and trained on the purity training data of the source domain of the transfer. A small number of hybrid plant images were collected from the seed breeding field of the variety to be managed, and corresponding pixel-level purity labels were labeled to form a fine-tuning training dataset. The standard morphological feature description is transformed into a morphological feature constraint vector, and the deviation between the hybrid morphological features predicted by the model and the morphological feature constraint vector is calculated as the morphological feature constraint loss term. The original instance segmentation loss function is added to the morphological feature constraint loss term to obtain the total loss function; The parameters of the feature extraction layer of the pre-trained purity identification model are frozen. The parameters of the classification output layer and mask output layer of the pre-trained purity identification model are fine-tuned using the fine-tuning training dataset. The parameters of the classification output layer and mask output layer are updated by backpropagation based on the total loss function until the preset convergence condition is met. After fine-tuning and convergence, a purity identification model for the variety to be managed is obtained.

[0077] Furthermore, digitalized planting management methods for improving seed breeding also include: Acquire field image data of the seed breeding field of the variety to be managed at least two growth stages, namely the seedling stage, flowering stage, and maturity stage, to form a multi-growth-stage field image sequence; The multi-growth-stage field image sequence is input into the purity identification model for the variety to be managed. The purity identification model for the variety to be managed identifies hybrid plants in the field images of each growth stage and outputs the location coordinates of the hybrid plants in the seed production field and the morphological characteristics of the hybrid plants. A heatmap of hybrid plant distribution is generated based on the location coordinates of the hybrid plants. Based on the heatmap of hybrid plant distribution and the morphological characteristics of the hybrid plants, a strategy for removing hybrids and preserving purity is generated for the current breeding generation.

Claims

1. A digital cultivation management method for hybrid breeding, characterized by, The method includes: Using breeding varieties as nodes, parent-offspring relationships as genetic edges, propagation associations from original varieties to improved varieties as generational edges, and morphological similarity among varieties as characteristic edges, a crop pedigree knowledge graph is constructed. Obtain the variety identifier and current breeding generation of the variety to be managed, and use the variety identifier as the query key to locate the variety node corresponding to the variety to be managed in the crop pedigree knowledge graph; Centered on the variety node, a graph traversal with a limited number of jumps is performed along the genetic edge and the feature edge to extract a set of closely related varieties. At the same time, the variety node corresponding to the current breeding generation is retrieved along the generation edge to extract the associated standard morphological feature description. Based on the phylogenetic distance between each closely related variety in the set of closely related varieties and the variety to be managed, varieties with labeled purity training data are selected as migration source domains according to the closeness priority strategy. Using the pre-trained purity identification model corresponding to the migration source domain as the base network, and employing the standard morphological feature description for migration adaptation, a purity identification model for the variety to be managed is obtained.

2. The digitalized cultivation management method for hybrid breeding according to claim 1, wherein, Using breeding varieties as nodes, parent-offspring relationships as genetic edges, propagation relationships from original strains to improved varieties as generational edges, and morphological similarity among varieties as characteristic edges, a crop pedigree knowledge graph is constructed, including: Obtain the variety name, variety approval number, parent combination information, and generational information from the variety approval information; Each variety name is associated with a variety entity as a node, and each variety node is assigned variety attributes, which include at least variety type, morphological characteristics, approval year, breeding unit, and generation level. Based on the parental combination information, a genetic edge is established between two variety nodes that have a parent-offspring relationship, with the direction of the genetic edge pointing from the parent node to the offspring node. Based on the generation level information, generation edges are established between adjacent generation variety nodes that have a propagation association. The generation edges connect the original seed node with the original seed node, and the original seed node with the improved seed node. Calculate the morphological similarity between any two variety nodes, and establish feature edges between two variety nodes whose morphological similarity exceeds a preset similarity threshold; The variety nodes, genetic edges, generation edges, and feature edges are stored in a graph database to form a crop pedigree knowledge graph.

3. The digitalized cultivation management method for hybrid breeding according to claim 2, wherein, Calculate the morphological similarity between any two variety nodes, and establish feature edges between two variety nodes whose morphological similarity exceeds a preset similarity threshold, including: Obtain the morphological feature description field associated with the two variety nodes, input the morphological feature description field into the pre-trained feature encoding model, and output the initial morphological similarity by the feature encoding model; Obtain breeding log text, extract morphological comparison description statements related to the two variety nodes from the breeding log text, input the morphological comparison description statements into a pre-trained semantic scoring model, and output a morphological similarity correction score from the semantic scoring model. The initial morphological similarity score and the morphological similarity correction score are weighted and fused together, and the fused value is used as the morphological similarity between the two variety nodes. Obtain the current breeding generation of the variety to be managed, and determine the corresponding generation purity requirement based on the current breeding generation; A preset similarity threshold is dynamically determined based on the generation purity requirement, wherein the higher the generation purity requirement, the higher the preset similarity threshold. The morphological similarity is compared with the preset similarity threshold, and a feature edge is established between two variety nodes whose morphological similarity exceeds the preset similarity threshold.

4. The digital farming management method for hybrid breeding according to claim 3, characterized by, The training process of a feature encoding model includes: Collect morphological feature description fields from variety approval information and construct a morphological description corpus; An initial feature encoding model is selected, which is a language representation model pre-trained by a masked language modeling task; The morphological description corpus is input into the initial feature encoding model, and the initial feature encoding model is pre-trained a second time using the masked language modeling task to obtain the pre-trained initial feature encoding model. Construct a similarity-labeled dataset, which includes multiple pairs of morphological feature description fields for varieties and their corresponding similarity labels; The pre-trained initial feature encoding model is then fine-tuned using the similarity-labeled dataset to obtain the trained feature encoding model.

5. The digitalized cultivation management method for hybrid breeding according to claim 3, wherein, The training process of the semantic scoring model includes: Collect morphological comparison descriptions from breeding log texts and construct a morphological comparison corpus; An initial semantic scoring model is selected, and a Sigmoid output layer is added on top of the initial semantic scoring model, wherein the initial semantic scoring model is a language representation model pre-trained by a masked language modeling task; The morphological comparison corpus is input into the initial semantic scoring model, and the initial semantic scoring model is pre-trained a second time using the masked language modeling task to obtain the pre-trained initial semantic scoring model. Construct a rating labeling dataset, which contains multiple morphological comparison description statements and corresponding similarity rating labels; The pre-trained initial semantic scoring model is fine-tuned using the aforementioned scoring annotation dataset to obtain the trained semantic scoring model.

6. The digitalized cultivation management method for hybrid breeding according to claim 1, wherein, Centered on the variety node, a graph traversal with a limited number of hops is performed along the genetic edge and the feature edge to extract a set of closely related varieties. Simultaneously, the variety node corresponding to the current breeding generation is retrieved along the generation edge, and the associated standard morphological feature description is extracted, including: Using the variety node obtained from the location as the starting node, initialize the set of closely related varieties as an empty set, and preset the maximum number of traversal jumps; Starting from the starting node, a breadth-first traversal is performed along the genetic edges and the feature edges. For each variety node visited during the traversal, the number of genetic edges and the number of feature edges traversed from the starting node to the variety node are recorded, and the variety node is added to the set of closely related varieties until the maximum number of traversal jumps is reached. For each closely related variety node in the set of closely related varieties, the number of genetic edges traversed to reach the closely related variety node is multiplied by a first preset weight, and then the number of feature edges traversed is multiplied by a second preset weight to obtain the kinship distance between the closely related variety node and the starting node. Search along the generation edge for the variety node corresponding to the current breeding generation, and extract the standard plant height range, standard leaf shape description, standard ear shape description and standard grain shape description that the variety to be managed should present under the current breeding generation from the variety attributes associated with the variety node, to form a standard morphological feature description.

7. The digitalized cultivation management method for hybrid breeding according to claim 1, wherein, Based on the phylogenetic distance between each closely related variety in the aforementioned set of closely related varieties and the variety to be managed, varieties with labeled purity training data are selected as migration source domains according to a close-relative priority strategy, including: Candidate varieties with labeled purity training data are selected from the set of closely related varieties. The labeled purity training data includes hybrid plant images from the seed breeding field of the candidate varieties and corresponding pixel-level purity labels. The candidate varieties are sorted in ascending order of their kinship distance; The candidate variety with the smallest phylogenetic distance was selected as the migration source region.

8. The digitalized cultivation management method for hybrid breeding according to claim 1, wherein, Using the pre-trained purity identification model corresponding to the migration source domain as the base network, and employing the standard morphological feature description for transfer adaptation, a purity identification model for the variety to be managed is obtained, including: Obtain the pre-trained purity assessment model corresponding to the source domain of the transfer, wherein the pre-trained purity assessment model is an instance segmentation network that has converged and trained on the purity training data of the source domain of the transfer. A small number of hybrid plant images were collected from the seed breeding field of the variety to be managed, and corresponding pixel-level purity labels were labeled to form a fine-tuning training dataset. The standard morphological feature description is transformed into a morphological feature constraint vector, and the deviation between the hybrid morphological features predicted by the model and the morphological feature constraint vector is calculated as the morphological feature constraint loss term. The original instance segmentation loss function is added to the morphological feature constraint loss term to obtain the total loss function; The parameters of the feature extraction layer of the pre-trained purity identification model are frozen. The parameters of the classification output layer and mask output layer of the pre-trained purity identification model are fine-tuned using the fine-tuning training dataset. The parameters of the classification output layer and mask output layer are updated by backpropagation based on the total loss function until the preset convergence condition is met. After fine-tuning and convergence, a purity identification model for the variety to be managed is obtained.

9. The digitalized cultivation management method for hybrid breeding according to claim 1, wherein, The method further includes: Acquire field image data of the seed breeding field of the variety to be managed at least two growth stages, namely the seedling stage, flowering stage, and maturity stage, to form a multi-growth-stage field image sequence; The multi-growth-stage field image sequence is input into the purity identification model for the variety to be managed. The purity identification model for the variety to be managed identifies hybrid plants in the field images of each growth stage and outputs the location coordinates of the hybrid plants in the seed production field and the morphological characteristics of the hybrid plants. A heatmap of hybrid plant distribution is generated based on the location coordinates of the hybrid plants. Based on the heatmap of hybrid plant distribution and the morphological characteristics of the hybrid plants, a strategy for removing hybrids and preserving purity is generated for the current breeding generation.

10. A digital plantation management system oriented to improved breeding, characterized by The method for implementing the digital planting management method for improved seed breeding as described in any one of claims 1-9 includes: The knowledge graph construction module is used to construct a crop pedigree knowledge graph with breeding varieties as nodes, parent-offspring relationships as genetic edges, propagation associations from original varieties to improved varieties as generational edges, and morphological similarity between varieties as feature edges. The variety node query module is used to obtain the variety identifier and current breeding generation of the variety to be managed, and to locate the variety node corresponding to the variety to be managed in the crop pedigree knowledge graph using the variety identifier as the query key. The associated feature extraction module is used to perform a graph traversal with a limited number of jumps along the genetic edge and the feature edge, centered on the variety node, to extract a set of closely related varieties, and at the same time, to search for the variety node corresponding to the current breeding generation along the generation edge and extract the associated standard morphological feature description. The migration source domain selection module is used to select varieties with labeled purity training data as migration source domains according to the phylogenetic distance between each closely related variety in the set of closely related varieties and the variety to be managed, and according to the closeness priority strategy. The identification model acquisition module is used to obtain a purity identification model for the variety to be managed by using the pre-trained purity identification model corresponding to the migration source domain as the base network and performing migration adaptation using the standard morphological feature description.