Sugarcane hybrid combination selection method and system

By constructing the path nodes and relationship network of sugarcane parents and combining it with a multi-dimensional scoring mechanism, the problem of traditional sugarcane hybrid selection relying on experience was solved, and the efficiency and success rate of sugarcane breeding were improved.

CN120787799APending Publication Date: 2025-10-17GUANGXI ZHUANG AUTONOMOUS REGION ACAD OF AGRI SCI
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510867714.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-26
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Traditional sugarcane hybridization and selection methods rely on manual experience, are inefficient, time-consuming and labor-intensive.

Method used

Based on various information of sugarcane parents, path nodes are created and path relationships are constructed. The optimal hybrid combination is selected through a multidimensional scoring mechanism, including quantitative network analysis of information such as parent type, gene type, and growth environment.

Benefits of technology

It significantly improved the efficiency and success rate of sugarcane hybrid selection, provided clear implementation steps and quantitative standards, and improved the scientific nature and efficiency of breeding practices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120787799A_ABST
    Figure CN120787799A_ABST
Patent Text Reader

Abstract

The invention provides a sugarcane hybrid combination selecting and matching method and system, and belongs to the technical field of sugarcane cultivation, in the method, path nodes and path relations among the path nodes are constructed based on multiple pieces of parent information of sugarcane parents, and each path relation corresponds to a path score. And for any group of sugarcane hybrid combination, obtaining a parent information set of the male parent and the female parent, and determining a candidate path based on the parent information set and the path relationship. And calculating a comprehensive score of the candidate path based on the path score and the reward score corresponding to the candidate path. And determining a sugarcane hybrid combination to be cultivated based on the comprehensive score of the candidate path. According to the method, through quantitative analysis and networked modeling, the sugarcane hybridization matching efficiency and feasibility can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of sugarcane breeding, and particularly relates to a sugarcane hybrid combination selection method and system. BACKGROUND

[0002] As an important sugar and energy crop in the world, sugarcane hybrid breeding is a key means to improve yield, quality and stress resistance. However, the traditional sugarcane hybrid selection method mainly relies on manual experience to screen parents, and the hybrid performance is verified through a large number of field tests, which is low in efficiency. SUMMARY

[0003] The technical problem to be solved by the present application is to provide a sugarcane hybrid combination selection method, which can improve the efficiency of sugarcane hybrid selection.

[0004] The technical scheme for solving the above technical problem is as follows:

[0005] On the one hand, the present application provides a sugarcane hybrid combination selection method, which creates a plurality of path nodes based on a plurality of parent information corresponding to sugarcane parents. The plurality of parent information includes parent type information, gene type information, trait type information, growth environment type information, metabolic pathway type information, and offspring performance information; one path node is used to represent one parent information. The path relationship between the path node corresponding to the first parent information and the path node corresponding to the second parent information is constructed. The first parent information and the second parent information are any two parent information in the plurality of parent information that have an associated relationship; one path relationship corresponds to one path score. For any group of sugarcane hybrid combinations, a parent information set consisting of parent information corresponding to a male sugarcane parent and parent information corresponding to a female sugarcane parent in any group of sugarcane hybrid combinations is obtained. Based on the parent information set and each path relationship, at least one candidate path is determined, wherein for any candidate path in the at least one candidate path, any candidate path includes at least one path relationship. Based on the path score corresponding to each path relationship in the at least one path relationship and the reward score corresponding to any candidate path, a comprehensive score corresponding to any candidate path is determined. The reward score corresponding to any candidate path is determined based on the environmental information of the region where any group of sugarcane hybrid combinations is to be planted, the planting performance information of the historical sugarcane hybrid combinations corresponding to the path relationship in the at least one path relationship, and the number of path nodes included in any candidate path. The highest comprehensive score of the candidate path corresponding to any group of sugarcane hybrid combinations is determined as the ultimate score corresponding to any group of sugarcane hybrid combinations. Based on the ultimate score corresponding to each sugarcane hybrid combination, the sugarcane hybrid combination to be bred is determined.

[0006] On the basis of the above technical scheme, the present application can also be improved as follows.

[0007] Further, the parent type information comprises identification data for characterizing the identity of the corresponding sugarcane parent; the gene type information comprises SNP gene data and / or InDel gene data of the corresponding sugarcane parent; the trait type information comprises at least one of high-sugar trait data, disease resistance trait data, and drought tolerance trait data of the corresponding sugarcane parent; the growth environment type information comprises at least one of temperature, humidity, daily irrigation amount, soil pH value, soil nitrogen content, soil phosphorus content, and soil potassium content of the growth environment where the corresponding sugarcane parent is located; the metabolic pathway type information comprises at least one of sugar metabolic pathway data, disease resistance related pathway data, stress resistance related pathway data, secondary metabolite synthesis pathway data, nutrient use efficiency pathway data, hormone regulation pathway data, photosynthesis related pathway data, and amino acid metabolic pathway data; the offspring representation information comprises the parent type information, the gene type information, the trait type information, the growth environment type information, and the metabolic pathway type information of the offspring sugarcane parent obtained by crossing the corresponding sugarcane parent.

[0008] Further, the association relationship between the parent type information and the gene type information is a carrying relationship; the association relationship between the gene type information and the trait type information is a regulation relationship; the association relationship between the trait type information and the growth environment type information is a dependence relationship; the association relationship between the gene type information and the metabolic pathway type information is a regulation relationship; the association relationship between the metabolic pathway type information and the trait type information is an influence relationship; the association relationship between the parent type information and the offspring representation information is a crossing relationship; the association relationship between the offspring representation information and the trait type information is a performance relationship; and the crossing relationship comprises obtaining the offspring sugarcane parent by crossing the corresponding sugarcane parent.

[0009] Further, based on the plurality of sugarcane parents, a plurality of sugarcane crossing pairs are obtained. For any sugarcane crossing pair in the plurality of sugarcane crossing pairs, the sugarcane crossing pair comprises a first sugarcane parent and a second sugarcane parent, and the first sugarcane parent and the second sugarcane parent are any two sugarcane parents included in the plurality of sugarcane parents. Based on the gene type information corresponding to the first sugarcane parent and the gene type information corresponding to the second sugarcane parent, a genetic distance between the first sugarcane parent and the second sugarcane parent is determined. Based on the genetic distance between the first sugarcane parent and the second sugarcane parent being less than a genetic distance threshold, the sugarcane crossing pair is determined as a candidate sugarcane crossing combination. Based on each candidate sugarcane crossing combination determined from the plurality of sugarcane crossing pairs, at least one group of sugarcane crossing combinations is determined, and the at least one group of sugarcane crossing combinations comprises any group of sugarcane crossing combinations.

[0010] Further, the environmental information comprises precipitation information, temperature information, and soil information. An environmental stress factor is determined. The environmental stress factor comprises a drought environmental stress factor, a high-temperature environmental stress factor, and a soil environmental stress factor. A standardized value corresponding to each environmental stress factor is obtained. The standardized value corresponding to the drought environmental stress factor is determined based on the precipitation information of the region where any one group of sugarcane hybrid combinations is to be planted. The standardized value corresponding to the high-temperature environmental stress factor is determined based on the temperature information of the region where any one group of sugarcane hybrid combinations is to be planted. The standardized value corresponding to the soil environmental stress factor is determined based on the soil information of the region where any one group of sugarcane hybrid combinations is to be planted. Based on the standardized value corresponding to each environmental stress factor and the weight, an environmental stress index of the region where any one group of sugarcane hybrid combinations is to be planted is determined. Historical hybrid data is obtained. The historical hybrid data comprises a genetic distance distribution of a plurality of historical sugarcane hybrid combinations in a non-stress environment. The non-stress environment is an environment with a preset environmental stress index. Based on the genetic distance distribution of the plurality of historical sugarcane hybrid combinations, a standard genetic distance threshold is determined. The standard genetic distance threshold is determined based on the genetic distance with the highest frequency in the genetic distance distribution of the plurality of historical sugarcane hybrid combinations. Based on the environmental stress index of the region where any one group of sugarcane hybrid combinations is to be planted, the environmental information of the region where any one group of sugarcane hybrid combinations is to be planted, and the historical hybrid data, an environmental adjustment coefficient of the region where any one group of sugarcane hybrid combinations is to be planted is determined. Based on the standard genetic distance threshold, the environmental adjustment coefficient of the region where any one group of sugarcane hybrid combinations is to be planted, and the environmental stress index of the region where any one group of sugarcane hybrid combinations is to be planted, a genetic distance threshold is determined.

[0011] Further, the Euclidean distance between the growth environment type information corresponding to any candidate path and the environmental information of the region where any one group of sugarcane hybrid combinations is to be planted is obtained. Based on the Euclidean distance between the growth environment type information corresponding to any candidate path and the environmental information of the region where any one group of sugarcane hybrid combinations is to be planted, an environmental reward score corresponding to any candidate path is determined. Based on the planting performance information of the historical sugarcane hybrid combinations corresponding to each path relationship in the at least one path relationship, a historical performance reward score corresponding to any candidate path is determined. The planting performance information comprises sugar content information, yield information, disease resistance information, and drought survival rate information of the offspring sugarcane corresponding to the historical sugarcane hybrid combinations. Based on the number of path nodes included in any candidate path, a path reward score corresponding to any candidate path is determined. Based on the environmental reward score, the historical performance reward score, and the path reward score corresponding to any candidate path, a reward score corresponding to any candidate path is determined.

[0012] Further, based on the parent information set, at least one target path node is determined in the plurality of path nodes. Wherein, for any target path node in the at least one target path node, the parent information set comprises parent information corresponding to any target path node. Based on the at least one target path node, at least one candidate path is determined in each path relationship. Wherein, any candidate path in the at least one candidate path comprises at least one path relationship. For any path relationship in the at least one path relationship, there is a path node in the at least one target path node for constructing any path relationship.

[0013] Further, the parent information of the hybrid offspring of the sugarcane hybrid combination to be bred is obtained; and based on the parent information of the hybrid offspring, the path scores respectively corresponding to each path relationship in the at least one path relationship are adjusted.

[0014] In another aspect, the present application provides a sugarcane hybrid combination matching system, comprising:

[0015] The path node generation module creates a plurality of path nodes based on a plurality of parent information corresponding to sugarcane parents. Wherein, the plurality of parent information comprises parent type information, gene type information, trait type information, growth environment type information, metabolic pathway type information, and offspring performance information; one path node is used to represent one parent information.

[0016] The path relationship construction module constructs a path relationship between the path node corresponding to the first parent information and the path node corresponding to the second parent information. Wherein, the first parent information and the second parent information are any two parent information in the plurality of parent information that have an association relationship; one path relationship corresponds to one path score.

[0017] The parent information fusion module obtains, for any set of sugarcane hybrid combinations, a parent information set composed of parent information corresponding to a male sugarcane parent and parent information corresponding to a female sugarcane parent in any set of sugarcane hybrid combinations.

[0018] The candidate path selection module determines at least one candidate path based on the parent information set and each path relationship. Wherein, for any candidate path in the at least one candidate path, any candidate path comprises at least one path relationship.

[0019] The comprehensive score calculation module determines a comprehensive score corresponding to any candidate path based on the path score respectively corresponding to each path relationship in the at least one path relationship and the reward score corresponding to any candidate path. Wherein, the reward score corresponding to any candidate path is determined based on the environmental information of the region where any set of sugarcane hybrid combinations is to be planted, the planting performance information of the historical sugarcane hybrid combinations corresponding to each path relationship in the at least one path relationship, and the number of path nodes included in any candidate path.

[0020] The ultimate score calculation module determines the highest comprehensive score of the candidate path corresponding to any group of sugarcane hybrid combinations as the ultimate score corresponding to any group of sugarcane hybrid combinations.

[0021] The selection module determines the sugarcane hybrid combination to be bred based on the ultimate score corresponding to each sugarcane hybrid combination.

[0022] On the basis of the above technical solutions, the application can also be improved as follows.

[0023] Further, the parent type information includes identification data for characterizing the identity of the corresponding sugarcane parent; the gene type information includes SNP gene data and / or InDel gene data of the corresponding sugarcane parent; the trait type information includes at least one of high-sugar trait data, disease resistance trait data, and drought tolerance trait data of the corresponding sugarcane parent; the growth environment type information includes at least one of temperature, humidity, daily irrigation amount, soil pH value, soil nitrogen content, soil phosphorus content, and soil potassium content of the growth environment where the corresponding sugarcane parent is located; the metabolic pathway type information includes at least one of sugar metabolic pathway data, disease resistance-related pathway data, stress resistance-related pathway data, secondary metabolite synthesis pathway data, nutrient use efficiency pathway data, hormone regulation pathway data, photosynthesis-related pathway data, and amino acid metabolic pathway data; the offspring representation information includes the parent type information, the gene type information, the trait type information, the growth environment type information, and the metabolic pathway type information of the offspring sugarcane parent obtained by crossing the corresponding sugarcane parent.

[0024] The association relationship between the parent type information and the gene type information is a carrying relationship; the association relationship between the gene type information and the trait type information is a regulation relationship; the association relationship between the trait type information and the growth environment type information is a dependence relationship; the association relationship between the gene type information and the metabolic pathway type information is a regulation relationship; the association relationship between the metabolic pathway type information and the trait type information is an influence relationship; the association relationship between the parent type information and the offspring representation information is a crossing relationship; the association relationship between the offspring representation information and the trait type information is a performance relationship; the crossing relationship includes obtaining the offspring sugarcane parent by crossing the corresponding sugarcane parent.

[0025] The application has the beneficial effects that the application converts the complex parent information of the sugarcane parent into a network structure, and solves the problems of low efficiency and dependence on experience in traditional sugarcane hybrid selection and matching by combining with a multi-dimensional scoring mechanism. Moreover, the application has clear implementation steps and quantitative standards, and can be widely applied to sugarcane breeding practice, and significantly improves the selection and matching efficiency and success rate.

[0026] In a third aspect, the present application provides an electronic device, comprising: a memory, one or more processors; the memory and the processor are coupled; wherein the memory has computer program codes stored therein, the computer program codes comprising computer instructions, when the computer instructions are executed by the processor, the electronic device executes the method of any one of the first aspect.

[0027] In a fourth aspect, a computer readable storage medium is provided, comprising computer instructions, when the computer instructions are run on an electronic device, the electronic device executes the method of any one of the first aspect.

[0028] In a fifth aspect, a computer program product is provided, when the computer program product is run on a computer, the computer executes the method of any one of the first aspect.

[0029] It can be understood that the beneficial effects achieved by the system of the second aspect, the electronic device of the third aspect, the computer readable storage medium of the fourth aspect, and the computer program product of the fifth aspect can refer to the beneficial effects of the first aspect and any possible design of the first aspect, which will not be repeated here. BRIEF DESCRIPTION OF DRAWINGS

[0030] Figure 1 A flowchart of a sugarcane hybrid combination selection method provided by the present application is shown in the figure.

[0031] Figure 2 A structural diagram of a sugarcane hybrid combination selection system provided by the present application is shown in the figure. DETAILED DESCRIPTION

[0032] The technical solutions in the embodiments of the present application will be described below with reference to the drawings in the embodiments of the present application. In the description of the present application, unless otherwise specified, " / " represents an "or" relationship between the objects before and after the " / " symbol, for example, A / B can represent A or B; in the present application, "and / or" is only a description of the relationship between the associated objects, which means that there can be three relationships, for example, A and / or B, which can represent the following three cases: A exists alone, A and B exist together, and B exists alone, where A and B can be singular or plural. In the description of the present application, unless otherwise specified, "multiple" means two or more than two. "At least one of the following" or similar expressions means any combination of the items, including any combination of single item or multiple items. For example, at least one of a, b, or c can represent a, b, c, a-b, a-c, b-c, or a-b-c, where a, b, and c can be single or multiple. In addition, in order to clearly describe the technical solutions of the embodiments of the present application, in the embodiments of the present application, "first", "second", and the like are used to distinguish the same items or similar items with basically the same function and effect. Those skilled in the art can understand that "first", "second", and the like do not limit the quantity and execution order, and "first", "second", and the like do not necessarily mean different. At the same time, in the embodiments of the present application, "exemplary" or "for example" means to serve as an example, illustration or description. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as more preferred or more advantageous than other embodiments or design schemes.

[0033] The traditional sugarcane hybrid selection method relies on manual experience to screen parents, and the hybrid performance of the selected parents is verified through a large number of field tests, which is time-consuming and labor-intensive. In view of this problem, the present application provides a sugarcane hybrid combination selection method and system, which can improve the efficiency and implementability of sugarcane hybrid selection.

[0034] Referring to Figure 1 A flowchart of a sugarcane hybrid combination selection method provided by the present application includes the following steps S101-S107:

[0035] S101: Based on a plurality of parent information corresponding to sugarcane parents, a plurality of path nodes are created.

[0036] The plurality of parent information includes parent type information, gene type information, trait type information, growth environment type information, metabolic pathway type information, and offspring performance information; and one path node is used to represent one parent information.

[0037] In some embodiments, the parent type information includes identification data for characterizing the identity of the corresponding sugarcane parent; the genetic type information includes SNP genetic data and / or InDel genetic data of the corresponding sugarcane parent; the trait type information includes at least one of high-sugar trait data, disease resistance trait data, and drought tolerance trait data of the corresponding sugarcane parent; the growth environment type information includes at least one of temperature, humidity, daily irrigation amount, soil pH, soil nitrogen content, soil phosphorus content, and soil potassium content of the growth environment in which the corresponding sugarcane parent is located; the metabolic pathway type information includes at least one of sugar metabolic pathway data, disease resistance related pathway data, stress resistance related pathway data, secondary metabolite synthesis pathway data, nutrient use efficiency pathway data, hormone regulation pathway data, photosynthesis related pathway data, and amino acid metabolic pathway data; the offspring representation information includes the parent type information, the genetic type information, the trait type information, the growth environment type information, and the metabolic pathway type information corresponding to the offspring sugarcane parent obtained by crossing the corresponding sugarcane parent.

[0038] It should be noted that the SNP genetic data is the genetic data of the sugarcane at the SNP site, and the SNP site is a single nucleotide polymorphism site that is significantly related to the traits of the sugarcane. For example, a certain SNP site of the sugarcane can be related to a key enzyme gene in the sugar metabolic pathway.

[0039] It should be noted that the InDel genetic data is the genetic data associated with the InDel marker of the sugarcane, and the InDel marker is an insertion / deletion variation related to the traits of the sugarcane. For example, a certain InDel marker of the sugarcane can be located in the promoter region of a disease resistance gene, affecting the expression level of the disease resistance gene.

[0040] In some embodiments, a sugarcane sample library can be constructed, and the sugarcane sample library includes a plurality of representative sugarcane parents. The representative sugarcane parents can be sugarcane parents with high-sugar traits, sugarcane parents with high disease resistance traits, sugarcane parents with drought tolerance traits, and the like. Parent information of each sugarcane parent can be collected, and the collected parent information can be preprocessed. The preprocessing can include noise removal, missing value filling, and standardization processing. Then, the preprocessed parent information is fused to obtain a multi-dimensional parent information matrix as shown in Table 1.

[0041] Table 1

[0042]

[0043] Based on the multi-dimensional parent information matrix, a plurality of path nodes can be created, each path node corresponding to the information corresponding to a data bit in the multi-dimensional parent information matrix. For example, the path nodes can include a path node for indicating that the parent type information is P001, a path node for indicating that the parent type information is P002, a path node for indicating that the gene type information is an InDel marker, a path node for indicating that the gene type information is no InDel marker, a path node for indicating that the SNP point in the gene type information is position 0, a path node for indicating that the SNP point in the gene type information is position 1, and the like.

[0044] S102: Construct a path relationship between the path node corresponding to the first parent information and the path node corresponding to the second parent information.

[0045] Among them, the first parent information and the second parent information are any two parent information with an associated relationship in the plurality of parent information; one path relationship corresponds to one path score.

[0046] In some embodiments, the association relationship between the parent type information and the gene type information is a carrying relationship; the association relationship between the gene type information and the trait type information is a regulation relationship; the association relationship between the trait type information and the growth environment type information is a dependence relationship; the association relationship between the gene type information and the metabolic pathway type information is a regulation relationship; the association relationship between the metabolic pathway type information and the trait type information is an influence relationship; the association relationship between the parent type information and the offspring representation information is a hybridization relationship; the association relationship between the offspring representation information and the trait type information is a performance relationship; the hybridization relationship includes a parent sugarcane parent obtained by hybridizing the corresponding sugarcane parent.

[0047] For example, the path relationship can include: parent type information-carries-gene type information, gene type information-regulates-trait type information, trait type information-depends on-growth environment type information, gene type information-regulates-metabolic pathway type information, metabolic pathway type information-affects-trait type information, parent type information-hybridizes-offspring representation information, offspring representation information-performs-trait type information.

[0048] The knowledge graph can be constructed by taking each path node as an entity and each path relationship as an edge, and the subsequent step S104 can be implemented based on the knowledge graph.

[0049] S103: For any set of sugarcane hybrid combinations, obtain a parent information set composed of the parent information corresponding to the male sugarcane parent and the parent information corresponding to the female sugarcane parent in any set of sugarcane hybrid combinations.

[0050] In some embodiments, based on the plurality of sugarcane parents, a plurality of sugarcane cross pairs is obtained. For any sugarcane cross pair in the plurality of sugarcane cross pairs, the any sugarcane cross pair includes a first sugarcane parent and a second sugarcane parent, the first sugarcane parent and the second sugarcane parent being any two sugarcane parents included in the plurality of sugarcane parents. Based on the genetic type information corresponding to the first sugarcane parent and the genetic type information corresponding to the second sugarcane parent, a genetic distance between the first sugarcane parent and the second sugarcane parent is determined. Based on the genetic distance between the first sugarcane parent and the second sugarcane parent being less than a genetic distance threshold, the any sugarcane cross pair is determined as a candidate sugarcane cross combination. Based on the candidate sugarcane cross combinations determined from the plurality of sugarcane cross pairs, at least one set of sugarcane cross combinations is determined, the at least one set of sugarcane cross combinations including any set of sugarcane cross combinations.

[0051] It should be noted that the genetic distance between the first sugarcane parent and the second sugarcane parent can be calculated based on the improved Nei's algorithm, and the embodiments of the present application do not limit the algorithm for calculating the genetic distance between the first sugarcane parent and the second sugarcane parent.

[0052] In some embodiments, the environmental information includes precipitation information, temperature information, and soil information. An environmental stress factor is determined. The environmental stress factor includes a drought environmental stress factor, a high-temperature environmental stress factor, and a soil environmental stress factor. A standardized value corresponding to each environmental stress factor is obtained. The standardized value corresponding to the drought environmental stress factor is determined based on the precipitation information of the region where any set of sugarcane cross combinations is to be planted, the standardized value corresponding to the high-temperature environmental stress factor is determined based on the temperature information of the region where any set of sugarcane cross combinations is to be planted, and the standardized value corresponding to the soil environmental stress factor is determined based on the soil information of the region where any set of sugarcane cross combinations is to be planted. Based on the standardized value corresponding to each environmental stress factor and a weight, an environmental stress index of the region where any set of sugarcane cross combinations is to be planted is determined. Historical cross data is obtained. The historical cross data includes a genetic distance distribution of a plurality of historical sugarcane cross combinations in a stress-free environment, the stress-free environment being an environment with a preset environmental stress index. Based on the genetic distance distribution of the plurality of historical sugarcane cross combinations, a standard genetic distance threshold is determined. The standard genetic distance threshold is determined based on the genetic distance with the highest frequency in the genetic distance distribution of the plurality of historical sugarcane cross combinations. Based on the environmental stress index of the region where any set of sugarcane cross combinations is to be planted, the environmental information of the region where any set of sugarcane cross combinations is to be planted, and the historical cross data, an environmental adjustment coefficient of the region where any set of sugarcane cross combinations is to be planted is determined. Based on the standard genetic distance threshold, the environmental adjustment coefficient of the region where any set of sugarcane cross combinations is to be planted, and the environmental stress index of the region where any set of sugarcane cross combinations is to be planted, a genetic distance threshold is determined.

[0053] Exemplarily, the genetic distance threshold can be determined based on the following formula:

[0054] D1=D2*(1+αE);

[0055] wherein, D1 represents the genetic distance threshold, D2 represents the standard genetic distance threshold, a represents the environmental adjustment coefficient, and E represents the environmental stress index.

[0056] The environmental adjustment coefficient can be determined based on the following formula:

[0057]

[0058] wherein, D3 represents the optimal genetic distance, which refers to the genetic difference degree between parents that can maximize the performance of target traits (such as yield, sugar content, stress resistance) of hybrid offspring under preset environmental conditions.

[0059] The environmental stress index can be determined based on the following formula:

[0060]

[0061] wherein, S i represents the standardized stress value of the i-th factor (such as drought ). W i represents the weight corresponding to S i , which can be determined based on expert experience or model training.

[0062] S104: determining at least one candidate path based on the parent information set and each path relationship.

[0063] wherein, for any candidate path in the at least one candidate path, any candidate path includes at least one path relationship.

[0064] In some embodiments, at least one target path node is determined in a plurality of path nodes based on the parent information set. Wherein, for any target path node in the at least one target path node, the parent information corresponding to any target path node is included in the parent information set. At least one candidate path is determined in each path relationship based on the at least one target path node. Wherein, any candidate path in the at least one candidate path includes at least one path relationship. For any path relationship in the at least one path relationship, there is a path node in the at least one target path node for constructing any path relationship.

[0065] Exemplarily, the parent information set can include {{P001, SNP-0, InDel-0, high sugar……}{P002, SNP-1, InDel-0, disease resistance……}}; based on the parent information set, the candidate path can be determined as:

[0066] Candidate Path 1: P001-carrying-SNP-0-increase-sugar-suppressed-drought environment;

[0067] Candidate Path 2: P001-carrying-SNP-0-increase-sugar-suppressed-drought environment;

[0068] Candidate Path 3: P001-hybrid-P011-carrying-SNP-12-increase-drought tolerance-suppressed-low temperature environment;

[0069] Candidate Path 4: P002-carrying-SNP-1-regulate-sugar metabolism pathway-increase-sugar-suppressed-drought environment;

[0070] Candidate Path 5: ……

[0071] S105: Based on the path score of each path relationship in the at least one path relationship, and the reward score corresponding to any candidate path, determine the comprehensive score corresponding to any candidate path.

[0072] Wherein, the reward score corresponding to any candidate path is determined based on the environmental information of the region where any group of sugarcane hybrid combination is to be planted, the planting performance information of the historical sugarcane hybrid combination corresponding to each path relationship in the at least one path relationship, and the number of path nodes included in any candidate path.

[0073] In some embodiments, the average value of the performance of the offspring of the similar past hybrid combinations (gene-trait-environment relationship similar) retrieved from the foregoing knowledge graph as the prediction benchmark based on the current candidate path. For example, if the current candidate path contains gene A→ sugar trait→ drought environment, and there are 5 similar hybrid combinations in history, the average sugar of their offspring is 16%±1%, then the prediction value is set to 16%. Then, based on the prediction benchmark, the prediction error distribution of each path relationship in the similar historical candidate path (such as 90% of the similar candidate path error <8%) can be calculated, and the historical average error is taken as the estimated error of the corresponding path relationship. Based on the estimated error of the corresponding path relationship obtained, the path score of the corresponding path relationship can be determined. For example, the estimated error of the corresponding path relationship is less than 20%, the path score of the corresponding path relationship is 5; the estimated error of the corresponding path relationship is greater than 70%, the path score of the corresponding path relationship is 1, etc.

[0074] In some embodiments, the Euclidean distance between the growth environment type information corresponding to any candidate path and the environmental information of the region where any group of sugarcane hybrid combination is to be planted is obtained. Based on the Euclidean distance between the growth environment type information corresponding to any candidate path and the environmental information of the region where any group of sugarcane hybrid combination is to be planted, the environmental reward score corresponding to any candidate path is determined.

[0075] Based on the planting performance information of the historical sugarcane hybrid combination corresponding to each path relationship in the at least one path relationship, a historical performance reward score corresponding to any candidate path is determined. The planting performance information includes sugar content information, yield information, disease resistance information, and drought tolerance survival rate information of the offspring sugarcane corresponding to the historical sugarcane hybrid combination.

[0076] Based on the number of path nodes included in any candidate path, a path reward score corresponding to any candidate path is determined.

[0077] Based on the environmental reward score, the historical performance reward score, and the path reward score corresponding to any candidate path, a reward score corresponding to any candidate path is determined.

[0078] As can be seen, the higher the final score of the candidate path corresponding to the sugarcane hybrid combination, the better the offspring performance (e.g., sugar content, disease resistance, drought tolerance, etc.) of the sugarcane hybrid combination, and the more suitable the sugarcane hybrid combination is for the environment of the region where the sugarcane hybrid combination is to be planted. Therefore, based on the final scores of the candidate paths corresponding to the sugarcane hybrid combinations, high-quality (high success rate of cultivation, good offspring performance) sugarcane hybrid combinations can be screened out.

[0079] In some embodiments, the Euclidean distance between the growth environment type information corresponding to any candidate path and the environmental information of the region where any group of sugarcane hybrid combinations is to be planted can be converted into a corresponding similarity score (0-1) based on a Gaussian kernel function. Based on the similarity score, the environmental reward score corresponding to the any candidate path can be determined.

[0080] In some embodiments, since excessive path length or inclusion of redundant relationships can lead to accumulation of prediction errors, candidate paths that are concise and have high confidence can be preferentially selected to reduce invalid calculations.

[0081] Specifically, the path reward score can be determined based on the path length L and the path node diversity Q, where the node diversity Q can be determined based on the following formula:

[0082]

[0083] For example, in the path "Gene X -> Trait Y -> Gene Z", the unique path node types are {Gene, Trait}, the total number of nodes is 3, and therefore Q = 1-2 / 3 ≈ 0.33.

[0084] Based on the node diversity Q, the path reward score P can be calculated as:

[0085]

[0086] where L maxrepresents a preset maximum path length; U1 and U2 are both adjustment weights (for example, U1=0.7, U2=0.3), which can be optimized by training data.

[0087] For example, the path length L=4, the preset maximum path length L max =5, and the node diversity Q=0.2, then the path reward score P=0.7x54+0.3x0.2=0.62.

[0088] S106: Determine the highest comprehensive score of the candidate path corresponding to any sugarcane hybrid combination as the ultimate score corresponding to the any sugarcane hybrid combination.

[0089] Specifically, any sugarcane hybrid combination can correspond to multiple candidate paths, and different candidate paths correspond to different comprehensive scores. The highest comprehensive score can be determined as the ultimate score of the any sugarcane hybrid combination.

[0090] S107: Determine the sugarcane hybrid combination to be cultivated based on the ultimate score corresponding to each sugarcane hybrid combination.

[0091] For example, the ultimate score corresponding to each sugarcane hybrid combination can be displayed to the user, so that the user can evaluate which sugarcane hybrid combination to prioritize based on the ultimate score corresponding to each sugarcane hybrid combination.

[0092] In some embodiments, a threshold value of the ultimate score can be set, and the sugarcane hybrid combination to be cultivated can be determined based on the sugarcane hybrid combination whose ultimate score is greater than the threshold value of the ultimate score.

[0093] In some embodiments, the parent information of the hybrid offspring of the sugarcane hybrid combination to be cultivated is obtained, and the path score corresponding to each path relationship in at least one path relationship is adjusted based on the parent information of the hybrid offspring.

[0094] In some embodiments, the path nodes and path relationships described above can be updated based on the parent information of the hybrid offspring of the sugarcane hybrid combination to be cultivated. When new data conflicts with old knowledge, high-weight knowledge can be retained according to confidence.

[0095] In some embodiments, referring to Figure 2 The application also provides a sugarcane hybrid combination matching system, which comprises:

[0096] A path node generation module creates multiple path nodes based on multiple parent information corresponding to sugarcane parents. The multiple parent information includes parent type information, gene type information, trait type information, growth environment type information, metabolic pathway type information, and offspring representation information. One path node is used to represent one parent information.

[0097] A path relationship construction module constructs a path relationship between a path node corresponding to the first parent information and a path node corresponding to the second parent information. The first parent information and the second parent information are any two parent information in the plurality of parent information that have an association relationship. One path relationship corresponds to one path score.

[0098] A parent information fusion module obtains a parent information set composed of parent information corresponding to a father sugarcane parent and parent information corresponding to a mother sugarcane parent in any one of the sugarcane hybrid combinations.

[0099] A candidate path selection module determines at least one candidate path based on the parent information set and the path relationships. For any candidate path in the at least one candidate path, the any candidate path includes at least one path relationship.

[0100] A comprehensive score calculation module determines a comprehensive score corresponding to any candidate path based on the path scores corresponding to the path relationships in the at least one path relationship and a reward score corresponding to the any candidate path. The reward score corresponding to the any candidate path is determined based on environmental information of a region where the any one of the sugarcane hybrid combinations is to be planted, planting performance information of historical sugarcane hybrid combinations corresponding to the path relationships in the at least one path relationship, and a number of path nodes included in the any candidate path.

[0101] A final score calculation module determines a highest comprehensive score of the candidate paths corresponding to any one of the sugarcane hybrid combinations as a final score corresponding to the any one of the sugarcane hybrid combinations.

[0102] A selection module determines the sugarcane hybrid combination to be cultivated based on the final scores corresponding to the sugarcane hybrid combinations.

[0103] It can be seen that the core innovation points of the present application include:

[0104] Path node and path relationship construction: based on various information of sugarcane parents (such as gene type, trait type, growth environment, etc.), path nodes are created, and path relationships between nodes are constructed. Each path relationship corresponds to a path score. This method converts complex parent information into quantifiable network structure, facilitating systematic analysis and optimization.

[0105] Comprehensive score of candidate path: the comprehensive score of the candidate path is determined by the path score, the reward score (based on environmental information, historical planting performance, etc.), and the optimal hybrid combination is finally selected to avoid blind hybridization. This multi-dimensional scoring mechanism combines genetics, environmental adaptability, and historical data, significantly improving the scientificity and efficiency of selection and pairing.

[0106] Dynamic adjustment and optimization: dynamically adjust the path score according to the performance of the hybrid offspring, which embodies the adaptability and iterative optimization capability of the method.

[0107] In some schemes, multiple embodiments of the present application can be combined, and the combined scheme can be implemented. Optionally, some operations in the flow of each method embodiment are optionally combined, and / or the order of some operations is optionally changed. Moreover, the execution order between the steps of each flow is only exemplary and does not constitute a limitation on the execution order between the steps, and other execution orders between the steps can also be used. It is not intended to indicate that the execution order is the only execution order in which these operations can be performed. Those skilled in the art can think of various ways to reorder the operations described herein. In addition, it should be pointed out that the process details involved in a certain embodiment herein are also applicable in a similar manner to other embodiments, or different embodiments can be combined for use.

[0108] In addition, some steps in the method embodiment can be equivalently replaced by other possible steps. Alternatively, some steps in the method embodiment can be optional and can be deleted in some use scenarios. Alternatively, other possible steps can be added to the method embodiment. Moreover, each method embodiment can be implemented individually or in combination.

[0109] From the above description of the embodiments, those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above functional modules is taken as an example for illustration, and in actual application, the above functions can be completed by different functional modules according to needs, that is, the internal structure of the system is divided into different functional modules to complete all or part of the functions described above.

[0110] In several embodiments provided in the present application, it should be understood that the disclosed system and method can be implemented in other ways. For example, the system embodiments described above are only illustrative, for example, the division of the modules or units is only a logical function division, and in actual implementation, there can be another division way, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the coupling or direct coupling or communication connection between the displayed or discussed each other can be through some interface, indirect coupling or communication connection between the system or unit, which can be electrical, mechanical or other forms.

[0111] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of software functional unit.

[0112] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a readable storage medium. Based on such understanding, the technical solutions of the embodiments of the present application can be embodied in the form of a software product in essence or in the part that contributes to the present application, or the whole or part of the technical solutions can be embodied in the form of a software product stored in a storage medium, including a plurality of instructions for causing an apparatus (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the method described in the embodiments of the present application. The foregoing storage medium includes a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and various media that can store program codes.

[0113] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto, and any change or replacement within the technical scope disclosed in the present application should be covered in the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A method for selecting sugarcane hybrid combinations, characterized in that: include: Based on multiple parent information corresponding to sugarcane parents, multiple path nodes are created; the multiple parent information includes parent type information, gene type information, trait type information, growth environment type information, metabolic pathway type information, and offspring performance information; one path node is used to represent one parent information; Constructing a path relationship between a path node corresponding to the first parent information and a path node corresponding to the second parent information; the first parent information and the second parent information are any two pieces of parent information that have an associated relationship among the plurality of parent information; one path relationship corresponds to one path score; For any group of sugarcane hybrid combinations, a parent information collection consisting of parent information corresponding to the father sugarcane parent and parent information corresponding to the mother sugarcane parent in the group of sugarcane hybrid combinations is obtained; Determining at least one candidate path based on the parent information collection and each of the path relationships; for any candidate path among the at least one candidate path, the any candidate path includes at least one path relationship; determining a comprehensive score corresponding to any candidate path based on the path scores corresponding to each path relationship in the at least one path relationship and the reward score corresponding to any candidate path; wherein the reward score corresponding to any candidate path is determined based on environmental information of the region where any group of sugarcane hybrid combinations is to be planted, historical planting performance information of sugarcane hybrid combinations corresponding to each path relationship in the at least one path relationship, and the number of path nodes included in any candidate path; Determine the highest comprehensive score of the candidate paths corresponding to any group of sugarcane hybrid combinations as the final score corresponding to any group of sugarcane hybrid combinations; Based on the final scores corresponding to the sugarcane hybrid combinations, the sugarcane hybrid combinations to be cultivated are determined.

2. The method according to claim 1, characterized in that The parent type information includes identification data for characterizing the identity of the corresponding sugarcane parent; the gene type information includes SNP gene data and / or InDel gene data of the corresponding sugarcane parent; the trait type information includes at least one of the high sugar trait data, disease resistance trait data, and drought resistance trait data of the corresponding sugarcane parent; the growth environment type information includes at least one of the temperature, humidity, daily irrigation amount, soil pH, soil nitrogen content, soil phosphorus content, and soil potassium content of the growth environment of the corresponding sugarcane parent; the metabolic pathway type information includes at least one of the sugar metabolism pathway data, disease resistance-related pathway data, stress resistance-related pathway data, secondary metabolite synthesis pathway data, nutrient utilization efficiency pathway data, hormone regulation pathway data, photosynthesis-related pathway data, and amino acid metabolism pathway data; the offspring performance information includes the parent type information, gene type information, trait type information, growth environment type information, and metabolic pathway type information corresponding to the offspring sugarcane parent obtained based on hybridization of the corresponding sugarcane parents.

3. The method according to claim 2, characterized in that The association relationship between the parent type information and the gene type information is a carrying relationship; the association relationship between the gene type information and the trait type information is a regulatory relationship; the association relationship between the trait type information and the growth environment type information is a dependency relationship; the association relationship between the gene type information and the metabolic pathway type information is a regulatory relationship; the association relationship between the metabolic pathway type information and the trait type information is an influence relationship; the association relationship between the parent type information and the offspring performance information is a hybridization relationship; the association relationship between the offspring performance information and the trait type information is a performance relationship; the hybridization relationship includes obtaining offspring sugarcane parents based on hybridization of corresponding sugarcane parents.

4. The method according to claim 3, characterized in that Before obtaining the parent information collection consisting of the parent information corresponding to the father sugarcane parent and the parent information corresponding to the mother sugarcane parent in any group of sugarcane hybrid combinations, the method further includes: Based on a plurality of sugarcane parents, a plurality of sugarcane hybrid pairs are obtained; for any one of the plurality of sugarcane hybrid pairs, the any one sugarcane hybrid pair includes a first sugarcane parent and a second sugarcane parent; the first sugarcane parent and the second sugarcane parent are any two sugarcane parents included in the plurality of sugarcane parents; determining a genetic distance between the first sugarcane parent and the second sugarcane parent based on the genotype information corresponding to the first sugarcane parent and the genotype information corresponding to the second sugarcane parent; determining any one of the sugarcane hybrid pairs as a candidate sugarcane hybrid combination based on that the genetic distance between the first sugarcane parent and the second sugarcane parent is less than a genetic distance threshold; At least one group of sugarcane hybrid combinations is determined based on each candidate sugarcane hybrid combination determined from the plurality of sugarcane hybrid pairs; the at least one group of sugarcane hybrid combinations includes any one of the groups of sugarcane hybrid combinations.

5. The method according to claim 4, characterized in that The environmental information includes precipitation information, temperature information, and soil information; the method further includes: Determining environmental stress factors; the environmental stress factors include drought environmental stress factors, high temperature environmental stress factors, and soil environmental stress factors; Obtaining standardized values ​​corresponding to each environmental stress factor; the standardized values ​​corresponding to the drought environmental stress factor are determined based on precipitation information of the region where the any group of sugarcane hybrid combinations is to be planted; the standardized values ​​corresponding to the high temperature environmental stress factor are determined based on temperature information of the region where the any group of sugarcane hybrid combinations is to be planted; and the standardized values ​​corresponding to the soil environmental stress factor are determined based on soil information of the region where the any group of sugarcane hybrid combinations is to be planted; determining an environmental stress index for a region where any one of the sugarcane hybrid combinations is to be planted based on the standardized values ​​and weights corresponding to the environmental stress factors; Acquiring historical hybridization data; the historical hybridization data includes genetic distance distributions of multiple historical sugarcane hybrid combinations in a stress-free environment; the stress-free environment is an environment in which the environmental stress index is a preset value; Determining a standard genetic distance threshold based on the genetic distance distributions of the plurality of historical sugarcane hybrid combinations; the standard genetic distance threshold is determined based on the genetic distance that appears most frequently in the genetic distance distributions of the plurality of historical sugarcane hybrid combinations; determining an environmental adjustment coefficient for the region where the any group of sugarcane hybrid combinations are to be planted based on an environmental stress index of the region where the any group of sugarcane hybrid combinations are to be planted, environmental information of the region where the any group of sugarcane hybrid combinations are to be planted, and the historical hybridization data; The genetic distance threshold is determined based on the standard genetic distance threshold, the environmental adjustment coefficient of the area where any group of sugarcane hybrid combinations are to be planted, and the environmental stress index of the area where any group of sugarcane hybrid combinations are to be planted.

6. The method according to claim 5, characterized in that Also includes: Obtaining the Euclidean distance between the growth environment type information corresponding to any candidate path and the environment information of the area where any group of sugarcane hybrid combinations is to be planted; determining an environmental reward score corresponding to any candidate path based on the Euclidean distance between the growth environment type information corresponding to any candidate path and the environmental information of the region where any group of sugarcane hybrid combinations is to be planted; Determining a historical performance reward score corresponding to any candidate path based on planting performance information of historical sugarcane hybrid combinations corresponding to each path relationship in the at least one path relationship; the planting performance information includes sugar content information, yield information, disease resistance information, and drought tolerance survival rate information of offspring sugarcane corresponding to the historical sugarcane hybrid combination; Determining a distance reward score corresponding to any candidate path based on the number of path nodes included in any candidate path; The reward score corresponding to any candidate path is determined based on the environment reward score, the historical performance reward score, and the distance reward score corresponding to any candidate path.

7. The method according to claim 6, characterized in that The determining of at least one candidate path based on the parent information collection and each of the path relationships includes: Based on the parent information collection, determining at least one target path node from the plurality of path nodes; for any target path node from the at least one target path node, the parent information collection includes parent information corresponding to the any target path node; Based on the at least one target path node, at least one candidate path is determined in each of the path relationships; any candidate path in the at least one candidate path includes at least one path relationship; for any path relationship in the at least one path relationship, a path node for constructing the any one path relationship exists in the at least one target path node.

8. The method according to claim 7, characterized in that After determining the sugarcane hybrid combinations to be cultivated based on the final scores corresponding to the sugarcane hybrid combinations, the method further includes: Obtaining parental information of hybrid offspring of the sugarcane hybrid combination to be cultivated; Based on the parent information of the hybrid offspring, the path score corresponding to each path relationship in the at least one path relationship is adjusted.

9. A sugarcane hybrid combination selection system, characterized in that: include: A path node generation module creates multiple path nodes based on multiple parent information corresponding to sugarcane parents; the multiple parent information includes parent type information, gene type information, trait type information, growth environment type information, metabolic pathway type information, and offspring performance information; one path node is used to represent one parent information; A path relationship construction module is configured to construct a path relationship between a path node corresponding to a first parent information and a path node corresponding to a second parent information; the first parent information and the second parent information are any two pieces of parent information that have an associated relationship among the plurality of parent information; and one path relationship corresponds to one path score; A parent information fusion module, for any group of sugarcane hybrid combinations, obtains a parent information collection consisting of parent information corresponding to the father sugarcane parent and parent information corresponding to the mother sugarcane parent in the group of sugarcane hybrid combinations; a candidate path selection module, which determines at least one candidate path based on the parent information collection and each of the path relationships; for any candidate path among the at least one candidate path, the any candidate path includes at least one path relationship; a comprehensive score calculation module, which determines a comprehensive score corresponding to any candidate path based on the path scores corresponding to each path relationship in the at least one path relationship and the reward score corresponding to any candidate path; the reward score corresponding to any candidate path is determined based on environmental information of the area where the group of sugarcane hybrid combinations is to be planted, historical planting performance information of the sugarcane hybrid combinations corresponding to each path relationship in the at least one path relationship, and the number of path nodes included in any candidate path; a final score calculation module, which determines the highest comprehensive score of the candidate paths corresponding to any group of sugarcane hybrid combinations as the final score corresponding to the any group of sugarcane hybrid combinations; The selection module determines the sugarcane hybrid combination to be cultivated based on the final score corresponding to each sugarcane hybrid combination.

10. The system according to claim 9, characterized in that The parent type information includes identification data for characterizing the identity of the corresponding sugarcane parent; the gene type information includes SNP gene data and / or InDel gene data of the corresponding sugarcane parent; the trait type information includes at least one of high sugar trait data, disease resistance trait data, and drought tolerance trait data of the corresponding sugarcane parent; the growth environment type information includes at least one of the temperature, humidity, daily irrigation amount, soil pH, soil nitrogen content, soil phosphorus content, and soil potassium content of the growth environment of the corresponding sugarcane parent; the metabolic pathway type information includes at least one of sugar metabolism pathway data, disease resistance-related pathway data, stress resistance-related pathway data, secondary metabolite synthesis pathway data, nutrient use efficiency pathway data, hormone regulation pathway data, photosynthesis-related pathway data, and amino acid metabolism pathway data; the offspring performance information includes parent type information, gene type information, trait type information, growth environment type information, and metabolic pathway type information corresponding to the offspring sugarcane parent obtained based on hybridization of the corresponding sugarcane parents; The association relationship between the parent type information and the gene type information is a carrying relationship; the association relationship between the gene type information and the trait type information is a regulatory relationship; the association relationship between the trait type information and the growth environment type information is a dependency relationship; the association relationship between the gene type information and the metabolic pathway type information is a regulatory relationship; the association relationship between the metabolic pathway type information and the trait type information is an influence relationship; the association relationship between the parent type information and the offspring performance information is a hybridization relationship; the association relationship between the offspring performance information and the trait type information is a performance relationship; the hybridization relationship includes obtaining offspring sugarcane parents based on hybridization of corresponding sugarcane parents.