A method for matching and screening loofah rootstocks and bitter gourd scions

CN121128467BActive Publication Date: 2026-09-01CROP RES INST OF FUJIAN ACAD OF AGRI SCI +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511456730.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-10-13
Publication Date
2026-09-01
Estimated Expiration
2045-10-13

AI Technical Summary

Technical Problem

[0005]本发明要解决的技术问题是提供一种丝瓜砧木与苦瓜接穗匹配筛选方法以解决如何精准识别和筛选出最适合嫁接的砧木类型的问题

Benefits of technology

本发明通过对丝瓜砧木样本进行基因组测序,构建基因型特征数据库,分析遗传差异性并评估砧木匹配度,进一步提取抗病基因组信息,预测特定环境下的病害抗性值,并结合植株生长势和产量表现等数据,对嫁接后苦瓜的生长潜力和品质稳定性进行综合评估,最终利用信息整合技术确定优选砧木名单,并通过持续数据更新优化筛选依据,本发明实现了丝瓜砧木与苦瓜接穗的精准匹配,提高了嫁接成功率和植株生长表现,为苦瓜种植提供了科学的砧木选择方法。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121128467B_ABST
    Figure CN121128467B_ABST
Patent Text Reader

Abstract

This invention provides a method for matching and screening rootstocks for loofah and scions for bitter gourd, comprising the following steps: Step S101: Obtaining preliminary genetic difference classification results; Step S102: Determining a potential candidate set of rootstocks; Step S103: For each individual in the candidate rootstock set, obtaining the expression potential assessment results of the disease resistance genome; Step S104: Determining the disease resistance ranking of each rootstock individual based on the expression potential assessment results of the disease resistance genome; Step S105: Obtaining predictive indicators of plant growth vigor; Step S106: Determining the priority ranking of yield performance values; Step S107: Obtaining the assessment results of quality stability; Step S108: Determining the final list of preferred rootstocks; Step S109: Obtaining the basis for continuous optimization of rootstock selection. This invention achieves precise matching between loofah rootstocks and bitter gourd scions, improves grafting success rate and plant growth performance, and provides a scientific method for rootstock selection in bitter gourd cultivation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of agricultural production technology, and in particular to a method for matching and screening loofah rootstock and bitter gourd scion. Background Technology

[0002] In agricultural production, vegetable grafting technology has received widespread attention as an important means to improve crop yield and stress resistance. In particular, in the cultivation of cash crops such as bitter gourd, the application of grafting technology is of key significance to ensuring yield and quality.

[0003] Grafting can combine the superior traits of bitter gourd with the disease resistance and adaptability of rootstock, thereby achieving higher economic benefits. However, in actual production, the selection and matching of different rootstocks often involve many uncertainties, which directly affect the grafting effect and the final crop performance. Existing grafting methods mostly rely on experience or simple phenotypic observation to select rootstocks. This approach often falls short when facing complex environmental conditions and disease pressures, especially in terms of the genetic compatibility between rootstock and scion and the stability of resistance gene expression. This lack of in-depth scientific evidence leads to inconsistent yields and disease resistance in grafted plants, sometimes with significant fluctuations. This uncertainty not only increases production costs but also limits the wider application of grafting technology.

[0004] In this field, the core challenge lies in how to accurately identify and select the most suitable rootstock type for grafting to ensure the growth potential and disease resistance of bitter gourd plants after grafting. First, due to the significant differences in the genetic background of different loofah rootstocks, it is difficult to accurately determine their matching degree with bitter gourd scions in the early stages, which directly affects the grafting survival rate and the overall performance of the plants in the later stages. Second, this genetic difference further leads to uncertainty in the effective expression of resistance genes, making it impossible for some rootstocks to provide sufficient protection against specific diseases, thereby affecting the yield and quality of bitter gourd. Summary of the Invention

[0005] The technical problem to be solved by the present invention is to provide a method for matching and screening loofah rootstock and bitter gourd scion in order to solve the problem of how to accurately identify and screen the most suitable rootstock type for grafting.

[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: A method for matching and screening loofah rootstocks and bitter gourd scions includes the following steps: Step S101: By performing genome sequencing on loofah rootstock samples, genetic difference-related data are obtained, and a genotype characteristic database is constructed for different rootstock individuals to obtain preliminary genetic difference classification results; Step S102: Use data comparison method to analyze the rootstock matching characteristics of loofah rootstock and bitter gourd scion, and perform preliminary screening in combination with preset matching threshold to determine the potential candidate set of rootstocks. Step S103: For each individual in the rootstock candidate set, extract disease resistance genome-related information from the genotype feature database, analyze the distribution of disease resistance genes through information comparison technology, and obtain the expression potential assessment results of the disease resistance genome; Step S104: Based on the evaluation results of the expression potential of the disease-resistant genome and combined with environmental adaptability-related data, the disease resistance value of the rootstock under specific environments is predicted using data analysis methods, and the disease resistance ability ranking of each rootstock individual is determined. Step S105: Based on the disease resistance value ranking results, obtain historical data related to plant growth potential, and simulate the growth potential of grafted bitter gourd plants using data modeling technology to obtain predictive indicators of plant growth potential. Step S106: Based on the predicted indicators of plant growth vigor and combined with relevant data on yield performance, a comprehensive evaluation method is used to quantitatively analyze the post-grafting yield potential of each rootstock individual and determine the priority ranking of yield performance values. Step S107: Prioritize the yield performance values, obtain relevant data on quality stability, predict the quality fluctuation range of grafted bitter gourd through data comparison technology, and obtain the evaluation results of quality stability. Step S108: Based on the evaluation results of quality stability and combined with the relevant data on screening accuracy, use information integration technology to comprehensively score all candidate rootstocks and determine the final list of preferred rootstocks. Step S109: For the final list of preferred rootstocks, data recording technology is used to store the screening process and results in the database. The data of subsequent grafting experiments are supplemented through an automatic update mechanism to obtain continuously optimized rootstock screening criteria.

[0007] Preferably, step S101 specifically includes the following steps: S1011: By performing genome sequencing on loofah rootstock samples, the raw genome sequence data was obtained, and preliminary data collection was completed; S1012: Based on the obtained genome sequence data, a data extraction process is used to separate key fragments related to genetic differences, resulting in a set of differentially expressed fragments; S1013: For the set of differential fragments, the support vector machine algorithm is applied to perform feature selection, extract genotype features with significant distinguishability, and determine the feature identifier set; S1014: If the data distribution in the feature identifier set meets the preset threshold conditions, then it is entered into the pre-established feature database to complete the storage of genotype features. S1015: By performing cluster analysis on the genotype features in the feature database, the genetic differences of individuals from different rootstocks are grouped to obtain preliminary individual classification results; S1016: Based on the individual classification results and combined with the set of differential fragments, analyze the genetic difference patterns between groups and determine whether there are significant classification boundaries. S1017: If the classification boundaries are clear, the classification results are used to generate structured data for genetic difference analysis to complete the final classification mapping.

[0008] Preferably, step S102 specifically includes the following steps: S1021: Using genetic difference data, the initial feature sets of loofah rootstock and bitter gourd scion are obtained from the classification results. The data comparison method is used to extract features and obtain the preliminary distribution of matching degree features. S1022: Based on the preliminary distribution of matching degree features, and considering the genetic differences between loofah rootstock and bitter gourd scion, the support vector machine algorithm is used to classify features, determine the correlation weights between features, and obtain the feature combinations after classification. S1023: Obtain the classified feature combination, and perform quantitative analysis of the matching degree feature in combination with the preset threshold. If the matching degree feature value is higher than the preset threshold, it is determined to be a potentially suitable combination, and the preliminary screening of matching pairs is obtained. S1024: For the initially screened matching pairs, the genetic differences between the loofah rootstock and bitter gourd scion are compared in depth using data comparison methods to determine whether there are hidden matching degree features and obtain the matching pair set after in-depth analysis. S1025: Based on the matching pair set after deep analysis, cluster analysis is used to group the potential fit combinations, determine the matching degree feature distribution of rootstock candidates and scions in each group, and obtain the grouped candidate set; S1026: Based on the grouped candidate sets, for the matching degree feature distribution within each group, if the consistency of the feature distribution of a certain group is higher than the preset threshold, it is determined to be a rootstock candidate with high adaptability, and the final adaptation candidate set is obtained. S1027: Obtain the final candidate set of rootstocks, and sort each group of rootstock candidates by combining genetic differences and matching degree features. Determine the stability of the matching degree features in the sorting results and determine the optimal combination of rootstock candidates.

[0009] Preferably, step S103 specifically includes the following steps: S1031: By obtaining the genotypic characteristic data of rootstock candidates from the database, preliminary information is organized for each individual to obtain the basic dataset of individual genotypes; S1032: Based on the basic dataset, information comparison technology is used to analyze the distribution of disease resistance genes for each individual and determine the location and distribution information of disease resistance genes in the genome; S1033: If the distribution density of disease-resistant genes in the location distribution information is higher than the preset threshold, then the corresponding individuals are marked as high-potential individuals, and the genome information dataset of high-potential individuals is obtained. S1034: Analyze the association between disease resistance genes and genotype characteristics in genomic information datasets of high-potential individuals to determine the expression potential intensity of disease resistance genes; S1035: Based on the analysis results of expression potential intensity and combined with distribution data, the disease resistance genome of each high-potential individual is comprehensively scored to obtain a comprehensive potential assessment value; S1036: Based on the comprehensive potential assessment value, all high-potential individuals are sorted to determine the final priority list of disease resistance gene expression potential; S1037: The random forest algorithm is used to perform validation analysis on individuals in the priority list. Combined with the association data between genotype characteristics and disease resistance genes, the reliability of the evaluation results is judged, and the final disease resistance genome potential analysis dataset is output.

[0010] Preferably, step S104 specifically includes the following steps: S1041: By obtaining relevant data on disease resistance genes from the database and combining them with the original records of gene expression, a preliminary assessment of expression potential is compiled. S1042: Based on the preliminary assessment results of expression potential, construct a comprehensive impact factor dataset for specific environments by integrating data on environmental adaptation and adaptation information; S1043: The random forest algorithm is used to process the comprehensive impact factor dataset, calculate the predicted values ​​of disease resistance, and obtain the resistance value of each rootstock individual; S1044: For the resistance values ​​of each individual rootstock, perform data standardization processing to obtain quantitative indicators of individual capabilities; S1045: By using quantitative indicators of individual capabilities and combining data analysis methods, rank all rootstock individuals by capability and determine the ranking results; S1046: If there are rootstock individuals with similar resistance values ​​in the ranking results, the final individual ability ranking is determined by comparing environmental adaptation data under a specific environment. S1047: After obtaining the final individual ability ranking, generate the corresponding disease resistance distribution map to determine the disease resistance location of each rootstock individual in a specific environment.

[0011] Preferably, step S105 specifically includes the following steps: S1051: Construct an initial dataset by obtaining relevant data from the disease resistance ranking results; S1052: Data cleaning techniques are used to preprocess the sorting results, removing outliers and missing values ​​to obtain a cleaned resistant dataset. S1053: Based on the organized resistance dataset, extract historical data related to plant growth; S1054: Perform feature selection on historical data, filter out feature variables that are highly correlated with growth potential, and determine the feature dataset; S1055: Input data for growth simulation is constructed by combining feature datasets with grafting treatment information of bitter gourd plants; S1056: If the values ​​of some variables in the feature dataset are lower than the preset threshold, they are standardized to obtain normalized simulated input data; S1057: A random forest model is used to train the normalized simulated input data, and the growth potential of bitter gourd plants is simulated and calculated to obtain preliminary growth potential predictions. S1058: Based on the preliminary growth potential predictions and combined with the plant growth trends in historical data, the prediction results are corrected. S1059: If the deviation between the predicted value and the historical trend exceeds the preset range, the predicted value is smoothed to obtain the corrected growth potential value. S10510: A predictive index of plant growth potential is generated using the corrected growth potential value. S10511: Perform multi-dimensional analysis on the prediction indicators to determine their stability under different environmental conditions and obtain the final prediction results.

[0012] Preferably, step S106 specifically includes the following steps: S1061: By obtaining plant growth potential prediction index data and yield performance data from the database, an initial dataset is formed, and the initial data integration is completed. S1062: Based on the initial dataset, the grafting yield potential of each rootstock individual is quantified using a pre-established comprehensive evaluation model to obtain the potential evaluation results. S1063: Based on the potential assessment results, if the potential value of a certain rootstock individual exceeds the preset threshold, it will be marked as a high-potential individual, and the initial screening will be completed. S1064: By conducting in-depth analysis of the yield performance data of high-potential individuals, obtain their correlation coefficients with growth indicators, and identify key influencing factors; S1065: Based on key influencing factors, a random forest model is used to predict yield performance values, and the predicted yield ranking of each rootstock individual is obtained. S1066: For the predicted yield ranking, if a certain rootstock individual ranks at the top, its growth analysis data is further extracted to determine its stability characteristics. S1067: By combining stability characteristics with yield potential data, the final priority ranking result is generated, completing the full-process analysis.

[0013] Preferably, step S107 specifically includes the following steps: S1071: By obtaining yield and quality data of grafted bitter gourd from historical records, an initial dataset is constructed to obtain basic information for subsequent analysis; S1072: Based on the initial dataset, prioritize the output performance values, use a preset threshold to filter out the data subset with high performance values, and determine the key analysis objects; S1073: By standardizing the quality data in the high-performance data subset, the distribution characteristics of the quality data are obtained, and its initial fluctuation range is determined. S1074: If the initial fluctuation range exceeds the preset threshold, further data comparison and analysis of the quality data will be performed, and the support vector machine algorithm will be used to accurately predict the fluctuation range to obtain the predicted fluctuation value. S1075: Based on the predicted fluctuation value and combined with the historical quality stability record of grafted bitter gourd, obtain the intermediate index for stability assessment and determine the degree of influence of the fluctuation range on stability. S1076: By comparing intermediate indicators with preset stability evaluation standards, it is determined whether the quality stability of grafted bitter gourd meets expectations, and the final evaluation result is obtained.

[0014] Preferably, step S108 specifically includes the following steps: S1081: Obtain quality stability data and screening accuracy data of rootstock candidate individuals, and integrate relevant indicators from multiple sources through a pre-established data collection system to form an initial dataset; S1082: For the initial dataset, information integration technology is used to standardize the quality stability data and screening accuracy data to obtain a unified evaluation data matrix; S1083: Based on the evaluation data matrix, each candidate rootstock is comprehensively scored using a preset scoring system. If the score of a candidate is lower than the preset threshold, it is removed from the candidate list to obtain preliminary screening results. S1084: By analyzing the individuals in the preliminary screening results and combining data fusion technology, multiple indicators of the remaining rootstock candidates are weighted and calculated to determine the weighted comprehensive score. S1085: Based on the weighted comprehensive score, sort the remaining rootstock candidates. If the comprehensive score of a candidate is higher than the preset upper limit threshold, mark it as a priority candidate and obtain the sorted candidate list. S1086: For the sorted candidate list, apply the decision tree algorithm to perform final verification on the priority objects, determine whether they meet the requirements of the preferred list, and output the set of the verified preferred individuals. S1087: By comparing the data of the verified preferred individual set, information integration technology is used to generate the final list of preferred rootstocks, and the final result that meets the requirements of quality stability and screening accuracy is determined.

[0015] Preferably, step S109 specifically includes the following steps: S1091: Obtain relevant information about the rootstock list from the screening process using data acquisition tools, organize it in a structured format, and store it in a pre-established database to obtain the initial rootstock screening dataset. S1092: Based on the initial rootstock screening dataset, classify the screening process information recorded therein, extract key indicators using preset field rules, and determine the core basis data in the screening process. S1093: If the core data meets the preset threshold conditions, the latest experimental data will be obtained from the grafting experiment through an automatic update mechanism and added to the database to determine whether to adjust the existing rootstock list. S1094: If the indicators of some rootstocks change after the experimental data are supplemented, the random forest algorithm is used to predict and analyze the trend of change to obtain the potential optimization direction of the rootstock list. S1095: Based on the optimization direction obtained from the predictive analysis, compare and verify the subsequent experimental data stored in the database, obtain experimental results consistent with the optimization direction, and determine the adjustment plan for the rootstock list. S1096: Update the rootstock list in the database by adjusting the scheme, and in conjunction with the goal of continuous improvement, associate the updated list with the screening criteria to obtain the latest screening reference data; S1097: For the latest screening reference data, use an automated data recording method to save it to the database, forming a closed-loop data supplementation and optimization process to judge the continuous improvement effect of the rootstock screening criteria.

[0016] Compared with the prior art, the present invention has at least the following beneficial effects: This invention utilizes genome sequencing of loofah rootstock samples to construct a genotype characteristic database, analyzes genetic differences and assesses rootstock matching, further extracts disease-resistant genome information, predicts disease resistance values ​​under specific environments, and comprehensively evaluates the growth potential and quality stability of grafted bitter gourd by combining data such as plant growth vigor and yield performance. Finally, it uses information integration technology to determine a list of preferred rootstocks and continuously updates and optimizes the screening criteria. This invention achieves precise matching between loofah rootstocks and bitter gourd scions, improves grafting success rate and plant growth performance, and provides a scientific method for rootstock selection in bitter gourd cultivation. Attached Figure Description

[0017] The accompanying drawings, which are incorporated herein and form part of the specification, illustrate embodiments of the present disclosure and, together with the specification, further serve to explain the principles of the present disclosure and enable those skilled in the art to implement and use the present disclosure.

[0018] Figure 1 A flowchart illustrating the method for matching and screening loofah rootstocks and bitter gourd scions.

[0019] Figure 2 A flowchart illustrating step S102 of the method for matching and screening loofah rootstocks and bitter gourd scions.

[0020] Figure 3 A flowchart illustrating step S109 of the method for matching and screening loofah rootstocks and bitter gourd scions. Detailed Implementation

[0021] The following describes in detail a method for matching and screening loofah rootstock and bitter gourd scion provided by the present invention, with reference to the accompanying drawings and specific embodiments. It should be noted that, to make the embodiments more detailed, the following embodiments are the best and preferred embodiments; those skilled in the art can also use other alternative methods to implement some known technologies; and the accompanying drawings are only for more specific description of the embodiments and are not intended to specifically limit the present invention.

[0022] It should be noted that the use of terms such as "an embodiment," "an embodiment," "an exemplary embodiment," and "some embodiments" in the specification indicates that the described embodiment may include a specific feature, structure, or characteristic, but not every embodiment necessarily includes that specific feature, structure, or characteristic. Furthermore, when a specific feature, structure, or characteristic is described in connection with an embodiment, implementing such a feature, structure, or characteristic in conjunction with other embodiments (whether explicitly described or not) should be within the knowledge of those skilled in the art.

[0023] like Figure 1-3 As shown, an embodiment of the present invention provides a method for matching and screening loofah rootstock and bitter gourd scion, comprising the following steps: Step S101: By performing genome sequencing on loofah rootstock samples, genetic difference-related data are obtained, and a genotype characteristic database is constructed for different rootstock individuals to obtain preliminary genetic difference classification results; specifically, the following steps are included: S1011: By performing genome sequencing on loofah rootstock samples, the raw genome sequence data was obtained, and preliminary data collection was completed; S1012: Based on the obtained genome sequence data, a data extraction process is used to separate key fragments related to genetic differences, resulting in a set of differentially expressed fragments; S1013: For the set of differential fragments, the support vector machine algorithm is applied to perform feature selection, extract genotype features with significant distinguishability, and determine the feature identifier set; S1014: If the data distribution in the feature identifier set meets the preset threshold conditions, then it is entered into the pre-established feature database to complete the storage of genotype features. S1015: By performing cluster analysis on the genotype features in the feature database, the genetic differences of individuals from different rootstocks are grouped to obtain preliminary individual classification results; S1016: Based on the individual classification results and combined with the set of differential fragments, analyze the genetic difference patterns between groups and determine whether there are significant classification boundaries. S1017: If the classification boundaries are clear, the classification results are used to generate structured data for genetic difference analysis to complete the final classification mapping.

[0024] A specific application example of step S101: For example, when performing genome sequencing on a loofah rootstock sample, the raw sequence data can be obtained through high-throughput sequencing technology.

[0025] Assuming that a single sequencing operation covers more than 95% of the sample genome, yielding approximately 1 billion short read sequences, this data lays the foundation for subsequent analysis, ensuring the comprehensiveness and reliability of the data.

[0026] In the data extraction process, for the separation of genetically differentially related fragments, alignment tools can be used to match the sequence with a reference genome to screen out variant sites such as single nucleotide polymorphism fragments.

[0027] Suppose that about 5 million differential fragments are extracted from 1 billion sequences. This collection of fragments provides key material for subsequent feature selection.

[0028] When using the support vector machine algorithm for feature selection on a set of differential fragments, a classification model can be constructed to identify genotype features that have significant distinguishing power for genetic differences in rootstocks.

[0029] Suppose that 1,000 key features are selected from 5 million segments to form a feature identifier set.

[0030] This step effectively reduces the dimensionality of the data and improves analysis efficiency.

[0031] If the data distribution of the feature identifier set meets the preset threshold conditions, such as the correlation between features being less than 0.2, then it will be entered into the feature database.

[0032] This process ensures the standardized storage of data, facilitating subsequent retrieval and analysis.

[0033] By performing cluster analysis on the genotype characteristics in the feature database, the K-means algorithm can be used to divide individuals of different rootstocks into three groups with different genetic differences.

[0034] Assuming that the 100 individuals in the sample are divided into three groups, A, B, and C, with each group containing 40, 30, and 30 individuals respectively, the preliminary classification results provide a basis for subsequent differential pattern analysis.

[0035] Based on the individual classification results, when analyzing the genetic difference patterns between groups using differential fragment set analysis, it can be found that group A and group B have significant differences at certain gene loci, while group C shows an intermediate state.

[0036] If statistical analysis determines that the classification boundaries are clear, for example, if the significance level of the difference between groups is below 0.05, then the validity of the classification can be confirmed.

[0037] If the classification boundaries are clear, structured data for genetic difference analysis is generated, such as recording the distribution of characteristics and differential loci in each group in tabular form, and finally completing the classification mapping.

[0038] This result can be used to guide rootstock selection and improve breeding precision.

[0039] For example, from another perspective, when selecting features, if the support vector machine algorithm is combined with other machine learning methods such as random forest for verification, the accuracy of the feature label set can be further improved.

[0040] If the overlap rate of the screening results from the two methods reaches 80%, then the reliability of the features is enhanced.

[0041] For example, in cluster analysis, if the genetic distance matrix is ​​used for group validation, and the genetic distance between group A and group B is 0.3, while the distance between group A and group C is 0.15, then the classification results are more biologically meaningful.

[0042] These multi-faceted verifications collectively support the scientific validity of the classification, providing a solid foundation for subsequent applications, while significantly enhancing the practical value of genetic difference analysis.

[0043] Step S102: Based on the genetic difference classification results, the matching characteristics of the rootstock and scion of loofah are analyzed using data comparison methods. Preliminary screening is then performed using a preset matching threshold to determine a potential candidate set of rootstocks. Specifically, this includes the following steps: S1021: Using genetic difference data, the initial feature sets of loofah rootstock and bitter gourd scion are obtained from the classification results. The data comparison method is used to extract features and obtain the preliminary distribution of matching degree features. S1022: Based on the preliminary distribution of matching degree features, and considering the genetic differences between loofah rootstock and bitter gourd scion, the support vector machine algorithm is used to classify features, determine the correlation weights between features, and obtain the feature combinations after classification. S1023: Obtain the classified feature combination, and perform quantitative analysis of the matching degree feature in combination with the preset threshold. If the matching degree feature value is higher than the preset threshold, it is determined to be a potentially suitable combination, and the preliminary screening of matching pairs is obtained. S1024: For the initially screened matching pairs, the genetic differences between the loofah rootstock and bitter gourd scion are compared in depth using data comparison methods to determine whether there are hidden matching degree features and obtain the matching pair set after in-depth analysis. S1025: Based on the matching pair set after deep analysis, cluster analysis is used to group the potential fit combinations, determine the matching degree feature distribution of rootstock candidates and scions in each group, and obtain the grouped candidate set; S1026: Based on the grouped candidate sets, for the matching degree feature distribution within each group, if the consistency of the feature distribution of a certain group is higher than the preset threshold, it is determined to be a rootstock candidate with high adaptability, and the final adaptation candidate set is obtained. S1027: Obtain the final candidate set of rootstocks, and sort each group of rootstock candidates by combining genetic differences and matching degree features. Determine the stability of the matching degree features in the sorting results and determine the optimal combination of rootstock candidates.

[0044] Specific application example of step S102: For example, when studying the genetic compatibility of loofah rootstock and bitter gourd scion, one can start by obtaining genetic difference data and gradually delve into the screening and optimization of matching pairs.

[0045] To obtain the initial feature set, the genomic data of loofah rootstock and bitter gourd scion can be compared to extract key differences such as single nucleotide polymorphism sites, forming a preliminary set containing hundreds of features.

[0046] These characteristics may include differences in gene fragment length or variation frequency at specific sites, laying the foundation for subsequent analysis.

[0047] For example, in feature extraction and matching degree distribution analysis, we can assume that the extracted features are preliminarily statistically analyzed. Suppose that the mutation rate of a certain gene segment in the loofah rootstock sample A is 0.05, while the corresponding mutation rate in the bitter gourd scion sample B is 0.07. By comparing the results, the matching degree score between the two samples may be 0.85, which preliminarily indicates that they have a certain potential for adaptation.

[0048] This method helps to quickly filter out potentially suitable combinations.

[0049] For example, for feature classification using the support vector machine algorithm, weights can be assigned to the correlation between features. Suppose that the influence weight of a certain feature on the classification is 0.6, while that of another feature is only 0.2, thus prioritizing the feature with the higher weight.

[0050] This method can effectively distinguish which genetic differences have a greater impact on fitness, providing a basis for subsequent screening.

[0051] In the quantitative analysis of matching characteristics, a threshold such as 0.8 can be set. If the matching score of a certain rootstock and scion is 0.82, it will be included in the potential matching combination.

[0052] This quantitative approach helps reduce subjective judgment and improve screening efficiency.

[0053] For deep alignment and the mining of hidden features, more detailed genomic fragment alignment can be used to discover subtle differences that are not covered by the initial feature set, such as the insertion or deletion of a gene segment, and then the matching pair set can be adjusted.

[0054] This in-depth analysis can improve the comprehensiveness of the suitability assessment.

[0055] In cluster analysis and the determination of candidate groups, combinations with similar matching scores can be grouped together. If the matching scores of the five combinations in a group are all between 0.80 and 0.85 and have high consistency, then the group can be considered to have good fitting potential.

[0056] This grouping method helps to focus on highly adaptable combinations.

[0057] For the ranking and stability assessment of the final candidate set, the matching degree characteristics of each group of rootstock candidates can be evaluated in multiple dimensions. For example, combining the stability of genetic differences, if a certain combination has a matching degree fluctuation of only 0.02 in multiple tests, it is considered to have high stability and should be given priority as the optimal candidate.

[0058] This sorting method ensures that the final selected combination has high reliability.

[0059] Through the above multi-level analysis and screening, we can not only accurately identify the compatibility between loofah rootstock and bitter gourd scion from the perspective of genetic differences, but also provide data support for subsequent grafting experiments, significantly improving the compatibility efficiency and success rate of grafting combinations.

[0060] Step S103: For each individual in the rootstock candidate set, extract disease resistance genome-related information from the genotype feature database, analyze the distribution of disease resistance genes using information comparison technology, and obtain the expression potential assessment results of the disease resistance genome; specifically including the following steps: S1031: By obtaining the genotypic characteristic data of rootstock candidates from the database, preliminary information is organized for each individual to obtain the basic dataset of individual genotypes; S1032: Based on the basic dataset, information comparison technology is used to analyze the distribution of disease resistance genes for each individual and determine the location and distribution information of disease resistance genes in the genome; S1033: If the distribution density of disease-resistant genes in the location distribution information is higher than the preset threshold, then the corresponding individuals are marked as high-potential individuals, and the genome information dataset of high-potential individuals is obtained. S1034: Analyze the association between disease resistance genes and genotype characteristics in genomic information datasets of high-potential individuals to determine the expression potential intensity of disease resistance genes; S1035: Based on the analysis results of expression potential intensity and combined with distribution data, the disease resistance genome of each high-potential individual is comprehensively scored to obtain a comprehensive potential assessment value; S1036: Based on the comprehensive potential assessment value, all high-potential individuals are sorted to determine the final priority list of disease resistance gene expression potential; S1037: The random forest algorithm is used to perform validation analysis on individuals in the priority list. Combined with the association data between genotype characteristics and disease resistance genes, the reliability of the evaluation results is judged, and the final disease resistance genome potential analysis dataset is output.

[0061] Specific application example of step S103: For example, when studying the compatibility between loofah rootstock and bitter gourd scion, and obtaining genotypic characteristic data of rootstock candidates from the database, one can first focus on specific fragment information in the gene sequence.

[0062] Assuming the database contains 100 candidate individuals for loofah rootstock, and each individual's genotype data includes marker loci information on chromosomes, by organizing this data, a basic dataset containing genotype markers and individual numbers is formed, laying the foundation for subsequent analysis.

[0063] Specifically, when using information comparison technology to analyze the distribution of disease resistance genes, the location of disease resistance genes in the genome can be determined by comparing the genome sequence of each individual with a known database of disease resistance genes.

[0064] For example, five disease resistance gene markers were found in the genome of a certain loofah rootstock individual, distributed on chromosomes 2 and 5, with three markers on chromosome 2, resulting in a distribution density of 1.5 markers per million base pairs.

[0065] If the preset threshold is 1.2 markers per 1 million base pairs, the individual is marked as a high-potential individual, and its genomic information will be further extracted for in-depth research.

[0066] For example, when analyzing the association between disease-resistant genes and genotypic characteristics in high-potential individuals, attention can be paid to the characteristics of gene expression regulatory regions.

[0067] Suppose that there is an enhanced signal in the promoter region near the disease resistance gene of a high-potential individual, indicating that its expression potential is high.

[0068] By comparing the regulatory region characteristics of 10 high-potential individuals, it was found that 6 of them had similar enhanced signals and expressed potential scores of 8.5 or higher, while the other individuals scored around 6.0, indicating that the former has greater application value.

[0069] Specifically, the calculation of the comprehensive potential assessment value can combine distribution density and expression potential score.

[0070] For example, if the distribution density is set to 40% and the expression potential is set to 60%, an individual's comprehensive score is calculated to be 8.2, while another individual's score is 7.5. Finally, the individuals are ranked according to their scores, with the former having higher priority.

[0071] This method helps to screen for rootstock candidates with greater disease resistance potential.

[0072] For example, when using the random forest algorithm to validate a priority list, the reliability of the evaluation results can be judged by constructing a feature matrix that includes genotype markers and disease resistance gene association data.

[0073] Assuming the validation results show that the predicted results of 4 out of the top 5 individuals are highly consistent with the actual genotype characteristics, with a reliability score of over 85%, the evaluation system is considered to be relatively credible.

[0074] This verification method can improve the scientific validity of the screening results and provide a reliable basis for subsequent research on the compatibility of rootstock and scion.

[0075] Specifically, the output of the disease resistance genome potential analysis dataset can be presented in tabular form, including information such as individual ID, comprehensive score, and priority.

[0076] For example, the individual ranked first in the output dataset has a comprehensive score of 8.8, and its disease resistance gene distribution density and expression potential are both excellent, making it a key candidate rootstock.

[0077] This intuitive presentation method helps researchers quickly identify target individuals, improving research efficiency.

[0078] Step S104: Based on the evaluation results of the expression potential of the disease-resistant genome and combined with environmental adaptability-related data, the disease resistance value of rootstocks under specific environments is predicted using data analysis methods to determine the ranking of disease resistance ability for each individual rootstock; specifically, this includes the following steps: S1041: By obtaining relevant data on disease resistance genes from the database and combining them with the original records of gene expression, a preliminary assessment of expression potential is compiled. S1042: Based on the preliminary assessment results of expression potential, construct a comprehensive impact factor dataset for specific environments by integrating data on environmental adaptation and adaptation information; S1043: The random forest algorithm is used to process the comprehensive impact factor dataset, calculate the predicted values ​​of disease resistance, and obtain the resistance value of each rootstock individual; S1044: For the resistance values ​​of each individual rootstock, perform data standardization processing to obtain quantitative indicators of individual capabilities; S1045: By using quantitative indicators of individual capabilities and combining data analysis methods, rank all rootstock individuals by capability and determine the ranking results; S1046: If there are rootstock individuals with similar resistance values ​​in the ranking results, the final individual ability ranking is determined by comparing environmental adaptation data under a specific environment. S1047: After obtaining the final individual ability ranking, generate the corresponding disease resistance distribution map to determine the disease resistance location of each rootstock individual in a specific environment.

[0079] Specific application example of step S104: For example, when obtaining disease resistance gene-related data from a database, one can first focus on the genotype information in the rootstock candidate set and extract gene fragment data that are directly related to disease resistance.

[0080] Suppose the database stores the genetic information of 1,000 rootstock individuals, including specific gene loci data related to disease resistance. By screening out these loci, the number and distribution characteristics of disease resistance genes in each individual can be preliminarily identified, laying the foundation for subsequent expression potential assessment.

[0081] In one possible implementation, when combining raw records of gene expression, data on gene expression levels measured in the laboratory can be matched with database information.

[0082] For example, when examining a specific rootstock individual, it was found that the expression level of its disease resistance gene under specific disease stress was twice the normal level, indicating that it has strong potential disease resistance.

[0083] Such data integration helps to more accurately assess expressive potential.

[0084] For example, when constructing a comprehensive influencing factor dataset by integrating environmental adaptation information, the effects of environmental variables such as temperature and humidity on gene expression can be considered.

[0085] Assuming that the expression of disease resistance genes in a certain rootstock individual is suppressed under high temperature and high humidity conditions, but significantly enhanced under mild conditions, a correlation model between environment and disease resistance is constructed by recording these differential data, providing a basis for subsequent predictions.

[0086] In one possible implementation, when using the random forest algorithm to process the comprehensive influencing factor dataset, environmental variables and gene expression data can be used as input features to predict the disease resistance value of each rootstock individual.

[0087] Suppose that, based on algorithmic analysis, the predicted resistance value of a certain rootstock individual is 85, while that of another individual is 60. This provides a quantitative basis for subsequent ranking.

[0088] This method can effectively integrate multi-dimensional data and improve the reliability of predictions.

[0089] For example, when standardizing individual resistance values, the values ​​of all individuals can be mapped to a range of 0 to 100, making it easier to compare them intuitively.

[0090] Suppose that the original resistance value of a certain rootstock individual is 75, and after standardization it is 80. This processing method helps to eliminate the difference in data units and improve the fairness of the comparison.

[0091] In one possible implementation, when ranking the abilities of individual rootstocks, if two individuals are found to have the same resistance value of 82, their adaptation data in specific environments, such as drought resistance or low temperature resistance data, are further compared. If an individual performs more stably in high temperature environments, it is given higher priority.

[0092] This refined comparison ensures that the sorting results are more closely aligned with actual application scenarios.

[0093] For example, when generating a disease resistance distribution map, the resistance values ​​of all rootstock individuals can be displayed in the form of a bar chart, intuitively showing the ability positioning of each individual.

[0094] If a certain rootstock individual is in the top 10% of the distribution map, it indicates that it has outstanding disease resistance under specific conditions. This visualization method facilitates the rapid identification of potential individuals and provides a reference for subsequent breeding selection.

[0095] Through the above multi-dimensional analysis and processing, the comprehensiveness and practicality of the rootstock disease resistance assessment can be effectively improved.

[0096] Step S105: Based on the disease resistance value ranking results, historical data related to plant growth potential are obtained. Data modeling techniques are used to simulate the growth potential of grafted bitter gourd plants to obtain predictive indicators of plant growth potential. Specifically, this includes the following steps: S1051: Construct an initial dataset by obtaining relevant data from the disease resistance ranking results; S1052: Data cleaning techniques are used to preprocess the sorting results, removing outliers and missing values ​​to obtain a cleaned resistant dataset. S1053: Based on the organized resistance dataset, extract historical data related to plant growth; S1054: Perform feature selection on historical data, filter out feature variables that are highly correlated with growth potential, and determine the feature dataset; S1055: Input data for growth simulation is constructed by combining feature datasets with grafting treatment information of bitter gourd plants; S1056: If the values ​​of some variables in the feature dataset are lower than the preset threshold, they are standardized to obtain normalized simulated input data; S1057: A random forest model is used to train the normalized simulated input data, and the growth potential of bitter gourd plants is simulated and calculated to obtain preliminary growth potential predictions. S1058: Based on the preliminary growth potential predictions and combined with the plant growth trends in historical data, the prediction results are corrected. S1059: If the deviation between the predicted value and the historical trend exceeds the preset range, the predicted value is smoothed to obtain the corrected growth potential value. S10510: A predictive index of plant growth potential is generated using the corrected growth potential value. S10511: Perform multi-dimensional analysis on the prediction indicators to determine their stability under different environmental conditions and obtain the final prediction results.

[0097] Specific application example of step S105: For example, when constructing the initial dataset, the resistance values ​​and related environmental parameters of bitter gourd rootstocks can be extracted from the disease resistance ranking results. Assuming there are 100 rootstock samples, each containing resistance values ​​and environmental humidity data, a basic table containing multidimensional information can be formed.

[0098] This method ensures the integrity of the data in subsequent analyses.

[0099] In one possible implementation, data cleaning techniques can be applied to handle outliers and missing values.

[0100] If the resistance value of a certain rootstock sample is abnormally high (3 times higher than the average), it can be considered an outlier and removed. For missing environmental data, it can be filled in by averaging the values ​​of adjacent samples.

[0101] This preprocessing helps improve the reliability of the dataset.

[0102] For example, when extracting historical data related to plant growth, growth records of bitter gourd plants over the past three years can be retrieved from the database, including information such as plant height and number of leaves. These records can then be correlated with the resistance dataset to identify indicators that have a significant impact on growth.

[0103] This correlation analysis lays the foundation for subsequent feature selection.

[0104] In one possible implementation, feature selection can focus on variables that are highly correlated with growth potential, such as soil pH and light duration, assuming that five key features are selected to form a feature dataset.

[0105] This screening method can reduce the interference of irrelevant variables and improve the accuracy of simulation.

[0106] For example, when constructing growth simulation input data by combining grafting treatment information of bitter gourd plants, information such as grafting method and grafting time can be integrated with the feature dataset. Assuming that a certain batch of rootstocks adopts the top grafting method, the growth status of the rootstocks 30 days after grafting can be recorded as input parameters.

[0107] This integration approach helps to simulate growth scenarios that are closer to reality.

[0108] In one possible implementation, variables in the feature dataset that are below a threshold, such as samples with light exposure duration of less than 8 hours / day, can be adjusted to a uniform range through standardization.

[0109] This standardization ensures data consistency, providing a stable foundation for model training.

[0110] For example, when training a random forest model, the normalized input data can be divided into a training set and a test set to simulate the growth potential of bitter gourd plants under different environments and obtain the predicted value for each sample.

[0111] This method can capture complex relationships between data and improve the accuracy of predictions.

[0112] In one possible implementation, when correcting the prediction results, if the initial prediction shows that a certain rootstock has a growth potential of 20 cm per year, but the historical trend is 15 cm, it can be adjusted by a weighted average method.

[0113] This correction method makes the predictions more consistent with actual growth patterns.

[0114] For example, when smoothing predicted values, if the deviation exceeds a preset range of 10%, the abnormal fluctuations can be smoothed using the moving average method to obtain a more stable growth potential value.

[0115] This process helps reduce the impact of noise in predictions.

[0116] In one possible implementation, when generating plant growth potential prediction indicators, the corrected potential value can be converted into a grade indicator, such as high, medium, and low grades, for easy and intuitive judgment.

[0117] This index-based approach provides a clear reference for subsequent decision-making.

[0118] For example, when conducting multi-dimensional analysis, the growth stability of bitter gourd plants under different humidity and temperature conditions can be simulated. It is assumed that the growth potential decreases by 5% in an environment with 80% humidity, while the plant performs best at a temperature of 25 degrees Celsius.

[0119] This analysis helps identify key environmental factors that influence growth, providing a basis for crop optimization.

[0120] Step S106: Based on the predicted indicators of plant growth vigor and combined with relevant data on yield performance, a comprehensive evaluation method is used to quantitatively analyze the post-grafting yield potential of each rootstock individual and determine the priority ranking of yield performance values; specifically, this includes the following steps: S1061: By obtaining plant growth potential prediction index data and yield performance data from the database, an initial dataset is formed, and the initial data integration is completed. S1062: Based on the initial dataset, the grafting yield potential of each rootstock individual is quantified using a pre-established comprehensive evaluation model to obtain the potential evaluation results. S1063: Based on the potential assessment results, if the potential value of a certain rootstock individual exceeds the preset threshold, it will be marked as a high-potential individual, and the initial screening will be completed. S1064: By conducting in-depth analysis of the yield performance data of high-potential individuals, obtain their correlation coefficients with growth indicators, and identify key influencing factors; S1065: Based on key influencing factors, a random forest model is used to predict yield performance values, and the predicted yield ranking of each rootstock individual is obtained. S1066: For the predicted yield ranking, if a certain rootstock individual ranks at the top, its growth analysis data is further extracted to determine its stability characteristics. S1067: By combining stability characteristics with yield potential data, the final priority ranking result is generated, completing the full-process analysis.

[0121] Specific application example of step S106: For example, when retrieving plant growth potential prediction index data and yield performance data from the database, relevant records of bitter gourd rootstocks from the past two years can be retrieved first, including information such as plant height growth rate, leaf expansion speed, and fruit yield per plant.

[0122] Suppose there are 200 rootstock samples, each containing predicted growth potential and actual yield data, forming a multi-dimensional initial dataset.

[0123] This approach ensures the comprehensiveness of the data, laying the foundation for subsequent integration.

[0124] For example, after the initial data integration is completed, growth potential data and yield data can be matched according to individual rootstocks, and incomplete records can be eliminated.

[0125] Suppose that 10 samples in a batch of data are excluded due to missing records, and ensure that the integrated dataset contains 190 valid samples.

[0126] This integration approach provides reliable data support for subsequent evaluations.

[0127] For example, to quantify the grafting yield potential in a comprehensive evaluation model, growth potential indicators and historical yield performance can be combined and analyzed based on a pre-set weighting system.

[0128] Assuming a rootstock individual has a predicted growth potential of 85 and a historical average yield of 3.5 kg per plant, its potential value calculated by the model is 88, which exceeds the average level.

[0129] This quantification method helps to intuitively reflect an individual's potential.

[0130] For example, when marking high-potential individuals in the potential assessment results, a threshold of 80 can be set. If a rootstock has a potential value of 88, it will be marked as a high-potential individual.

[0131] We further screened out 30 high-potential individuals so that we could concentrate resources on subsequent analysis.

[0132] This screening mechanism can quickly identify high-quality samples.

[0133] For example, when conducting in-depth analysis of yield performance data for high-potential individuals, the correlation coefficient between the data and growth indicators can be calculated.

[0134] If the correlation coefficient between the growth rate of plant height and yield of a certain rootstock is 0.75, it indicates that plant height is a key influencing factor.

[0135] This analytical approach helps to clarify the direction for optimization.

[0136] For example, when using a random forest model to predict yield performance, key factors such as plant height and number of leaves can be selected as input parameters.

[0137] Suppose that 30 high-potential individuals are predicted, and a certain rootstock is predicted to yield 4.2 kg per tree, ranking among the top.

[0138] This prediction method can provide a scientific basis for ranking.

[0139] For example, when extracting growth analysis data from rootstock individuals that rank high in predicted yield to determine stability characteristics, attention can be paid to their performance under different environments.

[0140] If the plant height of a certain rootstock fluctuates by less than 5% under humidity changes, it indicates that its stability is relatively strong.

[0141] This stability analysis provides an important reference for the final decision.

[0142] For example, when generating priority ranking results by combining stability features with yield potential data, both predicted yield and stability performance can be considered.

[0143] If the predicted yield of a certain rootstock is 4.2 kg and the stability fluctuation is less than 5%, then it is ranked first in priority.

[0144] This end-to-end analysis approach ensures the comprehensiveness and reliability of the results.

[0145] Step S107: Prioritize yield performance values, obtain quality stability data, and predict the quality fluctuation range of grafted bitter gourd using data comparison technology to obtain the quality stability assessment result; specifically including the following steps: S1071: By obtaining yield and quality data of grafted bitter gourd from historical records, an initial dataset is constructed to obtain basic information for subsequent analysis; S1072: Based on the initial dataset, prioritize the output performance values, use a preset threshold to filter out the data subset with high performance values, and determine the key analysis objects; S1073: By standardizing the quality data in the high-performance data subset, the distribution characteristics of the quality data are obtained, and its initial fluctuation range is determined. S1074: If the initial fluctuation range exceeds the preset threshold, further data comparison and analysis of the quality data will be performed, and the support vector machine algorithm will be used to accurately predict the fluctuation range to obtain the predicted fluctuation value. S1075: Based on the predicted fluctuation value and combined with the historical quality stability record of grafted bitter gourd, obtain the intermediate index for stability assessment and determine the degree of influence of the fluctuation range on stability. S1076: By comparing intermediate indicators with preset stability evaluation standards, it is determined whether the quality stability of grafted bitter gourd meets expectations, and the final evaluation result is obtained.

[0146] Step S108: Based on the evaluation results of quality stability and combined with relevant data on screening accuracy, information integration technology is used to comprehensively score all candidate rootstocks to determine the final list of preferred rootstocks; specifically, this includes the following steps: S1081: Obtain quality stability data and screening accuracy data of rootstock candidate individuals, and integrate relevant indicators from multiple sources through a pre-established data collection system to form an initial dataset; S1082: For the initial dataset, information integration technology is used to standardize the quality stability data and screening accuracy data to obtain a unified evaluation data matrix; S1083: Based on the evaluation data matrix, each candidate rootstock is comprehensively scored using a preset scoring system. If the score of a candidate is lower than the preset threshold, it is removed from the candidate list to obtain preliminary screening results. S1084: By analyzing the individuals in the preliminary screening results and combining data fusion technology, multiple indicators of the remaining rootstock candidates are weighted and calculated to determine the weighted comprehensive score. S1085: Based on the weighted comprehensive score, sort the remaining rootstock candidates. If the comprehensive score of a candidate is higher than the preset upper limit threshold, mark it as a priority candidate and obtain the sorted candidate list. S1086: For the sorted candidate list, apply the decision tree algorithm to perform final verification on the priority objects, determine whether they meet the requirements of the preferred list, and output the set of the verified preferred individuals. S1087: By comparing the data of the verified preferred individual set, information integration technology is used to generate the final list of preferred rootstocks, and the final result that meets the requirements of quality stability and screening accuracy is determined.

[0147] Step S109: For the final list of preferred rootstocks, data recording technology is used to store the screening process and results in a database. An automatic update mechanism is used to supplement the data from subsequent grafting experiments, obtaining a continuously optimized basis for rootstock selection. Specifically, this includes the following steps: S1091: Obtain relevant information about the rootstock list from the screening process using data acquisition tools, organize it in a structured format, and store it in a pre-established database to obtain the initial rootstock screening dataset. S1092: Based on the initial rootstock screening dataset, classify the screening process information recorded therein, extract key indicators using preset field rules, and determine the core basis data in the screening process. S1093: If the core data meets the preset threshold conditions, the latest experimental data will be obtained from the grafting experiment through an automatic update mechanism and added to the database to determine whether to adjust the existing rootstock list. S1094: If the indicators of some rootstocks change after the experimental data are supplemented, the random forest algorithm is used to predict and analyze the trend of change to obtain the potential optimization direction of the rootstock list. S1095: Based on the optimization direction obtained from the predictive analysis, compare and verify the subsequent experimental data stored in the database, obtain experimental results consistent with the optimization direction, and determine the adjustment plan for the rootstock list. S1096: Update the rootstock list in the database by adjusting the scheme, and in conjunction with the goal of continuous improvement, associate the updated list with the screening criteria to obtain the latest screening reference data; S1097: For the latest screening reference data, use an automated data recording method to save it to the database, forming a closed-loop data supplementation and optimization process to judge the continuous improvement effect of the rootstock screening criteria.

[0148] This invention encompasses any substitutions, modifications, equivalent methods, and solutions made within the spirit and scope of this invention. To provide the public with a thorough understanding of this invention, specific details are described in detail in the following preferred embodiments; however, those skilled in the art will fully understand the invention even without these details. Furthermore, to avoid unnecessary misunderstanding of the essence of this invention, well-known methods, processes, procedures, components, and circuits are not described in detail.

[0149] Those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc.

[0150] The above description is only a preferred embodiment of the present invention. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principle of the present invention, and these improvements and modifications should also be considered within the scope of protection of the present invention.

Claims

1. A method for matching and screening loofah rootstock and bitter gourd scion, characterized in that, Includes the following steps: Step S101: By performing genome sequencing on loofah rootstock samples, genetic difference-related data are obtained, and a genotype characteristic database is constructed for different rootstock individuals to obtain preliminary genetic difference classification results; Step S102: Use data comparison method to analyze the rootstock matching characteristics of loofah rootstock and bitter gourd scion, and perform preliminary screening in combination with preset matching threshold to determine the potential candidate set of rootstocks. Step S103: For each individual in the rootstock candidate set, extract disease resistance genome-related information from the genotype feature database, analyze the distribution of disease resistance genes through information comparison technology, and obtain the expression potential assessment results of the disease resistance genome; Step S104: Based on the evaluation results of the expression potential of the disease-resistant genome and combined with environmental adaptability-related data, the disease resistance value of the rootstock under specific environments is predicted using data analysis methods, and the disease resistance ability ranking of each rootstock individual is determined. Step S105: Based on the disease resistance value ranking results, obtain historical data related to plant growth potential, and simulate the growth potential of grafted bitter gourd plants using data modeling technology to obtain predictive indicators of plant growth potential. Step S106: Based on the predicted indicators of plant growth vigor and combined with relevant data on yield performance, a comprehensive evaluation method is used to quantitatively analyze the post-grafting yield potential of each rootstock individual and determine the priority ranking of yield performance values. Step S107: Prioritize the yield performance values, obtain relevant data on quality stability, predict the quality fluctuation range of grafted bitter gourd through data comparison technology, and obtain the evaluation results of quality stability. Step S108: Based on the evaluation results of quality stability and combined with the relevant data on screening accuracy, use information integration technology to comprehensively score all candidate rootstocks and determine the final list of preferred rootstocks. Step S109: For the final list of preferred rootstocks, data recording technology is used to store the screening process and results in the database. The data of subsequent grafting experiments are supplemented through an automatic update mechanism to obtain continuously optimized rootstock screening criteria.

2. The method for matching and screening loofah rootstock and bitter gourd scion according to claim 1, characterized in that, Step S101 specifically includes the following steps: S1011: By performing genome sequencing on loofah rootstock samples, the raw genome sequence data was obtained, and preliminary data collection was completed; S1012: Based on the obtained genome sequence data, a data extraction process is used to separate key fragments related to genetic differences, resulting in a set of differentially expressed fragments; S1013: For the set of differential fragments, the support vector machine algorithm is applied to perform feature selection, extract genotype features with significant distinguishability, and determine the feature identifier set; S1014: If the data distribution in the feature identifier set meets the preset threshold conditions, then it is entered into the pre-established feature database to complete the storage of genotype features. S1015: By performing cluster analysis on the genotype features in the feature database, the genetic differences of individuals from different rootstocks are grouped to obtain preliminary individual classification results; S1016: Based on the individual classification results and combined with the set of differential fragments, analyze the genetic difference patterns between groups and determine whether there are significant classification boundaries. S1017: If the classification boundaries are clear, the classification results are used to generate structured data for genetic difference analysis to complete the final classification mapping.

3. The method for matching and screening loofah rootstock and bitter gourd scion according to claim 1, characterized in that, Step S102 specifically includes the following steps: S1021: Using genetic difference data, the initial feature sets of loofah rootstock and bitter gourd scion are obtained from the classification results. The data comparison method is used to extract features and obtain the preliminary distribution of matching degree features. S1022: Based on the preliminary distribution of matching degree features, and considering the genetic differences between loofah rootstock and bitter gourd scion, the support vector machine algorithm is used to classify features, determine the correlation weights between features, and obtain the feature combinations after classification. S1023: Obtain the classified feature combination, and perform quantitative analysis of the matching degree feature in combination with the preset threshold. If the matching degree feature value is higher than the preset threshold, it is determined to be a potentially suitable combination, and the preliminary screening of matching pairs is obtained. S1024: For the initially screened matching pairs, the genetic differences between the loofah rootstock and bitter gourd scion are compared in depth using data comparison methods to determine whether there are hidden matching degree features and obtain the matching pair set after in-depth analysis. S1025: Based on the matching pair set after deep analysis, cluster analysis is used to group the potential fit combinations, determine the matching degree feature distribution of rootstock candidates and scions in each group, and obtain the grouped candidate set; S1026: Based on the grouped candidate sets, for the matching degree feature distribution within each group, if the consistency of the feature distribution of a certain group is higher than the preset threshold, it is determined to be a rootstock candidate with high adaptability, and the final adaptation candidate set is obtained. S1027: Obtain the final candidate set of rootstocks, and sort each group of rootstock candidates by combining genetic differences and matching degree features. Determine the stability of the matching degree features in the sorting results and determine the optimal combination of rootstock candidates.

4. The method for matching and screening loofah rootstock and bitter gourd scion according to claim 1, characterized in that, Step S103 specifically includes the following steps: S1031: By obtaining the genotypic characteristic data of rootstock candidates from the database, preliminary information is organized for each individual to obtain the basic dataset of individual genotypes; S1032: Based on the basic dataset, information comparison technology is used to analyze the distribution of disease resistance genes for each individual and determine the location and distribution information of disease resistance genes in the genome; S1033: If the distribution density of disease-resistant genes in the location distribution information is higher than the preset threshold, then the corresponding individuals are marked as high-potential individuals, and the genome information dataset of high-potential individuals is obtained. S1034: Analyze the association between disease resistance genes and genotype characteristics in genomic information datasets of high-potential individuals to determine the expression potential intensity of disease resistance genes; S1035: Based on the analysis results of expression potential intensity and combined with distribution data, the disease resistance genome of each high-potential individual is comprehensively scored to obtain a comprehensive potential assessment value; S1036: Based on the comprehensive potential assessment value, all high-potential individuals are sorted to determine the final priority list of disease resistance gene expression potential; S1037: The random forest algorithm is used to perform validation analysis on individuals in the priority list. Combined with the association data between genotype characteristics and disease resistance genes, the reliability of the evaluation results is judged, and the final disease resistance genome potential analysis dataset is output.

5. The method for matching and screening loofah rootstock and bitter gourd scion according to claim 1, characterized in that, Step S104 specifically includes the following steps: S1041: By obtaining relevant data on disease resistance genes from the database and combining them with the original records of gene expression, a preliminary assessment of expression potential is compiled. S1042: Based on the preliminary assessment results of expression potential, construct a comprehensive impact factor dataset for specific environments by integrating data on environmental adaptation and adaptation information; S1043: The random forest algorithm is used to process the comprehensive impact factor dataset, calculate the predicted values ​​of disease resistance, and obtain the resistance value of each rootstock individual; S1044: For the resistance values ​​of each individual rootstock, perform data standardization processing to obtain quantitative indicators of individual capabilities; S1045: By using quantitative indicators of individual capabilities and combining data analysis methods, rank all rootstock individuals by capability and determine the ranking results; S1046: If there are rootstock individuals with similar resistance values ​​in the ranking results, the final individual ability ranking is determined by comparing environmental adaptation data under a specific environment. S1047: After obtaining the final individual ability ranking, generate the corresponding disease resistance distribution map to determine the disease resistance location of each rootstock individual in a specific environment.

6. The method for matching and screening loofah rootstock and bitter gourd scion according to claim 1, characterized in that, Step S105 specifically includes the following steps: S1051: Construct an initial dataset by obtaining relevant data from the disease resistance ranking results; S1052: Data cleaning techniques are used to preprocess the sorting results, removing outliers and missing values ​​to obtain a cleaned resistant dataset. S1053: Based on the organized resistance dataset, extract historical data related to plant growth; S1054: Perform feature selection on historical data, filter out feature variables that are highly correlated with growth potential, and determine the feature dataset; S1055: Input data for growth simulation is constructed by combining feature datasets with grafting treatment information of bitter gourd plants; S1056: If the values ​​of some variables in the feature dataset are lower than the preset threshold, they are standardized to obtain normalized simulated input data; S1057: A random forest model is used to train the normalized simulated input data, and the growth potential of bitter gourd plants is simulated and calculated to obtain preliminary growth potential predictions. S1058: Based on the preliminary growth potential predictions and combined with the plant growth trends in historical data, the prediction results are corrected. S1059: If the deviation between the predicted value and the historical trend exceeds the preset range, the predicted value is smoothed to obtain the corrected growth potential value. S10510: A predictive index of plant growth potential is generated using the corrected growth potential value. S10511: Perform multi-dimensional analysis on the prediction indicators to determine their stability under different environmental conditions and obtain the final prediction results.

7. The method for matching and screening loofah rootstock and bitter gourd scion according to claim 1, characterized in that, Step S106 specifically includes the following steps: S1061: By obtaining plant growth potential prediction index data and yield performance data from the database, an initial dataset is formed, and the initial data integration is completed. S1062: Based on the initial dataset, the grafting yield potential of each rootstock individual is quantified using a pre-established comprehensive evaluation model to obtain the potential evaluation results. S1063: Based on the potential assessment results, if the potential value of a certain rootstock individual exceeds the preset threshold, it will be marked as a high-potential individual, and the initial screening will be completed. S1064: By conducting in-depth analysis of the yield performance data of high-potential individuals, obtain their correlation coefficients with growth indicators, and identify key influencing factors; S1065: Based on key influencing factors, a random forest model is used to predict yield performance values, and the predicted yield ranking of each rootstock individual is obtained. S1066: For the predicted yield ranking, if a certain rootstock individual ranks at the top, its growth analysis data is further extracted to determine its stability characteristics. S1067: By combining stability characteristics with yield potential data, the final priority ranking result is generated, completing the full-process analysis.

8. The method for matching and screening loofah rootstock and bitter gourd scion according to claim 1, characterized in that, Step S107 specifically includes the following steps: S1071: By obtaining yield and quality data of grafted bitter gourd from historical records, an initial dataset is constructed to obtain basic information for subsequent analysis; S1072: Based on the initial dataset, prioritize the output performance values, use a preset threshold to filter out the data subset with high performance values, and determine the key analysis objects; S1073: By standardizing the quality data in the high-performance data subset, the distribution characteristics of the quality data are obtained, and its initial fluctuation range is determined. S1074: If the initial fluctuation range exceeds the preset threshold, further data comparison and analysis of the quality data will be performed, and the support vector machine algorithm will be used to accurately predict the fluctuation range to obtain the predicted fluctuation value. S1075: Based on the predicted fluctuation value and combined with the historical quality stability record of grafted bitter gourd, obtain the intermediate index for stability assessment and determine the degree of influence of the fluctuation range on stability. S1076: By comparing intermediate indicators with preset stability evaluation standards, it is determined whether the quality stability of grafted bitter gourd meets expectations, and the final evaluation result is obtained.

9. The method for matching and screening loofah rootstock and bitter gourd scion according to claim 1, characterized in that, Step S108 specifically includes the following steps: S1081: Obtain quality stability data and screening accuracy data of rootstock candidate individuals, and integrate relevant indicators from multiple sources through a pre-established data collection system to form an initial dataset; S1082: For the initial dataset, information integration technology is used to standardize the quality stability data and screening accuracy data to obtain a unified evaluation data matrix; S1083: Based on the evaluation data matrix, each candidate rootstock is comprehensively scored using a preset scoring system. If the score of a candidate is lower than the preset threshold, it is removed from the candidate list to obtain preliminary screening results. S1084: By analyzing the individuals in the preliminary screening results and combining data fusion technology, multiple indicators of the remaining rootstock candidates are weighted and calculated to determine the weighted comprehensive score. S1085: Based on the weighted comprehensive score, sort the remaining rootstock candidates. If the comprehensive score of a candidate is higher than the preset upper limit threshold, mark it as a priority candidate and obtain the sorted candidate list. S1086: For the sorted candidate list, apply the decision tree algorithm to perform final verification on the priority objects, determine whether they meet the requirements of the preferred list, and output the set of the verified preferred individuals. S1087: By comparing the data of the verified preferred individual set, information integration technology is used to generate the final list of preferred rootstocks, and the final result that meets the requirements of quality stability and screening accuracy is determined.

10. The method for matching and screening loofah rootstock and bitter gourd scion according to claim 1, characterized in that, Step S109 specifically includes the following steps: S1091: Obtain relevant information about the rootstock list from the screening process using data acquisition tools, organize it in a structured format, and store it in a pre-established database to obtain the initial rootstock screening dataset. S1092: Based on the initial rootstock screening dataset, classify the screening process information recorded therein, extract key indicators using preset field rules, and determine the core basis data in the screening process. S1093: If the core data meets the preset threshold conditions, the latest experimental data will be obtained from the grafting experiment through an automatic update mechanism and added to the database to determine whether to adjust the existing rootstock list. S1094: If the indicators of some rootstocks change after the experimental data are supplemented, the random forest algorithm is used to predict and analyze the trend of change to obtain the potential optimization direction of the rootstock list. S1095: Based on the optimization direction obtained from the predictive analysis, compare and verify the subsequent experimental data stored in the database, obtain experimental results consistent with the optimization direction, and determine the adjustment plan for the rootstock list. S1096: Update the rootstock list in the database by adjusting the scheme, and in conjunction with the goal of continuous improvement, associate the updated list with the screening criteria to obtain the latest screening reference data; S1097: For the latest screening reference data, use an automated data recording method to save it to the database, forming a closed-loop data supplementation and optimization process to judge the continuous improvement effect of the rootstock screening criteria.

Citation Information

Patent Citations

  • Multi-variety fusion grafting control method of China rose, server, medium and product

    CN120493153A

  • Engrafted plants resistant to viral diseases and methods of producing same

    WO2005079162A2