Disease-resistant gene typing result analysis method and system
By converting multi-source heterogeneous microphenotype data into a unified set of quantitative indicators and performing data dimensionality reduction, identifying resistance components, and combining genotyping and macrophenotype data for correlation analysis, the problem of difficulty in integrating microscopic data in the existing technology is solved, and precise breeding in cucumber disease-resistant breeding is achieved.
Patent Information
- Application Number
- CN202510845183.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-23
AI Technical Summary
The existing technology is difficult to effectively integrate multi-source heterogeneous microphenotype data, and cannot reveal the resistance mechanism in cucumber disease-resistant breeding, resulting in breeding decision-making relying on macro results and being unable to achieve precise breeding.
By converting multi-source heterogeneous microphenotype data into a unified set of quantitative indicators, multiple resistance components are identified through data dimensionality reduction processing, resistance component feature vectors are generated, and correlation analysis is performed by combining genotyping and macroscopic disease-resistant phenotype data.
A deep understanding of the pathogenesis mechanism is achieved, precise breeding based on resistance mechanisms is supported, and breeding efficiency and accuracy are improved.
Smart Images

Figure CN120356529A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of plant disease-resistant breeding analysis, and particularly to a method and system for analyzing the results of disease-resistant gene typing. Background Art
[0002] In the field of modern plant breeding, especially in disease-resistant breeding projects of important cash crops such as cucumbers, breeding units widely rely on molecular marker-assisted selection technology to improve breeding efficiency. The conventional technical process usually involves the following steps: First, leaf samples are collected from germplasm resources or genetic populations, genomic DNA is extracted, and molecular marker data of the whole genome or target regions are obtained using high-throughput genotyping technologies (such as gene chips or sequencing). At the same time, in order to evaluate the disease-resistant performance of these cucumber materials, breeding technicians will record the macroscopic disease-resistant phenotype data of the materials, such as incidence rate, disease index, lesion characteristics, and yield loss assessment, by artificially inoculating or naturally inducing the target disease in a greenhouse environment strictly controlled by humans or at field test sites with high disease incidence. Existing analysis systems or methods mainly integrate genotyping data and macroscopic phenotype data to perform genotype-phenotype association analysis and locate candidate gene intervals or molecular markers significantly related to disease-resistant traits for material screening.
[0003] However, breeders are no longer satisfied with merely locating disease-resistant genes and instead hope to understand their action mechanisms. Therefore, microscopic phenotype data is introduced, that is, the dynamic response data of plant cells, tissues, and physiological and biochemical processes during pathogen infection, specifically including: image information such as the invasion sites of pathogens in leaf tissues, hyphal expansion patterns, and ultrastructural changes in host cells observed and recorded through equipment such as confocal microscopes; time series data of a series of key defense indicators measured through physiological and biochemical analysis, such as the concentration change curves of defense signal molecules, the burst levels of reactive oxygen species, and the activity change spectra of defense enzymes related to disease resistance; and high-dimensional concentration value data of disease-resistant related secondary metabolites detected using metabolomics technology, such as the dynamic accumulation patterns of flavonoids and phenols.
[0004] In the practice of cucumber disease-resistant breeding, when breeders obtain genotyping data as the genetic basis, macroscopic disease-resistant phenotype data as the final evaluation basis, and multi-source heterogeneous microscopic phenotype data for revealing the disease-resistant mechanism, the core technical problem faced by existing analysis methods is: how to design a data processing method that does not rely on a preset biological pathway model to effectively integrate and quantify these microscopic phenotype data with different structures and dimensions, and associate them with genotype data, so as to not only evaluate the final disease-resistant result, but also reveal the internal resistance action mode determined by the genotype, hidden under the macroscopic appearance and characterized by different combinations of microscopic phenotypes, in order to solve the dilemma that the current system cannot analyze microscopic data, resulting in unclear resistance mechanisms, so that breeding decisions can only rely on macroscopic results and cannot achieve precise breeding based on resistance mechanisms.
[0005] In view of the above problems, there is an urgent need for improvement in the existing technology. Summary of the Invention
[0006] The purpose of the present invention is to solve the deficiencies in the existing technology and propose a method and system for analyzing disease-resistant gene typing results.
[0007] In the first aspect, the present invention provides a method for analyzing disease-resistant gene typing results, which is applied to the cucumber disease-resistant breeding analysis environment. The method is characterized by comprising the following steps:
[0008] Obtain multi-source heterogeneous microscopic phenotype data for cucumber disease-resistant breeding analysis, and convert the multi-source heterogeneous microscopic phenotype data into a unified set of quantitative indicators;
[0009] Based on the unified set of quantitative indicators, through data dimensionality reduction processing, identify multiple resistance components, each of the resistance components characterizing a resistance action mode defined by a combination of microscopic phenotype data, and generate a resistance component feature vector for each cucumber material, which includes its quantitative scores on each of the resistance components;
[0010] Obtain the genotyping data of the cucumber material, and perform a first association analysis based on the genotyping data and the resistance component feature vector to identify gene information related to specific resistance components;
[0011] Obtain the macroscopic disease-resistant phenotype data of the cucumber material, and perform a second association analysis based on the resistance component feature vector and the macroscopic disease-resistant phenotype data to evaluate the effects of different resistance components on the macroscopic disease-resistant phenotype.
[0012] The core innovation of this application lies in converting multi-source heterogeneous micro-phenotypic data into a unified set of quantitative indicators, and based on this, identifying multiple resistance components through data dimensionality reduction processing, thereby abstracting complex micro-resistance responses into quantifiable component feature vectors. Furthermore, associative analysis is performed on this feature vector with genotyping data and macro-disease resistance phenotypic data, briefly describing how the problem that existing technologies are difficult to effectively integrate micro-data and reveal resistance mechanisms is solved, achieving the effects of deeply understanding the disease resistance mechanism and realizing mechanism-based precision breeding.
[0013] In a second aspect, a system for analyzing disease resistance genotyping results for cucumber disease resistance breeding is provided. The system includes:
[0014] A data acquisition and conversion module for acquiring multi-source heterogeneous micro-phenotypic data for cucumber disease resistance breeding and converting the multi-source heterogeneous micro-phenotypic data into a unified set of quantitative indicators;
[0015] A resistance component identification and feature vector generation module for identifying multiple resistance components based on the unified set of quantitative indicators through data dimensionality reduction processing. Each resistance component represents a resistance action mode defined by a combination of micro-phenotypic data, and a resistance component feature vector containing the quantitative scores of each cucumber material on each of the resistance components is generated for each cucumber material;
[0016] A gene association analysis module for acquiring the genotyping data of the cucumber material and performing a first association analysis based on the genotyping data and the resistance component feature vector to identify gene information related to specific resistance components;
[0017] A macro-phenotype association analysis module for acquiring the macro-disease resistance phenotypic data of the cucumber material and performing a second association analysis based on the resistance component feature vector and the macro-disease resistance phenotypic data to evaluate the effects of different resistance components on the macro-disease resistance phenotype.
[0018] Compared with the prior art, the present invention has the following beneficial effects:
[0019] By integrating multi-source micro-phenotypic data, identifying internal resistance components, and performing association analysis with genotypes and macro-phenotypes, the problems of the prior art being unable to analyze micro-data and having unclear resistance mechanisms are solved. It has the advantages of being able to effectively integrate multi-source heterogeneous micro-phenotypic data, revealing the internal resistance action mode determined by genotypes, and providing a basis for precision breeding based on resistance mechanisms. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 It is a flowchart of the method of the present invention.
[0021] Figure 2 It is a structural diagram of the system of the present invention.
[0022] In the figure: 201, data acquisition and conversion module; 202, resistance component identification and eigenvector generation module; 203, gene association analysis module; 204, macroscopic phenotype association analysis module. Specific implementation manners
[0023] The following details the implementation manners of the present invention. Examples of the implementation manners are shown in the accompanying drawings, where the same or similar reference numerals represent the same or similar elements or elements with the same or similar functions throughout. The implementation manners described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as a limitation to the present invention.
[0024] The terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the said features. In the description of the present invention, "a plurality of" means two or more unless otherwise specifically defined.
[0025] As Figure 1 shown, a method for analyzing the genotyping results of disease-resistant genes is applied to the cucumber disease-resistant breeding analysis environment. The method includes the following steps:
[0026] Obtain multi-source heterogeneous microscopic phenotype data for cucumber disease-resistant breeding analysis, and convert the multi-source heterogeneous microscopic phenotype data into a unified set of quantitative indicators;
[0027] Based on the unified set of quantitative indicators, through data dimensionality reduction processing, identify a plurality of resistance components, each resistance component characterizing a resistance action pattern defined by a combination of microscopic phenotype data, and generate a resistance component eigenvector for each cucumber material, which includes its quantitative scores on each resistance component;
[0028] Obtain the genotyping data of cucumber materials, and perform a first association analysis based on the genotyping data and the resistance component eigenvectors to identify gene information related to specific resistance components;
[0029] Obtain the macroscopic disease-resistant phenotype data of cucumber materials, and perform a second association analysis based on the resistance component eigenvectors and the macroscopic disease-resistant phenotype data to evaluate the effects of different resistance components on the macroscopic disease-resistant phenotype.
[0030] Among them, multi-source heterogeneous microscopic phenotypic data refers to data from different detection technologies, with different data structures and dimensions, which characterize the response of cucumber plants to pathogens at the microscopic level such as cells, tissues, physiology and biochemistry. It can be obtained in various forms such as microscopic image data, physiological and biochemical indicator time series data, metabolomics data, etc., and is mainly used to reveal the intrinsic biological mechanism of cucumber disease resistance. A unified set of quantitative indicators refers to converting the above multi-source heterogeneous microscopic phenotypic data into a data set with a unified format and dimension after preprocessing, standardization, feature extraction and other steps, such as a matrix, in which rows represent different cucumber materials and columns represent quantified microscopic phenotypic indicators. It is mainly to eliminate the differences between different data sources and facilitate subsequent integrated analysis. Data dimensionality reduction processing refers to the use of mathematical or statistical methods, such as principal component analysis (PCA), factor analysis (FA), independent component analysis (ICA) or other nonlinear dimensionality reduction techniques, to map a high-dimensional set of quantitative indicators to a low-dimensional space, while retaining important information in the original data as much as possible. It is mainly to simplify the data structure and extract potential key features or patterns. Resistance components refer to potential factors or dimensions that characterize a specific resistance action mode identified from microscopic phenotypic data through data dimensionality reduction. Each component represents a specific combination or weight of multiple microscopic phenotypic indicators, which is mainly used to summarize and abstract complex microscopic resistance response processes. Resistance action mode refers to a biological function or mechanism defined by a specific combination of microscopic phenotypic indicators, such as a defense mode composed of cell wall thickening and reactive oxygen burst, which is mainly used to describe the specific way cucumber plants resist pathogen infection. Resistance component feature vector refers to a vector generated for each cucumber material, containing its quantitative scores on each identified resistance component. This vector is a representation of the microscopic resistance characteristics of the material in a low-dimensional space, which is mainly used to quantify the performance differences of different cucumber materials in various resistance action modes. The first association analysis refers to the statistical association analysis of the genotyping data of the cucumber material with the resistance component feature vector, such as genome-wide association analysis (GWAS) or quantitative trait locus (QTL) positioning analysis, which is mainly used to identify genomic regions or genes significantly associated with specific resistance components. The second association analysis refers to the statistical association analysis of the characteristic vector of the resistance component and the macroscopic disease resistance phenotype data of the cucumber material, such as regression analysis or correlation analysis, which is mainly used to evaluate the contribution of different resistance components to the final macroscopic disease resistance performance.
[0031] The solution of this application realizes the standardization and integration of microscopic data from different sources and in different formats by obtaining multi-source heterogeneous microscopic phenotype data for cucumber disease-resistant breeding analysis and converting it into a unified set of quantitative indicators. Based on this unified set of quantitative indicators, dimensionality reduction processing is performed. This process can extract the potential structure reflecting the internal resistance mechanism, that is, multiple resistance components, from the high-dimensional microscopic phenotype data. Each component represents a resistance action mode defined by a specific combination of microscopic phenotypes. Subsequently, for each cucumber material, its quantitative scores on these resistance components are calculated to generate a resistance component feature vector, which is a refined representation of the microscopic resistance characteristics of the material. Further, the genotyping data of the cucumber material is obtained, and the first association analysis is performed based on the genotyping data and the resistance component feature vector, so as to identify the gene information related to a specific resistance component (i.e., a specific resistance action mode) and reveal the genetic basis of the resistance mechanism. At the same time, the macroscopic disease-resistant phenotype data of the cucumber material is obtained, and the second association analysis is performed based on the resistance component feature vector and the macroscopic disease-resistant phenotype data, so as to evaluate the influence and contribution degree of different resistance components (i.e., different resistance action modes) on the final macroscopic disease-resistant performance. The entire process forms a complete analysis chain from microscopic mechanism to genotype to macroscopic phenotype, enabling breeders to go beyond simple gene-macroscopic phenotype association and deeply understand how genes determine the final disease resistance by affecting specific microscopic resistance action modes.
[0032] As an implementation manner of the present invention, based on the unified set of quantitative indicators, through dimensionality reduction processing, the steps of identifying multiple resistance components, each resistance component characterizing a resistance action mode defined by a combination of microscopic phenotype data, and generating a resistance component feature vector containing the quantitative scores of each cucumber material on each resistance component include:
[0033] Perform dimensionality reduction processing on the unified set of quantitative indicators of at least one initial cucumber material batch to identify multiple resistance components, each resistance component characterizing a resistance action mode defined by a combination of microscopic phenotype data, and based on this, determine a fixed conversion rule for generating a resistance component feature vector from the unified set of quantitative indicators;
[0034] Apply the fixed conversion rule to the unified set of quantitative indicators of at least one initial cucumber material batch and subsequent obtained cucumber material batches to generate a resistance component feature vector containing the quantitative scores of each cucumber material on each resistance component.
[0035] Among them, data dimensionality reduction processing refers to mapping high-dimensional data (such as a unified set of quantitative indicators) to a low-dimensional space through mathematical or statistical methods, while retaining important information or structure in the original data as much as possible. It can be achieved by using principal component analysis (PCA), independent component analysis (ICA), factor analysis (FA) and other techniques. Its purpose is to extract representative, independent or low-correlated potential resistance action modes from complex microscopic phenotypic data, reduce data dimensions, and facilitate subsequent analysis; resistance components refer to potential factors or principal components identified from microscopic phenotypic data through data dimensionality reduction processing. Each component represents a specific resistance action mode contributed by multiple microscopic phenotypes. Its purpose is to abstract complex microscopic resistance performance into a few quantifiable dimensions; fixed transformation rules refer to mathematical models or parameter sets determined based on the initial data batch for linearly or nonlinearly mapping a unified set of quantitative indicators to the characteristic vectors of resistance components. It can be a transformation matrix (such as the PCA loading matrix). Its purpose is to provide a stable and unified transformation standard for subsequent batches of data analysis.
[0036] The scheme of the present application first performs data dimensionality reduction processing based on a unified set of quantitative indicators of at least one initial batch of cucumber materials, thereby identifying multiple resistance components, and based on this, determining a fixed conversion rule for generating a resistance component feature vector from a unified set of quantitative indicators. This process uses a representative initial data batch to establish a stable analysis benchmark, extracts the core resistance action mode from high-dimensional microscopic phenotypic data through data dimensionality reduction, and solidifies this mode into a conversion rule. Subsequently, the fixed conversion rule is applied to a unified set of quantitative indicators of at least one initial batch of cucumber materials and a subsequent batch of cucumber materials to generate a resistance component feature vector containing its quantitative scores on each resistance component for each cucumber material. This means that whether it is the initial data used to establish the rules or the new data that is continuously added later, the same set of fixed rules are used for conversion. This method avoids the time-consuming data dimensionality reduction calculation every time new data is added, and significantly improves the processing efficiency. At the same time, since all batches of data are mapped to the same resistance component space defined by the initial data, the scores of different batches of cucumber materials on each resistance component are directly comparable, ensuring the reliability of the subsequent association analysis results. This strategy of separating resistance component identification from rule determination and applying fixed rules to all data makes the entire cucumber disease resistance breeding analysis method more efficient, stable and consistent when processing large-scale, batch data, thereby better supporting resistance mechanism analysis and precision breeding based on microscopic phenotypic data.
[0037] As an implementation manner of the present invention, the step of performing data dimensionality reduction processing based on a unified set of quantitative indicators of at least one initial cucumber material batch to identify multiple resistance components, where each resistance component characterizes a resistance action mode defined by a combination of microscopic phenotype data, includes:
[0038] Obtain reference information characterizing the overall characteristics of the cucumber population, where the reference information includes at least one of the germplasm resource diversity distribution information of the cucumber population and the expected variation characteristics of key microscopic phenotypes;
[0039] Compare the unified set of quantitative indicators of at least one initial cucumber material batch with the reference information to identify areas in the unified set of quantitative indicators that do not adequately represent the cucumber population;
[0040] For the areas that do not adequately represent the cucumber population, perform a complementary adjustment operation on the unified set of quantitative indicators to generate an adjusted set of quantitative indicators. The complementary adjustment operation enables the adjusted set of quantitative indicators to prompt the identified resistance components to reflect a more comprehensive resistance action mode of the cucumber population when used for subsequent data dimensionality reduction processing;
[0041] Perform data dimensionality reduction processing based on the adjusted set of quantitative indicators to identify multiple resistance components, where each resistance component characterizes a resistance action mode defined by a combination of microscopic phenotype data.
[0042] Among them, the reference information representing the overall characteristics of the cucumber population refers to data or models that can reflect the macroscopic characteristics of the cucumber population in terms of genetic background, phenotypic variation, etc., and its purpose is to provide a benchmark for evaluating the representativeness of the initial set of quantitative indicators; the information on the distribution of germplasm resources diversity of the cucumber population refers to the composition ratio or distribution pattern of the cucumber population in different geographical origins, genetic lineages, variety types, etc. Specifically, it can be the results of population structure analysis based on molecular marker data, the composition list of the core germplasm bank, the list of representative materials of different ecological types, etc. Its purpose is to evaluate the coverage of the initial batch of materials in terms of genetic diversity; the expected variation characteristics of key microscopic phenotypes refer to the variation range, distribution pattern or mutual relationship that the microscopic phenotype indicators closely related to disease resistance should have in the healthy or typical disease-susceptible / resistant state. Specifically, it can be the normal fluctuation range of each microscopic phenotype indicator, the ideal distribution curve, or the expected correlation pattern between different indicators established based on historical data, biological knowledge or expert experience. Its purpose is to evaluate the integrity and typicality of the initial batch of materials in terms of key phenotypic data; the area that fails to fully represent the cucumber population refers to the part of the initial set of quantitative indicators where sparse, missing, abnormally distributed data or failure to cover a specific germplasm resource / phenotypic variation range is found after comparison with the reference information. Specifically, it can be that there are too few data points for some microscopic phenotype indicators, insufficient material samples for some germplasm resources, or some key phenotypic combinations not reflected in the data, etc. Its purpose is to clarify the target range that needs to be adjusted complementarily; the complementary adjustment operation refers to the data processing process performed on the initial set of quantitative indicators, aiming to improve its representativeness of the overall characteristics of the cucumber population. Specifically, it can be achieved through data imputation, data weighting, introducing simulated data or selectively increasing samples, etc. Its purpose is to generate a more representative set of quantitative indicators to optimize the subsequent dimensionality reduction results; the adjusted set of quantitative indicators refers to the set of quantitative indicators obtained after the complementary adjustment operation. Compared with the initial set, this set has been improved in terms of data coverage, distribution characteristics or germplasm resource / phenotypic representativeness. Its role is to serve as the input for data dimensionality reduction processing to promote the identified resistance components to reflect a more comprehensive resistance action pattern of the cucumber population.
[0043] The solution of this application first obtains reference information representing the overall characteristics of the cucumber population as an evaluation benchmark, and then compares the unified set of quantitative indicators of at least one initial cucumber material batch with this reference information, thereby identifying the areas in the unified set of quantitative indicators that do not fully represent the cucumber population. Then, for these areas that do not fully represent the cucumber population, a supplementary adjustment operation is performed on the unified set of quantitative indicators to generate an adjusted set of quantitative indicators. It is precisely due to this targeted supplementary adjustment that the adjusted set of quantitative indicators has been significantly improved in terms of data coverage and representativeness. Finally, based on this more representative adjusted set of quantitative indicators, dimensionality reduction processing is performed, so that multiple resistance components that are more comprehensive and can better reflect the overall resistance action mode of the cucumber population can be identified. This strategy of introducing reference information and evaluating and adjusting the initial data effectively overcomes the problem of incomplete identification of resistance components that may occur when directly using limited initial data for dimensionality reduction, ensuring that subsequent analyses based on these resistance components (such as association analysis with genotypes or macroscopic phenotypes) can obtain more reliable and comprehensive results, and improving the effectiveness of the entire breeding analysis method.
[0044] As an implementation manner of the present invention, the step of performing a supplementary adjustment operation on the unified set of quantitative indicators to generate an adjusted set of quantitative indicators for the areas that do not fully represent the cucumber population includes:
[0045] Determine the initial covariance relationship between the quantitative indicators in the unified set of quantitative indicators;
[0046] During the process of performing a supplementary adjustment operation on the unified set of quantitative indicators for the areas that do not fully represent the cucumber population to change the data composition or values of the unified set of quantitative indicators, evaluate the changes in the initial covariance relationship caused by the supplementary adjustment operation;
[0047] When the change in the covariance relationship between the quantitative indicators caused by the supplementary adjustment operation exceeds the preset condition relative to the initial covariance relationship, adjust the execution method or execution amplitude of the supplementary adjustment operation so that the covariance relationship in the adjusted set of quantitative indicators generated based on the adjusted supplementary adjustment operation meets the preset condition.
[0048] Among them, determining the initial covariance relationship between the quantitative indicators in the unified set of quantitative indicators refers to the statistical association pattern existing between different quantitative indicators in the set before performing the supplementary adjustment operation on the unified set of quantitative indicators. For example, their correlation, variance, or covariance structure, which can be determined by calculating the correlation coefficient matrix, covariance matrix, or performing principal component analysis, etc. The purpose is to establish a benchmark for subsequent evaluation of the impact of data adjustment on the relationship between indicators; evaluating the changes in the initial covariance relationship caused by the supplementary adjustment operation means monitoring and quantifying the degree of difference between the covariance relationship between the quantitative indicators in the adjusted set of quantitative indicators and the initial covariance relationship during the execution of the supplementary adjustment operation. It can be achieved by comparing the differences in the covariance matrices before and after adjustment, calculating the change in the correlation coefficient between specific indicator pairs, or using statistical distance metrics, etc. The purpose is to understand the impact of data adjustment on the original data structure in real-time or regularly; the preset condition refers to the threshold or standard for judging whether the impact of the supplementary adjustment operation on the covariance relationship between quantitative indicators is acceptable. It can be set as the maximum allowable amplitude of the change in the covariance relationship, the allowable change range of the correlation coefficient of specific key indicator pairs, or the significance level based on statistical tests, etc. The purpose is to control the degree of data adjustment and avoid excessive damage to the original data structure; adjusting the execution method or execution amplitude of the supplementary adjustment operation means modifying the ongoing or subsequent supplementary adjustment operation according to the evaluation results. For example, the amount of data increase or modification can be reduced, the weighting coefficient can be changed, or a different adjustment algorithm can be adopted. The purpose is to correct the adjustment operation so that the covariance relationship of the adjusted data meets the preset conditions.
[0049] The solution of this application establishes a reference benchmark before data adjustment by determining the initial covariance relationship between the quantitative indicators in the unified set of quantitative indicators. Subsequently, during the process of performing a supplementary adjustment operation on the area that fails to fully represent the cucumber population to change the data composition or values, the changes caused by this operation to the initial covariance relationship are continuously evaluated. It is precisely because the monitoring and evaluation of the changes in the covariance relationship are introduced during the adjustment process that when the changes exceed the preset conditions, the execution method or execution amplitude of the supplementary adjustment operation can be adjusted in a timely manner. This feedback control mechanism ensures that while enhancing the data representativeness, the supplementary adjustment maximally retains the original covariance structure in the data that reflects biological associations. This combination enables the resistance components identified based on the adjusted data to not only reflect a more comprehensive resistance action pattern of the cucumber population but also ensure that these patterns are based on relatively real data associations, thereby improving the accuracy and reliability of subsequent breeding analysis.
[0050] As an implementation manner of the present invention, the steps for determining the initial covariance relationship between the quantitative indicators in the unified set of quantitative indicators include:
[0051] Obtain the growth stage information or environmental condition information corresponding to each cucumber material in the unified set of quantitative indicators;
[0052] According to the growth stage information or environmental condition information, divide the unified set of quantitative indicators into multiple data subsets, and each data subset corresponds to a specific growth stage or environmental condition;
[0053] For each data subset, calculate the conditional specific covariance relationship among the quantitative indicators within the data subset;
[0054] Determine a reference benchmark, which is used to characterize the covariance characteristics of the cucumber population under the reference conditions;
[0055] Combine the degree of association between the conditional specific covariance relationship of each data subset and the reference benchmark, and the data proportion of each data subset in the unified set of quantitative indicators, and perform an integration process on each conditional specific covariance relationship to generate the initial covariance relationship among the quantitative indicators in the unified set of quantitative indicators.
[0056] Among them, the growth stage information or environmental condition information refers to the developmental stage or external environmental conditions experienced by cucumber materials when collecting microscopic phenotypic data, which can be obtained through manual recording, sensor monitoring, image analysis, etc. The purpose is to group cucumber materials for analyzing phenotypic responses under different conditions. The data subset refers to the data grouping selected from a unified set of quantitative indicators according to specific growth stages or environmental conditions. Each data subset contains the quantitative indicator data of all cucumber materials collected under that specific condition. The purpose is to decompose the overall data into more homogeneous parts for condition-specific analysis. The condition-specific covariation relationship refers to the degree and pattern of mutual association or co-variation among quantitative indicators in a unified set of quantitative indicators under a specific growth stage or environmental condition, which can be characterized by calculating the covariance matrix, correlation coefficient matrix, or partial correlation network of the data subset. The purpose is to reveal the internal relationship between indicators under specific conditions. The reference benchmark refers to a pre-determined or extracted from representative data reference pattern for measuring and comparing covariation relationships under different conditions. It can be determined based on a preset model or representative sample. The purpose is to provide an objective comparison standard for evaluating the deviation degree of the condition-specific covariation relationship from the reference pattern. The reference condition refers to the specific growth stage and environmental conditions used to define the reference benchmark, which can be the standard conditions for normal growth and development of cucumbers or typical conditions with known resistance performance. The degree of association refers to the similarity or difference measure between the condition-specific covariation relationship of each data subset and the reference benchmark, which can be evaluated by calculating matrix distance, similarity index, or statistical test. The purpose is to quantify the closeness of the covariation pattern under specific conditions to the reference pattern. The data proportion refers to the proportion of the number of cucumber materials included in each data subset to the total number of cucumber materials in the unified set of quantitative indicators. The purpose is to reflect the representativeness or importance of the data subset in the overall data. The integration process refers to comprehensively considering the degree of association between multiple condition-specific covariation relationships and the reference benchmark and the data proportion of the corresponding data subsets, and fusing them into a single initial covariation relationship representing the overall population characteristics through a certain algorithm or model. Methods such as weighted average, Bayesian integration, and graph model fusion can be used. The purpose is to generate a more comprehensive and robust initial covariation relationship. The initial covariation relationship refers to the overall covariation pattern of quantitative indicators in a unified set of quantitative indicators at the level of the entire cucumber population obtained after integration processing. It serves as the input or reference for subsequent data processing. The purpose is to provide a relatively stable and accurate information on the inter-index association structure.
[0057] The solution of the present application divides the unified set of quantitative indicators into multiple data subsets by obtaining the growth stage information or environmental condition information of cucumber materials, so as to be able to analyze the data under different conditions. The condition-specific covariation relationships are calculated for each data subset, revealing the association patterns among the quantitative indicators at specific growth stages or under specific environmental conditions. By determining a reference benchmark, a unified standard for evaluating the covariation relationships under different conditions is provided. Finally, the condition-specific covariation relationships of each data subset are integrated by combining the degree of association between the condition-specific covariation relationship of each data subset and the reference benchmark, and the data proportion of each data subset in the unified set of quantitative indicators. This integration method comprehensively considers the specific covariation patterns under different conditions and their representativeness in the overall population, and refers to the covariation characteristics of the population under standard conditions, thus overcoming the inaccuracies caused by simply calculating the overall covariation relationship or ignoring the condition differences. The initial covariation relationship obtained thereby can more accurately reflect the overall covariation characteristics of the cucumber population at different growth stages and under different environmental conditions, providing a more reliable reference for performing complementary adjustment operations on regions that do not fully represent the cucumber population, so that when the adjusted set of quantitative indicators is used for data dimensionality reduction processing, the identified resistance components can reflect a more comprehensive resistance action pattern of the cucumber population.
[0058] As an implementation manner of the present invention, the step of determining a reference benchmark for characterizing the covariation characteristics of the cucumber population under reference conditions is achieved by one of the following methods:
[0059] Construct a reference benchmark according to a preset expected covariation pattern between indicators of cucumbers under reference conditions;
[0060] Alternatively, select one or more groups of representative cucumber materials that exhibit specific performance under reference conditions, obtain the unified set of quantitative indicators corresponding to the representative cucumber materials, analyze and determine the covariation characteristics among the common quantitative indicators based on the obtained unified set of quantitative indicators corresponding to the representative cucumber materials, and use the determined common covariation characteristics as the reference benchmark.
[0061] Among them, the reference benchmark refers to a set of covariant features used to provide a stable reference point, which can be represented by a covariance matrix, a set of correlation coefficients, or a statistical model, and its purpose is to provide a standard for subsequent covariant relationship evaluation and adjustment; characterizing the expected covariant pattern between indicators of cucumbers under reference conditions refers to the description of the correlation relationship between various quantitative indicators of cucumbers under specific reference conditions based on prior knowledge or theoretical models, which can be defined based on biological pathway models or expert experience, and its purpose is to capture the ideal or theoretical indicator correlations; constructing the reference benchmark refers to the process of establishing the reference benchmark according to the preset expected covariant pattern or extracting information from representative material data, which can include parameter setting or statistical calculation, and its purpose is to transform the expected pattern or common features into a usable reference standard; representative cucumber materials with specific properties refer to samples of cucumber materials that exhibit specific biological or agronomic properties under reference conditions, which can include known resistant or susceptible varieties, materials with specific genetic backgrounds, or materials obtained through preliminary screening, and its purpose is to provide actual data reflecting specific traits; the common covariant features between quantitative indicators refer to the correlation patterns between quantitative indicators that are commonly found among selected representative cucumber material data through statistical analysis, which can be determined by methods such as correlation analysis, principal component analysis, or cluster analysis, and its purpose is to extract representative indicator relationships from actual data.
[0062] The solution of this application solves the problem of selecting the reference benchmark by providing two flexible ways to determine the reference benchmark. The first way constructs the reference benchmark according to the preset expected covariant pattern, which utilizes existing biological knowledge or expert experience and can quickly establish a theoretically ideal covariant standard. The second way determines the common covariant features as the reference benchmark by analyzing the actual data of representative cucumber materials that exhibit specific properties under reference conditions, which makes the reference benchmark more objective and can reflect the resistance performance of real cucumber materials. The reference benchmark determined by these two ways, combined with the condition-specific covariant relationships of each data subset obtained in the previous steps, can more accurately evaluate the representativeness and importance of the correlations between indicators at different growth stages or under different environmental conditions. Based on this more accurate evaluation, the condition-specific covariant relationships can be effectively integrated and processed to generate an initial covariant relationship that can better represent the overall characteristics of the cucumber population and is more reliable. This reliable initial covariant relationship provides a stable reference for subsequent supplementary adjustment operations for data-deficient regions, ensuring that the adjusted data can more comprehensively reflect the diversity of the cucumber population while maintaining a reasonable covariant structure, and further promoting the more accurate characterization of the disease resistance mode of cucumbers by the identified resistance components.
[0063] As an implementation manner of the present invention, the step of selecting one or more groups of representative cucumber materials that exhibit specific properties under reference conditions includes:
[0064] Under reference conditions, phenotypic screening of target traits is carried out on candidate cucumber materials;
[0065] Based on the phenotypic screening results, representative cucumber materials showing specific performance are selected.
[0066] Among them, the reference conditions refer to the environmental state with clear control parameters set for the phenotypic evaluation of cucumber materials, which may include specific temperature, humidity, light intensity, soil type, nutrient supply level, and pathogen inoculation method and concentration, etc. The purpose is to ensure that the phenotypic performances of different cucumber materials are obtained under a unified and reproducible environmental background, so as to reduce the influence of environmental variation on the phenotypic evaluation results and improve the accuracy of screening; Candidate cucumber materials refer to cucumber plant or strain samples obtained from cucumber germplasm resources, breeding populations or natural populations, which have potential research or utilization value. The number of them can be large, covering a wide range of variations in the target traits of the cucumber population. The purpose is to provide a sufficient variety of selection objects for subsequent phenotypic screening; Target traits refer to the biological characteristics concerned in cucumber disease resistance breeding analysis and directly or indirectly related to disease resistance, which may include the resistance level to specific pathogens, the development speed of disease spots, the disease course progress curve, the intensity of defense response, etc. The purpose is to distinguish the disease resistance performance differences of different cucumber materials by evaluating these traits; Phenotypic screening refers to the process of observing, measuring or detecting the phenotypic traits of cucumber materials under specific conditions and identifying the materials that meet specific requirements according to preset standards or thresholds. It can adopt various technical means such as manual visual inspection, image analysis, physiological and biochemical detection, etc. The purpose is to preliminarily screen out the materials that show differences or reach a specific level in the target traits from a large number of candidate materials; Specific performance refers to the level or pattern of biological characteristics with representativeness or research value shown by cucumber materials in the target traits under reference conditions, which may include different resistance levels such as high resistance, medium resistance, and susceptibility, or specific defense response types or intensities. The purpose is to select the materials that can represent the key variation characteristics of the cucumber population in the target traits in order to construct a representative reference benchmark; Representative cucumber materials refer to a small number or multiple materials further selected from the materials qualified in phenotypic screening, which can reflect the variation range and characteristics of the target traits of the cucumber population under reference conditions. The selection can be based on factors such as the distribution of phenotypic data, the diversity of genetic background, or the deviation degree from the population average level, etc. The purpose is to use the data of these materials to construct a reference benchmark, so as to characterize the covariation characteristics of the cucumber population under reference conditions.
[0067] The solution of this application conducts phenotypic screening of target traits on candidate cucumber materials under reference conditions, which is to initially identify materials with differences in target traits from a wide range of candidate materials. Conducting screening under reference conditions ensures the comparability and reliability of phenotypic data and reduces the interference of environmental factors. Subsequently, based on the screening results, representative cucumber materials showing specific performance are selected. These selected materials are representative in terms of target traits and can reflect the variation characteristics of the cucumber population in this trait. By using the data of these representative materials to construct a reference benchmark, the covariation characteristics of the cucumber population under reference conditions can be more accurately characterized. This process provides a solid foundation for determining the initial covariation relationship between various quantitative indicators in the subsequent unified set of quantitative indicators, because the accuracy of the reference benchmark directly affects the effectiveness of the integration process of condition-specific covariation relationships. By selecting representative materials, it can be ensured that the reference benchmark can capture the core covariation patterns of the cucumber population in key traits, so that the finally generated initial covariation relationship is closer to the actual situation of the cucumber population, thereby optimizing subsequent disease-resistant breeding analyses, such as the identification of resistance components and gene association analysis.
[0068] As an implementation manner of the present invention, the steps of conducting phenotypic screening of target traits on candidate cucumber materials under reference conditions include:
[0069] Collect phenotypic images of candidate cucumber materials;
[0070] Analyze the phenotypic images to quantitatively evaluate the phenotypes of the target traits of the candidate cucumber materials and conduct phenotypic screening of the target traits based on the quantitative evaluation results.
[0071] Among them, the phenotypic image refers to a visual record that reflects the external manifestations of the target traits of cucumber materials under specific reference conditions, which can be obtained by devices such as high-resolution cameras, multispectral imaging devices, or 3D scanners. The purpose is to objectively and comprehensively capture phenotypic information such as the morphology, color, and texture of cucumber materials; analyzing the phenotypic image means using computer vision or image processing technology to extract feature information related to the target traits from the collected phenotypic images, which can include image preprocessing, image segmentation, and feature extraction. The purpose is to convert the visual information in the image into computable and analyzable data; quantitative evaluation refers to converting the performance of the target traits of cucumber materials into numerical indicators based on the feature information extracted from image analysis through a preset algorithm or model. It can be to calculate the proportion of the lesion area to the total leaf area as the disease index, or to measure the size, color intensity, etc. of specific parts. The purpose is to convert subjective observations into objective and accurate numerical values for subsequent comparison and screening.
[0072] The solution of this application obtains the visual phenotypic information of cucumber materials under reference conditions by collecting phenotypic images of candidate cucumber materials. The reason for using image collection is that images can record rich phenotypic details in a non-contact and high-throughput manner, overcoming the subjectivity and limitations of manual observation. It is precisely because objective and detailed phenotypic image data are obtained that subsequent in-depth processing of these images can be carried out using image analysis techniques to extract various features related to target traits, such as the morphology, color, distribution, etc. of disease spots. Based on these extracted features, quantitative evaluation of target traits can be carried out, converting complex phenotypic manifestations into precise numerical indicators, such as disease index, number of disease spots, etc. This quantitative evaluation process avoids the ambiguity and uncertainty of manual evaluation, enabling accurate capture and comparison of phenotypic differences between different materials. It is precisely because quantitative phenotypic evaluation results are obtained that scientific and objective phenotypic screening can be carried out based on these numerical values. By setting thresholds or sorting, representative cucumber materials that meet specific performance requirements can be quickly and accurately identified. This screening method based on image analysis and quantitative evaluation improves the efficiency and accuracy of screening, provides more reliable input for subsequent breeding analysis, thus helping to more accurately determine the reference benchmark, and further enhancing the effectiveness of the entire cucumber disease-resistant breeding analysis method.
[0073] As an implementation manner of the present invention, the step of performing data dimensionality reduction processing based on the adjusted set of quantitative indicators to identify multiple resistance components, where each resistance component characterizes a resistance action mode defined by a combination of microscopic phenotypic data, includes:
[0074] Obtain the adjusted set of quantitative indicators. The adjusted set of quantitative indicators is a matrix X_adj with n_adj rows and p columns, where n_adj is the number of adjusted cucumber material samples, and p is the number of microscopic phenotypic indicators after unified quantification;
[0075] Based on the matrix X_adj, analyze the covariance structure among the quantitative indicators in the set of quantitative indicators to identify k potential resistance action modes that jointly affect the combination of microscopic phenotypic data, where k is the number of identified resistance components, and k < p;
[0076] According to the identified resistance action modes, determine the composition method of each resistance component defined by the combination of microscopic phenotypic data, and quantify the mapping relationship between the microscopic phenotypic data and the resistance components to generate a load matrix L_fixed with p rows and k columns. The element l_jk of the load matrix L_fixed represents the association strength of the jth microscopic phenotypic indicator on the kth resistance component;
[0077] Based on the covariant characteristics of the load matrix \(L_{fixed}\) and the matrix \(X_{adj}\), a fixed transformation rule matrix \(W_{fixed}\) with \(p\) rows and \(k\) columns is calculated and determined. \(W_{fixed}\) is used to transform the unified quantification index set of any cucumber material into its quantification scores on each resistance component, that is, the resistance component feature vector \(S\) with \(n\) rows and \(k\) columns, where \(n\) is the number of samples in any batch of cucumber materials, and \(S = X*W_{fixed}\).
[0078] The solution of this application lays a data foundation for subsequent resistance component identification and quantification by obtaining the adjusted quantification index set matrix \(X_{adj}\) generated through complementary adjustment operations. Since \(X_{adj}\) is adjusted and can better represent the overall characteristics of the cucumber population, the analysis based on it is more representative. On this basis, the solution further analyzes the covariant structure among the quantification indexes based on the matrix \(X_{adj}\), rather than simply performing dimensionality reduction. By analyzing the covariant structure, \(k\) potential resistance action modes that jointly affect the combination of micro-phenotypic data can be identified from the complex micro-phenotypic data. These modes are reflections of the internal structure of the data and have potential biological significance. According to the identified resistance action modes, the solution determines the specific composition method of each resistance component, and quantifies the mapping relationship between the micro-phenotypic data and the resistance components by generating the load matrix \(L_{fixed}\). The load matrix \(L_{fixed}\) clearly shows the correlation strength of each micro-phenotypic index on different resistance components, making the resistance components more interpretable. Finally, based on the covariant characteristics of the load matrix \(L_{fixed}\) and the matrix \(X_{adj}\), the solution calculates and determines a fixed transformation rule matrix \(W_{fixed}\) with \(p\) rows and \(k\) columns. This \(W_{fixed}\) is a stable transformation model from micro-phenotypic data to resistance component scores. Through \(S = X*W_{fixed}\), the unified quantification index set \(X\) (including new data not used for training) of any cucumber material can be transformed into its resistance component feature vector \(S\). Since \(W_{fixed}\) is determined based on the more representative adjusted data \(X_{adj}\) and its covariant characteristics, it has better stability and generalization compared to the rules obtained based on unadjusted data or simple dimensionality reduction methods. This method of analyzing the covariant structure based on the adjusted data and determining the fixed transformation rule effectively solves the challenge of accurately extracting key resistance information and establishing a stable mapping relationship in high-dimensional and complex micro-phenotypic data, making the identified resistance components more biologically meaningful, and the transformation rule can be reliably applied to new batches of cucumber materials to achieve unified resistance evaluation.
[0079] As Figure 2 shown, a disease-resistant gene typing result analysis system for cucumber disease-resistant breeding analysis, the system includes:
[0080] The data acquisition and conversion module 201 is used to acquire multi-source heterogeneous microscopic phenotype data for cucumber disease-resistant breeding analysis and convert the multi-source heterogeneous microscopic phenotype data into a unified set of quantitative indicators;
[0081] The resistance component identification and eigenvector generation module 202 is used to identify multiple resistance components based on the unified set of quantitative indicators through data dimensionality reduction processing. Each resistance component represents a resistance action mode defined by the combination of microscopic phenotype data, and generate a resistance component eigenvector containing the quantitative scores of each cucumber material on each resistance component;
[0082] The gene association analysis module 203 is used to acquire the genotyping data of cucumber materials and perform a first association analysis based on the genotyping data and the resistance component eigenvector to identify gene information related to specific resistance components;
[0083] The macroscopic phenotype association analysis module 204 is used to acquire the macroscopic disease-resistant phenotype data of cucumber materials and perform a second association analysis based on the resistance component eigenvector and the macroscopic disease-resistant phenotype data to evaluate the effects of different resistance components on the macroscopic disease-resistant phenotype.
[0084] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. What is described in the above embodiments and the specification is only the principle of the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.
Claims
1. A method for analyzing the results of disease-resistant gene typing, which is applied to the analysis environment of cucumber disease-resistant breeding, is characterized in that The method includes the following steps: Obtain multi-source heterogeneous microscopic phenotype data for cucumber disease-resistant breeding analysis, and convert the multi-source heterogeneous microscopic phenotype data into a unified set of quantitative indicators; Based on the unified set of quantitative indicators, through data dimensionality reduction processing, identify multiple resistance components, each of the resistance components characterizing a resistance action mode defined by a combination of microscopic phenotype data, and generate a resistance component feature vector for each cucumber material, which includes its quantitative scores on each of the resistance components; Obtain the genotyping data of the cucumber material, and perform a first association analysis based on the genotyping data and the resistance component feature vector to identify gene information related to specific resistance components; Obtain the macroscopic disease-resistant phenotype data of the cucumber material, and perform a second association analysis based on the resistance component feature vector and the macroscopic disease-resistant phenotype data to evaluate the effects of different resistance components on the macroscopic disease-resistant phenotype.
2. The method for analyzing the genotyping results of disease-resistant genes according to claim 1, wherein The step of, based on the unified set of quantitative indicators, through data dimensionality reduction processing, identifying multiple resistance components, each of the resistance components characterizing a resistance action mode defined by a combination of the microscopic phenotype data, and generating a resistance component feature vector for each cucumber material, which includes its quantitative scores on each of the resistance components, includes: Perform data dimensionality reduction processing on the unified set of quantitative indicators of at least one initial cucumber material batch to identify multiple resistance components, each of the resistance components characterizing a resistance action mode defined by a combination of the microscopic phenotype data, and based on this, determine a fixed conversion rule for generating the resistance component feature vector from the unified set of quantitative indicators; Apply the fixed conversion rule to the unified set of quantitative indicators of the at least one initial cucumber material batch and subsequent cucumber material batches obtained, to generate the resistance component feature vector for each cucumber material, which includes its quantitative scores on each of the resistance components.
3. The method for analyzing the genotyping results of disease-resistant genes according to claim 2, wherein The step of, based on the unified set of quantitative indicators of at least one initial cucumber material batch, performing data dimensionality reduction processing to identify multiple resistance components, each of the resistance components characterizing a resistance action mode defined by a combination of the microscopic phenotype data, includes: Obtain reference information characterizing the overall characteristics of the cucumber population, the reference information including at least one of the germplasm resource diversity distribution information of the cucumber population and the expected variation characteristics of key microscopic phenotypes; Compare the unified set of quantitative indicators of the at least one initial cucumber material batch with the reference information to identify regions in the unified set of quantitative indicators that do not fully represent the cucumber population; For the regions that do not fully represent the cucumber population, perform a supplementary adjustment operation on the unified set of quantitative indicators to generate an adjusted set of quantitative indicators, the supplementary adjustment operation enabling the adjusted set of quantitative indicators to, when used for subsequent data dimensionality reduction processing, prompt the identified resistance components to reflect a more comprehensive resistance action mode of the cucumber population; Perform data dimensionality reduction processing based on the adjusted set of quantization metrics to identify the multiple resistance components, where each resistance component represents a resistance action pattern defined by the combination of the micro-phenotype data.
4. The method for analyzing the genotyping results of disease-resistant genes according to claim 3, characterized in that The step of performing a supplementary adjustment operation on the unified set of quantization metrics for the region that fails to adequately represent the cucumber population to generate an adjusted set of quantization metrics includes: Determine the initial covariance relationship among the quantization metrics in the unified set of quantization metrics; During the process of performing the supplementary adjustment operation on the unified set of quantization metrics for the region that fails to adequately represent the cucumber population to change the data composition or values of the unified set of quantization metrics, evaluate the change in the initial covariance relationship caused by the supplementary adjustment operation; When the change in the covariance relationship among the quantization metrics caused by the supplementary adjustment operation exceeds a preset condition relative to the initial covariance relationship, adjust the execution method or execution amplitude of the supplementary adjustment operation so that the covariance relationship in the adjusted set of quantization metrics generated based on the adjusted supplementary adjustment operation meets the preset condition.
5. A method for analyzing the results of disease-resistant gene typing according to claim 4, characterized in that The step of determining the initial covariance relationship among the quantization metrics in the unified set of quantization metrics includes: Obtain the growth stage information or environmental condition information corresponding to each cucumber material in the unified set of quantization metrics; According to the growth stage information or the environmental condition information, divide the unified set of quantization metrics into multiple data subsets, where each data subset corresponds to a specific growth stage or environmental condition; For each data subset, calculate the conditional specific covariance relationship among the quantization metrics within the data subset; Determine a reference benchmark, where the reference benchmark is used to characterize the covariance characteristics of the cucumber population under reference conditions; Integrate the conditional specific covariance relationships of each data subset in combination with the degree of association between the conditional specific covariance relationship of each data subset and the reference benchmark, and the data proportion of each data subset in the unified set of quantization metrics, to generate the initial covariance relationship among the quantization metrics in the unified set of quantization metrics.
6. The method for analyzing the genotyping results of disease-resistant genes according to claim 5, wherein The step of determining a reference benchmark, where the reference benchmark is used to characterize the covariance characteristics of the cucumber population under reference conditions, is achieved by one of the following methods: Construct the reference benchmark according to a preset expected covariance pattern among the indicators of the cucumber under the reference conditions; Alternatively, select one or more groups of representative cucumber materials that exhibit specific performance under the reference conditions, obtain the unified set of quantization metrics corresponding to the representative cucumber materials, analyze and determine the common covariance characteristics among the quantization metrics thereof based on the obtained unified set of quantization metrics corresponding to the representative cucumber materials, and use the determined common covariance characteristics as the reference benchmark.
7. A method for analyzing the results of disease-resistant gene typing according to claim 6, characterized in that, The step of selecting one or more groups of representative cucumber materials that exhibit specific performance under the reference conditions includes: Under the reference conditions, perform phenotypic screening of the target traits on the candidate cucumber materials. According to the phenotypic screening results, representative cucumber materials showing the specific performance are selected.
8. A method for analyzing the results of disease-resistant gene typing according to claim 7, characterized in that, The step of performing phenotypic screening of target traits on candidate cucumber materials under the reference conditions includes: Collecting phenotypic images of the candidate cucumber materials; Analyzing the phenotypic images to quantitatively evaluate the phenotypes of the target traits of the candidate cucumber materials, and performing phenotypic screening of the target traits based on the quantitative evaluation results.
9. The method for analyzing the genotyping results of disease-resistant genes according to claim 3, wherein The step of performing data dimensionality reduction processing based on the adjusted set of quantitative indicators to identify the multiple resistance components, where each resistance component characterizes a resistance action pattern defined by a combination of the microscopic phenotypic data includes: Obtaining the adjusted set of quantitative indicators, where the adjusted set of quantitative indicators is a matrix X_adj with n_adj rows and p columns, where n_adj is the number of adjusted cucumber material samples and p is the number of uniformly quantified microscopic phenotypic indicators; Based on the matrix X_adj, analyzing the covariance structure among the quantitative indicators in the set of quantitative indicators to identify k potential resistance action patterns that jointly affect the combination of the microscopic phenotypic data, where k is the number of identified resistance components and k < p; According to the identified resistance action patterns, determining the composition method of each resistance component defined by the combination of the microscopic phenotypic data, and quantifying the mapping relationship between the microscopic phenotypic data and the resistance components to generate a load matrix L_fixed with p rows and k columns, where the element l_jk of the load matrix L_fixed represents the association strength of the j-th microscopic phenotypic indicator on the k-th resistance component; Based on the covariance characteristics of the load matrix L_fixed and the matrix X_adj, calculating and determining a fixed transformation rule matrix W_fixed with p rows and k columns, where W_fixed is used to convert the set of uniformly quantified indicators of any cucumber material into its quantitative scores on each of the resistance components, that is, a resistance component eigenvector S with n rows and k columns, where n is the number of samples in any batch of cucumber materials, and S = X * W_fixed.
10. A disease resistance gene typing result analysis system for cucumber disease resistance breeding analysis, characterized in that, The system includes: A data acquisition and conversion module for acquiring multi-source heterogeneous microscopic phenotypic data for cucumber disease-resistant breeding analysis and converting the multi-source heterogeneous microscopic phenotypic data into a unified set of quantitative indicators; A resistance component identification and eigenvector generation module for identifying multiple resistance components based on the unified set of quantitative indicators through data dimensionality reduction processing, where each resistance component characterizes a resistance action pattern defined by a combination of microscopic phenotypic data, and generating a resistance component eigenvector containing its quantitative scores on each of the resistance components for each cucumber material; A gene association analysis module for obtaining the genotyping data of the cucumber materials and performing a first association analysis based on the genotyping data and the resistance component eigenvectors to identify gene information related to specific resistance components; A macroscopic phenotype association analysis module, which is used to obtain the macroscopic disease-resistant phenotype data of the cucumber material, and perform a second association analysis based on the resistance component feature vector and the macroscopic disease-resistant phenotype data to evaluate the effects of different resistance components on the macroscopic disease-resistant phenotype.
Citation Information
Patent Citations
Lemon disease-resistant gene map construction method based on data mining
CN118335187A
Bidirectional correlation analysis method for plant and pathogenic bacterium gene interaction and application
CN120108498A
Grape downy mildew resistance whole genome selective breeding method based on machine learning
CN120126571A
Statistical validation of candidate genes
US20100145624A1
Cited By
Watermelon resistance breeding optimization method and system based on cultivation data identification traceability
CN121503787A
A method and system for cultivating watermelon resistance breeding optimization with data identification traceability
CN121503787B