A method and system for analyzing disease resistance gene typing results
By converting multi-source heterogeneous microphenotype data into a unified set of quantitative indicators and reducing the dimensionality of data, identifying resistance components, and combining genotyping and macrophenotype data for correlation analysis, the problem of unclear resistance mechanism in cucumber disease-resistant breeding is solved, and precise breeding is achieved.
Patent Information
- Application Number
- CN202510845183.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-23
- Publication Date
- 2025-08-15
- Estimated Expiration
- 2045-06-23
AI Technical Summary
The existing technology is difficult to effectively integrate multi-source heterogeneous microphenotype data, and cannot reveal the resistance mechanism in cucumber disease-resistant breeding, resulting in breeding decisions relying on macro results and being unable to achieve precise breeding.
Multi-source heterogeneous microphenotype data are converted into a unified set of quantitative indicators, resistant components are identified through data dimensionality reduction processing, and resistance components are generated, and correlation analysis is performed by combining genotyping and macroscopic disease-resistant phenotype data.
A deep understanding of the pathogenesis mechanism is achieved, and precise breeding based on mechanisms is supported, which improves breeding efficiency and accuracy.
Smart Images

Figure CN120356529B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of plant disease resistance breeding analysis, and in particular to a disease resistance gene typing result analysis method and system. Background Art
[0002] In modern plant breeding, especially in disease-resistant breeding programs for important cash crops like cucumber, breeders rely extensively on molecular marker-assisted selection (MAS) to improve breeding efficiency. The conventional process typically involves the following steps: First, leaf samples are collected from germplasm resources or genetic populations, genomic DNA is extracted, and high-throughput genotyping technologies (such as microarrays or sequencing) are used to obtain molecular marker data for the entire genome or target regions. Simultaneously, to evaluate the disease resistance of these cucumber materials, breeders artificially inoculate or naturally induce target diseases in a strictly controlled greenhouse environment or at high-incidence field sites. Macroscopic disease resistance phenotypic data, such as disease incidence, disease index, lesion characteristics, and yield loss assessments, are recorded. Existing analysis systems or methods primarily integrate genotyping and macroscopic phenotypic data to conduct genotype-phenotype association analysis, identifying candidate gene regions or molecular markers significantly associated with disease resistance traits for use in material screening.
[0003] However, breeders are no longer satisfied with simply locating disease-resistance genes; they also want to understand their mechanisms of action. Therefore, microscopic phenotypic data have been introduced, namely, data on the dynamic responses of plant cells, tissues, and physiology and biochemistry during pathogen infection. Specifically, these include: image information such as the invasion sites of pathogens in leaf tissues, hyphal expansion patterns, and ultrastructural changes in host cells, observed and recorded using equipment such as confocal microscopy; time series data on a series of key defense indicators measured through physiological and biochemical analysis, such as concentration curves of defense signaling molecules, burst levels of reactive oxygen species, and activity profiles of defense enzymes associated with disease resistance; and high-dimensional concentration data on disease-resistance-related secondary metabolites, such as the dynamic accumulation patterns of flavonoids and phenolics, detected using metabolomics technology.
[0004] In the practice of cucumber disease-resistant breeding, when breeders simultaneously obtain genotyping data as the genetic basis, macroscopic disease-resistance phenotypic data as the basis for final evaluation, and multi-source heterogeneous microscopic phenotypic data used to reveal the disease-resistance mechanism, the core technical problem faced by existing analysis methods is: how to design a data processing method that does not rely on preset biological pathway models, effectively integrate and quantify these microscopic phenotypic data with different structures and dimensions, and associate them with genotypic data, so as to not only evaluate the final disease-resistance results, but also reveal the intrinsic resistance action mode determined by the genotype, hidden under the macroscopic appearance, and represented by different microscopic phenotypic combinations, so as to solve the dilemma that the current system cannot parse microscopic data, resulting in unclear resistance mechanisms, making breeding decisions rely on macroscopic results and unable to achieve precise breeding based on resistance mechanisms.
[0005] In view of the above problems, the existing technology is in urgent need of improvement. Summary of the Invention
[0006] The purpose of the present invention is to solve the shortcomings of the existing technology and to propose a disease resistance gene typing result analysis method and system.
[0007] In a first aspect, the present invention provides a method for analyzing disease resistance gene typing results, which is applied to the analysis environment of cucumber disease resistance breeding, characterized in that the method comprises the following steps:
[0008] Acquiring multi-source heterogeneous microscopic phenotypic data for cucumber disease resistance breeding analysis, and converting the multi-source heterogeneous microscopic phenotypic data into a unified set of quantitative indicators;
[0009] Based on the unified set of quantitative indicators, multiple resistance components are identified through data dimensionality reduction processing, each of which represents a resistance action mode defined by a combination of microscopic phenotypic data, and a resistance component feature vector containing its quantitative score on each resistance component is generated for each cucumber material;
[0010] Obtaining genotyping data of the cucumber material, and performing a first association analysis based on the genotyping data and the resistance component feature vector to identify gene information related to a specific resistance component;
[0011] The macroscopic disease resistance phenotype data of the cucumber material is obtained, and a second association analysis is performed based on the resistance component characteristic vector and the macroscopic disease resistance phenotype data to evaluate the effects of different resistance components on the macroscopic disease resistance phenotype.
[0012] The core innovation of this application lies in that by converting multi-source heterogeneous microscopic phenotypic data into a unified set of quantitative indicators, and based on this, multiple resistance components are identified through data dimensionality reduction processing, thereby abstracting the complex microscopic resistance response into a quantifiable component feature vector, and further correlating the feature vector with genotyping data and macroscopic disease resistance phenotypic data for analysis. It briefly describes how to solve the problem that existing technologies are difficult to effectively integrate microscopic data and reveal resistance mechanisms, and achieve the effect of in-depth understanding of disease resistance mechanisms and realizing mechanism-based precision breeding.
[0013] In a second aspect, a disease resistance gene typing result analysis system is provided for cucumber disease resistance breeding analysis, the system comprising:
[0014] A data acquisition and conversion module is used to acquire multi-source heterogeneous microscopic phenotypic data for cucumber disease resistance breeding analysis and convert the multi-source heterogeneous microscopic phenotypic data into a unified set of quantitative indicators;
[0015] a resistance component identification and feature vector generation module, configured to identify multiple resistance components based on the unified set of quantitative indicators through data dimensionality reduction processing, wherein each resistance component represents a resistance action mode defined by a combination of microscopic phenotypic data, and generate a resistance component feature vector for each cucumber material containing its quantitative score on each resistance component;
[0016] a gene association analysis module, configured to obtain genotyping data of the cucumber material, and perform a first association analysis based on the genotyping data and the resistance component feature vector to identify gene information associated with a specific resistance component;
[0017] The macro-phenotype association analysis module is used to obtain the macro-disease resistance phenotype data of the cucumber material, and perform a second association analysis based on the resistance component characteristic vector and the macro-disease resistance phenotype data to evaluate the effects of different resistance components on the macro-disease resistance phenotype.
[0018] Compared with the prior art, the present invention has the following beneficial effects:
[0019] By integrating multi-source micro-phenotypic data, identifying intrinsic resistance components, and analyzing their association with genotypes and macro-phenotypes, the problem that existing technologies are unable to analyze micro-data and the resistance mechanism is unclear is solved. It has the advantage of being able to effectively integrate multi-source heterogeneous micro-phenotypic data, revealing the intrinsic resistance action mode determined by genotype, and providing a basis for precision breeding based on resistance mechanisms. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] Figure 1 Flow chart of the method of the present invention.
[0021] Figure 2 This is a system structure diagram of the present invention.
[0022] In the figure: 201, data acquisition and conversion module; 202, resistance component identification and feature vector generation module; 203, gene association analysis module; 204, macro-phenotype association analysis module. DETAILED DESCRIPTION
[0023] The embodiments of the present invention are described in detail below, examples of which are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and are not to be construed as limiting the present invention.
[0024] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of the technical features being referred to. Thus, a feature specified as "first" or "second" may explicitly or implicitly include one or more of the specified features. In the description of the present invention, "plurality" means two or more, unless otherwise specifically defined.
[0025] like Figure 1 A disease resistance gene typing result analysis method is shown, which is applied to the cucumber disease resistance breeding analysis environment, and the method comprises the following steps:
[0026] Acquire multi-source heterogeneous microscopic phenotypic data for cucumber disease resistance breeding analysis and convert the multi-source heterogeneous microscopic phenotypic data into a unified set of quantitative indicators;
[0027] Based on a unified set of quantitative indicators, multiple resistance components were identified through data dimensionality reduction. Each resistance component represents a resistance action mode defined by a combination of microscopic phenotypic data. A resistance component feature vector containing the quantitative scores of each resistance component was generated for each cucumber material.
[0028] Acquiring genotyping data of the cucumber material, and performing a first association analysis based on the genotyping data and the resistance component feature vector to identify gene information related to the specific resistance component;
[0029] The macroscopic disease resistance phenotype data of cucumber materials were obtained, and a second association analysis was performed based on the resistance component characteristic vectors and the macroscopic disease resistance phenotype data to evaluate the effects of different resistance components on the macroscopic disease resistance phenotype.
[0030] Multi-source heterogeneous microphenotypic data refers to data derived from different detection technologies, with varying data structures and dimensions, that characterizes the response of cucumber plants to pathogens at the microscopic levels, such as cellular, tissue, and physiological and biochemical levels. This data can be obtained in various forms, including microscopic image data, time-series data of physiological and biochemical indicators, and metabolomics data. Its primary purpose is to uncover the underlying biological mechanisms of cucumber disease resistance. A unified quantitative indicator set involves converting this multi-source heterogeneous microphenotypic data into a dataset with a unified format and dimensions, such as a matrix, where rows represent different cucumber accessions and columns represent quantified microphenotypic indicators, after preprocessing, standardization, and feature extraction. This is primarily intended to eliminate discrepancies between data sources and facilitate subsequent integrated analysis. Data dimensionality reduction involves mapping a high-dimensional set of quantitative indicators into a lower-dimensional space using mathematical or statistical methods, such as principal component analysis (PCA), factor analysis (FA), independent component analysis (ICA), or other nonlinear dimensionality reduction techniques, while preserving as much important information as possible from the original data. This approach aims to simplify the data structure and extract potential key features or patterns. Resistance components are potential factors or dimensions that characterize a specific resistance mode, identified from microscopic phenotypic data through data dimensionality reduction. Each component represents a specific combination or weighting of multiple microscopic phenotypic indicators and is primarily used to summarize and abstract complex microscopic resistance response processes. A resistance mode is a biological function or mechanism defined by a specific combination of microscopic phenotypic indicators, such as a defense mode consisting of cell wall thickening and reactive oxygen species bursts. It primarily describes the specific way cucumber plants resist pathogen infection. A resistance component feature vector is a vector generated for each cucumber accession containing its quantitative scores for each identified resistance component. This vector represents the accession's microscopic resistance characteristics in a low-dimensional space and is primarily used to quantify differences in the performance of different cucumber accessions across various resistance modes. Primary association analysis involves statistical association analysis, such as genome-wide association studies (GWAS) or quantitative trait loci (QTL) mapping, between the genotyping data of the cucumber accessions and the resistance component feature vector. This analysis identifies genomic regions or genes significantly associated with specific resistance components. The second association analysis refers to the statistical association analysis of the resistance component characteristic vector and the macroscopic disease resistance phenotypic data of the cucumber material, such as regression analysis or correlation analysis, which is mainly used to evaluate the contribution of different resistance components to the final macroscopic disease resistance performance.
[0031] The solution of the present application achieves the standardization and integration of micro data from different sources and formats by obtaining multi-source heterogeneous micro phenotypic data for cucumber disease resistance breeding analysis and converting it into a unified set of quantitative indicators. Based on this unified set of quantitative indicators, data dimensionality reduction processing is performed. This process can extract the potential structure reflecting the intrinsic resistance mechanism from the high-dimensional micro phenotypic data, that is, multiple resistance components, each component representing a resistance action mode defined by a specific micro phenotypic combination. Subsequently, the quantitative score of each cucumber material on these resistance components is calculated to generate a resistance component feature vector, which is a refined representation of the micro resistance characteristics of the material. Furthermore, the genotyping data of the cucumber material is obtained, and a first association analysis is performed based on the genotyping data and the resistance component feature vector, so that the gene information related to the specific resistance component (that is, the specific resistance action mode) can be identified, revealing the genetic basis of the resistance mechanism. Simultaneously, macroscopic disease resistance phenotypic data for the cucumber material is obtained, and a secondary association analysis is performed based on the resistance component eigenvectors and macroscopic disease resistance phenotypic data. This allows for the assessment of the impact and contribution of different resistance components (i.e., different resistance modes of action) to the ultimate macroscopic disease resistance. This entire process forms a complete analytical chain from microscopic mechanisms to genotypes to macroscopic phenotypes, enabling breeders to go beyond simple gene-macroscopic phenotype associations and gain a deeper understanding of how genes determine ultimate disease resistance by influencing specific microscopic resistance modes of action.
[0032] As an embodiment of the present invention, based on a unified set of quantitative indicators, multiple resistance components are identified through data dimensionality reduction processing, each resistance component represents a resistance action mode defined by a combination of microscopic phenotypic data, and a resistance component feature vector containing quantitative scores for each resistance component is generated for each cucumber material, including the following steps:
[0033] Performing data dimensionality reduction based on a unified set of quantitative indicators for at least one initial batch of cucumber material to identify multiple resistance components, each resistance component representing a resistance action mode defined by a combination of microscopic phenotypic data, and determining a fixed conversion rule for generating a resistance component feature vector from the unified set of quantitative indicators based on the resistance component feature vector;
[0034] A fixed transformation rule is applied to a unified set of quantitative indicators of at least one initial cucumber material batch and subsequently obtained cucumber material batches to generate a resistance component feature vector for each cucumber material containing its quantitative score on each resistance component.
[0035] Among them, data dimensionality reduction refers to mapping high-dimensional data (such as a unified set of quantitative indicators) to a low-dimensional space through mathematical or statistical methods, while retaining the important information or structure in the original data as much as possible. It can be achieved using techniques such as principal component analysis (PCA), independent component analysis (ICA), and factor analysis (FA). Its purpose is to extract representative, mutually independent, or low-correlated potential resistance action patterns from complex microphenotypic data, reduce data dimensions, and facilitate subsequent analysis; resistance components refer to potential factors or principal components identified from microphenotypic data through data dimensionality reduction. Each component represents a specific resistance action pattern contributed by multiple microphenotypes. Its purpose is to abstract complex microresistance manifestations into a few quantifiable dimensions; fixed transformation rules refer to mathematical models or parameter sets determined based on the initial data batch for linearly or nonlinearly mapping a unified set of quantitative indicators to the characteristic vectors of resistance components. It can be a transformation matrix (such as the PCA loading matrix). Its purpose is to provide a stable and unified transformation standard for subsequent batches of data analysis.
[0036] The proposed approach first identifies multiple resistance components by performing data dimensionality reduction based on a unified set of quantitative indices for at least one initial batch of cucumber material. Based on this, a fixed transformation rule is then determined for generating resistance component feature vectors from the unified set of quantitative indices. This process establishes a stable analytical benchmark using a representative initial data batch. Through data dimensionality reduction, core resistance action patterns are extracted from high-dimensional microphenotypic data and solidified into a transformation rule. This fixed transformation rule is then applied to the unified set of quantitative indices for at least one initial batch of cucumber material and subsequent batches of cucumber material to generate a resistance component feature vector for each cucumber material, containing the quantitative scores for each resistance component. This means that both the initial data used to establish the rules and the subsequent new data are transformed using the same fixed set of rules. This approach avoids the time-consuming re-recalculation of data dimensionality reduction each time new data is added, significantly improving processing efficiency. Furthermore, because all batches of data are mapped to the same resistance component space defined by the initial data, the scores for each resistance component from different batches of cucumber material are directly comparable, ensuring the reliability of subsequent association analysis results. This strategy of separating resistance component identification from rule determination and applying fixed rules to all data makes the entire cucumber disease resistance breeding analysis method more efficient, stable and consistent when processing large-scale, batch data, thereby better supporting resistance mechanism analysis and precision breeding based on microscopic phenotypic data.
[0037] As an embodiment of the present invention, the steps of performing data dimensionality reduction processing based on a unified set of quantitative indicators of at least one initial batch of cucumber materials to identify multiple resistance components, each resistance component representing a resistance action mode defined by a combination of microscopic phenotypic data, include:
[0038] Obtaining reference information characterizing the overall characteristics of the cucumber population, the reference information including the diversity distribution information of the cucumber population's germplasm resources and at least one of the expected variation characteristics of key microscopic phenotypes;
[0039] comparing the unified set of quantitative indices for at least one initial batch of cucumber material to the reference information to identify areas of the cucumber population that are underrepresented in the unified set of quantitative indices;
[0040] For areas that fail to fully represent the cucumber population, a supplementary adjustment operation is performed on the unified quantitative indicator set to generate an adjusted quantitative indicator set. The supplementary adjustment operation enables the adjusted quantitative indicator set to be used in the subsequent data dimensionality reduction process to enable the identified resistance components to reflect a more comprehensive resistance action pattern of the cucumber population.
[0041] Data dimensionality reduction was performed based on the adjusted set of quantitative indicators to identify multiple resistance components, each of which represents a resistance action mode defined by a combination of microscopic phenotypic data.
[0042] Among them, the reference information that characterizes the overall characteristics of the cucumber population refers to data or models that can reflect the macroscopic characteristics of the cucumber population in terms of genetic background, phenotypic variation, etc., and its purpose is to provide a benchmark for evaluating the representativeness of the initial quantitative indicator set; the germplasm resource diversity distribution information of the cucumber population refers to the composition ratio or distribution pattern of the cucumber population in terms of different geographical origins, genetic lineages, variety types, etc., which can be specifically the population structure analysis results constructed based on molecular marker data, the composition list of the core germplasm bank, the list of representative materials of different ecological types, etc., and its purpose is to evaluate the coverage of the initial material batch in genetic diversity; the expected variation characteristics of key micro-phenotypes refer to the variation range, distribution pattern or mutual relationship of micro-phenotypic indicators closely related to disease resistance in a healthy or typical susceptible / resistant state, which can be the normal fluctuation range, ideal distribution curve of each micro-phenotypic indicator based on historical data, biological knowledge or expert experience, or the expected correlation pattern between different indicators, and its purpose is to evaluate the integrity and typicality of the initial material batch in key phenotypic data; failure to fully represent cucumbers The region of the population refers to the part of the initial quantitative indicator set where the data is sparse, missing, abnormally distributed, or fails to cover the range of specific germplasm resources / phenotypic variation after comparison with the reference information. Specifically, it may be that there are too few data points for certain micro-phenotypic indicators, insufficient material samples for certain germplasm resources, or certain key phenotypic combinations are not reflected in the data. Its purpose is to clarify the target range that requires supplementary adjustment; the supplementary adjustment operation refers to the data processing process of the initial quantitative indicator set, which aims to improve its representativeness of the overall characteristics of the cucumber population. Specifically, it can be achieved through data interpolation, data weighting, introduction of simulated data, or selective increase of samples. Its purpose is to generate a more representative set of quantitative indicators to optimize the subsequent dimensionality reduction results; the adjusted quantitative indicator set refers to the quantitative indicator set obtained after the supplementary adjustment operation. Compared with the initial set, this set has been improved in terms of data coverage, distribution characteristics or germplasm resources / phenotype representativeness. Its role is to serve as the input of the data dimensionality reduction process to enable the identified resistance components to reflect a more comprehensive resistance action pattern of the cucumber population.
[0043] The solution of this application first obtains reference information representing the overall characteristics of a cucumber population as an evaluation benchmark. A unified set of quantitative indices for at least one initial batch of cucumber material is then compared with this reference information. This identifies regions within the unified set of quantitative indices that underrepresent the cucumber population. Next, supplementary adjustments are performed on the unified set of quantitative indices to generate an adjusted set of quantitative indices for these underrepresented regions. This targeted supplementary adjustment significantly improves the data coverage and representativeness of the adjusted set of quantitative indices. Finally, data dimensionality reduction is performed based on this more representative adjusted set of quantitative indices, enabling the identification of multiple resistance components that more comprehensively reflect the overall resistance action patterns of the cucumber population. This strategy of incorporating reference information and evaluating and adjusting initial data effectively overcomes the problem of incomplete resistance component identification that can occur when directly using limited initial data for dimensionality reduction. This ensures that subsequent analyses based on these resistance components (e.g., association analysis with genotypes or macrophenotypes) yield more reliable and comprehensive results, thereby enhancing the effectiveness of the entire breeding analysis method.
[0044] As an embodiment of the present invention, for areas that fail to fully represent the cucumber population, performing a supplementary adjustment operation on the unified quantitative index set to generate an adjusted quantitative index set includes the following steps:
[0045] Determine the initial covariance relationship between the quantitative indicators in the unified quantitative indicator set;
[0046] In the process of performing supplementary adjustment operations on the unified quantitative indicator set to change the data composition or values of the unified quantitative indicator set for the area where the cucumber population is not fully represented, the changes caused by the supplementary adjustment operations on the initial covariation relationship are evaluated;
[0047] When the supplementary adjustment operation causes the covariation relationship between the quantitative indicators to change beyond the preset conditions relative to the initial covariation relationship, the execution method or execution range of the supplementary adjustment operation is adjusted so that the covariation relationship in the adjusted quantitative indicator set generated based on the adjusted supplementary adjustment operation meets the preset conditions.
[0048] Among them, determining the initial covariance relationship between each quantitative indicator in a unified set of quantitative indicators refers to the statistical association pattern between different quantitative indicators in the set before performing a supplementary adjustment operation on the unified set of quantitative indicators, such as the correlation, variance or covariance structure between them. It can be determined by calculating the correlation coefficient matrix, covariance matrix or performing principal component analysis, etc. Its purpose is to establish a benchmark for subsequent evaluation of the impact of data adjustment on the relationship between indicators; evaluating the changes caused by the supplementary adjustment operation on the initial covariance relationship refers to monitoring and quantifying the degree of difference between the covariance relationship between each quantitative indicator in the adjusted set of quantitative indicators and the initial covariance relationship in the process of performing the supplementary adjustment operation. It can be determined by comparing the difference between the covariance matrix before and after adjustment, calculating the change in the correlation coefficient between specific indicator pairs or using The purpose of adjusting the data is to understand the impact of data adjustment on the original data structure in real time or regularly; the preset condition refers to the threshold or standard used to judge whether the impact of the supplementary adjustment operation on the covariation relationship between quantitative indicators is acceptable, which can be set as the maximum allowable range of change in the covariation relationship, the allowable range of change of the correlation coefficient of a specific key indicator, or the significance level based on a statistical test, etc., with the purpose of controlling the degree of data adjustment to avoid excessive damage to the original data structure; adjusting the execution mode or execution range of the supplementary adjustment operation refers to modifying the ongoing or subsequent supplementary adjustment operation according to the evaluation results. For example, the amount of data addition or modification can be reduced, the weighting coefficient can be changed, or a different adjustment algorithm can be used. The purpose is to correct the adjustment operation so that the adjusted data covariation relationship meets the preset conditions.
[0049] The solution of the present application establishes a reference benchmark before data adjustment by determining the initial covariation relationship between each quantitative indicator in a unified set of quantitative indicators. Subsequently, in the process of performing supplementary adjustment operations to change the data composition or value for areas that fail to fully represent the cucumber population, the changes caused by the operation to the initial covariation relationship are continuously evaluated. It is precisely because of the introduction of monitoring and evaluation of changes in covariation relationships during the adjustment process that when the changes exceed the preset conditions, the execution method or execution range of the supplementary adjustment operation can be adjusted in a timely manner. This feedback control mechanism ensures that the supplementary adjustment enhances the representativeness of the data while retaining the original covariation structure in the data that reflects the biological association to the greatest extent. This combination enables the resistance components identified based on the adjusted data to reflect a more comprehensive resistance action pattern of the cucumber population and ensures that these patterns are based on relatively real data associations, thereby improving the accuracy and reliability of subsequent breeding analysis.
[0050] As an embodiment of the present invention, the step of determining the initial covariance relationship between the quantitative indicators in the unified quantitative indicator set includes:
[0051] Obtaining growth stage information or environmental condition information corresponding to each cucumber material in a unified quantitative index set;
[0052] Based on the growth stage information or environmental condition information, the unified quantitative indicator set is divided into multiple data subsets, each data subset corresponds to a specific growth stage or environmental condition;
[0053] For each data subset, calculate the condition-specific covariation relationship between the quantitative indicators within the data subset;
[0054] Determine a reference benchmark, which is used to characterize the covariation characteristics of the cucumber population under reference conditions;
[0055] Combining the degree of correlation between the condition-specific covariation relationship of each data subset and the reference benchmark, as well as the data proportion of each data subset in the unified quantitative indicator set, the condition-specific covariation relationships are integrated to generate the initial covariation relationship between the quantitative indicators in the unified quantitative indicator set.
[0056] Among them, growth stage information or environmental condition information refers to the developmental period of the cucumber material or the external environmental conditions it experienced when collecting microscopic phenotypic data. It can be obtained through manual recording, sensor monitoring or image analysis, etc. Its purpose is to group the cucumber materials so as to analyze the phenotypic responses under different conditions; data subset refers to the data grouping selected from a unified set of quantitative indicators according to a specific growth stage or environmental condition. Each data subset contains the quantitative indicator data of all cucumber materials collected under this specific condition. Its purpose is to decompose the overall data into more homogeneous parts for condition-specific analysis; condition-specific covariation relationship is It refers to the degree and pattern of mutual correlation or common variation among the quantitative indicators in a unified set of quantitative indicators under a specific growth stage or environmental conditions. It can be represented by calculating the covariance matrix, correlation coefficient matrix or partial correlation network of the data subset, and its purpose is to reveal the intrinsic connection between indicators under specific conditions; the reference benchmark refers to a reference pattern that is predetermined or extracted from representative data and is used to measure and compare covariation relationships under different conditions. It can be determined based on a preset model or representative samples, and its purpose is to provide an objective comparison standard to evaluate the degree of deviation of condition-specific covariation relationships from the reference pattern; the reference condition refers to the condition used to define the covariance matrix. The specific growth stage and environmental conditions of the reference benchmark can be the standard conditions for normal growth and development of cucumbers, or the typical conditions for a certain known resistance performance; the degree of association refers to the similarity or difference measure between the condition-specific covariation relationship of each data subset and the reference benchmark, which can be evaluated by calculating matrix distance, similarity index or statistical test, etc. Its purpose is to quantify the degree of closeness between the covariation pattern under specific conditions and the reference pattern; the data proportion refers to the proportion of the number of cucumber materials contained in each data subset to the total number of cucumber materials in the unified quantitative indicator set, and its purpose is to reflect the representativeness or importance of the data subset in the overall data; the integration processing is It refers to the comprehensive consideration of the degree of correlation between multiple condition-specific covariation relationships and the reference benchmark, as well as the data proportion of the corresponding data subset, and the fusion of them into a single initial covariation relationship representing the overall population characteristics through a certain algorithm or model. Methods such as weighted average, Bayesian integration, and graph model fusion can be used. The purpose is to generate a more comprehensive and robust initial covariation relationship; the initial covariation relationship refers to the overall covariation pattern of each quantitative indicator in the unified set of quantitative indicators obtained after integration processing at the level of the entire cucumber population. It serves as input or reference for subsequent data processing, and its purpose is to provide a relatively stable and accurate correlation structure information between indicators.
[0057] The solution of the present application obtains the growth stage information or environmental condition information of the cucumber material, and divides the unified quantitative indicator set into multiple data subsets, so that data under different conditions can be analyzed. The condition-specific covariation relationship is calculated for each data subset, revealing the correlation pattern between the quantitative indicators under specific growth stages or environmental conditions. By determining a reference benchmark, a unified standard for evaluating covariation relationships under different conditions is provided. Finally, the condition-specific covariation relationships of each data subset are integrated based on the degree of correlation between the condition-specific covariation relationship and the reference benchmark, as well as the data proportion of each data subset in the unified quantitative indicator set. This integration method comprehensively considers the specific covariation patterns under different conditions and their representativeness in the overall population, and refers to the covariation characteristics of the population under standard conditions, thereby overcoming the inaccuracy brought about by simply calculating the overall covariation relationship or ignoring the differences in conditions. The initial covariation relationship obtained in this way can more accurately reflect the overall covariation characteristics of the cucumber population under different growth stages and environmental conditions, and provide a more reliable reference for supplementary adjustment operations in areas that fail to fully represent the cucumber population. When the adjusted quantitative indicator set is used for data dimensionality reduction processing, it can enable the identified resistance components to reflect a more comprehensive resistance action pattern of the cucumber population.
[0058] As an embodiment of the present invention, the step of determining a reference benchmark for characterizing the covariation characteristics of the cucumber population under reference conditions is achieved by one of the following methods:
[0059] A reference benchmark was constructed based on the pre-defined expected covariation pattern among the indicators characterizing cucumber under the reference condition.
[0060] Alternatively, one or more groups of representative cucumber materials that exhibit specific properties under reference conditions are selected, and a unified set of quantitative indicators corresponding to the representative cucumber materials is obtained. Based on the obtained unified set of quantitative indicators corresponding to the representative cucumber materials, the common covariation characteristics between the quantitative indicators are analyzed and determined, and the determined common covariation characteristics are used as a reference benchmark.
[0061] Among them, the reference benchmark refers to a set of covariation characteristics used to provide a stable reference point, which can be represented by a covariance matrix, a set of correlation coefficients or a statistical model, and its purpose is to provide a standard for subsequent covariation relationship evaluation and adjustment; characterizing the expected covariation pattern between indicators of cucumber under reference conditions refers to the description of the correlation between various quantitative indicators of cucumber under specific reference conditions based on prior knowledge or theoretical models, which can be defined based on biological pathway models or expert experience, and its purpose is to capture ideal or theoretical indicator associations; constructing a reference benchmark refers to the process of establishing a reference benchmark based on a preset expected covariation pattern or extracting information from representative material data, which may include parameter setting or statistical calculations, Its purpose is to convert expected patterns or common characteristics into usable reference standards; representative cucumber materials with specific properties refer to cucumber material samples that exhibit specific biological or agronomic properties under reference conditions, which may include known resistant or susceptible varieties, materials with specific genetic backgrounds or materials obtained through preliminary screening. Its purpose is to provide actual data reflecting specific traits; common covariation characteristics between quantitative indicators refer to the association patterns between quantitative indicators that are commonly found among these materials through statistical analysis in the selected representative cucumber material data. It can be determined by correlation analysis, principal component analysis or cluster analysis. Its purpose is to extract representative indicator relationships from actual data.
[0062] The solution of this application solves the problem of reference benchmark selection by providing two flexible approaches to determining a reference benchmark. The first approach constructs a reference benchmark based on a pre-determined expected covariation pattern. This leverages existing biological knowledge or expert experience and allows for the rapid establishment of a theoretically ideal covariation standard. The second approach identifies shared covariation characteristics as a reference benchmark by analyzing actual data from representative cucumber materials exhibiting specific properties under reference conditions. This makes the reference benchmark more objective and reflects the resistance performance of actual cucumber materials. The reference benchmarks determined by these two approaches, combined with the condition-specific covariation relationships of each data subset obtained in the previous step, can more accurately assess the representativeness and importance of the associations between indicators under different growth stages or environmental conditions. Based on this more accurate assessment, the condition-specific covariation relationships can be effectively integrated to generate a more reliable initial covariation relationship that better represents the overall characteristics of the cucumber population. This reliable initial covariation relationship provides a stable reference for subsequent supplementary adjustments in data-deficient areas, ensuring that the adjusted data maintains a reasonable covariation structure while more comprehensively reflecting the diversity of the cucumber population, thereby enabling the identified resistance components to more accurately characterize the disease resistance mode of action of the cucumber.
[0063] As an embodiment of the present invention, the step of selecting one or more groups of representative cucumber materials that exhibit specific properties under reference conditions includes:
[0064] Under reference conditions, the candidate cucumber materials were phenotypically screened for target traits;
[0065] Based on the phenotypic screening results, representative cucumber materials showing specific properties were selected.
[0066] Among them, reference conditions refer to environmental conditions with clear control parameters set for phenotypic evaluation of cucumber materials, which may include specific temperature, humidity, light intensity, soil type, nutrient supply level, and pathogen inoculation method and concentration, etc. Its purpose is to ensure that the phenotypic performance of different cucumber materials is obtained under a unified and repeatable environmental background, thereby reducing the impact of environmental variation on phenotypic evaluation results and improving the accuracy of screening; candidate cucumber materials refer to cucumber plants or varieties with potential research or utilization value obtained from cucumber germplasm resource banks, breeding populations or natural populations. The number can be large, covering a wide range of variation in the target traits of the cucumber population, with the purpose of providing sufficiently diverse selection objects for subsequent phenotypic screening; target traits refer to biological characteristics that are directly or indirectly related to disease resistance and are of concern in cucumber disease resistance breeding analysis. They may include resistance levels to specific pathogens, lesion development speed, disease progression curve, defense response intensity, etc. Its purpose is to distinguish the differences in disease resistance of different cucumber materials by evaluating these traits; phenotypic screening refers to the use of observation, measurement or testing under specific conditions to evaluate the disease resistance of different cucumber materials. The process of measuring the phenotypic traits of cucumber materials and identifying those that meet specific requirements based on preset standards or thresholds. This process can use a variety of technical means such as manual visual inspection, image analysis, and physiological and biochemical testing. Its purpose is to preliminarily screen out materials that show differences or reach specific levels in the target traits from a large number of candidate materials. Specific performance refers to the level or pattern of biological characteristics of the target traits exhibited by cucumber materials under reference conditions that are representative or of research value. This can include different resistance levels such as high resistance, moderate resistance, and susceptible, or specific defense response types or intensities. The purpose is to select materials that can represent the key variation characteristics of the cucumber population in the target traits in order to construct a representative reference benchmark. Representative cucumber materials refer to a small number or multiple materials further selected from materials that pass the phenotypic screening and can reflect the variation range and characteristics of the target traits of the cucumber population under reference conditions. Their selection can be based on factors such as the distribution of phenotypic data, the diversity of genetic background, or the degree of deviation from the population average. The purpose is to use the data of these materials to construct a reference benchmark, thereby characterizing the covariation characteristics of the cucumber population under reference conditions.
[0067] The present invention employs a phenotypic screening of candidate cucumber materials for target traits under reference conditions to initially identify materials that differ in target traits from a broad pool of candidate materials. Screening under reference conditions ensures the comparability and reliability of phenotypic data and reduces interference from environmental factors. Subsequently, based on the screening results, representative cucumber materials exhibiting specific properties are selected. These selected materials are representative for the target traits and can reflect the variation characteristics of the cucumber population for these traits. By constructing a reference benchmark using data from these representative materials, the covariation characteristics of the cucumber population under reference conditions can be more accurately characterized. This process provides a solid foundation for subsequently determining the initial covariation relationships between quantitative indicators within a unified set of quantitative indicators, as the accuracy of the reference benchmark directly impacts the effectiveness of the integration of condition-specific covariation relationships. By selecting representative materials, the reference benchmark ensures that the core covariation patterns of the cucumber population for key traits are captured, making the resulting initial covariation relationships more realistic for the cucumber population, thereby optimizing subsequent disease resistance breeding analyses, such as the identification of resistance components and gene association analysis.
[0068] As an embodiment of the present invention, under reference conditions, the step of phenotypic screening of target traits of candidate cucumber materials includes:
[0069] Collect phenotypic images of candidate cucumber materials;
[0070] Analyze phenotypic images to quantitatively evaluate the phenotypes of target traits of candidate cucumber materials, and perform phenotypic screening of target traits based on the quantitative evaluation results.
[0071] Among them, phenotypic image refers to the visual record of the external manifestation of the target traits of cucumber materials under specific reference conditions. It can be obtained by high-resolution cameras, multispectral imaging equipment or three-dimensional scanners, etc. Its purpose is to objectively and comprehensively capture the phenotypic information of the cucumber materials such as morphology, color, texture, etc.; analyzing phenotypic images refers to the use of computer vision or image processing technology to extract characteristic information related to the target traits from the collected phenotypic images, which may include image preprocessing, image segmentation, and feature extraction. Its purpose is to convert the visual information in the image into computable and analyzable data; quantitative evaluation refers to the conversion of the target trait performance of the cucumber material into numerical indicators based on the characteristic information extracted by image analysis through a preset algorithm or model. It can be by calculating the ratio of the lesion area to the total leaf area as the disease index, or by measuring the size, color intensity, etc. of specific parts. Its purpose is to convert subjective observations into objective and accurate numerical values to facilitate subsequent comparison and screening.
[0072] The solution of this application acquires visual phenotypic information of candidate cucumber materials under reference conditions by collecting phenotypic images. Image acquisition is used because images can record rich phenotypic details in a non-contact, high-throughput manner, overcoming the subjectivity and limitations of manual observation. The acquisition of objective, detailed phenotypic image data enables subsequent in-depth processing of these images using image analysis techniques to extract various features related to the target trait, such as the morphology, color, and distribution of lesions. Based on these extracted features, a quantitative assessment of the target trait can be performed, converting complex phenotypic manifestations into precise numerical indicators, such as disease index and number of lesions. This quantitative assessment process avoids the ambiguity and uncertainty of manual assessment, allowing phenotypic differences between different materials to be accurately captured and compared. The acquisition of quantitative phenotypic assessment results enables scientific and objective phenotypic screening based on these numerical values. By setting thresholds or sorting, representative cucumber materials that meet specific performance requirements can be quickly and accurately identified. This screening method based on image analysis and quantitative evaluation improves the efficiency and accuracy of screening, provides more reliable input for subsequent breeding analysis, and thus helps to more accurately determine the reference benchmark, thereby improving the effectiveness of the entire cucumber disease resistance breeding analysis method.
[0073] As an embodiment of the present invention, the steps of performing data dimensionality reduction based on the adjusted set of quantitative indicators to identify multiple resistance components, each of which represents a resistance action mode defined by a combination of microscopic phenotypic data, include:
[0074] Obtain the adjusted quantitative index set, which is a matrix X_adj with n_adj rows and p columns, where n_adj is the adjusted number of cucumber material samples and p is the number of uniformly quantified microscopic phenotypic indicators;
[0075] Based on the matrix X_adj, the covariation structure between the quantitative indicators in the quantitative indicator set is analyzed to identify k potential resistance action modes that jointly affect the combination of microscopic phenotypic data, where k is the number of identified resistance components, and k <p;
[0076] According to the identified resistance action mode, the composition of each resistance component defined by the combination of microscopic phenotypic data is determined, and the mapping relationship between microscopic phenotypic data and resistance components is quantified to generate a loading matrix L_fixed with p rows and k columns. The element l_jk of the loading matrix L_fixed represents the association strength of the j-th microscopic phenotypic indicator on the k-th resistance component.
[0077] Based on the covariance characteristics of the load matrix L_fixed and the matrix X_adj, a fixed transformation rule matrix W_fixed with p rows and k columns is calculated and determined. W_fixed is used to convert the unified quantitative index set of any cucumber material into its quantitative score on each resistance component, that is, the resistance component characteristic vector S with n rows and k columns, where n is the number of samples in any batch of cucumber materials, and S=X*W_fixed.
[0078] The scheme of the present application lays a data foundation for the subsequent identification and quantification of resistance components by obtaining the adjusted quantitative index set matrix X_adj generated by the supplementary adjustment operation. Since X_adj is adjusted, it can better represent the overall characteristics of the cucumber population, and the analysis based on it is more representative. On this basis, the scheme further analyzes the covariation structure between quantitative indicators based on the matrix X_adj, rather than just performing a simple dimensionality reduction. By analyzing the covariation structure, k potential resistance action patterns hidden in the complex micro-phenotypic data that jointly affect the combination of micro-phenotypic data can be identified. These patterns are a reflection of the inherent structure of the data and have potential biological significance. Based on the identified resistance action patterns, the scheme determines the specific composition of each resistance component, and quantifies the mapping relationship between micro-phenotypic data and resistance components by generating the load matrix L_fixed. The load matrix L_fixed clearly shows the correlation strength of each micro-phenotypic indicator on different resistance components, making the resistance components more interpretable. Ultimately, based on the covariation properties of the load matrix L_fixed and the matrix X_adj, the scheme calculated and determined a fixed transformation rule matrix W_fixed with p rows and k columns. This W_fixed represents a stable transformation model from microphenotypic data to resistance component scores. Using the formula S = X * W_fixed, the unified quantitative index set X of any cucumber material (including new data not used for training) can be converted into its resistance component feature vector S. Because W_fixed is determined based on the more representative adjusted data X_adj and its covariation properties, it exhibits greater stability and generalizability than rules derived from unadjusted data or simple dimensionality reduction methods. This method, which analyzes the covariation structure of adjusted data and determines fixed transformation rules, effectively addresses the challenge of accurately extracting key resistance information and establishing stable mapping relationships from high-dimensional and complex microphenotypic data. This makes the identified resistance components more biologically meaningful, and the transformation rules can be reliably applied to new batches of cucumber material, achieving unified resistance assessment.
[0079] like Figure 2 A disease resistance gene typing result analysis system is shown, which is used for cucumber disease resistance breeding analysis. The system includes:
[0080] The data acquisition and conversion module 201 is used to acquire multi-source heterogeneous microscopic phenotypic data for cucumber disease resistance breeding analysis and convert the multi-source heterogeneous microscopic phenotypic data into a unified set of quantitative indicators;
[0081] The resistance component identification and feature vector generation module 202 is used to identify multiple resistance components based on a unified set of quantitative indicators through data dimensionality reduction processing, each resistance component representing a resistance action mode defined by a combination of microscopic phenotypic data, and generate a resistance component feature vector containing the quantitative score of each resistance component for each cucumber material;
[0082] Gene association analysis module 203, used to obtain genotyping data of cucumber materials, and perform a first association analysis based on the genotyping data and the resistance component feature vector to identify gene information related to a specific resistance component;
[0083] The macro-phenotype association analysis module 204 is used to obtain the macro-disease resistance phenotype data of the cucumber material and perform a second association analysis based on the resistance component feature vector and the macro-disease resistance phenotype data to evaluate the effects of different resistance components on the macro-disease resistance phenotype.
[0084] The basic principles, main features, and advantages of the present invention are shown and described above. Those skilled in the art should understand that the present invention is not limited to the above embodiments. The above embodiments and descriptions only describe the principles of the present invention. Various changes and improvements are possible without departing from the spirit and scope of the present invention, and such changes and improvements fall within the scope of the invention as claimed.
Claims
1. A method for analyzing disease resistance gene typing results, applied to cucumber disease resistance breeding analysis environment, characterized in that: The method comprises the following steps: Acquiring multi-source heterogeneous microscopic phenotypic data for cucumber disease resistance breeding analysis, and converting the multi-source heterogeneous microscopic phenotypic data into a unified set of quantitative indicators; Based on the unified set of quantitative indicators, multiple resistance components are identified through data dimensionality reduction processing, each of which represents a resistance action mode defined by a combination of microscopic phenotypic data, and a resistance component feature vector containing its quantitative score on each resistance component is generated for each cucumber material; Acquiring genotyping data of the cucumber material, and performing a first association analysis based on the genotyping data and the resistance component feature vector to identify gene information related to a preset resistance component; The macroscopic disease resistance phenotype data of the cucumber material is obtained, and a second association analysis is performed based on the resistance component characteristic vector and the macroscopic disease resistance phenotype data to evaluate the effects of different resistance components on the macroscopic disease resistance phenotype.
2. The method for analyzing disease resistance gene typing results according to claim 1, wherein: The steps of identifying multiple resistance components based on the unified set of quantitative indicators through data dimensionality reduction processing, each of the resistance components representing a resistance action mode defined by the combination of the microscopic phenotypic data, and generating a resistance component feature vector containing the quantitative score of each resistance component for each cucumber material include: Performing data dimensionality reduction based on the unified set of quantitative indicators of at least one initial batch of cucumber material to identify a plurality of resistance components, each of the resistance components representing a resistance action mode defined by the combination of the microscopic phenotypic data, and determining a fixed conversion rule for generating a characteristic vector of the resistance component from the unified set of quantitative indicators based on the resistance component; The fixed conversion rule is applied to the unified set of quantitative indicators of the at least one initial cucumber material batch and the subsequently obtained cucumber material batches to generate the resistance component feature vector for each cucumber material, which contains its quantitative score on each resistance component.
3. The method for analyzing disease resistance gene typing results according to claim 2, wherein: The step of performing data dimensionality reduction processing based on the unified set of quantitative indicators of at least one initial batch of cucumber materials to identify multiple resistance components, each of which represents a resistance action mode defined by the combination of the microscopic phenotypic data, comprises: Obtaining reference information characterizing overall characteristics of a cucumber population, wherein the reference information includes at least one of germplasm resource diversity distribution information and key micro-phenotype expected variation characteristics of the cucumber population; comparing the unified set of quantitative indices for the at least one initial batch of cucumber material with the reference information to identify areas of the cucumber population that are underrepresented in the unified set of quantitative indices; For the area that fails to fully represent the cucumber population, performing a supplementary adjustment operation on the unified quantitative index set to generate an adjusted quantitative index set, wherein the supplementary adjustment operation enables the adjusted quantitative index set to, when subsequently used in data dimensionality reduction processing, enable the identified resistance components to reflect a more comprehensive resistance action pattern of the cucumber population; Data dimensionality reduction is performed based on the adjusted quantitative indicator set to identify the multiple resistance components, each of which represents a resistance action mode defined by the combination of the microscopic phenotypic data.
4. The method for analyzing disease resistance gene typing results according to claim 3, wherein: The step of performing a supplementary adjustment operation on the unified quantitative index set for the area that fails to fully represent the cucumber population to generate an adjusted quantitative index set includes: Determining an initial covariance relationship between quantitative indicators in the unified quantitative indicator set; In the process of performing the supplementary adjustment operation on the unified quantitative indicator set to change the data composition or value of the unified quantitative indicator set for the area that fails to fully represent the cucumber population, evaluating the change caused by the supplementary adjustment operation to the initial covariation relationship; When the supplementary adjustment operation causes the change of the covariance relationship between the quantitative indicators relative to the initial covariance relationship to exceed the preset conditions, the execution method or execution range of the supplementary adjustment operation is adjusted so that the covariance relationship in the adjusted quantitative indicator set generated based on the adjusted supplementary adjustment operation meets the preset conditions.
5. The method for analyzing disease resistance gene typing results according to claim 4, characterized in that: The step of determining the initial covariance relationship between the quantitative indicators in the unified quantitative indicator set includes: Obtaining growth stage information or environmental condition information corresponding to each cucumber material in the unified quantitative index set; Dividing the unified quantitative indicator set into a plurality of data subsets according to the growth stage information or the environmental condition information, each data subset corresponding to a preset growth stage or environmental condition; For each of the data subsets, calculating the condition-specific covariation relationship between the quantitative indicators within the data subset; determining a reference benchmark, wherein the reference benchmark is used to characterize the covariation characteristics of the cucumber population under reference conditions; Based on the degree of correlation between the condition-specific covariation relationship of each data subset and the reference benchmark, as well as the data proportion of each data subset in the unified set of quantitative indicators, the condition-specific covariation relationships are integrated to generate the initial covariation relationship between the quantitative indicators in the unified set of quantitative indicators.
6. The method for analyzing disease resistance gene typing results according to claim 5, characterized in that: The step of determining a reference benchmark for characterizing the covariation characteristics of the cucumber population under reference conditions is achieved by one of the following methods: constructing the reference benchmark based on a predetermined expected covariation pattern among the indicators of the cucumber under the reference conditions; Alternatively, one or more groups of representative cucumber materials that exhibit preset performance under the reference conditions are selected, and the unified set of quantitative indicators corresponding to the representative cucumber materials are obtained. Based on the obtained unified set of quantitative indicators corresponding to the representative cucumber materials, the common covariation characteristics between the quantitative indicators are analyzed and determined, and the determined common covariation characteristics are used as the reference benchmark.
7. A method for analyzing disease resistance gene typing results according to claim 6, characterized in that: The step of selecting one or more groups of representative cucumber materials that exhibit predetermined performance under the reference conditions comprises: Under the reference conditions, phenotypic screening of target traits is performed on candidate cucumber materials; Based on the phenotypic screening results, representative cucumber materials that exhibit the preset properties are selected.
8. The method for analyzing disease resistance gene typing results according to claim 7, wherein: The step of performing phenotypic screening of target traits on candidate cucumber materials under the reference conditions comprises: collecting phenotypic images of the candidate cucumber material; The phenotypic image is analyzed to quantitatively evaluate the phenotype of the target trait of the candidate cucumber material, and phenotypic screening of the target trait is performed based on the quantitative evaluation result.
9. The method for analyzing disease resistance gene typing results according to claim 3, wherein: The step of performing data dimensionality reduction processing based on the adjusted quantitative indicator set to identify the multiple resistance components, each of which represents a resistance action mode defined by the combination of the microscopic phenotypic data, comprises: Obtaining the adjusted quantitative index set, the adjusted quantitative index set is a matrix X_adj with n_adj rows and p columns, where n_adj is the adjusted number of cucumber material samples, and p is the number of uniformly quantified microscopic phenotypic indicators; Based on the matrix X_adj, the covariation structure between the quantitative indicators in the quantitative indicator set is analyzed to identify k potential resistance action modes that jointly affect the combination of microscopic phenotypic data, where k is the number of identified resistance components, and k <p; According to the identified resistance action mode, determining the composition mode of each resistance component defined by the combination of the microscopic phenotypic data, and quantifying the mapping relationship between the microscopic phenotypic data and the resistance component to generate a load matrix L_fixed with p rows and k columns, wherein the element l_jk of the load matrix L_fixed represents the association strength of the j-th microscopic phenotypic indicator on the k-th resistance component; Based on the covariance characteristics of the load matrix L_fixed and the matrix X_adj, a fixed conversion rule matrix W_fixed with p rows and k columns is calculated and determined. The W_fixed is used to convert the unified quantitative index set of any cucumber material into its quantitative score on each of the resistance components, that is, the resistance component characteristic vector S with n rows and k columns, where n is the number of samples in any batch of cucumber materials, and S=X*W_fixed.
10. A disease resistance gene typing result analysis system for cucumber disease resistance breeding analysis, characterized in that: The system includes: A data acquisition and conversion module is used to acquire multi-source heterogeneous microscopic phenotypic data for cucumber disease resistance breeding analysis and convert the multi-source heterogeneous microscopic phenotypic data into a unified set of quantitative indicators; a resistance component identification and feature vector generation module, configured to identify multiple resistance components based on the unified set of quantitative indicators through data dimensionality reduction processing, wherein each resistance component represents a resistance action mode defined by a combination of microscopic phenotypic data, and generate a resistance component feature vector for each cucumber material containing its quantitative score on each resistance component; a gene association analysis module, configured to obtain genotyping data of the cucumber material, and perform a first association analysis based on the genotyping data and the resistance component feature vector to identify gene information associated with a preset resistance component; The macro-phenotype association analysis module is used to obtain the macro-disease resistance phenotype data of the cucumber material, and perform a second association analysis based on the resistance component characteristic vector and the macro-disease resistance phenotype data to evaluate the effects of different resistance components on the macro-disease resistance phenotype.
Citation Information
Patent Citations
Lemon disease-resistant gene map construction method based on data mining
CN118335187A
Grape downy mildew resistance whole genome selective breeding method based on machine learning
CN120126571A