A bidirectional association analysis method for plant-pathogen gene interactions and its application

Through the two-way whole genome association analysis method, candidate genes and factors are identified from wheat disease-resistant genes and powdery mildew effector effectors, and the problem of difficulty in screening wheat disease-resistant genes and effector factors is solved, and an efficient screening and analytical interaction mechanism is achieved to support disease-resistant breeding.

CN120108498BActive Publication Date: 2025-09-02INST OF GENETICS & DEVELOPMENTAL BIOLOGY CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510172093.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-17
Publication Date
2025-09-02
Estimated Expiration
2045-02-17

AI Technical Summary

Technical Problem

The study of the interaction mechanism between wheat and powdery mildew in the prior art has unclear identification rules for disease-resistant genes and effectors, and traditional strategies are difficult to efficiently screen potential disease-resistant genes and effectors, resulting in insufficient mining of disease-resistant genes.

Method used

Two-way whole genome association analysis method was used to conduct correlation analysis from the two directions of plant disease-resistant genes and pathogenic effector factors. Candidate peaks were located through the whole genome association model, candidate disease-resistant genes and effector factors were identified, and interaction relationships were determined based on haplotype analysis.

Benefits of technology

Efficiently and accurately screen out disease-resistant genes in plants and effector factors in pathogenic bacteria, analyze the interaction mechanism, and provide a basis for cultivating new varieties of lasting broad-spectrum disease-resistant plants.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108498B_ABST
    Figure CN120108498B_ABST
Patent Text Reader

Abstract

The present invention discloses a bidirectional association analysis method for plant-pathogen gene interactions and its application, belonging to the field of biotechnology. The method comprises the following steps: (1) obtaining plant genetic variation information, pathogen genetic variation information, and the plant's resistance-susceptibility phenotype to the pathogen; (2) performing forward whole-genome association analysis to locate the pathogen effector from the plant's disease-resistance genes; (3) performing reverse whole-genome association analysis to locate the plant's disease-resistance genes from the pathogen effector; (4) the disease-resistance genes and effector obtained from the forward whole-genome association analysis and the reverse whole-genome association analysis are identical, and the two are the gene pairs interacting between the plant and the pathogen. The method provided by the present invention can efficiently and accurately screen out the gene pairs interacting between plant disease-resistance genes and pathogen effector factors simultaneously, providing strong support for analyzing the resistance mechanism exerted by plant disease-resistance genes and laying the foundation for breeding new plant varieties with long-lasting and broad-spectrum disease resistance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of biotechnology, and in particular relates to a bidirectional correlation analysis method for plant and pathogen gene interactions and its application. Background Art

[0002] Obligate parasitic pathogens of wheat powdery mildew Blumeria graminis f. sp. tritici Powdery mildew (Bgt) is a major fungal disease that harms wheat production. Widespread outbreaks of wheat powdery mildew can lead to severe yield losses. Breeding and disseminating disease-resistant wheat varieties is considered the most cost-effective way to control powdery mildew.

[0003] Wheat Nucleotide-binding Leucine-rich Repeat (NLR) disease resistance proteins play a key role in disease resistance. The remarkable diversity of NLRs provides a broad opportunity for the exploration of disease resistance gene resources. Studies have shown that the number of NLR genes varies significantly among different wheat varieties. For example, in the common wheat variety Chinese Spring, approximately 3,400 NLR genes have been identified, of which approximately half are functional, expressed, and have complete open reading frames (Steuernage et al., 2020). However, high-quality reference genome sequences of 10 core wheat varieties assembled through the 10+ Genomes Project reveal that only approximately one-third of these NLR genes are shared across varieties, with the remainder being specific. Based on this, it is estimated that the total number of NLR genes in wheat populations may be as high as 7,000, while less than 10% of these wheat disease resistance genes have been identified.

[0004] Although research on NLR disease-resistance genes has made some progress, there are still two major bottlenecks in the interaction mechanism between wheat and powdery mildew: first, the specific recognition rules between disease-resistance genes and effector factors have not been fully understood; second, the study of effector factor cloning is limited by traditional strategies, that is, the corresponding disease-resistance genes must be obtained before reverse screening for their corresponding non-toxic factors. The currently reported powdery mildew effector factors mainly include AvrPm1a (Hewitt et al., 2021), AvrPm2 (Praz et al., 2017), AvrPm3a (Bourras et al., 2015), AvrPm3b2 / c2 and AvrPm3d3 (Bourras et al., 2019) and AvrPm17 (Mueller et al., 2021). These effectors are Pm1a 、 Pm2 、 Pm3a / 3f 、 Pm3b / c 、 Pm3d3 and Pm17 The specific identification of disease-resistance genes such as β-lactamase and β-lactamase fully confirms the strict applicability of the "gene-for-gene" hypothesis in the wheat-powdery mildew interaction system.

[0005] Although wheat NLR genes confer disease resistance that relies on specific recognition of pathogen effectors, a large number of potential disease-resistance genes and effectors remain undiscovered, and their interaction mechanisms are far from fully understood. According to the "gene-for-gene" hypothesis proposed by Flor in 1956, wheat disease-resistance genes specifically recognize effectors secreted by powdery mildew fungi, triggering an immune response that causes cell necrosis at the site of infection and prevents the spread of the pathogen. However, powdery mildew fungi evade recognition by rapidly mutating to produce new effectors, while wheat regains recognition capability by evolving new disease-resistance genes. This has led to a long-term "arms race" between the two, resulting in continuous co-evolution. The diversity and regulatory mechanisms of this recognition relationship remain unresolved, necessitating the development of high-throughput, rapid methods for discovering disease-resistance genes and their corresponding effectors.

[0006] Pm3 It is the first cloned wheat powdery mildew resistance gene, and its resistance mechanism has been studied most intensively. Pm3 Alleles ( Pm3a to Pm3g and Pm3k to Pm3t ) and two orthologous genes ( Pm8 and Pm17 ), their sequence similarity exceeds 97%, and both show resistance to specific powdery mildew strains (Bhullar et al., 2010; Stirnweis et al., 2014). Studies have shown that the strain-specific resistance of these alleles is genetically regulated by a combination of multiple effector factors and their inhibitors (Bourras et al., 2015). Further studies have found that haplotype variation in effector proteins in natural populations significantly affects recognition ability. For example, changes in a single amino acid in the AVRPM3A2 / F2 protein may destroy specific recognition, and only changes in a few amino acids can enhance the immune response (McNally et al., 2018). Studies have pointed out that, Pm3Strain-specific resistance of alleles depends on specific recognition of the structure, rather than the sequence, of the effector protein. Even with significant sequence divergence, similarities in protein structure can still support functional recognition (Bourras et al., 2019). This suggests that changes in protein structure, triggered by sequence variation between haplotypes, are important determinants of specific recognition and immune responses, thus highlighting the critical role of natural variation in genetic resistance mechanisms. Therefore, different haplotypes of the same gene play an important role in strain-specific resistance, and variation between haplotypes significantly influences the specific recognition of the effector by the resistance gene. Summary of the Invention

[0007] In response to the above-mentioned prior art, the present invention provides a bidirectional correlation analysis method and application of the interaction between plant and pathogen genes, which solves the problem of difficulty in screening plant disease-resistant genes and pathogen effector factors in the prior art.

[0008] In order to achieve the above object, the technical solution adopted by the present invention is to provide a bidirectional correlation analysis method for the interaction between plant and pathogen genes, comprising the following steps:

[0009] (1) Obtain information on plant genetic variation, pathogen genetic variation, and plant resistance and susceptibility phenotypes to pathogens;

[0010] (2) Forward genome-wide association analysis to locate pathogen effector factors based on plant disease resistance genes:

[0011] ① Perform genome-wide association analysis on the genetic variation information of plants and the resistance and susceptibility phenotypes of plants to pathogens, obtain the association results between the variation sites on the plant genome and the resistance and susceptibility phenotypes of plants to pathogens, draw Manhattan plots and identify candidate peaks;

[0012] ② Identify candidate disease resistance genes in plants from candidate peaks;

[0013] ③ Combined with the plant's resistance phenotype to pathogens, identify the disease resistance haplotype of the candidate disease resistance gene;

[0014] ④ Using the pathogen's infection phenotype on plants with disease-resistant haplotypes containing candidate disease-resistance genes as a trait, perform genome-wide association analysis on the trait and the pathogen's genetic variation information to locate candidate effector factors on the pathogen's genome;

[0015] (3) Reverse genome-wide association analysis to locate plant disease resistance based on pathogen effector factors:

[0016] ① Perform genome-wide association analysis on the genetic variation information of pathogens and the resistance and susceptibility phenotypes of plants to pathogens, obtain the association results between the variant sites on the pathogen genome and the resistance and susceptibility phenotypes of plants to pathogens, draw a Manhattan plot and identify candidate peaks;

[0017] ② Identify candidate effector factors in pathogens from candidate peaks;

[0018] ③ Combined with the plant's resistance-susceptibility phenotype to pathogens, identify the avirulent haplotype of the candidate effector;

[0019] ④ Using the plant's susceptible phenotype to a non-toxic haplotype containing a candidate effector as a trait, perform genome-wide association analysis on the trait and the plant's genetic variation information to locate candidate disease resistance genes on the plant genome;

[0020] (4) The candidate disease resistance gene obtained in step (2) and step (3) is the same as the candidate effector factor, and the two are the gene pairs that interact between plants and pathogens.

[0021] On the basis of the above technical solution, the present invention can also be improved as follows.

[0022] Furthermore, the genetic variation information of plants is single nucleotide polymorphism, and the genetic variation information of pathogens is single nucleotide polymorphism.

[0023] Furthermore, in step ① of the forward whole-genome association analysis and step ① of the reverse whole-genome association analysis, whole-genome association analysis was performed using association models MLM (mixed linear model), FarmCPU (fixed and random model improvement), MLMM (multivariate linear model), and BLINK (Bayesian information and linkage disequilibrium iterative nested key-slot), and at least three association models located the same peak region in the Manhattan plot as candidate peaks.

[0024] Furthermore, the specific operation of step ② of the forward genome-wide association analysis is as follows: check the functional annotations of genes near the significant peak in the Manhattan plot, and search for candidate disease resistance genes within 1 Mb upstream and downstream of the significant peak;

[0025] Candidate peaks must meet the following conditions simultaneously: (a) the p-value of the SNP within the candidate peak is less than the significance threshold of 1e-8; (b) the distance between two adjacent SNPs that meet the previous condition must be less than 1Mb; (c) the number of SNPs that meet the first two conditions in the candidate peak must be greater than 5; (d) considering linkage disequilibrium, the r² value must be greater than 0.8;

[0026] Candidate disease-resistance genes meet the following conditions simultaneously: (a) the gene function annotation is related to immune response, pathogen recognition or defense mechanism; (b) the gene has homologous genes in the known disease-resistance gene database; (c) the gene expression level in disease-resistant varieties is higher than that in susceptible varieties.

[0027] Furthermore, the specific operation of step ② of the reverse genome-wide association analysis is to check the functional annotations of genes near the significant peak in the Manhattan plot and search for candidate effector factors within 100 kb upstream and downstream of the significant peak;

[0028] Candidate peaks must meet the following conditions simultaneously: (a) the p-value of the SNP within the candidate peak is less than the significance threshold of 1e-5; (b) the distance between two adjacent SNPs that meet the first condition must be less than 10 kb; (c) the number of SNPs that meet the first two conditions in the candidate peak must be greater than 10; (d) considering linkage disequilibrium, the r² value must be greater than 0.8;

[0029] The candidate effectors meet the following conditions simultaneously: (a) have a signal peptide; (b) have no transmembrane domain; (c) have a molecular weight less than 35 kDa; and (d) contain a specific YxC motif.

[0030] Furthermore, the specific operation of step ③ of the forward genome-wide association analysis is as follows: adjacent SNP sites on the candidate disease-resistance gene are strung together to form the haplotype of the candidate disease-resistance gene, and different haplotypes of the gene are associated with the plant's resistance phenotype to the pathogen using chi-square test, logistic regression, analysis of variance and / or Kruskal-Wallis test to evaluate the significance of the association between the haplotype and the phenotype, and select the disease-resistance haplotype of the candidate disease-resistance gene;

[0031] The specific operation of step ③ of the reverse whole-genome association analysis is as follows: adjacent SNP sites on the candidate effector are strung together to form the haplotype of the candidate effector, and the different haplotypes of the effector are associated with the resistance phenotype of the pathogen-infected plant using chi-square test, logistic regression, variance analysis and / or Kruskal-Wallis test to evaluate the significance of the association between the haplotype and the phenotype, and select the non-toxic haplotype of the candidate effector.

[0032] Furthermore, in step ③ of the forward genome-wide association analysis, logistic regression was used to evaluate the disease resistance haplotypes that were significantly associated with the resistance trait: the plant resistance phenotype was considered as a binary classification, with a phenotypic value <= 2.3 as resistance and a phenotypic value > 2.3 as non-resistance. Resistance and non-resistance were used as dependent variables, resistance was recorded as 1, non-resistance was recorded as 0, and haplotype was used as the independent variable; if the regression coefficient of the haplotype was positive and p < 0.05, the haplotype was considered a disease resistance haplotype;

[0033] In step ③ of the reverse genome-wide association analysis, logistic regression was used to evaluate the avirulent haplotypes that were significantly associated with the haplotypes of the candidate effector and the toxicity trait: the pathogen toxicity phenotype was regarded as a binary classification, with a phenotypic value <= 2.3 as avirulent and a phenotypic value > 2.3 as toxic. Toxicity and non-toxicity were used as dependent variables, with toxicity as 1 and non-toxicity as 0. The haplotype was used as the independent variable. If the regression coefficient of the haplotype was negative and p < 0.05, the haplotype was an avirulent haplotype.

[0034] Furthermore, in step ④ of the forward whole-genome association analysis and step ④ of the reverse whole-genome association analysis, whole-genome association analysis was performed using the association model MLM.

[0035] Furthermore, the above-mentioned bidirectional association analysis method for plant and pathogen gene interactions is used in screening interacting gene pairs.

[0036] Furthermore, the bidirectional association analysis method of plant-pathogen gene interaction was applied in screening wheat disease resistance genes and corresponding effector factors in powdery mildew.

[0037] The present invention has the following beneficial effects: The method provided by the present invention, based on bidirectional genomic association analysis, can efficiently and accurately screen for both disease-resistance genes in plants and corresponding effector factors in pathogens. Screening for these interacting gene pairs can provide strong support for understanding the resistance mechanisms exerted by plant disease-resistance genes and lay the foundation for the development of new plant varieties with durable and broad-spectrum disease resistance. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 This is a schematic diagram of a forward GWAS, which locates the interacting pathogen Avr based on the plant disease resistance R gene;

[0039] Figure 2 This is a schematic diagram of reverse GWAS, which locates interacting plant R genes starting from the pathogen effector Avr;

[0040] Figure 3 Schematic diagram for determining gene haplotype (resistance haplotype of R or avirulence haplotype of Avr); plant resistance gene PmX (or pathogen effector AvrPm) has 3 SNP sites, and these three SNPs form 9 haplotypes. a is the relationship between the haplotype heat map of the gene and the phenotype. Each row is a plant variety in the plant population, and each column on the left is a SNP that makes up the haplotype. The green color on the right indicates that the plant has a disease-resistant phenotype, and the red color indicates a susceptible phenotype. From the schematic diagram, we can see that the phenotype of samples with haplotypes "222" or "202" is mostly resistant, and b also shows that the phenotype of samples with these two haplotypes is resistant. Therefore, "222" and "202" are the disease-resistant (or avirulent) haplotypes of this R gene (or effector factor Avr).

[0041] Figure 4 Forward GWAS example, from the disease resistance gene Pm1a Effector factors were located in the disease-resistant haplotype of wheat AvrPm1a ;

[0042] Figure 5 For the reverse GWAS example, from the effect factor AvrPm1a Disease resistance genes were located in the avirulent haplotype strain Pm1a ;

[0043] Figure 6 Transient expression validation in tobacco Pm1a and AvrPm1a interaction; EV represents empty load, BAX represents positive control, scale bar: 1 cm. DETAILED DESCRIPTION

[0044] Data basis for bidirectional GWAS:

[0045] What is needed is ① the variation information file (SNPs) generated by the whole genome sequencing data of the plant population; ② the whole genome sequencing data of the pathogen population and its variation information file (SNPs); ③ the resistance phenotypic data of the plant population to the pathogen strain population.

[0046] The core idea of ​​bidirectional genome-wide association study (bidirectional GWAS):

[0047] Bidirectional genome-wide association studies (GWAS) refer to the process of locating the plant resistance gene (R gene) to the pathogen effector (Avr gene) it recognizes in the forward direction, and simultaneously locating the R gene in the reverse direction from the Avr gene. Using genotypic and phenotypic data, genome-wide association studies (GWAS) are used to locate and identify plant resistance genes (such as Avr genes) that are resistant to specific pathogen strains in plant populations. Figure 1 (a) Identification of candidate resistance genes PmX Afterwards, for PmX Different haplotypes and corresponding phenotypic relationships (such as Figure 3 a and Figure 3 b), the plant population is divided into two groups: one group is genetic PmX The disease resistance haplotype group; the other group is gene PmX The susceptible haplotype group is lacking in plants that can identify the susceptible haplotype. AvrPm The functional disease resistance gene haplotype of PmX Different plant individuals with the disease resistance haplotype group of the gene were used to locate the candidate disease resistance gene of this plant in the pathogen population. PmX Specific recognition of pathogen effector factors AvrPm When the plant with disease-resistant haplotype group can always fight against the effector in the pathogen population, AvrPm Establish significant associations (e.g. Figure 1 b), it indicates PmX and AvrPm There is a specific recognition relationship between them, which triggers the immune response and thus resists the infection of pathogens.

[0048] Similarly, the process is reversed, starting from the avirulent effector Avr to locate the plant R gene that interacts with it. First locate the candidate effector in the pathogen population AvrPm (like Figure 2 a), and then the pathogen strains were separated according to the different haplotypes and phenotypic relationships of this effector (as shown in Figure 3 a and Figure 3 b), the pathogen strains were divided into avirulent haplotype group and virulent haplotype group. When the avirulent haplotype group strains could locate the same PmX When (such as Figure 2 b), indicating AvrPm yes PmX Specific recognition effector.

[0049] Analyze from two directions (from resistance gene to avirulent effect factor and from avirulent effect factor to resistance gene) PmX and AvrPm A strict, bidirectional, and specific recognition relationship exists. This method can accelerate the identification of plant disease resistance genes and the analysis of their interactions with effector factors. This not only helps fill gaps in current research but also provides important theoretical foundations and technical support for uncovering the molecular mechanisms of plant disease resistance, improving disease resistance breeding strategies, and cultivating plant varieties with broad-spectrum, durable disease resistance.

[0050] The specific implementation methods of the present invention are described in detail below with reference to the embodiments.

[0051] Example 1

[0052] Detailed operation of the bidirectional GWAS analysis method for the interaction between wheat and powdery mildew genes:

[0053] 1. Data:

[0054] ①Genetic variation information (usually single nucleotide polymorphisms, SNPs) of 581 wheat populations, data source: https: / / ngdc.cncb.ac.cn / gvm / getProjectFile?t=ada7a2d8; ②Genotypic variation information (SNPs) of 120 powdery mildew strains, data source: https: / / ngdc.cncb.ac.cn / gvm / getProjectFile?t=2848772b; ③Resistance-susceptibility phenotypes of 581 wheat accessions to 120 powdery mildew strains, as shown in Tables 1 and 2.

[0055] Table 1 Phenotypic grade data of wheat population (first column) resistance to powdery mildew strains (first row)

[0056]

[0057]

[0058]

[0059]

[0060]

[0061]

[0062]

[0063] Note: In the table, level 0 indicates super resistance, level 1 indicates medium resistance, level 2 indicates weak resistance, level 3 indicates medium sensitivity, and level 4 indicates super sensitivity.

[0064] Table 2 Phenotypic grade data of wheat population (first row) resistance to powdery mildew strains (first column)

[0065]

[0066]

[0067] Note: In the table, level 0 indicates super resistance, level 1 indicates medium resistance, level 2 indicates weak resistance, level 3 indicates medium sensitivity, and level 4 indicates super sensitivity.

[0068] 2. Forward GWAS: Mapping the powdery mildew effector Avr from the wheat R gene

[0069] (1) Batch identification of candidate powdery mildew resistance genes in candidate wheat: Genome-wide association study (GWAS) was performed on the genotype variation files (SNPs) of the wheat population and the phenotypes of 120 strains of the wheat population. The models used in GWAS included MLM (mixed linear model), FarmCPU (fixed and random model improvement), MLMM (multivariate linear model), and BLINK (Bayesian information and linkage disequilibrium iterative nested key slot). The parameters were default parameters. The association results between the variation sites on the genome of the wheat population and the phenotype of each strain were obtained (the phenotype of each strain in the wheat population was used as a trait to obtain an association result, for a total of 120 association results). A Manhattan plot was drawn to show the association situation on the genome. The same peak area located by at least three association models was regarded as a candidate peak.

[0070] The peak on the Manhattan plot represents a statistically significant association between a specific chromosome region in wheat and the powdery mildew resistance phenotype. The higher the peak (the smaller the p-value or the larger the -log10(p-value)), the stronger the significance of the association. At the same time, the region where the peak is located often contains candidate genes related to the phenotype.

[0071] (2) Select candidate disease resistance genes that are significantly associated with the trait and pointed to by the candidate peak of interest: The significant region of the peak often contains candidate disease resistance genes, which may be related to disease resistance mechanisms, such as encoding disease resistance proteins (such as the NLR gene family), transcription factors that regulate gene expression, or genes involved in signal transduction and defense responses. By checking the functional annotations of genes near the significant peak, candidate disease resistance genes of interest related to disease resistance can be identified.

[0072] A candidate peak in the wheat population was considered significant if it met the following conditions simultaneously: (a) the p-value of the SNP within the candidate peak was less than the significance threshold of 1e-8; (b) the distance between two adjacent SNPs that met the first condition was less than 1Mb; (c) the number of SNPs in the candidate peak that met the first two conditions was greater than 5; (d) considering linkage disequilibrium, the r² value was greater than 0.8. Candidate disease resistance genes were searched within 1Mb upstream and downstream of the significant peak, and genes with disease resistance annotations were selected from 10 genes within and upstream and downstream of the significant peak as candidate disease resistance genes.

[0073] Candidate disease-resistance genes meet the following conditions simultaneously: (a) the gene function annotation is related to immune response, pathogen recognition or defense mechanism; (b) the gene has homologous genes in the known disease-resistance gene database; (c) the gene expression level in disease-resistant varieties is significantly higher than that in susceptible varieties.

[0074] In this example, GWAS association models MLM, MLMM, FarmCPU and BLINK (corresponding to Figure 4Model 1-4 in a) used the resistance phenotype of wheat population to powdery mildew strain B029-A1 as a trait to associate with the genetic variation of wheat population. The four association models simultaneously obtained the powdery mildew resistance gene corresponding to the significant peak at the end of chromosome 7A. Pm1a (like Figure 4 a), Pm1a It is a cloned powdery mildew resistance gene.

[0075] (3) Identify haplotypes that are significantly associated with disease resistance traits, i.e., disease resistance haplotypes: candidate disease resistance genes Pm1a The SNPs above can have multiple combinations in the population. Each combination of SNPs is a haplotype. Combined with the resistance-susceptibility phenotype of the wheat population to the B029A1 strain, a haplotype with a high correlation with the disease-resistant phenotype was identified.

[0076] The specific operation is: Pm1a The association analysis between different haplotypes and the resistance-susceptibility phenotype of wheat population to powdery mildew strain B029-A1 was performed, and the statistical method logistic regression was used to evaluate the Pm1a Disease-resistant haplotypes that are significantly associated with the resistance phenotype: The phenotypic data are regarded as binary categories (phenotypic value <= 2.3 is resistance, phenotypic value > 2.3 is non-resistance), haplotype is used as the independent variable, resistance is used as the dependent variable, resistance (1) and non-resistance (0), and binary logistic regression analysis is performed. The resistance haplotype is judged by the regression coefficient and its significance level. If the regression coefficient of the haplotype is significant (p-value < 0.05) and the regression coefficient is positive, then the haplotype is positively correlated with the resistance phenotype, that is, the disease-resistant haplotype.

[0077] These disease-resistant haplotypes may be Pm1a The coding or regulatory regions of disease-resistance genes carry key functional mutations (such as non-synonymous mutations, splice site mutations or regulatory enhancer mutations), which confer wheat resistance to the powdery mildew strain B029A1.

[0078] (4) Contains Pm1a Disease-resistant haplotypes of wheat can be mapped to candidate effector factors AvrPm1a :contain Pm1a The wheat varieties with disease-resistant haplotypes include Huzhuhong (number B049), CItr14087 (number B103), PI350766 (number B104), PI525317 (number B145), etc., and the powdery mildew populations were respectively Pm1aThe infection phenotype of disease-resistant haplotype wheat was used as a trait, and a genome-wide association study (GWAS) was performed between the trait and the genetic variation of powdery mildew strains through mixed linear model (MLM) analysis, which effectively identified genetic variations related to the disease resistance phenotype on the powdery mildew genome and located candidate effector factors on the powdery mildew genome.

[0079] The results showed that these Pm1a The wheat infection phenotypes of the resistant haplotypes all mapped to a common candidate effector on the powdery mildew genome AvrPm1a (like Figure 4 b). This means that Pm1a The resistance phenotypes of different wheat varieties with disease-resistant haplotypes to powdery mildew are all affected by the same effector AvrPm1a Regulation of disease resistance genes Pm1a and effect factors AvrPm1a The interaction between them is a key factor in wheat disease resistance.

[0080] 3. Reverse GWAS: Mapping wheat R genes based on the powdery mildew effector Avr

[0081] (1) Batch identification of candidate effectors in powdery mildew: Similar to the forward GWAS method, we first performed a genome-wide association study (GWAS) on the genetic variation files (SNPs) of the powdery mildew population and the infection phenotypes of 581 wheat varieties in the powdery mildew population. The models used in the GWAS were MLM, FarmCPU, MLMM, and BLINK, and the parameters were the default parameters. The association results between the SNPs in the powdery mildew population and the resistance-susceptibility phenotype of each wheat variety were obtained. Here, the phenotype of each wheat variety in the powdery mildew population was used as a trait to obtain an association result, for a total of 581 association results. A Manhattan plot was drawn for each association result to show the degree of association between the SNPs and the phenotype on the powdery mildew genome. The same peak area located by at least three association models was regarded as a candidate peak.

[0082] (2) Identify candidate Avr genes associated with the phenotype in the region where the peak is located: Effector genes are genes in the pathogen genome that encode small molecular proteins. These proteins can be secreted into host cells and interfere with the host's immune response or metabolic process. In powdery mildew, these genes are generally considered to be key factors that interact with host resistance genes, regulating the host's resistance-susceptibility phenotype directly or indirectly.

[0083] A candidate peak in a powdery mildew pathogen population is considered a significant peak if it meets the following conditions simultaneously: (a) the p-value of the SNP within the candidate peak is less than the significance threshold of 1e-5; (b) the distance between two adjacent SNPs that meet the first condition must be less than 10 kb; (c) the number of SNPs in the candidate peak that meet the first two conditions must be greater than 10; (d) considering linkage disequilibrium, the r² value must be greater than 0.8; candidate effect factors are searched within 100 kb upstream and downstream of the significant peak, and genes that meet the characteristics of effect factors among the 10 genes within and upstream and downstream of the significant peak are candidate effect factors;

[0084] Candidate effectors must meet the following conditions simultaneously: (a) have a signal peptide to ensure protein secretion; (b) have no transmembrane domain to exclude membrane-localized proteins; (c) have a molecular weight of less than 300 amino acids (approximately 30-35 kDa), which is consistent with the typical characteristics of effectors; and (d) contain a specific YxC motif, which is a conserved characteristic of effectors.

[0085] In this example, the resistance phenotype of wheat variety PI350766 (number B104) to powdery mildew was used as the trait, and the GWAS association models MLM, MLMM, FarmCPU, and BLINK (corresponding to Figure 5 Model 1-4 in a) was significantly associated with the effect factor AvrPm1a ( Figure 5 a), AvrPm1a It is an important cloned powdery mildew effector factor located on chromosome 6 of the powdery mildew genome. Its function and variation directly affect wheat's resistance to powdery mildew.

[0086] (3) Identify avirulent haplotypes of effectors that are significantly associated with traits: The avirulent haplotype (Avr) of the pathogen effector enables the pathogen's proteins or metabolites to be sensed by the host's disease resistance genes, thereby stimulating the host's immune response. In contrast, the virulent haplotype (Vir) changes the protein structure of the effector through mutations at key sites, allowing the pathogen to evade immune recognition and successfully infect the host. The avirulent haplotype is the key to the pathogen effector's ability to be recognized by the host's disease resistance genes and trigger a resistance response.

[0087] The specific operation is: use logistic regression to evaluate the non-toxic haplotype with significant association between the haplotype of the candidate effector and the toxic trait: regard the pathogen toxicity phenotype as a binary classification (phenotypic value <= 2.3 is non-toxic, phenotypic value > 2.3 is toxic), use haplotype as the independent variable, toxicity (1) and non-toxicity (0) as the dependent variables, and perform binary logistic regression analysis. If the regression coefficient of the haplotype is negative and p < 0.05, then the haplotype is a non-toxic haplotype.

[0088] Combined with successful positioning AvrPm1aThe resistance and susceptible phenotype of wheat material PI350766 to powdery mildew populations were investigated, and 120 powdery mildew strains were classified into effector factors. AvrPm1a Avirulent haplotypes and virulent haplotypes. The avirulent haplotypes (powdery mildew strains include B016A3, B096A1, B029A1, and B073A1) were all phenotypically incapable of infection with PI350766 (i.e., wheat was resistant to these strains). There was no significant correlation between the phenotype and genotype of the virulent haplotypes against PI350766.

[0089] (4) Contains AvrPm1a Strains with avirulent haplotypes can be located to the same AvrPm1a :contain AvrPm1a The non-toxic haplotypes of powdery mildew strains include B016A3, B096A1, B029A1, B073A1, etc. The resistance phenotypes of wheat populations to these strains were used as traits and were successfully identified through GWAS. AvrPm1a Disease resistance genes Pm1a ( Figure 5 b) Further clarification of pathogen effector factors AvrPm1a Host disease resistance genes Pm1a Specific recognition between.

[0090] Example 2

[0091] Tobacco transient expression experiment verifies disease resistance gene Pm1a and effector AvrPm1a Interaction

[0092] Verification of wheat disease resistance genes using tobacco transient expression system Pm1a Powdery mildew effector AvrPm1a In the experiment, we used Pm1a and AvrPm1a The Agrobacterium resuspension was used to prepare different Agrobacterium combinations and the different combinations of Agrobacterium were injected into the mesophyll of Nicotiana tabacum (leaves grown for one month). Pm1a or AvrPm1a When the leaves were annotated with empty vector (EV), there was no spontaneous hypersensitive response (HR), and empty vector itself also had no HR response, but when Pm1a and AvrPm1a When injected into tobacco leaves simultaneously, a significant HR response occurred in the region where the two genes were expressed ( Figure 6 ), which is manifested by rapid cell death or browning in some areas, while BAX is used as a positive control. Pm1a Able to specifically recognize effector factors AvrPm1a , triggering the plant's immune response.

[0093] In summary, the present invention provides a bidirectional association analysis method for plant-pathogen gene interactions. This method, analyzing resistance genes and effector factors in two directions (from resistance genes to avirulent effector factors and vice versa), reveals a strict, bidirectional, and specific recognition relationship between resistance genes and effector factors. This method can be used to screen for interacting gene pairs between plants and pathogens. This method can accelerate the identification of plant disease-resistance genes and analyze their interactions with effector factors. This not only helps fill gaps in current research but also provides an important theoretical foundation and technical support for uncovering the molecular mechanisms of plant disease resistance, improving disease-resistance breeding strategies, and cultivating plant varieties with broad-spectrum, durable disease resistance.

[0094] Although the specific embodiments of the present invention have been described in detail in conjunction with the embodiments, this should not be construed as limiting the scope of protection of this patent. Within the scope described by the claims, various modifications and variations that can be made by those skilled in the art without creative work still fall within the scope of protection of this patent.

Claims

1. A bidirectional association analysis method for plant and pathogen gene interactions, characterized in that: The following steps are involved: (1) Obtaining genetic variation information of plants, genetic variation information of pathogens, and resistance and susceptibility phenotypes of plants to pathogens; the genetic variation information of plants is single nucleotide polymorphisms, and the genetic variation information of pathogens is single nucleotide polymorphisms; (2) Forward genome-wide association analysis to locate pathogen effector factors based on plant disease resistance genes: ① Performing genome-wide association analysis on the genetic variation information of the plant and the resistance and susceptibility phenotype of the plant to the pathogen, obtaining the association results between the variation sites on the plant genome and the resistance and susceptibility phenotype of the plant to the pathogen, drawing a Manhattan plot and identifying candidate peaks; ② Identify candidate disease resistance genes in plants from candidate peaks; ③ Identify the resistance haplotype of the candidate disease-resistance gene based on the plant's resistance-susceptibility phenotype to the pathogen. Specifically, concatenate adjacent SNP sites on the candidate disease-resistance gene to form the haplotype of the candidate disease-resistance gene. Perform association analysis between the different haplotypes of the gene and the plant's resistance-susceptibility phenotype to the pathogen using chi-square tests, logistic regression, analysis of variance, and / or Kruskal-Wallis tests. Assess the significance of the association between the haplotype and the phenotype, and select the disease-resistance haplotype of the candidate disease-resistance gene. ④ Taking the infection phenotype of the pathogen on the plant containing the disease-resistant haplotype of the candidate disease-resistant gene as a trait, performing genome-wide association analysis on the trait and the genetic variation information of the pathogen to locate the candidate effector factor on the pathogen genome; (3) Reverse genome-wide association analysis to locate plant disease resistance based on pathogen effector factors: ① Performing genome-wide association analysis on the genetic variation information of the pathogen and the resistance and susceptibility phenotype of the plant to the pathogen, obtaining the association results between the variation sites on the pathogen genome and the resistance and susceptibility phenotype of the plant to the pathogen, drawing a Manhattan plot and identifying candidate peaks; ② Identify candidate effector factors in pathogens from candidate peaks; ③ Combined with the plant's resistance-susceptibility phenotype to pathogens, identify the avirulent haplotype of the candidate effector; ④ Taking the plant's susceptible phenotype to a strain with an avirulent haplotype containing a candidate effector as a trait, performing genome-wide association analysis on the trait and the plant's genetic variation information to locate candidate disease resistance genes on the plant genome; (4) The candidate disease resistance gene obtained in step (2) and step (3) is the same as the candidate effector factor, and the two are the gene pair that interacts between the plant and the pathogen.

2. The method for bidirectional association analysis of plant-pathogen gene interactions according to claim 1, wherein: In step ① of the forward whole-genome association analysis and step ① of the reverse whole-genome association analysis, whole-genome association analysis was performed using association models MLM, FarmCPU, MLMM, and BLINK, and at least three association models located the same peak region in the Manhattan plot as a candidate peak.

3. The method for bidirectional association analysis of plant-pathogen gene interactions according to claim 2, characterized in that: The specific operation of step ② of the forward genome-wide association analysis is to check the functional annotations of genes near the significant peak in the Manhattan plot and search for candidate disease resistance genes within 1 Mb upstream and downstream of the significant peak; Candidate peaks must meet the following conditions simultaneously: (a) the p-value of the SNP within the candidate peak is less than the significance threshold of 1e-8; (b) the distance between two adjacent SNPs that meet the first condition must be less than 1Mb; (c) the number of SNPs that meet the first two conditions in the candidate peak must be greater than 5; (d) considering linkage disequilibrium, the r² value must be greater than 0.8; Candidate disease-resistance genes meet the following conditions simultaneously: (a) the gene function annotation is related to immune response, pathogen recognition or defense mechanism; (b) the gene has homologous genes in the known disease-resistance gene database; (c) the gene expression level in disease-resistant varieties is higher than that in susceptible varieties.

4. The method for bidirectional association analysis of plant and pathogen gene interactions according to claim 2, characterized in that: The specific operation of step ② of the reverse genome-wide association analysis is to check the functional annotations of genes near the significant peak in the Manhattan plot and search for candidate effector factors within 100 kb upstream and downstream of the significant peak; Candidate peaks must meet the following conditions simultaneously: (a) the p-value of the SNP within the candidate peak is less than the significance threshold of 1e-5; (b) the distance between two adjacent SNPs that meet the first condition must be less than 10 kb; (c) the number of SNPs that meet the first two conditions in the candidate peak must be greater than 10; (d) considering linkage disequilibrium, the r² value must be greater than 0.8; The candidate effectors meet the following conditions simultaneously: (a) have a signal peptide; (b) have no transmembrane domain; (c) have a molecular weight less than 35 kDa; and (d) contain a specific YxC motif.

5. The method for bidirectional association analysis of plant-pathogen gene interactions according to claim 1, characterized in that: The specific operation of step ③ of the reverse whole-genome association analysis is as follows: adjacent SNP sites on the candidate effector are strung together to form the haplotype of the candidate effector, and the different haplotypes of the effector are associated with the resistance phenotype of the pathogen-infected plant using chi-square test, logistic regression, variance analysis and / or Kruskal-Wallis test to evaluate the significance of the association between the haplotype and the phenotype, and select the non-toxic haplotype of the candidate effector.

6. The method for bidirectional association analysis of plant and pathogen gene interactions according to claim 5, characterized in that: In step ③ of the forward genome-wide association analysis, logistic regression is used to evaluate the disease-resistant haplotypes that are significantly associated with the haplotypes of the candidate disease-resistant genes and the resistance traits: the plant resistance phenotype is regarded as a binary classification, with a phenotypic value <= 2.3 being resistance and a phenotypic value > 2.3 being non-resistance, with resistance and non-resistance as dependent variables, resistance being recorded as 1 and non-resistance being recorded as 0, and haplotype as the independent variable; if the regression coefficient of the haplotype is positive and p < 0.05, the haplotype is a disease-resistant haplotype; In step ③ of the reverse genome-wide association analysis, logistic regression is used to evaluate the avirulent haplotype that is significantly associated with the haplotype of the candidate effector and the toxic trait: the pathogen toxicity phenotype is regarded as a binary classification, the phenotypic value <= 2.3 is avirulent, the phenotypic value > 2.3 is toxic, toxicity and non-toxicity are used as dependent variables, toxicity is recorded as 1, non-toxicity is recorded as 0, and haplotype is used as the independent variable. If the regression coefficient of the haplotype is negative and p < 0.05, the haplotype is avirulent haplotype.

7. The method for bidirectional association analysis of plant-pathogen gene interactions according to claim 1, wherein: In step ④ of the forward whole-genome association analysis and step ④ of the reverse whole-genome association analysis, whole-genome association analysis was performed using the association model MLM.

8. Use of the bidirectional association analysis method for plant-pathogen gene interactions according to any one of claims 1 to 7 in screening interacting gene pairs.

9. The use according to claim 8, characterized in that: The bidirectional association analysis method of the interaction between plant and pathogen genes is used in screening wheat disease resistance genes and corresponding effect factors in powdery mildew.

Citation Information

Patent Citations

  • Method for screening cattle high altitude hypoxia adaptation molecular markers and application thereof

    CN109994153A

  • Group of SNP loci remarkably associated with wheat stripe rust resistance and application thereof in genetic breeding

    CN112481410A