Method for detecting nucleoplasm gene loci having interaction effect with agronomic traits
By calculating the nucleocytoplasmic genetic linkage disequilibrium coefficient and constructing a mixed linear model, nucleocytoplasmic gene loci that interact with agronomic traits were screened out, overcoming the shortcomings of existing technologies in detecting nucleocytoplasmic interaction effects and achieving efficient and accurate breeding guidance.
Patent Information
- Application Number
- CN202511111294.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-08
- Publication Date
- 2025-11-28
AI Technical Summary
Existing technologies lack methods for detecting the interaction effects of nucleocytoplasmic gene loci in crops, making it impossible to accurately assess the association between nucleocytoplasmic interaction loci and agronomic traits, and failing to meet the requirements for rapid agricultural testing or breeding guidance.
By calculating the genetic linkage disequilibrium coefficient between nucleocytoplasm and nucleocytoplasm, nucleocytoplasmic interaction pairs with abnormally strong linkage were screened out. A mixed linear model was constructed for GWAS analysis to screen out nucleocytoplasmic gene loci that have interaction effects with agronomic traits.
It has achieved accuracy and timeliness in detecting nucleoplasmic gene loci, simplified the analysis steps, reduced the workload of molecular genetic research on crop agronomic traits, and provided precise breeding guidance.
Smart Images

Figure CN121023072A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of plant genetic breeding, and in particular to a method for detecting a nucleo-cytoplasmic genetic locus with interaction effects on agronomic traits. BACKGROUND
[0002] The quality of nucleo-cytoplasmic co-evolution in plants can affect the performance of plants in important life activities such as photosynthesis, energy storage and metabolism, and disease resistance. Establishing a method for detecting the interaction effects of nucleo-cytoplasmic genetic loci on agronomic traits in crops can help breeders quickly identify potential and phenotypically superior nucleo-cytoplasmic interaction genotypes, providing guidance for breeding work.
[0003] However, detecting the interaction effects of nucleo-cytoplasmic genetic loci in crops faces several challenges, the core of which is the lack of evaluation of nucleo-cytoplasmic genotype interaction effects in currently disclosed technologies. In genetics, the linkage disequilibrium (LD) in the cell nucleus can be calculated to evaluate the simultaneous occurrence of alleles belonging to two or more loci at different sites in the cell nucleus at a frequency higher than random occurrence. For example, the construction method of a linkage disequilibrium analysis model for a natural population of homologous tetraploids disclosed in Chinese patent CN103699815A includes randomly selecting n homologous tetraploid individuals from the natural population to obtain the number of individuals with different genotypes; calculating the corresponding gamete frequency according to the haplotype frequency, and calculating the corresponding genotype frequency and haplotype gene frequency according to the gamete frequency, obtaining the estimated frequency of alleles at two sites through the haplotype gene frequency, and calculating the linkage disequilibrium coefficient between each two sites through the haplotype gene frequency; however, this method only evaluates the coupling between the sites in the cell nucleus, and the coupling between the two sites in the cell nucleus does not exclude physical linkage.
[0004] In addition, the current disclosed technologies lack the interaction between the cell nucleus and the cytoplasmic genome (hereinafter referred to as "nucleo-cytoplasmic") and the linkage with agronomic traits. In genetics, large-scale sequencing and agronomic trait identification of crop populations can be used to quickly correlate agronomic traits and genotypes using the GWAS (genome-wide association study) model. For example, the quantitative trait multi-locus oscillation search whole genome association analysis system and method disclosed in Chinese patent CN119580831A includes: data preparation, preparing SNP matrix files and phenotype matrix files; preprocessing: removing the entered SNP matrix and normalizing by column, and normalizing the phenotype matrix by column; association analysis: selecting SNP sites related to the phenotype from the SNP matrix, i.e. PseQTNs, through an oscillation search strategy. However, this method only identifies major genes or minor genes on the cell nucleus that are associated with agronomic traits. It lacks the identification of cytoplasmic genes, and even more lacks the identification of nucleo-cytoplasmic coupling associated with agronomic properties.
[0005] The existing method lacks detection of the effect of nucleo-cytoplasmic gene locus interaction in crops, and lacks correlation between the nucleo-cytoplasmic interaction locus and the phenotype, which cannot meet the requirements of rapid detection or breeding guidance in agriculture. Therefore, it is urgent to develop a detection method for the effect of nucleo-cytoplasmic gene locus interaction with high accuracy and good stability, which can simplify the analysis steps, shorten the running time, and improve the accuracy and timeliness of nucleo-cytoplasmic gene locus detection. SUMMARY
[0006] In order to solve the above problems, the present application provides a method for detecting nucleo-cytoplasmic gene loci having interaction effect with agronomic traits.
[0007] The present application provides a method for detecting nucleo-cytoplasmic gene loci having interaction effect with agronomic traits, comprising the following steps:
[0008] S1, downloading or measuring the data of the to-be-detected agronomic traits of each individual in the to-be-detected crop population from a public database;
[0009] S2, downloading or sequencing the nucleome SNP data and the mitochondrial genome SNP data of each individual in the to-be-detected crop population from a public database;
[0010] S3, pairing the nucleome SNP data and the mitochondrial genome SNP data obtained in S2 two by two, calculating the genetic linkage disequilibrium coefficient between the nucleo-cytoplasm of each mitochondrial genome SNP, and screening out the nucleome SNP and the mitochondrial genome SNP pair with abnormal linkage strength as a candidate nucleo-cytoplasm interaction coupling pair, taking each nucleome SNP as the core;
[0011] S4, taking the typing result of each nucleome SNP corresponding to the candidate nucleo-cytoplasm interaction coupling pair obtained in S3 as nucle SNP i, the typing result of each mitochondrial genome SNP as cyto SNP j, and the typing result of each candidate nucleo-cytoplasm interaction coupling pair as nucle SNP i x cyto SNP j, constructing a mixed linear model to implement GWAS analysis;
[0012] S5, introducing the mixed linear model constructed in S4 into the R package GAPIT program for running, using the R package CMplot to perform significance test on the association of each candidate nucleo-cytoplasm interaction coupling pair with the to-be-detected agronomic traits, and obtaining the p value of the association of each candidate nucleo-cytoplasm interaction coupling pair with the to-be-detected agronomic traits;
[0013] S6, using the R package CMplot to make a Q-Q plot and a Manhattan plot of the p value obtained in S5, determining a significance threshold according to the Q-Q plot, and screening out the nucleo-cytoplasm interaction coupling pair with a-log 10 p higher than the significance threshold, and the corresponding nucleome SNP and mitochondrial genome SNP are the nucleo-cytoplasmic gene loci having interaction effect with the agronomic traits.
[0014] Preferably, the step S3 is: pairing the nuclear genomic SNP data and the mitochondrial genomic SNP data obtained in S2 two by two, taking each nuclear genomic SNP as the core, calculating the genetic linkage disequilibrium coefficient between it and each mitochondrial genomic SNP, taking the calculation result of each nuclear genomic SNP as an independent analysis group, performing statistical test in each analysis group by Bootstrap self-help method, setting the threshold value as 5%, and identifying and screening out the nuclear genomic SNP and the mitochondrial genomic SNP pair with abnormal prominent linkage intensity in the group as the candidate nuclear-mitochondrial interaction coupling pair.
[0015] Preferably, in the step S3, the formula for calculating the genetic linkage disequilibrium coefficient between the nucleus and the mitochondria is:
[0016] ;
[0017] wherein,
[0018] : A represents a locus located on the nuclear genome, and represents an allele at the A locus in the population;
[0019] : B represents a locus located on the mitochondrial or chloroplast genome, and represents an allele at the B locus in the population;
[0020] : if and are completely independent, the frequency of simultaneous occurrence is expected;
[0021] : quantifies the degree to which a nuclear allele can predict the corresponding mitochondrial genome type; the value range is [0, 1].
[0022] Preferably, in the step S3, the calculation formula is:
[0023]
[0024] wherein,
[0025] : the frequency of observing ;
[0026] : the frequency of observing ;
[0027] : simultaneously observing and the frequency of simultaneous occurrence when
[0028] if and both are completely independent;
[0029] a core index measuring the degree of association between a specific nuclear allele and a specific cytonucleic genome ; .
[0030] Preferably, the mixed linear model in S4 is:
[0031] phenotype: nuclear SNP i + cytonucleic SNP j + nuclear SNP i x cytonucleic SNP j + cluster + (1|individual ID);
[0032] wherein,
[0033] phenotype: data of the to-be-tested agronomic traits of the to-be-tested crop population obtained in S1;
[0034] nuclear SNP i: each nuclear SNP i has two types;
[0035] cytonucleic SNP j: each cytonucleic SNP j has two types;
[0036] nuclear SNP i x cytonucleic SNP j: each candidate nuclear-cytonucleic interaction coupling pair has four types;
[0037] cluster: the clustering information of the to-be-tested crop population obtained by PCA analysis or t-SNE analysis;
[0038] 1|individual ID: the kinship matrix of the to-be-tested crop population obtained by R package GAPIT analysis.
[0039] Preferably, in the step S6, the method for determining the significance threshold according to the Q-Q plot is: the points far away from the diagonal line in the upper right corner of the Q-Q plot represent the real genetic loci related to the to-be-tested agronomic traits, and the position where the -log 10 p begins to deviate from the diagonal line is the significance threshold.
[0040] Preferably, in S1, in order to ensure data quality, the individual records with missing values in the data of the to-be-tested agronomic traits are deleted.
[0041] Preferably, in S2, to ensure data quality, the nuclear genome SNP data and the mitochondrial genome SNP data are filtered, and the filtering criteria are: removing non-bi-allele sites; the SNP site is located within the gene region and the physical distance of 2kb upstream and downstream thereof; the average MAPQ of the aligned reads is at least 20; the sequencing depth is at least 10; the minimum allele frequency is at least 5%; the minimum MAC is at least 5; and the heterozygous genotype frequency is less than 0.5.
[0042] Preferably, in S2, if the public database does not provide mitochondrial genome SNP data, the sequencing data of the crop population to be tested are used, the data are aligned to the mitochondrial reference genome of the corresponding crop by using BWA software, and the mitochondrial genome SNP data are obtained by using professional software.
[0043] The method for detecting nuclear-cytoplasmic gene loci having interaction effects on agronomic traits provided by the application is applied in crop genetic breeding.
[0044] The beneficial effects of the application are as follows:
[0045] The application provides a method for detecting nuclear-cytoplasmic gene loci having interaction effects on agronomic traits, comprising the following steps: 10 p is higher than a significant threshold value, and the corresponding nuclear genome SNP and mitochondrial genome SNP are nuclear-cytoplasmic gene loci having interaction effects on agronomic traits.
[0046] The application can evaluate the effects of non-random combinations of SNPs on all nuclear-cytoplasmic gene loci by distinguishing the effects of nuclear-cytoplasmic gene loci interaction from the effects of single nucleus or cytoplasm based on GWAS analysis of a mixed linear model, and can greatly reduce the workload of molecular genetic research on crop agronomic traits by screening nuclear-cytoplasmic gene loci having high correlation with the agronomic traits to be tested by a Bootstrap statistical method.
[0047] The method of the application can obtain results by software analysis, and has simple analysis steps, short time consumption, accurate detection of nuclear-cytoplasmic gene loci, and wide application range. BRIEF DESCRIPTION OF DRAWINGS
[0048] Figure 1 Q-Q plot for performing significance test on the correlation between each nuclear-cytoplasmic interaction coupling pair and the agronomic traits to be tested in Example 1;
[0049] Figure 2Manhattan plot for significance test of association between each nuclear-cytoplasmic interaction pair and the agronomic trait of interest in Example 1.
[0050] Figure 3 Distribution of genomic regions strongly associated with the plant height agronomic trait detected by the common GWAS method in Example 2. DETAILED DESCRIPTION
[0051] The application will be further described in conjunction with the following examples.
[0052] If not specifically indicated, the technical means used in the examples are the conventional means well known to those skilled in the art.
[0053] Example 1:
[0054] This example aims to establish a precise and efficient method by integrating genomics, population genetics and statistical analysis to identify the interaction effect sites between the nuclear and cytoplasmic genomes that affect key agronomic traits in crops.
[0055] This example uses a complete set of wheat germplasm resources from more than 30 countries, including 827 A.E. Watkins wheat landraces (Wheatlandraces) and 220 global modern varieties, a total of 1047 materials. This example includes the following steps:
[0056] S1, Obtain the plant height agronomic trait dataset. The plant height agronomic trait data of all materials are downloaded from the public database wwwg2b.com. To ensure data quality, individual records with missing values in the dataset were deleted.
[0057] S2, Obtain nuclear genome and mitochondrial genome SNP data. For nuclear genome SNP data, the original dataset containing 61659890 nuclear genome sites was directly downloaded from the above database. To ensure data quality, the dataset was strictly filtered, and finally 61109627 high-quality nuclear genome SNPs were obtained for subsequent analysis. The filtering criteria are as follows:
[0058] (i) Remove non-bi-allele sites;
[0059] (ii) SNP sites should be located within the physical distance of 2kb upstream and downstream of the gene region;
[0060] (iii) The average MAPQ of the aligned reads is at least 20;
[0061] (iv) The sequencing depth is at least 10;
[0062] (v) The minimum allele frequency is at least 5%;
[0063] (vi) the minimum MAC is at least 5;
[0064] (vii) the hybrid genotype frequency < 0.5;
[0065] In this embodiment, for the mitochondrial genome SNP data, since it is not directly provided in the database, the inventors downloaded the original resequencing data of the population from the database, aligned the data to the "Chinese Spring Wheat" mitochondrial reference genome using BWA software to obtain a SAM file, then converted the SAM file to a BAM file using samtools software, and then used the GATKMutect2 process for SNP variant detection. The preliminary obtained SNP set was subjected to the same filtering standard as the nuclear genome SNP for quality control, and finally 1549 high-quality mitochondrial genome SNPs were obtained.
[0066] S3, pair the nuclear genome SNP data and the mitochondrial genome SNP data obtained in S2 two by two, take each nuclear genome SNP as the core, calculate the nuclear-mitochondrial linkage disequilibrium coefficient thereof with each mitochondrial genome SNP, and the calculation result of each nuclear genome SNP is an independent analysis group. In each analysis group, statistical test is performed by Bootstrap self-help method, the threshold is set to 5%, and the nuclear genome SNP and the mitochondrial genome SNP pair with abnormal prominent linkage intensity in the group are identified and screened as candidate nuclear-mitochondrial interaction coupling pairs.
[0067] In this embodiment, the formula for calculating the genetic linkage disequilibrium coefficient between nuclear-mitochondrial sites is:
[0068] ;
[0069] wherein,
[0070] : A represents a locus located on the nuclear genome, and represents an allele at the A locus in the population;
[0071] : B represents a locus located on the mitochondrial or chloroplast genome, and represents an allele at the B locus in the population;
[0072] : If and are completely independent, the frequency of simultaneous occurrence is expected;
[0073] : Quantifies to what extent a nuclear allele can predict the corresponding mitochondrial genome type; its value range is [0, 1].
[0074]
[0075] wherein,
[0076] : the frequency (observed value) of observing
[0077] : the frequency (observed value) of observing
[0078] : the frequency (observed value) of observing and
[0079] : the frequency (expected value) of observing and if they are completely independent;
[0080] : the core index measuring the degree of association between a specific nuclear allele and a specific mitochondrial genome . .
[0081] In addition to the calculation formula of r 2 , other commonly used standardized calculation formulas of genetic linkage disequilibrium coefficients in the art, including D', D * , x 2 df , x 2’ , etc. can also be used for the calculation of this step.
[0082] This step initially obtains 61109627 x 1549 pairwise pairing results. In order to accurately screen candidate nuclear-mitochondrial interaction coupling pairs, the inventors use a grouping test strategy: taking each nuclear genome SNP as the core, the r² value of each nuclear genome SNP and all 1549 mitochondrial SNPs is taken as an independent analysis group. In each group, statistical test is performed by Bootstrap self-help method for distribution analysis, and the sampling number is set to 1000 times and the test threshold is set to 5%. The "nuclear-mitochondrial" pairs with abnormal linkage strength with high genetic linkage disequilibrium coefficient value higher than the set detection threshold are screened out, that is, the nuclear genome SNP and mitochondrial genome SNP pairs are taken as candidate nuclear-mitochondrial interaction coupling pairs to enter the next step of analysis. Through this step of screening, 4705441585 "nuclear-mitochondrial" pairs are actually obtained to enter the next step of analysis.
[0083] S4. Constructing a mixed linear model to perform GWAS analysis using the genotyping results of each candidate nuclear-mitochondrial interaction pair as the nuclear SNP i, the mitochondrial SNP j, and the nuclear-mitochondrial interaction SNP i x j, respectively.
[0084] The mixed linear model in S4 is:
[0085] Phenotype ~ nuclear SNP i + mitochondrial SNP j + nuclear-mitochondrial interaction SNP i x j + cluster + (1 | individual ID).
[0086] Wherein,
[0087] Phenotype: the data of the agronomic traits of the test crop population obtained in S1; the value of the agronomic trait to be analyzed, which in the present example corresponds to the plant height data of the sample population.
[0088] Nuclear SNP i: each nuclear SNP i has two genotypes; the candidate nuclear genome SNP is a fixed effect, representing its own main effect.
[0089] Mitochondrial SNP j: each mitochondrial SNP j has two genotypes; the candidate mitochondrial genome SNP is a fixed effect, representing its own main effect.
[0090] Nuclear-mitochondrial interaction SNP i x j: each candidate nuclear-mitochondrial interaction pair has four genotypes; the interaction term of the nuclear SNP and the mitochondrial SNP, as a fixed effect, is the core of the present example. The significance of this term represents the influence of the interaction effect of the nuclear-mitochondrial site on the phenotype.
[0091] Cluster: the clustering information of the test crop population obtained by PCA analysis or t-SNE analysis is used as a fixed effect covariate to correct the population structure. The clustering data in the present example directly uses the clustering results provided by the database, and 1048 wheat samples are divided into seven components.
[0092] 1 | individual ID: the kinship matrix of the test crop population is obtained by R package GAPIT analysis. The kinship matrix constructed by individual ID is used as a random effect to control the background noise caused by the recessive genetic relationship between individuals. The kinship matrix in the present example directly uses the kinship matrix results provided by the database.
[0093] The mixed linear model provided by this invention distinguishes the effects of nuclear genomic SNPs, cytoplasmic genomic SNPs, and the interaction effects of nuclear and cytoplasmic SNPs. This distinguishes the effects of SNP genotyping at nuclear loci on agronomic traits for each candidate nucleocytoplasmic interaction pair, the effects of SNP genotyping at each cytoplasmic locus on agronomic traits, and the interaction effects of SNP genotyping at each nuclear locus and cytoplasmic locus on agronomic traits. Furthermore, based on the actual conditions of the materials used in this example, it is necessary to control for population structure and inter-individual kinship relationships within the mixed linear model.
[0094] S5. The mixed linear model constructed in S4 is introduced into the R package GAPIT program and run. The R package CMplot is used to perform a significance test on the association between each candidate nucleocytoplasmic interaction pair and the agronomic trait to be tested, and the p-value of the association between each candidate nucleocytoplasmic interaction pair and the agronomic trait to be tested is obtained.
[0095] S6. Use the R package CMplot to plot a Q-plot of the p-values obtained in S5. Figure 1 The significance threshold is determined based on the QQ plot. The QQ plot is used to assess whether the p-value distribution matches expectations and to detect the influence of population structure or other confounding factors on the results. Ideally, if the model fits well and there is no population structure influence, all points should be closely distributed along the diagonal. If the curve is observed to deviate upwards from the diagonal, it may indicate the presence of uncorrected population structure. Figure 1 The points in the upper right corner of the QQ graph, away from the diagonal, represent the actual genetic loci associated with the agronomic trait being tested. Figure 1 It can be seen that when -log 10 When the p-value is around 5, the upper right corner starts to move away from the diagonal, so -log is chosen. 10 When p = 5, it is used as the preset significance threshold. If -log 10 When p reaches the preset significance threshold of 5, it is determined that the interaction effect of the nucleoplasmic locus is significantly associated with the agronomic trait of plant height.
[0096] Use the R package CMplot to plot the p-values obtained from S5 in a Manhattan plot. Figure 2 ),exist Figure 2 In the middle, filter out -log 10 Nucleo-cytoplasmic pairings with p values higher than the screening threshold of 5 are called nucleo-cytoplasmic interaction pairs, as shown in Table 1. A total of 10 nucleo-cytoplasmic interaction pairs were selected, which are the nucleo-cytoplasmic gene loci that have interaction effects with wheat plant height agronomic traits obtained in this embodiment.
[0097] Table 1. Nucleocytoplasmic gene loci that interact with wheat plant height agronomic traits.
[0098]
[0099] From the present embodiment, it can be seen that the method of the present application can accurately screen out the nucleo-cytoplasmic interaction coupling pair having interaction effect with the to-be-tested agronomic trait of the to-be-tested crop, and the corresponding nuclear genome SNP and mitochondrial genome SNP are the nucleo-cytoplasmic gene sites having interaction effect with the agronomic trait, thereby greatly reducing the workload of molecular genetic research of crop agronomic traits.
[0100] Example 2:
[0101] Example 2 is to mine the candidate genes related to the plant height agronomic trait in the wheat local variety by using the existing whole genome association analysis (GWAS), and the sample material is the same as that in Example 1, and the operation steps S1 and S2 are the same as those in Example 1. The data of the to-be-tested agronomic trait of each individual in the to-be-tested crop population and the nuclear genome SNP data and the mitochondrial genome SNP data of each individual in the to-be-tested crop population are obtained, all SNPs are introduced into a mixed linear model for GWAS analysis,
[0102] The mixed linear model is: phenotype ~ SNP i + grouping + (1|individual ID),
[0103] Wherein SNP i represents the effect of all SNPs (nuclear genome SNP data and mitochondrial genome SNP data),
[0104] phenotype, grouping, 1|individual ID are the same as those in Example 1.
[0105] Figure 3 is the distribution of the genomic region strongly related to the plant height agronomic trait obtained by using the ordinary GWAS method for analysis in Example 2. Figure 3 Also from the public database wwwg2b.com, according to the introduction of the database website, the figure is obtained after the p value of SNPi is called peak, the genomic fragment significantly related to the plant height trait is obtained, the green fragment in the figure is the gene fragment significantly related to the plant height trait, and the x-axis coordinate of the green segment represents the chromosome number and specific coordinates of the fragment on the genome. Then, according to these fragments, the candidate genes related to the trait are selected.
[0106] Through statistics, 180 nuclear genomic fragments which are strongly related to the plant height agronomic traits of wheat are obtained by the method of Example 2, and 23,541 candidate genes are involved, 0 mitochondrial genomic fragments which are strongly related to the plant height agronomic traits of wheat are obtained, and 0 candidate genes are involved. The range of the candidate genes with research potential obtained by the general GWAS method of Example 2 is too wide, and the specific sites involved are massive data, which is a huge workload for genetic breeding research. Moreover, neither the nuclear genome nor the mitochondrial genome has overlapping results with the results obtained in Example 1, which shows that the interaction effect of the nuclear-cytoplasmic gene locus in crops cannot be detected by the general GWAS method, and also shows that the method of the present application can accurately analyze the influence of the interaction effect of the nuclear-cytoplasmic locus on the agronomic traits in crops, and can effectively distinguish the interaction effect of the nuclear-cytoplasmic locus from the independent effect of the nuclear locus or the cytoplasmic locus, avoiding confusion of the detection results. The method of the present application can accurately screen the gene loci of the nucleus and cytoplasm (mitochondria) which have the interaction effect of the nuclear-cytoplasmic locus on the measured agronomic traits, greatly reducing the workload of breeding research, and can provide more accurate references for breeders.
[0107] It is apparent to those skilled in the art that the present application is not limited to the details of the foregoing exemplary embodiments, and that the present application can be implemented in other particular forms without departing from the spirit or essential characteristics of the present application. Therefore, the embodiments should be considered in all respects as illustrative and not restrictive, and the scope of the present application should be defined by the appended claims rather than the above description, and it is intended to include all changes falling within the meaning and range of equivalents of the claims.
[0108] In addition, it should be understood that although the present specification is described in terms of embodiments, not every embodiment contains only one independent technical solution, and the description of the specification is only for the sake of clarity, and those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can be combined appropriately to form other embodiments that those skilled in the art can understand.
Claims
1. A method for detecting nucleocytoplasmic gene loci that interact with agronomic traits, characterized in that, Includes the following steps: S1. Download or measure data on the agronomic traits to be tested for each individual in the crop population to be tested from public databases; S2. Obtain nuclear genome SNP data and mitochondrial genome SNP data for each individual in the crop population to be tested by downloading from public databases or sequencing. S3. Pair the nuclear genome SNP data and mitochondrial genome SNP data obtained in S2. Using each nuclear genome SNP as the core, calculate the nucleocytoplasmic linkage disequilibrium coefficient between it and each mitochondrial genome SNP. Screen out nuclear genome SNP and mitochondrial genome SNP pairs with abnormally strong linkage as candidate nucleocytoplasmic interaction coupling pairs. S4. The genotyping results of each nuclear genome SNP corresponding to the candidate nucleocytoplasmic interaction pair obtained in S3 are taken as nuclear SNP, the genotyping results of each mitochondrial genome SNP are taken as cytoplasmic SNPj, and the genotyping results of each candidate nucleocytoplasmic interaction pair are taken as nuclear SNP×cytoplasmic SNPj. A mixed linear model is constructed to perform GWAS analysis. S5. The mixed linear model constructed in S4 is introduced into the R package GAPIT program and run. The R package CMplot is used to perform a significance test on the association between each candidate nucleocytoplasmic interaction pair and the agronomic trait to be tested, and the p-value of the association between each candidate nucleocytoplasmic interaction pair and the agronomic trait to be tested is obtained. S6. Use the R package CMplot to create a QQ plot and a Manhattan plot for the p-values obtained in S5. Determine the significance threshold based on the QQ plot, and filter out -log values from the Manhattan plot. 10 Nucleocytoplasmic interaction pairs with p values higher than the significance threshold are nuclear genome SNPs and mitochondrial genome SNPs that are nucleocytoplasmic gene loci that interact with agronomic traits.
2. The method for detecting nucleocytoplasmic gene locus interaction effects in crops according to claim 1, characterized in that, Step S3 is as follows: the nuclear genome SNP data and mitochondrial genome SNP data obtained in S2 are paired in pairs. Taking each nuclear genome SNP as the core, the nucleocytoplasmic linkage disequilibrium coefficient between it and each mitochondrial genome SNP is calculated. The calculation results of each nuclear genome SNP are used as an independent analysis group. Within each analysis group, statistical tests are performed using the Bootstrap method. A threshold of 5% is set to identify and screen nuclear genome SNP and mitochondrial genome SNP pairs with abnormally strong linkage strength within the group as candidate nucleocytoplasmic interaction coupling pairs.
3. The method for detecting nucleocytoplasmic gene locus interaction effects in crops according to claim 1, characterized in that, In step S3, the formula for calculating the nucleocytoplasmic linkage disequilibrium coefficient is as follows: ; in, A represents a locus located on the genome in the cell nucleus, while Represents the allele at locus A in the population; B represents a gene locus located in the mitochondrial or chloroplast genome, while Represents the allele at locus B in the population; :like and The expected frequency of simultaneous occurrence when the two are completely independent; To what extent a nuclear allele can predict the corresponding genomic genome type; Its value range is [0,1].
4. The method for detecting nucleocytoplasmic gene locus interaction effects in crops according to claim 3, characterized in that, In step S3 The calculation formula is: in, Observed The frequency; Observed The frequency; Simultaneous observation and The frequency; :like and The expected frequency of simultaneous occurrence when the two are completely independent; : Measuring specific nuclear alleles and specific genomes The core indicator of the degree of correlation between them; .
5. The method for detecting nucleocytoplasmic gene locus interaction effects in crops according to claim 1, characterized in that, The hybrid linear model in S4 is: Phenotype ~ nucleus SNPj + cytogen SNPj + nucleus SNPj × cytogen SNPj + cluster + (1 | individual ID); in, Phenotype: Data on the agronomic traits of the crop population to be tested obtained in S1; Core SNPi: Each core SNPi has two fractals; Quality SNPj: Each quality SNPj has two subtypes; Nucleus SNPJ: Each candidate nucleocytic interaction pair has four subtypes; Clustering: Obtain clustering information of the crop population under test through PCA analysis or t-SNE analysis; 1 | Individual ID: The kinship matrix of the crop population under test was obtained by analysis using the R package GAPIT.
6. The method for detecting nucleocytoplasmic gene locus interaction effects in crops according to claim 1, characterized in that, In step S6, the method for determining the significance threshold based on the QQ plot is as follows: the points in the upper right corner of the QQ plot, away from the diagonal, represent the true genetic loci associated with the agronomic trait to be tested, and -log is selected. 10 The position where p begins to deviate from the diagonal is the significance threshold.
7. The method for detecting nucleocytoplasmic gene locus interaction effects in crops according to claim 1, characterized in that, In step S1, to ensure data quality, individual records with missing values in the data of the agronomic traits to be tested are deleted.
8. The method for detecting nucleocytoplasmic gene locus interaction effects in crops according to claim 1, characterized in that, In S2, to ensure data quality, nuclear genome SNP data and mitochondrial genome SNP data are filtered. The filtering criteria are: non-bi-allele sites are removed; SNP sites must be located within a 2kb physical distance of the gene region and its upstream and downstream regions; the average MAPQ of the aligned reads is at least 20; the sequencing depth is at least 10; the minimum allele frequency is at least 5%; the minimum MAC is at least 5; and the heterozygous genotype frequency is <0.
5.
9. The method for detecting nucleocytoplasmic gene locus interaction effects in crops according to claim 1, characterized in that, In step S2, if the publicly available database does not provide mitochondrial genome SNP data, the sequencing data of the crop population to be tested is used, and the data is aligned to the mitochondrial reference genome of the corresponding crop using BWA software, and the mitochondrial genome SNP data is obtained using specialized software.
10. The application of the method for detecting nucleocytoplasmic gene loci that interact with agronomic traits as described in any one of claims 1-9 in crop genetic breeding.
Citation Information
Patent Citations
Construction method for linkage disequilibrium analysis model of autotetraploid natural population
CN103699815A
Quantitative character multi-site oscillation search whole genome association analysis system and method
CN119580831A