Computer device and method for detecting target phenotype related genes or candidate genes based on low-depth sequencing data and application

By constructing a specific k-mer database and a genotype-to-phenotype prediction model, the problem of insufficient annotation of gliadin-encoding genes in the wheat genome was solved, thereby improving the processing quality of wheat. Genes in untested regions were identified and utilized, improving breeding efficiency and accuracy.

CN120913636APending Publication Date: 2025-11-07CHINA AGRI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410554692.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-05-07
Publication Date
2025-11-07

AI Technical Summary

Technical Problem

Existing technologies cannot effectively utilize gliadin-encoding genes in wheat quality breeding, resulting in insufficient gene annotation and limiting the improvement of wheat processing quality. In particular, due to the large number and complex structure of gliadin-encoding genes, there are many untested regions, making cloning and association analysis difficult.

Method used

By constructing computer devices and methods, using a specific k-mer database to screen and detect low-depth sequencing data, we can identify and clone wheat grain storage protein genes, establish a genotype-to-phenotype prediction model, and achieve efficient annotation of the wheat genome and mining of genes related to phenotypic traits.

Benefits of technology

This study enabled efficient cloning and gene annotation of untested regions in the wheat genome, identified candidate genes related to processing quality, and improved the breeding efficiency and accuracy of wheat processing quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure FT_1
    Figure FT_1
  • Figure FT_2
    Figure FT_2
  • Figure FT_3
    Figure FT_3
Patent Text Reader

Abstract

The invention discloses a computer device and method for detecting target phenotype related genes or candidate genes based on low-depth sequencing data and application. The present invention uses a 29-mer sequence representing each SSP gene at the generic genome level to reveal gene diversity that has not been explored between different wheat varieties and modern cultivated varieties. The SSP gene related to the processing quality is identified based on the whole genome association study of k-mer so as to improve the wheat processing quality. The device or the method provided by the invention can be applied to screening, mining and cloning of phenotypic character related candidate genes with insufficient gene annotation caused by numerous coding genes, complex structures, existence of a large number of long and short repetitive sequences in the genes and existence of a large number of non-communicated regions in coding sequences and flanking sequences; and annotation of related gene functions, correlation analysis and molecular marker development are carried out.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of bioinformatics, in particular to a computer device, method and application for detecting a target phenotype-related gene or candidate gene based on low-depth sequencing data. BACKGROUND

[0002] At present, improving the processing quality of wheat flour is one of the main breeding objectives of breeders. However, it is time-consuming and difficult to improve the processing quality of wheat, which limits the improvement of the processing quality of wheat. Previous studies have shown that gluten and prolamin are important factors affecting the processing quality of wheat, but the high molecular weight gluten subunit coding gene is mainly used in quality breeding. The reason for this phenomenon is that there are many prolamin coding genes with complex structure, and there are a large number of long short repeat sequences in the genes, which leads to a large number of untested regions (gaps) in the reference genome of prolamin coding sequences and flanking sequences, resulting in insufficient gene annotation, limiting the cloning of prolamin, association analysis and molecular marker development.

[0003] The SDS sedimentation value of wheat refers to the sedimentation reaction value of the gluten in the wheat flour to sodium dodecyl sulfate (SDS) under certain conditions. This value reflects the strength and quantity of the gluten in the wheat flour and is one of the important indicators for evaluating the quality of wheat. The SDS sedimentation value is related to the content and quality of the gluten in the wheat flour, which further affects the processing and baking characteristics of wheat. SUMMARY

[0004] The technical problem to be solved by the present application is how to mine wheat quality phenotype-related candidate genes and / or how to detect grain storage protein alleles in the wheat genome and / or how to annotate plant genotypes based on low-depth sequencing data and / or how to predict phenotypes based on plant genotypes and / or how to clone homologous prolamin and / or how to detect plant grain storage protein genes at the pan-genome level.

[0005] To solve the above technical problems, the present application first provides a computer device comprising a memory, a processor and a computer program stored in the memory, wherein the processor can execute the computer program to realize steps S1-S4:

[0006] S1, data receiving: receiving sequence data of all known genes related to the target phenotype of the plant to be tested;

[0007] S2, data grouping: grouping the sequence data based on sequence similarity to obtain similar gene sequence groups, and extracting the longest gene in each similar gene sequence group as a non-redundant phenotype-related gene;

[0008] S3, determining specific k-mers of non-redundant phenotype-related genes: using sequence analysis software to screen specific k-mers of each gene of all the non-redundant phenotype-related genes, and using the sequence analysis software to generate a database containing specific k-mers of all the non-redundant phenotype-related genes;

[0009] S4, detecting candidate genes in sequencing data of a plant to be tested: using the database to scan sequencing data of the plant to be tested, calculating a detection rate of specific k-mers of each of the non-redundant phenotype-related genes in the plant to be tested, and obtaining genes similar to the non-redundant phenotype-related genes or genes with variations in the plant to be tested based on the detection rate, wherein the genes similar to the non-redundant phenotype-related genes or the genes with variations are the phenotype-related genes or candidate genes of interest in the plant to be tested.

[0010] In the above computer device, the steps after S4 can further include C5 and / or C6.

[0011] To solve the above technical problems, the present application also provides a device for detecting phenotype-related genes or candidate genes in plants, which can include the following modules:

[0012] A1, a data receiving module: used for receiving sequence data of all known genes related to a phenotype of interest of a plant to be tested;

[0013] A2, a data grouping module: used for grouping the sequence data based on sequence similarity to obtain similar gene sequence groups, and extracting a longest gene in each of the similar gene sequence groups as a non-redundant phenotype-related gene;

[0014] A3, a specific k-mer database module for determining non-redundant phenotype-related genes: used for using sequence analysis software to screen specific k-mers of each gene of all the non-redundant phenotype-related genes, and using the sequence analysis software to generate a database containing specific k-mers of all the non-redundant phenotype-related genes;

[0015] A4, a candidate gene screening module for detecting sequencing data of a plant to be tested: used for using the database to scan sequencing data of the plant to be tested, calculating a detection rate of specific k-mers of each of the non-redundant phenotype-related genes in the plant to be tested, and obtaining genes similar to the non-redundant phenotype-related genes or genes with variations in the plant to be tested based on the detection rate, wherein the genes similar to the non-redundant phenotype-related genes or the genes with variations are the phenotype-related genes or candidate genes of interest in the plant to be tested.

[0016] The k-mer described above is a fixed-length nucleotide string, which refers to a short DNA sequence of length k. For example, a gene with a length of 1000 bp is broken into 29 bp short fragments, and the 29 bp short fragments after breaking are called 29-mers, and (1000-29+1) k-mer sequences can be obtained. The specific k-mer refers to a unique 29-mer that specifically exists in each gene among all candidate genes. The k can be less than or equal to the length of the sequencing reads in the sequencing data, such as k can be 29 bp.

[0017] To solve the above technical problems, the present application also provides a method for mining breeding candidate genes related to a target phenotype trait in plants, which can include the following steps:

[0018] C1, data receiving: receiving sequence data of all known genes related to a target phenotype of a plant to be tested;

[0019] C2, data grouping: grouping the sequence data based on sequence similarity to obtain similar gene sequence groups, and extracting the longest gene in each similar gene sequence group as a non-redundant phenotype-related gene;

[0020] C3, determining specific k-mer database of non-redundant phenotype-related genes: using sequence analysis software to screen and obtain the specific k-mer of each gene in all the non-redundant phenotype-related genes, and using the sequence analysis software to generate a database containing the specific k-mer of all the non-redundant phenotype-related genes;

[0021] C4, detecting sequencing data of a plant to be tested to screen candidate genes: using the database to scan the sequencing data of the plant to be tested, calculating the detection rate of the specific k-mer of each non-redundant phenotype-related gene in the plant to be tested, and obtaining genes similar to the non-redundant phenotype-related genes or genes with variations in the plant to be tested based on the detection rate, the similar genes or genes with variations are the target phenotype-related genes or candidate genes in the plant to be tested;

[0022] C5, obtaining a phenotype significantly related gene of interest: receiving the measured value of the phenotype of interest of each plant in the plant population to be tested, for each specific k-mer, dividing the plant population to be tested into two groups according to the presence or absence of the specific k-mer, performing a statistical test on the measured value of the phenotype of interest of the plants in the two groups, taking the p value of the statistical test as the correlation effect, calculating the correlation score, screening the specific k-mer according to the correlation score and the correlation effect to obtain a phenotype related specific k-mer, and screening the similar gene or the gene with variation according to the phenotype related specific k-mer to obtain a phenotype significantly related gene of interest;

[0023] The formula for calculating the correlation score is as follows:

[0024] Correlation score = -log 10 (p value) x |test statistic| / test statistic) Formula 1;

[0025] C6, obtaining a plant phenotype related breeding candidate gene of interest: detecting the percentage of the presence of the phenotype significantly related gene in the cultivated variety and wild variety of the plant to be tested to obtain the enrichment result of the phenotype significantly related gene in the cultivated variety and local variety, and screening to obtain a plant phenotype related breeding candidate gene of interest according to the enrichment result.

[0026] The statistical test described above can be a t-test, and the test statistic can be a t-test statistic.

[0027] The meaning of the presence or absence described above can be that the sequence consistent with the specific k-mer is detected in the plant to be tested, which is considered as presence, and the sequence inconsistent with the specific k-mer is not detected, which is considered as absence.

[0028] The similarity of the gene sequence in the same group described above is greater than or equal to 99%.

[0029] The variation described above can include presence (Presence) and / or absence (Absence) variation.

[0030] In order to solve the above technical problems, the present application also provides a device for constructing a model for predicting the phenotype of interest of a plant based on genotype, which can include the following modules:

[0031] B1, data receiving module: for receiving a database of specific k-mers and the measured value of the phenotype of interest of each plant in the plant population to be tested; the database of specific k-mers is obtained by a method comprising the following steps:

[0032] obtaining all known gene sequence data related to a target phenotype of a plant to be tested; grouping the gene sequence data based on sequence similarity to obtain similar gene sequence groups, and extracting the longest sequence gene in each of the similar gene sequence groups as a non-redundant phenotype-related gene; using sequence analysis software to screen specific k-mers of each gene in all the non-redundant phenotype-related genes, and using the sequence analysis software to generate a database containing specific k-mers of all the non-redundant phenotype-related genes;

[0033] B2, model construction module: for clustering specific k-mers in the database of specific k-mers according to the presence or absence in the plant population to be tested to generate 1000 clusters, randomly selecting the specific k-mers in each cluster to form an initial candidate set, using the initial candidate set and the target phenotype trait measurement value as input data, using the target phenotype trait measurement value as output data, and using a random forest algorithm to train a model for predicting plant phenotypes based on genotypes;

[0034] To solve the above technical problems, the present application also provides a method for plant breeding or quality improvement, which can be A or B;

[0035] The A can include the following steps: obtaining a target phenotype-related breeding candidate gene of a plant to be tested using the device described above, and selecting the plant to be tested containing the target phenotype-related candidate gene as a parent for breeding or quality improvement;

[0036] The B can include the following steps: predicting a target phenotype of a plant to be tested based on the model described above to obtain a phenotype prediction result, selecting a parent based on the phenotype prediction result, and using the parent for breeding or quality improvement related to the target phenotype.

[0037] To solve the above technical problems, the present application also provides a computer-readable storage medium storing a computer program, which can make a computer execute the following steps:

[0038] C1, data receiving: receiving sequence data of all known genes related to a target phenotype of a plant to be tested;

[0039] C2, data grouping: grouping the sequence data based on sequence similarity to obtain similar gene sequence groups, and extracting the longest sequence gene in each of the similar gene sequence groups as a non-redundant phenotype-related gene;

[0040] C3, determining specific k-mers of non-redundant phenotype-related genes: using sequence analysis software to screen specific k-mers of each gene of all the non-redundant phenotype-related genes, and using the sequence analysis software to generate a database containing specific k-mers of all the non-redundant phenotype-related genes;

[0041] C4, screening candidate genes from sequencing data of a plant to be tested: using the database to scan the sequencing data of the plant to be tested, calculating the detection rate of specific k-mers of each of the non-redundant phenotype-related genes in the plant to be tested, and obtaining genes similar to the non-redundant phenotype-related genes or genes with variations in the plant to be tested based on the detection rate, the similar genes or genes with variations being the phenotype-related genes or candidate genes of interest in the plant to be tested;

[0042] C5, obtaining phenotype-related genes of interest: receiving the measured values of the phenotype of interest of each plant in the plant population to be tested, for each specific k-mer, dividing the plant population to be tested into two groups according to the presence or absence of the specific k-mer, performing statistical tests on the measured values of the phenotype of interest of the plants in the two groups, taking the p value of the statistical test as the correlation effect, calculating the correlation score, screening the specific k-mers to obtain phenotype-related specific k-mers according to the correlation score and the correlation effect, and screening the similar genes or genes with variations to obtain phenotype-related genes of interest according to the phenotype-related specific k-mers;

[0043] The formula for calculating the correlation score is as follows:

[0044] Correlation score = -log 10 (p value) x |test statistic| / test statistic) Formula 1;

[0045] C6, obtaining plant phenotype-related breeding candidate genes of interest: detecting the percentage of the phenotype-related genes of interest in the cultivated varieties and wild varieties of the plant to be tested to obtain the enrichment results of the phenotype-related genes of interest in the cultivated varieties and local varieties, and screening to obtain plant phenotype-related breeding candidate genes of interest according to the enrichment results.

[0046] The k-mer described above is a fixed-length nucleotide string, which refers to a short DNA sequence of length k. For example: a certain gene is 1000 bp long, and the 1000 bp is broken into 29 bp short fragments, and the 29 bp short fragments after breaking are called 29-mers, and (1000-29+1) k-mer sequences can be obtained. The specific k-mer refers to a unique 29-mer among all the 29-mers of the candidate genes.

[0047] To solve the above technical problems, the present application also provides a computer program product comprising a computer program which, when executed by a processor, implements steps S1-S4:

[0048] S1, data receiving: receiving sequence data of all known genes related to a target phenotype of a plant to be tested;

[0049] S2, data grouping: grouping the sequence data based on sequence similarity to obtain similar gene sequence groups, and extracting the longest gene in each similar gene sequence group as a non-redundant phenotype-related gene;

[0050] S3, determining specific k-mer database of non-redundant phenotype-related genes: using sequence analysis software to screen for specific k-mers of each gene in all non-redundant phenotype-related genes, and using the sequence analysis software to generate a database containing specific k-mers of all non-redundant phenotype-related genes;

[0051] S4, detecting candidate genes in sequencing data of a plant to be tested: using the database to scan the sequencing data of the plant to be tested, calculating the detection rate of specific k-mers of each non-redundant phenotype-related gene in the plant to be tested, and based on the detection rate, obtaining genes similar to the non-redundant phenotype-related genes or genes with variations in the plant to be tested, which are the target phenotype-related genes or candidate genes in the plant to be tested.

[0052] The steps in the above computer program can further comprise C5 and / or C6 after S4.

[0053] The above computer device and / or any of the following applications of the above device also belong to the protection scope of the present application:

[0054] P1, application in breeding of a target phenotype trait of a plant;

[0055] P2, application in breeding or quality improvement of a target phenotype trait of a plant;

[0056] P3. Application in plant gene function annotation.

[0057] The above computer-readable storage medium and / or any of the following applications of the above computer program product also belong to the protection scope of the present application:

[0058] P1, application in breeding of a target phenotype trait of a plant;

[0059] P2, application in breeding or quality improvement of a target phenotype trait of a plant.

[0060] 10. Use according to claim 8 or 9, wherein the plant is any one of:

[0061] D1 ) a dicotyledonous plant;

[0062] D2) a monocotyledonous plant,

[0063] D3) a plant of the order Poales,

[0064] D4) a plant of the family Poaceae,

[0065] D5) a plant of the genus Triticum,

[0066] D6) a plant of the species Triticum aestivum.

[0067] The phenotype as described above can be the strength and amount of grain gluten or the SDS sedimentation value; the k-mer can be a 29-mer.

[0068] The plant as described above can be any one of:

[0069] D1 ) a dicotyledonous plant;

[0070] D2) a monocotyledonous plant,

[0071] D3) a plant of the order Poales,

[0072] D4) a plant of the family Poaceae,

[0073] D5) a plant of the genus Triticum,

[0074] D6) a plant of the species Triticum aestivum.

[0075] The sequencing depth of the sequencing data as described above can be low depth sequencing data of equal or greater than 4x, such as 4x.

[0076] The statistical test as described above can be a t-test, and the test statistic can be a t-test statistic.

[0077] The meaning of the presence or absence as described above can be that the detection of a sequence identical to the specific k-mer in the plant to be tested is the presence, and the non-detection of a sequence identical to the specific k-mer is the absence.

[0078] To overcome these obstacles and efficiently identify superior wheat grain Seed Storage Protein (SSP) alleles, the present application developed a PanSK (Pan-SSP k-mer) based on SSP gene resources for genotype-to-phenotype prediction. PanSK uses 29-mer sequences representing each SSP gene at the pan-genomic level to reveal unexplored genetic diversity among different wheat varieties and modern cultivars. Genome-wide association studies based on k-mers identified SSP genes associated with processing quality to improve wheat processing quality. And through machine learning-based prediction inspired by PanSK, the present application can accurately predict the processing quality phenotype only by genotype, for high-quality strong gluten wheat breeding.

[0079] Advantages of the present application:

[0080] The device or method of the present application can be applied to the screening, mining and cloning of phenotype-related candidate genes with numerous coding genes, complex structures and a large number of long short repeat sequences in genes, a large number of unmeasured gaps in coding sequences and flanking sequences, and insufficient gene annotation, as well as annotation, correlation analysis and molecular marker development of related gene functions. BRIEF DESCRIPTION OF DRAWINGS

[0081] Figure 1 For the establishment of PanSK.

[0082] Figure 2 For the analysis of alcohol-soluble protein PAV.

[0083] Figure 3 For association analysis and superior haplotype identification.

[0084] Figure 4 For genotype-to-phenotype prediction model construction. DETAILED DESCRIPTION

[0085] The present application will be further described in detail below with specific embodiments. The examples provided below are only intended to illustrate the present application, and are not intended to limit the scope of the present application. The examples provided below can serve as a guide for further improvement by those skilled in the art, and do not in any way constitute a limitation on the present application.

[0086] In the following examples, the experimental methods are conventional methods, and are performed according to the techniques or conditions described in the literature in the art or according to the product instructions, unless otherwise specified. The materials, reagents, etc. used in the following examples can be obtained commercially, unless otherwise specified.

[0087] Terminology:

[0088] The k-mer in the present application is a fixed-length nucleotide string, which refers to a short DNA sequence with a length of k. For example, a certain gene has a length of 1000 bp, and the 1000 bp is broken into short fragments with a length of 29 bp. The short fragments with a length of 29 bp after breaking are called 29-mers, and (1000-29+1) k-mer sequences can be obtained. The specific k-mer in the present application refers to a specific 29-mer that is unique among all candidate genes.

[0089] Example 1, Detection of wheat grain storage protein (SSP) alleles at the pan-genomic level based on low-depth sequencing data

[0090] 1. Obtaining of non-redundant grain storage protein sequences

[0091] The present application aims to detect and identify k-mers unique to each grain storage protein gene, and to represent the grain storage protein by k-mers. Further, k-mers are used to directly scan the second-generation sequencing data to quickly and accurately determine genetic variations, such as presence / absence variations and nucleotide polymorphisms in each variety. Accordingly, the present application develops a k-mer-based process, which is named PanSK, for detecting wheat grain storage protein genes by scanning raw second-generation sequencing data.

[0092] In order to obtain enough abundant grain storage protein sequences, a pan-genome level grain storage protein collection is established, and 649 SSP gene sequences are collected in the assembled reference genome of wheat, NCBI database and laboratory cloned grain storage protein sequences, of which 515 genes are from the annotation genes in 20 Triticeae genome assemblies (related literatures: 1: Chromosome-scale genome assembly of the transformation-amenable common wheat cultivar 'Fielder'; 2: Shifting the limits in wheat research and breeding using a fully annotated reference genome; 3: Comparative genomic and transcriptomic analyses uncover the molecular basis of high nitrogen-use efficiency in the wheat cultivar Kenong 9204; 4: Origin and adaptation to high altitude of Tibetan semi-wild wheat; 5: Optical maps refine the bread wheat Triticum aestivum cv.Chinese Spring genome assembly; 6: Multiple wheat genomes reveal global variation in modern breeding; 7: Durum wheat genome highlights past domestication signatures and future improvement targets; 8: Genome sequence of the progenitor of wheat A subgenome Triticum urartu: 9: A high-quality genome assembly highlights rye genomic characteristics and agronomically important genes: 10: Genome sequence of the progenitor of the wheat D genome Aegilops tauschii; 11: Improved Genome Sequence of Wild Emmer Wheat Zavitan with the Aid of Optical Maps), 125 genes were derived from Iso-Seq data collected from Nongda 3672 (Sequence Read Archive (SRA) accession number PRJNA1079223) and cultivated variety Xiaoyan 81 (related literature: gliadins, the dominant carriers of celiac disease epitopes), 9 published SSP sequences were obtained by Sanger sequencing; and the Markov clustering algorithm (MCL) was used to group SSP gene sequences with sequence similarity greater than 99% together, obtaining 139 groups of SSP similar gene sequence groups; then the longest SSP gene sequence in each group was selected as the representative non-redundant grain storage protein coding gene sequence, and finally 139 non-redundant grain storage protein sequences (rSSP) were obtained. Figure 1 Chinese A).

[0093] 2. Screening specific k-mers representing rSSP

[0094] All representative k-mers of the non-redundant SSPs (rSSPs) obtained at each step 1 were identified using the sequence analysis software JELLYFISH 2.2.10 (parameters '-t 15 -m k -s 4G-C', see A fast, lock-free approach for efficient parallel counting of occurrences of k-mers). Different k-mer sizes (k = 17, 19, 21, 23, 25, 27, 29, 31, 33) were tested in order to balance the computational complexity and the specificity of the k-mers. The 29 bp was chosen as the best k-mer length for the subsequent analysis. By using the JELLYFISH subcommand 'dump -L 1 -U 1' and selecting the specific 29-mers, a specific k-mer database of 139 rSSPs was generated: one database containing 40,453 specific 29-mers (Database 1) for representing the 139 non-redundant SSP genes. Figure 1

[0095] 3. Construction of the PanSK tool: detection of rSSP genes in resequencing data based on specific 29-mers

[0096] Based on the specific 29-mers obtained at step 2, the present invention constructed a PanSK tool for rapidly scanning the high-throughput raw sequencing data of the wheat to be tested using the database containing 40,453 specific 29-mers obtained at step 2, estimating the presence / absence of the SSP genes by calculating the relative detection ratio of the representative k-mers of each rSSP sequence detectable in the high-throughput raw sequencing data (detection ratio = number of k-mers present / total number of k-mers), evaluating the presence / absence of the grain storage proteins in the wheat to be tested, analyzing the variation of the prolamin proteins in the wheat to be tested, and obtaining the genes similar to each rSSP sequence or the sequences with variations in the wheat to be tested. Figure 1

[0097] 4. Determination of the parameters of the PanSK tool

[0098] ​​To evaluate the performance of PanSK in identifying presence / absence variants (PAVs) of SSPs, to determine the threshold of detection ratio for judging the presence of SSP genes, the present application used PanSK to scan representative specific k-mers of SSPs from simulated resequencing data of genome assemblies from 11 wheat varieties (Related literature: 1: Shifting the limits in wheat research and breeding using a fully annotated reference genome; 2: Optical maps refine the bread wheat Triticum aestivum cv. Chinese Spring genome assembly; 3: Multiple wheat genomes reveal global variation in modern breeding). The names and genome information of the 11 wheat varieties are as follows: Chinese Spring (IWGSC v1.0), Fielder, Norin 61, Stanley, SY Mattis, Landmark, Mace, Lancer, Arina LrFor, and Jagger. By comparing with annotated SSPs in the genome assemblies of these wheat varieties, the results of scanning using PanSK after simulating resequencing data showed that the detection ratio of k-mers for SSP genes increased with the increase of sequencing depth. Even in the sequencing data of low sequencing depth of 1x, the distribution of k-mer detection ratio of SSP genes obtained by PanSK scanning was distinguishable for the absence and presence of SSP genes, which indicated that the detection ratio of k-mers obtained by PanSK tool established by the present application could be used as a reliable indicator for inferring the presence (Presence) / absence (Absence) of SSP genes in the wheat to be tested Figure 1 In the middle B), detecting the presence or absence of grain storage proteins (SSPs) in the wheat to be tested.

[0099] In addition, by calculating the F-score (Precision x Recall / (Precision + Recall)) for evaluating the accuracy of PanSK, the results showed that it could accurately distinguish the presence / absence of SSP genes at a sequencing depth of 4x Figure 1 In the middle C).

[0100] The above analysis results found that in identifying grain storage proteins, PanSK can detect genome annotation errors, assembly errors or completely assembled grain storage protein genes based on short read or long read assembly ( Figure 1 in the middle D). Therefore, PanSK can directly and efficiently and quickly detect the presence / absence variation of wheat SSP genes, association analysis, etc. in high-throughput raw sequencing data without relying on reference genome mapping ( Figure 1 in the middle E).

[0101] Example 2, application example of PanSK tool of the present application

[0102] 1. Constructing SSP fingerprint based on PanSK analysis results

[0103] Using PanSK, the present application determines the similar genes or genes with presence variation of the 139 non-redundant SSP genes in step 1 of example 1 by the presence / absence variation in 365 resequenced wheat varieties (relevant literature: 1: Triticum population sequencing provides insights into wheat adaptation; 2: Resequencing of 145 Landmark Cultivars Reveals Asymmetric Sub-genome Selection and Strong Founder Genotype Effects on Wheat Breeding in China; 3: Frequent intra-and inter-species introgression shapes the landscape of genetic variation in bread wheat; 4: Origin and adaptation to high altitude of Tibetan semi-wild wheat), and establishes a SSP gene fingerprint ( Figure 2 in the middle A). The number of different types of SSP genes predicted in each variety varies greatly, with the number of members of alpha-gliadin (Gli-alpha) in one variety ranging from ≥10 to ≤30 ( Figure 2 in the middle B).

[0104] Among the 139 wheat rSSP genes, 8 grain storage proteins (rSSP) were present in more than 95% of the population (core genes) in 346 wheat materials. 26 rSSP genes were present in 80-95% of the population (secondary core genes) in 292-363 wheat materials. 76 rSSP genes were present in 5-80% of the population (non-core genes) in 18-293 wheat materials. 29 rSSP genes were present in less than 5% of the population (unique genes) in 0-17 wheat materials. The omega-gliadin gene has the highest proportion of unique genes and non-core genes (95%), reflecting its variability. Because non-core genes usually exist in the form of tandem repeats. Therefore, the higher variability count of gliadin genes at the population level contributes to the potential genetic plasticity of wheat processing quality Figure 2 C) in the middle.

[0105] 2. Association analysis based on PanSK analysis results and mining of new SSP candidate genes related to processing quality

[0106] In order to identify grain storage proteins related to processing quality, the present application uses a natural population with known SDS sedimentation value (SDS-SV) (Table 1), and relies on PanSK-based association analysis to mine new grain storage proteins related to quality. The specific method is that for each k-mer, the natural population is divided into two groups according to the presence or absence of the k-mer (detection of a sequence consistent with the specific k-mer is considered present, and detection of a sequence inconsistent with the specific k-mer is considered absent), and t-test statistical test is performed on the difference significance of SDS-SV of each group, and the p value is used as the association effect, and the positive and negative values of the t-test statistic are used to determine the positive and negative effects on quality, and finally the association score is calculated. Since the 1Dx+1Dy subunit pair of HMW-GSs and the secalin protein (grain storage protein in wheat-rye 1BL / 1RS translocation lines) have a significant impact on quality, the present application selects 103 wheat varieties without 1BL / 1RS translocation, which carry the same HMW-GS type (1Dx2+1Dy12) for association analysis. Among them, the calculation formula of the association score is as follows formula 1:

[0107] Association score = -log 10 (p value) x |t-test statistic| / t-test statistic Formula 1;

[0108] According to the association score and the association effect, the specific k-mer is screened, and the non-redundant phenotype-related gene is screened according to the specific k-mer. Finally, the present application identifies 336 k-mers from 23 genes that are associated with SDS-SV (|association score|≥7, p<1x10 -7) significantly associated genes, including one allele encoding HMW-GS 1Ax, 4 LMW-GS and 4 a-gliadin, 8 g-gliadin genes and 6 w-gliadin genes. Among them, 2 SSP genes are known to play a role in wheat processing quality, including 1 HMW-GS gene (TraesCS1A02G317311) and 2 g-gliadin genes Gli-g-1D-3

[0109] (TraesFLD1D01G005600) and Gli-g-1B-4 (TraesFLD1B01G010600). The remaining 20 SSP genes are new candidate genes for end-use quality Figure 3 A). Therefore, the new candidate genes obtained using the method of the present application provide a rich gene resource for wheat processing quality breeding.

[0110] Table 1. Wheat population and its corresponding SDS sedimentation value

[0111]

[0112]

[0113]

[0114] 3. Further screening of candidate genes based on SSP gene haplotypes

[0115] Further, the present application utilizes PansSK to clarify the variation of these genes by distinguishing the haplotypes of 23 SSP genes. A total of 63 haplotypes were identified (Table 2), including nucleotide polymorphisms and presence / absence variations. Among them, 31 haplotypes showed positive effects on quality, and 23 haplotypes showed negative effects on quality. In order to understand the utilization of these haplotypes in the breeding process, the present application calculates the percentage of each haplotype in landraces and cultivars to clarify whether the haplotype is enriched in cultivars or landraces (wild). Five haplotypes associated with high SDS-SV are enriched in cultivars, indicating that they have been selected and used to improve wheat processing quality in the breeding process. Four haplotypes associated with low SDS-SV are enriched in landraces, indicating that these genes may be eliminated in the breeding process. In addition, there are 25 haplotypes associated with high SDS-SV, but they have not been selected in the modern breeding process Figure 3 B), which are new candidate genes that can be used to improve wheat processing quality in the future.

[0116] Table 2. Haplotypes of 23 SSP genes and SDS-SV and their breeding selection scores have statistical significance

[0117]

[0118]

[0119]

[0120] As an example with Gli-γ-1B-3 encoding γ- alcohol-soluble protein, the present application extracts reads containing Gli-γ-1B-3 specific k-mers in the resequencing data, then uses BWA 0.7.15 to align the extracted reads to the Gli-γ-1B-3 sequence using Gli-γ-1B-3 sequence as the reference sequence, and then uses GATK v3.8

[0121] 'HaplotypeCaller' and 'FastaAlternateReferenceMaker' modules identify variations and assemble complete sequences. Finally, PanSK identifies that there are three haplotypes in the gene: Gli-γ-1B-3 h1 , Gli-γ-1B-3 h2 , and Gli-γ-1B-3 h3 . Resources carrying Gli-γ-1B-3 h2 or Gli-γ-1B-3 h3 haplotype have higher SDS-SV than resources carrying Gli-γ-1B-3 h1 (Gli-γ-1B-3 Figure 3 C). Gli-γ-1B-3 h2 and Gli-γ-1B-3 h3 haplotypes are dominant in improved varieties (34% and 63%), and the frequency of Gli-γ-1B-3 h3 haplotype in breeding varieties is higher than that in local varieties, indicating that Gli-γ-1B-3 h3 has been selected in breeding Figure 3 (D). The frequency of Gli-γ-1B-3 h1 in improved varieties is lower than that in local varieties, indicating that this haplotype has been lost. In order to compare the sequence variations among the three Gli-γ-1B-3 haplotypes, we reassembled their complete coding sequences based on k-mers. Gli-γ-1B-3 h1 and Gli-γ-1B-3 h2 differ by 19 SNPs, while Gli-γ-1B-3 h2 and Gli-γ-1B-3 h3 only differ by one SNP). Compared with Gli-γ-1B-3 h1 , Gli-γ-1B-3 h2 / 3A C to T change at 277 base pairs from the translation start site results in a premature stop codon Figure 3 E) These results suggest that loss-of-function alleles of Gli-γ-1B-3 contribute to improved processing quality in wheat.

[0122] 4. Construction of genotype-phenotype prediction model KPPer and its application in assisted breeding

[0123] The k-mers identified by PanSK can be used to study the presence / absence variation and allelic variation of SSP genes. Therefore, the present application develops a k-mer-based genotype-to-phenotype prediction model based on PanSK, and then realizes the precise prediction of processing quality using the presence / absence variation information of k-mers Figure 4 A) The present application first uses kmeans clustering to cluster the presence / absence of 40,453 specific k-mers in the population obtained in step 2 of Example 1 (detection of sequences consistent with specific k-mers is considered as present, and detection of sequences inconsistent with specific k-mers is considered as absent) into 1000 clusters, and randomly selects k-mers in each cluster to form an initial candidate set. Next, based on the initial candidate set and the SDS-SV phenotype determination value as input data, and the SDS-SV phenotype prediction value as output data, a random forest-based model is trained with the presence / absence variation Profile of k-mers as genotype and SDS-SV as phenotype, i.e. a genotype-based phenotype prediction model. In order to reduce the number of k-mers used as much as possible to facilitate the application of breeding, the present application starts from 1 k-mer, calculates the prediction performance of all k-mers, retains the best performing k-mers, and repeats the cycle. Finally, the present application found that the prediction performance of using 90 k-mers is sufficient to achieve stable high accuracy (Pearson correlation coefficient = 0.64) Figure 4 in B, C, D). Therefore, the present application develops a processing quality predictor called "KPPer" (k-mer-based phenotype predictor) by using the genetic variation of 90 selected k-mers (Table 3) as genotype.

[0124] Table 3. 90 selected k-mers in k-mer-based phenotype predictor

[0125]

[0126]

[0127] KPPer can predict phenotype from genotype, thus alleviating the limitations of traditional phenotyping methods and having great potential in breeding. Here, the present application proposes a new breeding strategy for improving processing quality using KPPer. First, the k-mers contained in germplasm resources are detected and evaluated, and parents that can complement each other are selected for combination. Then, the detection of 90 selected k-mers in the sequencing data of the wheat to be tested in the early stage of the offspring (such as F2 population) after the two parents are crossed, the prediction result of the processing quality phenotype of the wheat to be tested is obtained by KPPer, and the excellent individuals carrying more SSP genes with positive effects on processing quality are selected for the next generation screening. Finally, homozygous lines are selected in F3 or F4 generation. By detecting the composition of different k-mers in different varieties, the method of the present application can select complementary parent combinations to aggregate k-mers with positive effects on quality and exclude k-mers with negative effects on quality Figure 4 E). This will help to breed new high-quality, high-quality varieties.

[0128] The above describes the present application in detail. For those skilled in the art, the present application can be implemented within a wider range under equivalent parameters, concentrations and conditions without departing from the spirit and scope of the present application and without unnecessary experiments. Although the present application gives a special example, it should be understood that the present application can be further improved. In summary, according to the principle of the present application, the present application intends to include any change, use or improvement of the present application, including changes made by conventional techniques known in the art, which deviates from the scope disclosed in the present application.

Claims

1. A computer apparatus comprising a memory, a processor, and a computer program stored on the memory, wherein, The processor executes the computer program to implement steps S1-S4: S1, data receiving: receiving sequence data of all known genes related to a target phenotype of a plant to be tested; S2, data grouping: grouping the sequence data based on sequence similarity to obtain similar gene sequence groups, and extracting the longest sequence gene in each similar gene sequence group as a non-redundant phenotype-related gene; S3, determining a specific k-mer database of non-redundant phenotype-related genes: using sequence analysis software to screen and obtain the specific k-mer of each gene in all the non-redundant phenotype-related genes, and using the sequence analysis software to generate a database containing the specific k-mer of all the non-redundant phenotype-related genes; S4, detecting candidate genes in sequencing data of the plant to be tested: using the database to scan the sequencing data of the plant to be tested, calculating the detection rate of the specific k-mer of each non-redundant phenotype-related gene in the plant to be tested, and based on the detection rate, obtaining genes similar to the non-redundant phenotype-related genes or genes with variations in the plant to be tested, which are the target phenotype-related genes or candidate genes in the plant to be tested.

2. An apparatus for detecting target phenotype-related genes or candidate genes in a plant, comprising the following modules: A1, a data receiving module: for receiving sequence data of all known genes related to a target phenotype of a plant to be tested; A2, a data grouping module: for grouping the sequence data based on sequence similarity to obtain similar gene sequence groups, and extracting the longest sequence gene in each similar gene sequence group as a non-redundant phenotype-related gene; A3, a specific k-mer database determination module for obtaining non-redundant phenotype-related genes: for using sequence analysis software to screen and obtain the specific k-mer of each gene in all the non-redundant phenotype-related genes, and using the sequence analysis software to generate a database containing the specific k-mer of all the non-redundant phenotype-related genes; A4, a candidate gene detection module for scanning the sequencing data of the plant to be tested: for using the database to scan the sequencing data of the plant to be tested, calculating the detection rate of the specific k-mer of each non-redundant phenotype-related gene in the plant to be tested, and based on the detection rate, obtaining genes similar to the non-redundant phenotype-related genes or genes with variations in the plant to be tested, which are the target phenotype-related genes or candidate genes in the plant to be tested.

3. A method for mining breeding candidate genes related to a target phenotype trait in a plant, comprising the following steps: C1, data receiving: receiving sequence data of all known genes related to a target phenotype of a plant to be tested; C2, data grouping: grouping the sequence data based on sequence similarity to obtain similar gene sequence groups, and extracting the longest sequence gene in each similar gene sequence group as a non-redundant phenotype-related gene; C3, determining specific k-mer database of specific k-mer of non-redundant phenotype-related genes: using sequence analysis software to screen specific k-mer of each gene of all the non-redundant phenotype-related genes, and using the sequence analysis software to generate a database containing specific k-mer of all the non-redundant phenotype-related genes; C4, screening candidate genes by detecting sequencing data of the plant to be tested: using the database to scan the sequencing data of the plant to be tested, calculating the detection rate of specific k-mer of each non-redundant phenotype-related gene in the plant to be tested, and obtaining genes similar to the non-redundant phenotype-related genes or genes with variations in the plant to be tested based on the detection rate, wherein the similar genes or genes with variations are the purpose phenotype-related genes or candidate genes in the plant to be tested; C5, obtaining purpose phenotype significantly related genes: receiving the purpose phenotype trait measurement value of each plant in the plant population to be tested, for each specific k-mer, dividing the plant population to be tested into two groups according to the presence or absence of the specific k-mer, performing statistical test on the purpose phenotype trait measurement value of the plants in the two groups, taking the p value of the statistical test as the correlation effect, calculating the correlation score, screening the specific k-mer to obtain phenotype-related specific k-mer according to the correlation score and the correlation effect, and screening the similar genes or genes with variations to obtain purpose phenotype significantly related genes according to the phenotype-related specific k-mer; The formula of the correlation score is as follows: Association score = -log 10 (p value) x |test statistic| / test statistic) Equation 1; C6, obtaining plant purpose phenotype trait-related breeding candidate genes: detecting the percentage of the phenotype significantly related genes in the cultivated varieties and wild varieties of the plant to be tested to obtain the enrichment result of the phenotype significantly related genes in the cultivated varieties and local varieties, and screening to obtain plant purpose phenotype trait-related breeding candidate genes according to the enrichment result.

4. A model construction device for predicting plant purpose phenotypes based on genotypes, the device comprising the following modules: B1, data receiving module: used for receiving a database of specific k-mers and purpose phenotype trait measurement values of each plant in a plant population to be tested; the database of specific k-mers is obtained by a method comprising the following steps: obtaining all known gene sequence data related to the purpose phenotype of the plant to be tested; grouping the gene sequence data based on sequence similarity to obtain similar gene sequence groups, and extracting the longest sequence gene in each similar gene sequence group as a non-redundant phenotype-related gene; using sequence analysis software to screen specific k-mer of each gene of all the non-redundant phenotype-related genes, and using the sequence analysis software to generate a database containing specific k-mer of all the non-redundant phenotype-related genes; B2, model construction module: for clustering specific k-mers in the database of specific k-mers according to the presence or absence in the plant population to be tested to generate 1000 clusters, randomly selecting specific k-mers in each cluster to form an initial candidate set, using the initial candidate set and the measured value of the target phenotype as input data, using the measured value of the target phenotype as output data, and using a random forest algorithm to train a model for predicting plant phenotypes based on genotypes.

5. The method for plant breeding or quality improvement, which is A or B; A comprises the following steps: obtaining a breeding candidate gene related to the target phenotype of the plant to be tested using the device of claim 2, and selecting the plant to be tested containing the candidate gene related to the target phenotype as a parent for breeding or quality improvement; B comprises the following steps: predicting the phenotype of the plant to be tested based on the model of claim 3 to obtain a phenotype prediction result, selecting a parent based on the phenotype prediction result, and using the parent for breeding or quality improvement related to the target phenotype.

6. A computer readable storage medium storing a computer program, wherein the computer program causes a computer to execute the following steps: C1, data receiving: receiving sequence data of all known genes related to the target phenotype of the plant to be tested; C2, data grouping: grouping the sequence data based on sequence similarity to obtain similar gene sequence groups, and extracting the longest gene in each similar gene sequence group as a non-redundant phenotype-related gene; C3, determining specific k-mers of non-redundant phenotype-related genes to obtain a specific k-mer database: using sequence analysis software to screen and obtain specific k-mers of each gene in all non-redundant phenotype-related genes, and using the sequence analysis software to generate a database containing specific k-mers of all non-redundant phenotype-related genes; C4, detecting sequencing data of the plant to be tested to screen candidate genes: using the database to scan the sequencing data of the plant to be tested, calculating the detection rate of specific k-mers of each non-redundant phenotype-related gene in the plant to be tested, and based on the detection rate, obtaining genes similar to the non-redundant phenotype-related genes or genes with variations in the plant to be tested, which are the target phenotype-related genes or candidate genes in the plant to be tested; C5. Obtaining phenotype significantly associated genes: receiving the measured value of the phenotype trait of each plant in the plant population to be tested, for each specific k-mer, dividing the plant population to be tested into two groups according to the presence or absence of the specific k-mer, performing statistical test on the measured value of the phenotype trait of the plants in the two groups, taking the p value of the statistical test as the association effect, calculating the association score, screening the specific k-mer according to the association score and the association effect to obtain phenotype associated specific k-mer, and screening the similar genes or genes with variations according to the phenotype associated specific k-mer to obtain phenotype significantly associated genes; The calculation formula of the association score is as follows: Association score = -log 10 (p value) x |test statistic| / test statistic) Equation 1; C6. Obtaining plant phenotype trait associated breeding candidate genes: detecting the percentage of the phenotype significantly associated genes in the cultivated varieties and wild varieties of the plant to be tested to obtain the enrichment results of the phenotype significantly associated genes in the cultivated varieties and local varieties, and screening to obtain plant phenotype trait associated breeding candidate genes according to the enrichment results.

7. A computer program product comprising a computer program, characterized in that, The computer program is executed by the processor to realize the steps including S1-S4: S1. Data receiving: receiving the sequence data of all known genes associated with the phenotype of the plant to be tested; S2. Data grouping: grouping the sequence data based on sequence similarity to obtain similar gene sequence groups, and extracting the longest sequence gene in each similar gene sequence group as a non-redundant phenotype associated gene; S3. Determining specific k-mer database of non-redundant phenotype associated genes: using sequence analysis software to screen the specific k-mer of each gene in all the non-redundant phenotype associated genes, and using the sequence analysis software to generate a database containing the specific k-mer of all the non-redundant phenotype associated genes; S4. Detecting the sequencing data of the plant to be tested to screen candidate genes: using the database to scan the sequencing data of the plant to be tested, calculating the detection rate of the specific k-mer of each non-redundant phenotype associated gene in the plant to be tested, and obtaining similar genes or genes with variations in the plant to be tested based on the detection rate, which are the phenotype associated genes or candidate genes in the plant to be tested.

8. The computer device of claim 1 and / or the device of claim 2 or 3 is applied to: P1. In the breeding of plant phenotype traits; P2. In the selection or quality improvement of plant phenotype traits; P3. In the application.

9. The computer readable storage medium of claim 6 and / or the computer program product of claim 7 is applied to: P1. In the breeding of plant phenotype traits; P2. In the selection or quality improvement of plant phenotype traits.

10. Use according to claim 8 or 9, characterized in that: The plant is any of the following: D1) dicotyledonous plants; D2) monocotyledonous plants, D3) plants of the order Poales, D4) plants of the family Poaceae, D5) a plant of the genus Triticum, D6) wheat; The phenotypes of interest are the strength and quantity of grain gluten; the k-mers are 29-mers.