A method and system for rapid annotation of genome sequencing data based on site mapping
Through a site mapping-based method, an index file is established to quickly annotate DNA methylation sequencing data, solving the problem of inefficient annotation in the prior art and achieving efficient genome sequencing data annotation.
Patent Information
- Application Number
- CN202211165115.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-09-23
- Publication Date
- 2025-06-06
- Estimated Expiration
- 2042-09-23
AI Technical Summary
Existing bioinformatics software lacks consistent and rapid approaches when annotating DNA methylation sequencing data, especially in the promoter region, exon region, intron region and CpG Island region.
Using a site mapping-based approach, genome sequencing data is quickly annotated by establishing an index file. The method includes calculating the mapping values for each site and building an index file to quickly search and annotate the pending sites.
It significantly improves the annotation efficiency. For example, on Human chromosome 1, the search efficiency can be increased by 26,804 times, which is suitable for annotation of multiple chromosomes and multiple samples.
Smart Images

Figure CN115455920B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of bioinformatics, and in particular, relates to a method and system for rapid annotation of genome sequencing data based on site mapping. Background Art
[0002] Next-generation sequencing (NGS), also known as high-throughput sequencing, is a sequencing-by-synthesis technology developed based on PCR and gene chips. The main characteristics of high-throughput sequencing technology are: short sequencing read length, high throughput, and high accuracy. Compared with first-generation sequencing, high-throughput sequencing greatly reduces costs while maintaining high accuracy and greatly reducing sequencing time. Currently, high-throughput sequencing has been widely used in the whole-omics field. For example: transcriptome sequencing, resequencing, DNA methylation sequencing, m6A methylation sequencing, single-cell sequencing, etc.
[0003] DNA methylation is the main form of epigenetic modification. It can change genetic expression without changing the DNA sequence, and plays an important role in regulating gene expression and chromatin conformation. DNA methylation mainly forms 5-methylcytosine (5-mC) and a small amount of N6-methylpurine (N6-mA) and 7-methylguanine (7-mG). Generally, methylated DNA mainly refers to 5-methylcytosine (5mC). Methylation in mammalian cells mainly occurs on cytosine of CG dinucleotides, while a large proportion of non-CG (CHH, CHG, H represents A, C, T) methylation exists in plant cells. 5-Methylcytosine (5-mC) is converted to 5-methylcytosine (mC) by DNA methyltransferase (DNMT) catalyzing S-adenosylmethionine (SAM) as a methyl donor.
[0004] Whole-genome bisulphite sequencing (WGBS) combines bisulfite conversion with next-generation sequencing technology to efficiently detect the methylation status of whole-genome DNA at a single-base resolution level. Bisulfite treatment can deaminize unmethylated cytosine in DNA into uracil, while methylated cytosine remains unchanged; for the required fragments amplified by PCR, all uracil is converted into thymine. High-throughput sequencing of PCR products and comparison with reference sequences can determine whether CpG / CHG / CHH sites are methylated. Whole-genome methylation sequencing can comprehensively and accurately detect the methylation status of whole-genome DNA, laying the foundation for more in-depth epigenetic regulation analysis.
[0005] CpG Island in the gene promoter region is usually in a demethylated state, promoting gene transcription, while abnormal methylation can lead to transcriptional inactivation. Generally speaking, CpG Island methylation can lead to gene silencing. DNA methylation plays an important role in genomic imprinting. Hypermethylation of one of the biallelic genes can lead to monoallelic expression.
[0006] Current bioinformatics software does not have a consistent and fast annotation method for DNA methylation sequencing data in gene structural regions such as promoters, exons, introns, and intergenics, as well as CpGIsland regions. Summary of the invention
[0007] In order to solve at least one of the above technical problems, the technical solution adopted by the present invention is as follows:
[0008] The first aspect of the present invention provides a method for rapid annotation of genome sequencing data based on site mapping, comprising the following steps:
[0009] S1, create index file:
[0010] The starting and ending sites of the functional component region of the species from which the sequencing sample originated are obtained. For each site, the mapping value is obtained using formula (1):
[0011]
[0012] Among them, G i Represents the mapping value of the i-th site, INT represents the rounding operation, S i represents the value of the ith site, N is the value determined according to the chromosome length of the source species, and L iRepresents the number of digits at the i-th site. If L i ≤N then L i -N=1,
[0013] The mapping values of the start and end sites of all functional component regions are obtained, and the index file is constructed in the following format:
[0014] Chr SE se function
[0015] Among them, Chr represents the chromosome location information of the functional component region, S represents the mapping value of the starting site of the functional component region, E represents the mapping value of the ending site of the functional component region, s represents the starting site of the functional component region, E represents the ending site of the functional component region, and function represents the category of the functional component region;
[0016] S2, obtaining the mapping value of the site to be annotated: the site value is Q, and the mapping value G of the site to be annotated is obtained by using formula (1);
[0017] S3, searching the mapping value G obtained in step S2 in the second and third columns of the index file, if for a certain functional component region j, G satisfies S j ≤G≤E j , further determine whether Q satisfies s j ≤Q≤e j , if satisfied, the site to be annotated can be annotated to be located in the jth functional component region.
[0018] In the present invention, the functional components and functional elements have equivalent meanings.
[0019] In the present invention, the index value and the mapping value have the same meaning.
[0020] In some embodiments of the present invention, the method for determining N is as follows:
[0021] (1) Obtain the length CL and gene number GN of each chromosome, and calculate CL / GN;
[0022] (2) Obtain the representative number MN of all chromosomes CL / GN, divide it by the value q, and the integer number of the result of MN / q is the value N, where q = 1 to 100.
[0023] Here, the acquisition of the N value is an unexpected discovery of the present invention that can make the post-processing annotation more efficient. Those skilled in the art can also select the N value in other ways, as long as they do not violate the core idea of the present invention, they should be deemed to fall within the scope of protection of the present invention.
[0024] For example, the present invention can obtain the representative number of all chromosome lengths, take the square root of the representative number, and obtain the integer number of the result as the N value.
[0025] In some embodiments of the present invention, the representative number is selected from one of the median, mode, and mean.
[0026] In some embodiments of the present invention, the source species is a mammal. Preferably, the source species is a human.
[0027] In some embodiments of the present invention, the gene sequencing data refers to DNA methylation sequencing data.
[0028] In some embodiments of the present invention, the functional component regions include promoter regions, exon regions, intron regions, promoter CGIs, intragenic CGIs, 3' transcript CGIs, intergenic CGIs, repeat regions and miRNA regions.
[0029] Among them, promoter CGIs, intragenic CGIs, 3'transcript CGIs, and intergenic CGIs are defined based on the gene location to which the CGI belongs:
[0030] promoter CGIs -1000bp TSS to +300bp TSS intragenic CGIs +300bp TSS to +300bp TES 3'transcript CGIs -300bp TES to +300bp TES intergenic CGIs -300bp TES to-1000bp next gene's promoter
[0031] The second aspect of the present invention provides a rapid annotation system for genome sequencing data based on site mapping, comprising the following modules:
[0032] The index library module is used to store index files, wherein the index file is constructed as follows:
[0033] The starting and ending sites of the functional component region of the species from which the sequencing sample originated are obtained. For each site, the mapping value is obtained using formula (1):
[0034]
[0035] Among them, represents the mapping value of the th site, INT represents the rounding operation, S i represents the value of the ith site, N is the value determined according to the chromosome length of the source species, and L i Represents the number of digits at the i-th site. If L i ≤N then L i -N=1,
[0036] The mapping values of the start and end sites of all functional component regions are obtained, and the index file is constructed in the following format:
[0037] Chr SE se function
[0038] Among them, Chr represents the chromosome location information of the functional component region, S represents the mapping value of the starting site of the functional component region, E represents the mapping value of the ending site of the functional component region, s represents the starting site of the functional component region, E represents the ending site of the functional component region, and function represents the category of the functional component region.
[0039] The input module is used to receive sequencing data, obtain the sites to be annotated, and calculate the index value of the sites to be annotated using formula (1).
[0040] The search module is connected to the input module and the index library module respectively, and is used to search the index value of the site to be annotated obtained by the input module in the second column and the third column of the index file. If for a certain functional component region j, G satisfies S j ≤G≤E j , further determine whether Q satisfies s j ≤Q≤e j , if satisfied, the site to be annotated can be annotated to be located in the jth functional component region,
[0041] The result output module is used to output the annotation results.
[0042] In some embodiments of the present invention, the method for determining N is as follows:
[0043] (1) Obtain the length CL and gene number GN of each chromosome, and calculate CL / GN;
[0044] (2) Obtain the representative number MN of all chromosomes CL / GN, divide it by the value q, and the integer number of the result of MN / q is the value N, where q = 1 to 100.
[0045] As above, here, the acquisition of the N value is an unexpected discovery of the present invention that can make the post-processing annotation more efficient. Those skilled in the art can also select the N value in other ways, as long as they do not violate the core idea of the present invention, they should be deemed to fall within the protection scope of the present invention.
[0046] For example, the present invention can obtain the representative number of all chromosome lengths, take the square root of the representative number, and obtain the integer number of the result as the N value.
[0047] In some embodiments of the present invention, the representative number is selected from one of the median, mode, and mean.
[0048] In some embodiments of the present invention, the source species is a mammal. Preferably, the source species is a human.
[0049] In some embodiments of the present invention, the gene sequencing data refers to DNA methylation sequencing data.
[0050] In some embodiments of the present invention, the functional component regions include promoter regions, exon regions, intron regions, promoter CGIs, intragenic CGIs, 3' transcript CGIs, intergenic CGIs, repeat regions and miRNA regions.
[0051] Among them, promoter CGIs, intragenic CGIs, 3'transcript CGIs, and intergenic CGIs are defined according to the gene location to which the CGI belongs:
[0052] promoter CGIs -1000bp TSS to +300bp TSS intragenic CGIs +300bp TSS to +300bp TES 3'transcript CGIs -300bp TES to +300bp TES intergenic CGIs -300bp TES to-1000bp next gene's promoter
[0053] Beneficial effects of the present invention
[0054] Compared with the prior art, the present invention has the following beneficial effects:
[0055] By using the method and system of the present invention, by mapping the positions of functional components, the method is simple and easy to operate, and can greatly improve the efficiency of search annotation. Taking human chromosome 1 as an example, the search efficiency can be improved by 26804 times. For the annotation of multiple chromosomes and multiple samples, the effect is more obvious. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 The gene locations to which the CGIs belong are indicated.
[0057] Figure 2 A schematic diagram showing a site (10540) located in a promoter region ([10300, 13000]).
[0058] Figure 3 A schematic diagram of a rapid annotation system for genome sequencing data based on site mapping in an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0059] Unless otherwise indicated, implied from the context, or customary in the prior art, all parts and percentages in this application are based on weight, and the test and characterization methods used are all current as of the filing date of this application. Where applicable, the contents of any patent, patent application or publication referred to in this application are fully incorporated herein by reference, and their equivalent patent families are also incorporated by reference, especially the definitions of synthesis techniques, product and processing designs, polymers, comonomers, initiators or catalysts disclosed in these documents in the art. If the definition of a specific term disclosed in the prior art is inconsistent with any definition provided in this application, the definition of the term provided in this application shall prevail.
[0060] Numerical ranges in this application are approximate values, so unless otherwise specified, they may include numerical values outside the range. Numerical ranges include all numerical values from the lower limit to the upper limit increased by 1 unit, provided that there is an interval of at least 2 units between any lower value and any higher value. For example, if recorded is 100 to 1000, it means that all single numerical values are clearly listed, such as 100, 101, 102, etc., and all subranges, such as 100 to 166, 155 to 170, 198 to 200, etc. For a range containing a numerical value less than 1 or containing a fraction greater than 1 (such as 1.1, 1.5, etc.), 1 unit is appropriately regarded as 0.0001, 0.001, 0.01 or 0.1. For a range containing a single digit less than 10 (such as 1 to 5), 1 unit is generally regarded as 0.1. These are only specific examples of what is intended to be expressed, and all possible combinations of numerical values between the lowest and highest values listed are considered to be clearly recorded in this application.
[0061] The terms "comprising", "including", "having" and their derivatives do not exclude the presence of any other components, steps or processes, and are irrelevant to whether these other components, steps or processes are disclosed in this application. To eliminate any doubt, unless explicitly stated, all compositions using the terms "comprising", "including", or "having" in this application may include any additional additives, adjuvants or compounds. On the contrary, except for those necessary for operating performance, the term "essentially consisting of..." excludes any other components, steps or processes from the scope of any description of the term below. The term "consisting of..." does not include any components, steps or processes that are not specifically described or listed. Unless explicitly stated, the term "or" refers to the listed individual members or any combination thereof.
[0062] In order to make the technical problems, technical solutions and beneficial effects solved by the present invention more clearly understood, the present invention is further described in detail below in conjunction with embodiments.
[0063] Example
[0064] The following examples are used to demonstrate preferred embodiments of the present invention. It will be appreciated by those skilled in the art that the techniques disclosed in the following examples represent techniques discovered by the inventors that can be used to implement the present invention and therefore can be considered as preferred embodiments of the present invention. However, it will be appreciated by those skilled in the art based on this specification that many modifications may be made to the specific embodiments disclosed herein and still achieve the same or similar results without departing from the spirit or scope of the present invention.
[0065] Unless defined otherwise, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention belongs, and the disclosure and materials cited therein are hereby incorporated by reference.
[0066] Those skilled in the art will recognize, or be able to ascertain using no more than routine experimentation, many technical equivalents to the specific embodiments of the invention described herein. Such equivalents are intended to be encompassed by the claims.
[0067] The experimental methods in the following examples are conventional methods unless otherwise specified. The instruments and equipment used in the following examples are conventional laboratory instruments and equipment unless otherwise specified; the experimental materials used in the following examples are purchased from conventional biochemical reagent stores unless otherwise specified.
[0068] Example 1 Rapid annotation method for functional components based on DNA methylation sequencing
[0069] 1. Classification and definition of genome functional components
[0070] A promoter is a DNA sequence to which proteins bind to initiate transcription of a single RNA transcript from the DNA downstream of the promoter. The RNA transcript may encode a protein (mRNA) or may have its own function (such as tRNA or rRNA). The promoter is located near the gene transcription start site, upstream of the DNA (towards the 5' region of the sense strand). The promoter is about 100-1000 base pairs in length, and its sequence is highly dependent on the gene and transcript, the type or class of RNA polymerase recruited to the site, and the species of organism. The promoter region is the binding region for RNA polymerase, and its structure is directly related to the efficiency of transcription.
[0071] The transcription start site (TSS) refers to the base on the DNA chain corresponding to the first nucleotide of the newly generated RNA chain, usually a purine. The sequence before the starting point, i.e. the 5' end, is often called upstream, while the sequence after it, i.e. the 3' end, is called downstream. When describing the position of the base, it is generally represented by numbers, with the TSS starting point being +1, the downstream direction being +2, +3, etc., and the upstream direction being -1, -2, -3, etc.
[0072] In this embodiment, the inventors uniformly defined the promoter region based on TSS as a unified annotation standard for subsequent DNA methylation sequencing analysis, as shown in Table 1:
[0073] Table 1 Definition of promoter region
[0074] Promoter region promoters -2,200 to +500 bp Proximal Proximal(P) -200 to +500 bp Mid-range Intermediate(I) -200 to -1,000 bp remote Distal(D) -1,000 to -2,200 bp
[0075] The relative position and CpG content in the promoter sequence are important factors affecting the degree of promoter methylation, namely the O / E ratio (observed-to-expected CpG ratio). According to this ratio, the inventors divided the O / E values of CpG in the promoter region into three categories: low, medium, and high (LCP, ICP, and HCP). The calculation formula is as follows:
[0076]
[0077] Among them, Num of CpG represents the number of CpG, Num of C represents the number of base C in the sequence, Num of G represents the number of base G in the sequence, and Total number of Nucleotides in the sequence expresses the total number of bases in the sequence.
[0078] The inventors defined CpG islands (CGIs) according to the gene locations to which they belong as described in Table 2:
[0079] Table 2 CGI definition
[0080] promoter CGIs -1000bp TSS to +300bp TSS intragenic CGIs +300bp TSS to +300bp TES 3'transcript CGIs -300bp TES to +300bp TES intergenic CGIs -300bp TES to -1000bp next gene promoter
[0081] The gene location to which the CGI belongs is defined as Figure 1 shown.
[0082] 2. Use site translation method to annotate the functional component region of the sequence
[0083] According to the results of DNA methylation sequencing, the inventors aligned the sequencing reads to the genome, and the alignment results are usually output in SAM format. The SAM format includes site position information, POS: the leftmost position on the alignment, that is, the position of the first base on the genome where the reads are aligned. Based on the position of this alignment, people in this field need to quickly annotate which functional region of the gene this site is in (such as the promoter region, exon region, intron region), and also need to confirm whether it is in the CGI / CGI shores region.
[0084] Here, taking the human genome as an example, the length of human chromosome 1 is 249250621 bases. Assuming that some reads are aligned to the 10540th base position of the chromosome, how can we quickly find out whether the position (10540) belongs to the promoter region through site search?
[0085] Assume that the site is located in the promoter region ([10300,13000]) of one of the genes (e.g. Figure 2 As shown in FIG. 1 , if the single site translation method is used to traverse and search for annotations, it will take 10,540 cycles (starting from the first base) to find that this site is located in the promoter region of the genome. This method is very inefficient to execute.
[0086] 3. Site mapping method to annotate the functional component region of the sequence
[0087] In order to improve the search efficiency of functional annotation, the inventors created a rapid annotation site mapping search method, the detailed process is as follows:
[0088] (1) First, the inventors used the following formula to create a mapping of the functional components of the genome:
[0089]
[0090] Among them, G i represents the mapping value of the ith site, S i represents the value of the ith site, N is the value determined according to the chromosome length of the source species, and L i Represents the number of digits at the i-th site. If L i ≤N then L i -N=1.
[0091] The value of N is selected according to the following method:
[0092] 1. Obtain the length (i.e., number of bp) CL and number of genes GN of each chromosome, and calculate CL / GN;
[0093] 2. Obtain the median MN of the CL / GN of all chromosomes, divide it by the value q, and the integer digit of the result of MN / q is the N value (for example: 10000, ie N=4), where q=1~100.
[0094] For humans, the calculation results are shown in Table 3:
[0095] Table 3 Chromosome length information
[0096] Chromosome number CL(Mbp) GN CL / GN 1 250 4778 5232.31 2 242 3600 6722.22 3 198 2780 7122.30 4 190 2251 8440.69 5 182 2393 7605.52 6 171 2828 6046.68 7 159 2598 6120.09 8 145 2003 7239.14 9 138 2132 6472.80 10 134 2042 6562.19 11 135 2806 4811.12 12 133 2398 5546.29 13 114 1327 8590.81 14 107 1951 5484.37 15 102 1721 5926.79 16 90 1845 4878.05 17 83 2326 3568.36 18 80 938 8528.78 19 59 2420 2438.02 20 64 1243 5148.83 21 47 726 6473.83 22 51 1129 4517.27 X 156 1973 7906.74 Y 57 496 11491.94
[0097] The median MN of CL / GN is 6296.44. If q=1, then MN / q=6296.44, N=4; if q=10, then MN / q=629.6, N=3; if q=100, then MN / q=62.96, N=2.
[0098] The mapping values of the start and end sites of all functional component regions are obtained, and the index file is constructed in the following format:
[0099] Chr SE se function
[0100] Among them, Chr represents the chromosome location information of the functional component region, S represents the mapping value of the starting site of the functional component region, E represents the mapping value of the ending site of the functional component region, s represents the starting site of the functional component region, E represents the ending site of the functional component region, and function represents the category of the functional component region.
[0101] When the mapping value of the site to be annotated is obtained, the site value is Q, and the mapping value G of the site to be annotated is obtained by the above formula; then the mapping value G is searched in the second and third columns of the index file. If for a certain functional component region j, G satisfies S j ≤G≤E j , further determine whether Q satisfies s j ≤Q≤e j , if satisfied, the site to be annotated can be annotated as being located in the jth functional component region, thereby quickly completing the annotation.
[0102] Example 2 Application of annotation method based on site mapping.
[0103] In order to quickly annotate functional components using the method established in Example 1, the inventors selected q=10, so N=3.
[0104] Taking the above promoter region [10300, 13000] as an example, the start site of the promoter region is (10300), S i =10300,L i =5, G i =10300 mod 103 +5×10 2 =510.
[0105] Similarly, the index value of the promoter region termination site (13000) is: G j =513.
[0106] The inventors used this method to create the following index for the promoter region for storage, and the format is as follows:
[0107] Chr S E s e Function Chr1 510 513 10300 13000 promoter
[0108] Among them, Chr represents the chromosome location information of the functional component region, S represents the mapping value of the starting site of the functional component region, E represents the mapping value of the ending site of the functional component region, s represents the starting site of the functional component region, E represents the ending site of the functional component region, and function represents the category of the functional component region. The categories of functional component regions include but are not limited to: promoter region, exon region, intronic region, promoter CGIs, intragenic CGIs, 3'transcript CGIs, intergenic CGIs, repeat region and miRNA region.
[0109] A similar index conversion is performed on all known functional components to obtain an index file. It should be noted that these functional components are stored in the complete position interval of the starting position and the ending position of the functional component on the genome.
[0110] Through this method, the inventors compressed the original sites. The corresponding original sites of the interval [510,513] are: [10000,13999]. That is, 4 sites are used to store the original 4000 sites. The search efficiency can be improved by 1000 times. For large genomic data, the effect is very obvious.
[0111] When annotating, you only need to annotate the site and convert the corresponding index value before annotating:
[0112]
[0113] When searching, according to G k Go to the index file to search for the second column (S) and the second column (E) data. At the same time, the inventor checks whether the queried location falls within the interval of the third column and the fourth column. If S≤G k ≤E, and s≤Q k ≤e then Q k This site can be annotated into the corresponding functional component.
[0114] For example, for the above-mentioned site 10540 (Q k ), use the above method to convert the corresponding index value, the converted value is as follows:
[0115] G k =10540 mod 10 3 +5×10 2 =510
[0116] This site satisfies 510≤510≤513, and 10300≤10540≤13000, so site Q k Annotated as a promoter.
[0117] Example 3 Application of the rapid annotation method of functional components based on DNA methylation sequencing in the whole chromosome
[0118] This example uses chromosome 1 as an example. The sequence of chromosome 1 has a total of 249,250,621 bases. If the traditional site traversal method is used, it takes at most 249,250,621 traversals to find the annotated site. According to the site mapping search method of Example 1, N=4, which will be divided into 9249 index values at most (about 50 functional components), that is, at most 9299 searches can find the site, and the search efficiency is improved by 26,804 times.
[0119] The specific efficiency improvement comparison is shown in Table 3:
[0120] Table 3 Efficiency comparison
[0121]
[0122] It can be seen that the site mapping search method of Example 1 can greatly improve the search performance.
[0123] Example 4 Rapid annotation system for genome sequencing data based on site mapping
[0124] like Figure 3 As shown, this embodiment provides a system to implement the above-mentioned fast annotation method, and the system includes:
[0125] Index library module, used to store the index files built above
[0126] The input module is used to receive sequencing data, obtain the site to be annotated Q, and calculate the index value G of the site to be annotated using the above formula.
[0127] The search module is connected to the input module and the index library module respectively, and is used to search the index value of the site to be annotated obtained by the input module in the second column and the third column of the index file. If for a certain functional component region j, G satisfies Sj ≤G≤E j , further determine whether Q satisfies s j ≤Q≤e j , if satisfied, the site to be annotated can be annotated to be located in the jth functional component region,
[0128] The result output module is used to output the annotation results.
[0129] All documents mentioned in the present invention are cited as references in this application, just as each document is cited as reference individually. In addition, it should be understood that after reading the above teachings of the present invention, those skilled in the art can make various changes or modifications to the present invention, and these equivalent forms also fall within the scope defined by the claims attached to this application.
Claims
1. A rapid annotation method for genome sequencing data based on site mapping, It is characterized in that The following steps are involved: S1, create index file: The starting and ending sites of the functional component region of the species from which the sequencing sample originated are obtained. For each site, the mapping value is obtained using formula (1): Among them, G i Represents the mapping value of the i-th site, INT represents the rounding operation, S i represents the value of the ith site, N is the value determined according to the chromosome length of the source species, and L i Represents the number of digits at the i-th site. If L i ≤N then L i -N=1, The mapping values of the start and end sites of all functional component regions are obtained, and the index file is constructed in the following format: Chr SE se function Among them, Chr represents the chromosome location information of the functional component region, S represents the mapping value of the starting site of the functional component region, E represents the mapping value of the ending site of the functional component region, s represents the starting site of the functional component region, E represents the ending site of the functional component region, and function represents the category of the functional component region; S2, obtaining the mapping value of the site to be annotated: the site value is Q, and the mapping value G of the site to be annotated is obtained by using formula (1); S3, searching the mapping value G obtained in step S2 in the second and third columns of the index file, if for a certain functional component region j, G satisfies S j ≤G≤E j , further determine whether Q satisfies s j ≤Q≤e j , if satisfied, the site to be annotated can be annotated to be located in the jth functional component region.
2. The method for rapid annotation of genome sequencing data according to claim 1, It is characterized in that The method for determining N is specifically as follows: (1) Obtain the length CL and gene number GN of each chromosome, and calculate CL / GN; (2) Obtain the representative number MN of all chromosomes CL / GN, divide it by the value q, and the integer digit of the result of MN / q is the value N, where q = 1 to 100.
3. The method for rapid annotation of genome sequencing data according to claim 2, It is characterized in that The representative number is selected from one of the median, mode and mean.
4. The method for rapid annotation of genome sequencing data according to claim 1, It is characterized in that The source species is a mammal.
5. The method for rapid annotation of genome sequencing data according to claim 4, It is characterized in that The functional component regions include promoter regions, exon regions, intron regions, promoter CGIs, intragenic CGIs, 3'transcript CGIs, intergenic CGIs, repeat regions and miRNA regions.
6. A rapid annotation system for genome sequencing data based on site mapping, It is characterized in that Includes the following modules: The index library module is used to store index files, wherein the index file is constructed as follows: The starting and ending sites of the functional component region of the species from which the sequencing sample originated are obtained. For each site, the mapping value is obtained using formula (1): Among them, G i represents the mapping value of the ith site, S i represents the value of the ith site, N is the value determined according to the chromosome length of the source species, and L i Represents the number of digits at the i-th site. If L i ≤N then L i -N=1, The mapping values of the start and end sites of all functional component regions are obtained, and the index file is constructed in the following format: Chr SE se function Among them, Chr represents the chromosome location information of the functional component region, S represents the mapping value of the starting site of the functional component region, E represents the mapping value of the ending site of the functional component region, s represents the starting site of the functional component region, E represents the ending site of the functional component region, and function represents the category of the functional component region. The input module is used to receive sequencing data, obtain the sites to be annotated, and calculate the index value of the sites to be annotated using formula (1). The search module is connected to the input module and the index library module respectively, and is used to search the index value of the site to be annotated obtained by the input module in the second column and the third column of the index file. If for a certain functional component region j, G satisfies S j ≤G≤E j , further determine whether Q satisfies s j ≤Q≤e j , if satisfied, the site to be annotated can be annotated to be located in the jth functional component region, The result output module is used to output the annotation results.
7. The genome sequencing data rapid annotation system according to claim 6, It is characterized in that The method for determining N is specifically as follows: (1) Obtain the length CL and gene number GN of each chromosome, and calculate CL / GN; (2) Obtain the representative number MN of all chromosomes CL / GN, divide it by the value q, and the integer digit of the result of MN / q is the value N, where q = 1 to 100.
8. The genome sequencing data rapid annotation system according to claim 7, It is characterized in that The representative number is selected from one of the median, mode and mean.
9. The genome sequencing data rapid annotation system according to claim 6, It is characterized in that The source species is a mammal.
10. The genome sequencing data rapid annotation system according to claim 9, It is characterized in that The functional component regions include promoter regions, exon regions, intron regions, promoter CGIs, intragenic CGIs, 3'transcript CGIs, intergenic CGIs, repeat regions and miRNA regions.
Citation Information
Patent Citations
Method and system for quickly annotating genome sequencing data based on sliding window mapping
CN115631801A