Specific sequence hierarchical database and construction method thereof, and method and device for screening specific sequences
By segmenting genome sequences into K-mer sequences and establishing mapping relationships, a specific sequence hierarchical database is constructed, which solves the problems of slow speed and time-consuming computational resources in traditional methods, and realizes rapid and efficient specific sequence screening across all species.
Patent Information
- Application Number
- CN202510951689.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-11-11
AI Technical Summary
Traditional methods are slow and time-consuming when screening for specific DNA sequences, and are difficult to perform efficiently across all species, especially when screening for specific sequences between different biological taxa, which requires massive computational resources and time.
By dividing the genome sequence into continuous K-mer sequences according to a preset length K, a mapping relationship between each K-mer sequence and the species is established to determine its specific classification level, and a specific sequence classification database is constructed. This database is then used to quickly screen specific sequences.
It enables rapid screening of highly specific sequences across all species, reducing computational resources and time requirements, and improving screening efficiency, especially the speed of screening specific sequences across different biological classifications.
Smart Images

Figure CN120932747A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of bioinformatics technology, specifically relating to a specific sequence hierarchical database and its construction method, a method for screening specific sequences, and an apparatus. Background Technology
[0002] Genome sequences from different biological classifications (phylum, class, order, family, genus, species) share some similarities but also have differences. For example, the genomes of humans, chimpanzees, and mice have a certain degree of similarity; some genome sequences are shared by all three, some are found in humans and chimpanzees but not in mice, and some are unique to humans. Screening for specific sequences from different classifications can be used to design PCR primers for targeted amplification of specific class of genome sequences. For example, pathogen tNGS detection involves designing multiplex PCR primers targeting specific sequences of several pathogenic microorganisms, and then using PCR amplification and sequencing to identify the presence of the corresponding pathogenic microorganism in the sample.
[0003] The complexity of DNA sequences increases exponentially with sequence length. For example, there are 10^24 (4^5) possible sequences for a 5-nt sequence, over 1 trillion (4^20) possible sequences for a 20-nt sequence, and an even greater number for a 30-nt sequence. 18 In terms of order of magnitude, a 30-nt sequence generally has a certain degree of specificity. Typically, DNA sequences of 100 nt or more are needed for screening specific DNA sequences used in PCR, while PCR primers are usually 18-25 nt in length.
[0004] Traditional methods for screening specific sequences involve multiple sequence alignment to find specific target sequences. This method is slow and time-consuming, and can only be used for comparisons within a small range (e.g., within the same family / genus). If comparisons are to be made across all species, enormous computational resources and time are required. Summary of the Invention
[0005] Based on this, one embodiment of this application provides a specific sequence classification database and a method for constructing it, as well as a method and apparatus for screening specific sequences. One aspect of this application provides a method for constructing a specific sequence classification database, comprising the following steps:
[0006] The genome sequences of each species within or between target levels according to biological classification are divided into continuous K-mer sequences according to a preset length K.
[0007] Establish a mapping relationship between each K-mer sequence and its corresponding species;
[0008] The specific classification level of each K-mer sequence is determined based on the mapping relationship between each K-mer sequence and the species. The specific classification level is the nearest common classification level of all species corresponding to the K-mer sequence in the target level. When all species corresponding to the K-mer sequence are outside the classification of the target level, it is recorded as no classification.
[0009] In one embodiment, the target level is the species classification level.
[0010] In one embodiment, the species within or between the target hierarchies of the biological classification include all known species within or between the target hierarchies.
[0011] Optionally, each species includes the genome sequence of individuals with different phenotypes.
[0012] In one embodiment, the preset length K is an odd number and has a range of 31≤K≤45.
[0013] In one embodiment, when determining the specific taxonomic level of each K-mer sequence based on the mapping relationship between each K-mer sequence and a species, the number of species covered by each K-mer sequence is recorded at the same time.
[0014] This application also provides a specific sequence classification database, which contains the correspondence between K-mer sequences and species-specific classification levels, and is constructed by the above-described construction method.
[0015] This application also provides a database system, including:
[0016] Application programming interface (API) is used to receive user search requests and feedback results;
[0017] A database is used to match the data features of user retrieval requests from the application interface and output the matching results to the application interface. The database is constructed using the construction method described above.
[0018] In one embodiment, the following steps are included:
[0019] The genome sequence of the target organism is divided into continuous K-mer sequences to be screened according to a preset length K, wherein the preset length K is the same as the preset length K in the above-mentioned specific sequence classification database or the above-mentioned database system.
[0020] After searching the K-mer sequences to be screened in the aforementioned specific sequence classification database or the aforementioned database system, the screened specific sequences and their corresponding classification levels are obtained.
[0021] Another aspect of this application provides an apparatus for screening specific sequences, comprising:
[0022] The segmentation module is used to segment the genome sequence of the target organism into continuous K-mer sequences to be screened according to a preset length K, wherein the preset length K is the same as the preset length K in the above-mentioned specific sequence classification database or the above-mentioned database system.
[0023] The retrieval module is used to retrieve the K-mer sequence to be screened from the aforementioned specific sequence classification database or the aforementioned database system, and obtain the screened specific sequences and their corresponding classification levels.
[0024] In another aspect, this application provides a computer device including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the method described above.
[0025] In another aspect, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the method described above.
[0026] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the method described above.
[0027] This application provides a method for constructing a specific sequence classification database, including constructing a K-mer specific taxonomic level database. Once the K-mer specific taxonomic level database is established, it can be repeatedly used in subsequent steps. The step of obtaining the target specific taxonomic level sequence using the above database can be automated by a program, with a fast running speed. Taking a bacterial genome size (on the order of millions to tens of millions of bases) as an example, the specific sequence of that species can be obtained in just a few minutes.
[0028] Furthermore, this application enables rapid screening across all species, and due to the small granularity of K-mers, sequences formed by continuous K-mer-level specific combinations have high specificity. Attached Figure Description
[0029] To more clearly illustrate the technical solutions in the embodiments of this application and to more completely understand this application and its beneficial effects, the drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0030] Figure 1This is an example of a cutting method using a 35bp K-mer.
[0031] Figure 2 Examples of the most recent common classification level for humans and chimpanzees are family: Hominidae, the most recent common classification level for humans / chimpanzees / mice is class: Mammalia, and humans / chimpanzees / mice / E. coli exceed the highest level of classification (i.e., there is no most recent common classification level).
[0032] Figure 3 This application provides a sequence-specific alignment to the QueryCover and Identity values of Escherichia coli in one embodiment of the present application.
[0033] Figure 4 This application provides a sequence-specific alignment to the QueryCover and Identity values of Escherichia coli in one embodiment of the present application.
[0034] Figure 5 This application provides a sequence-specific alignment to the QueryCover and Identity values of Escherichia coli in one embodiment of the present application.
[0035] Figure 6 This application provides a sequence-specific alignment to the QueryCover and Identity values of Escherichia coli in one embodiment of the present application.
[0036] Figure 7 This application provides a sequence-specific alignment to the QueryCover and Identity values of Escherichia coli in one embodiment of the present application.
[0037] Figure 8 This application provides a sequence-specific alignment to the QueryCover and Identity values of Escherichia coli in one embodiment of the present application.
[0038] Figure 9 This is a sequence-specific alignment to the QueryCover and Identity values of Escherichia coli in one embodiment of this application. Detailed Implementation
[0039] The present application will be further described in detail below with reference to the embodiments and examples. It should be understood that these embodiments and examples are for illustrative purposes only and are not intended to limit the scope of the present application. The purpose of providing these embodiments and examples is to enable a more thorough and comprehensive understanding of the disclosure of the present application. It should also be understood that the present application can be implemented in many different forms and is not limited to the embodiments and examples described herein. Those skilled in the art can make various modifications or alterations without departing from the spirit of the present application, and the equivalent forms obtained also fall within the protection scope of the present application. Furthermore, numerous specific details are set forth in the following description to provide a fuller understanding of the present application. It should be understood that the present application can be implemented without one or more of these details.
[0040] Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains.
[0041] the term
[0042] Unless otherwise stated or in case of contradiction, the terms or phrases used herein shall have the following meanings:
[0043] The terms "and / or," "or / and," and "and / or" as used herein include any one of two or more of the related listed items, as well as any and all combinations of the related listed items. These arbitrary and all combinations include any two related listed items, any more related listed items, or a combination of all related listed items. It should be noted that when at least three items are connected by at least two conjunctions selected from "and / or," "or / and," and "and / or," it should be understood that in this application, the technical solution undoubtedly includes technical solutions connected by "logical AND," and also undoubtedly includes technical solutions connected by "logical OR." For example, "A and / or B" includes three parallel solutions: A, B, and A+B. For example, the technical solution of "A, and / or, B, and / or, C, and / or, D" includes any one of A, B, C, and D (that is, a technical solution that is connected by "logical OR"), as well as any and all combinations of A, B, C, and D, that is, combinations of any two or three of A, B, C, and D, and also combinations of all four of A, B, C, and D (that is, a technical solution that is connected by "logical AND").
[0044] In this application, the terms "multiple", "various", "multiple times", "multi-dimensional", etc., unless otherwise specified, refer to a quantity greater than or equal to 2. For example, "one or more" means one or more than or equal to two.
[0045] The terms “combinations of,” “any combination of,” and “any combination of” used in this article include all suitable combinations of any two or more of the listed items.
[0046] In this document, the term "suitable" as used in phrases such as "suitable combination," "suitable method," and "any suitable method" refers to the ability to implement the technical solution of this application, solve the technical problem of this application, and achieve the expected technical effect of this application.
[0047] In this application, terms such as "further," "even more," and "particularly" are used for descriptive purposes and to indicate differences in content, but should not be construed as limiting the scope of protection of this application.
[0048] In this application, "optionally," "optionally," and "optional" mean that something is optional, that is, it means that it is selected from either "with" or "without." If there are multiple "optional" entries in a technical solution, unless otherwise specified, and there are no contradictions or mutual constraints, each "optional" entry shall be independent.
[0049] In this application, the technical features described in an open-ended manner include both closed technical solutions consisting of the listed features and open technical solutions that include the listed features.
[0050] In this application, numerical intervals (i.e., numerical ranges) are involved. Unless otherwise specified, the selected numerical distributions within the aforementioned numerical intervals are considered continuous and include the two endpoints (i.e., the minimum and maximum values) of the numerical range, as well as every value between these two endpoints. Unless otherwise specified, when a numerical interval refers only to integers within that interval, it includes the two endpoint integers of the numerical range, as well as every integer between the two endpoints. In this document, this is equivalent to directly listing every integer. For example, if t is an integer selected from 1 to 10, it means that t is any integer selected from the group of integers consisting of 1, 2, 3, 4, 5, 6, 7, 8, 9, and 10. Furthermore, when multiple ranges are provided to describe features or characteristics, these ranges can be merged. In other words, unless otherwise specified, the ranges disclosed herein should be understood to include any and all subranges to which they are included.
[0051] Unless otherwise specified, the temperature parameters in this application are permitted to be either constant-temperature treatment or variations within a certain temperature range. It should be understood that the constant-temperature treatment allows temperature fluctuations within the precision range of the instrument control, such as ±5℃, ±4℃, ±3℃, ±2℃, or ±1℃.
[0052] In this application, % (w / w) and wt% both represent weight percentage, % (v / v) refers to volume percentage, and % (w / v) refers to mass-volume percentage.
[0053] All references to documents mentioned in this application are incorporated herein by reference as if each document were individually incorporated herein by reference. Unless they conflict with the inventive purpose and / or technical solution of this application, all cited documents are incorporated herein by reference in their entirety and for all purposes. When citing documents in this application, the definitions of relevant technical features, terms, nouns, phrases, etc., are also incorporated herein by reference. When citing documents in this application, examples and preferred embodiments of the cited technical features may also be incorporated herein by reference, but only to the extent that they enable the implementation of this application. It should be understood that when the cited content conflicts with the description in this application, this application shall prevail or modifications shall be made adaptably to the description in this application.
[0054] The term "specific sequence" used in this application does not necessarily refer to a gene; it can also refer to a non-gene sequence. The data source is the genome sequences of all known species.
[0055] In bioinformatics, the term "K-mer" refers to a collection of subsequences of length K that are continuously divided into segments of length K from a sequence (such as the nucleotide base sequence of DNA).
[0056] To ensure that the reverse complementary sequence of each K-mer sequence is different from itself (i.e. to avoid confusion between the positive and negative chains), its length K is usually set to an odd number.
[0057] The term "sequence alignment" refers to aligning two or more sequences together to indicate their similarities. Spaces (usually indicated by a hyphen "-") can be inserted into the sequences. Corresponding identical or similar symbols (A, T (or U), C, G in nucleic acids, and single-letter representations of amino acid residues in proteins) are arranged in the same column. This method is commonly used to study sequences that evolved from a common ancestor, especially biological sequences such as protein or DNA sequences. In alignment, mismatches correspond to mutations, while gaps correspond to insertions or deletions. Sequence alignment can also be used in studies such as language evolution or textual similarity.
[0058] This application provides a method that enables the retrieval and screening of highly specific sequences across all species, while significantly reducing the required computational resources and time.
[0059] This application provides a method for constructing a specific sequence hierarchical database, comprising the following steps:
[0060] The genome sequences of each species within or between target levels according to biological classification are divided into continuous K-mer sequences according to a preset length K.
[0061] Establish a mapping relationship between each K-mer sequence and its corresponding species;
[0062] The specific classification level of each K-mer sequence is determined based on the mapping relationship between each K-mer sequence and the species. The specific classification level is the nearest common classification level of all species corresponding to the K-mer sequence in the target level. When all species corresponding to the K-mer sequence are outside the classification of the target level, it is recorded as no classification.
[0063] In one embodiment, the target level is the species classification level.
[0064] This mainly includes: kingdom, phylum, class, order, family, genus, and species, and also includes other sub-levels, such as subspecies, subfamilies, and superfamilies, all of which are applicable to the construction method of this application. In one embodiment, the species within or between the target levels of the biological classification include all known species within or between the target levels; that is, when a target level is specified, the species covered by that level include the set of all known species in that taxonomic group and all its sub-levels.
[0065] For example, when the target level is a boundary, the species include the kingdom and all known species within the kingdom. Similarly, when the target level is a phylum, the species include the phylum and all known species within the phylum.
[0066] In one embodiment, each species includes the genome sequence of its different phenotypes.
[0067] In one embodiment, 31 ≤ K ≤ 45, and is generally set to an odd number. For example, K can be 31, 33, 35, 37, 39, 41, 43, or 45, or any value in between.
[0068] In one embodiment, when determining the specific taxonomic level of each K-mer sequence based on the mapping relationship between each K-mer sequence and a species, the number of species covered by each K-mer sequence is recorded at the same time.
[0069] In a specific example, the method for constructing a specific sequence classification database includes the following steps:
[0070] 1. Establish the K-mer set of species.
[0071] The genomic DNA sequences of all known species are continuously divided into sequences of length K (K-mer). Usually, a K-mer length of 35 is chosen. Longer K-mers require more computational memory. If memory is insufficient, a smaller K-mer can be chosen, but the minimum length should not be less than 30.
[0072] P.S.: K-mer is a commonly used sequence processing method in bioinformatics. For example... Figure 1 The image shows an example of a 35bp K-mer.
[0073] 2. Establish a list of species corresponding to each K-mer.
[0074] Read the species and their K-mer sets one by one from the above steps, and then reverse the process to obtain the species list corresponding to each K-mer.
[0075] For example:
[0076] Humans (Homo sapiens) have {K-mer1, K-mer2, K-mer3, K-mer4}.
[0077] Chimpanzees (Pan troglodytes) have {K-mer2, K-mer3, K-mer5, K-mer7}.
[0078] Mice (Mus musculus) have {K-mer3, K-mer5, K-mer6, K-mer7}.
[0079] Escherichia coli contains {K-mer7, K-mer8, K-mer9}, therefore:
[0080] The species list corresponding to K-mer1 and K-mer4 is [human].
[0081] The species list corresponding to K-mer2 is [human, chimpanzee].
[0082] The species list corresponding to K-mer3 is [human, chimpanzee, mouse].
[0083] The species list corresponding to K-mer5 is [chimpanzee, mouse].
[0084] The species list corresponding to K-mer6 is [mouse].
[0085] The species list corresponding to K-mer7 is [chimpanzee, mouse, Escherichia coli].
[0086] The species list corresponding to K-mer8 and K-mer9 is [Escherichia coli].
[0087] 3. Determine the specific taxonomic level of each K-mer based on the list of species corresponding to it, and record the number of species it covers.
[0088] The most recent common classification level of all species in the species list corresponding to a K-mer is the specific classification level of that K-mer.
[0089] like Figure 2As shown, for example, the most recent common classification level between humans and chimpanzees is family: Hominidae, the most recent common classification level between humans / chimpanzees / mice is class: Mammalia, and humans / chimpanzees / mice / E. coli exceed the highest level of classification (i.e., there is no most recent common classification level).
[0090] In the example above:
[0091] K-mer1 and K-mer4 correspond to the specific classification level of species: Homo sapiens (human), covering 1 species.
[0092] K-mer2 corresponds to the specific taxonomic level of family: Hominidae, covering 2 species.
[0093] K-mer3 corresponds to the specific taxonomic level of class: Mammalia, covering 3 species.
[0094] K-mer5 corresponds to the specific taxonomic level of class: Mammalia, covering 2 species.
[0095] K-mer6 corresponds to the specific classification level of species: Mus musculus (mouse), covering 1 species.
[0096] K-mer7 corresponds to the unranked classification level, covering 3 species.
[0097] K-mer8 and K-mer9 correspond to the specific classification level of species: Escherichia coli, covering 1 species.
[0098] The comprehensiveness of the species used in establishing the species K-mer set in step 1 has a significant impact on this step. For example, if the mouse genome sequence is not used in step 1, then the species list corresponding to K-mer3 in step 2 will be [human, chimpanzee], and the species list corresponding to K-mer5 will be [chimpanzee]. In this step, the specific taxonomic level corresponding to K-mer3 will be family: Hominidae, and the specific taxonomic level corresponding to K-mer5 will be species: Pan troglodytes (chimpanzee).
[0099] In this process, multiple reference genomes can be provided for the same species (e.g., multiple different strains of the same bacterium) to perform step 1. A more comprehensive individual genome can make the specific taxonomic level determined in this step more accurate.
[0100] This application also provides a specific sequence classification database, which contains the correspondence between K-mer sequences and species-specific classification levels, and is constructed by the above-described construction method.
[0101] This application also provides a database system, including:
[0102] Application programming interface (API) is used to receive user search requests and feedback results;
[0103] A database is used to match the data features of user retrieval requests from the application interface and output the matching results to the application interface. The database is constructed using the construction method described above.
[0104] This application also addresses the application of the species-specific gene grading database or the database system described herein in screening for species-specific sequences.
[0105] This application also provides a method for screening specific sequences, comprising the following steps: dividing the genome sequence of a target organism into continuous K-mer sequences to be screened according to a preset length K, wherein the preset length K is the same as the preset length K in the aforementioned specific sequence classification database or the aforementioned database system.
[0106] After searching the K-mer sequences to be screened in the aforementioned specific sequence classification database or the aforementioned database system, the screened specific sequences and their corresponding classification levels are obtained.
[0107] In one embodiment, when the K-mer sequence to be screened is retrieved from the specific sequence classification database or the database system, the K-mer sequence to be screened is considered to be the selected specific sequence when the matching degree between the K-mer sequence to be screened and the K-mer sequence in the specific sequence classification database or the database system reaches a preset value.
[0108] It is understandable that the specificity of the screening can be adjusted according to the actual situation.
[0109] For example, to find a specific sequence of length 300 bp, the optimal choice is that all K-mers comprising this 300 bp are target-specific classification levels. However, in reality, such a fragment may not exist. For instance, when the longest consecutive K-mers of target-specific classification levels that can be found correspond to a sequence length of 280 bp, theoretically, at least one 300 bp fragment can be found where the proportion of target-specific classification level K-mers in this fragment is higher than a certain percentage X (for example, when the K-mer length is 35, 280 bp corresponds to 246 K-mers, and 300 bp corresponds to 266 K-mers, then a 300 bp sequence can be obtained by extending this 280 bp by 20 bp in both directions, and the proportion of target-specific classification level K-mers in this 300 bp is at least X = 246 / 266 = 92.5%). The screening threshold can be adjusted to be 5% smaller than X.
[0110] In the example above, the threshold is set to have a K-mer proportion of target-specific classification level higher than 87.5% (the purpose is to select target-length specific sequences by sacrificing a small amount of specificity).
[0111] In a specific example, obtaining screening-specific sequences using the above database includes:
[0112] 1. Divide the species genomic DNA sequence (which may be multiple) under the target specific taxonomic level into continuous K-mers, and match the specific taxonomic level of each K-mer in the above database in turn.
[0113] A genomic DNA sequence of length N can be segmented into N-k+1 K-mer sequences of length k (numbered 0, 1, 2...Nk), and each K-mer is matched with the specific classification level of all K-mers.
[0114] 2. Retrieve consecutive target-specific classification-level K-mer sets and match them back to genomic DNA sequences.
[0115] Identify consecutive K-mers at the target-specific classification level (different consecutive lengths can be selected as needed) and their corresponding genomic DNA sequences (K-mer numbered ij corresponds to the (i+1)th to (j+k)th bases in the genomic sequence).
[0116] For example, if the target specific classification level is species: Escherichia coli, the specific classification level of the K-mers with a length of 35 cut from the E. coli genomic DNA sequence (denoted as S) is E. coli, and the K-mers numbered 8-172, 260-271, and 590-815 are all E. coli. The genomic DNA sequences corresponding to these consecutive K-mers are the 9th-207th (199 bp), 261-306th (46 bp), and 591-850th (260 bp) bases of S.
[0117] Another aspect of this application provides an apparatus for screening specific sequences, comprising:
[0118] The segmentation module is used to segment the genome sequence of the target organism into continuous K-mer sequences to be screened according to a preset length K, wherein the preset length K is the same as the preset length K in the above-mentioned specific sequence classification database or the above-mentioned database system.
[0119] The retrieval module is used to retrieve the K-mer sequence to be screened from the aforementioned specific sequence classification database or the aforementioned database system, and obtain the screened specific sequences and their corresponding classification levels.
[0120] This application, in another aspect, provides a computer device including a memory and a processor. The memory stores a computer program, and the processor, when executing the computer program, implements the steps of a method for screening specific sequences. The steps in the method for screening specific sequences described herein can be steps from one of the methods for screening specific sequences described in the various embodiments above.
[0121] This application, in another aspect, provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of a method for screening specific sequences. The steps in the method for screening specific sequences described above may be steps from one of the methods for screening specific sequences described in the various embodiments above.
[0122] This application, in another aspect, provides a computer program product, including a computer program that, when executed by a processor, implements the steps of a method for screening specific sequences. The steps in the method for screening specific sequences described herein may be steps from one of the methods for screening specific sequences described in the various embodiments above.
[0123] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0124] The databases involved in the various embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the various embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these. The implementation schemes of this application will be described in detail below with reference to embodiments. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of this application. For experimental methods in the following embodiments where specific conditions are not specified, priority should be given to the guidelines given in this application, or experimental manuals or conventional conditions in the art may be followed, or conditions recommended by the manufacturer may be followed, or experimental methods known in the art may be referenced.
[0125] The embodiments of this application will be described in detail below with reference to examples. It should be understood that these embodiments are for illustrative purposes only and are not intended to limit the scope of this application. For experimental methods in the following embodiments where specific conditions are not specified, please refer to the guidelines given in this application, or follow experimental manuals or conventional conditions in the art, or follow the conditions recommended by the manufacturer, or refer to experimental methods known in the art.
[0126] In the specific embodiments described below, the measurement parameters involving raw material components may have slight deviations within the weighing accuracy range unless otherwise specified. Temperature and time parameters are subject to acceptable deviations due to instrument testing accuracy or operational precision.
[0127] It should be understood that in the various embodiments of this application, the order of the above-mentioned processes does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application.
[0128] Example 1
[0129] This embodiment provides a method for constructing a specific sequence classification database.
[0130] To facilitate intuitive demonstration, this example uses the exon 1 sequence of the HBA1 gene from humans, chimpanzees, and mice as examples, selecting a K-mer length of 5, to demonstrate the steps for establishing this database. The sequences from the three species are as follows:
[0131] Human (Homo sapiens):
[0132] ACTCTTCTGGTCCCCACAGACTCAGAGAGAACCCACCATGGTGCTGTCTCCTGCCG ACAAGACCAACGTCAAGGCCGCCTGGGGTAAGGTCGGCGCACGCTGGCGAGTATGG TGCGGAGGCCCTGGAGAG (SEQ ID NO. 1).
[0133] Chimpanzee (Pan troglodytes):
[0134] ACTCTTCTGGTCCCCACAGACTCAGAAAGAACCCACCATGGTGCTGTCTCCTGCCG ACAAGACCAACGTCAAGGCCGCCTGGGGTAAGGTCGGCGCACGCTGGCGAGTATGG TGCGGAGGCCCTGGAGAG (SEQ ID NO. 2).
[0135] Mice (Mus musculus):
[0136] GACACTTCTGATTCTGACAGACTCAGGAAGAAACCATGGTGCTCTCTGGGGAAGA CAAAAGCAACATCAAGGCTGCCTGGGGGAAGATTGGTGGCCATGGTGCTGAATATGGA GCTGAAGCCCTGGAAAG (SEQ ID NO. 3).
[0137] 1. Construct a K-mer set of length 5 for the above sequences of the three species.
[0138] Human (Homo sapiens):
[0139] {'ACTCT','CTCTT','TCTTC','CTTCT','TTCTG','TCTGG','CTGGT','TGGTC','GGTCC','GTCCC','TCCCC','CCCCA','CCCAC','CCACA','CACAG' ,'ACAGA','CAGAC','AGACT','GACTC','ACTCA','CTCAG','TCAGA','CAGAG','AGAGA','GAGAG','GAGAA','AGAAC','GAACC','AACCC','ACCCA' ,'CCACC','CACCA','ACCAT','CCATG','CATGG','ATGGT','TGGTG','GGTGC','GTGCT','TGCTG','GCTGT','CTGTC','TGTCT','GTCTC','TCTCC' ,'CTCCT','TCCTG','CCTGC','CTGCC','TGCCG','GCCGA','CCGAC','CGACA','GACAA','ACAAG','CAAGA','AAGAC','AGACC','GACCA','ACCAA', 'CCAAC','CAACG','AACGT','ACGTC','CGTCA','GTCAA','TCAAG','CAAGG','AAGGC','AGGCC','GGCCG','GCCGC','CCGCC','CGCCT','GCCTG', 'CCTGG','CTGGG','TGGGG','GGGGT','GGGTA','GGTAA','GTAAG','TAAGG','AAGGT','AGGTC','GGTCG','GTCGG','TCGGC','CGGCG','GGCGC', 'GCGCG','CGCGC','GCGCA','CGCAC','GCACG','CACGC','ACGCT','CGCTG','GCTGG','CTGGC','TGGCG','GGCGA','GCGAG','CGAGT','GAGTA', 'AGTAT','GTATG','TATGG','GTGCG','TGCGG','GCGGA','CGGAG','GGAGG','GAGGC','GGCCC','GCCCT','CCCTG','CTGGA','TGGAG','GGAGA'}.
[0140] Chimpanzee (Pan troglodytes):
[0141] {'ACTCT','CTCTT','TCTTC','CTTCT','TTCTG','TCTGG','CTGGT','TGGTC','GGTCC','GTCCC','TCCCC','CCCCA','CCCAC','CCACA','CACAG',' ACAGA','CAGAC','AGACT','GACTC','ACTCA','CTCAG','TCAGA','CAGAA','AGAAA','GAAAG','AAAGA','AAGAA','AGAAC','GAACC','AACCC','AC CCA','CCACC','CACCA','ACCAT','CCATG','CATGG','ATGGT','TGGTG','GGTGC','GTGCT','TGCTG','GCTGT','CTGTC','TGTCT','GCTTC','TCTC C','CTCCT','TCCTG','CCTGC','CTGCC','TGCCG','GCCGA','CCGAC','CGA','DO','DO','DO','DO','DO','DO','DO', 'CCAAC','CAACG','AACGT','ACGTC','CGTCA','GTCAA','TCAAG','CAAGG','AAGGC','AGGCC','GGCCG','GCCGC','CCGCC','CGCCT','GCCTG','C CTGG','CTGGG','TGGGG','GGGGT','GGGTA','GGTAA','GTAAG','TAAGG','AAGGT','AGGTC','GGTCG','GTCGG','TCGGC','CGGCG','GGCGC',' CG','CGGCGC','GGCCA','CGCAC','GCCACG','CAGCG','ACGCT','CGCTG','GCTGG','CTGGC','TGGCG','GGCGA','GCGAG','CGAGT','GAGTA','AGTAT ','GTATG','TATGG','GTGCG','TGCGG','GCGGA','CGGAG','GGAGG','GAGGC','GGCCC','GCCCT','CCCTG','CTGGA','TGGAG','GGAGA','。GAGAG'}
[0142] Mouse (Mus musculus):
[0143] {'GACAC', 'ACACT', 'CACTT', 'ACTTC', 'CTTCT', 'TTCTG', 'TCTGA', 'CTGAT', 'TGATT', 'GATTC', 'ATTCT', 'CTGAC', 'TGACA', 'GACAG', 'ACAGA', 'CAGAC', 'AGACT', 'GACTC', 'ACTCA', 'CTCAG', 'TCAGG', 'CAGGA', 'AGGAA', 'GGAAG', 'GAAGA', 'AAGAA', 'AGAAA', 'GAAAC', 'AAACC', 'AACCA', 'ACCAT', 'CCATG', 'CATGG', 'ATGGT', 'TGGTG', 'GGTGC', 'GTGCT', 'TGCTC', 'GCTCT', 'CTCTC', 'TCTCT', 'CTCTG', 'TCTGG', 'CTGGG', 'TGGGG', 'GGGGA', 'GGGAA', 'AAGAC', 'AGACA', 'GACAA', 'ACAAA', 'CAAAA', 'AAAAG', 'AAAGC', 'AAGCA', 'AGCAA', 'GCAAC', 'CAACA', 'AACAT', 'ACATC', 'CATCA', 'ATCAA', 'TCAAG', 'CAAGG', 'AAGGC', 'AGGCT', 'GGCTG', 'GCTGC', 'CTGCC', 'TGCCT', 'GCCTG', 'CCTGG', 'GGGGG', 'AAGAT', 'AGATT', 'GATTG', 'ATTGG', 'TTGGT', 'GGTGG', 'GTGGC', 'TGGCC', 'GGCCA', 'GCCAT', 'TGCTG', 'GCTGA', 'CTGAA', 'TGAAT', 'GAATA', 'AATAT', 'ATATG', 'TATGG', 'ATGGA', 'TGGAG', 'GGAGC', 'GAGCT', 'AGCTG', 'TGAAG', 'GAAGC', 'AAGCC', 'AGCCC', 'GCCCT', 'CCCTG', 'CTGGA', 'TGGAA', 'GGAAA', 'GAAAG'}。
[0144] 2. Establish a list of species corresponding to the above K-mers
[0145] One of the K-mer spellings of the species['Homo sapiens','Pan troglodytes']:ACTCT、CTCTT、TCTTC、CTGGT、TGGTC、GGTCC、GTCCC、TCCCC、CCC CA、CCCAC、CCACA、CACAG、TCAGA、GAGAG、AGAAC、GAACC、AACCC、ACCCA、CCACC、CAC CA、GCTGT、CTGTC、TGTCT、GTCTC、TCTCC、CTCCT、TCCTG、CCTGC、TGCCG、GCCGA、CC GAC、CGACA、ACAAG、CAAGA、AGACC、GACCA、ACCAA、CCAAC、CAACG、AACGT、ACGTC、CG TCA、GTCAA、AGGCC、GGCCG、GCCGC、CCGCC、CGCCT、GGGGT、GGGTA、GGTAA、GTAAG、T AAGG、AAGGT、AGGTC、GGTCG、GTCGG、TCGGC、CGGCG、GGCGC、GCGCG、GCGCCA、GCGCAC GCAC、GCACG、CACGC、ACGCT、CGCTG、GCTGG、CTGGC、TGGCG、GGCGA、GCGAG、CGAGT GAGTA、AGTAT、GTATG、GTGCG、TGCGG、GCGGA、CGGAG、GGAGG、GAGGC、GGCCC、GGAGA.
[0146] One of the K-mer spellings of the species['Homo sapiens','Pan troglodytes','Musmusculus']:CTTCT, TTCTG, TCTGG, ACAGA, CAGAC, AGACT, GACTC, ACTCA, CTCAG, ACCAT, CCATG, CATGG, ATGGT, TGGT G、GGTGC、GTGCT、TGCTG、CTGCC、GACAA、AAGAC、TCAAG、CAAGG、AAGGC GCCTG, CCTGG, CTGGG, TGGGG, TATGG, GCCCT, CCCTG, CTGGA, TGGAG.
[0147] One of the K-mer snake species ['Pan troglodytes','Mus musculus']:AGAAA、GAAAG、AAGAA.
[0148] The following list of species corresponding to K-mer is ['Homo sapiens']: CAGAG, AGAGA, GAGAA.
[0149] The following list of species corresponding to K-mer is ['Pan troglodytes']: CAGAA, AAAGA.
[0150] The following species list corresponding to K-mer is ['Mus musculus']: GACAC, ACACT, CACTT, ACTTC, TCTGA, CTGAT, TGATT, GATTC, ATTCT, CTGAC, TGACA, GACAG, TCAGG, CAGGA, AGGAA, GGAAG, GA AGA, GAAAC, AAACC, AACCA, TGCTC, GCTCT, CTCTC, TCTCT, CTCTG, GGGGA, GGGAA, AGACA, ACAAA, CAAAA, AAAAG, AAAGC, AAGCA, AGCAA, GCA AC, CAACA, AACAT, ACATC, CATCA, ATCAA, AGGCT, GGCTG, GCTGC, TGCCT, GGGGG, AAGAT, AGATT, GATTG, ATTGG, TTGGT, GGTGG, GTGGC, TGGC C. GGCCA, GCCAT, GCTGA, CTGAA, TGAAT, GAATA, AATAT, ATATG, ATGGA, GGAGC, GAGCT, AGCTG, TGAAG, GAAGC, AAGCC, AGCCC, TGGAA, GGAAA.
[0151] 3. Determine the specific taxonomic level of the corresponding K-mer based on the relevant species list.
[0152] The corresponding species list ['Homo sapiens', 'Pan troglodytes'] K-mer is specifically classified at the family: Hominidae (Hominidae), covering 2 species.
[0153] The corresponding species list ['Homo sapiens', 'Pan troglodytes', 'Mus musculus'] has a specific taxonomic level of class: Mammalia, covering 3 species.
[0154] The corresponding species list ['Pan troglodytes', 'Mus musculus'] has a specific taxonomic level of class: Mammalia, covering 2 species.
[0155] The corresponding species list is ['Homo sapiens'] with a specific taxonomic level of species: Homosapiens (human), covering 1 species.
[0156] The corresponding species list is ['Pan troglodytes'] with a K-mer specific classification level of species: Pan troglodytes (chimpanzee), covering 1 species.
[0157] The corresponding species list is ['Mus musculus'] with a specific classification level of species: Musmusculus (mouse), covering 1 species.
[0158] Example 2 This example provides a method for screening specific sequences.
[0159] Based on the local server established using a large number of microbial sequences in this application embodiment, the following steps are demonstrated using Escherichia coli as an example:
[0160] 1. Using the reference genome sequence (U00096.3) of Escherichia coli K-12 strain, which is 4,641,652 bytes long, 4,641,618 K-mers of length 35 (numbered 0-4641617) were extracted. These K-mers were then retrieved from the established K-mer-specific classification database to obtain their corresponding specific classification levels. Some results are shown in Table 1 below:
[0161] Table 1
[0162]
[0163]
[0164] 2. Extract K-mer numbers for Escherichia coli with a specific taxonomic level, select consecutive K-mers that meet the criteria based on the length of the target specific sequence, and then map them back to the genome sequence.
[0165] For example, if the desired target-specific sequence length is greater than 100 bp, the following consecutive K-mer number intervals meet the criteria after screening: 84028-84097, 285164-285246, 576937-577182, 2558667-2558854, 2565301-2565467, 2784537-2784623, 3469919-3470011.
[0166] These consecutive K-mers correspond to the 84029-84132 bases (104bp) of the genome sequence (U00096.3): AATATTTTGATCAATTAATGTTAAGAATTAATGCATTAAATATATAAATTAATTATTAAATA AGCACATTTAATCCATTTTGTAGATGATTGAGTATTCGCGG (SEQ ID NO.4).
[0167] Bases 285165-285281 (117bp):
[0168] TTAATTATTTTAAGAGATAAAACCGTCTGCGGAATATTTCCCCGCAGACGGCTTTGTTGTTT TTGAAATTTATTAATTTAAAACAATTAGTTGAGATATATCGTTGGCGTCACAAAAGC (SEQ ID NO. 5).
[0169] Bases 576938-577217 (280bp):
[0170] TGAAGTAGATCCTATTTTTATCTGAACTTTTTTCTATCGAATCCTATTCATGGCTCTTGGC
[0171] TGAATAAAAATAAATCTATTAGCCAATTTATATTAACGGCTGTTATTTATAAGTGCTCTATA
[0172] ATTTGAAGGTTCAATTTAAACCGGCTAAAAATAACACTGGAAATTATTTTTTGGTTATTTG
[0173] TTGAGATTTGCTTATGTATTTGTAGTGGTGTTTTCAATACTCGGTAGCATTCTCTCAAATATCATTTAGTGGTTTACGTACGTAAAAAATTGGTT(SEQ ID NO.6).
[0174] Base pairs 2558668 - 2558889 (222bp):
[0175] CTTTAGTTTCATAAGTCGTTCCCTCAGGAAGAATCGATGATTGGCATTTTTCACAGCATTC
[0176] TAACAATCACGTTTCATCGTCAGCCCTTGTCGTGTAAGGTGGTTGCCTAAACACGCCCGT
[0177] TATTCATCACGCCGAACGCGCCGGATACATGATCGGGGTTATCCAGTCGTTAAATCAAGGTATCCGGTTTTGAGCAAGACACCACTCACAGCAAAGGCCAT(SEQ ID NO.7).
[0178] Base pairs 2565302 - 2565502 (201bp):
[0179] CAACTGGATTGATAAGCTGGGAGACGAATAAACCAGCCTTCAACCCCATCTCATCAATC
[0180] AACGCCCGGCCCCGCTGCCGGGTTTTTGCTATGCACCACAATTACCCCAACCGGATACA
[0181] CAGCCGGATACAATTCCACCAGCACCCAGCCACCCAGCGCCACCGCTGGCGAATACCGCATTCAGGAAGGAAATGCGAGTGAT(SEQ ID NO.8).
[0182] Base pairs 2784538 - 2784658 (121bp):
[0183] TCTTTTAGCGCATTCAAAAAACTGATCGGCATTATTTTTATTCGATAATTTTTTAGTTTCAG
[0184] AAAACACATTTTCATTGTTTTCCAGCTTTAGTTTAATGAGAAGATTTTCCCAGACCTGC (SEQ ID NO. 9).
[0185] Bases 3469920-3470046 (127bp):
[0186] AGAACCCTTCAATATGAATTAAATTACGGCATTAAAAATAAGAAAAAAGCCTGACAAAT
[0187] GAAGCATTTTAAAAACAGAAACATTCATATTTAAAATGTTAAATTGAATTGATATTTTAAA TATGAAT (SEQ ID NO. 10).
[0188] 3. Align the above 7 sequences with the NT library using BLAST to verify their specificity.
[0189] The blast results are shown below. Figures 3-9 The results showed that all 7 sequences specifically aligned to Escherichia coli (different alignments indicated different strains of Escherichia coli), and the Query Cover and Identity values of all alignments were 100% (meaning all 7 sequences were full-length and mismatch-free alignments), confirming that these 7 sequences are species-specific to Escherichia coli and have very high specificity.
[0190] The embodiments described above are merely illustrative of several implementation methods of this application, intended to facilitate a detailed understanding of the technical solutions of this application, but should not be construed as limiting the scope of protection of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the scope of protection of this application. Furthermore, it should be understood that after reading the above teachings of this application, those skilled in the art can make various alterations or modifications to this application, and the equivalent forms obtained also fall within the scope of protection of this application. It should also be understood that technical solutions obtained by those skilled in the art based on the technical solutions provided in this application through logical analysis, reasoning, or limited experimentation are all within the scope of protection of the appended claims. Therefore, the scope of protection of this patent application should be determined by the content of the appended claims, and the specification can be used to interpret the content of the claims.
Claims
1. A method for constructing a specific sequence hierarchical database, characterized in that, Includes the following steps: The genome sequences of each species within or between target levels according to biological classification are divided into continuous K-mer sequences according to a preset length K. Establish a mapping relationship between each K-mer sequence and its corresponding species; The specific classification level of each K-mer sequence is determined based on the mapping relationship between each K-mer sequence and the species. The specific classification level is the nearest common classification level of all species corresponding to the K-mer sequence in the target level. When all species corresponding to the K-mer sequence are outside the classification of the target level, it is recorded as no classification.
2. The method for constructing a specific sequence hierarchical database as described in claim 1, characterized in that, The target level is the species classification level.
3. The method for constructing a specific sequence hierarchical database as described in claim 1, characterized in that, Each species within or between the target levels of the biological classification includes all known species within or between the target levels. Optionally, each species includes the genome sequence of individuals with different phenotypes.
4. The method for constructing a specific sequence hierarchical database as described in claim 1, characterized in that, The preset length K is an odd number and its range is 31≤K≤45.
5. The method for constructing a specific sequence hierarchical database as described in any one of claims 1 to 4, characterized in that, When determining the specific taxonomic level of each K-mer sequence based on the mapping relationship between each K-mer sequence and the species, the number of species covered by each K-mer sequence is also recorded.
6. A specific sequence classification database containing the correspondence between K-mer sequences and species-specific classification levels, which is constructed by the construction method described in any one of claims 1 to 5.
7. A database system, characterized in that, include: Application programming interface (API) is used to receive user search requests and feedback results; A database is used to match the data features of user retrieval requests from the application interface and output the matching results to the application interface. The database is constructed using the construction method described in any one of claims 1 to 5.
8. The application of the specific sequence classification database as described in claim 6 or the database system as described in claim 7 in screening specific sequences of target species.
9. A method for screening specific sequences, characterized in that, Includes the following steps: The genome sequence of the target organism is divided into continuous K-mer sequences to be screened according to a preset length K, wherein the preset length K is the same as the preset length K in the specific sequence classification database of claim 6 or the database system of claim 7. After searching the K-mer sequence to be screened in the specific sequence classification database of claim 6 or the database system of claim 7, the screened specific sequences and their corresponding classification levels are obtained.
10. An apparatus for screening specific sequences, characterized in that, include: The segmentation module is used to segment the genome sequence of the target organism into continuous K-mer sequences to be screened according to a preset length K, wherein the preset length K is the same as the preset length K in the specific sequence classification database of claim 6 or the database system of claim 7. The retrieval module is used to retrieve the K-mer sequence to be screened from the specific sequence classification database of claim 6 or the database system of claim 7, and obtain the screened specific sequences and their corresponding classification levels.
11. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method of claim 9.
12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the method of claim 9.
13. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method of claim 9.
Citation Information
Patent Citations
Microbial genome database construction method and application thereof
CN112992277A
Sequence screening method and device, computer equipment and storage medium
CN117577185A
Metagenome data classification method, device and equipment and medium
CN118645149A