Biological sequencing
By using fingerprint data string repository to process biological sequencing reads, the problems of sequencing error and sequence reconstruction complexity in the prior art are solved, and a more efficient and accurate biological sequencing process is achieved.
Patent Information
- Application Number
- CN202080017929.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-08
- Filing Date
- 2020-02-07
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2040-02-07
AI Technical Summary
The existing sequencing errors in biological sequencing lead to complexity of sequence reconstruction, and the inability to effectively verify sequence accuracy or identify structural variants. The existing methods are unable to effectively construct pan-genomic maps, resulting in error accumulation and calculation difficulties.
By using a repository of fingerprint data strings, processing reads of biopolymers or biopolymer fragments, searching for featured biological subsequences and verifying or rejecting reads, the set of featured biological subsequences of fingerprint data strings is used to reduce complexity and error and improve sequencing accuracy.
It achieves the reduction of sequencing complexity and error, improves sequencing speed and accuracy, and can analyze and verify sequences in real time during the sequencing process, reducing error accumulation.
Smart Images

Figure CN113519029B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the processing of biological sequence information, and more particularly, to generating such biological sequence information, for example, by sequencing and / or sequence assembly. Systems and methods are provided for generating biological sequence information during a sequencing process. Background Art
[0002] In the past few decades, biological sequencing has advanced at an astonishing pace, making the Human Genome Project possible, which achieved a complete sequencing of the human genome more than 15 years ago. To drive this development, a large number of technological advancements have been required, from advancements in sample preparation and sequencing methods to data acquisition, processing, and analysis. At the same time, new scientific fields have emerged and developed, including genomics, proteomics, and bioinformatics.
[0003] Driven by the emphasis on data acquisition in the post-genomic era, this development has led to the accumulation of a large amount of sequence data. However, the ability to organize, analyze, and interpret this sequence to extract biologically relevant information has lagged behind. This problem is further complicated by the fact that a large amount of new sequence information is still generated every day. Muir et al. observed that this has triggered a paradigm shift and commented on the resulting changes in the sequencing cost structure and other related obstacles (MUIR, Paul et al., The real cost of sequencing: scaling computation to keep pace with data generation. Genome biology, 2016, 17.1:53.).
[0004] Currently, the most commonly used sequencing method is the so-called "high-throughput" or "next-generation sequencing" (NGS). Compared with first-generation sequencing, NGS is typically characterized by high scalability, allowing the entire genome to be sequenced at once. Generally, this is achieved by fragmenting larger sequences into smaller fragments, randomly sampling the fragments, and sequencing them. After sequencing different fragments, sequence assembly can be used to reconstruct the original sequence, in which the sequence fragments are aligned and merged based on their overlapping regions.
[0005] However, sequencers are not perfect, and sequencing errors (such as insertions, substitutions, and deletions) can always occur, especially when seeking high throughput. If the sequence fragments to be assembled contain errors, this will obviously complicate the reconstruction of the original sequence because the corresponding regions may no longer overlap. In addition, errors can also propagate into the final sequence, for example, resulting in incorrect variant identification. Some strategies have been developed to handle these sequencing errors, such as those disclosed by Shmilovici et al. (SHMILOVICI, Armin; BEN-GAL, Irad. Using a VOM model for reconstructing potential coding regions in EST sequences. Computational Statistics, 2007, 22.1: 49-69.). However, there is currently no effective method to directly verify whether a (fragment) sequence is correct or whether it contains one or more sequence errors.
[0006] Genome maps are used as a reference for sequence reconstruction from single reads, which are typically short DNA or RNA sequences. Thus, a linear reference is a representation of a single genome. For a complete representation, multiple genomes need to be combined to discover all the variations that a species may have.
[0007] Multiple problems arise when constructing a pan-genome map correctly. First, even the best assembled reference genomes contain deletions and errors. Second, it is not possible to find a suitable graphical representation to enclose all the necessary information to counter problems that will arise later during the process where graphical mapping will be performed. Neither De Bruijn graphs, directed graphs, nor bidirected graphs can accurately represent strands. Third, it seems possible to create reference groups using current technology, but due to the lack of structural coordinates, the constructed groups are essentially not useful in practice.
[0008] In addition, the curve graphs lack operation site definitions. Due to logarithmic complexity, repetitive regions are even more difficult to represent using known k-mer-based techniques. In conclusion, since it is impossible to maintain all the necessary data using existing technology, it is almost impossible to construct a variation group in the graphical structure of one species, and even less possible to construct variation groups for all biological species.
[0009] Structural variants play an important role in the development of cancer and other diseases and are less studied compared to single nucleotide variants, in part due to the lack of reliable identification from read data. When using k-mer techniques, the detection window for changes is by definition smaller than the total length of the k-mer. Using algorithms that overcome the k-mer window problem cannot effectively identify structural variants. High coverage is required to find evidence of just one structural change. Thus, the use of k-mers requires a large pool to effectively identify true changes from noise and read errors. Due to the lack of dynamic algorithms to align k-mers, many k-mers lead to computational challenges. The problem arises due to the use of dynamic programming; the infeasibility of this dynamic approach leads to the inherent use of heuristics. This in turn indicates the need for heuristics or parameterization to narrow the search space. However, the latter leads to inevitable error accumulation, suggesting that k-mers are not effective unified space patterns. Currently, this is only addressed in a strictly one-dimensional syntactic manner.
[0010] Due to the NP-hard nature of the mapping and assembly processes, greedy algorithms are commonly used to solve these problems, whereby an extended matrix is used to calculate relevant results based on a certain input.
[0011] Dynamic programming has been used, but the problem associated with it is that the source data (parameters such as location, read ID, etc.) is lost and cannot be traced back.
[0012] All of the above problems make effective and accurate graph folding almost impossible. This results in the inability to provide the necessary accuracy or location data required to construct a usable pan-genome graph. In addition, the use of k-mers lacks the specificity to distinguish multi-dimensional parameters in genetic information. This further adds to the inefficient construction of current genome graphs, which is demonstrated by the inability to identify structural variations, biases, or effectively enclose highly repetitive regions.
[0013] Therefore, there is still a need in the art for further improvement in sequencing and sequence assembly. SUMMARY OF THE INVENTION
[0014] The object of the present invention is to provide a good method for generating biological sequence information. This object is achieved by the method, device, and data structure according to the present invention.
[0015] In a first aspect, the present invention relates to a method for sequencing a biopolymer or a fragment of a biopolymer, taking into account information contained in a repository of fingerprint data strings, the method comprising: (a) obtaining, using a sequencer, at least one read of the biopolymer or the fragment of the biopolymer, and (b) processing the read by the following computer-implemented steps: (b1) searching the read for the occurrence of one or more of the characteristic biopolymer subsequences represented by the fingerprint data strings, and (b2) validating or rejecting the read by determining, at each occurrence, whether the sequence units contiguous to the characteristic biopolymer subsequence are consistent with the combined data in the repository, and / or (b1') searching the head and / or tail of the read for the occurrence of one of the characteristic biopolymer subsequences represented by the fingerprint data strings, and (b2') predicting one or more contiguous sequence units of the read from the combined data in the repository. Here, the repository of fingerprint data strings is used for a biological sequence database, each fingerprint data string representing a characteristic biopolymer subsequence composed of sequence units, each characteristic biopolymer subsequence having, in the biological sequence database, a combination number less than the total number of different sequence units available to it, the combination number of the biopolymer subsequence being defined as the number of different sequence units that occur as contiguous sequence units of the biopolymer subsequence in the biological sequence database, the repository further comprising combined data representing the different sequence units that occur as contiguous sequence units of the corresponding characteristic biopolymer subsequence in the biological sequence database.
[0016] Advantages of embodiments of the present invention are that systems and methods are obtained that provide reduced complexity.
[0017] Advantages of embodiments of the present invention are that deterministic systems and methods are obtained, i.e., ones that produce a given solution.
[0018] Advantages of embodiments of the present invention are that sequencing of biopolymers and fragments of biopolymers can be improved (e.g., by reducing the likelihood of errors or by speeding up the process) by relying on information contained in a repository of fingerprint data strings.
[0019] Advantages of embodiments of the present invention are that tentatively proposed biological sequences can be validated or rejected. Advantages of embodiments of the present invention are that errors occurring during sequencing can be reduced.
[0020] Advantages of embodiments of the present invention are that the speed of sequencing can be increased by predicting the next unit in the sequence or by limiting the number of its options.
[0021] Advantages of embodiments of the present invention are that the systems and methods have a deterministic character, i.e., the methods and systems lead to a specific solution for identifying / characterizing the sequence of a biopolymer or a fragment of a biopolymer.
[0022] Advantages of embodiments of the present invention are that the system and method allow tracking the ID of reads. The system and method allow, for example, backtracking, such as backtracking the error or uncertainty of reads.
[0023] Advantages of embodiments of the present invention are that, compared with at least most of the prior art systems, in embodiments of the present invention, even when the sequencing is still running, each generated read can be immediately analyzed. In this way, according to at least some embodiments of the present invention, data processing, such as the construction of subgraphs, can start immediately when the first read is received during the start of sequencing, and thus such data processing can be a progressive process. It can be executed in parallel with the set of reads. Another advantage of embodiments of the present invention is that when it is determined that the biological sequence being sequenced has been sufficiently identified before the completion of the full sequencing, the sequencing can be terminated in advance.
[0024] In a second aspect, the present invention relates to a sequencer adapted to perform step a of the method according to any embodiment of the first aspect.
[0025] In a third aspect, the present invention relates to a system comprising: (i) a sequencer according to the second aspect, and (ii) a data processing system adapted to receive reads from the sequencer and process the reads by performing step b of the method according to any embodiment of the first aspect.
[0026] Advantages of embodiments of the present invention are that, depending on the application, the steps of the method can be implemented by a variety of systems and devices, such as computer-based systems or sequencers. Another advantage of embodiments of the present invention is that the method can be implemented by computer-based systems (including cloud-based systems).
[0027] In a fourth aspect, the present invention relates to a computer program comprising instructions that, when executed by a computer, cause the computer to perform the method according to any embodiment of the first aspect.
[0028] In a fifth aspect, the present invention relates to a computer-readable medium comprising instructions that, when executed by a computer, cause the computer to perform the method according to any embodiment of the first aspect.
[0029] Specific aspects and preferred aspects of the present invention are set forth in the appended independent claims and dependent claims. Features from the dependent claims can be appropriately combined with the features of the independent claims and with the features of other dependent claims, not just as explicitly set forth in the claims.
[0030] Although the devices in this field have been continuously improved, changed and evolved, the concept of the present invention is considered to represent a substantially new and novel improvement, including improvements that depart from previous practices, which results in providing a more effective, stable and reliable device with such properties.
[0031] The above and other features, characteristics and advantages of the present invention will become apparent from the following detailed description in conjunction with the accompanying drawings, which illustrate the principles of the present invention by way of example. This description is given for the purpose of illustration only and does not limit the scope of the present invention. The reference figures cited below refer to the accompanying drawings. Brief Description of the Drawings
[0033] Figure 1 and Figure 2 is a graph showing the expected progress achieved by an embodiment of the present invention.
[0034] Figures 3 to 6 is a diagram depicting a system according to an embodiment of the present invention.
[0035] Figure 7 A schematic overview of the processing steps that can be performed in a method for sequencing according to an embodiment of the present invention.
[0036] Figures 8 to 11 is a schematic representation of several steps that can be used in an embodiment according to the present invention.
[0037] Figures 12 to 16 is a chart showing various metrics of the analysis of a processed Protein Data Bank (PDB) according to an embodiment of the present invention.
[0038] Figure 17 is a chart plotting against each other the number of HYFT TM matches found in the PDB database using two different matching strategies.
[0039] Figure 18 and Figure 21 is a graph comparing the total length of the search results using a prior art method (dashed line) on the one hand and a method according to an exemplary embodiment of the present invention (solid line) on the other hand.
[0040] Figure 19 and Figure 22 is a graph comparing the Levenshtein distance of the search results using a prior art method (dashed line) on the one hand and a method according to an exemplary embodiment of the present invention (solid line) on the other hand.
[0041] Figure 20 and Figure 23It is a graph comparing the longest common substrings of search results using prior art methods (dashed lines) on the one hand and methods according to exemplary embodiments of the present invention (solid lines) on the other hand.
[0042] In different figures, the same reference numerals refer to the same or similar elements. Detailed Description
[0043] The present invention will be described with respect to specific embodiments and with reference to certain drawings, but the present invention is not limited thereto and is only limited by the claims. The described drawings are merely illustrative and not restrictive. In the drawings, for illustrative purposes, the sizes of some elements may be enlarged and not necessarily drawn to scale. The dimensions and relative dimensions do not correspond to the actual reduction in the practice of the present invention.
[0044] Furthermore, the terms first, second, third, etc. in the specification and claims are used to distinguish similar elements and not necessarily to describe an order in terms of time, space, arrangement, or any other manner. It should be understood that such terms may be interchangeable where appropriate, and the embodiments of the present invention described herein are capable of operating in an order different from that described or illustrated herein.
[0045] Furthermore, the terms before, after, etc. in the specification and claims are used for descriptive purposes and not necessarily to describe a relative position. It should be understood that such terms may be interchanged with their antonyms where appropriate, and the embodiments of the present invention described herein are capable of operating in other orientations different from those described or illustrated herein.
[0046] It should be noted that the term "comprising" as used in the claims should not be construed as being limited to the means listed thereafter; it does not exclude other elements or steps. Thus, it is construed as specifying the presence of the stated features, integers, steps, or components, but does not exclude the presence or addition of one or more other features, integers, steps, or components or groups thereof. Thus, the term "comprising" encompasses both the case where only the stated features are present and the case where these features and one or more other features are present. Thus, the scope of the expression "a device comprising devices A and B" should not be construed as being limited to a device consisting only of components A and B. This means that, for the purposes of the present invention, the only relevant components of the device are A and B.
[0047] References to "an embodiment" or "embodiments" in the course of this specification mean that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the invention. Thus, the appearances of the phrases "in an embodiment" or "in embodiments" in various places throughout this specification are not necessarily all referring to the same embodiment, but may refer to the same embodiment. In addition, in one or more embodiments, the particular features, structures, or characteristics may be combined in any suitable manner, as will be apparent to those of ordinary skill in the art from the present disclosure.
[0048] Similarly, it should be understood that in the description of exemplary embodiments of the invention, various features of the invention are sometimes combined in a single embodiment, figure, or description thereof to simplify the disclosure and aid in understanding one or more of the various aspects of the invention. However, the method of the present disclosure should not be construed as reflecting an intention that the invention as claimed requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, aspects of the invention lie in less than all of the features of a single foregoing disclosed embodiment. Thus, the claims following the detailed description are hereby expressly incorporated into this detailed description, with each claim standing on its own as a separate embodiment of the invention.
[0049] In addition, as those skilled in the art will understand, although some embodiments described herein include some features and not other features included in other embodiments, combinations of features of different embodiments are meant to be within the scope of the invention and form different embodiments. For example, in the following claims, any of the embodiments claimed may be used in any combination.
[0050] In addition, some embodiments are described herein as methods or combinations of method elements that can be implemented by a processor of a computer system or by other devices that perform functions. Thus, a processor having the necessary instructions for executing such methods or method elements forms a means for executing the methods or method elements. In addition, the elements of the apparatus embodiments described herein are examples of means for performing the functions performed by the elements for the purpose of carrying out the invention.
[0051] In the description provided herein, numerous specific details are set forth. However, it should be understood that embodiments of the invention may be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this description.
[0052] The following terms are provided solely to assist in understanding the invention.
[0053] As used herein, a biological sequence is a sequence of a biopolymer that defines at least the primary structure of the biopolymer. The biopolymer can be, for example, deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or protein. Biopolymers are generally polymers of biomonomers (such as nucleotides or amino acids), but in some cases may further include one or more synthetic monomers.
[0054] As used herein, a "sequence unit" in a biological sequence is an amino acid when the biological sequence is related to a protein, and a codon when the biological sequence is related to DNA or RNA.
[0055] As used herein, a biological subsequence is a part of a biological sequence that is less than the complete biological sequence. A biological subsequence can have, for example, a total length of 100 sequence units or fewer, preferably 50 or fewer, and even more preferably 20 or fewer.
[0056] As used herein, a distinction is made between a "feature biological subsequence" (or "(HYFT TM ) fingerprint"), a "(HYFT TM ) fingerprint data string", and a "(HYFT TM ) fingerprint marker". The first is a subsequence with specific features as explained in more detail below. The second is a data representation of such a HYFT TM fingerprint - optionally combined with additional data (see below) - which can be stored, for example, in a corresponding repository. In some embodiments, one HYFT TM fingerprint data string can represent multiple equivalent HYFT TM fingerprints (e.g., equivalent by encoding the same result, such as in the case of multiple codons encoding the same amino acid, or by translation equivalence; see below). The third is a pointer to a HYFT TM fingerprint, such as a memory address that can locate the HYFT TM fingerprint or a reference that allows finding the HYFT TM fingerprint in a repository of fingerprint data strings. However, given their close relationship, in cases where a strict distinction between these three terms is not required, or where the meaning is clear from the context, these may simply be referred to as "HYFTs TM " in this document.
[0057] As used herein, a distinction is made between a "biological sequence" and a "processed biological sequence". The former is a biological sequence well - known in the art, while the latter is a biological sequence that includes the reconstruction / rewriting of fingerprint markers associated with the HYFT TM fingerprints of the present invention.
[0058] Obviously, HYFT TMNeither a fingerprint data string, a processed biological sequence, nor a repository storing these can be considered cognitive data and is not directed to a (human) user. Instead, it is intended to be used as functional data in various computer-implemented methods by a computer (or similar technological system) and is constructed to that effect. For example, the repository can be structured like a relational database (e.g., based on SQL) or a NoSQL database (e.g., a document-oriented database such as an XML database). Similarly, HYFT TM A fingerprint data string and / or a processed biological sequence can be constructed as suitable entries for such a database.
[0059] As used herein, some concepts will be illustrated by examples related to proteins, and it will be assumed that the possible monomer sequence units are 20 canonical (or "standard") amino acids. However, it is clear that this is merely for simplicity of illustration, and similar embodiments can equally be formulated with an extended number of amino acids (e.g., adding non-canonical amino acids or even synthetic compounds), or be related to DNA or RNA. In the case of DNA or RNA, the link between DNA or RNA and a protein can be easily established through the correspondence between codons and amino acids.
[0060] As used herein, "secondary / tertiary / quaternary" means "secondary and / or tertiary and / or quaternary".
[0061] In the present invention, it has been surprisingly recognized that in the case where the primary structure of a biological sequence was previously assumed to consist of sequence units selected essentially independently, such that in principle there are, for example, m biological sequences of length n based on m possible sequence units, (e.g., 20 n biological sequences based on 20 standard amino acids), (e.g., 20 n), which is not actually observed in essence. In fact, it has been found that starting from a certain length, not all theoretical combinations can be seen. Just give one example: The protein subsequence "MCMHNQA" is not found in any protein in the public database. It has been taken into account that this is not just an interruption in the database, but this absence has a physical origin and / or a chemical origin. Without being bound by theory, just list one possible effect, the steric hindrance of adjacent amino acids (such as "MCMHNQ" in the above example) can prevent one or more other amino acids (such as "A" in the above example) from binding to it. Therefore, once the missing subsequence has been identified, computational studies can be used to verify whether this subsequence is likely to occur, or whether its existence is physically impossible (or unlikely, for example because it is chemically unstable). The "certain length" mentioned above depends on the dataset under consideration, but for example corresponds to about 5 or 6 amino acids in the publicly available protein sequence database (which basically reflects the total diversity that is essentially visible). For a more limited set (for example, a set filtered based on specific criteria or for a specific biological sequence database, such as a set formulated for a specific domain), for a length of about 4 or 5, it has been found that there are less than m n theoretical maximum number of combinations.
[0062] At the same time, because the subsequence "MCMHNQA" does not exist, the subsequence "MCMHNQ" is not only a random combination of 5 amino acids, but also acquires an additional meaning; such subsequences will be further referred to as "characteristic biological subsequences" or "(HYFT TM ) fingerprints". Due to the additional meaning or significance of these HYFT TM fingerprints, it can be considered that the present invention processes biological sequence information in a more semantic way. Generally speaking, a characteristic subsequence is characterized in that it has fewer possible options (i.e., a lower combination number) for its consecutive sequence units (i.e., the sequence units that directly follow or precede it) than the maximum number of sequence units (i.e., the total number of different sequence units available for it; for example, less than 20 standard amino acids); in other words, at least one of the sequence units cannot follow (or precede) it. However, a more stringent definition can be selected: for example, only those subsequences with 15 or fewer sequence units may follow it, or 10 or fewer, 5 or fewer, 3, 2 or even 1. In addition, each such subsequence can be optionally regarded as a HYFT TM fingerprint, or only those subsequences that do not already contain another HYFT TM fingerprint are regarded as HYFT TM fingerprints (i.e., non-redundant). For example: taking "MCMHNQ" as a HYFT TMFor a fingerprint, there will be longer subsequences that include "MCMHNQ" and also have fewer sequence units that can follow (or precede) it than the theoretical number; in such cases, it is possible to choose to consider both the longer subsequence and "MCMHNQ" as HYFTs TM fingerprints, or consider only "MCMHNQ" as a HYFT TM fingerprint. The latter approach is generally likely to be preferred in order to control the size of the HYFT TM repository of data strings while accelerating the methods associated therewith. In fact, as the string length increases, searching for a match to the string in a biological sequence generally becomes more resource-intensive and slower. Additionally, as the size of the HYFT TM repository of data strings increases, searching for and retrieving a specific HYFT TM data string generally takes longer. In this non-redundant approach, longer subsequences with limited combinatorial possibilities can still be identified, but will then be identified as HYFTs TM patterns (with or without spacing). Thus, the advantages provided by this method do not necessarily result in a corresponding loss of information. Nevertheless, note that the former approach is still possible and doing so is still superior to the prior art
[0063] Subsequently, it was surprisingly found that a finite set of characteristic biological subsequences can be identified. Additionally, it was observed that these characteristic biological subsequences strike a balance between being specific enough such that not every characteristic biological subsequence is found in every biological sequence, and being common enough such that known biological subsequences generally include at least one of these HYFT TM fingerprints
[0064] From the narrative provided above, a protocol can be formulated for identifying HYFT TM fingerprints and constructing the corresponding repository of HYFT TM data strings (or "HYFT TM repository"). In fact, since the aim is to identify those subsequences in a biological sequence database that have limited combinatorial possibilities, it is sufficient to mine for subsequences that do not occur in the biological sequence database. Once such non-occurring subsequences are identified (e.g., "MCMHNQA"), the subsequence that is one sequence unit shorter (e.g., "MCMHNQ") corresponds to a HYFT TM fingerprint (provided the shorter subsequence actually occurs). Once identified, additional data about the HYFT TM fingerprint can be deduced. For example, the identified HYFT can be searched for in the biological sequence database TMCombinations of fingerprints with other sequence units (e.g., replacing the "A" in "MCMHNQA" with one of the other possible amino acids each time) and counting the number of combinations that occur to obtain the number of combinations. Optionally, non-found combinations can also be stored separately; these combinations can be used, for example, for error detection. Additionally, since the correspondence between DNA, RNA, and proteins is generally known through applicable codon tables, once a specific type of HYFT TM fingerprint (e.g., a protein HYFT TM ) is identified, it can be translated into corresponding HYFTs TM of different types (e.g., DNA and / or RNA HYFTs TM ). By repeating the above process and storing at least the identified HYFTs TM —optionally together with any additional data and translated HYFTs TM —a repository of HYFT TM fingerprint data strings can be constructed. As an alternative or supplement, at least some HYFT TM fingerprints can be discovered by experimental or computational methods, e.g., by synthesizing or simulating various subsequences and then identifying those subsequences that cannot—or are unlikely to—occur in the context of the biological sequence database under consideration.
[0065] In the above, the biological sequence database can be a publicly available database, such as the Protein Data Bank (PDB), or a proprietary database. In an embodiment, the biological sequence database can be a combination of multiple individual databases. For example, a repository of HYFT TM fingerprint data strings can be formulated from a biological sequence database that combines as many (trusted) biological sequence databases as are accessible, thereby seeking a general repository of HYFT TM fingerprint data strings that essentially represent all biological sequences that have been found in nature. Conversely, in a specific domain, it can be shown to be productive to construct a specific repository of HYFT TM fingerprint data strings based on the biological sequence database representing the specific domain. In an embodiment, this specific repository can contain HYFTs TM that are not present in the general repository because they do exist in nature but not within this specific domain. Similarly, a repository of HYFT TM fingerprint data strings can be constructed for synthetic sequences with their own specific content.
[0066] Based on the above findings, new methods for processing biological sequence information at all different but interrelated stages can be formulated. These methods can be considered analogous to performing more lexical analysis on sequences. The results are schematically depicted inFigure 1 which shows the complexity scaling of biological sequence information as the number (n) of sequence units increases. This complexity can be the total number of possible combinations of sequence units, but it is also related to the computational effort (e.g., time and memory) required to process it (e.g., perform a similarity search). The solid curve depicts the number of theoretical combinations assuming that all sequence units are independently selected, scaled by m n which also corresponds to the scaling of currently known algorithms. The dashed curve depicts the number of actual combinations essentially found (as observed in the present invention), where the curve deviates from m at around 5 or 6 sequence units n and flattens out asymptotically at high n. The dashed line shows the number of sequences that first correspond to a characteristic sequence, for which the number of sequence units that can follow it equals 1; "first" here means that if a longer sequence contains a HYFT TM fingerprint that has already been counted, it is never counted again. Thus, when its definition is chosen to have only 1 possible sequence unit following it and not to include another (shorter) HYFT TM fingerprint subsequence, the latter corresponds to the number of HYFT TM fingerprints of length n (as observed in the present invention) (see above).
[0067] Figure 2 depicts the predicted benefits of the present invention over time, where the markings on the bottom axis depict the present. Curve 1 shows Moore's Law as a reference. Curve 2 shows the total amount of sequencing data collected. Curve 3 shows the total cost of processing and maintaining the sequencing data. By processing biological sequence information as proposed in the present invention, it is expected that the total required storage for sequencing data and the total cost of data processing and maintenance will decrease, as depicted in Curves 4 and 5 respectively.
[0068] Note that although a repository of HYFT TM fingerprint data strings is typically constructed for a specific biological sequence database (or a combination thereof), this does not mean that HYFT TM fingerprint data strings are only applicable to processing biological sequences in the specific biological sequence database. In fact, a general repository of HYFT TM fingerprint data strings can be used, for example, to process more specific biological sequences. In other cases, a specific repository of HYFT TM fingerprint data strings can be used in the context of biological sequences that fall outside the databases used to formulate the repository. In both cases, favorable results can still be obtained. In any case, one can always determine through trial and error whether an existing HYFT TM fingerprint data string repository can be used for a specific application, or use a dedicated HYFTTM whether a repository of fingerprint data strings yields better results. Similarly, HYFT TM the repository of fingerprint data strings does not strictly need to cover all HYFT TM fingerprints that can be found in a biological sequence database. In fact, partial repositories have already yielded useful results. Such a partial repository can be, for example, a repository associated with HYFT TM fingerprints of a selected length (i.e., as opposed to HYFT TM fingerprints of any length).
[0069] The present invention utilizes a repository of fingerprint data strings. Accordingly, a repository of fingerprint data strings for a biological sequence database is described, each fingerprint data string representing a characteristic biological subsequence composed of sequence units, each characteristic biological subsequence having a combination number less than the total number of different sequence units available in the biological sequence database, the combination number of the biological subsequence being defined as the number of different sequence units that appear as consecutive sequence units of the biological subsequence in the biological sequence database. In Figure 4 FIG. 100 schematically depicts a repository (e.g., a database) of fingerprint data strings, which will be discussed in more detail below.
[0070] An advantage of embodiments of the present invention is that a repository of fingerprint data strings corresponding to characteristic biological subsequences can be provided. Another advantage of embodiments of the present invention is that the biological subsequences do not have to be of a single length, such as in the case of k-mers.
[0071] An advantage of embodiments of the present invention is that other data, such as metadata, can be included in the repository, such as data about sequence units that can be contiguous with the characteristic biological subsequence (i.e., directly after or directly before the characteristic biological subsequence), data about the secondary / tertiary / quaternary structure of the characteristic biological subsequence (e.g., when the characteristic biological subsequence is present in a biopolymer), data about the relationships between fingerprints (e.g., data related to the relationship between the characteristic biological subsequence and one or more other characteristic biological subsequences), etc.
[0072] In an embodiment, the repository can include at least a first fingerprint data string representing a first characteristic biological subsequence of a first length and a second fingerprint data string representing a second characteristic biological subsequence of a second length, wherein the first length and the second length are equal to 4 or greater, and wherein the first length and the second length are different from each other.
[0073] In an embodiment, the length may correspond to the number of sequence units. In an embodiment, the length may be up to 500 or less, such as up to 100 or less, preferably 50 or less, and still more preferably 20 or less. In an embodiment, the first length and the second length may be equal to or greater than 5, preferably equal to or greater than 6. In an embodiment, the length of the characteristic biological subsequence may be between 4 and 20, preferably between 5 and 15, and still more preferably between 6 and 12.
[0074] In an embodiment, the repository of fingerprint data strings may include at least 3 fingerprint data strings of different lengths from each other, preferably at least 4, still more preferably at least 5, and most preferably at least 6. Since the characteristic biological subsequence is not defined by its length, but by the number of possible sequence units that follow (or precede) it, the set of characteristic biological subsequences typically advantageously includes subsequences of different lengths. The repository of fingerprint data strings in the present invention differs from, for example, a set of k-mers (as known in the art) in that it includes biological subsequences of different lengths. In addition, a set of k-mers typically includes every permutation of a fixed length k (i.e., every possible combination of sequence units); this is not the case for the current repository of fingerprint data strings.
[0075] In an embodiment, the fingerprint data string may be a protein fingerprint data string, a DNA fingerprint data string, an RNA fingerprint data string, or a combination thereof. In an embodiment, the characteristic biological subsequence may be a characteristic protein subsequence, a characteristic DNA subsequence, or a characteristic RNA subsequence. In an embodiment, the repository of fingerprint data strings may include protein fingerprint data strings, DNA fingerprint data strings, RNA fingerprint data strings, or a combination of one or more of these (e.g., consisting of protein fingerprint data strings, DNA fingerprint data strings, RNA fingerprint data strings, or a combination of one or more of these). In an embodiment, the characteristic protein subsequence may be translated into a characteristic DNA or RNA subsequence, and vice versa. Such translation may be based on well-known DNA and RNA codon tables. Similarly, a protein fingerprint data string may be translated into a DNA or RNA fingerprint data string. In an embodiment, the repository of DNA or RNA fingerprint data strings may include information about equivalent codons (i.e., codons that encode the same amino acid). Such information about equivalent codons may equally be included in the fingerprint data string or stored separately from this in the repository. In a particular embodiment, the fingerprint data string may be in a sequence-independent format; this means that the fingerprint data string and the surrounding systems and processes enable it to be quickly compared with DNA, RNA, and protein sequences. This may be achieved, for example, by enabling the method using the fingerprint data string to perform the necessary translation on the fly. Such fingerprint data strings advantageously allow for the establishment of a single repository that is generally applicable to data strings across sequence types.
[0076] In an embodiment, a repository of fingerprint data strings may further include additional data for at least one of the fingerprint data strings. In a preferred embodiment, the data may be included in the fingerprint data string. In an alternative embodiment, the data may be stored separately from the fingerprint data string. In an embodiment, the additional data may include one or more of combinatorial data, structural data, relational data, positional data, and orientation data.
[0077] In an embodiment, combinatorial data may be data related to one or more sequence units that may be contiguous with a characteristic biological subsequence (e.g., may truly directly appear before or after it, e.g., those stable combinations) when the characteristic biological subsequence is present in a biological sequence. In an embodiment, the combinatorial data may include the number of possible sequence units, the possible sequence units themselves, the likelihood (e.g., probability) of each sequence unit, etc.
[0078] In an embodiment, structural data may be structural information and / or spatial shape information embedded in the fingerprint data string, e.g., data related to the secondary / tertiary / quaternary structure of a characteristic biological subsequence when the characteristic biological subsequence is present in a biopolymer. In an embodiment, the structural data may include the number of possible structures, the possible structures themselves, the likelihood (e.g., probability) of each structure, etc. In the case of multiple possible secondary / tertiary / quaternary structures for a given characteristic biological subsequence, in an embodiment, the repository may include a separate entry for each combination of the characteristic biological subsequence and the associated secondary / tertiary / tertiary structure. In an alternative embodiment, the repository may include one entry that includes the characteristic biological subsequence and multiple associated secondary / tertiary / quaternary structures. In an embodiment, secondary / tertiary / quaternary structures may be more relevant for proteins compared to DNA and RNA - especially quaternary structures.
[0079] In an embodiment, relational data is data related to the relationship between a characteristic biological subsequence and one or more other characteristic biological subsequences. In an embodiment, the relational data may include other characteristic biological subsequences that typically appear in its vicinity, the likelihood that the other characteristic biological subsequences appear in its vicinity, the specific significance (e.g., biological relevance, e.g., trait or secondary / tertiary / quaternary structure) of these characteristic biological subsequences appearing close to each other, etc. In an embodiment, the relationship may be expressed in the form of a path between two or more characteristic biological subsequences. In an embodiment, the relationship may include the order and / or the spacing of the characteristic biological subsequences. In an embodiment, the additional data may further include metadata for constructing the path.
[0080] In an embodiment, positional data may be data related to the spacing relative to the fingerprint data string (e.g., between the characteristic biological sequences it represents).
[0081] In an embodiment, the orientation data can be data related to the orientation (e.g., the inherent orientation) of the fingerprint data string (e.g., the characteristic biological sequence it represents).
[0082] In some embodiments, additional data may have been retrieved from a known dataset; for example, the secondary / tertiary / quaternary structures of several biological sequences are available in the art. In other embodiments, additional data can be extracted from the processed biological sequences described below or from a repository of the processed biological sequences described below. For example, after processing the biological sequences as described below (or constructing a repository of the processed biological sequences as described below), the relationships between the characteristic biological subsequences (e.g., paths) can be extracted and added to the repository of fingerprint data strings; this is schematically depicted in Figure 4 by the dashed arrows from the processed biological sequence 210 and the repository of the processed biological sequences 220 to the repository of fingerprint data strings 100.
[0083] In an embodiment, the fingerprint data string can be inherently oriented. In an embodiment, the fingerprint data string can include an orientation (i.e., can explicitly include an orientation). Since HYFT TM fingerprints are defined based on the actual segments that occur in a biopolymer or biopolymer fragment, the inherent physical, chemical, and structural limitations that inherently occur for the combinatorial possibilities that occur in a biopolymer are inherently present in HYFTs TM ; where "inherently present" is understood to mean that such information is (or at least can be) implicitly associated with the HYFT TM even if it is not explicitly included as additional data in the repository. Thus, since biological sequences themselves generally have an inherent orientation (i.e., according to the 5' to 3' direction in DNA / RNA and the N-terminus to C-terminus in proteins), this same orientation is inherently present in HYFTs TM . This link to the actual segments further defines the limit on the maximum number of biopolymer fragments that can follow after the last character or before the first character in a HYFT TM . The latter can also be explicitly expressed by a parameter (i.e., the number of combinations) representing the total amount of possible subsequent or previous combinations. This also results in HYFT TM having an inherent (strict) orientation.
[0084] In an embodiment, the fingerprint data string can include position information. The characters in HYFTs TM and the characters between HYFTs TM are syntactically related to each other and thus can define the relationships between them or different HYFTs TMThe spacing between. Such positions or spacings can inherently exist in HYFTs TM in the position information.
[0085] In an embodiment, the fingerprint data string may further include structural and / or spatial shape information. The possible structures and / or spatial shapes of certain HYFTs TM or combinations of HYFTs TM are also limited due to inherent physical, chemical, and structural constraints. Such information also inherently exists in HYFTs TM or related HYFTs TM sets.
[0086] In a first aspect, the present invention relates to a method for sequencing a biopolymer or a biopolymer fragment, taking into account the information contained in a repository of fingerprint data strings, the method comprising: (a) obtaining, using a sequencer, at least one read of the biopolymer or biopolymer fragment, and (b) processing the read through the following computer-implemented steps: (b1) searching for the occurrence of one or more of the characteristic biological subsequences represented by the fingerprint data string in the read and (b2) validating or rejecting the read by determining whether the sequence units contiguous to the characteristic biological subsequence are consistent with the combined data in the repository, and / or (b1') searching for the occurrence of one of the characteristic biological subsequences represented by the fingerprint data string in the head and / or tail of the read and (b2') predicting one or more contiguous sequence units of the read from the combined data in the repository. Here, the repository of fingerprint data strings is used for a biological sequence database, each fingerprint data string represents a characteristic biological subsequence composed of sequence units, each characteristic biological subsequence has a combination number in the biological sequence database that is less than the total number of different sequence units available to it, the combination number of the biological subsequence is defined as the number of different sequence units that appear as contiguous sequence units of the biological subsequence in the biological sequence database, and the repository further includes combined data representing the different sequence units that appear as contiguous sequence units of the corresponding characteristic biological subsequence in the biological sequence database. Figure 3 Schematically shows a sequencer 350 that uses the information contained in a repository 100 of fingerprint data strings to sequence a biopolymer (fragment) 500.
[0087] In an embodiment, the obtained read may be an initial (e.g., temporary or partial) biological sequence.
[0088] In an embodiment, the search in step b1 and / or step b1' may be as described in step b of the method for processing biological sequences described below.
[0089] Relative to step b2, since the repository contains information about what can be in a HYFT TMCombined data of sequence units that occur after the fingerprint (e.g., before or after), so this information can be advantageously used to verify whether the read is consistent with it. If not, the temporary biological sequence can be rejected and redone. Alternatively, the same purpose can be achieved by directly matching it with an undiscovered biological sequence rather than matching the read with the HYFT TM fingerprint itself (see above). Alternatively, such consistency verification can be combined with the use of additional data such as, for example, structural data, relational data, position data, and / or orientation data (see above). Such combinations can, for example, allow the rejection of reads that are indeed consistent with a known HYFT TM fingerprint but are not in the context set by the additional data.
[0090] Relative to step b2', based on the same combined data, some HYFT TM fingerprints (or combinations of HYFT TM fingerprints) have very limited combination possibilities (i.e., corresponding to a low combination number). For example, in the case of a HYFT TM fingerprint with a combination number of 1, the next sequence unit is known. This information can be advantageously used to accelerate sequencing by directly appending the said sequence unit to the read; thus allowing the actual sequencing to skip the said sequence unit. In an embodiment, the repository may contain data on a series of two, three, or more sequence units that together are the only possible options that occur after a particular HYFT TM fingerprint. In this case, the entire series can be advantageously appended directly to the read; thus allowing the actual sequencing to skip these units. Similarly, if the repository indicates that for an observed HYFT TM fingerprint, a limited number (but more than 1) of options are available as other sequence units (e.g., two or three options), then this information can still allow the sequencer to more quickly identify the specific sequence unit in this instance. Additionally, for such HYFT TM fingerprints with a low combination number, by combining the combined data with the use of additional data, the number of possibilities in the current case can be reduced to 1 (or at least the possibilities can exceed a predetermined threshold). Similarly, this combination can set a context that allows the rejection of some combination possibilities, thus, for example, reducing the remaining number to 1 and thereby revealing the subsequent sequence unit.
[0091] In an embodiment, step b2 and / or b2' can thus include using one or more of structural data, relational data, position data, and orientation data; as described above with respect to the repository of fingerprint data strings.
[0092] In an embodiment, sequencing may include obtaining a plurality of reads for the biopolymer or biopolymer fragment using a sequencer (e.g., a sequencing system). In an embodiment, step b may be initiated before all reads of the biopolymer or biopolymer fragment are obtained.
[0093] In an example, step b may include parsing the reads (e.g., using the information of the repository of fingerprint data strings); for example, according to the methods for processing biological sequences described below. In an embodiment, step b may include parsing at least one of the plurality of reads before all reads of the biopolymer or biopolymer fragment are obtained.
[0094] In an embodiment, the method may include another step of aligning (e.g., matching) the processed reads (e.g., included in step b); for example, by aligning and / or assembling according to the methods for comparing biological sequences described below. In an embodiment, the alignment may include using the characteristic biological subsequences identified in step b1 and / or b1'. In an embodiment, the fingerprint data strings may be inherently oriented and may include position information. In an embodiment, the alignment may include aligning the processed reads with an oriented graph. In an embodiment, the method may include aligning at least one of the plurality of processed reads before all reads of the biopolymer or biopolymer fragment are obtained. In at least some embodiments, the alignment may be aligning the processed reads with an oriented acyclic graph.
[0095] In some embodiments, Navarro-Levenshtein matching may be used to perform the alignment. A more detailed description of Navarro-Levenshtein matching can be found, for example, in Navarro, Theoretical Computer Science 237(2000)455-463. Based on the results in one or more of the data processing steps described above, feedback information regarding the sequencing may be generated. Such information may be used to control the sequencing process or to control the corresponding data processing. Such control may include terminating the sequencing process, for example, if sufficient information is available, then identifying one or more reads as incorrect and ignoring these in further data processing...
[0096] However, in the prior art, the assembly step typically can only start after the sequencing is completely finished, because this sequencing, for example, defines the necessary k-mer table, and the k-mer table can only be constructed when all read information is available. The methods and systems according to embodiments of the present invention allow for a progressive and parallel process of constructing subgraphs and obtaining reads. In this way, advantageously, an immediate analysis of the generated reads can be performed, even if, for example, the sequencing is still running and not all reads are available. The latter allows for a dynamic analysis of the data, whereby data analysis is performed during data generation, such as sequence data generation. In some embodiments, sequence data analysis will be able to be performed synchronously with data generation, such as sequence data generation. Nevertheless, it should be noted that data analysis can alternatively be performed separately from data generation.
[0097] The above principles lead to fast data analysis systems and methods. The above principles further allow for the direct incorporation of sequence analysis into a sequencer (i.e., a sequencing machine; see below), thus allowing for fast sequence data generation and analysis, and even optionally online analysis. In this way, relevant outputs may already be generated in the sequencer. Alternatively, similar advantages can be achieved by connecting the sequencer to a data processing system via a streaming data connection (see below); for example, in a distributed computing environment.
[0098] In an embodiment, the method may further include identifying changes in the assembled biological sequence; such as indel mutations, deletions, insertions, and / or duplications.
[0099] In an embodiment, the method may further include folding the processed reads by sorting them. It should be noted that the folding step in the embodiments of the present invention is not based on dynamic programming. Each HYFT TM has a specific number of bits, which can be reduced / optimized by Shannon entropy. HYFTs TM and additional reads can be sorted or classified according to the amount of information (bits) they possess. Since this is not equal for each HYFT TM because the next combination number can reach n - 1, there will be HYFTs TM and corresponding read patterns with a very small number of bits, as well as HYFTs TM and read patterns that require a higher number of bits. Therefore, in the sorting mechanism, a ready global bit threshold can be set to optimize the amount of bits used at each moment during the calculation process. And at most, fully maximize the hardware that must be used through parallelization in order to perform these given tasks. In this way, parallelization can be performed, which results in acceleration and true optimization. In some embodiments, the sorting can be performed based on length. In an embodiment, the classification can be performed based on the position of the HYFT TM in the read.
[0100] In an embodiment, the method may further include converting the obtained data into sub-read graphs and / or read graphs.
[0101] In an embodiment, the method may further include removing dead ends and / or loops.
[0102] In an embodiment, the method may include dynamically adjusting the sequencing based on information obtained from the processing and / or alignment. In an embodiment, the dynamic adjustment may include providing feedback on the number of reads that need to be obtained using the sequencing system. In an embodiment, the dynamic adjustment may include providing feedback on reads that will be ignored as erroneous reads based on information obtained from the processing and / or alignment.
[0103] In an embodiment, the method may include backtracking towards or until a read. In an embodiment, the method may further include capturing metadata, such as a read ID and maintaining the read ID throughout the process. This may advantageously facilitate backtracking, such as backtracking errors or uncertainties in reads.
[0104] According to an embodiment of the present invention, the construction of subgraphs and corresponding processing may be performed in separate threads. This may be further facilitated, for example, by an auto-completion function that may be inherently introduced in embodiments of the present invention. If a certain confidence threshold (equivalent to sufficient coverage) is reached in graph or subgraph construction, no other read information is required to complete the original string reconstruction. Such information may be used as feedback, and based on such information, it may be decided to terminate the sequencing. The latter may be performed based on human intervention, but may also be automated, and the controller may use feedback from the system to determine when the sequencing should be terminated.
[0105] According to an embodiment of the present invention, the method may include the steps of generating feedback information and controlling the sequencing based on the feedback information. Such steps of controlling the sequencing may include one or more of the following: deciding when the sequencing can be terminated based on sufficient information obtained, deciding not to use certain reads in view of detected errors, deciding to collect other or different types of reads...
[0106] In a second aspect, the present invention relates to a sequencer adapted to perform step a of the method according to any embodiment of the first aspect. Figure 3 Sequencer 350 is schematically shown, which sequences a biopolymer (fragment) 500 using information contained in a repository 100 of fingerprint data strings.
[0107] In some embodiments, the sequencer may be adapted to perform the method according to any embodiment of the first aspect (e.g., perform steps a and b).
[0108] In other embodiments, the sequencer may be adapted to transfer reads to a data processing system (e.g., for performing step b). In an embodiment, the sequencer may further be adapted to receive feedback from the data processing system (e.g., after or during step b). The received feedback may be, for example, the output of the data processing system or may be instructions for the sequencer. The instructions may be according to a dynamic adaptation sequencing method (see above) and may include, for example, feedback on whether to terminate sequencing, whether to reacquire certain reads, etc.
[0109] In an embodiment, the sequencer may be a DNA, RNA sequencer, or protein sequencer or a combination thereof. In an embodiment, the sequencer may be an array machine. For example, the sequencer may be a first-generation, next-generation, or third-generation DNA / RNA sequencer, a microarray, or a mass spectrometry device. In an embodiment, the sequencer may combine multiple sequencing techniques, such as in a gene expression array.
[0110] Sequencers are generally more specialized devices and typically may include other technical means for performing sequencing. However, this does not exclude that the sequencer may also be configured to perform one or more other methods (e.g., sequence assembly); in such a case, the sequencer may also be referred to as a sequence assembler, for example. Similarly, the sequencer may be part of a distributed computing environment (see below), where, for example, a client sequencer performs physical sequencing and communicates with a cloud-based data processing system.
[0111] Such a sequencer may be such a sequencer, or a combination of a sequencer and a sequence assembler. In an embodiment, the sequencer may be adapted to obtain reads of a biopolymer or a biopolymer fragment and to analyze the reads before all reads; for example, simultaneously with receiving additional reads. In an embodiment, the sequencer may thus include a processor for processing incoming reads while obtaining additional reads. In addition, the sequencer in some embodiments may include a controller for controlling either the reception of reads and / or data processing according to the obtained results. Thus, the controller may include a feedback loop for controlling the sequencer based on feedback obtained from the processor for processing incoming reads.
[0112] In a third aspect, the invention relates to a system comprising: (i) a sequencer according to the second aspect, and (ii) a data processing system adapted to receive reads from the sequencer and to process the reads by performing step b of the method according to any embodiment of the first aspect.
[0113] In an embodiment, the data processing system may be located on-site (e.g., in the same room) or off-site (e.g., in the cloud) relative to the sequencer.
[0114] In a fourth aspect, the invention relates to a computer program comprising instructions which, when executed by a computer, cause the computer to perform the method according to any embodiment of the first aspect.
[0115] In a fifth aspect, the invention relates to a computer-readable medium comprising instructions which, when executed by a computer, cause the computer to perform the method according to any embodiment of the first aspect.
[0116] A computer-implemented method for constructing and / or updating a repository of fingerprint data strings as described above is also described, comprising: (a) identifying characteristic biological subsequences in a biological sequence database, the characteristic biological subsequences having a number of combinations less than the total number of different sequence units available to them, the number of combinations of the biological subsequences being defined as the number of different sequence units that occur as consecutive sequence units of the biological subsequence in the biological sequence database; (b) optionally, translating the identified characteristic biological subsequences into one or more additional characteristic biological subsequences; and (c) populating the repository with one or more fingerprint data strings representing the identified characteristic biological subsequences and / or the one or more additional characteristic biological subsequences.
[0117] A computer-implemented method for processing biological sequences is also described, comprising: (a) retrieving one or more fingerprint data strings from a repository of fingerprint data strings as described above, (b) searching for the occurrence of characteristic biological subsequences represented by the one or more fingerprint data strings in a biological sequence, and (c) constructing a processed biological sequence comprising, for each occurrence in step b, a fingerprint marker associated with the fingerprint data string representing the occurring characteristic biological subsequence. Figure 4 The sequence processing unit 310 is schematically shown, which processes the biological sequence 200 using the repository 100 of fingerprint data strings to obtain a processed biological sequence 210.
[0118] Advantages of embodiments of the invention are that biological sequences can be processed relatively easily and efficiently. Another advantage of embodiments of the invention is that biological sequences can be analyzed in a lexical or even semantic manner.
[0119] An advantage of embodiments of the invention is that a processed biological sequence can be constructed by replacing the identified characteristic biological subsequences therein with markers associated with the corresponding fingerprint data strings.
[0120] Advantages of embodiments of the present invention are that parts of a biological sequence that do not correspond to one of the characteristic biological subsequences can be processed in various ways. Another advantage of some embodiments is that the biological sequence can be processed in a completely lossless manner (i.e., no information is lost due to processing). Another advantage of alternative embodiments of the present invention is that the biological sequence can be processed in a way that extracts more important information in a more compressed format.
[0121] An advantage of embodiments of the present invention is that the processed biological sequence can be compressed so that it occupies less storage space than the unprocessed counterpart.
[0122] An advantage of embodiments of the present invention is that the matching of parts of a biological sequence to characteristic biological subsequences is not limited to the primary structure, but can also consider secondary / tertiary / quaternary structures.
[0123] An advantage of embodiments of the present invention is that the secondary / tertiary / quaternary structure of a biological subsequence can be at least partially elucidated based on the known secondary / tertiary / quaternary structure of the characteristic biological subsequence contained therein. Another advantage of embodiments of the present invention is that it can assist or facilitate the design of biological sequences (such as protein design).
[0124] In an embodiment, the biological sequence to be processed can be the biological sequence of a biological polymer fragment that can be obtained by a method for sequencing according to the first aspect.
[0125] In some embodiments, the label can be a reference string. Such a reference string can, for example, point to a corresponding fingerprint data string in a repository. In other embodiments, the label can be the fingerprint data string itself or a part thereof.
[0126] In an embodiment, the biological sequence can include: (i) one or more first parts, each first part corresponding to one of the characteristic biological subsequences represented by one or more fingerprint data strings, and (ii) one or more second parts, each second part not corresponding to any of the characteristic biological subsequences represented by one or more fingerprint data strings. In an embodiment, constructing the processed biological sequence in step c can include replacing at least one first part with a corresponding label. In an embodiment, constructing the processed biological sequence in step c can further include adding position information about the first part to the processed biological sequence (for example, attaching it to the label). In an embodiment, constructing the processed biological sequence in step c can include leaving at least one second part unchanged, and / or replacing at least one second part with an indication of the length of the second part, and / or completely removing at least one second part. When leaving the second part unchanged, it is advantageously possible to process the biological sequence in a completely lossless manner.
[0127] In an embodiment, the processed biological sequence may be formulated in a compressed format. For example, by replacing a characteristic biological subsequence (i.e., the first part) with a reference string and / or by replacing the second part with an indication of its length or completely removing the second part, a processed biological sequence is obtained that requires less storage space than the original (i.e., unprocessed) biological sequence. Additional data compression can be achieved by exploiting paths that can represent multiple fingerprints by their mutual relationships.
[0128] In an embodiment, one or more fingerprint data strings may be in a biological format different from the biological sequence (e.g., protein vs. DNA vs. RNA sequence information), and step b may further include translating or transcribing the characteristic biological subsequence prior to the search.
[0129] In an embodiment, the search in step b may comprise searching for partial matches or equivalent matches (e.g., equivalent codons, or different amino acids that produce the same secondary / tertiary / quaternary structure). In an embodiment, the search in step b may take into account the secondary / tertiary / quaternary structure of the characteristic biological subsequence. Secondary, tertiary, and quaternary structures are generally more evolutionarily conserved and often undergo changes in the primary structure that do not alter the function of the biopolymer, e.g., because the secondary / tertiary / quaternary structure of its active site is substantially conserved. Thus, the secondary / tertiary / quaternary structure can reveal relevant information about the biopolymer that would be lost when strictly searching for an exact match in the primary structure.
[0130] In a preferred embodiment, the occurrences of the characteristic biological subsequence may be searched for in step b in a specific order. In an embodiment, the order may be based on the length and the number of combinations of the characteristic biological subsequence. In an embodiment, the search may be performed in an order starting with the longest characteristic biological subsequence having the lowest number of combinations and ending with the shortest characteristic biological subsequence having the highest number of combinations. In a preferred embodiment, the order may be from the longest characteristic biological subsequence to the shortest characteristic biological subsequence, and - for characteristic biological subsequences of the same length - from the lowest number of combinations to the highest number of combinations. In other embodiments, the order may be from the lowest number of combinations to the highest number of combinations, and - for characteristic biological subsequences having the same number of combinations - from the longest characteristic biological subsequence to the shortest characteristic biological subsequence. In an embodiment, the order may further take into account additional data (e.g., to determine the order within a set of characteristic biological subsequences having the same length and the same number of combinations), such as context data.
[0131] In an embodiment, the method may include another step d of at least partially inferring the secondary / tertiary / quaternary structure of the processed biological subsequence based on the structural data as described above after step c. Such at least partial elucidation of the secondary / tertiary / quaternary structure may assist and / or facilitate biological sequence design. In embodiments where the single primary structure of the characteristic biological subsequence is linked to multiple secondary or tertiary or quaternary structures, the secondary / tertiary / quaternary structure may be disambiguated based on the context in which the characteristic biological subsequence is found, e.g., the characteristic biological subsequence it surrounds. For example, the information required for such disambiguation may be found in a repository of fingerprint data strings, in the form of data (e.g., relationship data) related to the relationship in terms of the secondary / tertiary / quaternary structure between the characteristic biological subsequence and one or more other characteristic biological subsequences, as described above. For example, it may be known that a particular first HYFT TM fingerprint adopts a helix or turn configuration as the secondary structure, but when a particular second HYFT TM fingerprint is present within a certain spacing from the first HYFT TM it always adopts a helix configuration. In this case, the HYFT TM pattern of the fingerprint - if observed - may be used to disambiguate the secondary structure of the first HYFT TM TM .
[0132] In embodiments where the fingerprint data string is inherently oriented and includes position information, step c may include constructing the processed biological sequence as a directed graph. In an embodiment, the directed graph may be a directed acyclic graph. It should be noted that when referring to an acyclic graph, this does not mean that loops cannot occur, but rather that the entire graph is not cyclic. The resulting graphical representation of the reconstructed sequence as obtained in an embodiment of the present invention may be referred to as a HYFT TM graph. This HYFT TM graph may allow for a general genomic graphical representation.
[0133] In an embodiment, constructing the processed biological sequence may include considering the spacing between different fingerprint data strings, and / or may include considering the orientation (e.g., inherent orientation) of the fingerprint data string to construct a directed graph.
[0134] In an embodiment, constructing the processed biological sequence may include considering the structural and / or spatial shape information embedded in the fingerprint data string for constructing a directed graph, and / or may include considering the syntactic information embedded in the fingerprint data string.
[0135] In an embodiment, the search in step b may consider any one of the positional information, spacing information, secondary and / or tertiary and / or quaternary structures of different elements of the characteristic biological sequence, and / or the structural variations of the characteristic biological subsequence.
[0136] By way of illustration, and not limitation of the embodiments of the present invention, an example of how to search for a certain sequence is shown below. The method includes identifying HYFT present in the sequence to be searched in a first step TM . The method then further includes querying a reference database by searching all sequences in the reference database that also contain said HYFT TM . The different sequences found are then sorted, for example by length, and the positions of the HYFT TM in the sequences are identified. In addition, an alignment is performed. In some embodiments, the Navarro-Levenshtein match may be used to perform the alignment. A more detailed description of the Navarro-Levenshtein match can be found, for example, in Navarro, Theoretical Computer Science 237(2000)455-463. The alignment may be performed using a directed graph, such as a directed acyclic graph. The latter may be a general genomic reference graph, but the embodiments are not limited thereto. The alignment may include identifying variations of a specific sequence. To perform the above steps, the sequences may be further processed, whereby, for example, dead ends and loops may be removed.
[0137] A processed biological sequence is also described, which can be obtained by a computer-implemented method for processing biological sequences as described above. Figure 4 The processed biological sequence 210 is schematically depicted in
[0138] A computer-implemented method for constructing and / or updating a repository of processed biological sequences is also described, which includes populating the repository with the processed biological sequences as described above. Figure 4 A repository construction unit 320 for storing the processed biological sequence 210 into a repository 220 of processed biological sequences is schematically shown.
[0139] An advantage of the embodiments of the present invention is that a repository of processed biological sequences can be constructed and stored.
[0140] A repository of processed biological sequences is also described, which can be obtained by a computer-implemented method for constructing and / or updating a repository of processed biological sequences as described above. Figure 4 The repository 220 is schematically depicted in
[0141] One advantage is that repositories of processed biological sequences can be searched and navigated quickly. Another advantage is that, compared to known databases, the storage size of the repository can be relatively small by populating the repository with compressed processed biological sequences.
[0142] In an embodiment, a repository of processed biological sequences can be combined with a repository of fingerprint data strings.
[0143] In an embodiment, the repository can be a repository of processed biological fragment sequences (i.e., processed biological sequences of biopolymer fragments).
[0144] In an embodiment, the repository can be a database. In some embodiments, the repository of processed biological sequences can be an index repository. For example, the repository can be indexed based on fingerprint markers (corresponding to characteristic biological subsequences) present in each processed biological sequence. In other embodiments, the repository can be a graphical repository.
[0145] Also described is a computer-implemented method for comparing a first biological sequence with a second biological sequence, comprising: (a) processing the first biological sequence by a computer-implemented method as described above to obtain a processed first biological sequence, or retrieving the processed first biological sequence from a repository of processed biological sequences as described above, (b) processing the second biological sequence by a computer-implemented method as described above to obtain a processed second biological sequence, or retrieving the processed second biological sequence from a repository of processed biological sequences as described above, and (c) comparing at least the fingerprint markers in the processed first biological sequence with the fingerprint markers in the processed second biological sequence. Figure 5 The comparison unit 330 is schematically shown, which compares at least the first biological sequence 211 with the second biological sequence 212 to output a result 400.
[0146] Advantages of embodiments of the present invention are that the comparison of biological sequences can be changed from an NP-complete problem or an NP-hard problem to a polynomial-time problem. Another advantage of embodiments of the present invention is that the comparison can be performed in a greatly reduced time and can scale well with an increase in complexity (e.g., an increase in the length or number of biological sequences). Yet another advantage of embodiments of the present invention is that the required computing power and storage space can be reduced.
[0147] Advantages of embodiments of the present invention are that the degree of similarity between biological sequences can be calculated. Another advantage of embodiments of the present invention is that they can be ranked based on the degree of similarity of multiple biological sequences.
[0148] An advantage of embodiments of the present invention is that sequence similarity searches can be performed quickly and easily (e.g., in polynomial time).
[0149] An advantage of embodiments of the present invention is that the compared biological sequences can be aligned easily and quickly (e.g., in polynomial time).
[0150] An advantage of the embodiments is that multiple sequences can also be compared and aligned easily and quickly. Another advantage of the embodiments is that there is no error accumulation during alignment, as is the case in currently known methods (e.g., progressive alignment).
[0151] An advantage of embodiments of the present invention is that the sequences of biological polymer fragments can be aligned and combined easily and quickly to reconstruct the original biological polymer sequence.
[0152] By using characteristic biological subsequences (fingerprint markers in the processed biological sequences) according to embodiments of the present invention, the problem of comparing sequences is advantageously reformulated from an NP-complete or NP-hard problem to a polynomial time problem. In fact, identifying fingerprints in a sequence and then comparing sequences based on these fingerprints (which can be considered a lexical method) is computationally much simpler than currently used algorithms (which compare entire sequences, e.g., based on a sliding window method). Thus, the comparison can be performed significantly faster, even when requiring less computational power and storage space, and scales well with increasing complexity (e.g., an increase in the length or number of biological sequences).
[0153] In an embodiment, the second biological sequence can be a reference sequence.
[0154] In an embodiment, step c may include identifying whether one or more characteristic biological subsequences (represented by fingerprint markers) in the processed first biological sequence correspond (e.g., match) to one or more characteristic biological subsequences (represented by fingerprint markers) in the processed second biological sequence. In an embodiment, step c may include identifying whether the corresponding characteristic biological subsequences occur in the same order in the processed first biological sequence as in the processed second biological sequence. In an embodiment, step c may include identifying whether one or more pairs of characteristic biological subsequences in the processed first biological sequence and one or more pairs of corresponding characteristic biological subsequences in the processed second biological sequence have the same or similar (e.g., differ by less than 1000 sequence units, e.g., less than 100 sequence units, preferably less than 50 sequence units, more preferably less than 20 sequence units, most preferably less than 10 sequence units) spacing.
[0155] In an embodiment, step c may further include comparing one or more second portions of the processed first biological sequence with one or more second portions of the processed second biological sequence. In an embodiment, comparing the one or more second portions may include comparing corresponding second portions (i.e., the second portion between adjacent pairs of characteristic biological subsequences present in the processed first biological sequence with the second portion between corresponding adjacent pairs of characteristic biological subsequences present in the processed first biological sequence).
[0156] In an embodiment, step c may further include calculating a measure representing the degree of similarity (e.g., Levenshtein distance) between the first biological sequence and the second biological sequence. In an embodiment, the degree of similarity may be calculated based on a plurality of variables, such as combining a measure of syntactic similarity with a measure of structural similarity.
[0157] In an embodiment, the method may be used in sequence similarity searches by comparing a query sequence with one or more other biological sequences (e.g., corresponding to a sequence database to be searched, e.g., in the form of a repository of processed biological sequences). In an embodiment, the degree of similarity for each of the other biological sequences may be calculated. In an embodiment, the method may include another step of ranking the biological sequences (e.g., by decreasing degree of similarity). In an embodiment, the method may include filtering the biological sequences. The filtering may be performed before and / or after step c. By way of example, the filtering may be performed by selecting only those biological sequences from the database that meet certain criteria, such as based on the organism or group of organisms from which they are derived (e.g., plants, animals, humans, microorganisms, etc.), whether the secondary / tertiary / quaternary structure is known, their length, etc. Alternatively, the filtering may be performed after the comparison based on the same criteria or based on the calculated degree of similarity (e.g., only those sequences that exceed a certain similarity threshold may be selected). Contrary to sequence similarity searches in the prior art, where an alignment step is typically required, followed by establishing a measure of similarity therefrom, according to the embodiments, alignment is not strictly necessary for similarity searches. In fact, in the absence of alignment, similar sequences may already be found by simply searching for sequences with the same fingerprint (optionally also considering their order and their spacing); this in turn allows the search speed to be further accelerated. Nevertheless, the alignment according to the embodiments (see below) is also computationally simplified, such that the alignment may be performed in any way, even without strict requirements.
[0158] The method thus allows the similarity between the first biological sequence and the second biological sequence to be determined (and optionally measured). Such comparisons are also the cornerstone of other methods, such as methods for alignment and assembly (see below).
[0159] In an embodiment, the method can be used to align a first biological sequence with a second biological sequence. In an embodiment, step c can further include aligning the fingerprint markers in the processed first biological sequence with the fingerprint markers in the processed second biological sequence. Figure 5 Schematically shows the output result 400 from the comparison unit 330 (which is better referred to as the "alignment unit 330" in this case), where the biological sequences are aligned by their fingerprint markers.
[0160] Thus, alignment is also simplified in an embodiment because a good alignment may already be obtained by simply aligning the fingerprints. Again, this significantly reduces the computational complexity of the problem. Additionally, in prior art methods, such as those based on progressive alignment, there is an accumulation of alignment errors because a misalignment in one of the earlier sequences typically propagates and causes additional misalignments in later sequences. In contrast, since the same discrete set of fingerprint markers is aligned (or at least attempted to be aligned) within one (or more) alignment each time, there is no such error propagation.
[0161] In an embodiment, the method can further include subsequently aligning corresponding second parts. For example, one of the alignment methods known in the prior art can be used to perform the alignment of the second part. In fact, since the "skeleton" of the alignment has been provided by aligning the fingerprint markers, only the alignment between these markers remains to be filled in. Since each of these second parts is typically relatively short compared to the total biological sequence length, known methods can generally perform such alignments relatively quickly and efficiently.
[0162] In an embodiment, the method can be used to perform a multiple sequence alignment (i.e., the method can include aligning three or more biological sequences). In an embodiment, the method can include aligning the fingerprint markers in the processed third (or fourth, etc.) biological sequence with the fingerprint markers in the processed first and / or second biological sequences. This is schematically depicted in Figure 5 where the alignment unit 330 can also compare and align any number of further processed biological sequences 213 to 216.
[0163] In an embodiment, the method can be used for variant identification. In the case of a sequence alignment between two biological sequences, variant identification can identify variants (such as mutations) between a query sequence and a reference sequence. In the case of a multiple sequence alignment, variant identification can identify possible changes in a set of related sequences (which can include determining their frequency of occurrence); optionally relative to a reference sequence. Additionally, variants can be identified based on the primary structure, but secondary / tertiary / quaternary structures can also be considered. Thus, variants can be based on the primary structure, based on the secondary / tertiary / quaternary structure, and also based on the HYFT in the sequence TM related or relative to the next or previous HYFTTM each possible mutual relationship of distances related to distance information to identify variants. Identifying variants can also be based on changes in the codon table, thus allowing immediate information on DNA, RNA, and amino acid changes to be collected in the same variant analysis.
[0164] In an embodiment, the method can be used to perform sequence assembly. In an embodiment, the method can include: (a) providing a first biological sequence, which is the biological sequence of a first biological polymer fragment, (b) providing a second biological sequence, which is the biological sequence of a second biological polymer fragment or a reference biological sequence, (c) aligning the first biological sequence with the second biological sequence as described above, and (d) merging the first biological sequence and the second biological sequence to obtain an assembled biological sequence. Figure 6 Schematically shows a sequence assembly unit 340, which outputs an assembled biological sequence 510 by first aligning (by its fingerprint label) and then merging any number of biological sequences 500, including at least a first biological sequence 501 and a second biological sequence 502.
[0165] In an embodiment, the method steps a to d can be repeated to align and merge any number of biological polymer fragments.
[0166] For ease of sequencing, longer biological polymers can be fragmented because sequencing of individual fragments is faster and easier (e.g., it can be sequenced in parallel); as is known in the art. Then sequence assembly is typically used to align and merge the fragment sequences to reconstruct the original sequence; this can also be referred to as "read mapping", where the "reads" from the fragment sequences are "mapped" to a second biological polymer sequence. Depending on the type of sequence assembly being performed, e.g., de novo assembly versus mapping assembly, the second biological polymer sequence can be selected as a second biological polymer fragment or a reference sequence as appropriate. Here, de novo assembly is a de novo assembly without using a template (e.g., a scaffold sequence). In contrast, mapping assembly is an assembly by mapping one or more biological polymer fragment sequences to an existing scaffold sequence (e.g., a reference sequence), which is typically similar (but not necessarily identical) to the sequence to be reconstructed. The reference sequence can be based on, for example, a complete genome or transcriptome (a part thereof), or can be obtained from an earlier de novo assembly.
[0167] In an embodiment, the method can include another step e of aligning the assembled biological sequence with the second biological sequence as described above after step d. This additional alignment can be used to perform variant identification of the assembled biological sequence relative to the second biological sequence (e.g., a reference sequence).
[0168] In an embodiment, the fingerprint data string can be inherently oriented and include position information.
[0169] In an embodiment, the method may further include detecting variations such as, but not limited to in this embodiment, insertion-deletion mutations, deletions, insertions, and / or duplications.
[0170] In an embodiment, providing the first biological sequence and / or the second biological sequence may be performed using the methods described above.
[0171] A storage device is also described that includes a repository of fingerprint data strings as described above and / or a repository of processed biological sequences as described above.
[0172] A processing system is further described that includes such a storage device and further includes a processor adapted to obtain a fingerprint data string from the storage device and / or adapted to store a fingerprint data string in the storage device and / or search for a fingerprint data string in the storage device.
[0173] A data processing system is also described that is adapted to (e.g., includes means for) perform any of the computer-implemented methods described above.
[0174] The system can generally take different forms depending on the method it is intended to perform. In an embodiment, the system can be or include a sequence processing unit, a variant identification unit, a repository construction unit, a comparison unit, an alignment unit, or a sequence assembly unit. In an embodiment, a general-purpose data processing device (e.g., a personal computer or a smart phone) or a distributed computing environment (e.g., a cloud-based system) can be configured to perform one or more of these functions. The distributed computing environment can include, for example, server devices and networked client devices. Herein, the server device can perform most of one or more methods, including repositories for storing fingerprint data strings and processed biological sequences. On the other hand, the networked client devices can communicate instructions (e.g., inputs such as query sequences, and settings such as search preferences) to the server device and can receive method outputs.
[0175] A computer program (product) including instructions is also described that, when the program is executed by a computer (system), cause the computer to perform any of the computer-implemented methods described above.
[0176] A computer program product including instructions is further described that, when the program is executed by a computer system, cause the computer system to perform obtaining, searching, or storing fingerprint data strings from, in, or to a repository of fingerprint data strings, respectively.
[0177] A computer-readable medium including instructions is also described that, when executed by a computer (system), cause the computer to perform any of the computer-implemented methods described above.
[0178] Also described is the use of a repository of fingerprint data strings as described above for one or more of the following: sequencing a biopolymer or biopolymer fragment; performing sequence assembly; processing biological sequences; constructing a repository of processed biological sequences; comparing a first biological sequence with a second biological sequence; aligning a first biological sequence with a second biological sequence; performing a multiple sequence alignment; performing a sequence similarity search; and performing variant identification.
[0179] Also described is the use of a processed biological sequence as described above or a repository of processed biological sequences as described above for one or more of the following: comparing a first biological sequence with a second biological sequence; aligning a first biological sequence with a second biological sequence; performing a multiple sequence alignment; performing a sequence similarity search; and performing variant identification.
[0180] In an embodiment, any feature of any embodiment of any of the above aspects may be independently described as appropriate for any other aspect or any embodiment of the other described subject matter.
[0181] Aspects of several embodiments will now be described by a detailed description of several embodiments. Obviously, other embodiments of the present invention can be configured according to the knowledge of those skilled in the art without departing from the true technical teachings of the present invention, and the present invention is limited only by the terms of the appended claims.
[0182] Example 1: Sequencing according to an embodiment of the present invention
[0183] By way of illustration, embodiments of the present invention are not limited thereto, Figure 7 Examples of possible sequencing implementations are shown. The drawings show possible different method steps of a sequencing method according to an embodiment of the present invention. The method includes, after obtaining at least a first read of a biopolymer or biopolymer fragment and generally during further receipt of reads of the biopolymer or biopolymer fragment to be sequenced, parsing incoming, e.g., received, reads with fingerprints, called HYFTs TMAfter parsing, alignment (e.g., matching) can be performed to obtain a graph representing the sequence of a biopolymer or biopolymer fragment. Alignment can be performed by aligning with an oriented graph, such as an oriented acyclic graph. The latter can be a general genomic reference graph, but the embodiments are not limited thereto. Alignment can include identifying changes in a specific sequence. However, intermediate steps can be performed, such as constructing an overview graph, whereby the processed (e.g., parsed) sequences are grouped around one or more common or linked fingerprints among the processed sequences, and the data is folded, for example, by sorting in the overview graph. Such folding can be performed one character at a time, and nodes can be split when the characters are different. The method can also include forming a sub-read graph, whereby dead ends or bubbles are generally removed in this step. It should be noted that removing dead ends and / or bubbles can alternatively or additionally be performed in other steps of the method. The method can also include forming a read graph, wherein the sub-read graphs are combined. As a further illustration, the embodiments of the present invention are not limited thereto, Figures 8 to 11 shows different steps. Figure 7 illustrates the use of HYFTs TM to parse incoming reads. It should be noted that the portions of the sequences shown in the figures do not themselves form part of the present invention, but are introduced only to illustrate the processing of such data. Identify a certain fingerprint of the repository, i.e., HYFT TM in the read. Figure 8 illustrates the construction of an overview graph, whereby different processed sequences are grouped around the found linked HYFT TM for grouping. Figure 9 illustrates folding the constructed overview graph by sorting. The latter can be performed one character at a time, and is performed by splitting nodes when the characters are different. In addition, the sequences covering the nodes can be tracked. Generally, it can start from the HYFT TM fingerprint and generally move in one direction (e.g., to the right). Figure 11 illustrates the cleaning step, where loose ends are removed. Alternatively or in addition, bubbles or small inner loops can also be resolved.
[0184] Example 2: Processing of protein databases
[0185] Example 2a: Analysis of the protein database regarding HYFT TM fingerprints found in the protein database
[0186] To illustrate the prevalence of HYFT TM fingerprints in biological sequence databases, the protein database (PDB) is taken as an example of a large, generally available biological sequence database, and the repository of fingerprint data strings obtained as described above is processed according to the present invention. The results are analyzed regarding various metrics, and their selection is given below.
[0187] Figure 12 and Figure 13 show the HYFT TM coverage (in %) of processed protein sequences up to lengths of 50 and up to lengths over 5000, respectively. Here, the coverage is the part of the total sequence length for which the sequence units belong to the HYFT TM fingerprint. In other words, the coverage is the combined length of one or more first parts divided by the total sequence length.
[0188] For cases up to lengths over 5000, the inverse statistic is shown in Figure 14 , i.e., the part of the total sequence length not covered by the HYFT TM fingerprint (or the combined length of one or more second parts divided by the total sequence length).
[0189] In connection with the above, Figure 15 an overview of the number of HYFTs retrieved for each processed sequence is given in the form of a frequency distribution. TM of the number.
[0190] It is noted that these charts show that at least one HYFT TM fingerprint is found in each processed biological sequence; in fact, no PDB sequence is not covered by one or more HYFTs TM . Further, the HYFT TM patterns widely cover long sequences, where the coverage generally decreases with increasing sequence length. On average, a coverage close to 80% is achieved.
[0191] The typical spacing observed is shown in Figure 16 , which depicts the frequency distribution of the lengths of the second parts occurring before and after the HYFT TM fingerprint.
[0192] Overall, the above results support that virtually every protein sequence (and by extension DNA and / or RNA sequences) can be rewritten as a string of one or more HYFTs TM based on the repository of HYFT TM fingerprint data strings according to the present invention. TM In addition, due to the generally good coverage achieved, the processed sequences still retain the essential characteristics of their unprocessed counterparts; especially when not only the identified HYFTs TM are retained, but they are also extended with additional data (see above), such as the spacing (i.e., the lengths of the second parts) before, between, and after the identified HYFTs TM . It is possible to achieve a HYFT TMHigh-performance indexing for the mode - with near-perfect retrieval rates.
[0193] Example 2b: Effect of the matching strategy employed
[0194] Since different strategies can be employed when processing biological sequences according to the present invention, the differences between two different methods were studied. In the first method, HYFT TM fingerprint occurrences in biological sequences in the PDB database were searched for, including overlapping HYFTs TM , such that the order of the HYFT TM fingerprints became irrelevant. In the second method, a more stringent way of searching the PDB database for biological sequences was used, where the search was performed in the order from the longest HYFT TM fingerprint to the shortest HYFT TM fingerprint, and - in the case of the same length - from the lowest combination number to the highest combination number, and where overlapping of HYFTs TM was not allowed (i.e., where the part found corresponding to HYFT TM was excluded from then on for searching other HYFTs TM ). The goal of the second method was to identify the smallest number of HYFTs TM to describe the processed biological sequence, while still ensuring good coverage of the sequence by not allowing overlap and by favoring more stringent HYFTs TM (i.e., shorter lengths and higher combination numbers) over less stringent HYFTs TM (i.e., longer lengths and lower combination numbers).
[0195] In Figure 17 , the number of different matches found for each biological sequence was plotted relative to each other. It was observed that for the second method, which is more stringent than the first method, there was a roughly linear relationship where approximately 5 times fewer matches were found. These fewer matches corresponded to increased processing time - identifying HYFT TM fingerprints and subsequently using the processed sequences in other methods - and required storage space; nevertheless, the entire sequence was fully characterized. Therefore, the second method was considered to achieve the best balance and was generally preferred.
[0196] Nonetheless, it was noted that the number and nature of the matches found using the first method were fewer and better than comparable k-mer methods. Thus, although the second method was generally superior to the first method, the first method was still superior to methods of the prior art.
[0197] Example 3: Comparison between sequence searches known in the prior art and those described herein
[0198] Example 3a: Using short search strings
[0199] Two separate searches are performed based on the search string "AVFPSIVGRPRHQGVMVGMGQKDSY". This corresponds to a relatively short protein sequence of 25 sequence units in length, which can be, for example, a protein fragment in protein sequencing. Such a search can be used, for example, after fragment sequencing as part of identifying a suitable reference sequencing to be used in sequence assembly together with the fragment.
[0200] The first search is performed using BLAST (Basic Local Alignment Search Tool); more specifically, using "Protein BLAST" (available at: https: / / blast.ncbi.nlm.nih.gov / Blast.cgi?PROGRAM= blastp&PAGE_TYPE=BlastSearch&LINK_LOC=blasthome ). The following search parameters are used: database = Protein Data Bank proteins (pdb); algorithm = Blastp (protein-protein BLAST); maximum target sequences = 1000; short query = automatically adjusts parameters for short input sequences; expectation threshold = 20000; word size = 2; matrix = PAM30; composition adjustment = no adjustment. BLAST takes more than 30 seconds to perform this search and then returns 604 search results.
[0201] On the other hand, based on the principles of the present invention, it is determined that "IVGRPRHQGVM" is a characteristic biological subsequence (i.e., "HYFT TM fingerprint") included in the above short protein sequence. Accordingly, a second search is performed based on the search string "IVGRPRHQGVM" in a repository of processed biological sequences. This repository is based on the same protein database as used in BLAST (i.e., Protein Data Bank; PDB), which has previously been processed using a repository of fingerprint data strings (see above); i.e., identifying and labeling characteristic biological subsequences represented by fingerprint data strings in a collection of publicly available biological sequences. This search returns 661 results. In this case, the time frame required is only 196 milliseconds compared to BLAST. Thus, even for such relatively short sequences, it is observed that the present method is able to reduce the required time by more than 150 times compared to known technical methods.
[0202] Now refer to Figure 18 , Figure 19 and Figure 20 , which show these two searches (BLAST = dashed line; present method = solid line) in terms of their total length ( Figure 18 ), their Levenshtein distance ( Figure 19 ) and longest common substring ( Figure 20) Results in terms of. For each figure, the search results are presented in ascending order with respect to the drawing parameters (i.e., total length, Levenshtein distance, or longest common substring). Additionally, one of the search results, namely the protein sequence 5NW4_V (i.e., the first result listed by BLAST), is selected as the reference for calculating the Levenshtein distance and the longest common substring. As can be observed in these figures, the present method produces a smaller variation in the total length (characterized by a relatively flat segment spanning a significant portion of the results), a significantly lower Levenshtein distance, and a significantly larger longest common substring across the entire range of the search results; compared to the BLAST results. The combination of these indicates that the method of the present invention is capable of identifying results that are more relevant to the search performed.
[0203] Example 3b: Using a Longer Protein as the Search String
[0204] Repeat the previous example, but this time search for the complete protein sequence, 3MN5_A (length 359 sequence units).
[0205] The first search, using BLAST, returns 88 search results.
[0206] On the other hand, based on the principles of the present invention, it is determined that six characteristic biological subsequences (i.e., "HYFT TM fingerprints") can be found in the sequence 3MN5_A; these are represented as:
[0207] +4641474444415052415646_1, +495647525052485147564d_1, +4949544e5744444d454b49_1, +494d464554464e5650414d_1, +494b454b4c435956414c44_1, and +49474d4553414749484554_1,
[0208] where, for example, "49474d4553414749484554" corresponds to the corresponding subsequence in hexadecimal format. Thus, a second search is performed in the same repository of processed biological sequences as in the previous example to find those protein sequences that include the same six characteristic biological subsequences in the same order. This search returns 661 results.
[0209] Now refer to Figure 21 , Figure 22 and Figure 23 , which show these two searches (BLAST = dashed line; the present method = solid line) in terms of their total length ( Figure 21 ), their Levenshtein distance ( Figure 22) and the longest common substring ( Figure 23 ) The results in terms of. For each graph, the search results are presented in ascending order with respect to the plotting parameters (i.e., total length, Levenshtein distance, or longest common substring). In this case, the Levenshtein distance and the longest common substring are calculated with respect to the original query sequence 3MN5_A. As can be observed in these graphs, the characteristics of the search results of the two methods are relatively comparable in extreme cases. However, the method of the present invention produces a stable result in the intermediate range, with a relatively small variation in total length, a lower Levenshtein distance, and a relatively high longest common substring. The combination of these indicates that the method of the present invention is capable of identifying a larger number of relevant results.
[0210] It should be understood that although preferred embodiments, specific structures and configurations, and materials have been discussed herein with respect to the device according to the present invention, various changes or modifications can be made in form and detail without departing from the scope and technical teachings of the present invention. For example, any formula given above merely represents a procedure that can be used. Functions can be added or removed from the block diagrams, and operations can be interchanged between functional blocks. Steps can be added or removed from the methods described within the scope of the present invention. Sequence Listing <110> BioClue NV <120>Biological Sequencing <130> 20023VTr00WO / dw / ac / av <140> EPPCTNYK <141> 2020-02-07 <150> EP19190900.1 <151> 2019-08-08 <150> EP19156086.1 <151> 2019-02-07 <160> 10 <170> BiSSAP 1.3.6 <210> 1 <211> 7 <212> PRT <213> Unknown <220> <223> Unknown <400> 1 Met Cys Met His Asn Gln Ala 1 5 <210> 2 <211> 6 <212> PRT <213> Unknown <220> <223> Unknown <400> 2 Met Cys Met His Asn Gln 1 5 <210> 3 <211> 25 <212> PRT <213> Unknown <220> <223> Unknown <400> 3 Ala Val Phe Pro Ser Ile Val Gly Arg Pro Arg His Gln Gly Val Met 1 5 10 15 Val Gly Met Gly Gln Lys Asp Ser Tyr 20 25 <210> 4 <211> 11 <212> PRT <213> Unknown <220> <223> Unknown <400> 4 Ile Val Gly Arg Pro Arg His Gln Gly Val Met 1 5 10 <210> 5 <211> 11 <212> PRT <213> Unknown <220> <223> Unknown <400> 5 Phe Ala Gly Asp Asp Ala Pro Arg Ala Val Phe 1 5 10 <210> 6 <211> 11 <212> PRT <213> Unknown <220> <223> Unknown <400> 6 Ile Val Gly Arg Pro Arg His Gln Gly Val Met 1 5 10 <210> 7 <211> 11 <212> PRT <213> Unknown <220> <223> Unknown <400> 7 Ile Ile Thr Asn Trp Asp Asp Met Glu Lys Ile 1 5 10 <210> 8 <211> 11 <212> PRT <213> Unknown <220> <223> Unknown <400> 8 Ile Met Phe Glu Thr Phe Asn Val Pro Ala Met 1 5 10 <210> 9 <211> 11 <212> PRT <213> Unknown <220> <223> Unknown <400> 9 Ile Lys Glu Lys Leu Cys Tyr Val Ala Leu Asp 1 5 10 <210> 10 <211> 11 <212> PRT <213> Unknown <220> <223> Unknown <400> 10 Ile Gly Met Glu Ser Ala Gly Ile His Glu Thr 1 5 10
Claims
1. A method for sequencing a biopolymer or a biopolymer fragment (500), taking into account information contained in a repository (100) of fingerprint data strings for a biological sequence database, where each fingerprint data string represents a characteristic biological subsequence composed of sequence units, which are amino acids when the characteristic biological subsequence is related to a protein and codons when the characteristic biological subsequence is related to DNA or RNA, where each characteristic biological subsequence has in the biological sequence database a combination number less than the total number of different sequence units available to it, the combination number of the biological subsequence being defined as the number of different sequence units that appear as consecutive sequence units of the biological subsequence in the biological sequence database, the repository further comprising combination data representing the different sequence units that appear as consecutive sequence units of the corresponding characteristic biological subsequence in the biological sequence database; the method comprising: a. obtaining at least one read of the biopolymer or biopolymer fragment using a sequencer, and b. processing the read through the following computer-implemented steps: b1. searching for the occurrence of one or more of the characteristic biological subsequences represented by the fingerprint data strings in the read, and b2. validating or rejecting the read by determining at each occurrence whether the sequence units consecutive to the characteristic biological subsequence are consistent with the combination data in the repository, and / or b1'. searching for the occurrence of one of the characteristic biological subsequences represented by the fingerprint data strings in the head and / or tail of the read, and b2'. predicting one or more consecutive sequence units of the read from the combination data in the repository.
2. The method according to claim 1, wherein the repository comprises at least - a first fingerprint data string representing a first characteristic biological subsequence of a first length; and - a second fingerprint data string representing a second characteristic biological subsequence of a second length, where the first length and the second length are equal to 4 or greater than 4, and where the first length and the second length are different from each other.
3. The method according to claim 1 or 2, wherein step a comprises obtaining a plurality of reads of the biopolymer or biopolymer fragment, and wherein step b starts before all the reads of the biopolymer or biopolymer fragment are obtained.
4. The method according to claim 1 or 2, wherein step b2 and / or b2' comprises using the following: - data related to the secondary and / or tertiary and / or quaternary structure of the characteristic biological subsequence when the characteristic biological subsequence is present in a biopolymer; and / or - data related to the relationship between the characteristic biological subsequence and one or more other characteristic biological subsequences; and / or - data related to the spacing relative to the fingerprint data string; and / or - data related to the orientation of the fingerprint data string.
5. The method according to claim 1 or 2, wherein the fingerprint data string is inherently oriented and includes position information, and the method includes another step of aligning the processed reads with an orientation map using the characteristic biological subsequences identified in step b1 and / or b1'.
6. The method according to claim 5, wherein the alignment includes identifying possible sequence variations.
7. The method according to claim 1 or 2, wherein the method further includes collapsing the processed reads by sorting the processed reads.
8. The method according to claim 1 or 2, wherein the method further includes converting the obtained data into a sub-read map and / or a read map.
9. The method according to claim 1 or 2, wherein the method further includes removing any one of dead ends and / or loops.
10. The method according to claim 1 or 2, wherein the method includes dynamically adjusting the sequencing based on the information obtained from the processing and / or alignment.
11. The method according to claim 10, wherein the dynamic adjustment includes providing feedback on the number of reads that need to be obtained using the sequencing system and / or includes providing feedback on reads that will be ignored as erroneous reads based on the information obtained from the processing and / or alignment, and / or wherein the method includes backtracking towards or until a read.
12. A sequencer (350) adapted to perform steps a and b of the method according to any one of claims 1-11.
13. A system comprising: i. A sequencer (350) adapted to - perform step a of the method according to any one of claims 1-11, and - transmit reads to a data processing system; and ii. A data processing system adapted to - receive the reads from the sequencer (350), and - process the reads by performing step b of the method according to any one of claims 1 to 11.
14. A computer program or computer-readable medium comprising instructions that, when executed by a computer, cause the computer to perform the method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Methods, computer-accessible medium, and systems for score-driven whole-genome shotgun sequence assembly
US20130317755A1