Bioinformatics processing
By using the fingerprint data string repository method, identifying characteristic biological subsequences (HYFTTM fingerprints) and searching for information in its associated information repository, the problem of inefficient biological information processing in the prior art is solved, and fast, deterministic and efficient biological information analysis is achieved.
Patent Information
- Application Number
- CN202080012591.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-08-08
- Filing Date
- 2020-02-07
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2040-02-07
AI Technical Summary
Existing bioinformatics processing technologies are difficult to effectively process and analyze large amounts of biological sequence data, especially inefficient in identifying structural variants and targets, and the lack of dynamic algorithms has led to an exponential increase in the demand for computing resources.
Using the repository method of fingerprint data strings, the computational complexity and recognition efficiency are reduced by identifying characteristic biological subsequences (HYFTTM fingerprints) and searching for related information in their associated information repository.
It realizes rapid and deterministic processing of biological information, reduces sequencing errors, improves data analysis speed, can effectively link biological information fragments from different sources, and accelerates the sequencing process by predicting the next unit in the sequence.
Smart Images

Figure CN113454726B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the processing of biological information, and more particularly, to retrieving and / or correlating said biological information. Background Art
[0002] In the past few decades, biological sequencing has advanced at an astonishing pace, making the Human Genome Project possible, which achieved the complete sequencing of the human genome over 15 years ago. To drive this development, a large number of technological advancements have been required, from advancements in sample preparation and sequencing methods to data acquisition, processing, and analysis. At the same time, new scientific fields have emerged and developed, including genomics, proteomics, and bioinformatics.
[0003] Driven by the emphasis on data acquisition in the post-genomic era, this development has led to the accumulation of a large amount of biological (e.g., sequence) data. However, the ability to organize, analyze, and interpret this sequence to extract biologically relevant information has lagged behind. This problem has been further complicated by the fact that a large amount of new sequence information is still generated every day. Muir et al. observed that this has triggered a paradigm shift and commented on the resulting changes in the sequencing cost structure and other related obstacles (MUIR, Paul et al., The real cost of sequencing: scaling computation to keep pace with data generation. Genome biology, 2016, 17.1:53.).
[0004] Accessing, analyzing, or using sequence information in a meaningful way typically requires some form of sequence alignment and similarity search. A large number of computer software are commercially available to perform such alignments and sequence similarity searches, such as BLAST, PSI-BLAST; SSEARCH, FASTA, HMMER3. However, the known algorithms lack the speed or practical ability to process the large amount of existing data. Hardware optimizations have also been attempted, such as those disclosed in US2006020397A1, but have not brought about the necessary breakthroughs. At the core of this struggle is the fact that the problems being addressed are of NP-hard or NP-complete nature (NP = non-deterministic polynomial time); thus, as the task difficulty increases (e.g., as the sequence length increases or as the number of sequences to be compared increases), the resources required grow exponentially.
[0005] Structural variants play an important role in the development of cancer and other diseases and are less studied compared to single nucleotide variants, in part due to the lack of reliable identification from read data. When using k-mer technology, the detection window for variations is by definition smaller than the total length of the k-mer. Using algorithms that overcome the k-mer window problem cannot effectively identify structural variants. High coverage is required to find evidence of just one structural change. Therefore, the use of k-mers requires a large pool to effectively identify true changes from noise and read errors. Due to the lack of dynamic algorithms for aligning k-mers, many k-mers lead to computational challenges. This indicates the need for heuristics or parameterization to narrow the search space. However, the latter leads to inevitable error accumulation, which suggests that k-mers are not effective unified spatial patterns. Currently, this is only addressed in a strictly one-dimensional syntactic manner.
[0006] It is generally believed that the vast storage of biological data contains many secrets yet to be discovered, but the currently available tools do not allow the data to be combed through in a sufficiently convenient manner - for example to identify targets for treating specific pathologies. Thus, current efforts often boil down to the proverbial "searching for a needle in a haystack". Therefore, there is a great need and a search for new and unique ways that allow biological data from different sources to be correlated, thus providing new insights and revealing hidden patterns therein.
[0007] Accordingly, there is still a need in the art for further improvement in bioinformatics processing. Summary of the Invention
[0008] The object of the present invention is to provide a good method for processing biological information. This object is achieved by the method, device and data structure according to the present invention.
[0009] In a first aspect, the present invention relates to a computer-implemented method for obtaining information about a biological entity based on at least one biological sequence, comprising: (a) providing a repository of fingerprint data strings for a biological sequence database, each fingerprint data string representing a characteristic biological subsequence composed of sequence units, each characteristic biological subsequence having a combination number in the biological sequence database that is less than the total number of different sequence units available to it, the combination number of the biological subsequence being defined as the number of different sequence units that appear as consecutive sequence units of the biological subsequence in the biological sequence database; (b) determining one or more fingerprint data strings representative of the biological entity; (c) searching in a repository comprising information associated with the fingerprint data strings for information associated with the one or more representative fingerprint data strings; and (d) processing the information.
[0010] Advantages of embodiments of the present invention are that systems and methods are obtained that provide reduced complexity.
[0011] Advantages of embodiments of the present invention are that, for example, different pieces of common information from different sources can be linked together by common anchors. Another advantage of embodiments of the present invention is that common anchors can be collected in a repository of fingerprint data strings, which itself has many advantages (see below).
[0012] Advantages of embodiments of the present invention are that the same method can be used to improve the sequencing of biopolymers, and biopolymer fragments can be improved (e.g., by reducing the likelihood of errors or by speeding up the process), i.e., by relying on information contained in a repository of fingerprint data strings.
[0013] Advantages of embodiments of the present invention are that a temporarily proposed biological sequence can be verified or rejected. Advantages of embodiments of the present invention are that errors occurring during sequencing can be reduced.
[0014] Advantages of embodiments of the present invention are that the speed of sequencing can be increased by predicting the next unit in a sequence or by limiting the number of its options.
[0015] Advantages of embodiments of the present invention are that the system and method have deterministic properties, i.e., the method and system lead to a specific solution for determining the sequence for identifying / characterizing a biopolymer or a biopolymer fragment.
[0016] Advantages of embodiments of the present invention are that the system and method allow tracking of the ID of reads. The system and method allow, for example, backtracking, e.g., backtracking errors or uncertainties in reads.
[0017] Advantages of embodiments of the present invention are that, compared with at least most prior art systems, in embodiments of the present invention, fast and deterministic sequence generation can be obtained.
[0018] Advantages of embodiments of the present invention are that fast data analysis systems and methods can be developed.
[0019] In a second aspect, the present invention relates to a computer-implemented method for associating information with one or more fingerprint data strings as defined in any one of the foregoing technical solutions, comprising: (a) providing a biological sequence of a biological entity, the biological entities sharing equivalent information; (b) searching for equivalent characteristic biological subsequences in the biological sequence; and (c) associating the equivalent information with a fingerprint data string representing the equivalent characteristic biological subsequence.
[0020] Advantages of embodiments of the present invention are that links between different pieces of biological information can be found and discovered in ways not explored heretofore.
[0021] Advantages of embodiments of the present invention are that a repository of fingerprint data strings and / or a repository of processed biological sequences can be annotated with biological information.
[0022] Advantages of embodiments of the present invention are that information can be retrieved from different information sources, including public databases, proprietary databases, clinical records, and / or scientific literature. Another advantage of embodiments of the present invention is that these different information sources can be linked together through a central repository.
[0023] In a third aspect, the present invention relates to a data processing system adapted to execute a computer-implemented method according to any embodiment of the first or second aspect.
[0024] Advantages of embodiments of the present invention are that, depending on the application, steps of the method can be implemented by a variety of systems and devices, such as computer-based systems or sequencers. Another advantage of embodiments of the present invention is that the method can be implemented by computer-based systems (including cloud-based systems).
[0025] In a fourth aspect, the present invention relates to a computer program including instructions that, when executed by a computer, cause the computer to execute a method according to any embodiment of the first or second aspect.
[0026] In a fifth aspect, the present invention relates to a computer-readable medium including instructions that, when executed by a computer, cause the computer to execute a method according to any embodiment of the first or second aspect.
[0027] Specific and preferred aspects of the present invention are set forth in the appended independent and dependent claims. Features from the dependent claims can be appropriately combined with features of the independent claims and with features of other dependent claims, not only as explicitly set forth in the claims.
[0028] Although devices in the art have been continuously improved, changed, and evolved, the concept of the present invention is considered to represent substantially new and novel improvements, including improvements departing from prior practices, which result in providing a more effective, stable, and reliable device having such properties.
[0029] The above and other characteristics, features, and advantages of the present invention will become apparent from the following detailed description in conjunction with the accompanying drawings, which illustrate the principles of the present invention by way of example. This description is given for purposes of illustration only and does not limit the scope of the present invention. The reference figures cited below refer to the accompanying drawings. Brief Description of the Drawings
[0031] Figure 1 and Figure 2 are graphs showing the expected progress achieved by embodiments of the present invention.
[0032] Figures 3 to 6It is a diagram depicting a system according to an embodiment of the present invention.
[0033] Figure 7 Schematically depicts the results observed in a proof - of - concept according to the present invention.
[0034] Figure 8 Schematic overview of the processing steps that can be performed in a method for sequencing according to an embodiment of the present invention.
[0035] Figures 9 to 12 Is a schematic representation of several steps that can be used in an embodiment according to the present invention.
[0036] Figures 13 to 17 Is a graph showing various metrics regarding the analysis of a processed Protein Data Bank (PDB) according to an embodiment of the present invention.
[0037] Figure 18 Is a graph of the number of HYFT TM matches found in the PDB database plotted against each other using two different matching strategies.
[0038] Figure 19 and Figure 22 Is a graph comparing the total length of search results using prior - art methods (dashed line) on the one hand and methods according to an exemplary embodiment of the present invention (solid line) on the other hand.
[0039] Figure 20 and Figure 23 Is a graph comparing the Levenshtein distance of search results using prior - art methods (dashed line) on the one hand and methods according to an exemplary embodiment of the present invention (solid line) on the other hand.
[0040] Figure 21 and Figure 24 Is a graph comparing the longest common substring of search results using prior - art methods (dashed line) on the one hand and methods according to an exemplary embodiment of the present invention (solid line) on the other hand.
[0041] In different figures, the same reference numerals refer to the same or similar elements. Detailed Description
[0042] The present invention will be described with respect to specific embodiments and with reference to certain figures, but the present invention is not limited thereto and is only limited by the claims. The described figures are merely schematic and not restrictive. In the figures, for illustrative purposes, the size of some elements may be enlarged and not necessarily drawn to scale. The dimensions and relative dimensions do not correspond to the actual reduction in the practice of the present invention.
[0043] Furthermore, the terms first, second, third, etc. in the specification and claims are used to distinguish similar elements and not necessarily to describe a sequence in time, space, arrangement, or any other manner. It is to be understood that the terms so used are interchangeable under appropriate circumstances, and that the embodiments of the invention described herein are capable of operation in sequences other than those described or illustrated herein.
[0044] Furthermore, the terms "before...", "after...", etc. in the description and claims are used for descriptive purposes and not necessarily for describing relative positions. It is understood that the terms so used are interchangeable with their antonyms where appropriate, and that the embodiments of the invention described herein are capable of operation in other orientations than described or illustrated herein.
[0045] It should be noted that the term "comprising" used in the claims should not be interpreted as being limited to the means listed thereafter; it does not exclude other elements or steps. It should therefore be interpreted as specifying the presence of the stated features, integers, steps, or components mentioned, but not excluding the presence or addition of one or more other features, integers, steps, or components, or groups thereof. The term "comprising" therefore encompasses both the presence of only the stated features as well as the presence of these features and one or more other features. Thus, the scope of the expression "a device comprising means A and B" should not be interpreted as being limited to a device consisting solely of components A and B. This means that, for the purposes of the present invention, the only relevant components of the device are A and B.
[0046] Reference throughout this specification to "one embodiment" or "an embodiment" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present invention. Thus, appearances of the phrases "in one embodiment" or "in an embodiment" in various places throughout this specification are not necessarily all referring to the same embodiment, but may. Furthermore, the particular features, structures, or characteristics may be combined in any suitable manner in one or more embodiments, as will be apparent to one of ordinary skill in the art from this disclosure.
[0047] Similarly, it should be understood that in the description of exemplary embodiments of the invention, various features of the invention are sometimes grouped together in a single embodiment, figure, or description thereof to simplify the disclosure and aid in understanding one or more of the various inventive aspects. However, this method of disclosure should not be interpreted as reflecting an intention that the claimed invention requires more features than are expressly recited in each claim. Rather, as reflected in the following claims, aspects of the invention lie in fewer than all the features of a single preceding disclosed embodiment. Accordingly, the claims following the detailed description are hereby expressly incorporated into this detailed description, with each claim standing on its own as a separate embodiment of the invention.
[0048] In addition, as those skilled in the art will understand, although some of the embodiments described herein include some features and not other features included in other embodiments, the combination of features of different embodiments is meant to be within the scope of the present invention and forms different embodiments. For example, in the following claims, any of the claimed embodiments can be used in any combination.
[0049] In addition, some embodiments are described herein as methods or combinations of method elements that can be implemented by a processor of a computer system or by other devices that perform functions. Accordingly, a processor having the necessary instructions for executing such methods or method elements forms a means for executing the methods or method elements. In addition, the elements of the apparatus embodiments described herein are examples of means for performing the functions performed by the elements for the purpose of carrying out the present invention.
[0050] In the description provided herein, numerous specific details are set forth. However, it should be understood that embodiments of the present invention may be practiced without these specific details. In other instances, well-known methods, structures, and techniques have not been shown in detail so as not to obscure the understanding of this description.
[0051] The following terms are provided solely to assist in understanding the present invention.
[0052] As used herein, a biological sequence is a sequence of a biopolymer that at least defines the primary structure of the biopolymer. The biopolymer can be, for example, deoxyribonucleic acid (DNA), ribonucleic acid (RNA), or protein. Biopolymers are generally polymers of biological monomers (e.g., nucleotides or amino acids), but in some cases may further include one or more synthetic monomers.
[0053] As used herein, a "sequence unit" in a biological sequence is an amino acid when the biological sequence is related to a protein, and a codon when the biological sequence is related to DNA or RNA.
[0054] As used herein, a biological subsequence is a part of a biological sequence that is less than the complete biological sequence. A biological subsequence can have, for example, a total length of 100 sequence units or less, preferably 50 or less, and even more preferably 20 or less.
[0055] As used herein, a distinction is made between a "characteristic biological subsequence" (or "(HYFT TM ) fingerprint"), a "(HYFT TM ) fingerprint data string", and a "(HYFT TM ) fingerprint tag". The first is a subsequence having specific characteristics as explained in more detail below. The second is such a HYFT TMThe data representation of a fingerprint — optionally in combination with additional data (see below) — which can be stored, for example, in a corresponding repository. In some embodiments, one HYFT TM A fingerprint data string can represent multiple equivalent HYFTs simultaneously TM fingerprints (equivalent, for example, by encoding the same result, such as in the case of multiple codons encoding the same amino acid, or by translation equivalence; see below). The third is a pointer to the HYFT TM fingerprint, which can, for example, locate the HYFT TM the memory address of the fingerprint or allow finding the HYFT in the repository of fingerprint data strings TM reference of the fingerprint. However, given their close relationship, where a strict distinction between these three terms is not required, or where the meaning is clear in the context, these may simply be referred to herein as "HYFTs TM ".
[0056] As used herein, a distinction is made between a "biological sequence" and a "processed biological sequence". The former is a biological sequence well-known in the art, while the latter is a biological sequence that includes the reconstruction / rewriting of fingerprint markers associated with the HYFT TM fingerprint of the present invention.
[0057] It is clear that the HYFT TM fingerprint data string, the processed biological sequence, or the repository storing these cannot be regarded as cognitive data and is not targeted at (human) users. Instead, it is intended to be used as functional data in various computer-implemented methods by a computer (or similar technical system) and is constructed for this effect. For example, the repository can be structured like a relational database (e.g., based on SQL) or a NoSQL database (e.g., a document-oriented database, such as an XML database). Similarly, the HYFT TM fingerprint data string and / or the processed biological sequence can be constructed as suitable entries in such a database.
[0058] As used herein, some concepts will be illustrated by examples related to proteins, and it will be assumed that the possible monomer sequence units are 20 canonical (or "standard") amino acids. However, it is clear that this is merely for simplicity of illustration, and similar embodiments can equally be formulated with an extended number of amino acids (e.g., adding non-canonical amino acids or even synthetic compounds), or related to DNA or RNA. In the case of DNA or RNA, the link between DNA or RNA and proteins can be easily established through the correspondence between codons and amino acids.
[0059] As used herein, "secondary / tertiary / quaternary" means "secondary and / or tertiary and / or quaternary".
[0060] In the present invention, it has surprisingly been recognized that, in the case where the primary structure of a biological sequence was previously assumed to consist of sequence units that are essentially independently selected, such that in principle there exist, for example, m biological sequences of length n based on m possible sequence units (e.g., 20 based on 20 standard amino acids), which in fact are essentially not observed. In fact, it has been found that from a certain length onwards, not all of the theoretical combinations can be seen. Just to give one example: the protein subsequence "MCMHNQA" is not found in any protein in the public databases. It has been taken into account that this is not just an interruption in the databases, but that this absence has a physical and / or chemical origin. Without being bound by theory, just to list one possible effect, the steric hindrance of adjacent amino acids (e.g., "MCMHNQ" in the above example) can prevent one or more other amino acids (e.g., "A" in the above example) from binding to it. Thus, once a missing subsequence has been identified, computational studies can be used to verify whether this subsequence is possible or whether its existence is physically impossible (or unlikely, e.g., because it is chemically unstable). The above-mentioned "certain length" depends on the data set considered, but for example corresponds to about 5 or 6 amino acids of the publicly available protein sequence databases (which essentially reflects the total diversity that is essentially visible). For more restricted sets (e.g., sets filtered based on specific criteria or for a specific biological sequence database, e.g., a set developed for a specific domain), for a length of about 4 or 5, less than the theoretical maximum of m combinations has been found. n n ) n combinations.
[0061] At the same time, since the subsequence "MCMHNQA" does not exist, the subsequence "MCMHNQ" not only becomes more than a random combination of 5 amino acids, but also acquires an additional meaning; such subsequences will further be referred to as "characteristic biological subsequences" or "(HYFT TM ) fingerprints". Due to the additional meaning or significance of these HYFT TM fingerprints, the present invention can be considered to handle biological sequence information in a more semantic way. Generally speaking, characteristic subsequences are characterized in that for the sequence units directly following (or preceding) them, there are fewer possible options (i.e., a lower number of combinations) than the maximum number of sequence units (i.e., the total number of different sequence units available for it; e.g., less than 20 standard amino acids); in other words, at least one of the sequence units cannot follow (or precede) it. However, a more restrictive definition can be chosen: for example, only those subsequences with 15 or fewer sequence units may follow it, or 10 or fewer, 5 or fewer, 3, 2 or even 1. In addition, each such subsequence can be optionally considered as a HYFT TMfingerprints, or only consider those subsequences that do not yet contain another HYFT TM fingerprint as a HYFT TM fingerprint (i.e., non-redundant). For example: taking "MCMHNQ" as a HYFT TM fingerprint, there will be longer subsequences that include "MCMHNQ" and also have fewer than the theoretical number of sequence units that can follow (or precede) it; in this case, it is possible to choose to consider both the longer subsequence and "MCMHNQ" as HYFT TM fingerprints, or only consider "MCMHNQ" as a HYFT TM fingerprint. The latter method is usually likely to be preferred in order to control the size of the repository of HYFT TM data strings, while accelerating the methods associated with it. In fact, as the string length increases, searching for a match to the string in a biological sequence generally becomes more resource-intensive and slower. Additionally, as the size of the repository of HYFT TM data strings increases, searching for and retrieving a specific HYFT TM data string generally takes longer. In this non-redundant method, longer subsequences with limited combinatorial possibilities can still be identified, but then are identified as HYFTs TM patterns (with or without spacing). Thus, the advantages provided by this method do not necessarily result in a corresponding loss of information. Nevertheless, note that the former method is still possible and doing so is still superior to the prior art.
[0062] Subsequently, it was surprisingly found that a finite set of characteristic biological subsequences can be identified. Additionally, it was observed that these characteristic biological subsequences achieve a balance between being specific enough such that not every characteristic biological subsequence is found in every biological sequence, and being common enough such that known biological subsequences generally include at least one of these HYFT TM fingerprints.
[0063] From the narrative provided above, a protocol can be formulated for identifying HYFT TM fingerprints and constructing the corresponding repository of HYFT TM data strings (or "HYFT TM repository"). In fact, since the aim is to identify those subsequences with limited combinatorial possibilities in a biological sequence database, it is sufficient to mine for subsequences that do not appear in the biological sequence database. Once such non-appearing subsequences are identified (e.g., "MCMHNQA"), the subsequence that is one sequence unit shorter (e.g., "MCMHNQ") corresponds to a HYFT TM fingerprint (provided that the shorter subsequence does indeed appear). Once identified, information about the HYFT TMAdditional data of the fingerprint. For example, the identified HYFT can be recognized by searching in a biological sequence database TM The number of combinations can be obtained by combining the fingerprint with other sequence units (e.g., replacing "A" in "MCMHNQA" with one of the other possible amino acids each time) and counting the number of combinations found. Optionally, the non-found combinations can also be stored separately; these combinations can be used, for example, for error detection. In addition, since the correspondence between DNA, RNA, and proteins is usually known through an applicable codon table, once a specific type of HYFT TM fingerprint (e.g., protein HYFT TM ) is identified, it can be translated into corresponding HYFTs TM of different types (e.g., DNA and / or RNA HYFT TM ). By repeating the above process and storing at least the identified HYFTs TM —optionally together with any additional data and translated HYFTs TM —a repository of HYFT TM fingerprint data strings can be constructed. As an alternative or supplement, at least some HYFT TM fingerprints can be discovered by experimental or computational methods, for example, by synthesizing or simulating various subsequences and then identifying those subsequences that cannot—or are less likely to—occur in the context of the biological sequence database under consideration.
[0064] In the above, the biological sequence database can be a publicly available database, such as the Protein Data Bank (PDB), or a proprietary database. In an embodiment, the biological sequence database can be a combination of multiple individual databases. For example, a repository of HYFT TM fingerprint data strings can be formulated from a biological sequence database that combines as many (reliable) biological sequence databases as possible that are accessible, thereby seeking a general repository of HYFT TM fingerprint data strings that essentially represent all biological sequences that have been found in nature. Conversely, in a specific domain, it can be shown that constructing a specific repository of HYFT TM fingerprint data strings based on a biological sequence database representing the specific domain is productive. In an embodiment, this specific repository can contain HYFTs TM that do not exist in the general repository because they do indeed exist in nature but are not within this specific domain. Similarly, a repository of HYFT TM fingerprint data strings can be constructed for synthetic sequences with their own specific content.
[0065] Based on the above findings, new methods for processing biological sequence information can be developed at all different but interrelated stages. These methods can be considered analogous to performing more lexical analysis on the sequences. The results are schematically depicted in Figure 1 which shows the complexity scaling of biological sequence information as the number (n) of sequence units increases. This complexity can be the total number of possible combinations of sequence units, but it is also related to the computational effort (such as time and memory) required to process it (e.g., perform a similarity search). The solid curve depicts the number of theoretical combinations assuming that all sequence units are independently selected, scaled by m n which also corresponds to the scaling of currently known algorithms. The dashed curve depicts the number of actual combinations essentially found (as observed in the present invention), where the curve deviates from m at around 5 or 6 sequence units n and flattens out asymptotically at high n. The dashed line shows the number of sequences corresponding to the first characteristic sequence for which the number of sequence units that can follow it equals 1; the "first" here means that if a longer sequence contains a HYFT TM fingerprint that has already been counted, it is never counted again. Thus, when its definition is chosen to have only 1 possible sequence unit following it and does not yet include another (shorter) HYFT TM fingerprint subsequence, the latter corresponds to the number of HYFT TM fingerprints of length n (as observed in the present invention) (see above).
[0066] Figure 2 depicts the predicted benefits when using a repository of fingerprint data strings as described herein, where the markings on the bottom axis depict the present. Curve 1 shows Moore's Law as a reference. Curve 2 shows the total amount of sequencing data collected. Curve 3 shows the total cost of processing and maintaining the sequencing data. By processing biological sequence information as described herein, it is expected that the total storage required for sequencing data and the total cost of data processing and maintenance will decrease, as depicted in curves 4 and 5 respectively.
[0067] Note that although a repository of HYFT TM fingerprint data strings is typically constructed for a specific biological sequence database (or a combination thereof), this does not mean that HYFT TM fingerprint data strings are only applicable to processing biological sequences in the specific biological sequence database. In fact, a general repository of HYFT TM fingerprint data strings can be used, for example, to process more specific biological sequences. In other cases, HYFT TMA specific repository of fingerprint data strings can be used in the context of biological sequences that fall outside the database used to formulate the repository. In both cases, favorable results can still be obtained. In any case, one can always determine the existing HYFT through trial and error TM whether a repository of fingerprint data strings can be used for a specific application, or use a repository dedicated to HYFT TM to see if better results can be obtained. Similarly, HYFT TM a repository of fingerprint data strings does not strictly need to cover all HYFT TM fingerprints that can be found in biological sequence databases. In fact, partial repositories have already produced beneficial results. This partial repository can be, for example, a repository related to HYFT TM fingerprints of a selected length (i.e., as opposed to HYFT TM fingerprints of any length).
[0068] The present invention utilizes a repository of fingerprint data strings. Thus, a repository of fingerprint data strings for a biological sequence database is described, each fingerprint data string representing a characteristic biological subsequence composed of sequence units, each characteristic biological subsequence having a combination number less than the total number of different sequence units available in the biological sequence database, the combination number of the biological subsequence being defined as the number of different sequence units that appear as consecutive sequence units of the biological subsequence in the biological sequence database. In Figure 4 FIG. 100 schematically depicts a repository (e.g., a database) of fingerprint data strings, which will be discussed in more detail below.
[0069] An advantage of an embodiment of the present invention is that a repository of fingerprint data strings corresponding to characteristic biological subsequences can be provided. Another advantage of an embodiment of the present invention is that the biological subsequences do not have to be of a single length, such as in the case of k-mers.
[0070] An advantage of an embodiment of the present invention is that other data, such as metadata, can be included in the repository, such as data about sequence units that can be contiguous with the characteristic biological subsequence (i.e., directly after or directly before the characteristic biological subsequence), data about the secondary / tertiary / quaternary structure of the characteristic biological subsequence (e.g., when the characteristic biological subsequence is present in a biopolymer), data about the relationships between fingerprints (e.g., data related to the relationship between the characteristic biological subsequence and one or more other characteristic biological subsequences), etc.
[0071] In an embodiment, the repository may include at least a first fingerprint data string representing a first characteristic biological subsequence of a first length and a second fingerprint data string representing a second characteristic biological subsequence of a second length, wherein the first length and the second length are equal to 4 or greater, and wherein the first length and the second length are different from each other.
[0072] In an embodiment, the length may correspond to the number of sequence units. In an embodiment, the length may be up to 1000 or less, such as up to 100 or less, preferably 50 or less, and still more preferably 20 or less. In an embodiment, the first length and the second length may be equal to or greater than 5, preferably equal to or greater than 6. In an embodiment, the length of the characteristic biological subsequence may be between 4 and 20, preferably between 5 and 15, and still more preferably between 6 and 12.
[0073] In an embodiment, the repository of fingerprint data strings may include at least 3 fingerprint data strings of different lengths from each other, preferably at least 4, still more preferably at least 5, and most preferably at least 6. Since the characteristic biological subsequence is not defined by its length, but by the number of possible sequence units following (or preceding) it, the set of characteristic biological subsequences typically advantageously includes subsequences of different lengths. The repository of fingerprint data strings in the present invention differs from, for example, a set of k-mers (as known in the art) in that it includes biological subsequences of different lengths. In addition, a set of k-mers typically includes every permutation of a fixed length k (i.e., every possible combination of sequence units); this is not the case for the current repository of fingerprint data strings.
[0074] In an embodiment, the fingerprint data string can be a protein fingerprint data string, a DNA fingerprint data string, an RNA fingerprint data string, or a combination thereof. In an embodiment, the characteristic biological subsequence can be a characteristic protein subsequence, a characteristic DNA subsequence, or a characteristic RNA subsequence. In an embodiment, the repository of fingerprint data strings can include protein fingerprint data strings, DNA fingerprint data strings, RNA fingerprint data strings, or a combination of one or more of these (e.g., consisting of protein fingerprint data strings, DNA fingerprint data strings, RNA fingerprint data strings, or a combination of one or more of these). In an embodiment, the characteristic protein subsequence can be translated into a characteristic DNA or RNA subsequence, and vice versa. Such translation can be based on well-known DNA and RNA codon tables. Similarly, a protein fingerprint data string can be translated into a DNA or RNA fingerprint data string. In an embodiment, the repository of DNA or RNA fingerprint data strings can include information about equivalent codons (i.e., codons that encode the same amino acid). Such information about equivalent codons can be contained in the fingerprint data string as well, or stored separately from it in the repository. In a particular embodiment, the fingerprint data string can be in a sequence-independent format; this means that the fingerprint data string and the surrounding systems and processes enable it to be quickly compared with DNA, RNA, and protein sequences. This can be achieved, for example, by having the method using the fingerprint data string perform the necessary translation on the fly. Such fingerprint data strings advantageously allow for the creation of a single repository that is generally applicable to data strings across sequence types.
[0075] In an embodiment, the repository of fingerprint data strings can further include additional data for at least one of the fingerprint data strings. In a preferred embodiment, the data can be contained in the fingerprint data string. In an alternative embodiment, the data can be stored separately from the fingerprint data string. In an embodiment, the additional data can include one or more of combinatorial data, structural data, relational data, positional data, and orientation data.
[0076] In an embodiment, the combinatorial data can be data related to one or more sequence units that can be contiguous with the characteristic biological subsequence when the characteristic biological subsequence is present in a biological sequence (e.g., can truly directly appear before or after it, such as those stable combinations). In an embodiment, the combinatorial data can include the number of possible sequence units, the possible sequence units themselves, the likelihood (e.g., probability) of each sequence unit, etc.
[0077] In an embodiment, the structural data can be structural information and / or spatial shape information embedded in the fingerprint data string, such as data related to the secondary / tertiary / quaternary structure of a characteristic biological subsequence when the characteristic biological subsequence is present in a biopolymer. In an embodiment, the structural data can include the number of possible structures, the possible structures themselves, the likelihood (e.g., probability) of each structure, etc. In the case of multiple possible secondary / tertiary / quaternary structures for a given characteristic biological subsequence, in an embodiment, the repository can include a separate entry for each combination of the characteristic biological subsequence and the associated secondary / tertiary / tertiary structure. In an alternative embodiment, the repository can include one entry that includes the characteristic biological subsequence and multiple associated secondary / tertiary / quaternary structures. In an embodiment, the secondary / tertiary / quaternary structure may be more relevant for protein alignment with DNA and RNA - especially the quaternary structure.
[0078] In an embodiment, the relationship data is data related to the relationship between a characteristic biological subsequence and one or more additional characteristic biological subsequences. In an embodiment, the relationship data can include additional characteristic biological subsequences that typically occur in its vicinity, the likelihood of the additional characteristic biological subsequences occurring in its vicinity, the specific significance (e.g., biological relevance, such as a trait or secondary / tertiary / quaternary structure) of these characteristic biological subsequences occurring close to each other, etc. In an embodiment, the relationship can be expressed in the form of a path between two or more characteristic biological subsequences. In an embodiment, the relationship can include the order and / or the spacing of the characteristic biological subsequences. In an embodiment, the additional data can also include metadata for constructing the path.
[0079] In an embodiment, the position data can be data related to the spacing relative to the fingerprint data string (e.g., between the characteristic biological sequences it represents).
[0080] In an embodiment, the orientation data can be data related to the orientation (e.g., the inherent orientation) of the fingerprint data string (e.g., the characteristic biological sequence it represents).
[0081] In some embodiments, the additional data may have been retrieved from a known dataset; for example, the secondary / tertiary / quaternary structures of several biological sequences are available in the art. In other embodiments, the additional data can be extracted from the processed biological sequences described below or from a repository of the processed biological sequences described below. For example, after processing the biological sequences as described below (or constructing a repository of the processed biological sequences as described below), the relationships between characteristic biological subsequences (e.g., paths) can be extracted and added to the repository of the fingerprint data string; this is in Figure 4is schematically depicted by the dashed arrows from the processed biological sequences 210 and the repository 220 of processed biological sequences to the repository 100 of fingerprint data strings.
[0082] In an embodiment, the fingerprint data string may be inherently oriented. In an embodiment, the fingerprint data string may include an orientation (i.e., may explicitly include an orientation). Due to HYFT TM fingerprints are defined based on the actual segments present in a biopolymer or biopolymer fragment, there are inherently present in HYFTs the inherent physical, chemical, and structural limitations that occur for the combinatorial possibilities present in a biopolymer TM ; where "inherently present" is understood such that such information is (or at least can be) implicitly associated with the HYFT TM even if it is not explicitly included as additional data in the repository. Thus, since biological sequences themselves typically have an inherent directionality (i.e., according to the 5' to 3' direction in DNA / RNA and the N-terminus to C-terminus in proteins), this same directionality inherently exists in HYFTs TM . This link to the actual segments further defines the limit on the maximum number of biopolymer segments that can follow after the last character or precede the first character in a HYFT TM . The latter can also be explicitly expressed by a parameter (i.e., the number of combinations) representing the total amount of possible subsequent or prior combinations. This also results in HYFTs TM having an inherent (strict) orientation.
[0083] In an embodiment, the fingerprint data string may include position information. The characters in HYFTs TM and the characters between HYFTs TM are syntactically related to each other and thus can define the spacing between them or between different HYFTs TM . Such positions or spacings belong to the position information that can inherently exist in HYFTs TM .
[0084] In an embodiment, the fingerprint data string may further include structural and / or spatial shape information. The possible structures and / or spatial shapes of certain HYFTs TM or combinations of HYFTs TM are also limited due to the inherent physical, chemical, and structural limitations. Such information also inherently exists in HYFTs TM or related sets of HYFTs TM .
[0085] In a first aspect, the present invention relates to a computer-implemented method for obtaining information about a biological entity based on at least one biological sequence, comprising: (a) providing a repository of fingerprint data strings for a biological sequence database, each fingerprint data string representing a characteristic biological subsequence composed of sequence units, each characteristic biological subsequence having a combination number less than the total number of different sequence units available in the biological sequence database, the combination number of the biological subsequence being defined as the number of different sequence units that appear as consecutive sequence units of the biological subsequence in the biological sequence database; (b) determining one or more fingerprint data strings representative of the biological entity; (c) searching a repository comprising information associated with the fingerprint data strings for information associated with the one or more representative fingerprint data strings; and (d) processing the information.
[0086] It has surprisingly been found in the present invention that HYFTs TM can also be effectively used to associate different pieces of biological information (e.g., from different sources), where HYFTs TM serve as an anchor between the different pieces. Thus, the practical method of obtaining biological information becomes to select an entry point (e.g., entered by the user), determine the HYFT(s) TM representing the entry point (e.g., automatically determined by a computer), and then retrieve the information associated with the representative HYFT(s) TM . Here, the retrieved information can be substantially all the information associated with the HYFT(s) TM contained in the repository, or it can be a selection thereof (e.g., filtered based on the type of information the user has entered that he / she wants to retrieve). Optionally, the processing in step d can go beyond simple retrieval, as will be outlined below.
[0087] In an embodiment, the repository comprising information associated with the fingerprint data strings can be the repository of fingerprint data strings as described herein, the repository of processed biological sequences as described herein, or any other repository containing information associated with the fingerprint data strings.
[0088] In an embodiment, one or more of the fingerprint data strings representative of the biological entity include a fingerprint data string representing the longest characteristic biological subsequence found in at least one biological sequence, or - if more than one longest characteristic biological subsequence is found - include the characteristic biological subsequence among the longest characteristic biological subsequences having the lowest combination number. It has surprisingly been found that the representative HYFT(s) TM often contain the most stringent HYFT TM (i.e., the longest HYFT having the lowest combination number present in at least one biological sequence TM ). Furthermore, in many cases, the most stringent HYFT can be selected only byTM Acting as a representative and performing a search based thereon to obtain very useful information.
[0089] In an embodiment, a biological entity based on at least one biological sequence can be anything from a biological subsequence (e.g., a sequencing read) to an organism or species, as long as a representative HYFT TM can be associated with the entity; including but not limited to proteins, protein active sites or domains, genes, genomes, membranes, organelles, cells, bacteria, viruses, organs, etc.
[0090] In an embodiment, the information can include one or more of a medical condition, a biological function, a spatial structure, combinatorial information, or related information. Various different "exit points" (i.e., the categories of information to be obtained) can be advantageously accessed.
[0091] In an embodiment, processing the information can include retrieving the information itself or using the information to achieve different effects. For example, the information can be used to improve the processing of sequencing reads as described below. Other examples include target or biomarker identification, correlating changes (e.g., structural changes and / or indel mutations) across multiple genes and / or proteins with a specific outcome (e.g., a disease mechanism), etc.
[0092] In a particular set of embodiments, the method can be used to process sequencing reads of a biopolymer or a biopolymer fragment in consideration of the information contained in a repository of fingerprint data strings. In an embodiment, the information associated with the fingerprint data strings included in the repository can contain combinatorial data representing different sequence units that appear as consecutive sequence units of a corresponding characteristic biological subsequence in a biological sequence database. In some embodiments, step b can include searching for the occurrence of one or more of the characteristic biological subsequences represented by the fingerprint data strings in the read, and step d can include validating or rejecting the read by determining whether the sequence units consecutive to the characteristic biological subsequence conform to the combinatorial data in the repository. In alternative or additional embodiments, step b can include searching for the occurrence of one of the characteristic biological subsequences represented by the fingerprint data strings at the head and / or tail of the read, and step d includes predicting one or more consecutive sequence units of the read from the combinatorial data in the repository. Figure 3 Sequencing system 350 is schematically shown, which uses the information contained in repository 100 of fingerprint data strings to sequence biopolymer (fragment) 500.
[0093] In an embodiment, the read can be an initial (e.g., provisional or partial) biological sequence. In an embodiment, the method can be performed on a batch of reads. In an embodiment, the reads may have been obtained using a sequencer (e.g., a sequencing system). In an embodiment, the method can be started after obtaining the batch of reads.
[0094] In an embodiment, the search in step b can be as described for step b of the method for processing biological sequences described below.
[0095] Relative to the first type of embodiment, since the repository contains combined data about sequence units that occur after (e.g., before or after) the HYFT TM fingerprint, this information can be advantageously used to verify whether the reads are consistent with it. If not, the provisional biological sequence can be rejected and redone. Alternatively, the same purpose can be achieved by directly matching it with an undiscovered biological sequence rather than matching the reads with the HYFT TM fingerprint itself (see above). Alternatively, such consistency verification can be combined with the use of additional data such as, for example, structural data, relational data, positional data, and / or orientation data (see above). Such combinations can, for example, allow the rejection of reads that are indeed consistent with a known HYFT TM fingerprint but are not in the context set by the additional data.
[0096] Relative to the second type of embodiment, based on the same combined data, it is known that some HYFT TM fingerprints (or combinations of HYFT TM fingerprints) have very limited combination possibilities (i.e., correspond to a low number of combinations). For example, in the case of a HYFT TM fingerprint with a combination number of 1, the next sequence unit is known. This information can be advantageously used to accelerate sequencing by directly appending the sequence unit to the reads; thus allowing the actual sequencing to skip the sequence unit. In an embodiment, the repository can contain data about a series of two, three, or more sequence units that together are the only possible options that occur after a particular HYFT TM fingerprint. In this case, the entire series can be advantageously appended directly to the reads; thus allowing the actual sequencing to skip these units. Similarly, if the repository indicates that for an observed HYFT TM fingerprint, a limited number (but more than 1) of options are available as other sequence units (e.g., two or three options), this information can still allow the sequencer to more quickly identify the specific sequence unit in this instance. Additionally, for such HYFT TM fingerprints with a low number of combinations, by combining the combined data with the use of additional data, the number of possibilities in the current case can be reduced to 1 (or at least the possibilities can exceed a predetermined threshold). Similarly, this combination can set a context that allows the rejection of some combination possibilities, thus, for example, reducing the remaining number to 1 and thereby revealing the subsequent sequence unit.
[0097] In an embodiment, the information associated with the fingerprint data string included in the repository may further include one or more of structural data, relational data, spatial data, and orientation data; as described above with respect to the repository of fingerprint data strings. In an embodiment, the information processed in step d may include one or more of them.
[0098] In an embodiment, the method (e.g., steps b to d) may include parsing the reads (e.g., using the information of the repository of fingerprint data strings); for example, according to the method of processing biological sequences described below. In an embodiment, the method (e.g., steps b to d) may include parsing the batch of reads (e.g., after obtaining the batch of reads).
[0099] In an embodiment, the method may include another step of aligning (e.g., matching) the processed reads (e.g., included in step d); for example, by aligning and / or assembling according to the method for comparing biological sequences described below. In an embodiment, the alignment may include using the characteristic biological subsequences identified in step b. In an embodiment, the fingerprint data string may be inherently oriented and may include position information. In an embodiment, the alignment may include aligning the processed reads with an orientation graph. In an embodiment, the method may include aligning the batch of processed reads after the batch of reads has been obtained. In at least some embodiments, the alignment may be aligning the processed reads with an oriented acyclic graph.
[0100] In some embodiments, Navarro-Levenshtein matching may be used to perform the alignment. A more detailed description of Navarro-Levenshtein matching can be found, for example, in Navarro, Theoretical Computer Science 237(2000)455-463. Based on the results in one or more of the data processing steps described above, feedback information regarding identifying one or more reads as errors may be generated and these... may be ignored in further data processing.
[0101] In an embodiment, the method (e.g., alignment) may further include identifying variations; such as indel mutations, deletions, insertions, and / or duplications.
[0102] In an embodiment, the method may further include folding the processed reads by sorting them. It should be noted that the folding step in the embodiments of the present invention is not based on dynamic programming. Each HYFT TM has a specific number of bits and can be reduced / optimized by Shannon entropy. HYFTs TM and additional reads may be sorted or classified according to the amount of information (bits) they possess. Since this is for each HYFT TMThey are not equal because the next combination number can reach n - 1, so there will be HYFTs TM and corresponding read segment patterns with very few bits and HYFTs TM and read segment patterns that require a higher number of bits. Thus, in the sorting mechanism, the ready global bit threshold can be made to optimize the amount of bits used at each moment during the calculation process. And at most, fully maximize the hardware that must be used through parallelization in order to perform these given tasks. In this way, parallelization can be performed, which results in acceleration and true optimization. In some embodiments, sorting can be performed based on length. In an embodiment, classification can be performed based on the position of HYFT TM in the read segment.
[0103] In an embodiment, the method may further include converting a plurality of processed read segments into a sub-read segment graph and / or a read segment graph.
[0104] In an embodiment, the method may further include removing dead ends and / or loops.
[0105] In an embodiment, the method may include obtaining feedback on read segments that will be ignored as incorrect read segments based on information obtained from the processing and / or alignment and / or other processing of the read segments.
[0106] In an embodiment, the method may include backtracking towards or until the read segment. In an embodiment, the method may further include capturing metadata, such as a read segment ID and maintaining the read segment ID throughout the process. This can advantageously facilitate backtracking, such as backtracking the error or uncertainty of the read segment.
[0107] According to an embodiment of the present invention, the construction of the subgraph and the corresponding processing can be performed in separate threads. This can be further facilitated, for example, by an automatic completion function that can be inherently introduced in an embodiment of the present invention. If a certain confidence threshold (equivalent to sufficient coverage) is reached in the graph or subgraph construction, then no other read segment information is required to complete the original string reconstruction.
[0108] According to an embodiment of the present invention, the method may include the step of generating feedback information on read segments that will be ignored.
[0109] In a second aspect, the present invention relates to a computer-implemented method for associating information with one or more fingerprint data strings as defined in any one of the foregoing technical solutions, which includes: (a) providing a biological sequence of a biological entity, where the biological entities share equivalent information; (b) searching for equivalent characteristic biological subsequences in the biological sequence; and (c) associating the equivalent information with a fingerprint data string representing the equivalent characteristic biological subsequence.
[0110] In an embodiment, equivalent information can be information shared among biological entities, provided that the information is appropriately transformed (e.g., translated, transcribed, transposed, etc.) when needed. For example, a DNA string and a protein string (or a DNA string and an RNA string) can have sequences that are equivalent through translation (or transcription). Similarly, two species can have common traits, provided that the necessary changes are made when making the comparison. In an embodiment, equivalent information can include one or more of a medical condition, a biological function, a spatial structure, or combinatorial information.
[0111] In an embodiment, the method can further include another step a' of searching a data pool for biological entities sharing equivalent information before step a. In an embodiment, the data pool can include sequencing data or a biological sequence database. The sequencing data can include, for example, reads and / or assembled sequences (e.g., obtained by aligning reads). In an embodiment, the data pool can be from a public database, a proprietary database, clinical records, and / or scientific literature. In an embodiment, step a' can include using machine learning to identify equivalent information.
[0112] In an embodiment, step c can include annotating a repository of fingerprint data strings as described herein, a repository of processed biological sequences as described herein, or any other repository with the equivalent information.
[0113] In a third aspect, the invention relates to a data processing system adapted to execute a computer-implemented method according to any embodiment of the first or second aspect.
[0114] Such a system can include, for example, a data processing device or a processor for processing a batch of reads of a biopolymer or a biopolymer fragment.
[0115] In an embodiment, the data processing system can include a distributed computing environment (e.g., a cloud-based system) or can be—or can be part of—a distributed computing environment (e.g., a cloud-based system). The distributed computing environment can include, for example, server devices (e.g., data processing systems) and networked client devices. Here, the server device can perform most of one of the described methods. On the other hand, the networked client devices can communicate instructions (e.g., inputs, such as queries, and / or settings, such as search preferences) to the server device and can receive method outputs. In an embodiment, the data processing system can be located on-site (e.g., in the same building) or off-site (e.g., in the cloud) relative to the client devices.
[0116] In a fourth aspect, the invention relates to a computer program including instructions that, when executed by a computer, cause the computer to execute a method according to any embodiment of the first or second aspect.
[0117] In a fifth aspect, the present invention relates to a computer-readable medium comprising instructions which, when executed by a computer, cause the computer to perform the method according to any embodiment of the first or second aspect.
[0118] A computer-implemented method for constructing and / or updating a repository of fingerprint data strings as described above is also described, which comprises: (a) identifying characteristic biological subsequences in a biological sequence database, the characteristic biological subsequences having a number of combinations less than the total number of different sequence units available to them, the number of combinations of a biological subsequence being defined as the number of different sequence units that occur as consecutive sequence units of the biological subsequence in the biological sequence database; (b) optionally, translating the identified characteristic biological subsequences into one or more additional characteristic biological subsequences; and (c) populating the repository with one or more fingerprint data strings representing the identified characteristic biological subsequences and / or the one or more additional characteristic biological subsequences.
[0119] A computer-implemented method for processing biological sequences is also described, which comprises: (a) retrieving one or more fingerprint data strings from a repository of fingerprint data strings as described above, (b) searching for occurrences of characteristic biological subsequences represented by the one or more fingerprint data strings in a biological sequence, and (c) constructing a processed biological sequence comprising, for each occurrence in step b, a fingerprint marker associated with the fingerprint data string representing the occurring characteristic biological subsequence. Figure 4 Schematically shows a sequence processing unit 310 which processes a biological sequence 200 using a repository 100 of fingerprint data strings to obtain a processed biological sequence 210.
[0120] Advantages of embodiments of the present invention are that biological sequences can be processed relatively easily and efficiently. Another advantage of embodiments of the present invention is that biological sequences can be analyzed in a lexical or even semantic way.
[0121] An advantage of embodiments of the present invention is that a processed biological sequence can be constructed by replacing the identified characteristic biological subsequences therein with markers associated with the corresponding fingerprint data strings.
[0122] Advantages of embodiments of the present invention are that parts of a biological sequence that do not correspond to one of the characteristic biological subsequences can be processed in various ways. Another advantage of some embodiments is that biological sequences can be processed in a completely lossless way (i.e., no information is lost due to processing). Another advantage of alternative embodiments of the present invention is that biological sequences can be processed in a way that extracts more important information in a more compressed format.
[0123] An advantage of embodiments of the present invention is that the processed biological sequences can be compressed so that they occupy less storage space than their unprocessed counterparts.
[0124] An advantage of the embodiments of the present invention is that the matching of a portion of a biological sequence with a characteristic biological subsequence is not limited to the primary structure, but can also take into account secondary / tertiary / quaternary structures.
[0125] An advantage of the embodiments of the present invention is that the secondary / tertiary / quaternary structure of a biological subsequence can be at least partially elucidated based on the known secondary / tertiary / quaternary structure of the characteristic biological subsequence contained therein. Another advantage of the embodiments of the present invention is that it can assist or facilitate the design of biological sequences (such as proteins).
[0126] In an embodiment, the biological sequence to be processed can be the biological sequence of a fragment of a biopolymer that can be obtained by a method for sequencing according to the first aspect.
[0127] In some embodiments, the label can be a reference string. Such a reference string can, for example, point to a corresponding fingerprint data string in a repository. In other embodiments, the label can be the fingerprint data string itself or a part thereof.
[0128] In an embodiment, the biological sequence can include: (i) one or more first parts, each first part corresponding to one of the characteristic biological subsequences represented by one or more fingerprint data strings, and (ii) one or more second parts, each second part not corresponding to any of the characteristic biological subsequences represented by one or more fingerprint data strings. In an embodiment, constructing the processed biological sequence in step c can include replacing at least one first part with a corresponding label. In an embodiment, constructing the processed biological sequence in step c can further include adding position information about the first part to the processed biological sequence (such as appending it to the label). In an embodiment, constructing the processed biological sequence in step c can include leaving at least one second part unchanged, and / or replacing at least one second part with an indication of the length of the second part, and / or completely removing at least one second part. When leaving the second part unchanged, it is advantageously possible to process the biological sequence in a completely lossless manner.
[0129] In an embodiment, the processed biological sequence can be formulated in a compressed format. For example, by replacing the characteristic biological subsequence (i.e., the first part) with a reference string and / or by replacing the second part with an indication of its length or completely removing the second part, a processed biological sequence is obtained that requires less storage space than the original (i.e., unprocessed) biological sequence. Additional data compression can be achieved by exploiting paths that can represent multiple fingerprints by their interrelationships.
[0130] In an embodiment, one or more fingerprint data strings can be in a biological format different from a biological sequence (e.g., protein vs. DNA vs. RNA sequence information), and step b can further include translating or transcribing the characteristic biological subsequence prior to the search.
[0131] In an embodiment, the search in step b can include searching for partial matches or equivalent matches (e.g., equivalent codons, or different amino acids that yield the same secondary / tertiary / quaternary structure). In an embodiment, the search in step b can take into account the secondary / tertiary / quaternary structure of the characteristic biological subsequence. Secondary, tertiary, and quaternary structures are generally more evolutionarily conserved and often undergo changes in the primary structure that do not alter the function of the biopolymer, e.g., because the secondary / tertiary / quaternary structure of its active site is substantially conserved. Thus, the secondary / tertiary / quaternary structure can reveal relevant information about the biopolymer that would be lost when strictly searching for an exact match in the primary structure.
[0132] In a preferred embodiment, the occurrences of the characteristic biological subsequence are searched for in step b in a particular order. In an embodiment, the order can be based on the length and the number of combinations of the characteristic biological subsequence. In an embodiment, the search can be performed in an order starting with the longest characteristic biological subsequence having the lowest number of combinations and ending with the shortest characteristic biological subsequence having the highest number of combinations. In a preferred embodiment, the order can be from the longest characteristic biological subsequence to the shortest characteristic biological subsequence and, for characteristic biological subsequences of the same length, from the lowest number of combinations to the highest number of combinations. In other embodiments, the order can be from the lowest number of combinations to the highest number of combinations and, for characteristic biological subsequences having the same number of combinations, from the longest characteristic biological subsequence to the shortest characteristic biological subsequence. In an embodiment, the order can further take into account additional data (e.g., to determine the order within a set of characteristic biological subsequences having the same length and the same number of combinations), such as context data.
[0133] In an embodiment, the method may include another step d, after step c, of at least partially inferring the secondary / tertiary / quaternary structure of the processed biological subsequence based on the structural data as described above. This at least partial elucidation of the secondary / tertiary / quaternary structure may assist and / or facilitate biological sequence design. In embodiments where the single primary structure of the characteristic biological subsequence is linked to multiple secondary or tertiary or quaternary structures, the secondary / tertiary / quaternary structure may be disambiguated based on the context in which the characteristic biological subsequence is found, e.g., the characteristic biological subsequence it surrounds. For example, the information required for such disambiguation may be found in a (annotated) repository of fingerprint data strings. As described above, this may be in the form of data (e.g., relationship data) related to the relationship in terms of secondary / tertiary / quaternary structure between the characteristic biological subsequence and one or more additional characteristic biological subsequences. For example, it may be known that a particular first HYFT TM fingerprint adopts a helix or turn configuration as its secondary structure, but when a particular second HYFT TM fingerprint is present within a certain spacing from the first HYFT TM it always adopts a helix configuration. In this case, the HYFT TM pattern of the HYFT TM fingerprint - if observed - may be used to disambiguate the secondary structure of the first HYFT TM . Similarly, the information used may be any type of data as described above; or any other information (e.g., medical data) that may further disambiguate - alone or in combination.
[0134] In an embodiment where the fingerprint data string is inherently directed and includes position information, step c may include constructing the processed biological sequence as a directed graph. In an embodiment, the directed graph may be a directed acyclic graph. It should be noted that when referring to an acyclic graph, this does not mean that cycles cannot occur, but rather that the entire graph is not cyclic. The resulting graphical representation of the reconstructed sequence as obtained in an embodiment of the present invention may be referred to as a HYFT TM graph. This HYFT TM graph may allow for a general genomic graphical representation.
[0135] In an embodiment, constructing the processed biological sequence may include considering the spacing between different fingerprint data strings, and / or may include considering the direction of the fingerprint data string (e.g., the inherent direction) to construct the directed graph.
[0136] In an embodiment, constructing the processed biological sequence may include considering the structural and / or spatial shape information embedded in the fingerprint data string for constructing the directed graph, and / or may include considering the syntactic information embedded in the fingerprint data string.
[0137] In an embodiment, the search in step b may consider any one of the positional information, spacing information, secondary and / or tertiary and / or quaternary structure of the characteristic biological subsequence, and / or structural variations of the characteristic biological subsequence among different elements of the characteristic biological sequence.
[0138] By way of illustration, and not limited to this embodiment of the present invention, an example of how to search for a certain sequence is shown below. The method includes identifying HYFT present in the sequence to be searched in a first step TM . The method then further includes querying a reference database by searching all sequences in the reference database that also contain said HYFT TM . The different sequences found are then sorted, for example by length, and the position of the HYFT TM in the sequence is identified. In addition, alignment is performed. In some embodiments, Navarro-Levenshtein matching may be used to perform the alignment. A more detailed description of Navarro-Levenshtein matching can be found, for example, in Navarro, Theoretical Computer Science 237(2000)455-463. A directed graph, such as a directed acyclic graph, may be used to perform the alignment. The latter may be a general genomic reference graph, but the embodiments are not limited thereto. The alignment may include identifying changes in a specific sequence. To perform the above steps, the sequence may be further processed, whereby, for example, dead ends and loops may be removed.
[0139] A processed biological sequence is also described, which can be obtained by a computer-implemented method for processing biological sequences as described above. Figure 4 The processed biological sequence 210 is schematically depicted in
[0140] A computer-implemented method for constructing and / or updating a repository of processed biological sequences is also described, which includes populating the repository with the processed biological sequences as described above. Figure 4 A repository construction unit 320 for storing the processed biological sequence 210 into a repository 220 of processed biological sequences is schematically shown.
[0141] An advantage of the embodiments of the present invention is that a repository of processed biological sequences can be constructed and stored.
[0142] A repository of processed biological sequences is also described, which can be obtained by a computer-implemented method for constructing and / or updating a repository of processed biological sequences as described above. Figure 4 The repository 220 is schematically depicted in
[0143] One advantage is that repositories of processed biological sequences can be searched and navigated quickly. Another advantage is that, compared to known databases, the storage size of the repository can be relatively small by populating the repository with compressed processed biological sequences.
[0144] In an embodiment, a repository of processed biological sequences can be combined with a repository of fingerprint data strings.
[0145] In an embodiment, the repository can be a repository of processed biological fragment sequences (i.e., processed biological sequences of biopolymer fragments).
[0146] In an embodiment, the repository can be a database. In some embodiments, the repository of processed biological sequences can be an index repository. For example, the repository can be indexed based on fingerprint markers (corresponding to characteristic biological subsequences) present in each processed biological sequence. In other embodiments, the repository can be a graphical repository.
[0147] A computer-implemented method for comparing a first biological sequence with a second biological sequence is also described, which includes: (a) processing the first biological sequence by the computer-implemented method described above to obtain a processed first biological sequence, or retrieving the processed first biological sequence from a repository of processed biological sequences described above, (b) processing the second biological sequence by the computer-implemented method described above to obtain a processed second biological sequence, or retrieving the processed second biological sequence from a repository of processed biological sequences described above, and (c) comparing at least the fingerprint markers in the processed first biological sequence with the fingerprint markers in the processed second biological sequence. Figure 5 The comparison unit 330 is schematically shown, which compares at least the first biological sequence 211 with the second biological sequence 212 to output a result 400.
[0148] The advantages of the embodiments of the present invention are that the comparison of biological sequences can be changed from an NP-complete problem or an NP-hard problem to a polynomial-time problem. Another advantage of the embodiments of the present invention is that the comparison can be performed in a greatly reduced time and can scale well with an increase in complexity (e.g., an increase in the length or number of biological sequences). Yet another advantage of the embodiments of the present invention is that the required computing power and storage space can be reduced.
[0149] The advantages of the embodiments of the present invention are that the degree of similarity between biological sequences can be calculated. Another advantage of the embodiments of the present invention is that they can be ranked based on the degree of similarity of multiple biological sequences.
[0150] An advantage of the embodiments of the present invention is that sequence similarity searches can be performed quickly and easily (e.g., in polynomial time).
[0151] An advantage of the embodiments of the present invention is that the compared biological sequences can be aligned easily and quickly (e.g., in polynomial time).
[0152] An advantage of the embodiments is that multiple sequences can also be compared and aligned easily and quickly. Another advantage of the embodiments is that there is no error accumulation during alignment, as is the case in currently known methods (e.g., progressive alignment).
[0153] An advantage of the embodiments of the present invention is that the sequences of biological polymer fragments can be aligned and merged easily and quickly to reconstruct the original biological polymer sequence.
[0154] By using the characteristic biological subsequences (fingerprint markers in the processed biological sequences) according to the embodiments of the present invention, the problem of comparing sequences is advantageously reformulated from an NP-complete or NP-hard problem to a polynomial time problem. In fact, identifying the fingerprints in the sequences and then comparing the sequences based on these fingerprints (which can be considered a lexical method) is computationally much simpler than the currently used algorithms (which compare the entire sequences, for example, based on the sliding window method). Thus, even when less computational power and storage space are required, the comparison can be performed significantly faster and can scale well with an increase in complexity (e.g., an increase in the length or number of biological sequences).
[0155] In an embodiment, the second biological sequence can be a reference sequence.
[0156] In an embodiment, step c may include identifying whether one or more characteristic biological subsequences (represented by fingerprint markers) in the processed first biological sequence correspond (e.g., match) to one or more characteristic biological subsequences (represented by fingerprint markers) in the processed second biological sequence. In an embodiment, step c may include identifying whether the corresponding characteristic biological subsequences occur in the same order in the processed first biological sequence as in the processed second biological sequence. In an embodiment, step c may include identifying whether one or more pairs of characteristic biological subsequences in the processed first biological sequence and one or more pairs of corresponding characteristic biological subsequences in the processed second biological sequence have the same or similar (e.g., differing by less than 1000 sequence units, e.g., less than 100 sequence units, preferably less than 50 sequence units, still more preferably less than 20 sequence units, most preferably less than 10 sequence units) spacing.
[0157] In an embodiment, step c may further include comparing one or more second portions of the processed first biological sequence with one or more second portions of the processed second biological sequence. In an embodiment, comparing the one or more second portions may include comparing corresponding second portions (i.e., the second portion between adjacent pairs of characteristic biological subsequences that appear in the processed first biological sequence with the second portion between corresponding adjacent pairs of characteristic biological subsequences that appear in the processed first biological sequence).
[0158] In an embodiment, step c may further include calculating a measure representing the degree of similarity (e.g., Levenshtein distance) between the first biological sequence and the second biological sequence. In an embodiment, the degree of similarity may be calculated based on a plurality of variables, such as combining a measure of syntactic similarity with a measure of structural similarity.
[0159] In an embodiment, the method may be used in sequence similarity searches by comparing a query sequence with one or more other biological sequences (e.g., corresponding to a sequence database to be searched, e.g., in the form of a repository of processed biological sequences). In an embodiment, the degree of similarity for each of the other biological sequences may be calculated. In an embodiment, the method may include another step of ranking the biological sequences (e.g., by decreasing degree of similarity). In an embodiment, the method may include filtering the biological sequences. The filtering may be performed before and / or after step c. For example, the filtering may be performed by selecting only those biological sequences from the database that meet specific criteria, such as based on the organism or group of organisms from which they are derived (e.g., plants, animals, humans, microorganisms, etc.), whether the secondary / tertiary / quaternary structure is known, their length, etc. Alternatively, the filtering may be performed after the comparison based on the same criteria or based on the calculated degree of similarity (e.g., only those sequences that exceed a certain similarity threshold may be selected). Contrary to sequence similarity searches in the prior art, where an alignment step is typically required and then a measure of similarity is established from it, according to the embodiment, alignment is not strictly necessary for similarity searches. In fact, in the absence of alignment, similar sequences may have been found by simply searching for sequences with the same fingerprint (optionally also considering their order and their spacing); this in turn allows the search speed to be further accelerated. Nevertheless, the alignment according to the embodiment (see below) is also computationally simplified, such that the alignment may be performed in any way, even without strict requirements.
[0160] The method thus allows the determination (and optionally measurement) of the similarity between the first biological sequence and the second biological sequence. Such comparisons are also the cornerstone of other methods, such as methods for alignment and assembly (see below).
[0161] In an embodiment, the method can be used to align a first biological sequence with a second biological sequence. In an embodiment, step c can further include aligning the fingerprint markers in the processed first biological sequence with the fingerprint markers in the processed second biological sequence. Figure 5 Schematically shows the output result 400 from a comparison unit 330 (which is better referred to as an "alignment unit 330" in this case), where biological sequences are aligned by their fingerprint markers.
[0162] Thus, alignment is also simplified in an embodiment because a good alignment may already be obtained by simply aligning the fingerprints. Again, this significantly reduces the computational complexity of the problem. Additionally, in prior art methods, such as those based on progressive alignment, there is an accumulation of alignment errors because a misalignment in one of the earlier sequences typically propagates and causes additional misalignments in later sequences. In contrast, since the same discrete set of fingerprint markers is aligned (or at least attempted to be aligned) within one (or more) alignment each time, there is no such error propagation.
[0163] In an embodiment, the method can further include subsequently aligning corresponding second parts. For example, one of the alignment methods known in the prior art can be used to perform the alignment of the second part. In fact, since the "skeleton" of the alignment has been provided by aligning the fingerprint markers, only the alignment between these markers remains to be filled in. Since each of these second parts is typically relatively short compared to the total biological sequence length, known methods can generally perform such alignments relatively quickly and efficiently.
[0164] In an embodiment, the method can be used to perform a multiple sequence alignment (i.e., the method can include aligning three or more biological sequences). In an embodiment, the method can include aligning the fingerprint markers in the processed third (or fourth, etc.) biological sequence with the fingerprint markers in the processed first and / or second biological sequences. This is schematically depicted in Figure 5 where the alignment unit 330 can also compare and align any number of further processed biological sequences 213 to 216.
[0165] In an embodiment, the method can be used for variant identification. In the case of a sequence alignment between two biological sequences, variant identification can identify variants (such as mutations) between a query sequence and a reference sequence. In the case of a multiple sequence alignment, variant identification can identify possible changes in a set of related sequences (which can include determining their frequency of occurrence); optionally relative to a reference sequence. Additionally, variants can be identified based on the primary structure, but secondary / tertiary / quaternary structures can also be considered. Thus, variants can be based on the primary structure, based on the secondary / tertiary / quaternary structure, and also based on the HYFT in the sequence TM related or relative to the next or previous HYFTTM Each possible correlation of distances related to distance information is used to identify variants. Variant identification can also be based on changes in the codon table, thus allowing immediate information on DNA, RNA, and amino acid changes to be collected in the same variant analysis.
[0166] In an embodiment, the method can be used to perform sequence assembly. In an embodiment, the method can include: (a) providing a first biological sequence, the first biological sequence being the biological sequence of a first biological polymer fragment, (b) providing a second biological sequence, the second biological sequence being the biological sequence of a second biological polymer fragment or a reference biological sequence, (c) aligning the first biological sequence with the second biological sequence as described above, and (d) merging the first biological sequence with the second biological sequence to obtain an assembled biological sequence. Figure 6 Schematically shows a sequence assembly unit 340 that outputs an assembled biological sequence 510 by first aligning (by its fingerprint marker) and then merging any number of biological sequences 500, including at least a first biological sequence 501 and a second biological sequence 502.
[0167] In an embodiment, method steps a to d can be repeated to align and merge any number of biological polymer fragments.
[0168] For ease of sequencing, longer biological polymers can be fragmented because sequencing of individual fragments is faster and easier (e.g., it can be sequenced in parallel); as is known in the art. Subsequently, sequence assembly is typically used to align and merge the fragment sequences to reconstruct the original sequence; this can also be referred to as "read mapping", where the "reads" from the fragment sequences are "mapped" to a second biological polymer sequence. Depending on the type of sequence assembly being performed, e.g., de novo assembly versus mapping assembly, the second biological polymer sequence can be selected as a second biological polymer fragment or a reference sequence as appropriate. In this context, de novo assembly is a de novo assembly without using a template (e.g., a scaffold sequence). In contrast, mapping assembly is an assembly by mapping one or more biological polymer fragment sequences to an existing scaffold sequence (e.g., a reference sequence), which is typically similar (but not necessarily identical) to the sequence to be reconstructed. The reference sequence can be, for example, based on a complete genome or transcriptome (a part thereof), or can be obtained from an earlier de novo assembly.
[0169] In an embodiment, the method can include another step e of aligning the assembled biological sequence with the second biological sequence as described above after step d. This additional alignment can be used to perform variant identification of the assembled biological sequence relative to the second biological sequence (e.g., a reference sequence).
[0170] In an embodiment, the fingerprint data string can be inherently oriented and include position information.
[0171] In an embodiment, the method may further include detecting variations such as, but not limited to in this embodiment, insertions, deletions, insertions and / or duplications.
[0172] In an embodiment, providing the first biological sequence and / or the second biological sequence may be performed using the methods described above.
[0173] A storage device is also described, which includes a repository of fingerprint data strings as described above and / or a repository of processed biological sequences as described above.
[0174] A processing system is further described, which includes such a storage device and further includes a processor adapted to obtain fingerprint data strings from the storage device and / or adapted to store fingerprint data strings to the storage device and / or search in the fingerprint data strings in the storage device.
[0175] A data processing system is also described, which is adapted to (e.g., includes means for) perform any of the computer-implemented methods described above.
[0176] The system may generally take different forms depending on the method it is intended to perform. In an embodiment, the system may be or include a sequence processing unit, a variant identification unit, a repository construction unit, a comparison unit, an alignment unit or a sequence assembly unit. In an embodiment, a general-purpose data processing device (e.g., a personal computer or a smart phone) or a distributed computing environment (e.g., a cloud-based system) may be configured to perform one or more of these functions. The distributed computing environment may include, for example, server devices and networked client devices. Herein, the server device may perform most of one or more methods, including a repository of fingerprint data strings and a repository of processed biological sequences. On the other hand, the networked client devices may communicate instructions (e.g., inputs, e.g., query sequences, and settings, e.g., search preferences) to the server device and may receive method outputs.
[0177] A computer program (product) including instructions is also described, which when the program is executed by a computer (system), causes the computer to perform any of the computer-implemented methods described above.
[0178] A computer program product including instructions is further described, which when the program is executed by a computer system, causes the computer system to perform obtaining, searching or storing fingerprint data strings from, in or to a repository of fingerprint data strings, respectively.
[0179] A computer-readable medium including instructions is also described, which when executed by a computer (system), causes the computer to perform any of the computer-implemented methods described above.
[0180] Also described is the use of a repository of fingerprint data strings as described above for one or more of the following: sequencing a biopolymer or biopolymer fragment; performing sequence assembly; processing biological sequences; constructing a repository of processed biological sequences; comparing a first biological sequence with a second biological sequence; aligning a first biological sequence with a second biological sequence; performing a multiple sequence alignment; performing a sequence similarity search; performing variant identification; and identifying a target or biomarker.
[0181] Also described is the use of a processed biological sequence as described above or a repository of processed biological sequences as described above for one or more of the following: comparing a first biological sequence with a second biological sequence; aligning a first biological sequence with a second biological sequence; performing a multiple sequence alignment; performing a sequence similarity search; performing variant identification; and identifying a target or biomarker.
[0182] In an embodiment, any feature of any embodiment of any of the above aspects can be independently described accordingly as for any other aspect or any embodiment of any other described subject matter.
[0183] Aspects of several embodiments will now be described by way of a detailed description of several embodiments. Clearly, other embodiments of the present invention can be configured according to the knowledge of those skilled in the art without departing from the true technical teachings of the present invention, and the present invention is limited only by the terms of the appended claims.
[0184] Example 1: Associated biological information according to the present invention
[0185] Example 1a: Finding Biological Sequences with Equivalent Biological Functions
[0186] For applications in the agricultural field, a proof-of-concept information retrieval was performed. In this proof-of-concept, HYFT TM The protein fingerprint "WIGLVFL" was identified as a fingerprint that occurs relatively frequently within the domain. All protein sequences containing "WIGLVFL" were then retrieved from the repository of processed biological sequences as described herein and the results were analyzed. Notably, after studying the biological functions of the retrieved sequences using public databases, it was found that most of them were related to photosynthesis and this was across different species. Thus, it was found that HYFT TM "WIGLVFL" is an anchor related to different but functionally related biological entities.
[0187] Example 1b: Finding the Link between Related Biological Sequences
[0188] As another proof of concept, a simple text search was performed on protein sequences whose name includes "fibroblast growth factor receptor 2". The corresponding results were retrieved from a repository of processed biological sequences as described herein. After analyzing the retrieved results, it was found that substantially all protein sequences had "WSLIMES" or "WIKHVEK" as the most stringent HYFT TM (i.e., the longest HYFT with the lowest combination number TM ). Based on this, the repository of processed biological sequences and / or the repository of fingerprint data strings can be annotated with this information so that whenever information about HYFTs TM "WSLIMES" and / or "WIKHVEK" is sought for representative biological entities, this information can be used.
[0189] Note that there may be different entry and exit points here, which are linked by HYFTs TM . For example, a text search such as the above can be performed to retrieve information about, for example, the species in which such sequences occur. In another example, a specific protein domain can be used, and then its representative HYFT TM can be determined and a list of proteins sharing a similar domain can be generated through it. Similarly, the representative HYFT of a drug target TM can be identified, and through the said HYFT TM the potential other targets of the drug can be revealed; for example, allowing prediction and / or rationalization of side effects.
[0190] Example 1c: Finding connections between patients with equivalent medical conditions
[0191] In yet another proof of concept, publicly available data from cancer studies of the BRCA1 gene from different subjects (WEIGELT, Britta et al., Diverse BRCA1 and BRCA2 reversion mutations in circulating cell-free DNA of therapy-resistant breast or ovarian cancer. Clinical Cancer Research, 2017, 23.21: 6708 - 6720.) was processed. It was found that most typically there was a specific pattern of four HYFTs TM . However, surprisingly, it was found that subjects reported to be resistant to chemotherapy lacked the second HYFT TM ("TKCDHIF" in the corresponding protein) in this pattern. This is schematically shown Figure 7 for the selection of some of the subjects.
[0192] Thus, the absence of such a "TKCDHIF" subsequence in the protein sequence - or the absence of the corresponding DNA sequence encoding said protein sequence in the BRCA1 gene of a subject - indicates the presence of chemoresistance. Thus, this knowledge can be used to rapidly identify patients in such cases for whom chemotherapy may be less effective and to provide them with adjusted treatment.
[0193] (The reference sequence of the BRCA1 protein is publicly available in the UniProt database under accession number P38398, sequence version 2, entry versi
[0194] Example 1d: Processing of sequencing reads
[0195] By way of illustration, the embodiments of the present invention are not limited thereto, Figure 8 Examples of possible sequencing implementations are shown in. The figures show possible different method steps of a sequencing method according to an embodiment of the present invention. The method includes, after obtaining at least a first read of a biopolymer or a biopolymer fragment, and typically during further receipt of reads of the biopolymer or biopolymer fragment to be sequenced, parsing incoming, e.g., received, reads with fingerprints, called HYFTs TM . After parsing, an alignment can be performed to obtain a graph representing the sequence of the biopolymer or biopolymer fragment. The alignment can be performed by aligning with a directed graph, e.g., a directed acyclic graph. The latter can be a general genomic reference graph, but the embodiments are not limited thereto. The alignment can include identifying changes in a specific sequence. However, other intermediate steps can also be performed, such as constructing an overview graph, whereby the processed (e.g., parsed) sequences are grouped around one or more fingerprints common or linked between the processed sequences, and the data is folded, e.g., by sorting in the overview graph. Such folding can be performed one character at a time, and nodes can be split when the characters are different. The method can also include forming a sub-read graph, whereby dead ends or bubbles are typically removed in said step. It should be noted that removing dead ends and / or bubbles can alternatively or additionally be performed in other steps of the method. The method can also include forming a read graph, wherein the sub-read graphs are combined. By way of further illustration, the embodiments of the present invention are not limited thereto, Figures 9 to 12 Different steps are shown in. Figure 9 Illustrates the use of HYFTs TM The step of parsing incoming reads. It should be noted that the portions of the sequences shown in the figures do not themselves form part of the present invention, but are introduced only for illustration of the processing of such data. Identifying a certain fingerprint of a repository, i.e., the occurrence of a HYFT TM in the read. Figure 10 Illustrates the construction of an overview graph, whereby different processed sequences are grouped around the found linked HYFT TMGrouping is performed. Figure 11 Illustrated is the construction of an overview map by sorting and folding. The latter can be performed one character at a time and, when the characters are different, by splitting nodes. Additionally, the sequence of covering nodes can be tracked. Typically, it can start from the HYFT TM fingerprint and typically move in one direction (e.g., to the right). Figure 12 Illustrated is a cleaning step in which loose ends are removed. Alternatively or in addition, bubbles or small internal loops can also be addressed.
[0196] Example 2: Processing of protein databases
[0197] Example 2a: Analysis of the protein database with respect to the HYFT TM fingerprint in the protein database
[0198] To illustrate the prevalence of the HYFT TM fingerprint in bioinformatics sources, the Protein Data Bank (PDB) is taken as an example of a large, generally available bio-sequence database, and the repository of fingerprint data strings obtained as described above is processed according to the present invention. The results are analyzed with respect to various metrics, and their selection is given below.
[0199] Figure 13 and Figure 14 show the HYFT TM coverage (in %) of processed protein sequences up to lengths of 50 and up to lengths over 5000, respectively. Here, the coverage is the portion of the total sequence length in which the sequence units belong to the HYFT TM fingerprint. In other words, the coverage is the combined length of one or more first parts divided by the total sequence length.
[0200] For cases up to lengths over 5000, the inverse statistic is shown in Figure 15 i.e., the portion of the total sequence length not covered by the HYFT TM fingerprint (or the combined length of one or more second parts divided by the total sequence length).
[0201] Associated with the above, Figure 16 an overview of the number of HYFTs TM retrieved for each processed sequence is given in the form of a frequency distribution.
[0202] Notably, these charts show that at least one HYFT TM fingerprint is found in each processed biological sequence; in fact, no PDB sequence is not covered by one or more HYFTs TM . Additionally, the HYFT TMThe mode widely covers long sequences, where the coverage usually decreases as the sequence length increases. On average, a coverage close to 80% is achieved.
[0203] The typical spacing observed is shown in Figure 17 which depicts the frequency distribution of the lengths of the second parts that occur before and after the HYFT TM fingerprint.
[0204] Overall, the above results support that in fact each protein sequence (and by extension DNA and / or RNA sequences) can be rewritten as one or more HYFTs TM from a repository of HYFT TM fingerprint data strings (i.e., HYFT TM patterns). In addition, due to the generally good coverage achieved, the processed sequences still retain the basic characteristics of their unprocessed counterparts; especially when not only the identified HYFTs TM are retained, but they are also extended with additional data (see above), such as the spacing (i.e., the length of the second part) before, between, and after the identified HYFTs TM . High-performance indexing based on HYFT TM patterns can be achieved - with a nearly perfect retrieval rate.
[0205] Example 2b: Effect of the matching strategy employed
[0206] Since different strategies can be employed when processing biological sequences according to the present invention, the differences between two different methods were studied. In the first method, all occurrences of the HYFT TM fingerprint were searched for in the biological sequences in the PDB database, including overlapping HYFTs TM , such that the order of the HYFT TM fingerprints became irrelevant. In the second method, the biological sequences in the PDB database were searched in a more stringent manner, where the search was performed in the order from the longest HYFT TM fingerprint to the shortest HYFT TM fingerprint, and - in the case of the same length - from the lowest combination number to the highest combination number, and where no overlapping of HYFTs TM was allowed (i.e., where the part found corresponding to the HYFT TM was excluded from then on for searching for other HYFTs TM ). The goal of the second method was to identify the minimum number of HYFTs TM to describe the processed biological sequence, while by not allowing overlapping and by supporting a more stringent HYFTs TM(i.e., shorter lengths and higher combination numbers) more stringent HYFTs TM (i.e., longer lengths and lower combination numbers), still ensuring good coverage of the sequence.
[0207] In Figure 18 the number of different matches found for each biological sequence was plotted relative to each other. It was observed that for a second method that was more stringent than the first method, there was a roughly linear relationship with actually approximately 5 times fewer matches found. These fewer matches corresponded to increased processing time - identifying HYFT TM fingerprints and subsequently using the processed sequences in other methods - and the storage space required; nonetheless, the entire sequence was fully and completely characterized. Thus, the second method was considered to achieve the best balance and was generally preferred.
[0208] Nonetheless, it was noted that the number and nature of the matches found using the first method were fewer and better than comparable k-mer methods. Thus, although the second method was generally superior to the first method, the first method was still superior to methods of the prior art.
[0209] Example 3: Comparison between sequence searches known in the prior art and those described herein
[0210] Example 3a: Using short search strings
[0211] Two separate searches were performed based on the search string "AVFPSIVGRPRHQGVMVGMGQKDSY". This corresponds to a relatively short protein sequence of length 25 sequence units, which could be, for example, a protein fragment in protein sequencing. Such searches could be used, for example, after fragment sequencing as part of identifying a suitable reference sequencing to be used with the fragment in sequence assembly.
[0212] The first search was performed using BLAST (Basic Local Alignment Search Tool); more specifically, "Protein BLAST" (available at: https: / / blast.ncbi.nlm.nih.gov / Blast.cgi?PROGRAM=blastp&PAGE_TYPE=BlastSearch&LINK_LOC=blasthome) was used. The following search parameters were used: database = Protein Data Bank proteins (pdb); algorithm = blastp (protein-protein BLAST); max target sequences = 1000; short query = automatically adjust parameters for short input sequences; expect threshold = 20000; word size = 2; matrix = PAM30; composition adjustment = no adjustment. BLAST took over 30 seconds to perform this search and then returned 604 search results.
[0213] On the other hand, based on the principles of the present invention, it is determined that "IVGRPRHQGVM" is a characteristic biological subsequence (i.e., "HYFT TM fingerprint") included in the above short protein sequence. Therefore, a second search is performed based on the search string "IVGRPRHQGVM" in the repository of processed biological sequences. This repository is based on the same protein database (i.e., the Protein Data Bank; PDB) used in BLAST, which has previously been processed using the repository of fingerprint data strings (see above); i.e., characteristic biological subsequences represented by fingerprint data strings are identified and labeled in the collection of publicly available biological sequences. This search returns 661 results. Compared with BLAST, the time frame required in this case is only 196 milliseconds. Thus, even for such relatively short sequences, it is observed that the present method is capable of reducing the required time by more than 150 times compared to the prior art method.
[0214] Now refer to Figure 19 , Figure 20 and Figure 21 , which show the results of these two searches (BLAST = dashed line; present method = solid line) in terms of their total length ( Figure 19 ), their Levenshtein distance ( Figure 20 ) and longest common substring ( Figure 21 ). For each graph, the search results are shown in ascending order with respect to the plotted parameter (i.e., total length, Levenshtein distance or longest common substring). In addition, one of the search results, i.e., the protein sequence 5NW4_V (i.e., the first result listed by BLAST), is selected as a reference for calculating the Levenshtein distance and longest common substring. As can be observed in these graphs, the present method produces a smaller variation in total length (characterized by a relatively flat segment spanning a significant portion of the results) across the entire range of search results, a significantly lower Levenshtein distance and a significantly larger longest common substring; compared to the BLAST results. The combination of these indicates that the method of the present invention is capable of identifying results that are more relevant to the search performed.
[0215] Example 3b: Using a longer protein as the search string
[0216] Repeat the previous example, but this time search for the complete protein sequence, 3MN5_A (length 359 sequence units).
[0217] The first search, using BLAST, returns 88 search results.
[0218] On the other hand, based on the principles of the present invention, it is determined that six characteristic biological subsequences (i.e., "HYFT TM fingerprint") can be found in the sequence 3MN5_A; these are represented as:
[0219] +4641474444415052415646_1, +495647525052485147564d_1,
[0220] +4949544e5744444d454b49_1, +494d464554464e5650414d_1,
[0221] +494b454b4c435956414c44_1 and +49474d4553414749484554_1,
[0222] Among them, for example, "49474d4553414749484554" corresponds to the corresponding subsequence in hexadecimal format. Therefore, a second search is performed in the repository of processed biological sequences that is the same as the previous instance to find those protein sequences that include the same six characteristic biological subsequences in the same order. This search returns 661 results.
[0223] We now refer to Figure 22 , Figure 23 and Figure 24 , which show the results of these two searches (BLAST = dashed line; this method = solid line) in terms of their total length ( Figure 22 ), their Levenshtein distance ( Figure 23 ) and longest common substring ( Figure 24 ). For each graph, the search results are shown in ascending order with respect to the plotted parameter (i.e., total length, Levenshtein distance or longest common substring). In this case, the Levenshtein distance and longest common substring are calculated with respect to the original query sequence 3MN5_A. As can be observed in these figures, the characteristics of the search results of the two methods are relatively comparable in extreme cases. However, this method produces a stable result in the intermediate range, with a smaller variation in total length, a lower Levenshtein distance and a relatively high longest common substring. The combination of these indicates that the method of the present invention is capable of identifying a larger number of relevant results.
[0224] It should be understood that although preferred embodiments, specific structures and configurations, and materials have been discussed herein with respect to the device according to the present invention, various changes or modifications can be made in form and detail without departing from the scope and technical teaching of the present invention. For example, any formula given above merely represents a program that can be used. Functions can be added or removed from the block diagram, and operations can be interchanged between functional blocks. Steps can be added or removed in the methods described within the scope of the present invention. Sequence Listing <110> BioStrand B.V.B.A. BioKey B.V.B.A. BioClue N.V. <120> Biological information processing <130> 20024VTr00WO / dw / ac / av <140> EPPCTNYK <141> 2020-02-07 <150> EP19190901.9 <151> 2019-08-08 <150> EP19190899.5 <151> 2019-08-08 <150> EP19190900.1 <151> 2019-08-08 <150> EP19156086.1 <151> 2019-0۲-۰۷ <150> EP19156085.3 <151> 2019-0۲-۰۷ <150> BE2019 / 5077 <151> 2019-0۲-۰۷ <160> 14 <170> BiSSAP 1.3.6 <210> 1 <211> 7 <212> PRT <213> Unknown <220> <223> Unknown <400> 1 Met Cys Met His Asn Gln Ala 1 5 <210> 2 <211> 6 <212> PRT <213> Unknown <220> <223> Unknown <400> 2 Met Cys Met His Asn Gln 1 5 <210> 3 <211> 25 <212> PRT <213> Unknown <220> <223> Unknown <400> 3 Ala Val Phe Pro Ser Ile Val Gly Arg Pro Arg His Gln Gly Val Met 1 5 10 15 Val Gly Met Gly Gln Lys Asp Ser Tyr 20 25 <210> 4 <211> 11 <212> PRT <213> Unknown <220> <223> Unknown <400> 4 Ile Val Gly Arg Pro Arg His Gln Gly Val Met 1 5 10 <210> 5 <211> 11 <212> PRT <213> Unknown <220> <223> Unknown <400> 5 Phe Ala Gly Asp Asp Ala Pro Arg Ala Val Phe 1 5 10 <210> 6 <211> 11 <212> PRT <213> Unknown <220> <223> Unknown <400> 6 Ile Val Gly Arg Pro Arg His Gln Gly Val Met 1 5 10 <210> 7 <211> 11 <212> PRT <213> Unknown <220> <223> Unknown <400> 7 Ile Ile Thr Asn Trp Asp Asp Met Glu Lys Ile 1 5 10 <210> 8 <211> 11 <212> PRT <213> Unknown <220> <223> Unknown <400> 8 Ile Met Phe Glu Thr Phe Asn Val Pro Ala Met 1 5 10 <210> 9 <211> 11 <212> PRT <213> Unknown <220> <223> Unknown <400> 9 Ile Lys Glu Lys Leu Cys Tyr Val Ala Leu Asp 1 5 10 <210> 10 <211> 11 <212> PRT <213> Unknown <220> <223> Unknown <400> 10 Ile Gly Met Glu Ser Ala Gly Ile His Glu Thr 1 5 10 <210> 11 <211> 7 <212> PRT <213> Unknown <220> <223> Unknown <400> 11 Trp Ile Gly Leu Val Phe Leu 1 5 <210> 12 <211> 7 <212> PRT <213> Unknown <220> <223> Unknown <400> 12 Trp Ser Leu Ile Met Glu Ser 1 5 <210> 13 <211> 7 <212> PRT <213> Unknown <220> <223> Unknown <400> 13 Trp Ile Lys His Val Glu Lys 1 5 <210> 14 <211> 7 <212> PRT <213> Unknown <220> <223> Unknown <400> 14 Thr Lys Cys Asp His Ile Phe 1 5
Claims
1. A computer-implemented method for obtaining information about a biological entity based on at least one biological sequence, comprising: a. providing a repository of fingerprint data strings for a biological sequence database, each fingerprint data string representing a characteristic biological subsequence composed of sequence units, where the sequence units are amino acids when the characteristic biological subsequence is related to a protein and codons when the characteristic biological subsequence is related to DNA or RNA, and each characteristic biological subsequence has a combination number less than the total number of different sequence units available in the biological sequence database, and the combination number of the biological subsequence is defined as the number of different sequence units that appear as consecutive sequence units of the biological subsequence in the biological sequence database; b. determining one or more fingerprint data strings representing the biological entity; c. searching in a repository containing information associated with the fingerprint data strings for information associated with the one or more fingerprint data strings representing the biological entity; and d. processing the information.
2. The computer-implemented method according to claim 1, wherein the one or more fingerprint data strings representing the biological entity include the fingerprint data string representing the longest characteristic biological subsequence found in the at least one biological sequence, or if more than one longest characteristic biological subsequence is found, the characteristic biological subsequence having the lowest combination number among the longest characteristic biological subsequences.
3. The computer-implemented method according to claim 2, which is used to process sequencing reads of a biopolymer or a biopolymer fragment while considering the information contained in the repository of fingerprint data strings; Among them, the information associated with the fingerprint data strings included in the repository may contain combination data representing different sequence units that appear as consecutive sequence units of the corresponding characteristic biological subsequence in the biological sequence database; and wherein step b includes searching for the occurrence of one or more of the characteristic biological subsequences represented by the fingerprint data strings in the read, and step d includes validating or rejecting the read by determining for each occurrence whether the sequence units consecutive to the characteristic biological subsequence conform to the combination data in the repository, and / or step b includes searching for the occurrence of one of the characteristic biological subsequences represented by the fingerprint data strings at the head and / or tail of the read, and step d includes predicting consecutive sequence units from the combination data in the repository.
4. The computer-implemented method according to claim 3, performing the method on a batch of reads.
5. The computer-implemented method according to claim 3 or 4, wherein the fingerprint data string is inherently directed and includes position information, and the method includes another step of aligning the processed read with an orientation map using the characteristic biological subsequence identified in step b.
6. A computer-implemented method for associating information with one or more fingerprint data strings as defined in any one of the preceding claims, comprising: a. Provide the biological sequences of biological entities that share equivalent information; b. Search for equivalent characteristic biological subsequences in the biological sequences; and c. Associate the equivalent information with the fingerprint data string representing the equivalent characteristic biological subsequence.
7. The computer-implemented method according to claim 6, further comprising the following step a' before step a: a'. Search for biological entities sharing equivalent information in a data pool.
8. The computer-implemented method according to claim 7, wherein the data pool includes sequencing data or a biological sequence database.
9. A data processing system adapted to execute the computer-implemented method according to any one of claims 1 to 8.
10. A computer-readable medium comprising instructions that, when executed by a computer, cause the computer to execute the computer-implemented method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Methods for nucleic acid and polypeptide similarity search employing content addressable memories
US20060020397A1