Method for compression of genome sequence data
The method addresses the inefficiencies of existing genomic sequencing data compression techniques by encoding reads based on their mapping to a reference sequence, resulting in high-speed, high-compression-ratio compression that preserves read order and prevents information loss.
Patent Information
- Application Number
- JP2025030444
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2019-09-11
- Filing Date
- 2025-02-27
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2040-09-11
AI Technical Summary
Existing compression methods for genomic sequencing data face challenges such as low compression ratio, slow compression and decompression speeds, and information loss due to reordering of reads during compression.
A computer-implemented method that encodes genomic sequence data based on whether reads are fully or partially mapped to a reference sequence, using separate encoding processes for each case to reduce information source entropy and preserve the initial order of reads.
The method achieves high-speed compression and decompression with a high compression ratio, preserves the initial order of reads to prevent information loss, and facilitates easier downstream analysis and consistency checking.
Smart Images

Figure 2025087760000001_ABST
Abstract
Description
Technical Field
[0001] This field generally relates to a method for displaying genomic sequencing data generated by a sequencing machine, and more particularly to a computer-implemented method for compressing such genomic sequencing data. The present disclosure provides a reference-based compression method that enables fast compression and decompression while not causing loss of information and has a high compression ratio.
Background Art
[0002] Next-generation sequencing machines today generate vast amounts of sequencing data at low cost. Recent systems generate over 6 billion 150-nucleotide-long sequences sufficient to sequence 20 entire human genomes in a single 36-hour run. This opens up many new prospects for the diagnosis of genetic diseases and the development of personalized medicine, which aims to adopt treatments based on people's genomic specificities.
Summary of the Invention
Problems to be Solved by the Invention
[0003] However, this also comes with new problems, particularly the cost associated with storing vast amounts of data. The most commonly used file format for raw (unaligned) sequence data is the FASTQ format, which holds sequence data (strings of A, C, T, G nucleotides, also called reads), quality values (the probability that the sequencing platform made a sequencing error for each nucleotide), and sequence names. This is typically a simple ASCII text file compressed with the general-purpose text compression scheme LZ (Lempel-Ziv scheme, implemented in the gzip software). However, the use of such compression methods has several problems, - Low compression ratio due to the fact that data redundancy is not fully utilized - Slow compression and decompression associated with it.
[0004] There also exist compression methods specialized for FASTQ encoding that are split in a reference-based or non-reference-based manner. However, a) reference-based methods have a good compression ratio but are slow, and b) non-reference-based methods are faster but have a lower compression ratio, so neither fully meets the requirements. An example of such a non-reference-based method is provided by the software SPRING, which is a reference-free compressor for FASTQ files (World Wide Web address: github.com / shubhamchandak94 / SPRING). However, the compression method provided by the software SPRING has a low compression ratio.
[0005] Among reference-based compression methods, several methods have been proposed that use array alignment and aim to be faster with a good compression ratio. However, such methods are troubled by several problems, and in particular, the main issue is that these problems do not completely disappear. Such known reference-based compression methods are described, for example, in International Publication No. 2018 / 068829 (A1) of a patent document. In the described method, after being aligned to one or more reference arrays, the nucleotide sequences are classified by matching the degree of accuracy (thereby creating classes of aligned reads), and then, for each layer where the data is divided, different source models and entropy coders are used to code them as syntax elements of multiple layers. Therefore, the classes of data are encoded separately and are composed of syntax elements of different layers, and each layer contains a descriptor that uniquely represents the classified and aligned reads of that layer. This method is intended to obtain separate information sources with the reduction of information entropy, thereby enabling an increase in compression performance and selective access to specific classes in the compressed data. However, such a compression method reorders the reads (i.e., the reads are reordered by these classes) in an order different from the order obtained at the end of the read alignment process. Then, some information is lost in the compression process, especially in the initial array ordering. Therefore, the reproducibility of some analysis results may be affected because some downstream analysis software may depend on the order of the reads. In addition, by decompressing the data in an order different from the initial order of the reads, it becomes much more difficult to confirm that the uncompressed file is the same as the initial file. Furthermore, such a compression method is relatively slow, especially when compared with the state-of-the-art non-reference-based compression methods. Means for Solving the Problems
[0006] The features of the following independent claims solve the problems of existing prior art solutions by providing a method for compressing genomic sequence data. In one aspect, a computer-implemented method for compressing genomic sequence data generated by a sequencing machine, wherein the genomic sequence data includes reads of a sequence of nucleotides or bases aligned to a reference sequence, thereby creating aligned reads, and the aligned reads are stored as a list of reads in an initial file, the computer-implemented method comprising: - for each aligned read, determining whether the read is fully or partially mapped to the reference sequence or not mapped to the reference sequence; - encoding the reads based on the determination, wherein reads determined to be fully mapped are encoded by a first encoding process and reads determined to be not mapped are encoded by a second encoding process; - the determining step includes comparing the number of mismatches between the read and the reference sequence with a threshold for each incompletely mapped read; - in the encoding step, reads determined to be incompletely mapped are encoded by a second encoding process or a third encoding process, and incompletely mapped reads are encoded by the second encoding process when the number of mismatches is greater than the threshold, and incompletely mapped reads are encoded by the third encoding process when the number of mismatches is less than the threshold; - in the second encoding process, each nucleotide or base of the read is encoded individually; - the first encoding process and the third encoding process include separate sets of descriptors, each set of descriptors uniquely representing the reads associated with the corresponding encoding process, and each of the first encoding process and the third encoding process is an encoding process for reducing information source entropy.
[0007] The present invention overcomes the drawbacks of conventional compression methods by enabling high-speed compression and decompression without causing information loss and by providing a high compression ratio. More specifically, the present invention focuses on encoding the most frequent cases in the most compact way, which means adopting a reduced encoding mode for rare and least frequent cases even when this is the case. This leads to a significant increase in compression performance. Furthermore, due to the genomic information representation format used in the present invention, the compression performed by the method according to the present invention is faster. Finally, the method according to the present invention preserves the initial order of the reads as such and does not reorder the reads according to their classes. As a result, no information is lost during the process, which enables easier downstream analysis and efficient consistency checking after the decompression step.
[0008] These and other features and advantages of the present invention will become more apparent from the accompanying drawings and the following detailed description of the invention. Additionally, although thresholds may be referred to herein as being exceeded or not exceeded, such thresholds can conceptually be adopted to determine whether such thresholds are met, conform to, or otherwise detected, regardless of whether the numbers or values used to perform their threshold evaluations are described using positive or negative values.
[0009] In accordance with one innovative aspect of the present disclosure, a method for compressing genomic sequence data is disclosed. In one aspect, the method can include performing one or more operations via execution of software instructions by one or more computers, the operations including obtaining, by one or more computers, read records; determining, by one or more computers, whether the read records correspond to reads that are fully mapped or incompletely mapped to a reference sequence; determining, by one or more computers, based on a determination that the read records correspond to reads that are incompletely mapped to the reference sequence, whether the number of mismatches of the incompletely mapped reads meets a predetermined number of mismatch thresholds; and encoding, by one or more computers, each mismatch of the incompletely mapped reads into a record having a size of one byte, based on a determination that the number of mismatches meets the predetermined number of mismatch thresholds.
[0010] Other aspects include corresponding systems, devices, and computer programs for performing the actions of the methods as disclosed herein, as defined by instructions encoded on a computer-readable storage device.
[0011] These and other versions may optionally include one or more of the following features. For example, in some implementations, determining, by one or more computers, whether the number of mismatches of the incompletely mapped reads meets a predetermined number of mismatch thresholds can include determining, by one or more computers, whether the number of mismatches of the incompletely mapped reads is greater than a predetermined number of mismatch thresholds.
[0012] In some implementations, each read record can include data indicating the absolute start position of the read aligned with respect to the reference array, data indicating the length of the read, data indicating whether the read was fully mapped or incompletely mapped, data indicating the number of mismatches identified within the read, and data indicating the relative position of each such possible mismatch within the read.
[0013] In some implementations, encoding each mismatch of an incompletely mapped read into a record having a size of 1 byte involves, for each particular mismatch, encoding the first 2 bits of the 1 byte to include data representing an alternative nucleotide or base present within the read instead of the corresponding reference nucleotide or base in the reference array by one or more computers, and encoding the remaining 6 bits of the 1 byte to include data representing the position of the mismatch in the reference array, where the position is calculated as an offset from the previous mismatch of the read.
[0014] In some implementations, the method can further include determining by one or more computers whether the offset is greater than the maximum encodable value, and inserting by one or more computers at least one false mismatch between a particular mismatch and a previous mismatch based on determining that the offset is greater than the maximum encoded value.
[0015] In some implementations, the method can further include encoding by one or more computers a list of positions in the reference array corresponding to each position of the mismatch with respect to the reference array using an entropy reduction encoding process based on determining that the number of mismatches does not meet a predetermined threshold number of mismatches.
[0016] In some implementations, the method may further include encoding at least a portion of the read record using information entropy reduction encoding by one or more computers based on determining that the read record corresponds to a read that is fully mapped to the reference array.
[0017] In some implementations, the one or more computers can include one or more hardware processors.
[0018] In some implementations, the one or more hardware processors can include one or more field programmable gate arrays (FPGAs).
[0019] In some implementations, the method for compressing genomic array data can be performed by one or more hardware processors. In such implementations, the hardware processor can include a hardware processing circuit configured to execute one or more operations. In one aspect, the operations can include obtaining a read record by the hardware processing circuit, determining by the hardware processing circuit whether the read record corresponds to a read that is fully mapped or incompletely mapped to the reference array, determining by the one or more computers based on determining that the read record corresponds to a read that is incompletely mapped to the reference array whether the number of mismatches of the incompletely mapped read meets a predetermined number of mismatch thresholds, and encoding each mismatch of the incompletely mapped read into a record having a size of one byte by the hardware processing circuit based on determining that the number of mismatches meets the predetermined number of mismatch thresholds.
[0020] These features and advantages of the present invention, as well as other features and advantages, will become more apparent from the accompanying drawings and the following detailed description of the invention.
[0021] In some implementations, each read record can include data indicating the absolute start position of the read aligned with the reference array, data indicating the length of the read, data indicating whether the read is fully mapped or incompletely mapped, data indicating the number of mismatches identified within the read, and data indicating the relative positions of the possible mismatches within the read.
[0022] In some implementations, determining by a hardware processing circuit whether the number of mismatches in an incompletely mapped read meets a predetermined threshold number of mismatches can include determining by the hardware processing circuit whether the number of mismatches in the incompletely mapped read is greater than a predetermined threshold number of mismatches.
[0023] In some implementations, encoding each mismatch in an incompletely mapped read into a record having a size of 1 byte can include, for each specific mismatch encoding, encoding the first 2 bits of the 1 byte to include data representing an alternative nucleotide or base present in the read instead of the corresponding reference nucleotide or base in the reference array, and encoding the remaining 6 bits of the 1 byte to include data representing the position of the mismatch in the reference array, where the position is calculated as an offset from the previous mismatch in the read.
[0024] In some implementations, the hardware processor circuit is further configured to determine, by one or more hardware processing circuits, whether an offset is greater than the maximum encodable value, and based on determining that the offset is greater than the maximum encoded value, insert, by the hardware processing circuit, at least one false mismatch between a particular mismatch and a previous mismatch.
[0025] In some implementations, the hardware processor circuit is further configured to execute, by the hardware processing circuit, an operation including encoding a list of positions of a reference array corresponding to each position of a mismatch with respect to the reference array using an entropy reduction encoding process, based on determining that the number of mismatches does not meet a predetermined number of mismatch thresholds.
[0026] In some implementations, the hardware processor circuit is further configured to execute, by the hardware processing circuit, an operation including encoding at least a portion of a read record using entropy reduction encoding, based on determining that the read record corresponds to a read that is fully mapped to the reference array.
[0027] In some implementations, the hardware processing circuit includes one or more field programmable gate arrays (FPGAs).
[0028] According to another innovative aspect of the present disclosure, a computer-implemented method for compressing genomic sequence data generated by a sequencing machine, the genomic sequence data including reads of a sequence of nucleotides or bases aligned to a reference sequence, thereby creating aligned reads, and the aligned reads being stored as a list of reads in an initial file. In one aspect, the method can include the actions of, for each aligned read, determining whether the read is fully mapped, incompletely mapped, or not mapped to the reference sequence, and encoding the read based on the determination. Reads determined to be fully mapped are encoded by a first encoding process, reads determined to be not mapped are encoded by a second encoding process, and the determining step includes comparing the number of mismatches between the read and the reference sequence to a threshold for each incompletely mapped read. In the encoding step, reads determined to be incompletely mapped are encoded by a second encoding process or a third encoding process. When the number of mismatches is greater than the threshold, the incompletely mapped reads are encoded by the second encoding process, and when the number of mismatches is less than the threshold, the incompletely mapped reads are encoded by the third encoding process. In the second encoding process, each nucleotide or base of the read is encoded individually. The first encoding process and the third encoding process each include a separate set of descriptors, each set of descriptors uniquely representing the reads associated with the corresponding encoding process, and each of the first encoding process and the third encoding process is an encoding process for reducing information source entropy.
[0029] Other aspects include corresponding systems, devices, and computer programs for performing the actions of the methods as disclosed herein, as defined by instructions encoded on a computer-readable storage device.
[0030] These and other versions may optionally include one or more of the following features. For example, in some implementations, the determining step may include determining that the read is incompletely mapped in the reference array and, when having a number of mismatches less than a threshold, making a further determination as to whether the read is globally mapped or locally mapped in the reference array. The third encoding process includes a first encoding sub-process and a second encoding sub-process. Reads determined to be globally mapped are encoded by the first encoding sub-process, and reads determined to be locally mapped are encoded by the second encoding sub-process. The first encoding sub-process and the second encoding sub-process include separate sets of descriptors, and each set of descriptors uniquely represents the reads associated with the corresponding encoding sub-process.
[0031] In some implementations, the descriptors of the first encoding sub-process can include the alignment start position in the reference array, the read length, and a list of mismatches due to symbol substitution. The descriptors of the second encoding sub-process include the local alignment start position in the reference array, the read length, a list of mismatches due to symbol substitution, and the length of the clipped portion of the read that is not part of the alignment.
[0032] In some implementations, in the encoding step, the clipped portions of the reads to be encoded by the second encoding sub-process are concatenated, and each nucleotide or base of the clipped portion is encoded individually.
[0033] In some implementations, in the encoding step, each mismatch of the incompletely mapped reads is encoded into 1 byte.
[0034] In some implementations, in the encoding process, for each mismatch of a partially mapped read, the two first bits of one byte are used to encode an alternative nucleotide or base present in the read instead of the corresponding reference nucleotide or base in the reference array, and the six last bits of one byte are used and encoded to encode the position of the mismatch in the reference array, which position is calculated as an offset from the previous mismatch of the read.
[0035] In some implementations, in the encoding process, if the offset calculated between a given mismatch and the previous mismatch is greater than the maximum encodable value, at least one false mismatch is inserted between the two mismatches until every offset between each of the mismatches and at least one false mismatch is lower than the maximum encodable value, and a false mismatch is defined as a mismatch in which bits of one byte are used to encode the mismatch or to encode a nucleotide or base equal to the corresponding reference nucleotide or base in the reference array.
[0036] In some implementations, the initial step of splitting a list of reads into blocks of reads begins with a header for each block that contains the information required to decode the block, and the compression method is performed for each block.
[0037] In some implementations, the blocks of reads have the same block size.
[0038] In some implementations, the final step of providing the compressed file includes a list of encoded reads, and the encoded reads are stored in the compressed file in the same order as the order of the reads stored in the initial file.
[0039] In some implementations, the threshold is equal to 31.
[0040] In some implementations, for each aligned read, a step of determining whether to include at least one mismatch corresponding to the case where the read could not be called by the sequencing machine for any base or nucleotide.
[0041] In some implementations, for each read including at least one mismatch corresponding to the case where the sequencing machine could not call any base or nucleotide, a step of determining the number of such mismatches, and a step of comparing the number with a reference threshold.
[0042] In some implementations, in the encoding step, when the number of such mismatches is greater than the reference threshold, each nucleotide or base of the read to be encoded by the second encoding process is individually encoded into 4 bits, and when the number of such mismatches is less than the reference threshold, each nucleotide or base of the read to be encoded by the second encoding process is individually encoded into 2 bits. The encoding step further includes encoding a list of positions along the reference sequence, and the positions correspond to the positions of such mismatches within the reference sequence.
Brief Description of the Drawings
[0043]
Figure 1
Figure 2
Figure 3
Figure 4
Embodiments for Carrying Out the Invention
[0044] The genomic sequences referred to in the present invention include, for example, but not limited to, nucleotide sequences, deoxyribonucleic acid (DNA) sequences, ribonucleic acid (RNA), and amino acid sequences. Although the description herein is rather detailed regarding genomic information in the form of nucleotide sequences, it will be understood that, although there are some variations, as will be understood by those skilled in the art, the compression method according to the present invention can be implemented for other genomic sequences.
[0045] Genomic sequencing information is generated by a sequencing machine in the form of a sequence of nucleotides (or more generally, bases) represented by a string from a defined vocabulary. The smallest vocabulary is represented by five symbols (A, C, G, T, N) that represent the four types of nucleotides present in DNA, namely adenine, cytosine, guanine, and thymine. In RNA, thymine is replaced by uracil (U). N indicates that the sequencing machine was unable to call any base, and thus the identity of the entity at that position is not determined.
[0046] The nucleotide sequences generated by an array determination machine are called "reads". Array reads can be from dozens to thousands of nucleotides in length. Some techniques generate array reads in pairs, where one read of the pair is from one DNA strand and the second read is from the other strand. Throughout this disclosure, a "reference sequence" is any sequence that aligns / maps the nucleotide or base sequences generated by an array determination machine. An example of such a reference sequence can actually be a reference genome, i.e., a sequence assembled by scientists as a representative example of a set of gene species. However, the reference sequence may also consist of a synthetic sequence made to simply improve the compressibility of the reads, taking into account further processing of the reads. An array determination machine can introduce errors into the array reads, and in particular, introduce the use of incorrect symbols to represent the nucleic acids or bases actually present in the sequenced sample (i.e., to represent different nucleic acids). This is usually called a substitution error or "mismatch".
[0047] The present invention is a reference-based compression method that receives as input reads of sequences of nucleotides or bases, and such reads create aligned reads by being already aligned to a reference sequence. The aligned reads are then stored as a list of reads in an initial file. The method of aligning the reads and storing the reads in the initial file once aligned is not important for the present invention and is not the object of this disclosure. Each read is then encoded as a list of positions on the reference sequence and differences from that reference sequence. Each read can then be reconstructed from the alignment encoding information and the reference sequence by appropriate decompression software configured according to the present invention.
[0048] Preferably, before providing the reads as input to the compression software and apparatus, alignment software that processes the reads and aligns the reads to a reference array does not consider certain types of errors introduced in the array reads, such as insertion errors or deletion errors. An insertion error is an insertion in an array read of one or more additional symbols that do not refer to any nucleic acid actually present. A deletion error is a deletion from an array read consisting of one or more symbols representing a nucleic acid actually present in the sequenced sample. More precisely, in the case of an insertion error or deletion error within a given array read, the alignment software will consider the resulting incorrect nucleic acid as a substitution error, also called a "mismatch". This preferential choice for the alignment software configuration enables faster subsequent coding and provides a better compromise between speed and compression ratio.
[0049] The alignment software provides, for each read, a corresponding read record to the compression software and apparatus. Each read record contains at least the following information: the absolute start position of the aligned read with respect to the reference array, the length of the read, the type of alignment of the read, the number of mismatches identified within the read, and, if necessary, the relative positions of the possible mismatches within the read.
[0050] Here, the compression method according to the present invention will be described with reference to FIG. 1. The method is executed, for example, by an apparatus 20 shown in FIG. 2. The apparatus comprises at least one processor 22 and one memory 24 operably coupled to the processor 22 to form a computing device. The memory 24 may store a computer program code or software 26 containing computer-executable instructions, which, when executed by the processor 22, cause the processor 22 to perform operations including the steps of the compression method according to the present invention.
[0051] The initial file in which the aligned reads are stored as a list of reads is stored, for example, in the memory of device 20. Returning to FIG. 1, the method preferably includes an initial step 2 of splitting the initial list of aligned reads into blocks of reads. Typically, the list of aligned reads is split into blocks of 50,000 reads, and this particular value is not to be construed as limiting the scope of the present invention, which can be applied as well as other values. Preferably, the blocks of reads have the same block size. Each block of reads begins with a header containing the information necessary to decode the block, such as, for example, the size in bytes of the contents of the block, and / or an identifier of the block or its contents, and / or the number of reads contained in the block. This enables support for concatenation of compressed files and streaming capabilities (each block of reads contains all the information necessary to decode the reads of the block). In addition, the compression method can then be run on several blocks, which also enables multi-threading in the blocks of reads, thereby enabling parallelization in the processing time and some resulting gain. If all the reads of a given block have the same length, the read length is also stored in the header; otherwise, a list of each read length is explicitly stored during the compression method.
[0052] Each read record contains information about the type of alignment of the read. Typically, two main types of alignment can be identified, namely, perfect alignment and imperfect alignment, as well as an additional type corresponding to "unmapped" reads. "Imperfect alignment" means that the read contains at least one mismatch other than N, while at least a part of the read matches a part of the reference sequence (according to this definition, an imperfectly mapped read may contain one or more Ns if it also contains one or more other mismatches). In an exemplary embodiment, each read record begins with the following bit flags, each bit flag having one value between two possible values: - A first bit flag indicating a forward or reverse direction with respect to the reference array, - A second bit flag indicating whether it is a complete alignment or not, - A third bit flag indicating whether the read contains at least one N, - A fourth bit flag indicating whether the position information is encoded in 16 bits or 32 bits.
[0053] The following steps 4 to 12 are executed one after another for the read blocks and one after another for the reads within the blocks.
[0054] The method includes the following step 4 of determining, for each aligned read, whether the read is fully mapped, incompletely mapped, or not mapped at all to the reference sequence. This determining step 4 can include, for each incompletely mapped read, comparing the number of mismatches between the read and the reference sequence with a threshold value (4A). In a preferred embodiment, although not to be construed as limiting the scope of the present invention, the threshold value is equal to 31. This specific value is intentionally selected to provide the best possible compromise for storing the number of mismatches in a sufficiently compact manner, as will be better understood later with respect to step 12. In fact, in most cases, it has been statistically observed that incompletely mapped reads have less than 31 mismatches. The rationale behind that selection is to encode the most frequent cases in the most compact way while still having some very rare degraded cases. If a read is determined to be incompletely mapped with a number of mismatches less than the threshold value, the determining step 4 also includes a further determination as to whether the read is globally mapped or locally mapped to the reference sequence. A "globally mapped read" is an incompletely mapped read where the entire sequence including the start and end of the read is incompletely mapped to the reference sequence. A "locally mapped read" is an incompletely mapped read that contains a segment of nucleotides or bases that are incompletely mapped to the reference sequence. Thus, the segment of nucleotides or bases corresponds to a part of the original read.
[0055] Preferably, the method further includes step 6 of determining, for each aligned read, whether the read contains at least one N, i.e., whether the read contains at least one mismatch corresponding to a case where the sequencing machine could not call any base or nucleotide. The method then includes step 8 of determining, for each read containing at least one N, the number of such N mismatches, and step 10 of comparing the number of N mismatches with a reference threshold. In a preferred embodiment, although not to be construed as limiting the scope of the present invention, the reference threshold is equal to 31.
[0056] Regardless of the result of step 4 of determining, the method includes step 12 of encoding the read at least by said determination. More precisely, reads determined to be fully mapped in the reference sequence, whether they contain no N or have a number of Ns smaller than the reference threshold, are encoded by a first encoding process. Reads determined not to be mapped, or reads determined to be fully mapped but having a number of Ns greater than the reference threshold, are encoded by a second encoding process in which each nucleotide or base is encoded individually, whether or not the nucleotide or base is aligned. Reads determined to be incompletely mapped are encoded by a second encoding process or a third encoding process. More precisely, reads determined to be incompletely mapped with a number of mismatches greater than the threshold are encoded by a second encoding process. If a read is determined to be incompletely mapped with a number of mismatches smaller than the threshold, and if the read contains no N or has a number of Ns smaller than the reference threshold, the read is encoded by a third encoding process. Otherwise, i.e., if the read has a number of Ns greater than the reference threshold, the read is encoded by a second encoding process.
[0057] Whether a given read is determined to be fully mapped, partially mapped, or unmapped, if the read contains at least one N and has a number of Ns that is less than a reference threshold, the encoding step 12 includes encoding a list of positions along the reference array, where the positions correspond to the positions of the Ns within the reference array. The list of positions is then stored in the memory of the computing device, which implements a compression method. If the read contains at least one N and has a number of Ns that is less than a reference threshold and is to be encoded by a second encoding process, each nucleotide or base of the read is individually encoded in 2 bits.
[0058] If the read contains at least one N and has a number of Ns that is greater than a reference threshold, the read is, in any case, encoded by a second encoding process, and each nucleotide or base of the read is individually encoded in 4 bits. In this case, the encoding step 12 does not include encoding and storing a list of the positions of the Ns within the reference array. In fact, each N mismatch is directly encoded by the second encoding process in much the same way as the other nucleotides or bases of the read.
[0059] The first encoding process and the third encoding process include separate sets of descriptors. Each set of descriptors uniquely identifies a read associated with the corresponding encoding process, and each of the first encoding process and the third encoding process is an encoding process for information entropy reduction. More precisely, the third encoding process includes a first encoding subprocess and a second encoding subprocess. The incompletely mapped reads determined to be globally mapped during step 4 are encoded by the first encoding subprocess. The incompletely mapped reads determined to be locally mapped during step 4 are encoded by the second encoding subprocess. The first encoding subprocess and the second encoding subprocess include separate sets of descriptors, and each set of descriptors uniquely identifies a read associated with the corresponding encoding subprocess.
[0060] Alignment information that encodes each read and enables reconstruction of the entire read array during data decompression then depends on the corresponding encoding process or subprocess used for that read. For example, the descriptors used in the first encoding process can be ○ The absolute start position of a read that is fully mapped with respect to a reference array encoded in 16 bits or 32 bits, and ○ The length of the read, encoded with differential coding with respect to the length of the preceding read, having a variable-length code in the range of 2 bits to 34 bits. and can be.
[0061] The descriptors used in the first encoding subprocess can be ○ The absolute start position of a read that is incompletely mapped with respect to a reference array encoded in 16 bits or 32 bits, ○ The length of the read, encoded with differential coding with respect to the length of the preceding read, having a variable-length code in the range of 2 bits to 34 bits, ○ A list of mismatches of the read, and can be.
[0062] The descriptors used in the second encoding sub-process are ○ The absolute start position of the partially unmapped portion of the read with respect to the reference array, also referred to as the local alignment start position (encoded in 16 bits or 32 bits), and ○ The length of the read (encoded with differential coding with respect to the length of the preceding read, having a variable-length code in the range of 2 bits to 34 bits), and ○ A list of read mismatches, and ○ The length of the clipped portion of the read that is not part of the alignment (encoded in 8 bits for each clipped portion), and can be.
[0063] Preferably, the list of mismatches encoded in the first and second sub-processes includes a header (a bit flag encoded in 1 byte). The first 5 bits of the 1 byte are used to encode the number of mismatches included in the read (in a preferred embodiment where the threshold is equal to 31, the number is in the range of [0 to 31]). Thereafter, 1 bit can be used to encode whether the partially unmapped read is globally mapped or locally mapped. Another bit can be used to encode whether the 2-bit mode is enabled for the second encoding process. The last bit can be used to encode whether the 4-bit mode is enabled for the second encoding process. Preferably, for each read encoded by the second encoding sub-process during step 12 of encoding, the clipped portions of the read (i.e., the portions that are not part of the local alignment) are concatenated, and each nucleotide or base of the clipped portion is encoded individually. In a preferred implementation, each nucleotide or base of such a clipped portion of the read is encoded individually in 2 bits.
[0064] In a preferred implementation form, each mismatch in the list of mismatches of incompletely mapped reads that is encoded (i.e., encoded by the first encoding sub-process or the second encoding sub-process) is encoded in one byte. More precisely, each mismatch of incompletely mapped reads that is to be encoded by the first encoding sub-process or the second encoding sub-process may be encoded as follows. ○ The two first bits of the one byte are used to encode the alternative nucleotide or base present in the read instead of the corresponding reference nucleotide or base in the reference array, ○ The six last bits are used to encode the position of the mismatch in the reference array, which is calculated as an offset from the previous mismatch of the read (the relative position of the mismatch, except for the first mismatch of the read for which the absolute position is encoded). Therefore, the range of this offset encoded in six bits is [0 - 63].
[0065] Figure 3 provides an example of the encoding of mismatches of a read by the first encoding sub-process. This read is an incompletely mapped read that is globally mapped with a reference array. This read has two mismatches, ○ a first mismatch located at the 12th position in the read, which is a substitution of an A nucleotide in the reference array by a T nucleotide in the read, and ○ a second mismatch located at the 21st position in the read, which is a substitution of a C nucleotide in the reference array by a G nucleotide in the read, having.
[0066] Then, the list of mismatches of the read is ○ <12,T>, i.e., the value "12" corresponding to the absolute position of the first mismatch in the read, and ○ <9,G>, i.e., the value "9" corresponding to the relative position of the second mismatch in the read, i.e., the offset between the second mismatch and the first mismatch, encoded as.
[0067] <12,T> may be converted, for example, to the value "51" (encoded in 1 byte), and <9,G> may be converted to the value "38" (encoded in 1 byte). Such 1-byte encoding is offset position × 4 + nucleotide value (A = 0, C = 1, G = 2, T = 3) obtained by
[0068] Preferably, for each incompletely mapped read encoded by the first encoding sub-process or the second encoding sub-process, if the offset calculated between a given mismatch of the read and a preceding mismatch is greater than the maximum encodable value, at least one "false" mismatch is inserted between the two mismatches until every offset between each of the two mismatches and at least one "false" mismatch is lower than the maximum encodable value. A "false" mismatch is defined as a mismatch where the bits of the byte used to encode the mismatch encode a nucleotide or base equal to the corresponding reference nucleotide or base in the reference sequence. In a preferred embodiment, although not to be construed as limiting the scope of the present invention, the maximum encodable value is equal to 63 and corresponds to the maximum value encodable in 6 bits.
[0069] Figure 4 provides an example of the encoding of read mismatches by the first encoding sub-process in a case where a "false" mismatch needs to be inserted. This read is an incompletely mapped read that is globally mapped in the reference sequence. This read has two mismatches, ○ a first mismatch located at position 22 in the read, which is a substitution of an A nucleotide in the reference sequence by a T nucleotide in the read, and ○ a second mismatch located at position 134 in the read, which is a substitution of a C nucleotide in the reference sequence by a G nucleotide in the read. has
[0070] The positional offset between the second mismatch and the first mismatch is 112, which is greater than the maximum encodable value of 63. Therefore, a "false" mismatch needs to be inserted between the two mismatches, such that any offset between each of the mismatches and the "false" mismatch is less than the maximum encodable value. The "false" mismatch using a T nucleotide (corresponding to the "actual" T nucleotide in the reference sequence) is inserted, for example, at the 85th position within the read. The positional offset calculated between the "false" mismatch and the first mismatch is 63, which corresponds to the maximum encodable value. The positional offset calculated between the second mismatch and the "false" mismatch is 49, which is less than 63.
[0071] The list of mismatches for the read is then ○ <22,T>, i.e., the value "22" corresponding to the absolute position of the first mismatch within the read, ○ <63,T>, i.e., the value "63" corresponding to the relative position of the "false" mismatch within the read, i.e., the offset between the "false" mismatch and the first mismatch, and ○ <49,G>, i.e., the value "49" corresponding to the relative position of the second mismatch within the read, i.e., the offset between the second mismatch and the "false" mismatch, encoded as such.
[0072] <22,T> may be converted, for example, to the value "91" (encoded in 1 byte), <63,T> may be converted to the value "255" (encoded in 1 byte), and <49,G> may be converted to the value "198" (encoded in 1 byte). Such 1-byte encoding is obtained by offset position × 4 + nucleotide value (A = 0, C = 1, G = 2, T = 3).
[0073] This method includes a final step 14 of providing a compressed file that includes a list of encoded reads. The encoded reads are stored in the compressed file in the same order as they are stored in the initial uncompressed file. Each read can then be reconstructed from the alignment encoding information and the reference sequence by appropriate decompression software and / or methods configured according to the present invention.
[0074] Although described with respect to an exemplary architecture of a computing device 20 (shown in FIG. 2 for illustrative purposes), the techniques of the present invention disclosed herein may be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, the computer program code may be stored in a computer medium and executed by a hardware processing unit including one or more processors, as in the case of using the device 20 of FIG. 2. As used herein, the term "processor" is intended to include one or more processing devices including a signal processor, a microprocessor, a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other type of processing circuit, and portions or combinations of such circuit elements. Also, as used herein, the term "memory" is intended to include any combination of electronic memories associated with a processor, such as random access memory (RAM), read-only memory (ROM), or other types of memory.
[0075] Accordingly, software instructions or code for performing the methodologies and protocols described herein may be stored in one or more associated memory devices, such as one of ROM, fixed or removable memory, and loaded into RAM and executed by a processor when ready for use.
[0076] The technology of the present disclosure can be implemented in a wide variety of devices or apparatuses, including, for example, mobile phones, computers, servers, tablets, and similar devices.
[0077] Exemplary embodiments of the present invention have been described herein with reference to the accompanying drawings, but the present invention is not limited to the exact embodiments shown in the drawings, and various other changes and modifications can be made by those skilled in the art without departing from the scope or spirit of the present invention.
[0078] Statistical and numerical examples of the compression method according to the present invention The following comparative examples were performed on an uncompressed data file containing 48 million reads or sequences of nucleotides. ○ Size of the uncompressed data file: 35,770 MB (megabytes) ○ Size of the file compressed with gzip software: 6,649 MB ○ Size of the file compressed with non-reference-based SPRING software: 1,402 MB ○ Size of the file compressed with the reference-based compression method according to the present invention: 1,179 MB ○ Compression time with non-reference-based SPRING software: 1,722 seconds ○ Compression time with the reference-based compression method according to the present invention: 181 seconds ○ Average size in bits / nucleotide of the uncompressed data file (ASCII encoded): 8 bits / nucleotide ○ Average size in bits / nucleotide of the file compressed with a coding adapted to the four possible characters A, T, C, G: 2 bits / nucleotide ○ Average size in bits / nucleotide of the file compressed with the reference-based compression method according to the present invention: 0.33 bits / nucleotide
[0079] The numerical examples shown above demonstrate that the present invention enables fast compression and decompression while providing a high compression ratio.
Claims
1. 1. A computer-implemented method for compression of genomic sequence data generated by a sequencing machine, the genomic sequence data including nucleotide or base sequence reads aligned to a reference sequence, thereby creating aligned reads, the aligned reads being stored as a list of reads in an initial file, the method comprising: For each aligned read, determining whether the read is completely or incompletely mapped to the reference sequence, or whether the read is not mapped to the reference sequence; encoding the reads according to the determination, wherein the reads determined to be fully mapped are encoded by a first encoding process and the reads determined to be unmapped are encoded by a second encoding process; The determining step comprises, for each incompletely mapped read, comparing a number of mismatches between the read and the reference sequence to a threshold; In the encoding step, the reads determined to be incompletely mapped are encoded by the second encoding process or a third encoding process, the incompletely mapped reads are encoded by the second encoding process when the number of mismatches is greater than the threshold, and the incompletely mapped reads are encoded by the third encoding process when the number of mismatches is less than the threshold; In the second encoding process, each nucleotide or base of the read is encoded individually; A method according to the present invention, wherein the first encoding process and the third encoding process include separate sets of descriptors, each set of descriptors uniquely representing a lead associated with the corresponding encoding process, and each of the first encoding process and the third encoding process is an information source entropy reduction encoding process.
2. 2. The method of claim 1, wherein the determining step comprises a further determination as to whether the read is globally or locally mapped with the reference sequence when the read is determined to be incompletely mapped with the reference sequence and has a number of mismatches less than the threshold, the third encoding process comprises a first encoding sub-process and a second encoding sub-process, the reads determined to be globally mapped are encoded by the first encoding sub-process and the reads determined to be locally mapped are encoded by the second encoding sub-process, the first encoding sub-process and the second encoding sub-process comprise separate sets of descriptors, each set of descriptors uniquely representing the reads associated with the corresponding encoding sub-process.
3. 3. The method of claim 2, wherein the descriptors of the first encoding sub-process include an alignment start position in the reference sequence, a read length, and a list of mismatches due to symbol substitution, and the descriptors of the second encoding sub-process include a local alignment start position in the reference sequence, a read length, a list of mismatches due to symbol substitution, and a length of a clipped portion of the read that is not part of an alignment.
4. 4. The method of claim 3, wherein in the encoding step, the clipped portions of the reads to be encoded by the second encoding sub-process are concatenated and each nucleotide or base of the clipped portions is encoded individually.
5. The method of any one of claims 1 to 4, wherein in the encoding step, each mismatch of an incompletely mapped read is encoded into one byte.
6. In the encoding step, each mismatch of an incompletely mapped read is the two first bits of said byte are used to encode an alternative nucleotide or base present in said read instead of a corresponding reference nucleotide or base in said reference sequence; The method of claim 5, wherein the six last bits of the byte are used to encode the position of the mismatch in the reference sequence, the position being calculated and encoded as an offset from the preceding mismatch of the read.
7. 7. The method of claim 6, wherein in the encoding step, if the offset calculated between a given mismatch and the previous mismatch is greater than the maximum codeable value, at least one false mismatch is inserted between the two mismatches until every offset between each of the mismatches and the at least one false mismatch is less than the maximum codeable value, a false mismatch being defined as a mismatch in which a bit of the one byte is used to code the mismatch or to code a nucleotide or base that is equal to the corresponding reference nucleotide or reference base in the reference sequence.
8. 8. The method according to claim 1, further comprising an initial step of dividing the list of reads into blocks of reads, each block beginning with a header containing information needed to decrypt the block, and wherein the compression method is performed block-by-block.
9. The method of claim 8 , wherein the blocks of reads have the same block size.
10. 10. The method of any one of claims 1 to 9, further comprising a final step of providing a compressed file comprising a list of encoded reads, said encoded reads being stored in said compressed file in the same order as the order of the reads stored in said initial file.
11. The method according to any one of claims 1 to 10, wherein said threshold value is equal to 31.
12. 12. The method of any one of claims 1 to 11, further comprising the step of determining, for each aligned read, whether the read contains at least one mismatch corresponding to a case where the sequencing machine was unable to call any base or nucleotide.
13. 13. The method of claim 12, further comprising: for each read containing at least one mismatch corresponding to a case where the sequencing machine was unable to call any base or nucleotide, determining the number of such mismatches and comparing said number to a reference threshold.
14. 14. The method of claim 13, wherein in the encoding step, if the number of such mismatches is greater than the reference threshold, each nucleotide or base of the read to be encoded by the second encoding process is individually coded into 4 bits, and if the number of such mismatches is less than the reference threshold, each nucleotide or base of the read to be encoded by the second encoding process is individually coded into 2 bits, and the encoding step further comprises encoding a list of positions along the reference sequence, the positions corresponding to the positions of such mismatches in the reference sequence.
15. A computer program product embodied in a computer readable storage medium, the computer program product comprising computer executable instructions which, when executed by a processor, cause the processor to perform operations including the steps of the method of any one of claims 1 to 14.
16. A computer-readable storage medium having computer-executable instructions which, when executed by a processor, cause the processor to perform operations including the steps of the method according to any one of claims 1 to 14.
17. 1. An apparatus comprising: A processor; An apparatus comprising: a memory operatively coupled to the processor to form a computing device, the memory storing processor-executable instructions that, when executed on at least the processor, cause the processor to perform operations including the steps of the method of claim 1.
18. 1. A method for compressing genomic sequence data, the method comprising: obtaining, by one or more computers, lead records; determining, by the one or more computers, whether the read records correspond to reads that are completely mapped to a reference sequence or that are incompletely mapped to the reference sequence; determining, by the one or more computers, whether a number of mismatches of the incompletely mapped reads meets a predetermined mismatch threshold number based on determining, by the one or more computers, that the read records correspond to reads that are incompletely mapped to the reference sequence; and encoding, by the one or more computers, each mismatch of the incompletely mapped read into a record having a size of one byte based on determining that the number of mismatches meets the predetermined mismatch threshold number.
19. Determining, by the one or more computers, whether the number of mismatches of the incompletely mapped reads meets a predetermined mismatch threshold number, 20. The method of claim 18, comprising determining, by the one or more computers, whether the number of mismatches in the incompletely mapped reads is greater than the predetermined threshold number of mismatches.
20. Each lead record is Data indicating the absolute start position of the aligned reads with respect to the reference sequence; Data indicating the length of the read; data indicating whether the read is completely or incompletely mapped; data indicating the number of mismatches identified in the reads; and data indicating the relative position of each of the possible mismatches in the read.
21. Encoding each mismatch of the incompletely mapped read into a record having a size of 1 byte includes, for each particular mismatch: encoding, by the one or more computers, the first two bits of the byte to contain data indicative of an alternative nucleotide or base present in the read in place of a corresponding reference nucleotide or base in the reference sequence; and encoding, by one or more computers, the remaining 6 bits of the byte to contain data indicating the position of a mismatch in the reference sequence, the position being calculated as an offset from a preceding mismatch in the read.
22. The method comprises: determining, by one or more computers, whether the offset is greater than a maximum codable value; 22. The method of claim 21, further comprising: inserting, by one or more computers, at least one false mismatch between the particular mismatch and the previous mismatch based on determining that the offset is greater than the maximum encoded value.
23. The method comprises: The method of claim 18, further comprising encoding, by one or more computers, a list of positions of the reference sequence corresponding to each position of the mismatch relative to the reference sequence using an information entropy reduction encoding process based on determining that the number of mismatches does not meet the predetermined mismatch threshold number.
24. The method comprises: The method of claim 18, further comprising encoding, by one or more computers, at least a portion of the read record using information entropy reduction encoding based on determining that the read record corresponds to a read that is completely mapped to the reference sequence.
25. The method of claim 18 , wherein the one or more computers comprise one or more hardware processors.
26. 26. The method of claim 25, wherein the one or more hardware processors include one or more field programmable gate arrays (FPGAs).
27. 1. A hardware processor comprising hardware processing circuitry configured to perform one or more operations, the one or more operations comprising: obtaining a lead record by the hardware processing circuit; determining, by the hardware processing circuitry, whether the read record corresponds to a read that is completely mapped to a reference sequence or that is incompletely mapped to the reference sequence; determining, by the one or more computers, whether a number of mismatches of the incompletely mapped reads meets a predetermined mismatch threshold number based on the hardware processing circuit determining that the read record corresponds to a read that is incompletely mapped to the reference sequence; and encoding, by the hardware processing circuitry, each mismatch of the incompletely mapped read into a record having a size of 1 bit based on determining that the number of mismatches meets the predetermined mismatch threshold number.
28. Each lead record is Data indicating the absolute start position of the aligned reads with respect to the reference sequence; Data indicating the length of the read; data indicating whether the read is completely or incompletely mapped; data indicating the number of mismatches identified in the reads; and data indicative of a relative location of the possible mismatch in the read.
29. Encoding each mismatch of the incompletely mapped read into a record having a size of 1 byte includes, for each particular mismatch: encoding, by the hardware processing circuitry, the first two bits of the byte to contain data indicative of an alternative nucleotide or base present in the read in place of a corresponding reference nucleotide or base in the reference sequence; and encoding, by the hardware processing circuitry, the remaining 6 bits of the byte to contain data indicating a position of a mismatch in the reference sequence, the position being calculated as an offset from a previous mismatch in the read.
30. The hardware processor circuit includes: determining, by the hardware processing circuitry, whether the offset is greater than a maximum codable value; 30. The hardware processor of claim 29, further configured to perform an operation including: inserting, by the hardware processing circuitry, at least one false mismatch between the particular mismatch and the previous mismatch based on determining that the offset is greater than the maximum encoded value.
31. The hardware processor circuit includes:
28. The hardware processor of claim 27, further configured to perform an operation, based on determining that the number of mismatches does not meet the predetermined mismatch threshold number, including: encoding, by the hardware processing circuitry, a list of positions of the reference sequence that correspond to each position of the mismatches relative to the reference sequence using an information entropy reduction encoding process.
32. The hardware processor circuit includes: The hardware processor of claim 27, further configured to perform, based on determining that the read record corresponds to a read that is fully mapped to the reference sequence, an operation including encoding, by the hardware processing circuitry, at least a portion of the read record using information entropy reduction encoding.
33. 25. The hardware processor of claim 24, wherein the hardware processing circuitry comprises one or more field programmable gate arrays (FPGAs).
34. Determining, by the hardware processing circuitry, whether a number of mismatches of the incompletely mapped reads meets a predetermined mismatch threshold number, 20. The hardware processor of claim 18, further comprising determining, by the hardware processing circuitry, whether the number of mismatches of the incompletely mapped reads is greater than the predetermined mismatch threshold number.
35. 1. A system for compressing genomic sequence data, the system comprising: one or more computers; and one or more storage devices that store instructions, which when executed by the one or more computers, cause the one or more computers to: obtaining, by the one or more computers, a lead record; determining, by the one or more computers, whether the read records correspond to reads that are completely mapped to a reference sequence or that are incompletely mapped to the reference sequence; determining, by the one or more computers, whether a number of mismatches of the incompletely mapped reads meets a predetermined mismatch threshold number based on determining, by the one or more computers, that the read records correspond to reads that are incompletely mapped to the reference sequence; The system is operable to cause the one or more computers to perform an operation including: encoding, based on determining that the number of mismatches meets the predetermined mismatch threshold number, each mismatch of the incompletely mapped read into a record having a size of one byte.
36. Each lead record is Data indicating the absolute start position of the aligned reads with respect to the reference sequence; Data indicating the length of the read; data indicating whether the read is completely or incompletely mapped; data indicating the number of mismatches identified in the reads; and data indicating a relative position of each of the possible mismatches in the read.
37. Encoding each mismatch of the incompletely mapped read into a record having a size of 1 byte includes, for each particular mismatch: encoding, by one or more computers, the first two bits of said byte to contain data indicative of an alternative nucleotide or base present in said read in place of a corresponding reference nucleotide or base in said reference sequence; and encoding, by one or more computers, the remaining 6 bits of the byte to contain data indicating the position of a mismatch in the reference sequence, the position being calculated as an offset from a preceding mismatch in the read.
38. The calculation is determining, by the one or more computers, whether the offset is greater than a maximum codable value; 38. The system of claim 37, further comprising: inserting, by one or more computers, at least one false mismatch between the particular mismatch and the preceding mismatch based on determining that the offset is greater than the maximum encoded value.
39. The calculation is The system of claim 35, further comprising, based on determining that the number of mismatches does not meet the predetermined mismatch threshold number, encoding, by one or more computers, a list of positions of the reference sequence corresponding to each position of the mismatch relative to the reference sequence using an information entropy reduction encoding process.
40. The calculation is The system of claim 35, further comprising encoding, by one or more computers, at least a portion of the read record using information entropy reduction encoding based on determining that the read record corresponds to a read that is completely mapped to the reference sequence.
41. 36. The system of claim 35, wherein the one or more computers include one or more hardware processors.
42. 42. The system of claim 41, wherein the one or more hardware processors include one or more field programmable gate arrays (FPGAs).
43. 1. A computer readable storage device having stored thereon instructions that, when executed by a data processing apparatus, cause the data processing apparatus to perform an operation for compressing genomic sequence data, the operation comprising: Obtaining a lead record; determining whether the read record corresponds to a read that is completely mapped to a reference sequence or that is incompletely mapped to the reference sequence; determining whether a number of mismatches of the incompletely mapped reads meets a predetermined mismatch threshold number based on determining that the read record corresponds to a read that is incompletely mapped to the reference sequence; and encoding each mismatch of the incompletely mapped read into a record having a size of 1 byte based on determining that the number of mismatches meets the predetermined mismatch threshold number.
44. Each lead record is Data indicating the absolute start position of the aligned reads with respect to the reference sequence; Data indicating the length of the read; data indicating whether the read is completely or incompletely mapped; data indicating the number of mismatches identified in the reads; and data indicative of a relative position of each of the possible mismatches in the read.
45. Encoding each mismatch of the incompletely mapped read into a record having a size of 1 byte includes, for each particular mismatch: encoding, by one or more computers, the first two bits of said byte to contain data indicative of an alternative nucleotide or base present in said read in place of a corresponding reference nucleotide or base in said reference sequence; and encoding, by one or more computers, the remaining 6 bits of the byte to contain data indicating a position of a mismatch in the reference sequence, the position calculated as an offset from a preceding mismatch in the read.
46. The calculation is determining, by the one or more computers, whether the offset is greater than a maximum codable value; 46. The computer-readable storage device of claim 45, further comprising: inserting, by the one or more computers, at least one false mismatch between the particular mismatch and the previous mismatch based on determining that the offset is greater than the maximum encoded value.
47. The calculation is 44. The computer-readable storage device of claim 43, further comprising: encoding a list of positions of the reference sequence corresponding to each position of the mismatches relative to the reference sequence using an information entropy reduction encoding process based on determining that the number of mismatches does not meet the predetermined mismatch threshold number.
48. The calculation is 44. The computer-readable storage device of claim 43, further comprising encoding at least a portion of the read record using information entropy reduction encoding based on determining that the read record corresponds to a read that is fully mapped to the reference sequence.
Citation Information
Patent Citations
Lossless compression of DNA sequences
US20150227686A1
Method and apparatus for compact representation of bioinformatics data
WO2018068829A1
Method and system for selective access of stored or transmitted bioinformatics data
WO2018071054A1
Method and apparatus for the access to bioinformatics data structured in access units
WO2018071078A1