Methods for compressing genomic sequence data

After comparing the genome sequencing data, the compression is carried out using different encoding methods according to the mapping between the read segment and the reference sequence, which solves the problems of low compression ratio and slow speed in the prior art, and achieves efficient and lossless data compression.

CN114402314BActive Publication Date: 2025-05-09ILLUMINA INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202080062683.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-11
Filing Date
2020-09-11
Publication Date
2025-05-09
Estimated Expiration
2040-09-11

AI Technical Summary

Technical Problem

The prior art when compressing genome sequencing data, the compression ratio is low, the speed is slow, and may lead to information loss, affecting the reliability of subsequent analysis.

Method used

By comparing the mapping between the read segment and the reference sequence, different encoding processes are used to encode the fully mapped, incomplete mapped and unmasked read segments, and the encoding method is determined using thresholds, so as to achieve rapid compression and decompression and maintain the initial order of the read segments.

Benefits of technology

It achieves a high compression ratio while maintaining data integrity and rapid compression and decompression process, which is suitable for the storage and analysis needs of massive genome sequencing data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114402314B_ABST
    Figure CN114402314B_ABST
Patent Text Reader

Abstract

The present invention relates to a reference-based method for compressing genomic sequence data generated by a sequencing machine. It is determined whether a sequence of nucleotides or bases that has been previously aligned with a reference sequence is completely mapped, incompletely mapped, or unmapped to the reference sequence; and then encoding is performed according to the determination. The determination step includes: for each incompletely mapped sequence, comparing the number of mismatches between the sequence and the reference sequence with a reference threshold, and encoding the incompletely mapped sequence according to different encoding processes based on the results of the comparison method for compressing the genomic sequence data generated by the sequencing machine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention generally relates to methods of representing genomic sequencing data generated by a sequencing machine, and more specifically to computer-implemented methods for compressing such genomic sequencing data. The present disclosure provides a reference-based compression method that allows fast compression and decompression without causing information loss and with a high compression ratio. Background Art

[0002] Next-generation sequencing machines now generate massive amounts of sequencing data at an affordable price. The latest systems generate more than 6 billion 150-nucleotide-long sequences in a single run of 36 hours, enough to sequence 20 complete human genomes. This opens up many new perspectives for the diagnosis of genetic diseases and the development of personalized medicine aimed at adapting treatments based on the specificity of the human genome.

[0003] However, this also brings new challenges, especially the costs associated with storing massive amounts of data. The most commonly used file format for raw (unaligned) sequence data is the FASTQ format, which saves sequence data (strings of A, C, T, G nucleotides, also called reads), quality values ​​(the probability that the sequencing platform will make a sequencing error for each nucleotide), and sequence names. This is a plain ASCII text file, usually compressed with the general text compression scheme LZ (Lempel-Ziv scheme, implemented in gzip software). However, using such compression methods brings several problems:

[0004] -Low compression ratio due to incomplete use of data redundancy

[0005] - Slow compression and decompression

[0006] There are also compression methods dedicated to FASTQ encoding, which are divided into reference-based methods or non-reference-based methods. However, none of them is completely satisfactory because: a) reference-based methods have good compression ratios, but are slow, and b) non-reference-based methods are faster, but have lower compression ratios. An example of such a non-reference-based method is provided by the software SPRING, which is a reference-free compressor for FASTQ files (World Wide Web address: github.com / shubhamchandak94 / SPRING). However, the compression ratio of the compression method provided by the software SPRING is low.

[0007] In reference-based compression methods, some methods that use sequence alignment and are intended to have faster speeds and good compression ratios have been proposed. However, such methods encounter several problems, one of the main problems worth noting is that they are not completely lossless. This known reference-based compression method is described, for example, in patent document WO 2018 / 068829 A1. In the described method, after alignment with one or more reference sequences, nucleotide sequences are classified according to the matching accuracy (thereby creating categories of aligned reads), and then these nucleotide sequences are encoded into multi-layer syntax elements for each layer in which the data is partitioned using different source models and entropy encoders. Therefore, the categories of data are encoded separately and constructed in different layers of syntax elements, each layer containing descriptors that univocally represent the classified and aligned reads of the layer. The method aims to obtain different information sources with simplified information entropy, thereby allowing improved compression performance and selective access to specific categories of compressed data. However, this compression method reorders the reads in an order different from the order obtained at the end of the read alignment step (i.e., the reads are reordered according to their categories). Some information is then lost during the compression process (especially the initial sequence ordering). As a result, the reproducibility of some analytical results may be affected, as some downstream analysis software may rely on the order of the reads. In addition, decompressing the data in an order different from the original order of the reads makes it more difficult to check whether the uncompressed file is identical to the original file. In addition, this compression method is relatively slow, especially when compared to state-of-the-art non-reference-based compression methods. Summary of the invention

[0008] The features of the following independent claims solve the problems of the prior art solutions by providing a method for compressing genomic sequence data. In one aspect, a computer-implemented method for compressing genomic sequence data generated by a sequencing machine is disclosed, the genomic sequence data comprising reads of sequences of nucleotides or bases that have been aligned to a reference sequence, thereby generating aligned reads, the aligned reads being stored as a read list in an initial file, the method comprising the following steps:

[0009] - for each aligned read, determining whether the read is completely or incompletely mapped to the reference sequence, or whether the read is unmapped to the reference sequence,

[0010] - encoding the reads according to the determination, wherein the reads determined to be fully mapped are encoded according to a first encoding process and the reads determined to be unmapped are encoded according to a second encoding process,

[0011] - wherein the determining step comprises, for each incompletely mapped read, comparing the number of mismatches between the read and the reference sequence to a threshold value,

[0012] - wherein, in the encoding step, the read segments determined to be incompletely mapped are encoded according to the second encoding process or the third encoding process, when the number of mismatches is greater than the threshold, the incompletely mapped read segments are encoded according to the second encoding process, and when the number of mismatches is lower than the threshold, the incompletely mapped read segments are encoded according to the third encoding process,

[0013] - wherein, during said second encoding process, each nucleotide or base of said read segment is encoded individually,

[0014] - wherein the first encoding process and the third encoding process include different sets of descriptors, each descriptor set univocally representing a read segment associated with the corresponding encoding process, each of the first encoding process and the third encoding process is a simplified information source entropy encoding process.

[0015] The present invention overcomes the shortcomings of existing compression methods by allowing fast compression and decompression without causing information loss and providing a high compression ratio. More specifically, the focus of the present invention is to encode the most frequently occurring cases in the most compact way, even if this means using a degraded encoding mode for the rare least frequently occurring cases. This leads to a huge improvement in compression performance. In addition, due to the genomic information representation format used in the present invention, the compression performed by the method according to the present invention is faster. Last but not least, the method according to the present invention maintains the initial order of the reads as such and does not reorder the reads according to the category of the reads. Therefore, no information is lost during the process, which makes it easier to perform downstream analysis, as well as effective consistency checks after the decompression step.

[0016] These and other features and advantages of the present invention will become more apparent from the accompanying drawings and the following detailed description. In addition, although thresholds may be referred to herein as being exceeded or not exceeded, it should be understood that such thresholds may be conceptually employed to determine whether such thresholds are met, met, or otherwise detected, regardless of whether the numbers or values ​​used to achieve those threshold evaluations are described using positive or negative values.

[0017] According to an innovative aspect of the present disclosure, a method for compressing genomic sequence data is disclosed. In one aspect, the method may include performing one or more operations by executing software instructions through one or more computers, wherein the operations include: obtaining a read record by the one or more computers; determining by the one or more computers whether the read record corresponds to a read that is completely mapped to a reference sequence or a read that is incompletely mapped to the reference sequence; based on the determination by the one or more computers that the read record corresponds to a read that is incompletely mapped to the reference sequence, determining by the one or more computers whether the number of mismatches of the incompletely mapped reads meets a predetermined mismatch threshold number, and based on determining that the number of mismatches meets the predetermined mismatch threshold number, encoding each mismatch of the incompletely mapped reads into a record having a size of 1 byte by the one or more computers.

[0018] Other aspects include corresponding systems, apparatus, and computer programs that perform the actions of the methods as disclosed herein, as defined by instructions encoded on a computer-readable storage device.

[0019] These and other versions may optionally include one or more of the following features. For example, in some implementations, determining by the one or more computers whether the number of mismatches of the incompletely mapped reads satisfies a predetermined threshold number of mismatches may include determining by the one or more computers whether the number of mismatches of the incompletely mapped reads is greater than the predetermined threshold number of mismatches.

[0020] In some embodiments, each read record may include: data indicating the absolute starting position of the aligned read relative to the reference sequence; data indicating the length of the read; data indicating whether the read is fully mapped or incompletely mapped; data indicating the number of mismatches identified in the read; and data indicating the relative position of each of the possible mismatches in the read.

[0021] In some embodiments, encoding each mismatch of the incompletely mapped read segment as a record having a size of 1 byte includes: for each specific mismatch, encoding the first two bits of the byte by the one or more computers to include data indicating an alternative nucleotide or base present in the read segment instead of the corresponding reference nucleotide or base in the reference sequence; and encoding the remaining six bits of the byte by the one or more computers to include data indicating the position of the mismatch in the reference sequence, wherein the position is calculated as an offset relative to a previous mismatch of the read segment.

[0022] In some specific implementations, the method may also include: determining, by one or more computers, whether the offset is greater than a maximum encodable value; and inserting, by one or more computers, at least one false mismatch between the specific mismatch and the previous mismatch based on determining that the offset is greater than the maximum encodable value.

[0023] In some specific implementations, the method may also include: based on determining that the number of mismatches does not meet the predetermined mismatch threshold number, one or more computers use a simplified information entropy coding process to encode a list of positions of the reference sequence corresponding to the position of each of the mismatches into the reference sequence.

[0024] In some implementations, the method may further include: encoding, by one or more computers, at least a portion of the read record using simplified information entropy coding based on determining that the read record corresponds to a read that is fully mapped to the reference sequence.

[0025] In some implementations, the one or more computers can include one or more hardware processors.

[0026] In some embodiments, the one or more hardware processors may include one or more field programmable gate arrays (FPGAs).

[0027] In some specific implementations, the method for compressing genomic sequence data can be performed by one or more hardware processors. In such specific implementations, the hardware processor may include a hardware processing circuit system configured to perform one or more operations. In one aspect, the operation may include: obtaining a read record by the hardware processing circuit system; determining by the hardware processing circuit system whether the read record corresponds to a read that is completely mapped to a reference sequence or a read that is incompletely mapped to the reference sequence; based on the determination by the hardware processing circuit system that the read record corresponds to a read that is incompletely mapped to the reference sequence, determining by the one or more computers whether the number of mismatches of the incompletely mapped reads meets a predetermined mismatch threshold number, and based on determining that the number of mismatches meets the predetermined mismatch threshold number, encoding each mismatch of the incompletely mapped reads into a record having a size of 1 byte by the hardware processing circuit system.

[0028] These and other features and advantages of the present invention will become more apparent from the accompanying drawings and the following detailed description.

[0029] In some embodiments, each read record may include: data indicating the absolute starting position of the aligned read relative to the reference sequence; data indicating the length of the read; data indicating whether the read is fully mapped or incompletely mapped; data indicating the number of mismatches identified in the read; and data indicating the relative position of the possible mismatches in the read.

[0030] In some specific implementations, determining by the hardware processing circuit system whether the number of mismatches of the incompletely mapped read segments meets a predetermined mismatch threshold number may include determining by the hardware processing circuit system whether the number of mismatches of the incompletely mapped read segments is greater than the predetermined mismatch threshold number.

[0031] In some specific implementations, encoding each mismatch of the incompletely mapped read segment as a record having a size of 1 byte may include: for each specific mismatch, encoding the first two bits of the byte by the hardware processing circuit system to include data indicating an alternative nucleotide or base present in the read segment instead of the corresponding reference nucleotide or base in the reference sequence; and encoding the remaining six bits of the byte by the hardware processing circuit system to include data indicating the position of the mismatch in the reference sequence, the position being calculated as an offset relative to a previous mismatch of the read segment.

[0032] In some specific implementations, the hardware processor circuit system is further configured to perform the following operations, which include: determining by the hardware processing circuit system whether the offset is greater than a maximum encodable value; and based on determining that the offset is greater than the maximum encodable value, inserting by the hardware processing circuit system at least one false mismatch between the specific mismatch and the previous mismatch.

[0033] In some specific implementations, the hardware processor circuit system is further configured to perform the following operations, which include: based on determining that the number of mismatches does not meet the predetermined mismatch threshold number, the hardware processing circuit system uses a simplified information entropy coding process to encode a list of positions of the reference sequence corresponding to the position of each of the mismatches into the reference sequence.

[0034] In some specific implementations, the hardware processor circuit system is further configured to perform the following operations, which include: based on determining that the read segment record corresponds to a read segment that is completely mapped to the reference sequence, the hardware processing circuit system encodes at least a portion of the read segment record using simplified information entropy coding.

[0035] In some implementations, the hardware processing circuitry includes one or more field programmable gate arrays (FPGAs).

[0036] According to another innovative aspect of the present disclosure, a computer-implemented method for compressing genomic sequence data generated by a sequencing machine is disclosed, the genomic sequence data including reads of sequences of nucleotides or bases that have been aligned to a reference sequence, thereby generating aligned reads, the aligned reads being stored as a read list in an initial file. In one aspect, the method may include the following actions for each aligned read: determining whether the read is fully mapped or incompletely mapped to the reference sequence, or whether the read is unmapped to the reference sequence; encoding the read according to the determination, wherein the read determined to be fully mapped is encoded according to a first encoding process, and the read determined to be unmapped is encoded according to a second encoding process, wherein the determination step includes, for each incompletely mapped read, comparing the number of mismatches between the read and the reference sequence with a threshold, wherein in the encoding step, the read is encoded according to the second encoding process or the third encoding process. The read segment determined to be incompletely mapped is encoded, when the number of mismatches is greater than the threshold, the incompletely mapped read segment is encoded according to the second encoding process, and when the number of mismatches is less than the threshold, the incompletely mapped read segment is encoded according to the third encoding process, wherein in the second encoding process, each nucleotide or base of the read segment is encoded separately, wherein the first encoding process and the third encoding process include different descriptor sets, each descriptor set univocally represents the read segment associated with the corresponding encoding process, and each of the first encoding process and the third encoding process is a simplified information source entropy encoding process.

[0037] Other aspects include corresponding systems, apparatus, and computer programs that perform the actions of the methods as disclosed herein, as defined by instructions encoded on a computer-readable storage device.

[0038] These and other versions may optionally include one or more of the following features. For example, in some implementations, the determining step may include a further determination when it is determined that the read is not completely mapped to the reference sequence and has a number of mismatches below the threshold, the further determination being about whether the read is globally mapped or locally mapped to the reference sequence, and wherein the third encoding process includes a first encoding subprocess and a second encoding subprocess, according to which the read determined to be globally mapped is encoded, and according to which the read determined to be locally mapped is encoded, the first encoding subprocess and the second encoding subprocess include different descriptor sets, each descriptor set univocally representing the read associated with the corresponding encoding subprocess.

[0039] In some specific implementations, the descriptor of the first encoding sub-process may include the alignment start position in the reference sequence, the read segment length, and a mismatch list represented by symbol replacement, and wherein the descriptor of the second encoding sub-process includes the local alignment start position in the reference sequence, the read segment length, a mismatch list represented by symbol replacement, and the length of the clipped portion of the read segment that is not part of the alignment.

[0040] In some specific implementations, in the encoding step, the clipped parts of the read segment to be encoded according to the second encoding sub-process are concatenated, and each nucleotide or base of the clipped part is encoded separately.

[0041] In some implementations, in the encoding step, each mismatch of an incompletely mapped read segment is encoded on 1 byte.

[0042] In some embodiments, in the encoding step, each mismatch of the incompletely mapped read is encoded as follows: the first two bits of the byte are used to encode the alternative nucleotide or base present in the read instead of the corresponding reference nucleotide or base in the reference sequence, and the last six bits of the byte are used to encode the position of the mismatch in the reference sequence, which is calculated as an offset relative to the previous mismatch of the read.

[0043] In some specific implementations, in the encoding step, if the offset calculated between a given mismatch and the previous mismatch is greater than the maximum encodable value, at least one false mismatch is inserted between the two mismatches until each offset between each of the mismatches and the at least one false mismatch is below the maximum encodable value, a false mismatch being defined as a mismatch for which bits of the byte are used to encode the mismatch, or to encode a nucleotide or base that is equal to a corresponding reference nucleotide or base in the reference sequence.

[0044] In some implementations, an initial step is to divide the read list into read blocks, where each block starts with a header containing information needed to decode the block, and where the compression method is performed block by block.

[0045] In some implementations, the read chunks have the same chunk size.

[0046] In some embodiments, the final step is to provide a compressed file containing a list of encoded reads stored in the compressed file in the same order as the reads were stored in the initial file.

[0047] In some implementations, the threshold is equal to 31.

[0048] In some implementations, for each aligned read, a step is provided of determining whether the read contains at least one mismatch corresponding to a situation where the sequencing machine cannot call any base or nucleotide.

[0049] In some implementations, for reads that each contain at least one mismatch corresponding to a case where the sequencing machine cannot call any base or nucleotide, a step of determining the number of such mismatches is provided, and a step of comparing the number to a reference threshold.

[0050] In some specific implementations, in the encoding step, if the number of such mismatches is greater than the reference threshold, each nucleotide or base of the read to be encoded according to the second encoding process is individually encoded with 4 bits, and if the number of such mismatches is less than the reference threshold, each nucleotide or base of the read to be encoded according to the second encoding process is individually encoded with 2 bits, and the encoding step also includes encoding a list of positions along the reference sequence, which positions correspond to the positions of such mismatches in the reference sequence. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] Figure 1 is a flow chart showing the steps of a compression method according to the present invention.

[0052] Figure 2 is a schematic diagram showing an apparatus for implementing the steps of the compression method according to the present invention.

[0053] Figure 3 A first example of reads globally mapped to a reference sequence is shown.

[0054] Figure 4 A second example of reads globally mapped to a reference sequence is shown where false mismatches must be inserted. DETAILED DESCRIPTION

[0055] The genomic sequences mentioned in the present invention include, for example, but not limited to, nucleotide sequences, deoxyribonucleic acid (DNA) sequences, ribonucleic acid (RNA) sequences, and amino acid sequences. Although the specific embodiments herein describe genomic information in the form of nucleotide sequences in considerable detail, it should be understood that, as understood by those skilled in the art, the compression method according to the present invention can also be implemented for other genomic sequences, although with some changes.

[0056] Genome sequencing information is generated by a sequencing machine in the form of a sequence of nucleotides (or more generally, bases) represented by a string of letters from a limited vocabulary. The minimum vocabulary is represented by five symbols: {A, C, G, T, N}, representing the four types of nucleotides present in DNA, namely adenine, cytosine, guanine and thymine. In RNA, thymine is replaced by uracil (U). N indicates that the sequencing machine cannot detect any base, so the true nature of the position is uncertain.

[0057] The nucleotide sequence produced by the sequencing machine is called a "read". The length of the sequence read can be between tens of nucleotides and thousands of nucleotides. Some technologies produce sequence reads in pairs, where one read comes from one DNA chain and the second read comes from another chain. Throughout this disclosure, a "reference sequence" is any sequence on which a nucleotide or base sequence produced by a sequencing machine is aligned / mapped. An example of such a reference sequence can actually be a reference genome, i.e., a sequence assembled by scientists as a representative example of a collection of genes for a species. However, the reference sequence can also be composed of synthetic sequences, which are conceived to simply improve the compressibility of the reads, given that they are to be further processed. Sequencing machines may introduce errors in sequence reads, and it is worth noting that incorrect symbols (i.e., representing different nucleic acids) may be used to represent nucleic acids or bases that are actually present in the sample being sequenced; this is generally referred to as substitution errors or "mismatches".

[0058] The present invention is a reference-based compression method that receives as input reads of a nucleotide or base sequence, such reads having been previously aligned with a reference sequence, thereby generating aligned reads. The aligned reads are then stored as a list of reads in an initial file. The aligned reads and the manner in which these reads are stored once aligned in the initial file are not critical to the present invention and are not the purpose of this disclosure. Each read is then encoded as a list of positions on a reference sequence and differences from the reference sequence. Each read can then be reconstructed from the alignment encoding information and the reference sequence by appropriate decompression software configured in accordance with the present invention.

[0059] Preferably, the alignment software that processes the reads and aligns them with the reference sequence before providing them as input to the compression software and device does not take into account certain types of errors introduced in the sequence reads, such as insertion errors or deletion errors. An insertion error consists in inserting one or more additional symbols that do not relate to any actually existing nucleic acid in a sequence read. A deletion error consists in missing one or more symbols representing nucleic acids actually present in the sequencing sample from a sequence read. More precisely, in the case of an insertion error or a deletion error in a given sequence read, the alignment software then treats the resulting erroneous nucleic acid as a substitution error, also referred to as a "mismatch". This preference for alignment software configuration allows faster subsequent encoding, particularly providing a better compromise between speed and compression ratio.

[0060] For each read, the alignment software provides a corresponding read record to the compression software and device. Each read record contains at least the following information: the absolute starting position of the aligned read relative to the reference sequence, the length of the read, the alignment type of the read, the number of mismatches identified in the read, and the relative position of the possible mismatches in the read (where appropriate).

[0061] Now refer to Figure 1 The compression method according to the invention is described. Figure 2 20 is executed. The device includes at least one processor 22 and a memory 24 operatively coupled to the processor 22 to form a computing device. The memory 24 may store computer program code or software 26, which contains computer executable instructions that, when executed by the processor 22, cause the processor 22 to perform operations including the steps of the compression method according to the present invention.

[0062] The initial file in which the aligned reads are stored as a read list is stored, for example, in a memory of the device 20. Return Figure 1, the method preferably includes an initial step 2, in which the initial list of aligned reads is divided into read blocks. Typically, the aligned read list is divided into 50,000 read blocks, and this specific value is not interpreted as limiting the scope of the present invention that can be applied in the same manner with other values. Preferably, the read blocks have the same block size. Each read block begins with a header containing the information required to decode the block, such as the size in bytes of the content of the block, and / or an identifier of the block or its content, and / or the number of reads contained in the block. This allows support for the concatenation of compressed files, as well as streaming capabilities (each read block contains all the information required to decode the reads of the block). In addition, since the compression method can then be performed block by block, this also allows multi-threading of the read blocks, thereby allowing parallelization and some resulting gains in processing time. If all reads of a given block have the same length, the read length is also stored in the header, otherwise a list of each read length is explicitly stored during the compression method.

[0063] Each read record contains information about the type of read alignment. In general, two main alignment types can be identified: complete alignment and incomplete alignment, plus an additional type corresponding to "unmapped" reads. "Incomplete alignment" means that the read contains at least one non-N mismatch, while at least a portion of the read matches a portion of the reference sequence (according to this definition, an incompletely mapped read may contain one or more Ns, provided that it also contains one or more other mismatches). In an exemplary embodiment, each read record begins with the following bit markers, each of which has a value between two possible values:

[0064] - a first marker which indicates the forward or reverse orientation relative to the reference sequence,

[0065] - a second flag indicating a complete alignment or an incomplete alignment,

[0066] - a third bit flag indicating whether the read segment contains at least one N,

[0067] - The fourth bit flag indicates whether the position information is encoded in 16 or 32 bits.

[0068] The following steps 4-12 are performed block by read, and within a block read by read.

[0069] The method includes a next step 4, which determines, for each aligned read, whether the read is completely mapped or incompletely mapped with the reference sequence, or whether the read is unmapped with the reference sequence. This determination step 4 includes, for each incompletely mapped read, comparing the number of mismatches between the read and the reference sequence with a threshold 4A. In a preferred embodiment, although it should not be interpreted as limiting the scope of the invention, the threshold is equal to 31. This particular value is deliberately selected so as to provide the best possible compromise for storing the number of mismatches in a sufficiently compact manner, as will be better understood later with respect to step 12. In fact, it has been statistically observed that in the vast majority of cases, the incompletely mapped reads have less than 31 mismatches. The principle behind this selection is to encode the most frequently occurring situations in the most compact manner, leaving some very few degraded situations. If the read is determined to be incompletely mapped with a number of mismatches below the threshold, the determination step 4 includes a further determination as to whether the read is globally mapped or locally mapped with the reference sequence. A "globally mapped read" is an incompletely mapped read whose entire sequence (including the start and end points of the read) is incompletely mapped to the reference sequence. A "locally mapped read" is an incompletely mapped read containing a segment of nucleotides or bases that is incompletely mapped to the reference sequence. Thus, the segment of nucleotides or bases corresponds to a portion of the initial read.

[0070] Preferably, the method further comprises a step 6 of determining, for each aligned read, whether the read comprises at least one N, i.e., whether the read comprises at least one mismatch corresponding to a case where the sequencing machine cannot detect any base or nucleotide. For each read comprising at least one N, the method then comprises a step 8 of determining the number of such N mismatches; and a step 10 of comparing the number of N mismatches to a reference threshold. In a preferred embodiment, although it should not be construed as limiting the scope of the invention, the reference threshold is equal to 31.

[0071] Regardless of the result of determining step 4, the method includes a next step 12, which encodes the read segment at least according to the determination. More precisely, the read segment determined to be completely mapped with the reference sequence is encoded according to the first encoding process, regardless of whether the read segment does not contain N or has a number of N below the reference threshold. The read segment determined to be unmapped or the read segment determined to be completely mapped but has a number of N greater than the reference threshold is encoded according to the second encoding process, wherein each nucleotide or base is encoded separately, regardless of whether the nucleotide or base is aligned or unaligned. The read segment determined to be incompletely mapped is encoded according to the second encoding process or the third encoding process. More precisely, the read segment determined to be incompletely mapped with a number of mismatches greater than the threshold is encoded according to the second encoding process. If the read segment is determined to be incompletely mapped with a number of mismatches below the threshold, if the read segment does not contain N or has a number of N below the reference threshold, the read segment is encoded according to the third encoding process. If this is not the case, ie if the reads have a number N greater than a reference threshold, the reads are encoded according to a second encoding process.

[0072] Regardless of whether a given read has been determined to be fully mapped, incompletely mapped, or unmapped, if the read contains at least one N but has a number of N below a reference threshold, the encoding step 12 includes encoding a list of positions along the reference sequence that correspond to the positions of N in the reference sequence. This list of positions is then stored in a memory of a computing device that implements the compression method. If the read contains at least one N but has a number of N below a reference threshold, it is encoded according to a second encoding process, and each nucleotide or base of the read is encoded separately with 2 bits.

[0073] If a read contains at least one N but has a number of N greater than the reference threshold, the read will in any case be encoded according to the second encoding process and each nucleotide or base of the read will be encoded individually with 4 bits. In this case, the encoding step 12 does not include encoding and storing a list of N positions in the reference sequence. In fact, each N mismatch is then directly encoded according to the second encoding process in much the same way as the other nucleotides or bases of the read.

[0074] The first encoding process and the third encoding process include different descriptor sets. Each descriptor set univocally represents a read associated with the corresponding encoding process, and each of the first encoding process and the third encoding process is a simplified information entropy encoding process. More precisely, the third encoding process includes a first encoding sub-process and a second encoding sub-process. According to the first encoding sub-process, the reads of the incomplete mapping determined as the global mapping during step 4 are encoded. According to the second encoding sub-process, the reads of the incomplete mapping determined as the local mapping during step 4 are encoded. The first encoding sub-process and the second encoding sub-process include different descriptor sets, each descriptor set univocally represents a read associated with the corresponding encoding sub-process.

[0075] The alignment information encoded for each read and enabling reconstruction of the entire read sequence during data decompression then depends on the corresponding encoding process or sub-process used for the read. For example, the descriptor for the first encoding process may be:

[0076] o the absolute starting position of the fully mapped read relative to the reference sequence (encoded in 16 or 32 bits), and

[0077] o The length of the read (encoded relative to the length of the previous read using differential encoding, where the variable length code is in the range of 2 to 34 bits).

[0078] The descriptor for the first encoding sub-process may be:

[0079] o the absolute starting position of the incompletely mapped read relative to the reference sequence (encoded in 16 or 32 bits),

[0080] o the length of the read (encoded with respect to the length of the previous read using differential encoding, with variable length codes ranging from 2 bits to 34 bits), and

[0081] o List of mismatches for the reads.

[0082] The descriptor for the second encoding sub-process may be:

[0083] o the absolute starting position of the incompletely mapped portion of the read relative to the reference sequence - also called the local alignment start position (encoded in 16 or 32 bits),

[0084] o the length of the read (encoded relative to the length of the previous read using differential encoding, where the variable length code is in the range of 2 bits to 34 bits),

[0085] o a list of mismatches for the reads, and

[0086] o The length of the clipped parts of the read that are not part of the alignment (encoded in 8 bits for each clipped part).

[0087] Preferably, the mismatch list encoded in the first subprocess and the second subprocess includes a header (a bit mark encoded on 1 byte). The first five bits of this byte are used to encode the number of mismatches contained in the read segment (in a preferred embodiment, where the threshold is equal to 31, the number is in the range [0-31]). Then one bit is used to encode whether the incompletely mapped read segment is a global mapping or a local mapping. Another bit is used to encode whether the 2-bit mode is activated for the second encoding process. The last bit is used to encode whether the 4-bit mode is activated for the second encoding process. Preferably, for each read segment encoded according to the second encoding subprocess during the encoding step 12, the cut portion of the read segment (i.e., those portions that are not part of the local alignment) is concatenated, and each nucleotide or base of the cut portion is encoded separately. In a preferred embodiment, each nucleotide or base of such a cut portion of the read segment is encoded separately with 2 bits.

[0088] In a preferred embodiment, each mismatch encoded in the mismatch list of an incompletely mapped read (i.e., encoded according to the first encoding sub-pass or the second encoding sub-pass) is encoded on 1 byte. More precisely, each mismatch of an incompletely mapped read to be encoded according to the first encoding sub-pass or the second encoding sub-pass may be encoded as follows:

[0089] The first two bits of the o-byte are used to encode the alternative nucleotide or base present in the read instead of the corresponding reference nucleotide or base in the reference sequence,

[0090] o The last six bits are used to encode the position of the mismatch in the reference sequence, which is calculated as an offset relative to the previous mismatch of the read (the relative position of the mismatch, except for the absolute position of the first mismatch of the read, which is encoded). Therefore, the range of this offset (encoded in 6 bits) is [0-63].

[0091] Figure 3 An example of encoding mismatches of a read according to a first encoding sub-process is provided. The read is an incompletely mapped read that is globally mapped to a reference sequence. The read has two mismatches:

[0092] o a first mismatch, located at position 12 in the read, consisting in replacing an A nucleotide in the reference sequence with a T nucleotide in the read, and

[0093] o The second mismatch, located at position 21 in the read, consists in replacing a C nucleotide in the reference sequence with a G nucleotide in the read.

[0094] The mismatch list for this read is then encoded as:

[0095] ο<12,T>, the value "12" corresponds to the absolute position of the first mismatch in the read segment, and

[0096] ο<9,G>, the value “9” corresponds to the relative position of the second mismatch in the read segment, ie, the offset between the second mismatch and the first mismatch.

[0097] For example, <12,T> can be converted to the value "51" (encoded on 1 byte), and <9,G> can be converted to the value "38" (encoded on 1 byte). This byte encoding is obtained as follows:

[0098] Offset position x4 + nucleotide value (where A=0, C=1, G=2, T=3)

[0099] Preferably, for each incompletely mapped read to be encoded according to the first encoding subprocess or the second encoding subprocess, if the offset calculated between a given mismatch of the read and the previous mismatch is greater than the maximum encodable value, at least one "false" mismatch is inserted between the two mismatches until each offset between each of the mismatches and the at least one "false" mismatch is below the maximum encodable value. A "false" mismatch is defined as a mismatch for which bits of a byte are used to encode the mismatch, or to encode a nucleotide or base that is equal to the corresponding reference nucleotide or base in the reference sequence. In a preferred embodiment, although it should not be construed as limiting the scope of the invention, the maximum encodable value is equal to 63, corresponding to the maximum value encoded in 6 bits.

[0100] Figure 4 An example of encoding mismatches of a read according to a first encoding sub-process is provided in the case where a "false" mismatch must be inserted. The read is an incompletely mapped read that is globally mapped to a reference sequence. The read has two mismatches:

[0101] o a first mismatch, located at position 22 in the read, consisting in replacing an A nucleotide in the reference sequence with a T nucleotide in the read, and

[0102] o The second mismatch, located at position 134 in the read, consists in replacing a C nucleotide in the reference sequence with a G nucleotide in the read.

[0103] The position offset between the second mismatch and the first mismatch is 112, which is greater than the maximum encodable value of 63. Therefore, a "false" mismatch must be inserted between the two mismatches so that each offset between each of the mismatches and the "false" mismatch is lower than the maximum encodable value. A "false" mismatch with a T nucleotide (corresponding to a "true" T nucleotide in the reference sequence) is inserted, for example, at position 85 in the read. The calculated position offset between this "false" mismatch and the first mismatch is 63, which corresponds to the maximum encodable value. The calculated position offset between the second mismatch and the "false" mismatch is 49, which is lower than 63.

[0104] The mismatch list for this read is then encoded as:

[0105] o<22,T>, the value "22" corresponds to the absolute position of the first mismatch in the read segment,

[0106] ο<63,T>, the value "63" corresponds to the relative position of the "false" mismatch in the read segment, i.e., the offset between the "false" mismatch and the first mismatch, and

[0107] ο<49,G>, the value “49” corresponds to the relative position of the second mismatch in the read segment, ie, the offset between the second mismatch and the “false” mismatch.

[0108] For example, <22,T> can be converted to the value "91" (encoded on 1 byte), <63,T> can be converted to the value "255" (encoded on 1 byte), and <49,G> can be converted to the value "198" (encoded on 1 byte). This byte encoding is obtained as follows:

[0109] Offset position x4 + nucleotide value (where A=0, C=1, G=2, T=3)

[0110] The method comprises a final step 14 of providing a compressed file containing a list of encoded reads. The encoded reads are stored in the compressed file in the same order as the reads stored in the initial uncompressed file. Each read can then be reconstructed from aligning the encoded information and the reference sequence by appropriate decompression software and / or methods configured according to the present invention.

[0111] Although reference is made to the exemplary architecture of computing device 20 (for illustrative purposes only) Figure 2 Although the present invention is described in detail, the technology disclosed herein can be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, the computer program code can be stored on a computer medium and executed by a hardware processing unit including one or more processors, which is similar to using Figure 2The same is true for the case of the device 20. It should be understood that the term "processor" as used herein is intended to include one or more processing devices, including signal processors, microprocessors, microcontrollers, application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other types of processing circuit systems, as well as parts or combinations of such circuit system elements. In addition, the term "memory" as used herein is intended to include electronic memory associated with a processor, such as random access memory (RAM), read-only memory (ROM), or other types of memory used in any combination.

[0112] Thus, software instructions or code for executing the methods and protocols described herein may be stored in one or more of the associated memory devices (e.g., ROM, fixed memory, or removable memory) and, when ready for use, loaded into RAM and executed by the processor.

[0113] The techniques of this disclosure may be implemented in a wide variety of devices or apparatuses, including, for example, mobile phones, computers, servers, tablet computers, and similar devices.

[0114] Although illustrative embodiments of the present invention have been described herein with reference to the accompanying drawings, it will be understood that the invention is not limited to those precise embodiments and that various other changes and modifications may be made by those skilled in the art without departing from the scope or spirit of the invention.

[0115] Statistical and numerical examples of the compression method according to the invention

[0116] The following comparative example was performed on an uncompressed data file containing 48 million nucleotide reads or sequences:

[0117] Uncompressed data file size: 35,770 MB (megabytes)

[0118] ο Size of the file compressed with gzip software: 6,649MB

[0119] ο Size of the file compressed with non-reference based SPRING software: 1,402MB

[0120] o Size of the file compressed using the reference-based compression method according to the present invention: 1,179 MB

[0121] ο Compression time using non-reference based SPRING software: 1,722s

[0122] o Compression time using the reference-based compression method according to the present invention: 181 s

[0123] o Average size of uncompressed data files in bits / nucleotide (ASCII encoding): 8 bits / nucleotide

[0124] o Average size in bits / nucleotide of files compressed with encoding suitable for 4 possible characters A, T, C, G: 2 bits / nucleotide

[0125] o Average size in bits / nucleotide of files that have been compressed using the reference-based compression method according to the invention: 0.33 bits / nucleotide

[0126] The numerical examples indicated above illustrate that the present invention allows fast compression and decompression while providing a high compression ratio.

Claims

1. A method for compressing genomic sequence data, the method comprising: Obtaining read records by one or more computers; determining, by the one or more computers, whether the read record corresponds to a read that is completely mapped to a reference sequence or a read that is incompletely mapped to the reference sequence; Based on a determination by the one or more computers that the read record corresponds to a read that is incompletely mapped to the reference sequence, determining by the one or more computers whether a number of mismatches for the incompletely mapped reads satisfies a predetermined threshold number of mismatches; as well as Based on determining that the number of mismatches satisfies the predetermined threshold number of mismatches, (i) the one or more computers obtain an offset relative to a previous mismatch that is less than a maximum encodable offset value, and (ii) the one or more computers encode each mismatch of the incompletely mapped read segment and the offset relative to the previous mismatch of the read segment into a record having a size of 1 byte.

2. The method of claim 1 , wherein determining, by the one or more computers, whether the number of mismatches of the incompletely mapped read segments satisfies a predetermined mismatch threshold number comprises: A determination is made, by the one or more computers, whether the number of mismatches for the incompletely mapped reads is greater than the predetermined threshold number of mismatches.

3. The method of claim 1, wherein each read segment record comprises: data indicating the absolute starting position of the aligned read relative to the reference sequence; data indicating the length of the read segment; data indicating whether the read segment is completely mapped or incompletely mapped; data indicating a number of mismatches identified in the read segment; as well as Data indicating a relative position of each of the mismatches in the read segment.

4. The method of claim 1 , wherein encoding each mismatch of the incompletely mapped read segment as a record having a size of 1 byte comprises: For each specific mismatch, encoding, by the one or more computers, first two bits of the byte to include data representing an alternative nucleotide or base present in the read instead of a corresponding reference nucleotide or base in the reference sequence; as well as The remaining six bits of the byte are encoded by one or more computers to include data representing the offset.

5. The method according to claim 4, further comprising: determining, by one or more computers, whether the offset is greater than a maximum encodable value; as well as Based on determining that the offset is greater than the maximum encodable value, at least one false mismatch is inserted by one or more computers between the particular mismatch and the previous mismatch.

6. The method according to claim 1, further comprising: Based on determining that the number of mismatches does not satisfy the predetermined threshold number of mismatches, a list of positions of the reference sequence corresponding to the position of each of the mismatches is encoded into the reference sequence by one or more computers using a simplified information entropy encoding process.

7. The method according to claim 1, further comprising: Based on determining that the read record corresponds to a read that is completely mapped to the reference sequence, encoding at least a portion of the read record using simplified information entropy coding by one or more computers.

8. The method of claim 1, wherein the one or more computers comprise one or more hardware processors.

9. The method of claim 8, wherein the one or more hardware processors include one or more field programmable gate arrays (FPGAs).

10. A hardware processor comprising hardware processing circuitry configured to perform one or more operations comprising: Obtaining a read segment record by the hardware processing circuit system; determining, by the hardware processing circuitry, whether the read record corresponds to a read that is completely mapped to a reference sequence or a read that is incompletely mapped to the reference sequence; Based on a determination by the hardware processing circuitry that the read record corresponds to a read that is incompletely mapped to the reference sequence, determining by the one or more computers whether a number of mismatches for the incompletely mapped reads satisfies a predetermined threshold number of mismatches; and Based on determining that the number of mismatches satisfies the predetermined mismatch threshold number, (i) the hardware processing circuit system obtains an offset relative to a previous mismatch that is less than a maximum encodable offset value, and (ii) the hardware processing circuit system encodes each mismatch of the incompletely mapped read segment and the offset relative to the previous mismatch of the read segment into a record having a size of 1 byte.

11. The hardware processor of claim 10, wherein each read segment record comprises: data indicating the absolute starting position of the aligned read relative to the reference sequence; data indicating the length of the read segment; data indicating whether the read segment is completely mapped or incompletely mapped; data indicating a number of mismatches identified in the read segment; and Data indicating a relative position of the mismatch in the read segment.

12. The hardware processor of claim 10, wherein encoding each mismatch of the incompletely mapped read segment as a record having a size of 1 byte comprises: For each specific mismatch, encoding, by the hardware processing circuitry, first two bits of the byte to include data representing an alternative nucleotide or base present in the read instead of a corresponding reference nucleotide or base in the reference sequence; and The remaining six bits of the byte are encoded by the hardware processing circuitry to include data representing the offset.

13. The hardware processor of claim 12, wherein the hardware processor circuitry is further configured to perform operations comprising: determining, by the hardware processing circuit system, whether the offset is greater than a maximum encodable value; Based on determining that the offset is greater than the maximum encodable value, at least one false mismatch is inserted, by the hardware processing circuitry, between the particular mismatch and the previous mismatch.

14. The hardware processor of claim 10, wherein the hardware processor circuitry is further configured to perform operations comprising: Based on determining that the number of mismatches does not satisfy the predetermined threshold number of mismatches, encoding, by the hardware processing circuitry, a list of positions of the reference sequence corresponding to the position of each of the mismatches into the reference sequence using a simplified information entropy encoding process.

15. The hardware processor of claim 10, wherein the hardware processor circuitry is further configured to perform operations comprising: Based on determining that the read record corresponds to a read that is completely mapped to the reference sequence, encoding, by the hardware processing circuitry, at least a portion of the read record using simplified information entropy encoding.

16. The hardware processor of claim 10, wherein the hardware processing circuitry comprises one or more field programmable gate arrays (FPGAs).

17. The hardware processor of claim 10, wherein determining, by the hardware processing circuitry, whether the number of mismatches of the incompletely mapped read segments satisfies a predetermined mismatch threshold number comprises: A determination is made, by the hardware processing circuitry, whether the number of mismatches for the incompletely mapped read segments is greater than the predetermined threshold number of mismatches.

18. A system for compressing genomic sequence data, the system comprising: One or more computers, and one or more storage devices storing instructions, wherein the instructions, when executed by the one or more computers, are operable to cause the one or more computers to perform the following operations, wherein the operations include: Obtaining, by the one or more computers, a read record; determining, by the one or more computers, whether the read record corresponds to a read that is completely mapped to a reference sequence or a read that is incompletely mapped to the reference sequence; Based on a determination by the one or more computers that the read record corresponds to a read that is incompletely mapped to the reference sequence, determining by the one or more computers whether a number of mismatches for the incompletely mapped reads satisfies a predetermined threshold number of mismatches; and Based on determining that the number of mismatches satisfies the predetermined threshold number of mismatches, (i) the one or more computers obtain an offset relative to a previous mismatch that is less than a maximum encodable offset value, and (ii) the one or more computers encode each mismatch of the incompletely mapped read segment and the offset relative to the previous mismatch of the read segment into a record having a size of 1 byte.

19. The system of claim 18, wherein each read record comprises: data indicating the absolute starting position of the aligned read relative to the reference sequence; data indicating the length of the read segment; data indicating whether the read segment is completely mapped or incompletely mapped; data indicating a number of mismatches identified in the read segment; as well as Data indicating a relative position of each of the mismatches in the read segment.

20. The system of claim 18, wherein encoding each mismatch of the incompletely mapped read segment as a record having a size of 1 byte comprises: For each specific mismatch: encoding, by one or more computers, first two bits of the byte to include data representing an alternative nucleotide or base present in the read instead of a corresponding reference nucleotide or base in the reference sequence; as well as The remaining six bits of the byte are encoded by one or more computers to include data representing the offset.

21. The system of claim 20, wherein the operations further comprise: determining, by the one or more computers, whether the offset is greater than a maximum encodable value; as well as Based on determining that the offset is greater than the maximum encodable value, at least one false mismatch is inserted by one or more computers between the particular mismatch and the previous mismatch.

22. The system of claim 18, the operations further comprising: Based on determining that the number of mismatches does not satisfy the predetermined threshold number of mismatches, a list of positions of the reference sequence corresponding to the position of each of the mismatches is encoded into the reference sequence by one or more computers using a simplified information entropy encoding process.

23. The system of claim 18, the operations further comprising: Based on determining that the read record corresponds to a read that is completely mapped to the reference sequence, encoding at least a portion of the read record using simplified information entropy coding by one or more computers.

24. The system of claim 18, wherein the one or more computers include one or more hardware processors.

25. The system of claim 24, wherein the one or more hardware processors include one or more field programmable gate arrays (FPGAs).

26. A computer-readable storage device having instructions stored thereon, which, when executed by a data processing device, cause the data processing device to perform operations for compressing genomic sequence data, the operations comprising: Obtain read segment records; determining whether the read record corresponds to a read that is completely mapped to a reference sequence or a read that is incompletely mapped to the reference sequence; Based on determining that the read record corresponds to a read that is incompletely mapped to the reference sequence, determining whether a number of mismatches for the incompletely mapped reads satisfies a predetermined threshold number of mismatches; and Based on determining that the number of mismatches satisfies the predetermined mismatch threshold number, (i) obtaining an offset relative to a previous mismatch that is less than a maximum encodable offset value, and (ii) encoding each mismatch of the incompletely mapped read segment and the offset relative to the previous mismatch of the read segment into a record having a size of 1 byte.

27. The computer readable storage device of claim 26, wherein each read segment record comprises: data indicating the absolute starting position of the aligned read relative to the reference sequence; data indicating the length of the read segment; data indicating whether the read segment is completely mapped or incompletely mapped; data indicating a number of mismatches identified in the read segment; as well as Data indicating a relative position of each of the mismatches in the read segment.

28. The computer readable storage device of claim 26, wherein encoding each mismatch of the incompletely mapped read segment as a record having a size of 1 byte comprises: For each specific mismatch, encoding, by one or more computers, first two bits of the byte to include data representing an alternative nucleotide or base present in the read instead of a corresponding reference nucleotide or base in the reference sequence; as well as The remaining six bits of the byte are encoded by one or more computers to include data representing the offset.

29. The computer readable storage device of claim 28, the operations further comprising: determining, by the one or more computers, whether the offset is greater than a maximum encodable value; as well as Based on determining that the offset is greater than the maximum encodable value, at least one false mismatch is inserted by one or more computers between the particular mismatch and the previous mismatch.

30. The computer readable storage device of claim 26, the operations further comprising: Based on determining that the number of mismatches does not satisfy the predetermined threshold number of mismatches, a list of positions of the reference sequence corresponding to the position of each of the mismatches is encoded into the reference sequence using a simplified information entropy encoding process.

31. The computer readable storage device of claim 26, the operations further comprising: Based on determining that the read record corresponds to a read that is completely mapped to the reference sequence, encoding at least a portion of the read record using simplified information entropy coding.

Citation Information

Patent Citations

  • Method and apparatus for compact representation of bioinformatics data

    WO2018068829A1

  • Method for compressing genomic data

    CN107851137A

  • Method and apparatus for compact representation of bioinformatics data

    CN110168649A