Method for compressing genomic sequence data

By encoding the mapping state between read segments and reference sequences, and employing simplified information entropy encoding and spurious mismatch insertion methods, the problems of low compression ratio and information loss in genome sequencing data compression are solved, achieving fast compression and high compression ratio while maintaining read sequence consistency.

CN114341988BActive Publication Date: 2025-10-28ILLUMINA INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202080062727.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2019-09-11
Filing Date
2020-09-11
Publication Date
2025-10-28
Estimated Expiration
2040-09-11

AI Technical Summary

Technical Problem

Existing methods for compressing genome sequencing data suffer from low compression ratios, slow speeds, and information loss. In particular, reference-based compression methods can cause changes in read order during compression, affecting downstream analysis and file consistency checks.

Method used

By determining the mapping state between the read segment and the reference sequence, different encoding processes are used to encode fully mapped and incompletely mapped read segments. A simplified information entropy encoding process is adopted to maintain the initial order of the read segments, and false mismatches are inserted during the encoding process to optimize compression, thereby achieving high compression ratio and fast compression.

Benefits of technology

It achieves fast compression and decompression while maintaining the initial order of read segments, provides a high compression ratio, and makes downstream analysis and consistency checks easier, avoiding information loss.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114341988B_ABST
    Figure CN114341988B_ABST
Patent Text Reader

Abstract

This invention relates to a reference-based method for compressing genomic sequence data generated by a sequencing machine. The method involves determining whether a sequence of nucleotides or bases previously aligned with a reference sequence is fully mapped, partially mapped, or unmapped; and then encoding based on this determination. The determination step includes: for each partially mapped sequence, comparing the number of mismatches between the sequence and the reference sequence to a reference threshold; and encoding the partially mapped sequences according to a different encoding process based on the results of the comparison method used to compress the genomic sequence data generated by the sequencing machine.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates generally to methods for representing genome sequencing data generated by sequencing machines, and more specifically to computer-implemented methods for compressing such genome sequencing data. This disclosure provides a reference-based compression method that allows for rapid compression and decompression without information loss and achieves a high compression ratio. Background Technology

[0002] Next-generation sequencing machines are now generating massive amounts of sequencing data at affordable prices. The latest systems produce over 6 billion 150-nucleotide sequences in a single 36-hour run, enough to sequence 20 complete human genomes. This opens up many new perspectives for the diagnosis of genetic diseases and the development of personalized medicine, aiming to accommodate treatments based on human genome specificity.

[0003] However, this also brings new challenges, particularly the costs associated with storing massive amounts of data. The most common file format for raw (unaligned) sequence data is FASTQ, which stores sequence data (strings of A, C, T, and G nucleotides, also known as reads), quality values ​​(the probability that the sequencing platform will cause a sequencing error for each nucleotide), and sequence names. This is a plain ASCII text file, typically compressed using the common text compression scheme LZ (Lempel-Ziv, implemented in gzip software). However, using such compression methods introduces several problems:

[0004] -Low compression ratio due to incomplete utilization of data redundancy

[0005] - Slow compression and decompression

[0006] There are also compression methods specifically designed for FASTQ encoding, categorized as either reference-based or non-reference-based. However, none are entirely satisfactory because: a) reference-based methods offer good compression ratios but are slow; b) non-reference-based methods are faster but have lower compression ratios. An example of a non-reference-based method is provided by the software SPRING, a non-reference compressor for FASTQ files (Web address: github.com / shubhamchandak94 / SPRING). However, the compression ratio provided by SPRING is low.

[0007] In reference-based compression methods, several approaches have been proposed that utilize sequence alignment and aim for faster speeds and good compression ratios. However, such methods encounter several problems, a notable major one being that they are not entirely lossless. One known reference-based compression method is described, for example, in patent document WO 2018 / 068829 A1. In this described method, after alignment with one or more reference sequences, nucleotide sequences are classified according to matching accuracy (thus creating categories of aligned reads), and these nucleotide sequences are then encoded into multi-level syntax elements for each layer in which the data is partitioned, using different source models and entropy encoders. Thus, the categories of data are encoded separately and constructed in different layers of syntax elements, each layer containing descriptors that unilaterally represent the classified and aligned reads of that layer. This method aims to obtain different information sources with simplified entropy, thereby allowing improved compression performance and selective access to compressed data of specific categories. However, this compression method reorders reads in an order different from the order obtained at the end of the read alignment step (i.e., reorders reads according to their categories). Consequently, some information is lost during the compression process (especially the initial sequence ordering). Therefore, the reproducibility of some analysis results may be affected, as some downstream analysis software may rely on the order of the read segments. Furthermore, decompressing data in a different order than the initial read segment order makes it more difficult to verify whether the uncompressed file is identical to the original file. Additionally, this compression method is relatively slow, especially compared to state-of-the-art non-reference-based compression methods. Summary of the Invention

[0008] This disclosure addresses the problems of prior art solutions by providing systems, methods, computer programs, and hardware circuitry for compressing genomic sequence data. In one aspect, a computer-implemented method is disclosed for compressing genomic sequence data generated by a sequencing machine, the genomic sequence data comprising reads of sequences of nucleotides or bases aligned to a reference sequence, thereby generating aligned reads, the aligned reads being stored as a read list in an initial file, the method comprising:

[0009] For each aligned read segment, determine whether the read segment is fully mapped to the reference sequence, partially mapped, or unmapped with the reference sequence.

[0010] - The read segments are encoded according to the determination, wherein the read segments determined to be fully mapped are encoded according to a first encoding process, and the read segments determined to be unmapped are encoded according to a second encoding process.

[0011] The determining step includes, for each incompletely mapped read, comparing the number of mismatches between the read and the reference sequence with a threshold.

[0012] In the encoding step, the read segments determined to be incompletely mapped are encoded according to the second encoding process or the third encoding process. When the number of mismatches is greater than the threshold, the incompletely mapped read segments are encoded according to the second encoding process; when the number of mismatches is less than the threshold, the incompletely mapped read segments are encoded according to the third encoding process.

[0013] -In the second encoding process, each nucleotide or base of the read segment is encoded individually.

[0014] -The first encoding process and the third encoding process include different sets of descriptors, each set of descriptors unilaterally representing a read segment associated with the corresponding encoding process, and each of the first encoding process and the third encoding process is a simplified information source entropy encoding process.

[0015] This disclosure overcomes the shortcomings of existing compression methods by allowing rapid compression and decompression without information loss and providing a high compression ratio. More specifically, this disclosure focuses on encoding the most frequent cases in the most compact way, even if it means using a degraded encoding pattern for the rarest and least frequent cases. This results in a significant improvement in compression performance. Furthermore, compression performed by the method described herein is faster due to the genomic information representation format used in this disclosure. Last but not least, this disclosure preserves the initial order of reads and does not reorder reads based on their category. Therefore, no information is lost during the process, which makes downstream analysis easier and allows for effective consistency checks after the decompression step.

[0016] These and other features and advantages of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Furthermore, although thresholds may be referred to herein as being exceeded or not exceeded, it should be understood that such thresholds can be conceptually employed such that determining whether such thresholds are met, satisfied, or otherwise detected is done regardless of whether the number or value used to implement those threshold assessments is described using positive or negative values.

[0017] According to one innovative aspect of this disclosure, a method for compressing genomic sequence data is disclosed. In one aspect, the method may include performing one or more operations via executing software instructions through one or more computers, wherein the operations include: obtaining read records by the one or more computers; determining by the one or more computers whether the read record corresponds to a read that is fully mapped to a reference sequence or a read that is not fully mapped to the reference sequence; based on the determination by the one or more computers that the read record corresponds to a read that is not fully mapped to the reference sequence, determining by the one or more computers whether the number of mismatches of the not fully mapped reads meets the predetermined mismatch threshold number; and based on the determination that the number of mismatches meets the predetermined mismatch threshold number, encoding each mismatch of the not fully mapped reads into a record of size 1 byte by the one or more computers.

[0018] Other aspects include corresponding systems, apparatuses, and computer programs that perform actions as disclosed herein, such as those defined by instructions encoded on a computer-readable storage device.

[0019] These and other versions may optionally include one or more of the following features. For example, in some implementations, determining by the one or more computers whether the number of mismatches in the incompletely mapped read segments meets a predetermined mismatch threshold number may include determining by the one or more computers whether the number of mismatches in the incompletely mapped read segments is greater than the predetermined mismatch threshold number.

[0020] In some implementations, each read segment record may include: data indicating the absolute start position of the read segment relative to the reference sequence; data indicating the length of the read segment; data indicating whether the read segment is fully mapped or partially mapped; data indicating the number of mismatches identified in the read segment; and data indicating the relative position of each of the possible mismatches in the read segment.

[0021] In some specific implementations, encoding each mismatch of the incompletely mapped read segment as a record of 1 byte size includes: for each specific mismatch, the one or more computers encoding the first two bits of the byte to include data indicating a substitute nucleotide or base present in the read segment instead of a corresponding reference nucleotide or base in the reference sequence; and the one or more computers encoding the remaining six bits of the byte to include data indicating the position of the mismatch in the reference sequence, the position being calculated as an offset relative to the previous mismatch of the read segment.

[0022] In some specific implementations, the method may further include: determining by one or more computers whether the offset is greater than the maximum coded value; and based on determining that the offset is greater than the maximum coded value, inserting at least one false mismatch between the specific mismatch and the previous mismatch by one or more computers.

[0023] In some specific implementations, the method may further include: based on determining that the number of mismatches does not meet the predetermined number of mismatch thresholds, one or more computers use a simplified information entropy encoding process to encode a list of positions of the reference sequence corresponding to the position of each of the mismatches into the reference sequence.

[0024] In some specific implementations, the method may further include: based on determining that the read segment record corresponds to a read segment that is fully mapped to the reference sequence, one or more computers use simplified information entropy coding to encode at least a portion of the read segment record.

[0025] In some specific implementations, the one or more computers may include one or more hardware processors.

[0026] In some implementations, the one or more hardware processors may include one or more field-programmable gate arrays (FPGAs).

[0027] In some implementations, the method for compressing genomic sequence data may be performed by one or more hardware processors. In such implementations, the hardware processor may include a hardware processing circuitry system configured to perform one or more operations. In one aspect, the operations may include: obtaining read records by the hardware processing circuitry system; determining by the hardware processing circuitry system whether the read record corresponds to a read that is fully mapped to a reference sequence or a read that is not fully mapped to the reference sequence; based on the determination by the hardware processing circuitry system that the read record corresponds to a read that is not fully mapped to the reference sequence, determining by the one or more computers whether the number of mismatches of the not fully mapped reads meets a predetermined mismatch threshold number; and based on the determination that the number of mismatches meets the predetermined mismatch threshold number, encoding each mismatch of the not fully mapped reads into a record of size 1 byte by the hardware processing circuitry system.

[0028] In some specific implementations, each read record may include: data indicating the absolute start position of the alignment read relative to the reference sequence; data indicating the length of the read; data indicating whether the read is fully mapped or partially mapped; data indicating the number of mismatches identified in the read; and data indicating the relative position of the possible mismatches in the read.

[0029] In some specific implementations, determining whether the number of mismatches in the incompletely mapped read segments meets a predetermined mismatch threshold number may include determining whether the number of mismatches in the incompletely mapped read segments is greater than the predetermined mismatch threshold number.

[0030] In some specific implementations, encoding each mismatch of the incompletely mapped read segment as a record of 1 byte size may include: for each specific mismatch, the hardware processing circuitry system encoding the first two bits of the byte to include data indicating that a substitute nucleotide or base is present in the read segment instead of a corresponding reference nucleotide or base in the reference sequence; and the hardware processing circuitry system encoding the remaining six bits of the byte to include data indicating the position of the mismatch in the reference sequence, the position being calculated as an offset relative to the previous mismatch of the read segment.

[0031] In some specific implementations, the hardware processing circuitry is further configured to perform the following operations: determining whether the offset is greater than a maximum coded value; and based on the determination that the offset is greater than the maximum coded value, inserting at least one false mismatch between the specific mismatch and the previous mismatch.

[0032] In some specific implementations, the hardware processing circuitry is further configured to perform the following operations: based on determining that the number of mismatches does not meet the predetermined number of mismatches, the hardware processing circuitry encodes a list of positions of the reference sequence corresponding to the position of each of the mismatches into the reference sequence using a simplified information entropy encoding process.

[0033] In some specific implementations, the hardware processing circuitry is further configured to perform the following operations: based on determining that the read segment record corresponds to a read segment that is fully mapped to the reference sequence, the hardware processing circuitry uses simplified information entropy coding to encode at least a portion of the read segment record.

[0034] In some implementations, the hardware processing circuitry system includes one or more field-programmable gate arrays (FPGAs).

[0035] According to another innovative aspect of this disclosure, a method for compressing genomic sequence data is disclosed. In one aspect, the method may include the following operations: one or more processors accessing a storage device storing the plurality of read records in a manner that preserves the sequence order of the plurality of read records generated by a mapping and alignment module; for each specific read record among the plurality of read records: the one or more processors obtaining the specific read record; the one or more processors determining whether the specific read record corresponds to a read that is fully mapped to a reference sequence or a read that is not fully mapped to the reference sequence; based on the determination by the one or more processors that the specific read record corresponds to a read that is not fully mapped to the reference sequence, the one or more processors determining whether the number of mismatches of the not fully mapped reads meets a predetermined mismatch threshold number; based on the determination that the number of mismatches meets the predetermined mismatch threshold number, the one or more processors encoding each mismatch of the not fully mapped reads into a compressed record having a predetermined compressed record size; and the one or more processors storing the compressed record in the storage device while maintaining the sequence order of the read records.

[0036] Other aspects include corresponding systems, apparatuses, and computer programs that perform actions as disclosed herein, such as those defined by instructions encoded on a computer-readable storage device.

[0037] These and other versions may optionally include one or more of the following features. For example, in some specific implementations, each of the plurality of read records may include: data indicating the absolute start position of the alignment read relative to the reference sequence; data indicating the length of the read; data indicating whether the read is fully mapped or partially mapped; data indicating the number of mismatches identified in the read; data indicating whether the read includes at least one undetermined base N; data indicating the number of undetermined base Ns in the read; data indicating whether the read is mapped or unmapped; data indicating the position of the read record in the read record sequence output by the mapping and alignment module; and data indicating the relative position of the possible mismatches in the read.

[0038] In some implementations, the predetermined compressed record size is one byte.

[0039] In some implementations, encoding each mismatch of the incompletely mapped read segment into a compressed record of one byte size may include: for each specific mismatch, having one or more processors encode the first two bits of the byte to include data indicating that an alternative nucleotide or base is present in the read segment instead of a corresponding reference nucleotide or base in the reference sequence; and having one or more processors encode the remaining six bits of the byte to include data indicating the position of the mismatch in the reference sequence, the position being calculated as an offset relative to the previous mismatch of the read segment.

[0040] In some specific implementations, the method may further include: determining by one or more processors whether the offset is greater than the maximum coded value; and based on determining that the offset is greater than the maximum coded value, inserting at least one false mismatch between the specific mismatch and the previous mismatch by one or more processors.

[0041] In some specific implementations, the method may further include: based on determining that the number of mismatches does not meet the predetermined number of mismatch thresholds, one or more processors use a simplified information entropy encoding process to encode a list of positions of the reference sequence corresponding to the positions of each of the mismatches into the reference sequence.

[0042] In some specific implementations, the method may further include: based on determining that the read segment record corresponds to a read segment that is fully mapped to the reference sequence, the one or more processors use simplified information entropy coding to encode at least a portion of the read segment record.

[0043] In some specific implementations, determining whether the number of mismatches in the incompletely mapped read segments meets a predetermined mismatch threshold number by the one or more computers may include determining whether the number of mismatches in the incompletely mapped read segments is greater than the reference threshold.

[0044] According to another aspect of this disclosure, a hardware processor is disclosed. In one aspect, the hardware processor may include a hardware processing circuitry system configured to perform one or more operations. In one aspect, the hardware processing circuitry is configured to perform operations including: accessing a storage device that stores the plurality of read segment records in a manner that preserves the sequential order of the read segment records generated by the mapping and alignment module; for each specific read segment record among the plurality of read segment records: obtaining the specific read segment record by the hardware processing circuitry; determining whether the specific read segment record corresponds to a read segment fully mapped to a reference sequence or a read segment partially mapped to the reference sequence; based on the determination by the hardware processing circuitry that the specific read segment record corresponds to a read segment partially mapped to the reference sequence, determining whether the number of mismatches of the partially mapped read segments meets a predetermined mismatch threshold number; based on the determination that the number of mismatches meets the predetermined mismatch threshold number, encoding each mismatch of the partially mapped read segments into a compressed record having a predetermined compressed record size by the hardware processing circuitry; and storing the compressed record in the storage device while maintaining the sequential order of the read segment records.

[0045] These and other versions may optionally include one or more of the following features. For example, in some specific implementations, each of the plurality of read segment records accessible by the hardware processing circuitry system may include: data indicating the absolute start position of the alignment read segment relative to the reference sequence; data indicating the length of the read segment; data indicating whether the read segment is fully mapped or partially mapped; data indicating the number of mismatches identified in the read segment; data indicating whether the read segment includes at least one undetermined base N; data indicating the number of undetermined base Ns in the read segment; data indicating whether the read segment is mapped or unmapped; data indicating the position of the read segment record in the read segment record sequence output by the mapping and alignment module; and data indicating the relative position of the possible mismatches in the read segment.

[0046] In some specific implementations, the predetermined size of the compressed record generated by the hardware processing circuitry system can be one byte.

[0047] In some specific implementations, encoding each mismatch of the incompletely mapped read segment into a compressed record of one byte size may include: for each specific mismatch, the hardware processing circuitry system encoding the first two bits of the byte to include data indicating that an alternative nucleotide or base is present in the read segment instead of a corresponding reference nucleotide or base in the reference sequence; and the hardware processing circuitry system encoding the remaining six bits of the byte to include data indicating the position of the mismatch in the reference sequence, the position being calculated as an offset relative to the previous mismatch of the read segment.

[0048] In some specific implementations, the hardware processor may be further configured to include a hardware processing circuitry system configured to perform the following operations: determining whether the offset is greater than a maximum coded value; and based on the determination that the offset is greater than the maximum coded value, inserting at least one false mismatch between the specific mismatch and the previous mismatch.

[0049] In some specific implementations, the hardware processor may be further configured to include a hardware processing circuit system configured to perform the following operations: based on determining that the number of mismatches does not meet the predetermined number of mismatches, the hardware processing circuit system encodes a list of positions of the reference sequence corresponding to the position of each of the mismatches into the reference sequence using a simplified information entropy encoding process.

[0050] In some specific implementations, the hardware processor may be further configured to include a hardware processing circuitry system configured to perform the following operations: based on determining that the read segment record corresponds to a read segment that is fully mapped to the reference sequence, the hardware processing circuitry system uses simplified information entropy coding to encode at least a portion of the read segment record.

[0051] In some specific implementations, the determination by the hardware processing circuit system of whether the number of mismatches in the incompletely mapped read segments meets a predetermined mismatch threshold number includes:

[0052] The hardware processing circuit system determines whether the number of mismatches in the incompletely mapped read segments is greater than the predetermined mismatch threshold number.

[0053] According to another innovative aspect of this disclosure, a computer-implemented method for compressing genomic sequence data generated by a sequencing machine is disclosed, the genomic sequence data comprising reads of sequences of nucleotides or bases aligned with a reference sequence, thereby generating aligned reads, which are stored as a read list in an initial file. In one aspect, the method may include the following actions for each aligned read: determining whether the read is fully mapped or partially mapped with the reference sequence, or whether the read is not mapped with the reference sequence; encoding the read according to the determination, wherein reads determined to be fully mapped are encoded according to a first encoding process, and reads determined to be unmapped are encoded according to a second encoding process, wherein the determination step includes, for each partially mapped read, comparing the number of mismatches between the read and the reference sequence with a threshold, wherein, in the encoding step, the second encoding process or a third encoding process is used... The reads identified as incompletely mapped are encoded. When the number of mismatches is greater than the threshold, the incompletely mapped reads are encoded according to the second encoding process, and when the number of mismatches is less than the threshold, the incompletely mapped reads are encoded according to the third encoding process. In the second encoding process, each nucleotide or base of the read is encoded individually. The first and third encoding processes include different sets of descriptors, each set of descriptors unilaterally representing the read associated with the corresponding encoding process. Each of the first and third encoding processes is a simplified information source entropy encoding process.

[0054] Other aspects include corresponding systems, apparatuses, and computer programs that perform actions as disclosed herein, such as those defined by instructions encoded on a computer-readable storage device.

[0055] These and other versions may optionally include one or more of the following features. For example, in some implementations, the determining step may include a further determination when it is determined that the read segment is not fully mapped to the reference sequence and has a mismatch number below the threshold, the further determination relating to whether the read segment is globally mapped or locally mapped to the reference sequence, and wherein the third encoding process includes a first encoding subprocess and a second encoding subprocess, the first encoding subprocess encoding the read segment determined to be globally mapped, and the second encoding subprocess encoding the read segment determined to be locally mapped, the first encoding subprocess and the second encoding subprocess including different sets of descriptors, each set of descriptors unilaterally representing the read segment associated with the corresponding encoding subprocess.

[0056] In some specific implementations, the descriptor of the first encoding sub-process may include the alignment start position, read length, and mismatch list represented by symbol substitution in the reference sequence, and wherein the descriptor of the second encoding sub-process includes the local alignment start position, read length, mismatch list represented by symbol substitution in the reference sequence, and the length of the cut portion of the read that is not part of the alignment.

[0057] In some specific implementations, during the encoding step, the cut portions of the read to be encoded according to the second encoding subprocess are concatenated, and each nucleotide or base of the cut portion is encoded individually.

[0058] In some specific implementations, during the encoding step, each mismatch of an incompletely mapped read segment is encoded on a single byte.

[0059] In some specific implementations, during the encoding step, each mismatch of an incompletely mapped read is encoded as follows: the first two bits of the byte are used to encode the alternative nucleotide or base present in the read instead of the corresponding reference nucleotide or base in the reference sequence, and the last six bits of the byte are used to encode the position of the mismatch in the reference sequence, the position being calculated as an offset relative to the previous mismatch of the read.

[0060] In some specific implementations, during the encoding step, if the offset calculated between a given mismatch and the previous mismatch is greater than the maximum coded value, at least one false mismatch is inserted between the two mismatches until each offset between each of the mismatches and the at least one false mismatch is less than the maximum coded value. A false mismatch is defined as a mismatch in which, for the mismatch, the bits of the byte are used to encode the mismatch or to encode a nucleotide or base that is equal to a corresponding reference nucleotide or base in the reference sequence.

[0061] In some implementations, the initial step is to divide the list of read segments into read segment blocks, where each block begins with a header containing information required to decode the block, and the compression method is performed block by block.

[0062] In some specific implementations, the read segments have the same block size.

[0063] In some implementations, the final step is to provide a compressed file containing a list of coded read segments stored in the compressed file in the same order as the read segments stored in the initial file.

[0064] In some specific implementations, the threshold is equal to 31.

[0065] In some implementations, for each aligned read, a step is provided to determine whether the read contains at least one mismatch that corresponds to a situation where the sequencing machine cannot detect any bases or nucleotides.

[0066] In some specific implementations, for reads containing at least one mismatch corresponding to a situation where the sequencing machine cannot detect any bases or nucleotides, a step is provided to determine the number of such mismatches, and a step to compare said number with a reference threshold.

[0067] In some specific implementations, during the encoding step, if the number of such mismatches is greater than the reference threshold, each nucleotide or base of the read to be encoded according to the second encoding process will be encoded individually with 4 bits, and if the number of such mismatches is less than the reference threshold, each nucleotide or base of the read to be encoded according to the second encoding process will be encoded individually with 2 bits, and the encoding step also includes encoding a list of positions along the reference sequence corresponding to the positions of such mismatches in the reference sequence. Attached Figure Description

[0068] Figure 1 This is a flowchart illustrating an example of the compression method described in this article.

[0069] Figure 1A It is shown Figure 1 A flowchart providing a more detailed example of the compression method.

[0070] Figure 2 This is an illustration showing an example of a system used to implement one or more compression methods described herein.

[0071] Figure 2A This is an illustration showing another example of a system used to implement the compression method described herein.

[0072] Figure 2B This is an illustration showing another example of a system used to implement the compression method described herein.

[0073] Figure 3 This is a schematic diagram illustrating a first example of a read segment globally mapped to a reference sequence.

[0074] Figure 4 This is a schematic diagram illustrating a second example of a read segment globally mapped to a reference sequence when a false mismatch must be inserted.

[0075] Figure 5 It can be used for implementation. Figure 1 and Figure 1A An illustration of an example of the computational component of a system using a compression method.

[0076] Figure 6 Several bar graphs illustrating the experimental results of this disclosure are depicted.

[0077] Figure 7 Several bar graphs illustrating additional experimental results of this disclosure are depicted.

[0078] Figure 8 Several bar graphs illustrating additional experimental results of this disclosure are depicted. Detailed Implementation

[0079] The genomic sequences mentioned in this disclosure include, for example, but not limited to, nucleotide sequences, deoxyribonucleic acid (DNA) sequences, ribonucleic acid (RNA) sequences, and amino acid sequences. Although this disclosure describes genomic information in nucleotide sequence form in considerable detail herein, it should be understood that, as those skilled in the art will appreciate, the compression method according to the invention can also be implemented for other genomic sequences, although with some variations.

[0080] Genome sequencing information is generated by a sequencing machine in the form of nucleotide (or more generally, base) sequences represented by strings of letters from a defined vocabulary. The minimal vocabulary is represented by five symbols: {A, C, G, T, N}, representing the four types of nucleotides present in DNA: adenine, cytosine, guanine, and thymine. In RNA, thymine is replaced by uracil (U). N indicates that the sequencing machine cannot detect any base, and therefore the true nature of that position is indeterminate. Therefore, for the purposes of this disclosure, the symbol “N” refers to an undetermined base, and the number of “N”s in a read indicates the number of undetermined bases in that read.

[0081] The nucleotide sequences generated by sequencing machines can be referred to as "reads". The length of a read can range from tens to thousands of nucleotides. Some techniques generate reads in pairs, where the first read of the pair comes from one DNA strand and the second read of the pair comes from another DNA strand. Throughout this disclosure, a "reference sequence" is any sequence to which a read, consisting of a nucleotide or base sequence generated by a sequencing machine, can be compared / mapped. An example of such a reference sequence could actually be a reference genome, a sequence assembled by scientists as a representative example of the gene set of a species. However, reference sequences can also consist of synthetic sequences, which are envisioned simply to improve the compressibility of the reads, given that they will be further processed.

[0082] In some cases, sequencing machines may introduce errors into sequence reads, and it's worth noting that error symbols (i.e., indicating different nucleic acids) may be used to represent nucleic acids or bases actually present in the sequenced sample. This type of substitution error may ultimately be identified as a "mismatch" by the mapping and alignment module. This is because when the read is aligned with the reference sequence, the substitution error in the read may not match the corresponding position in the reference sequence. However, the meaning of "mismatch" is not limited to this type of situation. Rather, a "mismatch" can be any of the following bases or nucleotides in a read detected by the sequencing device that do not match the corresponding position in the reference sequence when the read is aligned with the reference sequence at a threshold level of accuracy. Such mismatches can include candidate variants, variants, or other differences between the aligned read and the reference sequence position.

[0083] This disclosure relates to a reference-based compression method that receives reads of nucleotide or base sequences as input, such reads having been previously aligned to a reference sequence by a mapping and alignment module to produce aligned reads. In some embodiments, the previously aligned reads may include reads aligned using a software mapping and alignment module that performs the mapping and alignment of the received reads to a reference sequence. For example, in some embodiments, the software mapper may perform hash-based mapping and alignment of the received reads by executing software instructions using one or more processors (such as one or more central processing units (CPUs), one or more graphics processing units (GPUs), or any combination thereof). In other embodiments, the previously aligned reads may include reads aligned using a hardware mapping and alignment module that performs the mapping and alignment of the received reads to a reference sequence. For example, in some implementations, the hardware mapping and comparison module can perform hash-based mapping and comparison using one or more hardware processors (such as one or more field-programmable gate arrays (FPGAs)) having hard-wired digital logic circuitry configured to perform hash-based mapping and comparison of received read segments.

[0084] The aligned reads are then stored as a list of reads in the initial file. The alignment reads and the manner in which they are stored once aligned in the initial file are not essential to this invention and are not the purpose of this disclosure. Each read is then encoded as a position on a reference sequence and a list of differences from said reference sequence. Each read can then be reconstructed from the alignment encoding information and the reference sequence using appropriate decompression software configured as described herein with respect to this disclosure.

[0085] In some implementations, the compression module of this disclosure can be implemented via software instructions executed by one or more CPUs or GPUs, hardwired digital logic circuitry executing one or more hardware processors, or a combination of both, to process and compress alignment reads. Before compressing the reads, they can be aligned with a reference sequence, regardless of certain types of errors introduced into the sequence reads, such as insertion errors or deletion errors. An insertion error occurs when one or more extra symbols that do not involve any actually present nucleic acids are inserted into a sequence read. A deletion error occurs when one or more symbols representing nucleic acids actually present in the sequencing sample are missing from a sequence read. More precisely, in the case of insertion or deletion errors in a given sequence read, the alignment software subsequently treats the resulting erroneous nucleic acids as substitution errors, also known as "mismatches." This preference in alignment software configuration allows for faster subsequent encoding, particularly providing a better trade-off between speed and compression ratio.

[0086] For each aligned read segment, the mapping and alignment module can generate and provide a read segment record. In some implementations, each read segment record can be directly provided as input from the mapping and alignment module to the compression module. In other implementations, each read segment record generated by the mapping and alignment module can be output and stored in memory or other storage devices. In such implementations, the compression module can then access and compress the stored read segment records.

[0087] Each read record generated, provided, or stored by the mapping and alignment module includes data generated by the mapping and alignment module describing the read represented by the read record. Such read records may include at least the following information: the absolute start position of the aligned read relative to the reference sequence, the length of the read, the alignment type of the read (such as whether the read is a mapped read or an unmapped read), the number of mismatches identified in the read, an indication of whether the read is a fully mapped read or an incompletely mapped read, the relative position of the possible mismatches in the read, and so on.

[0088] Although the examples described herein indicate that the data in the read segment records and the data contained therein are generated by the mapping and alignment module, this disclosure is not limited thereto. Instead, other intermediate modules between the mapping and alignment module and the compression module can be used to generate the read segment records and the data contained therein.

[0089] In some implementations, read segment records provided or stored by the mapping and alignment module can be provided or stored in a manner that preserves the order in which the read segment records generated by the mapping and alignment module are ordered. In some implementations, for example, each read segment record may also contain data indicating the placement of the read segment records in order of order. Such data indicating the placement of the read segment records may include, for example, a sequence_id. In some implementations, the sequence_id may be, for example, a number starting with "1" for the first read segment record generated by the mapping and alignment module, which then increments for each subsequent read segment record generated by the mapping and alignment module. The compression module of this disclosure can then access these read segment records and compress them in their current order without needing to reorder the read segment records into clusters for compression. Compressing these read segment records in a manner that preserves the initial order in which they are generated by the mapping and alignment module provides an advantage over conventional methods by achieving lossless compression of the read segment records, since even the order in which the read segment records are ordered is preserved. Furthermore, preserving the order of the read segment records during compression also makes verification of the read segment record compression easier.

[0090] Now refer to Figure 1 The compression method of this disclosure is described. In some specific implementations, for example, the method may be... Figure 2 The device 20 shown herein performs the operation. Device 20 may include at least one processor 22 and at least one memory 24 operatively coupled to at least one processor 22 to form a computing device. Memory 24 may store computer program code or software 26 containing computer-executable instructions that, when executed by processor 22, cause processor 22 to perform operations of a compression module, including performing multiple stages of one or more compression methods described herein. However, this disclosure is not necessarily limited to implementation by device 20.

[0091] For example, in some specific implementations, the compression method of this disclosure can be... Figure 2AThe illustrated device 20A is implemented. Device 20A is similar to device 20 because it also includes a processor 22 and at least one memory 24 operatively coupled to the processor 22 to form a computing device. The memory 24 of device 20A also stores computer program code or software 26 containing computer-executable instructions that, when executed by the processor 22, cause the processor 22 to perform operations including multiple stages of one or more compression methods described herein. However, device 20A further includes computer program code or software 28 containing computer-executable instructions that, when executed by the processor 22, cause the processor 22 to perform operations to implement the functionality of a mapping and alignment module. This mapping and alignment module, whose functionality is implemented via the execution of computer software instructions, can generate one or more alignment read segments 29 and store these alignment read segments 29 in the memory 24. Then, processor 22 can execute software instructions 26 of the compression module to access one or more aligned reads 29, and compress the one or more aligned reads 29 using multiple stages of one or more compression methods described herein. In some embodiments, device 20A may be a nucleic acid sequencing device.

[0092] As another example, in some specific implementations, the compression method of this disclosure can be derived from... Figure 2BThe illustrated device 20B is implemented. Device 20B differs from device 20 because it includes one or more hardware processors 22B, such as one or more field-programmable gate arrays (FPGAs). In this example, the one or more hardware processors may implement the functionality of multiple stages of one or more compression methods described herein, as well as the functionality of multiple stages of a mapping and comparison module in the hardware circuitry of the one or more hardware processors 22B. For example, hardware processor 22B may include hard-wired digital logic circuitry 26B configured as a compression module to perform multiple stages of one or more compression methods described herein. Similarly, hardware processor 22B may include hard-wired digital logic circuitry 28B configured to perform the operations of a mapping and comparison module configured to generate a comparison read segment 29B and store the comparison read segment 29B in memory 24. The hard-wired digital logic circuitry 26B configured as a compression module to implement the functionality of multiple stages of one or more compression methods described herein can access the comparison read segment 29B from memory 24 and compress the comparison read segment 29B using the compression methods described herein. In some implementations, device 20B may be a nucleic acid sequencing device. An initial file storing aligned read records as a read list is stored, for example, in the memory of device 20. In some implementations, this read list may include the plurality of aligned read records stored in the device's memory in a manner that preserves the sequence order of the plurality of aligned read records generated by the mapping and alignment module. This sequence order of the aligned read records may be the same order obtained at the end of the mapping and alignment phase.

[0093] In some implementations, the initial list of alignment reads can be divided into read block blocks. For example, in some implementations, the list of alignment reads can be divided into blocks of 50,000 reads. However, this particular value of 50,000 read blocks should not be construed as limiting the scope of this disclosure, as other values ​​can be used to achieve the same result in different implementations.

[0094] In some implementations, read segments can have the same block size. However, in other implementations, read segments can have different block sizes. In any case, each read segment block can begin with a header containing information needed to decode the block, such as the size of the block's contents in bytes, and / or an identifier for the block or its contents, and / or the number of read segments contained in the block. This allows for concatenation of compressed files, as well as streaming capabilities (each read segment block contains all the information needed to decode the read segments of the block). Furthermore, since the compression method can then be performed block by block, this also allows for multi-threaded processing of read segments, thus allowing for parallelization and some gains in processing time. If all read segments of a given block have the same length, the read segment length is also stored in the header; otherwise, a list of each read segment length is explicitly stored during the compression method.

[0095] return Figure 1 The method preferably includes an initial phase 2, wherein the device obtains the comparison read segment records from the memory of device 20, 20A, or 20B. In some embodiments, this may include: the device accessing a memory or other storage device that stores the multiple read segment records in a sequential order, maintaining the multiple read segment records generated by the mapping and comparison module. For example, the device may determine the next read segment record for compression based on the sequence_id of the previous read segment record and the sequence_id of one or more other read segment records stored in memory. In some embodiments, the sequence_id may be a numerical value that increments for each subsequent read segment record generated by the mapping and comparison module, and the compression module may maintain a counter that increments for each subsequent read segment record generated by the mapping and comparison module. Figure 1 The compression process increments with each iteration and provides an indication of the next read record that should be accessed at stage 2.

[0096] Each read record contains information about the type of read alignment. Information about the alignment type can include any information describing the mapping and alignment level of the read relative to a reference genome. In some embodiments, the alignment type can include complete alignment, incomplete alignment, or “unmapped” read alignment. A “complete alignment” or “fully mapped read” can include a read in which each nucleotide is mapped to and aligned with a portion of the reference genome. In some embodiments, a “complete alignment” or “fully mapped read” may have zero mismatches and zero undetermined “N” bases. In other embodiments, a “complete alignment” or “fully mapped read” may have zero mismatches but may have one or more undetermined “N” bases. Generally, the definition of “incomplete alignment” or “incompletely mapped read” depends on the meaning of “fully mapped read” implemented in a specific embodiment of the compression method described herein. For example, if an embodiment in which a fully mapped read may contain zero mismatches and zero unidentified N bases is used, then “incomplete alignment” or “incompletely mapped read” means any read that matches at least a portion of a reference sequence and includes at least one mismatch or at least one N. However, if, for example, an embodiment in which a fully mapped read may contain zero mismatches caused by one or more N bases is used, then “incomplete alignment” or “incompletely mapped read” means any read having at least one mismatch other than an undisclosed N base, and at least a portion of said read matches a portion of a reference sequence (according to this alternative definition of an incompletely mapped read, an incompletely mapped read may contain one or more N bases, provided it also contains one or more other mismatches). Therefore, how any particular system or embodiment is configured to recognize fully mapped reads will determine the meaning of incompletely mapped reads for that embodiment. “Unmapped reads” may include reads that have not yet been mapped to or aligned with a reference genome.

[0097] In some implementations, each read segment record includes multiple bit markers describing the attributes of the read segment. In some implementations, these multiple bit markers may be stored using one or more fields at the beginning of the read segment record. However, in other implementations, other fields of the read segment record may be used to store the multiple bit markers. Each of the multiple bit markers may use one of a plurality of values ​​to indicate the value of its corresponding read segment attribute. In some implementations, the following bit markers may be used to indicate the value of the read segment attribute of the read segment record:

[0098] - The first marker indicates whether the orientation is forward or reverse relative to the reference sequence.

[0099] - The second marker indicates whether the alignment is complete or incomplete.

[0100] - The third marker indicates whether the read segment contains at least one N.

[0101] - The fourth bit indicates whether the position information is encoded in 16-bit or 32-bit.

[0102] - The fifth marker indicates whether the read segment is mapped or unmapped.

[0103] Stages 4 through 12 are executed for each of the multiple read segments. If the read segments are grouped into blocks, stages 4 through 12 are executed for each read segment of each read segment block.

[0104] The compression method of this disclosure may include a next stage 4, which uses devices 20, 20A, or 20B to determine for each aligned read whether the read is fully mapped to or partially mapped to a reference sequence, or whether the read is unmapped to a reference sequence. In some embodiments, devices 20, 20A, and 20B may determine whether a read is fully mapped, partially mapped, or unmapped based on information received from the mapping and alignment module. This information may include, for example, whether the read represented by the obtained read record is mapped or unmapped; whether the read represented by the read record is fully mapped or partially mapped; an indication of the total number of mismatches (such as variant or sequencing errors, unidentified bases); or any combination thereof. In some embodiments, this information may be included within the obtained read record itself.

[0105] In some specific implementations, devices 20, 20A, or 20B can first determine whether the comparison read segment is mapped or unmapped. If device 20, 20A, or 20B determines that the comparison read segment is unmapped, the device can continue execution at stage 6. Figure 1 The process. Alternatively, if device 20, 20A, or 20B determines that the read segment is mapped, then device 20, 20A, or 20B can determine whether the read segment is incompletely mapped or fully mapped.

[0106] In some implementations, devices 20, 20A, or 20B can determine whether a read segment is incompletely mapped or fully mapped by evaluating the total number of mismatches in the read segment. In some implementations, this total number of mismatches can be provided by the mapping and alignment module and obtained from the acquired read segment records. In such implementations, if devices 20, 20A, or 20B determine that the total number of mismatches is zero, then devices 20, 20A, or 20B can determine at stage 4 that the acquired aligned read segment is a fully mapped read segment and can continue execution at stage 6. Figure 1The process. Alternatively, if at stage 4, device 20, 20A, or 20B determines that the total number of mismatches is greater than zero, then device 20, 20A, or 20B may determine at stage 4 that the read segment corresponding to the read segment record is an incompletely mapped read segment, and device 20, 20A, or 20B may continue execution at stage 6. Figure 1 The process.

[0107] However, it should be noted that the above-described embodiments are merely examples of how devices 20, 20A, or 20B can determine whether a matched read record is fully mapped, partially mapped, or unmapped. For example, in some embodiments, this determination can be based on information contained in the obtained read record and without comparing the number of mismatches to a zero threshold. By way of example, the read record may retain bit flags in its header or other portions indicating whether the read is mapped, unmapped, fully mapped, or partially mapped, etc. In such embodiments, devices 20, 20A, or 20B can determine at stage 4 whether the matched read record is mapped, unmapped, fully mapped, or partially mapped based on the bit flags of the obtained read record and without comparing the number of mismatches to a zero threshold. Other embodiments also fall within the scope of this disclosure. For example, it is conceivable that the following specific implementation could be adopted: information stored in a data structure different from the obtained read segment record can be accessed and treated as read segment bit markers or other data to indicate whether a particular read segment record is mapped, unmapped, fully mapped, or partially mapped.

[0108] In some implementations, step 4 may further include, for each incompletely mapped read, comparing the number of mismatches between the read and the reference sequence to a threshold 4a. This may include a total number of mismatches, where the total number of mismatches includes the sum of any differences (including variants, sequencing errors, and unidentified base N) between the aligned read and the reference sequence. In some implementations, the number of mismatches may be provided by the mapping and alignment module and obtained from the read record.

[0109] In some embodiments, the threshold may be 31. This particular value can be chosen to provide the best possible trade-off for storing the number of mismatches in a sufficiently compact manner, as will be better understood later with respect to stage 12. Indeed, it has been statistically observed that, in the vast majority of cases, incompletely mapped reads have fewer than 31 mismatches. The principle behind this choice is to encode the most frequent cases in the most compact way, leaving a very small number of degraded cases. However, while a threshold of 31 mismatches may be used in some embodiments such as short-read embodiments (where a read length of approximately 150 nucleotides or bases may be advantageous), this disclosure is not limited to those embodiments in which the threshold is equal to 31. Rather, for other embodiments, it may be desirable to use a threshold higher than 31. For example, although aspects (e.g., a threshold of 31 mismatches) may be intended for compressing read records representing reads generated by a short-read sequencer, it is contemplated that the genomic data compression method of the present invention can be used in other embodiments, such as for compressing read records generated by a long-read sequencer. Therefore, in such specific implementations, where the read segment is represented by a read segment record that is significantly longer than 150 nucleotides or bases, the threshold can be set to a value higher than 31 to achieve the functionality of the compression method of this disclosure for long read systems.

[0110] If a read is determined to be incompletely mapped with a number of mismatches below a threshold, determination stage 4 may further include an additional determination as to whether the read is globally or locally mapped to the reference sequence. A “globally mapped read” is an incompletely mapped read whose entire sequence (including the start and end points of the read) is not fully mapped to the reference sequence. A “locally mapped read” is an incompletely mapped read containing segments of nucleotides or bases that are not fully mapped to the reference sequence. Therefore, the segments of nucleotides or bases correspond to a portion of the initial read.

[0111] In some embodiments, the compression method may further include a stage 6, which, for each aligned read, determines whether the read contains at least one undetermined base "N," i.e., whether the read contains at least one mismatch corresponding to a situation where the sequencing machine cannot detect any base or nucleotide. For each read containing at least one "N," the method then includes a stage 8, which determines the number of such undetermined base "N"; and a stage 10, which compares the number of undetermined base "N" with a reference threshold. In some embodiments, the reference threshold may be equal to 31. However, in other embodiments, the reference threshold may be set to other values.

[0112] Regardless of the outcome of stage 4, the method includes a next stage 12, which encodes the reads based on the determination. More specifically, reads determined to be perfectly mapped to a reference sequence are encoded according to a first encoding process, regardless of whether the read does not contain an undetermined base "N" or has a number of undetermined bases "N" below a reference threshold. Reads determined to be unmapped or determined to be perfectly mapped but have a number of undetermined bases "N" above a reference threshold are encoded according to a second encoding process, wherein each nucleotide or base is encoded individually, regardless of whether the nucleotide or base is aligned or unaligned. Reads determined to be incompletely mapped are encoded according to a second or third encoding process. More specifically, reads determined to be incompletely mapped with a mismatch number above a threshold are encoded according to a second encoding process. If a read is determined to be incompletely mapped with a mismatch number below a threshold, then if the read does not contain N or has a number of N below a reference threshold, the read is encoded according to a third encoding process. If not, that is, if the number of read segments is greater than N, then the read segments are encoded according to the second encoding process.

[0113] Regardless of whether a given read has been determined to be fully mapped, partially mapped, or unmapped, if the read contains at least one N but has a number of Ns below a reference threshold, then encoding stage 12 includes encoding a list of positions along a reference sequence corresponding to the positions of Ns in the reference sequence. This list of positions is then stored in the memory of a computing device that implements the compression method. If the read contains at least one N but has a number of Ns below the reference threshold, it is encoded according to a second encoding process, and each nucleotide or base of the read is encoded separately with 2 bits.

[0114] If a read contains at least one N but has a number of Ns greater than a reference threshold, the read will in any case be encoded according to a second encoding process, with each nucleotide or base of the read encoded individually in 4-bit increments. In this case, encoding stage 12 does not include encoding and storing a list of N positions in the reference sequence. In fact, each N mismatch is then directly encoded according to the second encoding process in exactly the same manner as the other nucleotides or bases of the read.

[0115] The first and third encoding processes each comprise different sets of descriptors. Each set of descriptors unilaterally represents a read associated with its corresponding encoding process, and each of the first and third encoding processes is a simplified information entropy encoding process. More precisely, the third encoding process comprises a first encoding subprocess and a second encoding subprocess. The first encoding subprocess encodes reads that are determined to be globally mapped but are not fully mapped during stage 4. The second encoding subprocess encodes reads that are determined to be locally mapped during stage 4. The first and second encoding subprocesses each comprise different sets of descriptors, each set unilaterally representing a read associated with its corresponding encoding subprocess.

[0116] Therefore, the alignment information for encoding each read segment and enabling the reconstruction of the entire read segment sequence during data decompression depends on the corresponding encoding process or subprocess for the read segment.

[0117] For example, in some specific implementations, the first set of descriptors used for the first encoding process may include:

[0118] The absolute start position of the fully mapped read segment relative to the reference sequence (encoded in 16-bit or 32-bit), and

[0119] o. The length of the read segment (encoded relative to the length of the previous read segment using differential coding, where the variable length code is in the range of 2 to 34 bits).

[0120] As another example, in some specific implementations, the second set of descriptors used for the first encoding subprocess may include:

[0121] o The absolute start position of the incompletely mapped read segment relative to the reference sequence (encoded in 16-bit or 32-bit).

[0122] The length of the read segment (encoded using differential coding relative to the length of the previous read segment, where the variable-length code ranges from 2 to 34 bits), and

[0123] List of mismatches for o-read segments.

[0124] As another example, in some specific implementations, the third set of descriptors used for the second encoding sub-process may include:

[0125] The incompletely mapped portion of the o-read segment relative to the absolute start position of the reference sequence – also known as the local alignment start position (encoded in 16-bit or 32-bit)

[0126] o. The length of the read segment (encoded using differential coding relative to the length of the previous read segment, where the variable length code ranges from 2 to 34 bits).

[0127] List of mismatches for o-read segments, and

[0128] The length of the cut portion of the read segment that is not part of the alignment (encoded in 8 bits for each cut portion).

[0129] Preferably, the mismatch list encoded in the first and second sub-processes may include a header. For example, in some embodiments, the header may be encoded using bit tags and on a single byte. In such embodiments, the first five bits of the one-byte header may be used to encode the number of mismatches contained in the read. In embodiments where the threshold is equal to 31, the number of mismatches may be in the range between 0 and 31. One bit of the one-byte header may be used to encode whether the incompletely mapped read is globally mapped or locally mapped. Another bit of the one-byte header may be used to encode whether a 2-bit pattern is activated for the second encoding process. The last bit of the one-byte header may be used to encode whether a 4-bit pattern is activated for the second encoding process. In some embodiments, for each read encoded according to the second encoding sub-process during encoding phase 12, the cut portions of the read (i.e., those portions that are not part of a local alignment) are concatenated, and each nucleotide or base of the cut portions is encoded individually. In some embodiments, each nucleotide or base of such cut portions of the read is encoded individually with 2 bits.

[0130] In some specific implementations, each mismatch encoded in the mismatch list of the incompletely mapped read segment (i.e., encoded according to the first or second encoding sub-procedure) can be encoded on one byte. More precisely, each mismatch of the incompletely mapped read segment to be encoded according to the first or second encoding sub-procedure can be encoded as follows:

[0131] The first two bits of the 'o' byte are used to encode substituted nucleotides or bases present in the read segment, rather than the corresponding reference nucleotides or bases in the reference sequence.

[0132] The last six bits are used to encode the position of a mismatch in the reference sequence, which is calculated as an offset relative to the previous mismatch in the read segment. This calculated position can be a relative position of the mismatch, except for the first mismatch in the read segment whose absolute position is encoded. Therefore, the range of this offset (encoded in 6 bits) can be [0-63].

[0133] Completed Figure 1 The encoded or compressed records generated during the process can be stored in the device's memory or other storage device. In some specific implementations, such encoded or compressed records can be stored in the device's memory or other storage device in a manner that maintains the sequential order of the read segment records. This helps ensure that the compression of the comparison read segment records is lossless, because even the initial sequential order of the comparison read segment records is preserved.

[0134] The comparison reading record was obtained at stage 102.

[0135] refer to Figure 1A The compression method 100A is described in more detail. Figure 1 Compression methods. Compression method 100A, performed by device 20, 20A, or 20B, can begin in an initial phase 102, which includes obtaining aligned read records (hereinafter also referred to as "obtained read records" or "unmapped reads" / "mapped reads" / "fully mapped reads" / "partially mapped reads," based on the subsequent classification of read records obtained during the execution of method 100A). In some embodiments, the aligned read records can be obtained from multiple aligned read records stored in a manner that preserves their initial order provided by the sequencing device. Therefore, the entire operation of the mapping and alignment module and the compression module can maintain the read records in their initial order provided by the sequencing device. In some embodiments, the aligned read records can be stored to preserve their initial order by using a sequence_id, which is stored with each aligned read record and increments with each aligned read record generated by the mapping and alignment module.

[0136] At stage 104, it is determined whether the read corresponding to the aligned read record is a complete mapping, an incomplete mapping, or... Unmapped

[0137] The compression method of this disclosure may include a next stage 104, which uses devices 20, 20A, or 20B to determine whether the acquired read record corresponds to a read that is fully mapped to a reference sequence, a read that is not fully mapped to a reference sequence, or a read that is not mapped to a reference sequence. In some embodiments, devices 20, 20A, and 20B may determine whether a read is fully mapped, partially mapped, or unmapped based on information received from a mapping and alignment module. This information may include, for example, whether the read represented by the acquired read record is mapped or unmapped; whether the read represented by the read record is fully mapped or partially mapped; an indication of the total number of mismatches (such as variant or sequencing errors, unidentified bases); or any combination thereof. In some embodiments, this information may be included within the read record itself.

[0138] In some specific implementations, devices 20, 20A, or 20B can first determine at stage 104 whether the comparison read segment is mapped or unmapped. If devices 20, 20A, or 20B determine that the comparison read segment is unmapped, the devices can continue execution at stage 120. Figure 1AThe process is 100A. Alternatively, if devices 20, 20A, or 20B determine that a read segment is mapped, then devices 20, 20A, or 20B may further determine during stage 104 whether the read segment is incompletely mapped or fully mapped.

[0139] In some embodiments, devices 20, 20A, or 20B can determine whether a read segment is incompletely mapped or fully mapped by evaluating the number of mismatches in the read segment during stage 104. In some embodiments, the number of mismatches can be provided by the mapping and alignment module and obtained from the read segment record. The number of mismatches can be recorded in different ways for different embodiments. In some embodiments, the number of mismatches at stage 104 may not include the number of undetermined N bases. In other embodiments, the number of mismatches determined at stage 104 may include the sum of the number of mismatches and the number of undetermined N bases.

[0140] exist Figure 1A In this example, it is assumed that the undetermined base N is not a mismatch. Therefore, a fully mapped read may include 0 mismatches and one or more undetermined base Ns. Thus, in this embodiment, an incompletely mapped read will need to have at least one mismatch and may or may not have any undetermined base Ns. However, in other embodiments, Figure 1A The process can be modified by assuming that the presence of N in the read segment may be a mismatch. In such an implementation, a read segment can only be identified as a fully mapped read segment if it is determined to have 0 mismatches and 0 undetermined N bases, while read segments with 0 mismatches and one or more undetermined N bases are classified as incompletely mapped read segments.

[0141] In the first specific implementation, at stage 104, if device 20, 20A, or 20B determines that the total number of mismatches is equal to zero and the total number of undetermined base Ns is zero or more, then device 20, 20A, or 20B can determine at stage 4 that the obtained alignment read is a fully mapped read and can continue execution at stage 116. Figure 1A Process 100A. Alternatively, in this first embodiment, if during stage 104, devices 20, 20A, or 20B determine that the total number of mismatches is greater than zero and the total number of undetermined base Ns is zero or more, then devices 20, 20A, or 20B may determine during stage 104 that the read corresponding to the obtained read record is an incompletely mapped read, and devices 20, 20A, or 20B may continue execution at stage 106. Figure 1A The process is 100A.

[0142] In the second and alternative embodiments, at stage 104, devices 20, 20A, or 20B will only determine that a read is a fully mapped read if the total number of mismatches is zero and the total number of undetermined base Ns is zero, and in such a case, devices 20, 20A, or 20B can continue execution at stage 116. Figure 1A Process 100A. Alternatively, in this second embodiment, if during stage 104, devices 20, 20A, or 20B determine that the total number of mismatches is greater than zero or the total number of undetermined base Ns is greater than zero, then devices 20, 20A, or 20B may determine during stage 104 that the read corresponding to the obtained read record is an incompletely mapped read, and devices 20, 20A, or 20B may continue execution at stage 106. Figure 1A The process is 100A.

[0143] However, it should be noted that the above-described embodiments are merely examples of how devices 20, 20A, or 20B can determine at stage 104 whether a read corresponding to the acquired read record is fully mapped, partially mapped, or unmapped. For example, in some embodiments, this determination may alternatively be based on information contained in the acquired read record and performed without comparing the number of mismatches to a threshold, without comparing the number of undetermined base Ns to a threshold, or without performing either of these comparisons. By way of example, the read record may retain bit markers in its header or other portions indicating whether the read is mapped, unmapped, fully mapped, or partially mapped, etc. In such embodiments, devices 20, 20A, or 20B may determine at stage 4 whether the aligned read record is mapped, unmapped, fully mapped, or partially mapped based on the bit markers of the read record and without comparing the number of mismatches or undisclosed base Ns to a threshold. Other embodiments also fall within the scope of this disclosure. For example, it is conceivable that the following specific implementation could be adopted: information stored in a data structure different from the read segment record can be accessed and treated as read segment bit markers or other data to indicate whether the read segment corresponding to a particular read segment record is mapped, unmapped, fully mapped, or partially mapped.

[0144] The "Incomplete Segment Mapping" branch in stage 104

[0145] If device 20, 20A, or 20B determines at stage 104 that the read corresponding to the acquired read record is an incompletely mapped read, then device 20, 20A, or 20B may determine at stage 106 whether the number of differences between the incompletely mapped read and the reference sequence exceeds a first threshold. This may include the total number of mismatches, where the total number of mismatches includes the sum of any differences (including variants, sequencing errors, and unidentified base N) between the aligned read and the reference sequence. In other embodiments, the number of differences at stage 106 may include only the number of mismatches without taking into account the number of unidentified base Ns. In some embodiments, the number of mismatches may be provided by the mapping and alignment module and obtained from the read record.

[0146] In some implementations, the first threshold may be 31. This particular value can be chosen to provide the best possible trade-off for storing the number of mismatches in a sufficiently compact manner, as will be better understood later regarding subsequent stages. Indeed, it has been statistically observed that, in the vast majority of cases, incompletely mapped reads have fewer than 31 mismatches. The principle behind this choice is to encode the most frequent cases in the most compact way, leaving a very small number of degraded cases. However, while using a first threshold of 31 can achieve certain advantages, this disclosure is not limited to those implementations where the first threshold is equal to 31. Rather, for other implementations, it may be desirable to use a threshold higher than 31. For example, although aspects (e.g., a threshold of 31 mismatches) may be intended for compressing read records representing reads generated by short read sequencers, it is contemplated that the genomic data compression method of the present invention can be used in other implementations, such as for compressing read records generated by long read sequencers. Therefore, in such specific implementations, where the read segment is represented by a read segment record that is significantly longer than 150 nucleotides or bases, the threshold can be set to a value higher than 31 to achieve the functionality of the compression method of this disclosure for long read systems.

[0147] The "Yes" branch in stage 106

[0148] If device 20, 20A, or 20B determines at stage 106 that the number of differences between the incompletely mapped read and the reference sequence exceeds a first threshold, the device may continue with process 100A at stage 114. At stage 114, device 20, 20A, or 20B may determine whether the number of undetermined bases "N" in the incompletely mapped read exceeds a second threshold. In some embodiments, the second threshold may also be equal to 31. However, similar to the first threshold, the second threshold of this disclosure is not limited to the value 31. Rather, any value (including values ​​higher than 31) may be used for the second threshold based on the length of the read disclosed in this embodiment. Furthermore, it is not necessary for the first and second thresholds to use the same threshold.

[0149] The "Yes" branch in stage 114

[0150] If device 20, 20A, or 20B determines that the number of undisclosed "N" bases in an incompletely mapped read exceeds a second threshold, then device 20, 20A, or 20B may determine to encode the incompletely mapped read using the second encoding module 110, thereby using a second encoding process to encode the incompletely mapped read. The second encoding process is related to the above... Figure 1 The second encoding process is the same, wherein each nucleotide or base is encoded individually, regardless of whether the nucleotide or base is aligned or not. In some specific embodiments, since devices 20, 20A, or 20B determine at stage 114 that the number of unidentified bases "N" exceeds a second threshold, devices 20, 20A, or 20B can use a second encoding module to encode the read segment as 4 bits 110a using the second encoding process 110. Once the read segment is encoded using 4 bits of encoding 110a using the second encoding process 110, devices 20, 20A, or 20B can store the encoded read segment in memory or other storage device at stage 122. Devices 20, 20A, or 20B can determine at stage 124 whether another sequentially ordered aligned read segment exists to be compressed. And, if another sequentially ordered aligned read segment exists to be compressed, devices 20, 20A, or 20B can perform the operation of stage 102 to obtain the next sequentially ordered aligned read segment record and execute process 100A again. Then, devices 20, 20A, or 20B continue to iteratively execute process 100A until no more sequentially ordered comparison read records are identified at stage 124. After such determination, process 100A can terminate at stage 126.

[0151] Phase 114 "No" branch

[0152] If device 20, 20A, or 20B determines during stage 114 that the number of undisclosed "N" bases in the incompletely mapped read does not exceed a second threshold, then device 20, 20A, or 20B can determine to encode the incompletely mapped read using the second encoding module 110, thereby using a second encoding process to encode the incompletely mapped read. The second encoding process is related to the above... Figure 1The second encoding process is the same, wherein each nucleotide or base is encoded individually, regardless of whether the nucleotide or base is aligned or not. In some specific embodiments, since devices 20, 20A, or 20B determine at stage 114 that the number of unidentified bases "N" does not exceed a second threshold, devices 20, 20A, or 20B can use a second encoding module to encode the read segment as 2 bits 110b using the second encoding process 110. Once the read segment is encoded using the 2-bit encoding 110b using the second encoding process 110, devices 20, 20A, or 20B can store the encoded read segment in memory or other storage device at stage 122. Devices 20, 20A, or 20B can determine at stage 124 whether another sequentially ordered aligned read segment exists to be compressed. And, if another sequentially ordered aligned read segment exists to be compressed, devices 20, 20A, or 20B can perform the operation of stage 102 to obtain the next sequentially ordered aligned read segment record and execute process 100A again. Then, devices 20, 20A, or 20B continue to iteratively execute process 100A until no more sequentially ordered comparison read records are identified at stage 124. After such determination, process 100A can terminate at stage 126.

[0153] The "No" branch in stage 106

[0154] If device 20, 20A, or 20B determines at stage 106 that the number of differences between the incompletely mapped read and the reference sequence does not exceed a first threshold, then device 20, 20A, or 20B may continue with process 100A at stage 108. At stage 108, device 20, 20A, or 20B may determine whether the incompletely mapped read includes more than a second threshold number of undetermined bases "N".

[0155] The "Yes" branch in stage 108

[0156] If device 20, 20A, or 20B determines at stage 108 that the number of undisclosed "N" bases in the incompletely mapped read exceeds a second threshold, then device 20, 20A, or 20B may determine at stage 108 to encode the incompletely mapped read using the second encoding module 110, thereby using a second encoding process to encode the incompletely mapped read. The second encoding process is related to the above... Figure 1The second encoding process is the same, wherein each nucleotide or base is encoded individually, regardless of whether the nucleotide or base is aligned or not. In some specific embodiments, since devices 20, 20A, or 20B determine at stage 108 that the number of unidentified bases "N" exceeds a second threshold, devices 20, 20A, or 20B can use a second encoding module to encode the read segment into 4 bits 110a using the second encoding process 110. Once the read segment is encoded using 4 bits of encoding 110a using the second encoding process 110, devices 20, 20A, or 20B can store the encoded read segment in memory or other storage device at stage 122. Devices 20, 20A, or 20B can determine at stage 124 whether another sequentially ordered aligned read segment exists to be compressed. And, if another sequentially ordered aligned read segment exists to be compressed, devices 20, 20A, or 20B can perform the operation of stage 102 to obtain the next sequentially ordered aligned read segment record and execute process 100A again. Then, devices 20, 20A, or 20B continue to iteratively execute process 100A until no more sequentially ordered comparison read records are identified at stage 124. After such determination, process 100A can terminate at stage 126.

[0157] Phase 108 "No" branch

[0158] If device 20, 20A, or 20B determines during stage 108 that an incompletely mapped read segment includes an undetermined number of undetermined bases “N” that do not meet the second threshold, then device 20, 20A, or 20B may use third encoding module 112 to encode the incompletely mapped read segment using a third encoding process. Figure 1A The third encoding process in the above reference Figure 1 The process is the same as the third encoding process described above, and uses the same descriptor as the third encoding process described above. Once the read segment is encoded using the third encoding process of the third encoding module 112, the device 20, 20A, or 20B can store the encoded read segment in memory or other storage device at stage 122. The device 20, 20A, or 20B can determine at stage 124 whether there is another sequentially ordered comparison read segment to be compressed. And, if there is another sequentially ordered comparison read segment to be compressed, the device 20, 20A, or 20B can perform the operation of stage 102 to obtain the next sequentially ordered comparison read segment record and execute process 100A again. Then, the device 20, 20A, or 20B continues to iteratively execute process 100A until no sequentially ordered comparison read segment record is identified at stage 124. After such determination, process 100A can terminate at stage 126.

[0159] Phase 104 "Complete Segment Mapping" branch

[0160] Alternatively, if it is determined at stage 104 that the read corresponding to the obtained read record is a fully mapped read, then devices 20, 20A, or 20B may determine at stage 116 whether the fully mapped read includes an undetermined number of unidentified bases "N" exceeding a second threshold. In some embodiments, the second threshold may also be equal to 31. However, similar to the first threshold, this disclosure is not limited to the second threshold 31. Rather, any value (including values ​​higher than 31) may be used for the second threshold based on the length of the read disclosed in this embodiment. Furthermore, it is not necessary for the first and second thresholds to use the same threshold.

[0161] The "No" branch in stage 116

[0162] If device 20, 20A, or 20B determines at stage 116 that the fully mapped read does not contain more than the second threshold number of undetermined bases "N", then device 20, 20A, or 20B may determine to encode the read using the first encoding process with the first encoding module 122. If the fully mapped read does not contain any undetermined bases "N", then the first encoding module 122 performs the same process as described above. Figure 1 The first encoding process described above is the same as the first encoding process described above, and uses the same descriptors as the first encoding process described above. Alternatively, if the fully mapped read segment includes one or more "N", then the first encoding module 122 uses the referenced above. Figure 1 The first encoding process uses the same descriptor used in the first encoding process described above to encode the fully mapped read segment. Furthermore, in a specific implementation where the fully mapped read segment includes one or more N (but less than a second threshold number of N), the first encoding module 118 may also store a list of undetermined positions of base N on the read segment.

[0163] Once the read segment is encoded using the first encoding module 118, devices 20, 20A, or 20B can store the encoded read segment in memory or other storage device at stage 122. Devices 20, 20A, or 20B can determine at stage 124 whether another sequentially ordered comparison read segment exists to be compressed. If such a segment exists, devices 20, 20A, or 20B can perform the operation of stage 102 to obtain the next sequentially ordered comparison read segment record and execute process 100A again. Devices 20, 20A, or 20B then continue iteratively executing process 100A until no more sequentially ordered comparison read segment records are identified at stage 124. After this determination, process 100A can terminate at stage 126.

[0164] The "Yes" branch in stage 116

[0165] However, if the device determines at stage 116 that the read segment does indeed include more than the number of undetermined bases “N” than the second threshold, then the device 20, 20A, or 20B can use the second encoding module 110 to encode the read segment as 4 bits 110a using the second encoding process. Figure 1A The second encoding process is similar to the one mentioned above. Figure 1 The second encoding process is the same, in which each nucleotide or base is encoded individually, regardless of whether the nucleotide or base is aligned or not. Once the read is encoded using 4-bit encoding 110a using the second encoding process 110, devices 20, 20A, or 20B can store the encoded read in memory or other storage device at stage 122. Devices 20, 20A, or 20B can determine at stage 124 whether another sequentially ordered aligned read exists to be compressed. And, if another sequentially ordered aligned read exists to be compressed, devices 20, 20A, or 20B can perform the operation of stage 102 to obtain the next sequentially ordered aligned read record and execute process 100A again. Devices 20, 20A, or 20B then continue to iteratively execute process 100A until no more sequentially ordered aligned read records are identified at stage 124. After such determination, process 100A can terminate at stage 126.

[0166] Phase 104 "Unmapped reads" branch

[0167] Alternatively, if at stage 104 it is determined that the read corresponding to the obtained read record is an unmapped read, then devices 20, 20A, or 20B may at stage 120 determine whether the unmapped read includes an undetermined number of unidentified bases "N" exceeding a second threshold. In some embodiments, the second threshold may also be equal to 31. However, similar to the first threshold, this disclosure is not limited to the second threshold 31. Rather, based on the length of the read disclosed in this embodiment, any value (including values ​​higher than 31) may be used for the second threshold. Furthermore, it is not necessary for the first and second thresholds to use the same threshold.

[0168] Phase 120 "No" branch

[0169] If device 20, 20A, or 20B determines at stage 120 that the unmapped read segment does not contain more than the number of undetermined bases "N", then device 20, 20A, or 20B can determine to encode the read segment using the second encoding module 110 and the second encoding process. The second encoding process is related to the above... Figure 1The second encoding process is the same, wherein each nucleotide or base is encoded individually, regardless of whether the nucleotide or base is aligned or not. In some specific embodiments, since devices 20, 20A, or 20B determine at stage 120 that the number of unidentified bases "N" does not exceed a second threshold, devices 20, 20A, or 20B can use a second encoding module to encode the read segment as 2 bits of 110b using the second encoding process 110. Once the read segment is encoded using 2 bits of 110b using the second encoding process 110, devices 20, 20A, or 20B can store the encoded read segment in memory or other storage device at stage 122. Devices 20, 20A, or 20B can determine at stage 124 whether another sequentially ordered aligned read segment exists to be compressed. And, if another sequentially ordered aligned read segment exists to be compressed, devices 20, 20A, or 20B can perform the operation of stage 102 to obtain the next sequentially ordered aligned read segment record and execute process 100A again. Then, devices 20, 20A, or 20B continue to iteratively execute process 100A until no more sequentially ordered comparison read records are identified at stage 124. After such determination, process 100A can terminate at stage 126.

[0170] The "Yes" branch in stage 120

[0171] However, if the device determines at stage 120 that the unmapped read segment does indeed include more than the number of undetermined bases “N” than the second threshold, then the device 20, 20A, or 20B may use the second encoding module 110 to encode the read segment as 4 bits 110a using the second encoding process. Figure 1A The second encoding process is similar to the one mentioned above. Figure 1 The second encoding process is the same, in which each nucleotide or base is encoded individually, regardless of whether the nucleotide or base is aligned or not. Once the read is encoded using 4-bit encoding 110a using the second encoding process 110, devices 20, 20A, or 20B can store the encoded read in memory or other storage device at stage 122. Devices 20, 20A, or 20B can determine at stage 124 whether another sequentially ordered aligned read exists to be compressed. And, if another sequentially ordered aligned read exists to be compressed, devices 20, 20A, or 20B can perform the operation of stage 102 to obtain the next sequentially ordered aligned read record and execute process 100A again. Devices 20, 20A, or 20B then continue to iteratively execute process 100A until no more sequentially ordered aligned read records are identified at stage 124. After such determination, process 100A can terminate at stage 126.

[0172] Figure 3An example is provided of encoding mismatches in a read segment according to the first encoding sub-process. This read segment is an incompletely mapped read segment, which is globally mapped to the reference sequence. This read segment has two mismatches:

[0173] The first mismatch, located at position 12 of the read, involves the substitution of a T nucleotide for an A nucleotide in the reference sequence.

[0174] The second mismatch occurs at position 21 of the read, where a G nucleotide replaces a C nucleotide in the reference sequence.

[0175] The mismatch list for this read segment is then encoded as:

[0176] o<12,T>, where the value "12" corresponds to the absolute position of the first mismatch in the read segment, and

[0177] o<9,G>, where the value “9” corresponds to the relative position of the second mismatch in the read segment, that is, the offset between the second mismatch and the first mismatch.

[0178] For example, <12,T> can be converted to the value "51" (encoded on 1 byte), and <9,G> can be converted to the value "38" (encoded on 1 byte). This byte encoding is obtained as follows:

[0179] Offset position x4 + nucleotide value (where A=0, C=1, G=2, T=3)

[0180] Preferably, for each incompletely mapped read segment to be encoded according to the first or second encoding sub-process, if the offset calculated between a given mismatch in the read segment and the previous mismatch is greater than the maximum coded value, at least one “false” mismatch is inserted between the two mismatches until each offset between each of the mismatches and the at least one “false” mismatch is less than the maximum coded value. A “false” mismatch is defined as a mismatch for which bits of a byte are used to encode the mismatch, or to encode a nucleotide or base that is equal to a corresponding reference nucleotide or base in the reference sequence. In some embodiments, the maximum coded value is equal to 63, corresponding to the maximum value that can be encoded in 6 bits. However, this disclosure is not limited to embodiments with a maximum coded value of 63. For embodiments with a maximum coded value greater than 63, additional bits may be used to encode the value. In such embodiments, this may require, for example, adjusting the length of other bits in the header used for the read segment, increasing the header size by more than one byte, or a combination of both. Therefore, the features of the algorithm disclosed herein are flexible for specific use cases, but specific implementations that may be subject to design changes may result in corresponding performance trade-offs, which may be acceptable and even beneficial in certain circumstances of any particular specific implementation.

[0181] Figure 4 An example is provided of encoding mismatches in a read segment according to the first encoding sub-process when "false" mismatches must be inserted. This read segment is an incompletely mapped read segment, which is globally mapped to the reference sequence. This read segment has two mismatches:

[0182] The first mismatch, located at position 22 of the read, involves the substitution of a T nucleotide for an A nucleotide in the reference sequence.

[0183] The second mismatch occurs at position 134 of the read, where a G nucleotide replaces a C nucleotide in the reference sequence.

[0184] The positional offset between the second mismatch and the first mismatch is 112, which is greater than the maximum codeable value of 63. Therefore, a "false" mismatch must be inserted between the two mismatches such that each offset between the false mismatch and each of the mismatches is less than the maximum codeable value. A false mismatch with a T nucleotide (corresponding to a "true" T nucleotide in the reference sequence) is inserted, for example, at position 85 of the read. The calculated positional offset between this false mismatch and the first mismatch is 63, which corresponds to the maximum codeable value. The calculated positional offset between the second mismatch and this false mismatch is 49, which is less than 63.

[0185] The mismatch list for this read segment is then encoded as:

[0186] o<22,T>, where the value "22" corresponds to the absolute position of the first mismatch in this segment.

[0187] o<63,T>, where the value "63" corresponds to the relative position of the "false" mismatch in this reading segment, i.e., the offset between the "false" mismatch and the first mismatch, and

[0188] o<49,G>, where the value “49” corresponds to the relative position of the second mismatch in this reading segment, that is, the offset between the second mismatch and the “false” mismatch.

[0189] For example, <22,T> can be converted to the value "91" (encoded on 1 byte), <63,T> can be converted to the value "255" (encoded on 1 byte), and <49,G> can be converted to the value "198" (encoded on 1 byte). This byte encoding is obtained as follows:

[0190] Offset position x4 + nucleotide value (where A=0, C=1, G=2, T=3)

[0191] The method includes a final step 14 of providing a compressed file containing a list of coded read segments. The coded read segments are stored in the compressed file in the same order as the read segments stored in the initial uncompressed file. Each read segment can then be reconstructed from the alignment of the coded information and a reference sequence using appropriate decompression software and / or methods configured according to the invention.

[0192] Although the exemplary architecture of the reference computing device 20 (for illustrative purposes) Figure 2 The invention described herein (as shown in the illustration) can be implemented in hardware, software, firmware, or any combination thereof. When implemented in software, the computer program code can be stored on a computer medium and executed by a hardware processing unit including one or more processors, which is similar to using... Figure 2 The same applies to device 20. It should be understood that, as used herein, the term "processor" is intended to include one or more processing devices, including signal processors, microprocessors, microcontrollers, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other types of processing circuitry systems, and portions or combinations of elements of such circuitry systems. Additionally, as used herein, the term "memory" is intended to include electronic memory associated with a processor, such as random access memory (RAM), read-only memory (ROM), or other types of memory used in any combination.

[0193] Therefore, the software instructions or code used to perform the methods and protocols described herein may be stored in one or more of the associated memory devices (e.g., ROM, fixed memory, or removable memory) and loaded into RAM and executed by the processor when ready for use.

[0194] The technology disclosed herein can be implemented in a wide variety of devices or equipment, including, for example, mobile phones, computers, servers, tablets, and similar devices.

[0195] Although illustrative embodiments of the invention have been described herein with reference to the accompanying drawings, it should be understood that the invention is not limited to those exact embodiments, and various other changes and modifications can be made by those skilled in the art without departing from the scope or spirit of the invention.

[0196] Figure 5 It can be used for implementation. Figure 1 and Figure 1A An illustration of an example of the computational component of a system using a compression method.

[0197] Computing device 500 is intended to represent various forms of digital computers, such as laptops, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframes, and other suitable computers. Computing device 550 is intended to represent various forms of mobile devices, such as personal digital assistants, mobile phones, smartphones, and other similar computing devices. Furthermore, computing device 500 or 550 may include a Universal Serial Bus (USB) flash drive. The USB flash drive may store an operating system and other applications. The USB flash drive may include input / output components, such as a wireless transmitter or USB connector that can be plugged into a USB port of another computing device. The components shown herein, their connections and relationships, and their functions are intended only as examples and are not intended to limit the specific implementations of the invention described and / or claimed in this document.

[0198] Computing device 500 includes a processor 502, a memory 504, a storage device 508, a high-speed interface 508 connected to the memory 504 and a high-speed expansion port 510, and a low-speed interface 512 connected to a low-speed bus 514 and the storage device 508. Each of components 502, 504, 508, 508, 510, and 512 is interconnected using various buses and may be mounted on a common motherboard or otherwise. Processor 502 can process instructions for execution within computing device 500, including instructions stored in memory 504 or on storage device 508, to display graphical information of a GUI on an external input / output device, such as a display 516 coupled to high-speed interface 508. In other embodiments, multiple processors and / or multiple buses may be used with multiple memories and various types of memory, depending on the situation. Additionally, multiple computing devices 500 may be connected, each providing some part of the necessary operation, for example, as a server library, a group of blade servers, or a multiprocessor system.

[0199] Memory 504 stores information within computing device 500. In one embodiment, memory 504 is one or more volatile memory cells. In another embodiment, memory 504 is one or more non-volatile memory cells. Memory 504 can also be another form of computer-readable medium, such as a magnetic disk or optical disk.

[0200] Storage device 508 provides massive storage for computing device 500. In one embodiment, storage device 508 may be or contain computer-readable media, such as floppy disk devices, hard disk devices, optical disk devices, magnetic tape devices, flash memory or other similar solid-state storage devices, or device arrays, including devices or other configurations in a storage area network. A computer program product may be tangibly embodied in an information carrier. The computer program product may also contain instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as memory 504, storage device 508, or memory on processor 502.

[0201] High-speed controller 508 manages bandwidth-intensive operations of computing device 500, while low-speed controller 512 manages less bandwidth-intensive operations. This functional allocation is merely illustrative. In one embodiment, high-speed controller 508 is coupled to memory 504, display 516, and high-speed expansion port 510, which can accept various expansion cards (not shown), for example, via a graphics processor or accelerator. In this embodiment, low-speed controller 512 is coupled to storage device 508 and low-speed expansion port 514. The low-speed expansion port (which may include various communication ports such as USB, Bluetooth, Ethernet, and wireless Ethernet) may be coupled to one or more input / output devices, such as keyboards, pointing devices, microphone / speaker pairs, scanners, or networking devices such as switches or routers, for example, via a network adapter. Computing device 500 can be implemented in many different forms, as shown. For example, the computing device may be implemented as a standard server 520, or multiple times in a group of such servers. It may also be implemented as part of a rack-mount server system 524. Furthermore, the computing device may be implemented in a personal computer such as laptop computer 522. Alternatively, components from computing device 500 may be combined with other components from mobile devices (not shown), such as device 550. Each of such devices may include one or more of computing devices 500, 550, and the entire system may consist of multiple computing devices 500, 550 communicating with each other.

[0202] The computing device 500 can be implemented in a variety of different forms, as shown in the figure. For example, the computing device can be implemented as a standard server 520, or multiple times in a group of such servers. It can also be implemented as part of a rack server system 524. Furthermore, the computing device can be implemented in a personal computer such as a laptop computer 522. Alternatively, components from the computing device 500 can be combined with other components from mobile devices (not shown), such as device 550. Each of these devices can contain one or more of the computing devices 500, 550, and the entire system can consist of multiple computing devices 500, 550 communicating with each other.

[0203] The computing device 550 includes a processor 552, a memory 564, and input / output devices such as a display 554, a communication interface 566, and a transceiver 568, as well as other components. The device 550 may also be provided with storage devices, such as microdrives or other devices, to provide additional storage. Each of components 550, 552, 564, 554, 566, and 568 is interconnected using various buses, and some of these components may be mounted on a common motherboard or otherwise.

[0204] Processor 552 can execute instructions within computing device 550, including instructions stored in memory 564. The processor can be implemented as a chipset comprising multiple independent analog and digital processors. Alternatively, the processor can be implemented using any of a variety of architectures. For example, processor 510 can be a CISC (Complex Instruction Set Computer) processor, a RISC (Reduced Instruction Set Computer) processor, or a MISC (Minimum Instruction Set Computer) processor. The processor can provide, for example, coordination with other components of device 550, such as control of the user interface, applications running by device 550, and wireless communications performed by device 550.

[0205] Processor 552 can communicate with the user via control interface 558 and display interface 556 coupled to display 554. Display 554 can be, for example, a TFT (Thin Film Transistor Liquid Crystal Display) or OLED (Organic Light Emitting Diode) display, or other suitable display technology. Display interface 556 can include suitable circuitry for driving display 554 to present graphics and other information to the user. Control interface 558 can receive commands from the user and translate these commands for submission to processor 552. Furthermore, an external interface 562 can be provided to communicate with processor 552 to enable near-field communication between device 550 and other devices. External interface 562 can provide wired communication in some embodiments, or wireless communication in others, and multiple interfaces can be used.

[0206] Memory 564 stores information within computing device 550. Memory 564 may be implemented as one or more computer-readable media, one or more volatile memory cells, or one or more non-volatile memory cells. Extended memory 574 may also be provided and connected to device 550 via an extended interface 572, which may include, for example, a SIMM (Single In-line Memory Module) card interface. Such extended memory 574 may provide additional storage space for device 550, or it may store applications or other information for device 550. Specifically, extended memory 574 may include instructions for performing or supplementing the above processes, and may also include security information. Thus, for example, extended memory 574 may be provided as a security module for device 550 and may be programmed with instructions that allow secure use of device 550. Furthermore, security applications may be provided via a SIMM card along with additional information, such as placing identification information on the SIMM card in an unbreakable manner.

[0207] The memory may include, for example, flash memory and / or NVRAM memory, as described below. In one embodiment, the computer program product is tangibly embodied in an information carrier. The computer program product contains instructions that, when executed, perform one or more methods, such as those described above. The information carrier is a computer-readable or machine-readable medium, such as memory 564, extended memory 574, or memory on processor 552 that can be received, for example, by transceiver 568 or external interface 562.

[0208] Device 550 can communicate wirelessly via communication interface 566, which may include a digital signal processing circuitry system if necessary. Communication interface 566 can provide communication under various modes or protocols, such as GSM voice calls, SMS, EMS or MMS messaging, CDMA, TDMA, PDC, WCDMA, CDMA2000, or GPRS. Such communication can occur, for example, via radio frequency transceiver 568. Furthermore, short-range communication can occur, such as using Bluetooth, Wi-Fi, or other such transceivers (not shown). Additionally, GPS (Global Positioning System) receiver module 570 can provide device 550 with additional navigation-related and location-related wireless data, which may be used by applications running on device 550 as appropriate.

[0209] Device 550 can also communicate audibly using audio codec 560, which can receive spoken information from a user and convert it into usable digital information. Audio codec 560 can also generate audible sound for the user, such as through a speaker (e.g., in the handheld terminal of device 550). This sound can include sounds from voice telephone calls, recorded sounds (e.g., voice messages, music files, etc.), and sounds generated by applications operating on device 550.

[0210] The computing device 550 can be implemented in a variety of different forms, as shown in the figure. For example, the computing device can be implemented as a mobile phone 580. The computing device can also be implemented as part of a smartphone 582, a personal digital assistant, or other similar mobile devices.

[0211] Various embodiments of the systems and methods described herein may be implemented in digital electronic circuits, integrated circuits, specially designed ASICs (Application-Specific Integrated Circuits), computer hardware, firmware, software, and / or combinations of such embodiments. These various embodiments may include implementations in one or more computer programs capable of executing and / or interpreting on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose processor coupled to receive data and instructions from a storage system, at least one input device, and at least one output device, and to send data and instructions to the storage system, at least one input device, and at least one output device.

[0212] These computer programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented in high-level procedural and / or object-oriented programming languages ​​and / or in assembly language / machine language. As used herein, the terms "machine-readable medium" and "computer-readable medium" refer to any computer program product, apparatus, and / or device used to provide machine instructions and / or data to a programmable processor, such as a disk, optical disk, memory, programmable logic device (PLD), including machine-readable media that receive machine instructions as machine-readable signals. The term "machine-readable signal" refers to any signal used to provide machine instructions and / or data to a programmable processor.

[0213] To provide interaction with the user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user, and a keyboard and pointing device (e.g., a mouse or trackball) that the user can use to provide input to the computer. Other types of devices may also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback, such as visual feedback, auditory feedback, or tactile feedback; and input from the user can be received in any form, including sound, speech, or tactile input.

[0214] The systems and technologies described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), middleware components (e.g., an application server), or front-end components (e.g., a client computer with a graphical user interface or a web browser). Users can interact with specific implementations of the systems and technologies described herein, or with any combination of such back-end, middleware, or front-end components, through this computing system. The components of the system can be interconnected via any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), and the Internet.

[0215] This computing system may include clients and servers. Clients and servers are typically located far apart and usually interact via a communication network. The client-server relationship is established by means of computer programs running on the respective computers and having a client-server relationship with each other.

[0216] Other implementation plans

[0217] Several embodiments have been described. However, it should be understood that various modifications can be made without departing from the spirit and scope of the invention. Furthermore, the logical flow shown in the figures does not require the specific order or sequence shown to achieve the desired result. Additionally, other steps may be provided in the flow, or steps may be eliminated, and other components may be added to or removed from the system. Therefore, other embodiments are also within the scope of the following claims.

[0218] Experimental results

[0219] Statistical and numerical examples of the compression method according to the present invention

[0220] The following comparative examples were performed on an uncompressed data file of 35,770 MB containing 48 million nucleotide reads or sequences. The results of this comparative example are also... Figure 6 It is depicted using graphics.

[0221] The following results indicate the size of the compressed version of the uncompressed file, which has 48 million read segments and a size of 35,770 MB, when compressed using each corresponding algorithm. These results are depicted in Figure 610.

[0222] Size of the file compressed using gzip: 6,649 MB (612)

[0223] o File size compressed using non-reference based Spring software: 1,402 MB (614)

[0224] The size of the file compressed using the reference-based compression method according to this disclosure is: 1,179 MB (6160 MB).

[0225] The following results indicate the amount of time taken to compare an uncompressed file of size 35,770 MB with 48 million read segments using each corresponding algorithm. These results are depicted in Figure 620.

[0226] Compression time using non-reference based SPRING software: 1,722 s (622).

[0227] Compression time using the reference-based compression method according to the present invention: 181s (624)

[0228] The following results indicate the average size, in bits / nucleotides, of the compressed version of the uncompressed file, which has 48 million reads and a size of 35,770 MB, when compressed using each corresponding algorithm. These results are depicted in Figure 630.

[0229] Average size of uncompressed data files in bits / nucleotides (ASCII encoding): 8 bits / nucleotide (630)

[0230] The average size of a file compressed using an encoding suitable for the four possible characters A, T, C, G, in bits / nucleotides: 2 bits / nucleotide (634)

[0231] The average size of the file compressed using the reference-based compression method according to the present invention, in bits / nucleotides: 0.33 bits / nucleotide (636).

[0232] Figure 7 Additional comparisons are presented on the compression of sample reads generated using WXS novaseq and the reference genome SRR8604734 using different compression algorithms.

[0233] Figure 710 shows a comparison of the resulting compressed sizes (MB) between gzip compression (712) of WXS novaseq reads using the reference genome SRR8604734, Spring compression (716) of WXS novaseq reads using the reference genome SRR8604734, and the compression method of this disclosure (716) of WXS novaseq reads using the reference genome SRR8604734. Measurements are in megabytes.

[0234] Figure 720 shows a comparison of the compression speed of the compression algorithm in 720 for compressing WXS novaseq reads using the reference genome SRR8604734. The Spring compression speed for compressing WXS novaseq reads using the reference genome SRR8604734 is shown in (722), and the speed of this disclosure for compressing WXSnovaseq reads using the reference genome SRR8604734 is shown in 724. Measurements are in seconds.

[0235] Figure 730 shows a comparison of memory usage during compression of WXS novaseq reads using reference genome SRR8604734 by the compression algorithm in 730. Spring compression used 13,428 MB of memory to compress WXS novaseq read 732 using reference genome SRR8604734, while this disclosure used 3,604 MB of memory to compress WXS novaseq read 734 using reference genome SRR8604734. Measurements are in megabytes.

[0236] Figure 8 Additional comparative results are presented, showing the compression of sample reads generated using different sequencers and different reference genomes at different compression ratios using different compression algorithms.

[0237] Figure 810 shows the original size of the read file generated using Novaseq in gigabytes (GB) (812). Using gzip to compress a 100GB Novaseq read, gzip compresses the original data to 17.7GB (814). This disclosure compresses the same original size file 812 of the 100GB Novaseq-generated read to a compressed file of 3.4GB (816). As shown in 810, the reference genome used for compression by both gzip and this disclosure is SRR6882909, and the compression ratio is 5.2x.

[0238] Figure 820 shows the original size of the read file generated using Hiseq X Ten in gigabytes (GB) (822). Using gzip to compress a 100GB Hiseq X Ten read, gzip compresses the original data to 24.9GB (824). This disclosure compresses the same original size file 822 of the 100GB Hiseq X Ten read to a compressed file of 8.1GB (826). As shown in 820, the reference genome used for compression by both gzip and this disclosure is SRR7725247, and the compression ratio is 3x.

[0239] Figure 830 shows the original size of the read file generated using Hiseq 2000 in gigabytes (GB) (832). Using gzip to compress a 100GB Hiseq 2000 read, gzip compresses the original data to 27.6GB (834). This disclosure compresses the same original size file 832 of the 100GB Hiseq2000-generated read to a compressed file of 11.3GB (836). As shown in 830, the reference genome used for compression by both gzip and this disclosure is ERR174324, and the compression ratio is 2.4x.

[0240] The numerical examples cited above illustrate that the present invention allows for rapid compression and decompression while providing a high compression ratio.

Claims

1. A method for compressing genomic sequence data, the method comprising: A storage device is accessed by one or more processors, the storage device storing the multiple read segment records in a sequential order that preserves the multiple read segment records generated by the mapping and alignment module, each of the multiple read segment records corresponding to a fully mapped read segment or a partially mapped read segment; For each specific read segment record among the multiple read segment records: The one or more processors obtain the specific read segment record generated based on the data output by the mapping and comparison module, wherein the specific read segment record includes data indicating whether the read segment corresponding to the specific read segment record is fully mapped or incompletely mapped; The one or more processors determine, based on the specific read segment record, whether the specific read segment record corresponds to a read segment that is fully mapped to the reference sequence or a read segment that is not fully mapped to the reference sequence; Based on the determination by the one or more processors that the specific read segment record corresponds to a read segment that is not fully mapped to the reference sequence, the one or more processors determine whether the number of mismatches of the incompletely mapped read segments does not exceed a predetermined mismatch threshold number; Based on the determination that the number of mismatches does not exceed the predetermined mismatch threshold number, the one or more processors encode each mismatch of the incompletely mapped read segment into a compressed record with a predetermined compressed record size, wherein the encoding includes an offset relative to the previous mismatch for a specific mismatch of the incompletely mapped read segment. as well as The compressed record is stored in the storage device by the one or more processors in a manner that preserves the sequence order of the plurality of read segment records generated by the mapping and comparison module.

2. The method according to claim 1, wherein each of the plurality of read segment records further comprises: The data indicating the absolute start position of the comparison read segment relative to the reference sequence. Data indicating the length of the read segment, Data indicating the number of mismatches identified in the read segment. Indicates whether the read segment includes data containing at least one undetermined base N. Data indicating the number of undetermined N bases in the read segment. Indicates whether the read segment is mapped or unmapped data. Data indicating the position of the read segment record within the read segment record sequence output by the mapping and comparison module, and Data indicating the relative positions of possible mismatches within the read segment.

3. The method of claim 1, wherein the predetermined compressed record size is one byte, and wherein encoding each mismatch of the incompletely mapped read segment into a compressed record having a size of one byte comprises: For each specific mismatch, The first two bits of the byte are encoded by one or more processors to include data representing alternative nucleotides or bases present in the read segment instead of corresponding reference nucleotides or bases in the reference sequence; as well as The remaining six bits of the byte are encoded by one or more processors to include data representing the position of the mismatch in the reference sequence, the position being calculated as an offset relative to the previous mismatch of the read segment.

4. The method according to claim 3, further comprising: One or more processors determine whether the offset is greater than the maximum coded value; Based on the determination that the offset is greater than the maximum coded value, one or more processors insert at least one false mismatch between the specific mismatch and the previous mismatch.

5. The method according to claim 1, wherein the method further comprises: Based on the determination that the number of mismatches does not meet the predetermined number of mismatches, one or more processors use a simplified information entropy encoding process to encode the position list of the reference sequence corresponding to the position of each of the mismatches into the reference sequence.

6. The method according to claim 1, wherein the method further comprises: Based on the determination that the read segment record corresponds to a read segment that is fully mapped to the reference sequence, the one or more processors use simplified information entropy coding to encode at least a portion of the read segment record.

7. The method of claim 1, wherein determining by the one or more processors whether the number of mismatches in the incompletely mapped read segments does not exceed a predetermined mismatch threshold number comprises: The one or more processors determine whether the number of mismatches in the incompletely mapped read segments is greater than a reference threshold.

8. A system for compressing genomic sequence data, the system comprising: One or more computers, and one or more storage devices storing instructions, which, when executed by the one or more computers, are operable to cause the one or more computers to perform operations including any of the preceding claims.

9. A computer-readable storage device having instructions stored thereon, the instructions, when executed by a data processing device, causing the data processing device to perform operations for compressing genomic sequence data, the operations comprising the method of any one of claims 1 to 7.

Citation Information

Patent Citations

  • Method and apparatus for compact representation of bioinformatics data

    WO2018068829A1

  • Lossless compression of DNA sequences

    US20150227686A1