A novel context-based framework for improved quality value compression in aligned sequencing data

By employing count-based and neural network prediction-based arithmetic coding with alignment information, the method addresses inefficiencies in genome sequencing data compression, achieving up to 6% improvement in nanopore and 2% in Illumina datasets by predicting quality values and reducing file size.

JP7794129B2Active Publication Date: 2026-01-06KONINKLIJKE PHILIPS NV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
JP2022547930
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Priority Date
2020-02-07
Filing Date
2021-01-27
Publication Date
2026-01-06
Estimated Expiration
2041-01-27

AI Technical Summary

Technical Problem

Existing methods for compressing quality values in genome sequencing data are inefficient, particularly due to their unpredictable nature and large alphabet size, which can account for up to 80% of the compressed file size, and fail to effectively utilize alignment information for improved compression.

Method used

A method and system that utilize count-based adaptive arithmetic coding and neural network prediction-based arithmetic coding to compress quality values in genome sequencing data, leveraging alignment information to generate new contexts such as matches, mismatches, and genomic coordinate correlations for improved compression.

Benefits of technology

The proposed method achieves significant improvements in compression efficiency, with up to 6% improvement in nanopore sequencing datasets and 2% in Illumina sequencing datasets, by effectively utilizing alignment information to predict quality values and reduce file size.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007794129000002
    Figure 0007794129000002
  • Figure 0007794129000003
    Figure 0007794129000003
  • Figure 0007794129000004
    Figure 0007794129000004
Patent Text Reader

Abstract

A method for compressing information includes accessing reads of genome sequencing data, aligning the reads to a reference, generating alignment data based on the alignment of the reads, obtaining a set of contexts based on the alignment data, and compressing quality values ​​corresponding to the alignment data based on the set of contexts, where the alignment data provides an indication of errors in the genome sequencing data, and each of the quality values ​​provides an indication of the probability of an error at one or more bases in the genome sequencing data.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] TECHNICAL FIELD

[0001] This disclosure relates generally to processing information, and more particularly, but not exclusively, to processing genome-related information. [Background technology]

[0002]

[0002] Genome sequencing typically generates large amounts of data in the form of reads (e.g., noisy substrings of the genome and corresponding quality values ​​that provide an indication of the certainty or reliability in the read sequence). However, existing methods for compressing the quality values ​​of genome sequencing data have drawbacks. Summary of the Invention

[0003]

[0003] A summary of various example embodiments is provided below. Some simplifications and omissions are made in the following summary, which highlights and introduces some aspects of various example embodiments but is not intended to limit the scope of the invention. Detailed descriptions of example embodiments sufficient to enable those skilled in the art to make and use the inventive concepts follow in later sections.

[0004] According to one or more embodiments, a method for compressing information includes: (a) accessing reads of genome sequencing data; (b) aligning the reads to a reference; (c) generating alignment data based on the alignment of the reads; (d) obtaining a set of contexts based on the alignment data; and (e) compressing quality values ​​corresponding to the alignment data based on the set of contexts, wherein the alignment data provides an indication of errors in the genome sequencing data, and each of the quality values ​​provides an indication of a probability of an error at one or more bases in the genome sequencing data. The set of contexts includes at least one context.

[0005] In (e), the aligned genome sequencing data is compressed based on count-based adaptive arithmetic coding. In (e), the aligned genome sequencing data is compressed based on neural network prediction-based arithmetic coding. The set of contexts includes matches between reads and reference bases. The set of contexts includes at least one of the presence of mismatches and the type of mismatch. The set of contexts includes multiple bases in the reference sequence surrounding one or more of the quality values. The set of contexts includes an average quality value across multiple bases at one or more genomic coordinates. The set of contexts includes errors at current and nearby bases measured using a pileup of reads mapping to the same genomic coordinate. Operation (d) includes selecting a set of contexts based on one or more criteria, wherein the one or more criteria include dataset type, dataset size, context size, predictive ability of the context, or amount of data to be compressed.

[0006] According to one or more embodiments, a system for compressing information includes a memory storing instructions and a processor, wherein the processor executes the instructions to (a) access reads of genome sequencing data, (b) align the reads to a reference, (c) generate alignment data based on the alignment of the reads, (d) obtain a set of contexts based on the alignment data, and (e) compress quality values ​​corresponding to the alignment data based on the set of contexts, wherein the alignment data provides an indication of errors in the genome sequencing data, and each of the quality values ​​provides an indication of a probability of an error at one or more bases in the genome sequencing data. The set of contexts includes at least one context.

[0007] The processor compresses the aligned genome sequencing data based on count-based adaptive arithmetic coding. The processor compresses the aligned genome sequencing data based on neural network prediction-based arithmetic coding in (e). The set of contexts includes matches between reads and reference bases. The set of contexts includes at least one of the presence of mismatches and the type of mismatch. The set of contexts includes multiple bases in the reference sequence surrounding one or more of the quality values. The set of contexts includes an average quality value across multiple bases at one or more genomic coordinates. The set of contexts includes errors at current and nearby bases measured using a pileup of reads mapping to the same genomic coordinate. Operation (d) includes selecting the set of contexts based on one or more criteria, the one or more criteria including dataset type, dataset size, context size, predictive ability of the context, or amount of data to be compressed.

[0008]

[0008] The accompanying drawings, in which like reference numerals refer to identical or functionally similar elements throughout the separate drawings, are incorporated into and form part of the specification, and together with the detailed description below, serve to illustrate example embodiments of the concepts found in the claims and to explain various principles and advantages of those embodiments.

[0009]

[0009] These and other more detailed and specific features are more fully disclosed in the following specification, in which reference is made to the accompanying drawings. [Brief explanation of the drawings]

[0010] [Figure 1]

[0010] An example of a sequence alignment map file is shown below. [Figure 2]

[0011] 1 shows an example of aligned genomic data with corresponding quality values. [Figure 3]

[0012] 1 shows an example of aligned genomic data with corresponding quality values. [Figure 4]

[0013] 1 illustrates one embodiment of a method for compressing genomic data. [Figure 5]

[0014] 1 illustrates one embodiment of a method for compressing genomic data. [Figure 6]

[0015] 1 illustrates one embodiment of an arithmetic coder for genomic data. [Figure 7]

[0016] 1 illustrates one embodiment of a system for compressing genomic data. DETAILED DESCRIPTION OF THE INVENTION

[0011]

[0017] It should be understood that the drawings are only schematic and not drawn to scale. It should also be understood that the same reference numerals will be used throughout the drawings to denote the same or similar parts.

[0012]

[0018] The description and drawings illustrate the principles of various example embodiments. Thus, it should be recognized that those skilled in the art can devise various configurations that embody the principles of the present invention and are within its scope, even though not explicitly described or shown herein. Furthermore, all examples detailed herein are intended expressly for educational purposes, primarily to aid the reader in understanding the inventive principles and concepts that contribute to the advancement of the art by the inventors, and should not be construed as being limited to the specifically detailed examples and conditions of such. Additionally, the term "or" as used herein means a non-exclusive or (i.e., and / or) unless otherwise specified (e.g., "or otherwise" or "or alternatively"). Furthermore, since some example embodiments can be combined with one or more other example embodiments to form new example embodiments, the various example embodiments described herein are not necessarily exclusive. Descriptors such as "first," "second," "third," etc., are not intended to limit the order of the elements described, but are used to distinguish one element from the next and are generally interchangeable. Values ​​such as maximum or minimum are predetermined and set to different values ​​based on the application.

[0013]

[0019] Two platforms for sequencing genomic data are (i) Illumina sequencing and (ii) Oxford Nanopore (ONT) sequencing. Illumina sequencing provides high-throughput, fixed-length, and short-read sequencing with very low error rates (<1%—mostly substitutions). ONT sequencing provides real-time, variable-length, and long-read sequencing with high error rates (10-15%—insertions, deletions, and substitutions).

[0014]

[0020] Raw sequencing data obtained from sequencers implementing one or both of the above platforms is aligned to a reference genome for further analysis, such as variant calling. Alignment is performed using standard tools that attempt to find the portion of the genome that is most similar to each sequenced read in terms of a similarity metric, such as Hamming distance or edit distance. Typical alignment tools include bwa for short-read Illumina sequencing data and minimap2 for Nanopore sequencing data. Both of these aligners use indexing strategies to enable quick searches for matches to sequenced reads in the genome.

[0015]

[0021] Aligned genomic data are represented using a file in sequence alignment map (SAM) format (or a compressed representation thereof). An example of a SAM file is shown in Figure 1. The file contains information about the sequence of nucleobases (A / C / G / T) of the reads, the position of the alignment, substitutions / insertions / deletions between the alignments, and associated quality values. The quality values ​​are expressed, for example, as ASCII characters, but are equated to integer values ​​(e.g., ranging from 0 to 40) that represent the probability of error on a logarithmic scale.

[0016]

[0022] Figure 2 shows an example of genomic data with corresponding quality values. In this example, the quality values ​​exhibit a dependency on genomic coordinates for the Klebsiella pneumoniae nanopore dataset. In Figure 2, rows represent reads and columns represent genomic locations. The shading of symbols representing nucleic acid bases (C, T, A, G) represents the quality value at the corresponding genomic location, with lighter shading representing higher quality. The quality values ​​are generated by the base calling process, which converts raw analog data from sequencing technologies into the most likely base sequence and associated confidence in the prediction (quality value).

[0017]

[0023] Figure 3 shows an example alignment of the Illumina MiSeq E. coli dataset, where rows represent reads and columns represent genomic locations. The shading of symbols representing nucleobases indicates different quality values ​​at the corresponding genomic locations. In this example, there is little correlation between genomic coordinates and quality values. Instead, correlations are mostly horizontal, e.g., within-read quality. Due to their unpredictable nature and large alphabet size, quality values ​​are difficult to compress and can account for up to 80% of the size of the compressed file after alignment.

[0018]

[0024] Techniques for compressing quality values ​​include lossy and lossless techniques. When all reads match at a genomic location, some types of lossy compression use alignment information to discard matching quality values. For Illumina sequencing, the error rate is relatively low, and therefore the quality values ​​have little impact on the analysis. For Oxford Nanopore sequencing, the error rate is relatively high, and therefore faithful preservation of quality values ​​is more important for downstream applications. It should be noted that nanopore sequencing offers several advantages over Illumina sequencing due to its long read length, which enables analysis of structural variations in the genome.

[0019]

[0025] Lossless compression of quality values ​​can involve, for example, arithmetic coding techniques or can be implemented using a general-purpose universal compressor such as gzip or bzip2. Arithmetic coding performs compression based on a (possibly adaptive) probabilistic model of the data. The better the model predicts the data, the better the compression. The model may incorporate various contexts that have a statistical correlation with the data being compressed. For arithmetic coding, a context of previous quality values ​​is used (e.g., order = number of previous quality values ​​used as context). The probabilistic model for each context is updated based on the data already seen by that context. The size of the context is selected according to the size of the data. Otherwise, there will be insufficient data per context, which will lead to poor probabilistic models and compression. In one or more embodiments, one or two orders of magnitude is sufficient for Illumina sequencing datasets. To take advantage of the fact that quality values ​​for Illumina sequencing become worse toward the end of the read, an effect that does not apply to nanopore sequencing, in addition to the previous quality value, the compressor also uses the position within the read as context.

[0020]

[0026] In view of the increasing size of genome sequencing data and the contribution of quality values ​​to the overall size, one or more embodiments described herein provide improved compression of quality values ​​of genome sequencing data. These embodiments include systems and methods for generating and / or selecting one or more new contexts for compressing quality values ​​of genome data subjected to an alignment process. Compression is performed, for example, using count-based adaptive arithmetic coding or neural network prediction-based arithmetic coding.

[0021]

[0027] Compressed quality values ​​based on alignment information provide improved or optimal results for a number of reasons. For example, alignment information provides an indication of errors present in the sequencing process, which directly corresponds to a quality value that measures the probability of an error at a given base. Thus, alignment provides "side information" for quality value compression, which leads to improved compression.

[0022]

[0028] For nanopore sequencing, there is a correlation between genomic coordinates and quality values. According to one or more embodiments, this correlation is used as a basis for improving lossy compression of quality values ​​based on alignment information. For example, alignment information is used to devise a new context for predicting quality values, which leads to improved prediction and compression using arithmetic coding.

[0023]

[0029] According to these or other embodiments, a neural network prediction-based arithmetic coding mode is provided that does not treat quality values ​​as separate symbols, but rather treats quality values ​​as related integers. For example, according to one embodiment, a neural network is provided that performs compression based on the similarity of quality values ​​that are adjacent or fall within a predetermined range of each other. Additionally or alternatively, compression is performed based on the similarity between base sequences (e.g., the nucleobase sequence ACGAT should be close to the nucleobase sequence AGGAT by the sequence GCCGA, where C is cytosine, T is thymine, A is adenine, and G is guanine). By using contexts in the right manner, in contrast to count-based adaptive arithmetic coding, which considers each context value independent, neural network prediction-based arithmetic coding performs more accurately based on multiple contexts, each of which has significantly less data compared to, for example, other types of context-adaptive arithmetic coding.

[0024]

[0030] FIG. 4 illustrates one embodiment of a method for compressing genomic data, including quality values, and FIG. 5 provides a conceptual diagram of the compression method. Referring to FIGS. 4 and 5, the method includes, at 410, accessing information from a file containing genomic information. This operation is preceded by, for example, operation 404, which includes accessing reads from genomic sequencing data, and operation 408, which includes aligning the reads to a reference. The file containing genomic data, for example, a SAM file generated from the aligned reads, includes a read identifier (id), an alignment position of the read, a CIGAR string representing the operation of converting the reference sequence to the read sequence, the read sequence, and a quality value. The CIGAR string indicates information about the alignment. For example, when aligning a read sequence (e.g., a sequence of nucleobases C, T, A, and G) to a reference, additional bases not in the reference are present and / or bases in the reference are missing. A CIGAR string is a sequence of base lengths and associated actions that are used to indicate which bases align to the reference (match or mismatch), are deleted from the reference, and / or are inserted if not in the reference. In other embodiments, the file is different from the SAM file but contains the same or similar information.

[0025]

[0031] At 420, a set of contexts 510 for the genome sequencing data is obtained based on information in the file (e.g., alignment data). In one embodiment, the contexts are obtained from a set of possible contexts already identified and stored in memory. In other embodiments, the set of contexts is generated by a processor, for example, based on one or more criteria.

[0026]

[0032] A context set includes one or more contexts, which refer to, for example, one or more conditions for correlating, organizing, collecting, or comparing aligned genome sensing data. These conditions are used as the basis for the model's quality value, which predicts the probability of the next quality value symbol, for example, used in arithmetic coding. Note that when obtaining a context set, the file data includes aligned sequencing data with corresponding quality values. Therefore, access to the genome coordinates to which each base is aligned is determined. The file data also includes information about read mapping to genome coordinates and the presence and type of errors in reads at the current and nearby genome coordinates.

[0027]

[0033] The set of contexts obtained in operation 420 are new contexts in that they have not been previously used to process genomic data (and in particular with respect to quality values ​​of genomic data), and the quality values ​​of this type of data have not been compressed based on any of the types of contexts in the set. The set of contexts can be obtained in various ways. For example, in the case of a SAM file, the file is parsed line by line, and for each quality value symbol, one or more contexts are generated based on the fields in the SAM file and the reference genome sequence. For each context, the number of possible values ​​(denoted as N) is determined, which is found to be relevant at least for count-based adaptive arithmetic coding.

[0028]

[0034] According to one or more embodiments, the set of new contexts includes one or more of the following: The first context is whether the lead base matches the reference base. Whether this condition is met is indicated by a binary value (N=2), where the binary value is set to 1 if the lead base exactly matches the reference base (as indicated by the CIGAR string), and 0 if there is no match.

[0029]

[0035] The second context corresponds to whether there is a mismatch and, if so, what the type of mismatch is. The type of mismatch includes one or more of an insertion, deletion, or substitution. This information is typically contained within the CIGAR string in the SAM file format, which has N=4.

[0030]

[0036] The third context is k bases in the reference sequence that surround the quality value. This is the alignment position and the reference sequence, N=4. k is obtained based on

[0031]

[0037] The fourth context is the average quality value across multiple bases at the current and nearby genomic coordinates. The multiple bases could be all or fewer of all bases at the current and nearby coordinates. The average quality value is obtained by collecting all reads with alignments that overlap with a specific base, and then calculating the average value from the respective quality values ​​for the specific base and nearby bases. In this case, N = (range of quality values). Since this context cannot be calculated directly without the quality values ​​themselves in certain situations, the average quality value is stored separately. For example, the average quality value at each genomic coordinate is stored separately after compressing them using a compressor, e.g., 7-zip. This is done so that the context is calculated in the decompressor and, in some cases, does not access the next quality value.

[0032]

[0038] The fifth context corresponds to errors at the current and nearby bases measured using a pileup of reads mapping to the same genome coordinate. To obtain this context, all reads with overlapping alignments to a specific base are collected. Then, their pileup information, representing the count of aligning bases at a given position, is taken and used as the basis for determining whether there is an error in the current read at a specific position or whether there is a discrepancy between the sequence data and the reference genome used for alignment. This context is similar to the first context, except for the additional consideration that discrepancies in the reads relative to the reference may be due to mutations rather than sequencing errors. Sixth and subsequent contexts also exist, for example, any field in the aligned data can be used as a context for quality value compression.

[0033]

[0039] In one embodiment, when using a set of contexts, e.g., c_1, c_2, ..., c_m, from the contexts in the list above, the set is considered to be a single context, denoted as a tuple c = (c_1, c_2, ..., c_m). If the number of possible values ​​for context c_i is N_i, then the number of possible values ​​for context c is N_1 x N_2 x ... x N_m. This context c is then used to perform encoding or to train a neural network model.

[0034]

[0040] In one embodiment, the context in step n must meet the forward quality q for the decompression to be successful. n In one case, we assume that the context takes on a finite set of values, although this is not necessary for some forms of coding, e.g., machine learning prediction-based arithmetic coding.

[0035]

[0041] The context obtained in operation 420 also includes other previously proposed contexts. A proposed context may be any of k, i.e., q n-1 , …, q n-k The quality values ​​for k in the past. This context is taken directly from the quality value string in the SAM file. In this case, N = (range of quality values). k where the quality values ​​range from 40 to 80 (the number of different quality values ​​possible for each base). Another proposed context is the position within the read. This corresponds to the position of the quality value symbol within the read, which it is a part of, with N = maximum read length. Another context is the k bases within the read that surround the quality value. This is obtained by selecting a k-length substring of the read sequence centered at the position within the read, where N = 4k.

[0036]

[0042] At 430, one or more contexts 520 in the set obtained in operation 420 are selected prior to compression (encoding). Context selection can be performed in a variety of ways. For example, contexts are selected based on one or more of the following criteria: type of dataset, size of dataset, size of context, predictive ability of the context, and amount of data to be compressed. In one embodiment, the selected context is determined based on the compression algorithm (e.g., encoding mode) used. For some forms of encoding (e.g., machine learning), context selection is performed based on a set of training data 525. The training data indicates that a first set of contexts should be selected for genomic data having a first set of features or characteristics, and that another set of contexts should be selected for genomic data having a second, different set of features or characteristics.

[0037]

[0043] At 440, an encoding mode is selected based on one or more predetermined criteria, including data size, predictive ability, processing efficiency, availability of training data, compatibility with other systems or applications, and / or one or more other criteria or trade-offs. In one embodiment, two possible entropy encoding modes are used: (1) Mode 1—Count-Based Adaptive Arithmetic Coding 540 and (2) Mode 2—Machine Learning Prediction-Based Arithmetic Coding 550. Each type of encoding has its own strengths and weaknesses. Once a mode is selected, a corresponding compression algorithm is applied to compress the quality values ​​530, which are input, for example, from a SAM file.

[0038]

[0044] Mode 1 is implemented very efficiently relative to Mode 2 and is useful for encoding large amounts of genomic data. However, Mode 1 has some limitations. For example, to achieve improved efficiency, the amount of available data must be significantly larger than the set of all possible contexts. Insufficient data leads to a count array that is sparsely filled and therefore fails to provide a true probability distribution. These considerations limit the number of contexts available for prediction, at least for some applications. Also, the encoding algorithm for Mode 1 cannot necessarily exploit similarities between context values ​​(e.g., numerically similar quality values ​​or similar sequences of bases). Because the count array counts occurrences of each context value separately, more data is needed to obtain a sufficient number of counts for each, even when similarities exist between them.

[0039]

[0045] Mode 2 addresses these limitations by providing a more powerful prediction framework that can efficiently utilize context and ignore irrelevant context. Mode 2 machine learning prediction-based arithmetic coding is performed based on a trained model 558 that is trained based on a set of training data 554. Mode 2 provides improved results over Mode 1 in some situations. Nevertheless, Mode 1 provides good, if not optimal, results, for example, when a large amount of genome sequencing data is available.

[0040]

[0046] Once the coding type is selected at 450, the quality values ​​of the alignment data to be compressed are input to the compressor to implement the selected coding mode. An example of the inputs and outputs of the selected coder 620 is shown in Figure 6. The input 610 contains the predicted probability of the quality value symbol (based on the context value) at each step, and the output 630 contains the compressed bitstream. In Figure 6, q n denotes the nth quality value symbol. The entire context is represented as a tuple containing one or more of the possible contexts selected in the previous operation. Once compression is performed, the selected coder 620 outputs the compressed file 560.

[0041]

[0047] When mode 1 is selected as the encoding type implemented by the compressor, a count-based probability calculation is performed, where the number of occurrences of each (context, quality) pair is stored to calculate the probability at each step. In one embodiment, encoding performed in mode 1 is performed as follows: First, initialize the array count[quality][context] to 1 for all (quality, context) pairs. Second, the size parameter is initialized to a zero value (in bits) to represent the compressed size during the compression procedure. Next, q in the list of quality values ​​is n For a context c, q nThen the probability for the context is calculated according to: Prob(:|c)=counts[:][c] / sum(counts[:][c]). Then the value q n is coded by arithmetic coding using Prob(:|c) as the probability distribution. Then, the size parameter is calculated as Size=Size+log2(1 / Prob(q n Finally, the count value is adjusted to be counts[q n ][c]+=1.

[0042]

[0048] When mode 2 is selected as the coding type implemented by the compressor, a predictive model is trained using the selected context as input. During the training procedure, the probability of each possible quality value is then output, and the loss function is the classification cross-entropy loss. The classification cross-entropy loss is a standard loss function used in classification tasks relevant here because it also represents the compression size when applying arithmetic coding using the predicted probabilities. The predictive model may be, for example, a machine learning model such as a decision tree, a neural network, one or more linear filters, or other types of models. For model input, the quality value is, for example, treated as a numerical variable (instead of a categorical variable), and other contexts are incorporated as categorical or numerical variables.

[0043]

[0049] When the compressor implements mode 2, it performs the following operations: First, the size parameter is initialized to zero (in bits) and represents the compressed size during the compression procedure. Second, the q n For, q n Context c is calculated for q. Then, the prediction model generates a probability Prob(:|c) by setting the input to c. The size is then adjusted as follows: Size=Size+log2(1 / Prob(q n |c)). In optional operation, adaptive training in predictive models is performed by (qn ,c). This adaptive training operation is used in both modes of the training procedure, giving a total of four possible modes of operation, as described in more detail below. In some cases, this operation increases computation time, but improves compression when training data is unavailable or when there is a mismatch between the training data and the data to be compressed.

[0044]

[0050] The model used in mode 2 encoding can be trained in various ways. For example, model training can be performed on the data to be compressed. In this case, the trained model parameters are included as part of the compressed file (e.g., after compression using a tool like 7-zip). Since this information is included in the file, a decompressor performs the same operations as a compressor to decompress the compressed data. Another training technique involves training the model on a dataset different from the data to be compressed. In this case, the model is shared between the encoder and decoder. Here, if the model parameters are already known to the decoder, there is no need to include them in the file. The first training procedure described above is useful, for example, when a similar dataset for training is not available. The second training procedure can be more efficient in terms of compression time and compressed size in some cases.

[0045]

[0051] Decompression is performed in a manner symmetrical to the compression performed to compress the aligned data, e.g., the arithmetic coder is replaced by an arithmetic decoder that performs the inverse operations of the mode 1 or mode 2 compression performed.

[0046]

[0052] 7 illustrates one embodiment of a system for compressing genome sequencing data, which may perform, for example, method embodiments described herein. The system includes a processor 710, a memory 720, and a database 730. The processor includes a controller 740, an aligner 750, a context selector 760, a mode 1 compressor 770, and a mode 2 compressor 780. The processor performs or controls the operation of method embodiments by, for example, executing instructions stored in memory 720, which may be a non-transitory computer-readable medium. Examples of memory 720 include read-only memory or random-access memory, including various types of these memories.

[0047]

[0053] The aligner 740 aligns reads of the genome sequencing data stored in the database 730 to a predetermined reference. The processor generates alignment data based on the alignment of the reads. The alignment data is stored in the database along with the genome sequencing data. The context selector 750 selects a set of contexts based on the alignment data, and the controller selects one of the mode 1 and mode 2 compressors based on criteria or other conditions described herein. The processor then receives results from the selected compressor and outputs a compressed file to a workstation or other terminal, or both, for storage in the database 730. The compressor is replaced by or configured to operate as a corresponding decompressor to decompress the output file at the next time. [Example]

[0048]

[0054] During testing, the above-described embodiment was applied to compress one Illumina dataset and two Nanopore datasets of genomic data. Table 1 shows the datasets used for evaluation. For simplicity, only aligned reads mapping to the forward strand were used in the experiments.

[0049] [Table 1]

[0050]

[0055] For the Illumina dataset, the results show a very slight improvement (0.6%) by incorporating additional context that signifies the presence of discrepancies, while other contexts are less effective. This is to be expected, as Illumina's error rate is very low, rendering most quality values ​​useless. Furthermore, there is little dependency between the read / reference sequence and the quality values ​​for Illumina sequencing.

[0051]

[0056] For the nanopore dataset, the results show even greater improvement. For the smaller lambda phage dataset, using the mismatch type and the average quality value at the genome coordinate as additional context (along with one previous quality value) improved compression by 2.4% in mode 1. Due to the small dataset size, using more context in mode 1 leads to worse results. On the other hand, compression in mode 2 allows the use of more context and provided an additional 4% improvement. The context set used in this case was two previous quality values, the average quality value at the genome coordinate, and five nearby bases in the read. Compression was performed using a neural network model with a three-hidden layer fully connected network with a width of 20. The model was trained (i.e., training performed on the data to be compressed) using the first training procedure described above for 20 epochs (ReLU nonlinearity, batch normalization, softmax activation used). Further improvement is possible with more powerful models such as RNNs and associated training procedures.

[0052]

[0057] In mode 1 compression experiments on a larger Klebsiella pneumoniae dataset, results showed an improvement of nearly 6% when using alignment information (with context of nearby bases given in the reference and mismatch type) and prior quality, versus using prior quality alone. Further improvement in mode 1 was obtained using context of nearby read bases and prior quality values, giving an additional 2% improvement.

[0053]

[0058] Additional experiments were performed on a subset of this dataset using mode 2 compression, a compression not optimized for speed, in the context of the dataset's size. Even when using a small fully connected neural network, the results show at least a 2% improvement over mode 1 compression.

[0054]

[0059] Therefore, using the additional context resulting from the alignment provides an improvement of about 5% or more for quality value compression. The choice of the set of contexts is important for mode 1 compression, but less important for mode 2 compression - here the results are even better, but at the expense of increased computation time. Further optimizations are possible for speed, for the choice of context, and for the neural network architecture and training procedure.

[0055]

[0060] Other embodiments include a computer-readable medium having stored thereon instructions for causing a processor to perform the operations of the embodiments described herein. Additional instructions are stored within the computer-readable medium to perform other operations of the system and method embodiments.

[0056]

[0061] The processors, controllers, compressors, decompressors, coders, decoders, selectors, aligners, and other information that generate, process, and calculate the features of the embodiments disclosed herein may be implemented in logic that includes, for example, hardware, software, or both. When implemented at least partially in hardware, the processors, controllers, compressors, decompressors, coders, decoders, selectors, aligners, and other information that generate, process, and calculate the features may be, for example, any one of a variety of integrated circuits, including, but not limited to, application specific integrated circuits, field programmable gate arrays, combinations of logic gates, systems on a chip, microprocessors, or other types of processing or control circuitry.

[0057]

[0062] When implemented at least in part in software, the processor, controller, compressor, decompressor, coder, decoder, selector, aligner, and other information that generates, processes, and calculates features may include, for example, memory or other storage for storing code or instructions executed by, for example, a computer, processor, microprocessor, controller, or other signal processing device. As algorithms that form the basis of the methods (or the operation of a computer, processor, microprocessor, controller, or other signal processing device) are detailed, the code or instructions for implementing the operations of method embodiments transform the computer, processor, controller, or other signal processing device into a special-purpose processor for performing the methods herein.

[0058]

[0063] While various exemplary embodiments have been described in detail with particular reference to certain exemplary aspects thereof, it should be understood that the invention is capable of other exemplary embodiments and its details are susceptible to modifications in various obvious respects. As will be apparent to those skilled in the art, variations and modifications may be effected without departing from the spirit and scope of the invention. Embodiments may be combined to form additional embodiments. Accordingly, the foregoing disclosure, description, and drawings are for illustrative purposes only and do not limit the invention in any manner.

Claims

1. A computer program that, when executed by a computer, causes the computer to perform a method for compressing genomic information, the method comprising: (a) accessing genome sequencing data reads; (b) aligning the reads to a reference; (c) generating alignment data after aligning the reads, the alignment data including alignment positions on the reference and characteristics of matches and mismatches of bases in the reads relative to bases in the reference at alignment positions; (d) arithmetically compressing quality values ​​of the reads, each of the quality values ​​providing an indication of the probability of error of a base in the genome sequencing data; and wherein the compressing step comprises: (e) selecting an arithmetic compression context; The arithmetic compression context is based on the alignment data, and the arithmetic compression context based on the alignment data is - whether the base in the read at the position of the quality value matches the base in the reference; the type of discrepancy at said position of said quality value, being an insertion, deletion or substitution, Selected from: Computer program.

2. The computer program of claim 1 , wherein the aligned genome sequencing data is compressed based on neural network prediction-based arithmetic coding based on multiple contexts.

3. 2. The computer program of claim 1, wherein in (e) the aligned genome sequencing data is compressed using a mode of arithmetic coding selected based on one or more criteria, the one or more criteria including data size, predictive power, processing efficiency, availability of training data, or compatibility with other systems or applications.

4. The computer program of claim 1, wherein the step of selecting an arithmetic compression context includes selecting a set consisting of a plurality of arithmetic compression contexts.

5. A computer program as described in claim 1 or 4, wherein step (e) includes selecting the context based on one or more criteria, the one or more criteria including dataset type, dataset size, context size or amount of data to be compressed.

6. 1. A system for compressing information, the system comprising: a memory for storing instructions; a processor, the processor comprising: (a) accessing genome sequencing data reads; (b) aligning the reads to a reference; (c) generating alignment data based on the alignment of the reads, the alignment data including alignment positions on the reference and match and mismatch characteristics of bases in the reads relative to bases in the reference at alignment positions; (d) obtaining a sequence of quality values ​​for the reads, each of the quality values ​​providing an indication of the probability of an error at a base in the genome sequencing data; (e) selecting an arithmetic compression context, the arithmetic compression context being based on the alignment data, the arithmetic compression context being based on the alignment data, whether the base in the read at the position of the quality value matches the base in the reference; the type of mismatch at said position of said quality value being an insertion, deletion or substitution; a step selected from: (f) arithmetically compressing the quality values ​​using the arithmetic compression context; and executing the instructions to perform the above steps.

7. 7. The system of claim 6, wherein the processor compresses the aligned genome sequencing data based on neural network prediction based arithmetic coding based on multiple contexts.

8. the processor compresses the aligned genome sequencing data based on arithmetic coding, with the arithmetic coding mode and training procedure selected based on one or more criteria, the one or more criteria including data size, predictive ability, processing efficiency, availability of training data, or compatibility with other systems or applications; The system of claim 6.

9. The system of claim 6, wherein the step of selecting an arithmetic compression context includes selecting a set consisting of a plurality of arithmetic compression contexts.

10. The system of claim 6, wherein step (e) includes selecting the context based on one or more criteria, the one or more criteria including dataset type, dataset size, context size or amount of data to be compressed.

Citation Information

Patent Citations

  • Compression Of Genomic Data File

    US20130166518A1

  • Quality score compression for improving downstream genotyping accuracy

    US20170147597A1

  • Deep learning-based variant classifier

    WO2019140402A1