Nanopore Sequencing Raw Signal Data Compression Method, Device, Equipment and Medium

By extracting the base sequence data in the nanopore sequencing file and using the SSDC compressor for lossless compression, the problem of poor compression in the prior art is solved, and more efficient compression of the original signal data of nanopore sequencing is achieved.

CN115798605BActive Publication Date: 2025-07-29SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211415179.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-11
Publication Date
2025-07-29
Estimated Expiration
2042-11-11

AI Technical Summary

Technical Problem

Existing nanopore sequencing raw signal data compressors such as Picopore and VBZ fail to effectively utilize the correlation between different data sets in Fast5 files, resulting in poor compression.

Method used

The base sequence data in the nanopore sequencing file is extracted, and the sequencing original signal data is lost-free compressed by combining the base sequence data. By generating the expected signal and performing end-to-end mapping processing to improve compression performance.

Benefits of technology

The compression performance has been significantly improved, with the compression effect being increased by about 33%, and 5% higher than that of Picopore.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115798605B_ABST
    Figure CN115798605B_ABST
Patent Text Reader

Abstract

The present invention provides a method, apparatus, device and medium for compressing raw signal data of nanopore sequencing. By obtaining a nanopore sequencing file, extracting the dataset part in the nanopore sequencing file, extracting the base sequence data in the first dataset, and calling a preset SSDC compressor to compress the second dataset in combination with the extracted base sequence data, compressed data is obtained. Compared with the traditional VBZ compression method, when implementing this technical solution, the base sequence data is extracted in advance. Since the base sequence data has a high correlation with the raw signal data of the sequencing in the second dataset, considering most of the characteristics of the raw signal data, when using the lossless compression SSDC compressor to compress the second dataset in combination with the base sequence data, the compression performance can be greatly improved and the compression quality can be guaranteed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of gene sequencing, and particularly relates to a method, device, equipment and medium for compressing raw signal data of nanopore sequencing. Background Art

[0002] In the past few decades, the continuous progress of sequencing technology has continuously reduced the cost of genome sequencing, and improved the sequencing throughput and accuracy. This provides important technical support for people to deconstruct complex genomic structures and answer the principles of phenotypic changes and disease occurrence caused by genomic variations.

[0003] Currently, existing compressors related to raw signal data of nanopore sequencing include Picopore and VBZ. The disadvantage of the compression method of the Picopore compressor is that its compression of the raw signal data in the downloaded Fast5 file only sets the Gzip compression level built into the Fast5 file dataset from the default lowest level to the highest level, without performing additional operations on the data itself. The disadvantage of the compression method of the VBZ compressor is that VBZ only realizes compression by using some characteristics of the raw signal data, without considering using the correlation between different datasets of Fast5 to improve the compression performance, especially the correlation between the raw signal data and the corresponding base sequence data. That is, the existing compression methods have the problem of poor compression effect.

[0004] Therefore, the prior art needs to be improved. Summary of the Invention

[0005] The main purpose of the present invention is to propose a method, system, device, equipment and medium for compressing raw signal data of nanopore sequencing, so as to at least solve the problem of poor compression effect of the compression method in the related art.

[0006] In the first aspect of the present invention, a method for compressing raw signal data of nanopore sequencing is provided, including:

[0007] Obtain a nanopore sequencing file, and extract the dataset part in the nanopore sequencing file; wherein, the dataset part includes a first dataset for storing sequencing base data and a second dataset for storing raw signal data of sequencing;

[0008] Extract the base sequence data in the first dataset;

[0009] Call a preset SSDC compressor to compress the second dataset in combination with the extracted base sequence data to obtain compressed data.

[0010] In the second aspect of the present invention, a custom application reinforcement security device is provided, including:

[0011] An acquisition module, configured to acquire a nanopore sequencing file and extract a dataset part from the nanopore sequencing file; wherein, the dataset part includes a first dataset for storing sequencing base data and a second dataset for storing sequencing raw signal data;

[0012] An extraction module, configured to extract base sequence data from the first dataset;

[0013] An invocation module, configured to invoke a preset SSDC compressor to perform compression processing on the base sequence data and the second dataset to obtain compressed data.

[0014] In a third aspect of the present invention, an electronic device is provided, including a memory, a processor, and a bus;

[0015] The bus is used to implement connection communication between the memory and the processor;

[0016] The processor is configured to execute a computer program stored on the memory;

[0017] When the processor executes the computer program, the steps in the nanopore sequencing raw signal data compression method provided in the first aspect are implemented.

[0018] In a fourth aspect of the present invention, a computer-readable storage medium is provided, on which a computer program is stored. The computer program, when executed by a processor, implements the steps in the nanopore sequencing raw signal data compression method provided in the first aspect.

[0019] The present invention provides a nanopore sequencing raw signal data compression method, device, equipment, and medium. By acquiring a nanopore sequencing file, extracting a dataset part from the nanopore sequencing file, extracting base sequence data from the first dataset, and invoking a preset SSDC compressor to perform compression processing on the second dataset in combination with the extracted base sequence data to obtain compressed data. That is, when implementing this technical solution, the base sequence data is extracted in advance. Since the base sequence data has a high correlation with the sequencing raw signal data in the second dataset, during the process of using the lossless compression SSDC compressor to compress the second dataset, the compression performance can be greatly improved and the compression quality can be ensured. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings without creative efforts.

[0021] Figure 1 It is a schematic diagram of the sequencing process of bases in the related art;

[0022] Figure 2 It is a schematic diagram of the sequencing result data in the related art;

[0023] Figure 3 It is a schematic diagram of the original signal data in the related art;

[0024] Figure 4 It is a schematic flowchart of the nanopore sequencing original signal data compression method provided in the embodiments of the present application;

[0025] Figure 5 It is a schematic diagram of calculating the target path with the smallest sum of elements through the distance matrix in the embodiments of the present application;

[0026] Figure 6 It is a schematic flowchart of compressing the base sequence data in the first dataset and the second dataset into the final compressed data in the embodiments of the present application;

[0027] Figure 7 It is a schematic diagram of the program module of the attack prevention and verification device provided in the third embodiment of the present application;

[0028] Figure 8 It is a schematic diagram of the structure of the electronic device provided in the fourth embodiment of the present application.

[0029] The realization, functional features and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. Detailed Embodiments

[0030] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0031] It should be noted that related terms such as "first", "second", etc. can be used to describe various components, but these terms do not limit the components. These terms are only used to distinguish one component from another. For example, without departing from the scope of the present invention, the first component can be called the second component, and similarly, the second component can also be called the first component. The term "and / or" means any one or more combinations of related items and described items.

[0032] Among the related technologies, nanopore sequencing, as a representative of the third-generation sequencing technology that can measure sequences with longer read lengths (up to 10,000 base pairs) and does not require a complex library construction process, has received widespread attention from the academic community in recent years. Its core is composed of a resistor membrane with a nanopore, in which a molecular linker is covalently bound. After the nanopore protein is fixed on the resistor membrane, the nucleic acid is pulled through the nanopore by a motor protein. When the nucleic acid passes through the nanopore, the charge changes, causing a change in the current on the resistor membrane. Because the diameter of the nanopore is very small, only a single nucleic acid polymer is allowed to pass through, and the charge properties of individual bases such as ATCG are different, different bases have different interferences on the current when passing through the protein nanopore. These current signals are sampled in real time to generate the raw signal data of the nanopore sequencer, and then this specific electrical signal sequence is translated into a base sequence using algorithm software such as Bonito and Chiron. This process is also called BaseCalling, that is, base recognition, to achieve sequencing, such as Figure 1 shown.

[0033] Since the research on BaseCalling technology is still under improvement and there is room for further improvement in accuracy, in addition to saving the Fastq data obtained by sequencing, Figure 2 As shown in Figure 2, it is also necessary to retain the original signal data of nanopore sequencing, such as Figure 3 The data set shown in the figure corresponds to the raw Fast5 file generated by nanopore sequencing, allowing for repeated analysis of sequencing data. However, the cost of storage hardware and data transmission has far outstripped the growth of genetic sequencing data, making data storage and transmission a bottleneck in the genetic sequencing industry. Highly efficient compression algorithms are an effective way to address this bottleneck. Furthermore, due to the continuity and high sampling rate of the raw current signal, the corresponding raw signal data requires an order of magnitude more space than the base sequence. Therefore, a more efficient compression algorithm is needed for raw signal data.

[0034] Currently available compressors for nanopore sequencing raw signal data include Picopore and VBZ. A drawback of the Picopore compressor is that it compresses the raw signal data in the downloaded Fast5 file by simply increasing the default Gzip compression level of the Fast5 file dataset from the lowest to the highest level, without performing any additional operations on the data itself. A drawback of the VBZ compressor is that it only achieves compression by exploiting some characteristics of the raw signal data and fails to consider improving compression performance by leveraging correlations between different Fast5 datasets, particularly between the raw signal data and the corresponding base sequence data. Consequently, existing compression methods suffer from poor compression performance.

[0035] Thus, to solve the technical problem of poor compression effect in the related art, please refer to Figure 4 , this application proposes a method for compressing raw signal data of nanopore sequencing, which specifically includes:

[0036] Step S401: Obtain a nanopore sequencing file and extract the dataset part in the nanopore sequencing file;

[0037] Specifically, the nanopore sequencing file FAST5 represents a file that can measure longer read sequences (up to 10,000 base pairs at most). By performing extraction operations on it according to a preset extraction rule, the dataset part Datasets in the nanopore sequencing file can be obtained. Generally, the dataset part Datasets includes a first dataset Fastq and a second dataset RawSignal. The first dataset is used to store sequencing base data, and the second dataset is used to store raw sequencing signal data.

[0038] Step S402: Extract the base sequence data in the first dataset;

[0039] Specifically, perform extraction processing on the first dataset to obtain the base sequence data BaseSequence in the first dataset; in this way, the base sequence data BaseSequence with high correlation with the raw sequencing signal data can be obtained.

[0040] More specifically, a high degree of correlation between the raw sequencing signal data and the base sequence data Base Sequence indicates that the base sequence data is derived from the raw sequencing signal data. This is mainly because the base sequence data is obtained from the raw sequencing signal data through base recognition technology (BaseCalling). Then, reverse thinking is used to generate a simulated raw sequencing signal data using the base sequence data, and then the difference is calculated with the original real data, and the data volume is reduced by storing the difference. For example: Let the base sequence data be x and the raw signal data be y, and Ay = x, where A represents the nanopore base recognition technology. Here, Bx = y', where B represents the official statistical data k-mer pore model provided by nanopore sequencing, which can approximately replace the inverse process of A, and y' represents the first expected signal.

[0041] Step S403: Call a preset SSDC compressor to compress the second dataset in combination with the extracted base sequence data to obtain compressed data.

[0042] Specifically, after obtaining the base sequence data and the second data set, since there is a high correlation (a high degree of correlation) between the base sequence data and the raw sequencing signal data in the second data set, better compression quality and effect can be achieved by compressing based on multiple highly correlated data. Then, a Simulation Signal Difference Compressor (SSDC for short) is called to compress the second data set in combination with the base sequence data. Since the SSDC compressor is a lossless compression tool for nanopore sequencing raw signal data, it can utilize the correlation between the raw signal data and the corresponding base sequence data to improve the compression performance. That is, when implementing this technical solution, the base sequence data is extracted in advance. Since the base sequence data has a high correlation with the raw sequencing signal data in the second data set and most of the characteristics of the raw signal data are considered, the lossless compression SSDC compressor is used to compress the second data set, thereby greatly improving the compression performance and ensuring the compression quality.

[0043] In some implementation manners of this embodiment, the step of calling a preset SSDC compressor to compress the second data set in combination with the extracted base sequence data to obtain compressed data specifically includes:

[0044] Step S201, call the pore model in the preset SSDC compressor to process the base sequence data to generate a first expected signal;

[0045] Specifically, the base sequence data can be processed by the k-mer pore model provided by the official nanopore sequencing to generate a first expected signal X e . Among them, the k-mer pore model in nanopore sequencing is used to model the expected current signal of a given k-mer nucleotide in a single-stranded DNA passing through a nanopore. For nanopore sequencing with R9.4 pore chemistry, k is equal to 6. This problem can be described as follows: Given a base sequence X = x1, x2,..., x L , where x i is a 4-state nucleotide base and can take a value from {A, T, C, G} of DNA. Then the corresponding expected current signal Y = y1, y2,..., y L-5 is generated, where y i is the expected current signal of the 6-mer starting from position i in X (e.g., 'ACCCGT'). In this way, the length of Y can be extended to the same length as X by filling the end of Y with the last value of Y (i.e., y L ) or the average expected signal value of all 6-mers.

[0046] In some embodiments of this embodiment, when generating the first expected signal, a simulator tool for nanopore sequencing can also be used, such as ReadSim, SiLiCO, NanoSim; they generate simulated data using the input base sequence and a set of parameters, where the parameters refer to, for example, insertion and deletion rates, substitution rates, read lengths, error rates, and quality scores. For example, ReadSim uses a fixed configuration file; Silico uses a user-provided configuration file; NanoSim uses user-provided empirical data to learn the configuration file that will be used in the simulation phase.

[0047] Step S202: Perform a mapping process on the first expected signal to obtain a second expected signal;

[0048] Specifically, since the frequency of current measurement is 8-10 times higher than the passing speed of the DNA sequence, the original signal sequence and the first expected signal sequence are of unequal length. Therefore, a second expected signal Xs of the same length as the original signal is obtained through mapping, and the mapping method can be an end-to-end mapping method. In this way, it is ensured that the length of the obtained second expected signal is equal to the length of the original sequencing signal data.

[0049] Step S203: Perform a compression process on the second expected signal and the original sequencing signal data to obtain compressed data.

[0050] Specifically, when the second expected signal and the original sequencing signal data are obtained, since the length of the second expected signal is equal to the length of the original sequencing signal data, based on the same length of the two signals, better compressed data with higher compression quality can be obtained when compressing these two signals.

[0051] In some embodiments of this embodiment, the step of performing a mapping process on the first expected signal to obtain a second expected signal specifically includes: determining the head data and tail data in the original sequencing signal data, and performing an end-to-end mapping process on the first expected signal according to the head data and tail data to obtain a second expected signal.

[0052] Specifically, the dynamic time warping (DTW) algorithm can be used to find the best end-to-end mapping from the first expected signal to the original sequencing signal data; for example, first determine the head data and tail data in the original sequencing signal data, and perform an end-to-end mapping process on the first expected signal according to the head data and tail data to obtain a second expected signal of the same length as the original signal.

[0053] In some embodiments of this embodiment, the step of performing a compression process on the second expected signal and the original sequencing signal data to obtain compressed data specifically includes:

[0054] Obtain a first time series according to the second expected signal, obtain a second time series according to the original sequencing signal data, calculate the first time series and the second time series to obtain a warping path and a warping path distance, and obtain compressed data according to the warping path and the warping path distance.

[0055] Specifically, since the obtained second expected signal represents multiple time periods and the signal wavelengths corresponding to the multiple time periods, the second expected signal can be defined in a sequence, and the corresponding first time series A can be obtained in chronological order; similarly, the corresponding second time series B can be obtained from the original sequencing signal data. Among them, they respectively correspond to the second expected signal and the expected signal in the original problem, and the length distributions are |A| and |B|. Then calculate the first time series A and the second time series B to obtain a warping path W and a warping path distance D, and obtain compressed data according to the warping path and the warping path distance.

[0056] In some implementation manners of this embodiment, the step of calculating the first time series and the second time series to obtain a warping path and a warping path distance specifically includes: establishing a distance matrix through the first time series and the second time series, calculating the target path with the smallest sum of elements in the distance matrix, using the target path as the warping path, and using the sum of the elements corresponding to the target path as the warping path distance.

[0057] Specifically, based on the preset form of the warping path is W = w1, w2,..., w k , corresponding to the warping path distance D = d1, d2,..., d k , where Max(|A|, |B|) <= K <= |A| + |B|. The form of w k is (i, j), where i represents the i coordinate in the second time series B, and j represents the j coordinate in the first time series A. The path length from the upper left corner to the lower right corner of the matrix has the following properties: (1) The current path length = the path length of the previous step + the size of the current element; (2) For a certain element (i, j) on the path, its previous element can only be one of the three, namely the adjacent element on the left (i, j - 1), the adjacent element above (i - 1, j), and the adjacent element in the upper left (i - 1, j - 1). Then based on the above properties, and since the distance matrix is composed of multiple elements, the initial position of the target path is the upper left corner of the distance matrix, and the end position of the target path is the lower right corner of the distance matrix, the warping path that can finally be obtained through the distance matrix is the warping path with the shortest distance, and this process can be solved using dynamic programming, such as Figure 5For the sample shown, for time series A = {1, 6, 3, 6, 4, 5} and B = {2, 4, 2, 5}, the final regularized path W = {(1,1), (2,2), (2,3), (3,3), (4,4), (4,5), (4,6)} and the regularized path distance D = {1, 2, 1, 1, 1, 1, 0} are obtained.

[0058] In some embodiments of this embodiment, the step of obtaining the compressed data according to the regularized path and the regularized path distance specifically includes: determining target coordinates in the regularized path whose occurrence times are greater than a preset occurrence threshold, deleting the target coordinates in the regularized path, and deleting the distance values corresponding to the target coordinates in the regularized path distance, to obtain a target regularized path and a target regularized path distance.

[0059] Specifically, because the ultimate goal is to be able to restore the original signal sequence through the expected signal sequence and the mapping relationship, so without affecting the ultimate goal, the target coordinates can be deleted in the regularized path, and the distance values corresponding to the target coordinates can be deleted in the regularized path distance, that is, first remove the w and the corresponding d values of the repeated j coordinates existing in the regularized path, that is, by determining the target coordinates whose occurrence times are greater than the preset threshold, and deleting the target coordinate w and the corresponding d value, and then only save the i coordinates in the regularized path after deduplication, so as to reduce the storage space. Finally, after deduplication, the target regularized path W1 = {1, 2, 2, 4, 4, 4} and the target regularized path distance D1 = {1, 2, 1, 1, 1, 0} are the data we need to store. Among them, the target regularized path and the target regularized path distance form the compressed data.

[0060] Among them, after obtaining the target regularized path and the target regularized path distance, further compression can be performed: for the target regularized path, it can be further compressed by a preset differential encoding method combined with a range encoding method; for the target regularized path distance, first use a preset linear predictive coding (LPC) method to utilize the context relationship to obtain a smaller linear prediction error, and then further compress it in combination with range encoding, so as to finally obtain the final compressed data CompressResults (see Figure 6 ).

[0061] It should be understood that the dynamic time warping algorithm will have a time complexity, which can be changed from the original O(N1*N2) of DTW to O(N*log N) where N = min(N1, N2), and N1 and N2 respectively correspond to the number of elements in two sequences. This dynamic time warping algorithm includes three key parts: continuous wavelet transform representation, context-dependent constrained DTW, and multi-level refinement. Specifically, as follows:

[0062] (1) Continuous wavelet transform representation: First, run the Continuous Wavelet Transform (CWT for short) on each input signal sequence to obtain a feature representation of the information. Subsequently, select the peaks and valleys to generate a low-resolution signal with a reduced length.

[0063] (2) Context-dependent constrained DTW: Use the warping path calculated at a lower resolution to determine the search boundary of the warping path calculated at a higher resolution.

[0064] (3) Multi-level refinement: Combine the low-resolution and high-resolution information from the CWT at different scales, and gradually refine the warping path as the level becomes finer until the final path reaches the original resolution of the input sequence.

[0065] Compared with the VBZ compression in the prior art that compresses the nanopore raw signal data by using variable-byte integer coding, the SSDC compression of the present invention utilizes the correlation between the base sequence and the corresponding raw signal sequence, establishes an end-to-end mapping by generating an expected signal and combining with the DTW algorithm, and transforms the problem of storing the raw signal into the way of storing the mapping result, so that the compression effect is improved by about 5%. Compared with the Picopore in the prior art that compresses the nanopore raw signal data by adjusting the Gzip compression level, the compression effect of the present invention is improved by about 33%.

[0066] Figure 7 Fig. shows a data compression device provided in the second embodiment of the present application, including:

[0067] An acquisition module 701, configured to acquire a nanopore sequencing file and extract the dataset part in the nanopore sequencing file; wherein, the dataset part includes a first dataset for storing sequencing base data and a second dataset for storing sequencing raw signal data;

[0068] An extraction module 702, configured to extract the base sequence data in the first dataset;

[0069] An invocation module 703, configured to invoke a preset SSDC compressor to compress the second dataset in combination with the extracted base sequence data to obtain compressed data.

[0070] When the above data compression device is implemented, the acquisition module 701 is used to acquire nanopore sequencing files and extract the dataset part from the nanopore sequencing files. The extraction module 702 is used to extract the base sequence data from the first dataset. The calling module 703 is used to call a preset SSDC compressor to compress the second dataset in combination with the extracted base sequence data to obtain compressed data. That is, when this technical solution is implemented, the base sequence data is extracted in advance. Since the base sequence data has a high correlation with the raw signal data in the second dataset and most of the characteristics of the raw signal data are considered, the lossless compression SSDC compressor is then used to compress the second dataset in combination with the base sequence data, thereby greatly improving the compression performance and ensuring the compression quality.

[0071] Figure 8 FIG. shows the electronic device provided in the fourth embodiment of the present invention, and this electronic device can be used to implement the nanopore sequencing raw signal data compression method in any of the foregoing embodiments. The electronic device includes:

[0072] A memory 801, a processor 802, a bus 803, and a computer program stored on the memory 801 and executable on the processor 802. The memory 801 and the processor 802 are connected through the bus 803. When the processor 802 executes this computer program, the nanopore sequencing raw signal data compression method in the foregoing embodiments is implemented. Among them, the number of processors can be one or more.

[0073] The memory 801 can be a high-speed random access memory (RAM, Random Access Memory) or a non-volatile memory, such as a disk memory. The memory 801 is used to store executable program codes, and the processor 802 is coupled to the memory 801.

[0074] Furthermore, an embodiment of the present application also provides a computer-readable storage medium. This computer-readable storage medium can be disposed in the electronic device in the foregoing embodiments, and this computer-readable storage medium can be a memory.

[0075] A computer program is stored on this computer-readable storage medium, and when this program is executed by a processor, the nanopore sequencing raw signal data compression method in the foregoing embodiments is implemented. Furthermore, this computer-readable storage medium can also be various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a RAM, a magnetic disk, or an optical disc that can store program codes.

[0076] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of devices or modules can be in electrical, mechanical or other forms.

[0077] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they may be located in one place, or may be distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0078] In addition, each functional module in various embodiments of the present application can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules.

[0079] If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The foregoing readable storage medium includes: USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical disks and other various media that can store program codes.

[0080] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily all essential to the present application.

[0081] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For the parts not elaborated in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0082] The above are only the preferred embodiments of the present invention, and do not limit the patent scope of the present invention accordingly. Any equivalent structural or equivalent process transformations made by using the content of the specification and drawings of the present invention, or directly or indirectly applied in other related technical fields, are equally included in the patent protection scope of the present invention.

Claims

1. A method for compressing raw signal data of nanopore sequencing, characterized in that, Including: Obtain a nanopore sequencing file, and extract the dataset part in the nanopore sequencing file; wherein, the dataset part includes a first dataset for storing sequencing base data and a second dataset for storing sequencing raw signal data; Extract the base sequence data in the first dataset; Call a preset SSDC compressor to compress the second dataset by combining the extracted base sequence data to obtain compressed data; The step of calling a preset SSDC compressor to compress the second dataset by combining the extracted base sequence data to obtain compressed data specifically includes: Call the pore model in the preset SSDC compressor to process the base sequence data to generate a first expected signal; Perform a mapping process on the first expected signal to obtain a second expected signal; wherein, the length of the second expected signal is equal to the length of the sequencing raw signal data; Perform a compression process on the second expected signal and the sequencing raw signal data to obtain compressed data; The step of performing a mapping process on the first expected signal to obtain a second expected signal specifically includes: Determine the head data and tail data in the sequencing raw signal data; Perform an end-to-end mapping process on the first expected signal according to the head data and tail data to obtain a second expected signal.

2. The nanopore sequencing raw signal data compression method according to claim 1, wherein The step of performing a compression process on the second expected signal and the sequencing raw signal data to obtain compressed data specifically includes: Obtain a first time series according to the second expected signal, and obtain a second time series according to the sequencing raw signal data; Perform calculations on the first time series and the second time series to obtain an alignment path and an alignment path distance; Obtain compressed data according to the alignment path and the alignment path distance.

3. The nanopore sequencing raw signal data compression method according to claim 2, wherein The step of performing calculations on the first time series and the second time series to obtain an alignment path and an alignment path distance specifically includes: Establish a distance matrix through the first time series and the second time series; wherein, the distance matrix is composed of multiple elements; Calculate a target path with the smallest sum of elements in the distance matrix; wherein, the initial position of the target path is the upper left corner of the distance matrix, and the end position of the target path is the lower right corner of the distance matrix; Take the target path as the alignment path, and take the sum of the elements corresponding to the target path as the alignment path distance.

4. The nanopore sequencing raw signal data compression method according to claim 3, wherein The step of obtaining compressed data according to the alignment path and the alignment path distance specifically includes: Determine target coordinates in the alignment path whose occurrence times are greater than a preset number threshold; Delete the target coordinates in the alignment path, and delete the distance values corresponding to the target coordinates in the alignment path distance to obtain a target alignment path and a target alignment path distance; Wherein, the target alignment path and the target alignment path distance form compressed data.

5. The nanopore sequencing raw signal data compression method according to claim 1, characterized in that The step of calling the pore model in the preset SSDC compressor to process the base sequence data to generate a first expected signal specifically includes: Process the base sequence data by calling the k-mer pore model in the preset SSDC compressor to generate a first expected signal.

6. A data compression device, characterized in that, It includes: An acquisition module for acquiring a nanopore sequencing file and extracting the dataset part in the nanopore sequencing file; wherein, the dataset part includes a first dataset for storing sequencing base data and a second dataset for storing sequencing raw signal data; An extraction module for extracting the base sequence data in the first dataset; A calling module for processing the base sequence data by calling the pore model in the preset SSDC compressor to generate a first expected signal; determining the head data and tail data in the sequencing raw signal data, and performing end-to-end mapping processing on the first expected signal according to the head data and tail data to obtain a second expected signal; performing compression processing on the second expected signal and the sequencing raw signal data to obtain compressed data; wherein, the length of the second expected signal is equal to the length of the sequencing raw signal data.

7. An electronic device, characterized in that, It includes a memory, a processor and a bus; The bus is used to realize the connection and communication between the memory and the processor; The processor is used to execute the computer program stored on the memory; When the processor executes the computer program, it realizes the steps in the nanopore sequencing raw signal data compression method described in any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it realizes the steps in the nanopore sequencing raw signal data compression method described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Compression method for next generation sequencing data

    CN105760706A

  • Multi-thread fast storage lossless compression method and system for FASTQ data

    CN106100641A