Sequence data compression method and device, electronic equipment and storage medium

By clustering and constructing similar maps of base sequences, determining the reference sequence and encoding the same and differential sequences, the problem of the lack of a universal reference sequence in multiple species base sequences is solved, and the compression ratio of the base sequence is improved.

CN120375931APending Publication Date: 2025-07-25MGI TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410102769.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-24
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The lack of a common reference sequence between base sequences of multiple different species leads to a problem of low base sequence compression ratio.

Method used

The base sequences in the target dataset are clustered using the preset clustering algorithm to construct a similar graph, determine that the base sequence corresponding to the central node is a reference sequence, and the same sequence and differential sequence are encoded through the preset encoding method.

Benefits of technology

The compression ratio of base sequences is improved, the problem of lack of reference sequences of base sequences is solved, and more efficient data compression is achieved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120375931A_ABST
    Figure CN120375931A_ABST
Patent Text Reader

Abstract

The invention provides a sequence data compression method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring a target data set, wherein the target data set comprises at least two base sequences; clustering the at least two base sequences by using a preset clustering algorithm to obtain at least one cluster; constructing a similar graph corresponding to each cluster; wherein nodes in the similar graph are base sequences in the clustering clusters, center nodes in the similar graph are at least one of the nodes, the base sequence corresponding to each center node in the similar graph is a reference sequence, and edges between the nodes in the similar graph are used for representing the similarity of the corresponding base sequences; determining the same sequence and the difference sequence between the base sequence corresponding to each central node and the base sequence corresponding to the node connected with the corresponding central node; and coding the same sequence and the difference sequence according to a preset coding method to obtain compressed data of the target data set.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of data compression technology, and in particular, to a method, apparatus, electronic device, and storage medium for compressing sequence data. Background Art

[0002] In the related art, due to the high sequence repeatability among the base sequences of a single species, it is convenient to compress the base sequences. In the related art, a compression algorithm based on a reference sequence or a compression algorithm without a reference sequence is usually used to compress the base sequences. Among them, the compression algorithm based on a reference sequence usually uses a general reference sequence to compress the base sequences; the compression algorithm without a reference sequence usually utilizes the self-similarity within the base sequences or the high repeatability among the base sequences to find a reference sequence, and then compresses the base sequences.

[0003] However, there may be a large number of variant sequences and unknown sequences among the base sequences of multiple different species. Therefore, there is usually a lack of a general reference sequence among the base sequences of multiple different species. Among them, the reference sequence is the same sequence in the base sequence. If there is no general same sequence in the base sequence, it means that there are many different sequences in the base sequence, which will lead to the problem of low compression ratio of the base sequence. Summary of the Invention

[0004] The present disclosure provides a method, apparatus, electronic device, and storage medium for compressing sequence data to solve the problems in the related art.

[0005] The first aspect embodiment of the present disclosure proposes a method for compressing sequence data, the method includes:

[0006] Obtain a target data set, where the target data set includes at least two base sequences;

[0007] Cluster at least two base sequences by using a preset clustering algorithm to obtain at least one clustering cluster; where each clustering cluster in the at least one clustering cluster includes at least one base sequence;

[0008] Construct a similarity graph corresponding to each clustering cluster; where the nodes in the similarity graph are the base sequences in the clustering cluster, the central nodes in the similarity graph are at least one of the nodes, and the base sequence corresponding to each central node in the similarity graph is the reference sequence, and the edges between the nodes in the similarity graph are used to represent the similarity of the corresponding base sequences;

[0009] Determine the same sequences and different sequences between the base sequence corresponding to each central node and the base sequence corresponding to the node connected to the corresponding central node;

[0010] Encode the same sequences and different sequences according to a preset encoding method to obtain the compressed data of the target data set.

[0011] In some embodiments of the present disclosure, the preset clustering algorithm is a hash-based clustering algorithm;

[0012] Using the preset clustering algorithm to cluster at least two base sequences to obtain at least one cluster, including:

[0013] According to the preset first length, each base sequence in the at least two base sequences is divided into at least one first short sequence;

[0014] Determine that any two base sequences with the number of identical first short sequences in the at least two base sequences greater than the preset first number belong to the same cluster.

[0015] In some embodiments of the present disclosure, constructing a similarity graph corresponding to each cluster, including:

[0016] According to the preset second length, each base sequence in each cluster is divided into at least one second short sequence;

[0017] Determine the number of identical second short sequences between the first base sequence in each cluster and the second base sequence in the corresponding cluster; the first base sequence refers to each base sequence in the cluster, and the second base sequence refers to other base sequences in the corresponding cluster except the corresponding first base sequence;

[0018] Determine that in each cluster, the first base sequence with the number of identical second short sequences greater than the preset second number is the central node, the corresponding second base sequence is the node, and the number of identical second short sequences between the first base sequence and the corresponding second base sequence is the edge, and construct a similarity graph corresponding to each cluster.

[0019] In some embodiments of the present disclosure, before determining the identical sequences and different sequences between the base sequence corresponding to each central node and the base sequence corresponding to the node connected to the corresponding central node, the method further includes:

[0020] Sort the edges between each central node and the node connected to the corresponding central node according to the number of identical second short sequences;

[0021] Delete the edges with a relatively late order corresponding to each central node according to the preset number of edges of the central node.

[0022] In some embodiments of the present disclosure, determining the identical sequences and different sequences between the base sequence corresponding to each central node and the base sequence corresponding to the node connected to the corresponding central node, including:

[0023] Determine that the same second-shortest sequence among the base sequences corresponding to each central node and the base sequences corresponding to the nodes connected to the corresponding central node is the identical sequence, and the sequences other than the second-shortest sequence are the differential sequences.

[0024] In some embodiments of the present disclosure, the preset encoding method includes dictionary encoding, transform encoding, or entropy encoding.

[0025] In some embodiments of the present disclosure, obtaining a target data set includes:

[0026] Obtain a data set, where the data set includes at least two base sequences;

[0027] Divide the data set into at least one target data set according to the data volume or the number of base sequences.

[0028] An embodiment of the second aspect of the present disclosure provides a sequence data compression device, which includes:

[0029] An obtaining unit, configured to obtain a target data set, where the target data set includes at least two base sequences;

[0030] A clustering unit, configured to cluster at least two base sequences by using a preset clustering algorithm to obtain at least one clustering cluster; wherein, each clustering cluster in the at least one clustering cluster includes at least one base sequence;

[0031] A similarity graph construction unit, configured to construct a similarity graph corresponding to each clustering cluster; wherein, the nodes in the similarity graph are the base sequences in the clustering cluster, the central node in the similarity graph is at least one of the nodes, and the base sequence corresponding to each central node in the similarity graph is a reference sequence, and the edges between the nodes in the similarity graph are used to represent the similarity of the corresponding base sequences;

[0032] A determination unit, configured to determine the identical sequences and differential sequences between the base sequences corresponding to each central node and the base sequences corresponding to the nodes connected to the corresponding central node;

[0033] An encoding unit, configured to encode the identical sequences and differential sequences according to a preset encoding method to obtain compressed data of the target data set.

[0034] An embodiment of the third aspect of the present disclosure provides an electronic device, including:

[0035] At least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in the first aspect embodiment of the present disclosure.

[0036] A fourth aspect embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, where the computer instructions are used to cause a computer to execute the method described in the first aspect embodiment of the present disclosure.

[0037] In summary, the present disclosure provides a sequence data compression method, apparatus, electronic device, and storage medium. The method includes: obtaining a target data set, where the target data set includes at least two base sequences; clustering the at least two base sequences by using a preset clustering algorithm to obtain at least one clustering cluster; where each clustering cluster in the at least one clustering cluster includes at least one base sequence; constructing a similarity graph corresponding to each clustering cluster; where the nodes in the similarity graph are the base sequences in the clustering cluster, the central nodes in the similarity graph are at least one of the nodes, and the base sequence corresponding to each central node in the similarity graph is a reference sequence, and the edges between the nodes in the similarity graph are used to represent the similarity of the corresponding base sequences; determining the identical sequences and different sequences between the base sequence corresponding to each central node and the base sequences corresponding to the nodes connected to the corresponding central node; encoding the identical sequences and different sequences according to a preset encoding method to obtain the compressed data of the target data set.

[0038] Through the solution provided by the present disclosure, at least two base sequences in the target data set are clustered by using a preset clustering algorithm to obtain at least one clustering cluster; according to the similarity between the base sequences in each clustering cluster, a similarity graph corresponding to each clustering cluster is constructed; according to the similarity graph corresponding to each clustering cluster, it is determined that the base sequence corresponding to the central node in the similarity graph is the reference sequence; the identical sequences and different sequences between the base sequence corresponding to each central node and the base sequences corresponding to the nodes connected to the corresponding central node are determined; the identical sequences and different sequences in each clustering cluster are encoded according to a preset encoding method to obtain the compressed data of the target data set. By determining that the base sequence corresponding to each central node in the similarity graph is the reference sequence, the problem that the base sequences in the target data set lack a reference sequence is solved, and then the identical sequences and different sequences in each clustering cluster are encoded according to a preset encoding method, improving the compression ratio of the base sequences.

[0039] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure and do not constitute an improper limitation of the present disclosure.

[0041] Figure 1 It is a flowchart of the sequence data compression method provided by the embodiment of the present disclosure;

[0042] Figure 2 Flow diagram of the method for determining clustering clusters provided by the embodiments of the present disclosure;

[0043] Figure 3 Flow diagram of the method for constructing a similarity graph corresponding to each clustering cluster provided by the embodiments of the present disclosure;

[0044] Figure 4 Flow diagram of the method for optimizing the similarity graph provided by the embodiments of the present disclosure;

[0045] Figure 5 Flow diagram of the sequence data compression method provided by the application example of the present disclosure;

[0046] Figure 6 Schematic diagram of the compression ratio of representative tools of different compression methods provided by the application example of the present disclosure on 10 gut metagenomic samples;

[0047] Figure 7 Schematic diagram of the compression time of representative tools of different compression methods provided by the application example of the present disclosure on 10 gut metagenomic samples;

[0048] Figure 8 Schematic diagram of the structure of the sequence data compression device provided by the embodiments of the present disclosure;

[0049] Figure 9 Schematic diagram of the hardware composition structure of the electronic device provided by the embodiments of the present disclosure. Detailed implementation manners

[0050] The embodiments of the present disclosure will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to explain the present disclosure, but should not be construed as limiting the present disclosure.

[0051] Data compression technology is a process of encoding the original data with less space. Data compression technology refers to a method of removing redundant data to reduce storage space and improve the transmission, storage, and processing efficiency of data without losing useful information. Data compression technology is further divided into lossy compression and lossless compression according to whether some data is sacrificed. Since the base sequence contains a large amount of biological information, the data compression technology used for the base sequence is usually lossless compression.

[0052] Currently, the compression methods applicable to base sequence data can generally be divided into reference-sequence-based compression methods and reference-sequence-free compression methods. Among them, the reference-sequence-based data compression method uses the difference information between the reference sequence and the base sequence for compression, and the reference-sequence-free data compression method uses the self-similarity within the base sequence or the high repeatability between base sequences to achieve compression. The reference-sequence-free data compression method can be further divided into context-based compression methods (representative tools Quip and Fqzcomp) and sequence-splicing-based compression methods (representative tool CoLoRd).

[0053] The reference-sequence-based data compression method can usually obtain a higher compression ratio. However, there may be a large number of variant sequences and unknown sequences between different species, and these characteristics make it lack a universal reference sequence for base sequences. Among them, the reference sequence is the same sequence in the base sequence. If there is no universal same sequence in the base sequence, it means that there are more different sequences in the base sequence, which will further lead to the problem of low compression ratio of the base sequence.

[0054] To solve the defects in the related art, the present disclosure uses a preset clustering algorithm to cluster at least two base sequences in a target data set to obtain at least one clustering cluster; constructs a similarity graph corresponding to each clustering cluster according to the similarity between the base sequences in each clustering cluster; determines the base sequence corresponding to the central node in the similarity graph as the reference sequence according to the similarity graph corresponding to each clustering cluster; determines the same sequences and different sequences between the base sequence corresponding to each central node and the base sequence corresponding to the node connected to the corresponding central node; encodes the same sequences and different sequences according to a preset encoding method to obtain the compressed data of the target data set. By determining the base sequence corresponding to each central node in the similarity graph as the reference sequence, the problem that the base sequence in the target data set lacks a reference sequence is solved, and then the same sequences and different sequences in each clustering cluster are encoded according to the preset encoding method, improving the compression ratio of the base sequence.

[0055] The following further describes the present disclosure in detail with reference to the accompanying drawings and specific embodiments.

[0056] As Figure 1 shown Figure 1 is a schematic flowchart of the sequence data compression method provided by an embodiment of the present disclosure. The sequence data compression method provided by an embodiment of the present disclosure includes the following steps:

[0057] Step 101, obtain a target data set, where the target data set includes at least two base sequences;

[0058] In one embodiment, metagenomic sequencing typically employs second-generation sequencing (Next-generation sequencing, NGS) technology. After metagenomic sequencing using NGS technology, short read sequences as shown in Table 1 are output.

[0059] Table 1

[0060]

[0061] The short read sequences shown in Table 1 are in the FASTQ file format. In a FASTQ file, the first line is the base sequence identifier, usually starting with the "@" character, followed by the unique identifier of the base sequence and an optional description; the second line is the base sequence line, storing the base sequence obtained by sequencing; the third line is a delimiter, usually starting with "+", often storing additional information; the fourth line is the quality value, composed of ASCII characters, with each character corresponding to the quality score of a base.

[0062] The base sequence in this disclosure refers to the short read sequence shown in the second line of Table 1 corresponding to the short read sequence.

[0063] In one embodiment, the target data set can be obtained from a nucleic acid sequence database, such as the European Nucleotide Archive (ENA).

[0064] In one embodiment, the target data set can be obtained from a metagenomic sequencing experiment, such as the target data set obtained when sequencing the gut metagenome using NGS technology.

[0065] Step 102: Cluster at least two base sequences using a preset clustering algorithm to obtain at least one cluster; wherein each cluster in the at least one cluster includes at least one base sequence.

[0066] In one embodiment, the preset clustering algorithm can be a hash-based clustering algorithm or a distance-based clustering algorithm.

[0067] In one embodiment, taking the preset clustering algorithm being a hash-based clustering algorithm as an example:

[0068] When clustering at least two base sequences using a hash-based clustering algorithm, each base sequence can be divided into at least one first short sequence of the same length; then, based on whether there are the same first short sequences in the at least one first short sequence corresponding to each base sequence as those in the at least one first short sequence corresponding to other base sequences, the at least two base sequences are clustered to obtain at least one cluster.

[0069] In one embodiment, taking the preset clustering algorithm being a distance-based clustering algorithm as an example:

[0070] When clustering at least two base sequences using a distance-based clustering algorithm, the edit distance between any two base sequences can be calculated; then, it is determined that two base sequences with an edit distance greater than a preset edit distance threshold belong to the same cluster, and finally, at least one cluster is obtained.

[0071] In one embodiment, in the target dataset, depending on the different clustering results of at least two base sequences, there may be one cluster, that is, all base sequences in the target dataset belong to one class.

[0072] In one embodiment, in the target dataset, depending on the different clustering results of at least two base sequences, there may be multiple clusters, that is, all base sequences in the target dataset belong to multiple different classes.

[0073] In the present disclosure, the number of clusters in the target dataset is not limited.

[0074] Step 103, construct a similarity graph corresponding to each cluster; wherein, the nodes in the similarity graph are the base sequences in the cluster, the central nodes in the similarity graph are at least one of the nodes, and the base sequence corresponding to each central node in the similarity graph is a reference sequence, and the edges between the nodes in the similarity graph are used to represent the similarity of the corresponding base sequences;

[0075] In one embodiment, the similarity is used to represent the similarity between all base sequences in each cluster.

[0076] In one embodiment, the similarity graph contains at least one central node and at least one ordinary node (node). Among them, the node and the central node are relative. For example, a node relative to the first central node may be the central node of other nodes. The edge between the central node and the node in the similarity graph represents the similarity between the base sequence corresponding to the central node and the base sequence corresponding to the node. Among them, the similarity can be measured by the distance between the nodes or by the length of the same sequence of the corresponding base sequences.

[0077] In one embodiment, the method for constructing the similarity graph can adopt any graph construction method for a multi-node network, and only need to replace the nodes in the graph with base sequences and replace the edges in the graph with the similarity between the base sequences corresponding to the nodes. In the present disclosure, the specific method for constructing the similarity graph is not limited.

[0078] In one embodiment, in the similarity graph corresponding to each cluster, there may be only one central node. That is, in the corresponding cluster, all other base sequences have a certain similarity to the base sequence corresponding to the central node.

[0079] In one embodiment, in the similarity graph corresponding to each cluster, there may be multiple central nodes. The present disclosure does not limit the number of central nodes included in each similarity graph.

[0080] Step 104: Determine the identical sequences and different sequences between the base sequence corresponding to each central node and the base sequence corresponding to the node connected to the corresponding central node.

[0081] In one embodiment, the base sequence corresponding to the central node and the base sequence corresponding to the node connected thereto must have a certain degree of similarity.

[0082] In one embodiment, taking the hash-based clustering algorithm in step 102 as an example, the identical sequences can be the same first short sequences.

[0083] In one embodiment, the number of the same first short sequences can be one or multiple. In the present disclosure, the number of the same first short sequences is not limited.

[0084] In one embodiment, the different sequences can refer to the sequences in the base sequence except for the corresponding identical first short sequences.

[0085] In one embodiment, the different sequences can be the differences between the base sequence and the base sequence corresponding to the corresponding central node, or the differences between the base sequence and the corresponding identical first short sequences.

[0086] Step 105: Encode the identical sequences and different sequences according to a preset encoding method to obtain the compressed data of the target data set.

[0087] In one embodiment, the preset encoding method includes dictionary encoding, transform encoding or entropy encoding.

[0088] In one embodiment, taking the preset encoding method as dictionary encoding as an example:

[0089] First, pre-define the encoding values corresponding to the first short sequences.

[0090] Then, according to the differences between the different sequences and the identical first short sequences, define the encoding rules for the different sequences. For example, if the different sequences are sequences obtained by events such as insertion and deletion from the corresponding identical first short sequences, then the encoding rules for events such as insertion and deletion can be preset, and then the different sequences can be encoded.

[0091] In one embodiment, after obtaining the compressed data, a preset decoding method corresponding to the preset encoding method can also be used to decode the compressed data to obtain each base sequence in the target data set.

[0092] In one embodiment, taking the preset clustering algorithm as the hash-based clustering algorithm as an example, such as Figure 2As shown, step 102 includes:

[0093] Step 201, according to a preset first length, divide each of at least two base sequences into at least one first short sequence;

[0094] In one embodiment, the preset first length can be any positive integer less than or equal to the length of each of at least two base sequences.

[0095] In one embodiment, the first short sequence refers to all base sequences with a length of the preset first length included in each of at least two base sequences.

[0096] In one embodiment, each of the at least one first short sequences may be the same or different.

[0097] In one embodiment, if the length of each of at least two base sequences is n and the preset first length is k, then each of at least two base sequences contains n - k + 1 first short sequences.

[0098] In one embodiment, preferably, the preset first length is a positive integer greater than 1.

[0099] Step 202, determine that any two base sequences with the number of identical first short sequences in at least two base sequences greater than the preset first number belong to the same cluster.

[0100] In one embodiment, compare whether the first short sequences included in each of at least two base sequences are the same to determine the number of identical first short sequences included in each of at least two base sequences.

[0101] In one embodiment, the preset first number refers to the critical number that can divide any two base sequences into the same cluster, that is, if the number of identical first short sequences in at least two base sequences is greater than or equal to this critical number, then determine that at least two base sequences belong to the same cluster.

[0102] In one embodiment, the critical number that can divide any two base sequences into the same cluster can be obtained by referring to relevant scientific literature on determining that two base sequences belong to the same cluster, or can be obtained according to experiments on determining that two base sequences belong to the same cluster.

[0103] In one embodiment, the same first short sequence in at least two base sequences may be at the same position or different positions in the two base sequences. For example, if the first base sequence and the second base sequence in the target dataset contain the same first short sequence, and if this first short sequence is at position A in the first base sequence, then this first short sequence may be at any position in the second base sequence, where any position includes position A and other positions except position A.

[0104] In one embodiment, position A in the first base sequence refers to the position of any one of the first short sequences in the first base sequence.

[0105] In one embodiment, as Figure 3 shown, step 103 includes:

[0106] Step 301, according to a preset second length, divide each base sequence in each cluster into at least one second short sequence;

[0107] In one embodiment, the preset second length may be any positive integer less than or equal to the length of each base sequence in each cluster.

[0108] In one embodiment, the preset second length may be the same as or different from the preset first length.

[0109] In one embodiment, the second short sequence refers to a base sequence with a length of the preset second length.

[0110] In one embodiment, if the preset second length is the same as the preset first length, then the first short sequence in step 201 may be directly used to replace the second short sequence in step 301.

[0111] Step 302, determine the number of the same second short sequences between the first base sequence in each cluster and the second base sequence in the corresponding cluster; the first base sequence refers to each base sequence in the cluster, and the second base sequence refers to other base sequences in the corresponding cluster except the corresponding first base sequence;

[0112] In one embodiment, the first base sequence may be any one of the base sequences in each cluster.

[0113] In one embodiment, compare whether the second short sequences between the first base sequence in each cluster and the second base sequence in the corresponding cluster are the same, and accumulate the number of the same second short sequences, so as to determine the number of the same second short sequences between the first base sequence in each cluster and the second base sequence in the corresponding cluster.

[0114] Step 303: Determine, for each cluster, that the first base sequences with the number of the same second-shortest sequences greater than a preset second number are central nodes, the corresponding second base sequences are nodes, and the number of the same second-shortest sequences between the first base sequence and the corresponding second base sequence is an edge, and construct a similarity graph corresponding to each cluster.

[0115] In one embodiment, the preset second number refers to the number of the same second-shortest sequences between the corresponding first base sequence and the second base sequence.

[0116] In one embodiment, one cluster corresponds to one similarity graph.

[0117] In one embodiment, as Figure 4 shown, before step 104, the sequence data compression method provided by the embodiments of the present disclosure further includes:

[0118] Step 401: Sort the edges between each central node and the nodes connected to the corresponding central node according to the number of the same second-shortest sequences.

[0119] In one embodiment, sort the edges between each central node and the nodes connected to the corresponding central node in each cluster in descending order according to the number of the same second-shortest sequences between each central node in each cluster and the nodes connected to the corresponding central node in the corresponding cluster.

[0120] In one embodiment, the edges between each central node and the nodes connected to the corresponding central node can be sorted from an association table of the common number of the second-shortest sequences between the base sequence corresponding to each central node and the base sequences corresponding to the nodes connected to the corresponding central node.

[0121] Step 402: Delete the edges with a lower sorting order corresponding to each central node according to the preset number of edges of the central node.

[0122] In one embodiment, the preset number of edges means that the number of edges between each central node and the nodes connected to the corresponding central node does not exceed a preset number.

[0123] In one embodiment, the preset number can be set according to the similarity of the edges corresponding to the number of edges between each central node and the nodes connected to the corresponding central node.

[0124] In one embodiment, the edges with a lower sorting order corresponding to each central node refer to the edges exceeding the preset number of edges in the sorting result corresponding to each central node.

[0125] In one embodiment, if the number of edges between each central node and the nodes connected to the corresponding central node exceeds a preset number, then according to the sorting result of the edges between each central node and the nodes connected to the corresponding central node, the first preset number of edges in the sorting result corresponding to each central node are retained, and the edges in the sorting result corresponding to each central node that are ranked after the preset number of edges are deleted.

[0126] In one embodiment, by deleting the edges with lower rankings corresponding to each central node, the optimization of the similarity graph corresponding to each clustering cluster is achieved. In the optimized similarity graph, each central node and the base sequence corresponding to each central node, that is, the reference sequence, can be found more intuitively.

[0127] In one embodiment, step 104 includes:

[0128] Determine that the same second-shortest sequence in the base sequence corresponding to each central node and the base sequence corresponding to the node connected to the corresponding central node is the same sequence, and the sequences other than the second-shortest sequence are the different sequences.

[0129] In one embodiment, the same sequence refers to a base sequence with the same length and the same base types (adenine A, thymine T, cytosine C, guanine G) at corresponding positions.

[0130] In one embodiment, the different sequence refers to the base sequence other than the same sequence in the base sequence corresponding to each central node and the base sequence corresponding to the node connected to the corresponding central node.

[0131] In one embodiment, determine the same sequence and different sequence of the base sequence corresponding to each central node and the base sequence corresponding to the node connected to the corresponding central node in each clustering cluster.

[0132] In one embodiment, the preset coding method includes dictionary coding, transform coding or entropy coding.

[0133] In one embodiment, dictionary coding starts from the first base of the base sequence and replaces the strings that have appeared in the dictionary with an index value to achieve the purpose of compression, such as Lempel Ziv (LZ) coding.

[0134] In one embodiment, transform coding realizes the compression of base sequence data by converting the same sequence and different sequence from one domain (time domain or spatial domain) to another domain (frequency domain or wavelet domain) and encoding in the new domain. Transform coding includes discrete cosine transform (converting from spatial domain to frequency domain), discrete Fourier transform (converting from time domain to frequency domain), etc.

[0135] In one embodiment, entropy encoding compresses sequence data by using shorter encodings to represent identical sequences and longer encodings to represent different sequences, effectively reducing the encoding length. Entropy encoding mainly includes arithmetic encoding, Huffman encoding, etc.

[0136] In one embodiment, the model can be trained using the identical information, different information, and corresponding encodings in the base sequence to obtain a trained encoding model. For subsequent sequences, if the input base sequence can match the base sequence in the model, the corresponding encoding can be directly output. If there are significant differences between the input base sequence and the base sequence in the model, the difference information can be recorded and the encoding model can be continuously updated to include these new difference information.

[0137] In one embodiment, step 101 includes:

[0138] Obtain a data set, where the data set includes at least two base sequences;

[0139] In one embodiment, the data set can be obtained from ENA.

[0140] In one embodiment, the data set can be obtained from a metagenomic sequencing experiment, such as a data set obtained when using NGS technology to sequence the gut metagenome.

[0141] Divide the data set into at least one target data set according to the data volume or the number of base sequences.

[0142] In one embodiment, the at least one target data set can be one, two, or more target data sets.

[0143] In one embodiment, the data set is divided into at least one target data set according to the data volume. For example, the value corresponding to the data volume is any value not exceeding the value corresponding to the total amount of data in the data set.

[0144] In one embodiment, the data set is divided into at least one target data set according to the number of base sequences. For example, the number of base sequences is any value not exceeding the number of all base sequences in the data set.

[0145] In summary, the solution provided by the present disclosure:

[0146] First, use a preset clustering algorithm to cluster at least two base sequences in a target data set to obtain at least one clustering cluster; construct a similarity graph corresponding to each clustering cluster according to the similarity between the base sequences in each clustering cluster; determine that the base sequence corresponding to the central node in the similarity graph is the reference sequence according to the similarity graph corresponding to each clustering cluster; determine the identical sequences and different sequences between the base sequence corresponding to each central node and the base sequence corresponding to the node connected to the corresponding central node; encode the identical sequences and different sequences according to a preset encoding method to obtain the compressed data of the target data set. By determining that the base sequence corresponding to each central node in the similarity graph is the reference sequence, the problem that the base sequences in the target data set lack a reference sequence is solved, and then the identical sequences and different sequences in each clustering cluster are encoded according to a preset encoding method, improving the compression ratio of the base sequences.

[0147] Secondly, sort the edges between each central node and the node connected to the corresponding central node according to the number of the same second-shortest sequences; delete the edges with a later sort order corresponding to each central node according to the preset number of edges of the central node. By deleting the edges with a later sort order corresponding to each central node according to the edge sorting result between each central node and the node connected to the corresponding central node, the similarity graph corresponding to each clustering cluster is divided into multiple clusters, facilitating the identification of identical sequences and different sequences between the base sequence corresponding to each central node and the base sequence corresponding to the node connected to the corresponding central node.

[0148] The following uses an application example to further illustrate the sequence data compression method provided by the present disclosure:

[0149] As Figure 5 shown, Figure 5 is a schematic flowchart of the sequence data compression method provided by the application example of the present disclosure. The sequence data compression method provided by the application example of the present disclosure includes the following steps:

[0150] Step 501, obtain a data set, where the data set includes at least two base sequences;

[0151] In one embodiment, the data set is obtained when sequencing the intestinal metagenome using NGS technology.

[0152] Step 502, divide the data set into at least one target data set according to the data volume or the number of base sequences;

[0153] Step 503, obtain a target data set, where the target data set includes at least two base sequences;

[0154] In one embodiment, 10 intestinal metagenome samples are selected from the data set.

[0155] Step 504: Divide each of the at least two base sequences into at least one first short sequence according to a preset first length;

[0156] Step 505: Determine that any two base sequences in which the number of identical first short sequences in the at least two base sequences is greater than a preset first number belong to the same clustering cluster;

[0157] Step 506: Divide each base sequence in each clustering cluster into at least one second short sequence according to a preset second length;

[0158] Step 507: Determine the number of identical second short sequences between a first base sequence and a corresponding second base sequence in each clustering cluster;

[0159] Step 508: Determine that, in each clustering cluster, a first base sequence with the number of identical second short sequences greater than a preset second number is a central node, the corresponding second base sequence is a node, and the number of identical second short sequences between the first base sequence and the corresponding second base sequence is an edge, and construct a similarity graph corresponding to each clustering cluster;

[0160] Step 509: Sort the edges between each central node and the nodes connected to the corresponding central node according to the number of identical second short sequences;

[0161] Step 510: Delete the edges with a later order corresponding to each central node according to the preset number of edges of the central node;

[0162] Step 511: Determine that, among the base sequences corresponding to each central node and the base sequences of the nodes connected to the corresponding central node, the identical second short sequences are identical sequences, and the sequences other than the second short sequences are different sequences;

[0163] Step 512: Encode the identical sequences and different sequences according to dictionary encoding, transform encoding or entropy encoding to obtain the compressed data of the target data set.

[0164] In one embodiment, in order to verify the compression effect of the sequence data compression method proposed in this application, the sequence data compression method proposed in this application is compared with the sequence data compression methods in the related art, and the performance of various sequence data compression methods is measured from two indicators of compression ratio and compression time.

[0165] In one embodiment, the intestinal metagenomic data is compressed by using a general compression method (representative tool Gzip), a reference-free compression method (representative tools Quip, Fqzcomp, CoLoRd and Genozip), and a reference-based compression algorithm (representative tool Genozip-ref) respectively.

[0166] In one embodiment, the compression ratios (the size of the sample before compression divided by the size after compression) and compression times using different compression methods are calculated, and the calculation results are as Figure 6 and Figure 7 shown.

[0167] Among them, as Figure 6 can be seen, when using CoLoRd compression, the compression ratios of 10 samples are from 13.73 to 15.11, all higher than those using Gzip, Quip, and Fqzcomp. The average compression ratio of all samples compressed by CoLoRd is 13.73, which is 2.90 times that of Gzip (4.74). Compared with Fqzcomp (8.62), Quip (8.13), and Genozip (7.65), it is increased by 59.29%, 68.88%, and 79.48% respectively. In the case of using a reference sequence, the average compression ratio of Genozip-ref reaches 10.82, which is increased by 41.55% compared with compression without a reference sequence, and the highest compression ratio can reach 19.10. These results indicate that the reference sequence has a significant improvement effect on the compression of metagenomic data. However, among the 10 test samples, the CoLoRd compression ratios of 6 samples are higher than those of Genozip-ref compression, and the compression ratios of 2 samples are similar, indicating that the reference-free sequence assembly compression method can also achieve comparable or even better effects than the reference-sequence-based compression method.

[0168] As Figure 7 can be seen, the Genozip compression mode shows extremely fast compression speed, followed by the CoLoRd and Genozip-ref compression modes. The average compression time of all samples compressed by CoLoRd is 14.85 minutes, which is similar to that of the Genozip-ref compression mode (14.63 minutes), and much lower than that of Gzip compression (46.71 minutes). These results show that the reference-free sequence assembly compression method can achieve a relatively fast compression speed without relying on a reference sequence, and at the same time avoid the time overhead of constructing a reference sequence and its index, featuring high compression ratio and fast compression speed.

[0169] To implement the sequence data compression method provided by the embodiments of the present disclosure, the embodiments of the present disclosure also provide a sequence data compression device, as Figure 8 shown. Figure 8 is a schematic structural diagram of the sequence data compression device provided by the embodiments of the present disclosure. The sequence data compression device 800 includes:

[0170] An acquisition unit 801, configured to acquire a target data set, where the target data set includes at least two base sequences;

[0171] A clustering unit 802 is configured to cluster at least two base sequences by using a preset clustering algorithm to obtain at least one clustering cluster; wherein, each clustering cluster in the at least one clustering cluster includes at least one base sequence;

[0172] A similarity graph construction unit 803 is configured to construct a similarity graph corresponding to each clustering cluster; wherein, the nodes in the similarity graph are the base sequences in the clustering cluster, the central nodes in the similarity graph are at least one of the nodes, and the base sequence corresponding to each central node in the similarity graph is a reference sequence, and the edges between the nodes in the similarity graph are used to represent the similarity of the corresponding base sequences;

[0173] A determination unit 804 is configured to determine the identical sequences and different sequences between the base sequence corresponding to each central node and the base sequence corresponding to the node connected to the corresponding central node;

[0174] An encoding unit 805 is configured to encode the identical sequences and different sequences according to a preset encoding method to obtain the compressed data of the target data set.

[0175] In one embodiment, the preset clustering algorithm is a hash-based clustering algorithm;

[0176] In one embodiment, the clustering unit 802 is specifically configured to:

[0177] According to a preset first length, each base sequence in the at least two base sequences is divided into at least one first short sequence;

[0178] Determine that any two base sequences in which the number of identical first short sequences in the at least two base sequences is greater than a preset first number belong to the same clustering cluster.

[0179] In one embodiment, the similarity graph construction unit 803 is specifically configured to:

[0180] According to a preset second length, each base sequence in each clustering cluster is divided into at least one second short sequence;

[0181] Determine the number of identical second short sequences between a first base sequence in each clustering cluster and a second base sequence in the corresponding clustering cluster; the first base sequence refers to each base sequence in the clustering cluster, and the second base sequence refers to other base sequences in the corresponding clustering cluster except the corresponding first base sequence;

[0182] Determine that in each clustering cluster, the first base sequence with the number of identical second short sequences greater than a preset second number is the central node, the corresponding second base sequence is the node, and the number of identical second short sequences between the first base sequence and the corresponding second base sequence is the edge, and construct a similarity graph corresponding to each clustering cluster.

[0183] In one embodiment, the sequence data compression device 800 further includes a deletion unit, and the deletion unit is configured to:

[0184] Sort the edges between each central node and the nodes connected to the corresponding central node according to the number of the same second-shortest sequences;

[0185] Delete the edges with a later order corresponding to each central node according to the preset number of edges of the central node.

[0186] In one embodiment, the determination unit 804 is specifically configured to:

[0187] Determine that the same second-shortest sequences in the base sequence corresponding to each central node and the base sequence corresponding to the node connected to the corresponding central node are the same sequences, and the sequences other than the second-shortest sequences are the differential sequences.

[0188] In one embodiment, the preset encoding method includes dictionary encoding, transform encoding or entropy encoding.

[0189] In one embodiment, the obtaining unit 801 is specifically configured to:

[0190] Obtain a data set, where the data set includes at least two base sequences;

[0191] Divide the data set into at least one target data set according to the data volume or the number of base sequences.

[0192] It should be noted that: when the sequence data compression device 800 provided in the above embodiment performs sequence data compression, only the above division of each program module is used for illustration. In practical applications, the above processing can be allocated to different program modules according to needs, that is, the internal structure of the sequence data compression device 800 is divided into different program modules to complete all or part of the above-described processing. In addition, the sequence data compression device 800 provided in the above embodiment and the sequence data compression method embodiment provided in the present disclosure belong to the same concept, and the specific implementation process can be seen in the method embodiment, which will not be elaborated here.

[0193] Figure 9 This is a schematic diagram of the hardware composition structure of the electronic device provided in the embodiment of the present disclosure. As Figure 9 shown, the electronic device 900 includes at least one processor 902; and a memory 901 communicatively connected to at least one processor 902; wherein, the memory 901 stores instructions executable by at least one processor 902, and the instructions are executed by at least one processor 902 to implement the steps of the sequence data compression method of the embodiment of the present disclosure.

[0194] Optionally, the electronic device may specifically be the sequence data compression device according to the embodiment of the present application, and the electronic device may implement the corresponding processes implemented by the sequence data compression device in each method of the embodiment of the present application. For the sake of brevity, details are not described herein again.

[0195] It can be understood that the electronic device further includes a communication interface 903. Each component in the electronic device is coupled together through a bus system 904. It can be understood that the bus system 904 is used to realize the connection and communication between these components. In addition to the data bus, the bus system 904 further includes a power bus, a control bus, and a status signal bus. However, for the sake of clear illustration, in Figure 9 all kinds of buses are labeled as the bus system 904.

[0196] It can be understood that the memory 901 can be a volatile memory or a non-volatile memory, or can include both volatile and non-volatile memories. Among them, the non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), a ferromagnetic random access memory (FRAM), a flash memory, a magnetic surface memory, an optical disc, or a compact disc read-only memory (CD-ROM); the magnetic surface memory can be a disk memory or a tape memory. The volatile memory can be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), synchronous static random access memory (SSRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), sync link dynamic random access memory (SLDRAM), direct rambus random access memory (DRRAM). The memory 901 described in the embodiments of the present invention is intended to include but not be limited to these and any other suitable types of memories.

[0197] The method disclosed in the embodiments of the present disclosure can be applied to or implemented by the processor 902. The processor 902 may be an integrated circuit chip with signal processing capabilities. During implementation, the steps of the above method can be completed by the integrated logic circuit in hardware or instructions in software form in the processor 902. The above-mentioned processor 902 may be a general-purpose processor, DSP, or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The processor 902 can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention. The general-purpose processor may be a microprocessor or any conventional processor, etc. Combining the steps of the method disclosed in the embodiments of the present invention can be directly embodied as being executed and completed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module may be located in a storage medium, and this storage medium is located in the memory 901. The processor 902 reads the information in the memory 901 and combines its hardware to complete the steps of the foregoing sequence data compression method.

[0198] In an exemplary embodiment, the electronic device can be implemented by one or more application-specific integrated circuits (ASICs, Application Specific Integrated Circuits), DSPs, programmable logic devices (PLDs, Programmable Logic Devices), complex programmable logic devices (CPLDs, Complex Programmable Logic Devices), FPGAs, general-purpose processors, controllers, MCUs, microprocessors (Microprocessors), or other electronic components, and is used to execute the foregoing method.

[0199] The embodiments of the present disclosure also provide a non-transitory computer-readable storage medium storing computer instructions, and the computer instructions are used to cause a computer to implement the steps of the sequence data compression method in the embodiments of the present invention when executed.

[0200] Optionally, the computer-readable storage medium can be applied to the sequence data compression device in the embodiments of the present application, and the computer instructions cause the computer to execute the corresponding processes implemented by the sequence data compression device in the various methods of the embodiments of the present application. For the sake of brevity, it will not be elaborated here.

[0201] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined, or can be integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the components shown or discussed with each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be electrical, mechanical, or other forms.

[0202] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units; some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0203] In addition, each functional unit in each embodiment of the present invention can be all integrated in a processing unit, or each unit can be separately used as a unit, or two or more units can be integrated in a unit; the above-mentioned integrated units can be implemented in the form of hardware, or in the form of hardware plus software functional units.

[0204] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it executes the steps including the above method embodiments; and the foregoing storage medium includes: removable storage devices, ROM, RAM, magnetic disks, or optical disks and other various media that can store program codes.

[0205] Alternatively, if the above-mentioned integrated units of the present invention are implemented in the form of software function modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of the present invention essentially or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the methods of the various embodiments of the present invention. And the foregoing storage medium includes: removable storage devices, ROM, RAM, magnetic disks, or optical disks and other various media that can store program codes.

[0206] The above are only the specific embodiments of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered within the protection scope of the present invention. Therefore, the protection scope of the present invention shall be subject to the protection scope of the claims.

Claims

1. A sequence data compression method, characterized in that, Including: Obtain a target data set, where the target data set includes at least two base sequences; Cluster the at least two base sequences by using a preset clustering algorithm to obtain at least one cluster; wherein, each of the at least one cluster includes at least one base sequence; Construct a similarity graph corresponding to each of the clusters; wherein, the nodes in the similarity graph are the base sequences in the cluster, the central nodes in the similarity graph are at least one of the nodes, and the base sequence corresponding to each central node in the similarity graph is a reference sequence, and the edges between the nodes in the similarity graph are used to represent the similarity of the corresponding base sequences; Determine the identical sequences and different sequences between the base sequence corresponding to each central node and the base sequence corresponding to the node connected to the corresponding central node; Encode the identical sequences and different sequences according to a preset encoding method to obtain the compressed data of the target data set.

2. The method according to claim 1, wherein The preset clustering algorithm is a hash-based clustering algorithm; The clustering the at least two base sequences by using a preset clustering algorithm to obtain at least one cluster includes: According to a preset first length, divide each base sequence in the at least two base sequences into at least one first short sequence; Determine that any two base sequences in the at least two base sequences with the number of identical first short sequences greater than a preset first number belong to the same cluster.

3. The method according to claim 1, characterized in that The constructing a similarity graph corresponding to each of the clusters includes: According to a preset second length, divide each base sequence in each of the clusters into at least one second short sequence; Determine the number of identical second short sequences between a first base sequence in each cluster and a second base sequence in the corresponding cluster; the first base sequence refers to each base sequence in the cluster, and the second base sequence refers to other base sequences in the corresponding cluster except the corresponding first base sequence; Determine that, in each cluster, the first base sequence with the number of identical second short sequences greater than a preset second number is a central node, the corresponding second base sequence is a node, and the number of identical second short sequences between the first base sequence and the corresponding second base sequence is an edge, and construct a similarity graph corresponding to each cluster.

4. The method according to claim 3, wherein Before the determining the identical sequences and different sequences between the base sequence corresponding to each central node and the base sequence corresponding to the node connected to the corresponding central node, the method further includes: Sort the edges between each central node and the node connected to the corresponding central node according to the number of identical second short sequences; Delete the edges with a lower ranking corresponding to each central node according to the preset number of edges of the central node.

5. The method according to claim 4, wherein The determining the identical sequences and different sequences between the base sequence corresponding to each central node and the base sequence corresponding to the node connected to the corresponding central node includes: Determine that, in the base sequence corresponding to each central node and the base sequence corresponding to the node connected to the corresponding central node, the identical second short sequences are identical sequences, and the sequences other than the second short sequences are different sequences.

6. The method according to claim 1, wherein The preset encoding method includes dictionary encoding, transform encoding, or entropy encoding.

7. The method according to claim 1, characterized in that, The obtaining of the target data set includes: Obtaining a data set, where the data set includes at least two of the base sequences; Dividing the data set into at least one of the target data sets according to the data volume or the number of the base sequences.

8. A device, characterized in that, It includes: An obtaining unit, configured to obtain a target data set, where the target data set includes at least two base sequences; A clustering unit, configured to cluster the at least two base sequences by using a preset clustering algorithm to obtain at least one clustering cluster; wherein, each of the at least one clustering clusters includes at least one base sequence; A similarity graph construction unit, configured to construct a similarity graph corresponding to each of the clustering clusters; wherein, the nodes in the similarity graph are the base sequences in the clustering cluster, the central node in the similarity graph is at least one of the nodes, and the base sequence corresponding to each central node in the similarity graph is a reference sequence, and the edges between the nodes in the similarity graph are used to represent the similarity of the corresponding base sequences; A determination unit, configured to determine the identical sequences and different sequences between the base sequence corresponding to each central node and the base sequence corresponding to the node connected to the corresponding central node; An encoding unit, configured to encode the identical sequences and different sequences according to a preset encoding method to obtain compressed data of the target data set.

9. An electronic device, characterized in that, It includes: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1 to 7.

10. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to execute the method according to any one of claims 1 to 7.