Sequencing data file compression method and device and storage medium
By using multiple matching components on the reference sequence to coordinate the reference fragments and comparing information classification encoding and compression, the problem of low compression rate of Fastq files is solved, and more efficient storage and transmission is achieved.
Patent Information
- Application Number
- CN202410087007.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-22
- Publication Date
- 2025-07-22
AI Technical Summary
The existing Fastq file compression scheme has low compression rate, resulting in high storage and transmission costs.
At least two matching components work together on the reference sequence, determine the reference fragment for the sequencing sequence, construct the alignment information between the sequencing sequence and the reference fragment, and perform classification encoding and compression according to the comparison information content type.
Improves the compression rate of sequencing data files and reduces storage and transmission costs.
Smart Images

Figure CN120356531A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data processing, and particularly to a method, device, and storage medium for compressing sequencing data files. Background Art
[0002] The Fastq format is a commonly used format for storing raw sequencing data in the bioinformatics field. A Fastq file stores biological sequences (usually nucleic acid sequences) and corresponding quality evaluation value sequences.
[0003] As the number of Fastq files generated in the bioinformatics field increases, new challenges are brought to the storage cost and transmission cost of Fastq files. Compressing Fastq files can greatly reduce the file size, thereby saving storage and transmission costs. However, the compression ratios achievable by current compression schemes for Fastq files are generally not high, resulting in poor compression effects for Fastq files.
[0004] Therefore, there is an urgent need for a solution that can provide a higher compression ratio for Fastq files. Summary of the Invention
[0005] Multiple aspects of this application provide a method, device, and storage medium for compressing sequencing data files to improve the compression ratio of sequencing data files.
[0006] An embodiment of this application provides a method for compressing a sequencing data file, including:
[0007] For any sequencing sequence in the file to be compressed, use at least two matching components to perform a matching operation on a reference sequence in cooperation to determine a reference fragment for the sequencing sequence;
[0008] Construct alignment information between the sequencing sequence and the reference fragment;
[0009] Based on the alignment information generated under each sequencing sequence in the file to be compressed, perform compression.
[0010] Another embodiment of this application provides a method for compressing a sequencing data file, including:
[0011] Based on the reference fragments matched by each sequencing sequence included in the file to be compressed on a reference sequence, respectively construct alignment information corresponding to each sequencing sequence, and the alignment information contains various types of information content;
[0012] Classify the information content in the alignment information generated under each sequencing sequence included in the file to be compressed according to the content type;
[0013] According to the correspondence between the content type and the coding rule, the information content under different content types is encoded respectively using the corresponding coding rule;
[0014] The results generated after encoding for different content types are compressed respectively to obtain the compression result of the sequencing sequence in the file to be compressed.
[0015] An embodiment of the present application further provides a computing device, including a memory and a processor;
[0016] The memory is used to store one or more computer instructions;
[0017] The processor is coupled with the memory and is used to execute the one or more computer instructions to execute the foregoing compression method of the sequencing data file.
[0018] An embodiment of the present application further provides a computer-readable storage medium storing computer instructions, which, when executed by one or more processors, cause the one or more processors to execute the foregoing compression method of the sequencing data file.
[0019] An embodiment of the present application further provides a computer program product, including a computer program / instructions, wherein when the computer program is executed by a processor, the processor is caused to implement the foregoing compression method of the sequencing data file
[0020] In an embodiment of the present application, a compression scheme for a sequencing data file is proposed. For each sequencing sequence included in the sequencing data file, it is proposed to use at least two matching components to cooperate to determine a reference fragment for the same sequencing sequence on a reference sequence; based on the determined reference fragment, the alignment information between the sequencing sequence and the reference fragment can be constructed, and compression is performed on each piece of alignment information constructed under the file to be compressed. In this way, by the cooperation of at least two matching components, not only can the efficiency of searching for reference fragments for sequencing sequences be improved, but also the matching results between the matching components can be mutually verified, which can effectively improve the matching accuracy of the reference fragments, thereby improving the compression rate for sequencing sequences, and further improving the compression rate of the sequencing data file. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The exemplary embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation to the present application. In the drawings:
[0022] Figure 1 is a schematic flowchart of a compression method for a sequencing data file provided by an exemplary embodiment of the present application;
[0023] Figure 2aSchematic diagram of the internal structure of a Fastq file provided by an exemplary embodiment of the present application;
[0024] Figure 2b Logical schematic diagram of a content segmentation method for a Fastq file provided by an exemplary embodiment of the present application;
[0025] Figure 3 Logical schematic diagram of an exemplary solution for extracting subsequences provided by an exemplary embodiment of the present application;
[0026] Figure 4 Flow chart of another compression method for sequencing data files provided by an exemplary embodiment of the present application;
[0027] Figure 5 Flow chart of yet another compression method for sequencing data files provided by an exemplary embodiment of the present application;
[0028] Figure 6 Flow chart of a compression method for sequencing data files provided by another exemplary embodiment of the present application;
[0029] Figure 7 Schematic diagram of the structure of a computing device provided by yet another exemplary embodiment of the present application. Detailed implementation manners
[0030] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Apparently, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0031] Before starting to elaborate on the technical solutions provided by the embodiments of the present application, several technical concepts involved in the present application are briefly explained as follows.
[0032] The sequencing data file can be understood as the sequencing result file output after using a sequencing instrument to perform DNA sequencing on a sample. The sequencing data file contains at least the sequencing sequences obtained by sequencing (which can also be called biological sequences, usually base sequences, or amino acid sequences).
[0033] A Fastq file is a sequencing data file in a typical format, used to store sequencing sequences and corresponding quality assessment sequences. A Fastq file contains multiple reads. Each read generally consists of four lines. The first line starts with '@' followed by the description information of the biological sequence. The second line is the biological sequence. The third line starts with '+' and can also be followed by the description information of the biological sequence. The fourth line is the quality assessment corresponding to the biological sequence in the second line (quality values, the quality assessment of sequencing). The number of elements in the quality assessment sequence in the fourth line is the same as the number of elements in the biological sequence in the second line.
[0034] A reference sequence refers to a reference genomic sequence used for comparison and analysis. It is a known genomic sequence obtained through sequencing and assembly, usually a representative sample of the genome of a certain species. Reference sequences generally come from public databases, which provide a large number of known genomic sequences for researchers to use in gene sequencing and bioinformatics research. At the same time, with the development of technology and the progress of new sequencing projects, reference sequences are constantly updated and improved to meet the needs of more species and more accurate research.
[0035] As mentioned in the background art, with the continuous progress of gene sequencing technology, the types of sequencing data files generated are increasing, bringing new challenges to storage and transmission. The inventors found in the research process that there are already some solutions for compressing sequencing data files in the current field, but the compression ratios that these solutions can achieve are generally not high, resulting in very little cost savings in storage and transmission for the compressed files.
[0036] For this reason, in this embodiment, a compression method for sequencing data files is proposed to improve the compression ratio of sequencing data files.
[0037] The following will detail the technical solutions provided by each embodiment of the present application in conjunction with the accompanying drawings.
[0038] Figure 1 It is a schematic flowchart of a compression method for a sequencing data file provided by an exemplary embodiment of the present application. This method can be executed by a file compression device, which can be implemented as software, hardware, or a combination of software and hardware, and can be integrated in a computing device. Refer to Figure 1 , this method may include:
[0039] Step 100: For any sequencing sequence in the file to be compressed, use at least two matching components to perform a matching operation on the reference sequence in cooperation to determine a reference fragment for the sequencing sequence;
[0040] Step 101: Construct alignment information between the sequencing sequence and the reference fragment;
[0041] Step 102: Perform compression based on the alignment information generated for each sequencing sequence in the file to be compressed.
[0042] In this embodiment, the sequencing data file to be compressed is described as the file to be compressed. It should be noted that in this embodiment, the file format of the file to be compressed is not limited. For example, the file to be compressed can be the Fastq file mentioned above, or other format files that can be used to store sequencing sequences currently or in the future. No more examples of file formats are given here. In this embodiment, the file to be compressed may at least contain sequencing sequences, and usually contains multiple sequencing sequences. In this embodiment, other file contents included in the file to be compressed are not limited. For different file formats, the other file contents included in the file to be compressed may not be exactly the same. For example, in the Fastq file mentioned above, in addition to containing sequencing sequences, it also contains quality evaluation sequences and identification information of the sequencing sequences, etc. No more examples of other file contents included in other format files are given.
[0043] The compression method for the sequencing data file provided in this embodiment can split the content of the file to be compressed to obtain each sequencing sequence included in the file to be compressed.
[0044] Figure 2a It is a schematic diagram of the internal structure of a Fastq file provided by an exemplary embodiment of the present application. Figure 2b It is a logical schematic diagram of a content splitting method for a Fastq file provided by an exemplary embodiment of the present application. Refer to Figure 2a , in this case, the file to be compressed contains multiple reads. Figure 2a Shown in Figure 2b are the 4 lines of content included in one read. Among them, the second line of the read is the sequencing sequence, the fourth line of the read is the quality evaluation sequence, and the first and third lines of the read are the identification information of the sequencing sequence. Based on this, refer to
[0045] On this basis, an improved solution is provided for compressing the sequencing sequences in the file to be compressed in this embodiment. Based on the improved solution provided in this embodiment, the sequencing sequence part cut out from the file to be compressed can be compressed separately to improve the compression rate of the sequencing sequence part, and then improve the compression rate of the file to be compressed. In this embodiment, it is further proposed that an improved solution can also be provided for compressing other parts cut out from the file to be compressed, which will be described in detail later.
[0046] Reference Figure 2b , in this embodiment, different file parts cut out from the file to be compressed can be compressed separately, and by improving the compression rate of a single file part, the overall compression rate of the file to be compressed can be improved. Among them, after the compression results generated by separately compressing each file part are combined, a compressed file corresponding to the file to be compressed can be generated. The compressed file will be used as the object for storage and transmission.
[0047] The following will detail each step in the file method provided in this embodiment. In this embodiment, the compression solution for the sequencing sequence part can at least include two links: the first is a preprocessing link with a single sequencing sequence as the object, and the second is an execution compression link with each sequencing sequence as the object. The following will detail the two links separately.
[0048] Regarding the first link, considering that the preprocessing logic implemented for each sequencing sequence is the same, for the convenience of description, any sequencing sequence in the file to be compressed will be taken as an example to detail the first link.
[0049] Reference Figure 1 , in step 100, for any sequencing sequence in the file to be compressed, at least two matching components are used to perform a matching operation on the reference sequence in cooperation to determine a reference fragment for the sequencing sequence. The matching component can be understood as a logical component with the function of matching a similar fragment for the sequencing sequence on the reference sequence. In step 100, at least two matching components can work together to determine a reference fragment for the same sequencing sequence on the reference sequence. The cooperation in this embodiment can be understood as a reasonable division of labor, mutual cooperation, and mutual verification among at least two matching components to determine a reference fragment for the same sequencing sequence by synthesizing the matching results of each matching component. As mentioned above, each matching component has the ability to match a similar fragment for the sequencing sequence on the reference sequence. Therefore, in step 100, the matching results obtained by each matching component for the same sequencing sequence can be mutually verified to finally determine the reference fragment corresponding to the sequencing sequence.
[0050] In a preferred implementation: At least two matching components can be used to match similar segments for the sequencing sequence on the reference sequence respectively based on subsequences extracted from the sequencing sequence. An exemplary logic for matching similar segments can be: Using the subsequence as a seed, search on the reference sequence to see if there is a segment identical to the subsequence; if there is, map the relative position relationship between the subsequence and the sequencing sequence to the reference sequence to locate the similar segment for the sequencing sequence. It should be understood that this is only exemplary, and the matching logic in the matcher component in this embodiment is not limited to this, and no more examples are given here.
[0051] The subsequence in this embodiment can be understood as a subset of the sequencing sequence. In this embodiment, it is proposed that the subsequences used in different matching components do not overlap with each other, which can effectively avoid the problem of repeatedly performing matching operations on the subsequence, so as to avoid affecting the matching efficiency. This embodiment provides an exemplary subsequence extraction scheme: Sliding window tools can be respectively set in at least two matching components. Based on this, non-overlapping sequence segments can be allocated to different sliding window tools on the sequencing sequence; control each sliding window tool to slide on the sequence segment allocated to itself to extract subsequences that meet the target requirements.
[0052] Figure 3 It is a schematic diagram of the logic of an exemplary scheme for extracting subsequences provided for an exemplary embodiment of the present application. Refer to Figure 3 , for example, two matching components can be used to match similar segments for the sequencing sequence, and the sequencing sequence can be divided into two sequence segments and allocated to the sliding window tools in the two matching components ( Figure 3 The two boxes shown in are the sliding windows). The starting positions and sliding directions of the two sliding window tools can also be configured. Based on this, the two sliding window tools can be controlled to slide on their respective sequence segments. It should be understood that the width of the sliding window should be the same as the length of the subsequence. Refer to Figure 3 , the two sliding window tools can slide from both ends of the sequencing sequence towards the middle, and what the sliding window hits is the subsequence.
[0053] In this exemplary subsequence extraction scheme, the aforementioned target requirements can be flexibly set according to actual needs. Exemplarily, the target requirements can include but are not limited to that the sequence elements at the ends of the subsequence are of a preset type, and / or the sequence elements at the ends of the subsequence are different from their adjacent sequence elements. Here, the sequence elements can be bases or amino acids, etc. In this way, the two matching components can independently extract the subsequences they use, and the two matching components will not extract the same subsequence.
[0054] In addition, in step 100, preferably, at least two matching components can run in parallel. From the perspective of a single matching component, after a similar segment is matched for the sequencing sequence on the reference sequence based on a certain subsequence extracted, it can pause to wait for the matching results in other matching components.
[0055] In the process of research, the inventors found that sequencing sequences generally can be divided into two categories: variant sequences and non-variant sequences. Among them, variant sequences can be understood as those with insertions or deletions of sequence elements relative to a certain segment on the reference sequence. That is to say, variant sequences are usually due to Insertion and Deletion (InDel) events, and InDel events are usually caused by processes such as gene mutations, chromosomal rearrangements, and gene recombinations. Correspondingly, non-variant sequences can be understood as sequencing sequences without InDel events.
[0056] In this embodiment, by using at least two matching components to match similar segments for the same sequencing sequence on the reference sequence, it is possible to efficiently identify whether the sequencing sequence belongs to a variant sequence or a non-variant sequence.
[0057] Continuing to refer to Figure 1 , as proposed in step 101, if at least two matching components match the same similar segment and the similar segment meets the preset similarity requirement, then the similar segment is determined as the reference segment corresponding to the sequencing sequence. In the process of research, the inventors found that for non-variant sequences, there are segments on the reference sequence that can be aligned with the non-variant sequences and have a high enough similarity, and such segments can be used as the reference segments for non-variant sequences. Among them, the reference segment can be understood as the segment on the reference sequence used as the compression reference for the sequencing sequence. Therefore, in this embodiment, for a sequencing sequence, if at least two matching components match the same similar segment and the similar segment meets the preset similarity requirement, then it can be determined that the sequencing sequence is a non-variant sequence. And the matched similar segment can be used as the reference segment corresponding to the sequencing sequence.
[0058] The inventors also found that in practical applications, based on the fact that most subsequences in non-variant sequences can match the final reference segment on the reference sequence, therefore, in this embodiment, at least two matching components generally only need to perform a small number of matching operations and then can match a similar segment and pause. In this way, in this embodiment, at least two matching components can be used to efficiently match a reference segment for non-variant sequences.
[0059] Moreover, since at least two matching components are used to perform matching operations for the same sequencing sequence in this embodiment, the matching results obtained by the at least two matching components can be mutually verified. As mentioned above, only when the at least two matching components match the same similar fragment and the similar fragment meets the preset similarity requirement, will the similar fragment be determined as the reference fragment. This can effectively improve the matching accuracy of the reference fragment and avoid matching an inappropriate reference sequence to a variant sequence.
[0060] In other words, in this embodiment, if it is determined that the sequencing sequence belongs to a non-variant sequence by using at least two matching components, the reference fragment corresponding to the sequencing sequence will be determined based on the matching results of the at least two matching components. If it is determined that the sequencing sequence belongs to a variant sequence by using at least two matching components, then the method will switch to other reference fragment determination methods more suitable for variant sequences to more accurately determine the reference fragment for the variant sequence.
[0061] In this embodiment, it is proposed that if it is detected that the at least two matching components match similar fragments at different positions, and / or the at least two matching components match the same similar fragment but the similarity between the similar fragment and the sequencing sequence does not meet the preset similarity requirement, then it is determined that the sequencing sequence is a variant sequence.
[0062] As mentioned above, for a variant sequence, insertions or deletions of sequence elements occur relative to the reference sequence. Therefore, there may be multiple subsequences on the variant sequence that can match similar fragments on the reference sequence. However, the positions of these similar fragments are usually different. For this reason, in this embodiment, when at least two matching components match different similar fragments for the same sequencing sequence, it can be determined that the sequencing sequence is a variant sequence. In addition, considering that the subsequences used in the at least two matching components may be on the same side of the inserted or deleted sequence elements in the variant sequence, in this case, if it is possible for the at least two matching components to match the same similar fragment for the variant sequence, therefore, in this embodiment, when the at least two matching components match the same similar fragment but the similarity between the similar fragment and the sequencing sequence does not meet the preset similarity requirement, it can also be determined that the sequencing sequence is a variant sequence.
[0063] Furthermore, in this embodiment, it is also proposed that after it is determined that the sequencing sequence is a variant sequence based on the above logic, the at least two matching components will no longer extract the remaining subsequences on the sequencing sequence. In this way, variant sequences can be identified more efficiently, and the at least two matching components do not need to traverse all the subsequences on the variant sequence.
[0064] Of course, in this embodiment, if it is detected that any matching component fails to match a similar segment for the sequencing sequence on the reference sequence after traversing all the subsequences it extracts, it can also be determined that the sequencing sequence is a variant sequence. Other matching components may no longer extract the remaining subsequences on the sequencing sequence.
[0065] In this way, in this embodiment, the sequencing sequences that may have mutated can be identified more comprehensively and efficiently. Thus, the determination method of the reference segment can be switched in a timely manner for such sequencing sequences, thereby ensuring the accuracy of the reference segments of such sequencing sequences.
[0066] Exemplarily, this embodiment proposes that local alignment technology can be used to determine reference segments for each variant sequence included in the file to be compressed on the reference sequence. Local alignment technology: It is to find a suitable matching region in two sequences for alignment to find a suitable matching scheme. Local alignment technology is usually used to align sequences with relatively low similarity. In this embodiment, the type of algorithm used in the local alignment technology is not limited. For example, the BWA (Burrows-Wheeler Aligner) algorithm or the Wave Front Alignment (WFA) algorithm can be used, etc. The principle of the local alignment technology will not be elaborated in detail here. Any algorithm in the local alignment technology can be used in this embodiment to determine the reference segment for the variant sequence.
[0067] Continuing to refer to Figure 1 , after determining the reference segment for the sequencing sequence, the alignment information between the sequencing sequence and the reference segment can be constructed. Among them, the alignment information is used to describe the differences between the sequencing sequence and the reference segment.
[0068] In this way, by implementing the above-mentioned preprocessing logic for each sequencing sequence included in the file to be compressed, the alignment information generated under each sequencing sequence in the file to be compressed can be obtained. As mentioned above, in this embodiment, the reference segment can be determined more accurately for the sequencing sequence, which can effectively reduce the amount of information in the alignment information generated by each sequencing sequence.
[0069] For the second link in the compression scheme for the sequencing sequence part, referring to Figure 1 , in step 103, compression can be performed based on the alignment information generated under each sequencing sequence in the file to be compressed. That is, each sequencing sequence in the file to be compressed will be converted into the corresponding alignment information. By performing compression on these alignment information, the compression of the sequencing sequence part cut out from the file to be compressed can be achieved. Since the reference segment can be determined more accurately in this embodiment, the amount of information in the alignment information can be effectively reduced, and thus an excellent compression ratio can be obtained in this second link.
[0070] In summary, in this embodiment, a compression scheme for sequencing data files is proposed. For each sequencing sequence included in the sequencing data file, it is proposed to use at least two matching components to cooperate to determine a reference segment for the same sequencing sequence on the reference sequence. Based on the determined reference segment, the alignment information between the sequencing sequence and the reference segment can be constructed, and compression is performed on each alignment information constructed under the file to be compressed. In this way, through the cooperation of at least two matching components, not only can the efficiency of searching for reference segments for sequencing sequences be improved, but also the matching results can be mutually verified between the matching components, effectively improving the matching accuracy of the reference segments, thereby increasing the compression rate for sequencing sequences and further improving the compression rate of the sequencing data file.
[0071] Figure 4 It is a schematic flowchart of another compression method for sequencing data files provided by an exemplary embodiment of the present application. Refer to Figure 4 , the method may include:
[0072] Step 401: For any sequencing sequence in the file to be compressed, use at least two matching components to perform a matching operation on the reference sequence in cooperation to determine a reference segment for the sequencing sequence;
[0073] Step 402: Construct the alignment information between the sequencing sequence and the reference segment;
[0074] Step 403: Classify the information content in the alignment information generated under each sequencing sequence included in the file to be compressed according to the content type;
[0075] Step 404: According to the correspondence between the content type and the coding rule, respectively use the corresponding coding rule to encode the information content under different content types;
[0076] Step 405: Compress the results generated after encoding under different content types respectively.
[0077] Among them, Steps 401 to 402 can refer to the relevant descriptions in the foregoing embodiments and will not be repeated here. In this embodiment, an optional implementation manner for compressing the alignment information generated in the file to be compressed is provided based on Steps 403 to 405. This optional implementation manner can be combined with other implementation manners in the above or below embodiments to generate various technical solutions.
[0078] Refer to Figure 4, in this embodiment, the alignment information corresponding to a single sequencing sequence may include various types of information content. The content types in this embodiment may include, but are not limited to, position type, strand attribute type, result type, etc. Among them, the position type information content can be used to describe the position of the reference segment corresponding to the sequencing sequence; the strand attribute type information content can be used to describe whether the sequencing sequence matches the positive strand or the negative strand in the reference sequence; the result type information content may include an array for marking the consistency of sequence elements and each sequence element marked as inconsistent.
[0079] In step 403, the information content in the alignment information generated under each sequencing sequence included in the file to be compressed can be classified according to the content type. It should be understood that since the file to be compressed contains multiple sequencing sequences, therefore, in this embodiment, the alignment information generated under the file to be compressed will also be multiple. Moreover, since there is usually an order relationship between the sequencing sequences in the file to be compressed, therefore, in this embodiment, it is proposed that the alignment information generated under the file to be compressed can inherit this order relationship. Based on this, in step 403, the information content in each alignment information can be divided into the corresponding content type, so that a group of information content arranged in order can be obtained under a single content type.
[0080] In a preferred solution, the alignment information under each sequencing sequence in the file to be compressed can be re - sorted according to the reference segment position; under the target content type, according to the order between the re - sorted alignment information, the information content belonging to the target content type is queried from the alignment information corresponding to each sequencing sequence in turn to determine the information content under the target content type; where the target content type is any one of the content types. That is, in this exemplary solution, the sequencing sequences in the file to be compressed are re - sorted, so that the order relationship generated between the re - sorted alignment information will be inherited under each content type. Based on this preferred solution, the reference segment positions in each alignment information can be adjacent to each other in the reference sequence, so that the result generated by performing encoding under the information class will be more conducive to compression.
[0081] Of course, this is only preferred. In this embodiment, re - sorting may not be done either. Whether re - sorted or not, the order relationship between the alignment information will be inherited under each content type according to the actual arrangement state.
[0082] Reference Figure 4 , after the classification of the information content is completed, it is proposed in step 404 that customized encoding rules are provided for different content types to obtain a better encoding effect under each content type. Based on this, according to the correspondence between the content type and the encoding rule, the information content under different content types can be encoded using the corresponding encoding rules respectively.
[0083] During the research process, the inventors found that the commonly used sequencing technologies in the current field include single - end sequencing technology and paired - end sequencing technology. Among them, single - end sequencing technology can be understood as sequencing unidirectionally from one end to the other end of a gene fragment. If sequencing is performed in two directions from both ends to the other ends respectively, it is called paired - end sequencing. Single - end sequencing technology generates one sequencing sequence for each gene fragment, while paired - end sequencing generates two sequencing sequences in opposite directions for each gene fragment. From the perspective of the sequencing data file, single - end sequencing technology will generate one sequencing data file for the same sample, while paired - end sequencing technology will generate a pair of sequencing data files for the same sample. For ease of description, in this embodiment, the two sequencing sequences generated for a single fragment in the paired - end sequencing file are collectively referred to as a sequencing data segment. It can be understood that a single sequencing data segment contains a pair of sequencing sequences, which are stored one - to - one in the aforementioned pair of sequencing data files. In addition, it is worth noting that the sequencing sequences between the pair of sequencing data files generated under paired - end sequencing technology are aligned. Under the aforementioned re - sorting scheme, it can be understood that it is the sequencing data segments that are re - sorted, and the sequencing sequences in the pair of sequencing data files will change their positions synchronously with the change of the position of the sequencing data segments.
[0084] Therefore, in this embodiment, it is proposed that for the files to be compressed generated by single - end sequencing technology and paired - end sequencing technology under some content types, different encoding rules can be adopted for information content encoding.
[0085] The following will separately expand on the exemplary encoding processes for different content types.
[0086] Position class information content
[0087] If the file to be compressed is a sequencing data file generated by single - end sequencing technology, an exemplary encoding scheme for position - type information content can be as follows: encode each position - type information content into a first marker array, where the elements in the first marker array correspond one - to - one to each sequencing sequence; successively determine for each sequencing sequence whether the distance difference between the reference fragment position and the reference fragment position of the adjacent preceding sequencing sequence exceeds a preset distance threshold; record the position - type information content corresponding to the sequencing sequences determined to exceed the preset distance threshold successively in the first position array, and mark the corresponding elements in the first marker array as values for pointing to the first position array; record the distance differences corresponding to the sequencing sequences determined not to exceed the preset distance threshold successively in the second position array, and mark the corresponding elements in the first marker array as values for pointing to the second position array. Among them, the element values in the first marker array can be 1 or 0.
[0088] For example, if the file to be compressed contains 5 sequencing sequences, after being reordered according to the reference fragment positions mentioned above, the 5 position-related information contents are: coordinate A, coordinate B, coordinate C, coordinate D, and coordinate E in sequence. In this embodiment, instead of directly compressing these coordinates, they are encoded. For example, the distance difference between coordinate B and coordinate A is 3, the distance difference between coordinate C and coordinate B is 7, the distance difference between coordinate D and coordinate C is 2, and the distance difference between coordinate E and coordinate D is 1. If the preset distance threshold is 5, the encoded first marker array will be [1, 0, 1, 0, 0], where the value 1 in the first marker array points to the first position array, and the value 0 points to the second position array. The encoded first position array will be [coordinate A, coordinate C], and the encoded second position array will be [3, 2, 1]. It can be seen that after encoding, only two original position-related information contents - coordinate A and coordinate C - are left, and other position-related information has been encoded into an array that is more conducive to compression.
[0089] That is to say, if the file to be compressed is one of the two sequencing data files generated for the same sample based on the paired-end sequencing technology, and the other sequencing data file has completed the encoding of the position-related information content according to the encoding rules set for the single-end sequencing technology, then for the file to be compressed, an exemplary encoding scheme for the position-related information content can be: after encoding each position-related information content into a first marker array, successively determine whether the position difference value between the reference fragment position and the reference fragment position of the corresponding sequencing sequence in the other sequencing data file exceeds the preset difference threshold for each sequencing sequence; record the position-related information content corresponding to each sequencing sequence determined to exceed the preset difference threshold in the first position array in sequence, and mark the corresponding element in the first marker array with the value pointing to the first position array; record the position difference values corresponding to each sequencing sequence determined not to exceed the preset difference threshold in the second position array in sequence, and mark the corresponding element in the first marker array with the value pointing to the second position array.
[0090] That is, for one of the two sequencing data files generated by the paired-end sequencing, the encoding scheme provided for the single-end testing technology can be used to independently complete the encoding of the position-related information content. For the other sequencing data file, the position-related information content is encoded according to the exemplary encoding scheme provided here. In this case, first, each position-related information content is encoded into a first marker array. Then, instead of comparing the reference fragment positions between the sequencing sequences within the file to be compressed, the reference fragment positions of each sequencing sequence in the file to be compressed are respectively compared with the corresponding sequencing sequences in the other paired sequencing data file. Then, the first position array and the second position array are encoded according to the comparison results, and the element values in the first marker array are determined.
[0091] Continuing with the above example, if the file to be compressed is paired with sequencing data file 1, then the coordinates A, B, C, D, and E under the file to be compressed will be compared with the coordinates of the reference fragments of the corresponding sequencing sequences in sequencing data file 1: coordinates A', B', C', D', and E', respectively. The position difference value between coordinate A and coordinate A' is 1, the position difference value between coordinate B and coordinate B' is 3, the position difference value between coordinate C and coordinate C' is 2, the position difference value between coordinate D and coordinate D' is 0, and the position difference value between coordinate E and coordinate E' is 1. If the preset difference threshold is 4, then the first marker array encoded will be [0, 0, 0, 0, 0]. Among them, the value 1 in the first marker array points to the first position array, and the value 0 points to the second position array. And the encoded first position array will be empty, and the encoded second position array will be [1, 3, 2, 0, 1]. It can be seen that after encoding, no original position type information content is left, and it has all been encoded into arrays that are more conducive to compression.
[0092] It should be understood that the encoding scheme for the position type information content provided under the paired-end sequencing technology above is exemplary, and this embodiment is not limited to this. Under the paired-end sequencing technology, it is also possible not to perform reference fragment position comparison between paired sequencing data files, but to perform independent comparison within a single sequencing data file without interference.
[0093] In this way, in this embodiment, for the file to be compressed, the position type information content will be encoded into three arrays: the first marker array, the first position array, and the second position array. These arrays are more conducive to compression compared to the position type information content itself, and a higher compression ratio can be obtained.
[0094] Chain attribute class information content
[0095] If the file to be compressed is a sequencing data file generated based on single-end sequencing technology, then an exemplary encoding method for the strand attribute type information content can be: encoding each strand attribute type information content into a second marker array, and the elements in the second marker array correspond to each sequencing sequence one by one; marking the corresponding element in the second marker array as the first value for the sequencing sequence that matches the positive strand; marking the corresponding element in the second marker array as the second value for the sequencing sequence that matches the negative strand. Among them, the element values in the second marker array can be 1 or 0. It is only required that the first value and the second value are different, and it is not limited whether the first value is 1 or 0.
[0096] For example, if the file to be compressed contains 5 sequencing sequences, and the 5 chain attribute information contents are: forward strand, forward strand, forward strand, reverse strand, reverse strand. In this embodiment, these chain attribute information contents are not directly compressed, but encoded into a second marker array which will be [1, 1, 1, 0, 0]. Among them, the element value 1 in the second marker array represents the forward strand, and the element value 0 represents the reverse strand.
[0097] If the file to be compressed is any one of the two sequencing data files generated for the same sample based on the paired-end sequencing technology, an exemplary encoding method for the chain attribute information content can be: configure a third marker array under the chain attribute class, and the elements in the third marker array correspond to each sequencing data segment one by one; for the first type of sequencing data segment containing sequencing sequences with different chain attributes, perform an exchange operation on the chain attribute information content as needed, so as to concentrate the same chain attribute information content under the same sequencing data file, and mark the elements corresponding to each first type of sequencing data segment in the third marker array as the target value; after the exchange operation is completed, perform the operation of encoding each chain attribute information content into a second marker array and element marking for the file to be compressed.
[0098] That is to say, in the case of paired-end sequencing technology, the sequencing sequence exchange operation can be first performed between the paired sequencing data files to try to concentrate the sequencing sequences with the same chain attribute into the same sequencing data file. After the exchange operation is completed, the two sequencing data files can then independently implement the foregoing encoding process for the second marker array.
[0099] Continuing with the above example, the 5 chain attribute information contents under the file to be compressed are: forward strand, forward strand, forward strand, reverse strand, reverse strand. If the 5 chain attribute information contents under the paired sequencing data file 1 of the file to be compressed are: reverse strand, reverse strand, forward strand, forward strand, forward strand, then the target chain attribute can be first determined for the file to be compressed. For example, the target chain attribute is set to the forward strand. Then, it can be detected that the 4th and 5th sequencing data segments need to perform the exchange operation. After performing the exchange operation accordingly, the 5 chain attribute information contents under the file to be compressed will be: forward strand, forward strand, forward strand, forward strand, forward strand; and the 5 chain attribute information contents under the sequencing data file 1 will be: reverse strand, reverse strand, forward strand, reverse strand, reverse strand. Correspondingly, the second marker array encoded for the file to be compressed will be [1, 1, 1, 1, 1], and the second marker array encoded for the sequencing data file 1 will be [0, 0, 1, 0, 0]. And the third marker array configured under the chain attribute class can be encoded as [0, 0, 0, 1, 1], where the element value 1 in the third array represents that the exchange operation has occurred, and the element value 0 represents that the exchange operation has not occurred.
[0100] It can be understood that after the foregoing interaction operations, the same chain attribute class information content is concentrated under the same sequencing data file. This makes it such that in the second marker arrays encoded under two paired sequencing data files, the element values will basically be the same. And the more elements with the same value in a single array, the higher the compression rate that can be obtained. Therefore, this exemplary encoding scheme can effectively improve the compression rate of the chain attribute class information content.
[0101] In this way, in this embodiment, for the file to be compressed, the chain attribute class information content will be encoded into 1 array: the second marker array. In the paired-end sequencing technology, a third marker array can also be provided for the two paired sequencing data files. These arrays are more conducive to compression compared to the chain attribute class information content itself and can obtain a higher compression rate.
[0102] Result class information content
[0103] As mentioned above, the result class information content in this embodiment may include an array for marking the consistency of sequence elements and each sequence element marked as inconsistent. Among them, the array for marking the consistency of sequence elements is already in the form of an array and can be retained in this embodiment. For each sequence element marked as inconsistent, this embodiment proposes that in the result class information content corresponding to a single sequencing sequence in the file to be compressed, each sequence element marked as inconsistent can be stored in the result array in order. That is, a result array is used to carry each sequence element marked as inconsistent.
[0104] So far, in this embodiment, the corresponding information content can be encoded into an array under different content types, and through the designed encoding rules, the encoded arrays can obtain a higher compression rate.
[0105] On this basis, continuing to refer to Figure 4 , in step 405, the results generated after encoding under different content types can be compressed respectively, which can further improve the compression rate.
[0106] In summary, in this embodiment, it is proposed to classify the information content in each comparison information generated under the file to be compressed according to the content type, so as to separate different types of information content. Moreover, a clever encoding rule is proposed for different content types, so that the information content under different content types can be encoded into an array that is more conducive to compression. In this way, the results generated after encoding under different content types can be compressed respectively, thereby effectively improving the compression rate of the comparison information.
[0107] Figure 5 It is a schematic flowchart of another compression method for a sequencing data file provided by an exemplary embodiment of the present application. Refer to Figure 5, the method may include:
[0108] Step 500, for any quality evaluation sequence in the file to be compressed, check whether there is a target evaluation value whose proportion exceeds a preset proportion threshold;
[0109] Step 501, if there is, after filtering the quality evaluation sequence based on the target evaluation value, perform an encoding operation;
[0110] Step 502, perform compression based on each encoded quality evaluation sequence in the file to be compressed.
[0111] In this embodiment, an improved solution for compressing the quality evaluation sequence part in the file to be compressed is provided. This improved solution can be combined with the improved solution for compressing the sequencing sequence part in the foregoing embodiment to further improve the compression ratio of the file to be compressed.
[0112] Reference Figure 5 , in step 500, in each quality evaluation sequence included in the file to be compressed, check whether there is a target evaluation value whose proportion exceeds a preset proportion threshold. Among them, the preset proportion threshold can be flexibly set as needed. For example, it can be set to 50% etc., and is not limited here.
[0113] For any quality evaluation sequence, if there is a target evaluation value therein, in step 501, after filtering the quality evaluation sequence based on the target evaluation value, perform an encoding operation. In this way, the duplicate content in the quality evaluation sequence can be removed, thereby reducing the amount of data to be compressed.
[0114] An exemplary encoding scheme can be:
[0115] Encode the quality evaluation sequence into a fourth marker array, where the elements in the fourth marker array correspond one-to-one to the elements in the quality evaluation sequence. Among them, the elements in the quality evaluation sequence that are different from the target evaluation value are encoded as a third value in the fourth marker array, and the elements that are the same as the target evaluation value are encoded as a fourth value in the fourth marker array;
[0116] Record the original element values of the elements in the quality evaluation sequence that are different from the target evaluation value in the first evaluation value array pointed to by the third value in order.
[0117] Of course, the target evaluation value can also be recorded in the second evaluation value array. Obviously, the second evaluation value array can only contain 1 element, and the element value is the target evaluation value.
[0118] In this exemplary coding scheme, it is proposed that each quality evaluation sequence included in the file to be compressed be encoded into a fourth marker array. Moreover, only two element values are included in the fourth marker array, which effectively ensures that the fourth marker array can achieve an excellent compression rate. Other evaluation values in the quality evaluation sequence that are different from the target evaluation value will be stored in the first evaluation value array. The inventor found during the research process that the proportion of the remaining evaluation values in the quality evaluation sequence that are different from the target evaluation value is usually relatively low. Therefore, the number of elements in the first evaluation value array will not be too large, usually much smaller than the number of elements in the quality evaluation sequence. In this way, for the quality evaluation sequence part divided from the file to be compressed, the amount of data to be compressed will be effectively reduced, and the array form is also more conducive to obtaining a higher compression rate.
[0119] On this basis, continue to refer to Figure 5 , in step 502, compression can be performed based on each encoded quality evaluation sequence in the file to be compressed. As mentioned above, compared with the quality evaluation sequence, the encoded result obtained after encoding has a smaller amount of data. Moreover, the above-mentioned exemplary coding scheme proposes to encode the quality evaluation sequence into an array form, which is also more convenient for compression, thereby improving the compression rate for the quality evaluation sequence part cut out from the file to be compressed.
[0120] In addition, referring to Figure 2b , it is further proposed in this embodiment that if the file to be compressed contains the identification information corresponding to each sequencing sequence, the identification information is deleted from the file to be compressed. That is, there is no need to compress the identification information part cut out from the file to be compressed, which can further reduce the amount of data to be compressed.
[0121] In this way, in this embodiment, referring to Figure 2, the compression results obtained by performing compression on the quality evaluation sequence and the sequencing sequence respectively can be combined to obtain the compressed file corresponding to the file to be compressed. As mentioned above, in this embodiment, the compression rates for the quality evaluation sequence and the sequencing sequence can be effectively improved. Therefore, the improvement of the compression rates in both aspects, plus the compression rate saved by deleting the identification information, can significantly improve the overall compression rate of the file to be compressed.
[0122] Considering that the identification information contained in the file to be compressed is deleted during the compression process in this embodiment, it will cause the identification information to be irrecoverable after decompressing the compressed file subsequently. For this reason, this embodiment further proposes that after decompressing the compressed file, identification information can be regenerated for each of the decompressed sequencing sequences to distinguish different sequencing sequences. The inventor found during the research process that the identification information in the file to be compressed is usually used to distinguish different sequencing sequences and does not refer to information such as the order between sequencing sequences. Therefore, the method of regenerating identification information will not affect the use of the sequencing data generated after decompression.
[0123] In summary, in this embodiment, a compression improvement scheme for several other exemplary file parts segmented from the file to be compressed is proposed. By deleting the identification information in the file to be compressed and no longer compressing it, a certain compression rate can be improved. Also, by filtering and encoding the quality evaluation sequence part, the data volume of the result generated after encoding can be made lower than the data volume of the quality evaluation sequence itself. Moreover, the result generated after encoding adopts an array form, which can further improve the compression rate for the quality evaluation sequence part. In this way, the compression rate of each file part segmented from the file to be compressed can be improved, so that the overall compression rate of the file to be compressed can be significantly improved.
[0124] Figure 6 It is a schematic flowchart of a compression method for a sequencing data file provided in another exemplary embodiment of this application. Refer to Figure 6 , the method may include:
[0125] Step 600: Based on the reference segments matched by each sequencing sequence included in the file to be compressed on the reference sequence, construct the alignment information corresponding to each sequencing sequence, and the alignment information includes various types of information content;
[0126] Step 601: Classify the information content in the alignment information generated under each sequencing sequence included in the file to be compressed according to the content type;
[0127] Step 602: According to the correspondence between the content type and the encoding rule, encode the information content under different content types respectively using the corresponding encoding rules;
[0128] Step 603: Compress the results generated after encoding under different content types respectively to obtain the compression result of the sequencing sequence in the file to be compressed.
[0129] Among them, the implementation method used to construct the alignment information corresponding to each sequencing sequence is not limited in this embodiment. An exemplary method may be: for any sequencing sequence in the file to be compressed, at least two matching components are used to match similar segments for the sequencing sequence on the reference sequence respectively based on the subsequences extracted from the sequencing sequence, and the subsequences used by different matching components do not repeat each other; if at least two matching components match the same similar segment and the similar segment meets the preset similarity requirement, the similar segment is determined as the reference segment corresponding to the sequencing sequence; the alignment information between the sequencing sequence and the reference segment is constructed. Regarding this implementation method, reference can be made to the relevant descriptions in the foregoing embodiments and will not be elaborated here.
[0130] In this embodiment, other implementation methods can be supported to construct the alignment information corresponding to each sequencing sequence.
[0131] Regardless of the implementation method used to construct the alignment information corresponding to each sequencing sequence, in this embodiment, an improved scheme for performing the compression process on the alignment information generated under each sequencing sequence in the file to be compressed can be provided to improve the compression ratio of the file to be compressed. Regarding the technical details in this improved scheme, reference can also be made to the relevant descriptions in the foregoing embodiments and will not be elaborated here.
[0132] In this embodiment, by improving the compression process of the alignment information generated under each sequencing sequence in the file to be compressed, the compression ratio of the alignment information can be effectively increased, and then the compression ratio of the file to be compressed can be improved.
[0133] It should be noted that in some of the processes described in the above embodiments and the accompanying drawings, a plurality of operations appear in a specific order, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or may be executed in parallel. The operation numbers such as 101 and 102 are only used to distinguish different operations, and the numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different arrays, values, etc., do not represent the order of precedence, and do not limit that "first" and "second" are of different types.
[0134] Figure 7 It is a schematic structural diagram of a computing device provided for another exemplary embodiment of the present application. As Figure 7 shown, the computing device includes: a memory 70, a processor 71, and a communication component 72.
[0135] The processor 71 is coupled to the memory 70 and the communication component 72 and is used to execute the computer program in the memory 70 for:
[0136] For any sequencing sequence in the file to be compressed, at least two matching components are used to perform a matching operation on the reference sequence in cooperation to determine a reference segment for the sequencing sequence;
[0137] Construct alignment information between the sequencing sequence and the reference segment;
[0138] Based on the alignment information generated for each sequencing sequence in the file to be compressed, perform compression.
[0139] In an alternative embodiment, when the processor 71 uses at least two matching components to perform a matching operation on the reference sequence in cooperation to determine a reference segment for any sequencing sequence in the file to be compressed, it can be specifically used for:
[0140] Use the at least two matching components to respectively match similar segments for the sequencing sequence on the reference sequence based on the subsequences extracted from the sequencing sequence, and the subsequences used by different matching components are non-repetitive;
[0141] If the at least two matching components match the same similar segment and the similar segment meets the preset similarity requirement, then determine the similar segment as the reference segment corresponding to the sequencing sequence.
[0142] In an alternative embodiment, the processor 71 can also be used for:
[0143] If it is detected that the at least two matching components match similar segments at different positions, and / or, the at least two matching components match the same similar segment but the similarity between the similar segment and the sequencing sequence does not meet the preset similarity requirement, then determine the sequencing sequence as a variant sequence;
[0144] Adopt local alignment technology to determine a reference segment for each variant sequence included in the file to be compressed on the reference sequence.
[0145] In an alternative embodiment, in the case where the sequencing sequence is determined to be a variant sequence, the at least two matching components no longer extract the remaining subsequences on the sequencing sequence.
[0146] In an alternative embodiment, a sliding window tool is respectively arranged in the at least two matching components, and the processor 71 can also be used for:
[0147] Allocate non-overlapping sequence segments on the sequencing sequence for different sliding window tools;
[0148] Control each sliding window tool to slide on the sequence segment allocated to itself to extract subsequences that meet the target requirements;
[0149] Among them, the target requirements include that the sequence elements at the ends of the subsequences are of a preset type, and / or the sequence elements at the ends of the subsequences are different from their adjacent sequence elements.
[0150] In an alternative embodiment, the alignment information corresponding to a single sequencing sequence includes multiple types of information content. When the processor 71 performs compression based on the alignment information generated under each sequencing sequence in the file to be compressed, it can be specifically used for:
[0151] Classify the information content in the alignment information generated under each sequencing sequence included in the file to be compressed according to the content type;
[0152] According to the correspondence between the content type and the encoding rule, encode the information content under different content types using the corresponding encoding rules respectively;
[0153] Compress the results generated after encoding under different content types respectively.
[0154] In an alternative embodiment, the content type includes a position type, and the position type information content is used to describe the position of the reference segment corresponding to the sequencing sequence; when the processor 71 classifies the information content in the alignment information generated under each sequencing sequence included in the file to be compressed according to the content type, it can be specifically used for:
[0155] Reorder the alignment information under each sequencing sequence in the file to be compressed according to the position of the reference segment;
[0156] Under the target content type, sequentially query the information content belonging to the target content type from the alignment information corresponding to each sequencing sequence according to the order between the reordered alignment information, so as to determine the information content under the target content type;
[0157] Among them, the target content type is any one of the content types.
[0158] In an alternative embodiment, if the file to be compressed is a sequencing data file generated based on single-end sequencing technology, then when the processor 71 encodes the position type information content, it can be specifically used for:
[0159] Encode each position type information content into a first marker array, and the elements in the first marker array correspond to each sequencing sequence one by one;
[0160] Sequentially determine whether the distance difference between the position of the reference segment and the position of the reference segment of the previous adjacent sequencing sequence exceeds a preset distance threshold for each sequencing sequence;
[0161] Record the position - type information content corresponding to each sequencing sequence determined to exceed the preset distance threshold in the first position array in sequence, and mark the corresponding element in the first marker array as a value for pointing to the value in the first position array;
[0162] Record the distance difference corresponding to each sequencing sequence determined not to exceed the preset distance threshold in the second position array in sequence, and mark the corresponding element in the first marker array as a value for pointing to the value in the second position array.
[0163] In an alternative embodiment, if the file to be compressed is one of two sequencing data files generated for the same sample based on the paired - end sequencing technology, a single sequencing data segment generated under the paired - end sequencing technology contains a pair of sequencing sequences, which are stored one - to - one in the file to be compressed and the paired other sequencing data file, and the other sequencing data file has completed the encoding of the position - type information content according to the encoding rules set for the single - end sequencing technology, then when the processor 71 encodes the position - type information content for the file to be compressed, it can be specifically used for:
[0164] After encoding each position - type information content into the first marker array, determine in sequence whether the position difference value between the reference fragment position of each sequencing sequence and the reference fragment position of the corresponding sequencing sequence in the other sequencing data file exceeds a preset difference threshold;
[0165] Record the position - type information content corresponding to each sequencing sequence determined to exceed the preset difference threshold in the first position array in sequence, and mark the corresponding element in the first marker array as a value pointing to the first position array;
[0166] Record the position difference value corresponding to each sequencing sequence determined not to exceed the preset difference threshold in the second position array in sequence, and mark the corresponding element in the first marker array as a value pointing to the second position array.
[0167] In an alternative embodiment, the content type includes strand - attribute type, and the strand - attribute type information content is used to describe whether the sequencing sequence matches the positive strand or the negative strand in the reference sequence. If the file to be compressed is a sequencing data file generated based on the single - end sequencing technology, then when the processor 71 encodes the strand - attribute type information content, it can be specifically used for:
[0168] Encode each strand - attribute type information content into a second marker array, where the elements in the second marker array correspond to each sequencing sequence one - to - one;
[0169] Mark the corresponding element of the sequencing sequence that matches the positive strand as the first value in the second marker array;
[0170] Mark the corresponding element in the second marker array with a second value for the sequencing sequence that matches the negative strand.
[0171] In an alternative embodiment, if the file to be compressed is any one of two sequencing data files generated for the same sample based on paired-end sequencing technology, and a single sequencing data segment generated under paired-end sequencing technology contains a pair of sequencing sequences, which are stored one-to-one in the file to be compressed and the paired other sequencing data file, then when the processor 71 encodes the content of the strand attribute class information, it can be specifically used for:
[0172] Configure a third marker array under the strand attribute class, where the elements in the third marker array correspond one-to-one to each sequencing data segment;
[0173] For the first type of sequencing data segments containing sequencing sequences with different strand attributes, perform an exchange operation on the content of the strand attribute class information as needed, so as to centralize the same strand attribute class information content under the same sequencing data file, and mark the corresponding elements of each first type of sequencing data segment in the third marker array with a target value;
[0174] After completing the exchange operation, perform an operation of encoding each content of the strand attribute class information into the second marker array and element marking for the file to be compressed.
[0175] In an alternative embodiment, the content type includes result class, and the result class information content includes an array for marking the consistency of sequence elements and each sequence element marked as inconsistent. When the processor 71 encodes the result class information content, it can be specifically used for:
[0176] In the result class information content corresponding to a single sequencing sequence in the file to be compressed, sequentially store each sequence element marked as inconsistent into the result array.
[0177] In an alternative embodiment, the processor 71 can also be used for:
[0178] For any quality evaluation sequence in the file to be compressed, check whether there is a target evaluation value whose proportion exceeds a preset proportion threshold;
[0179] If there is, after filtering the quality evaluation sequence based on the target evaluation value, perform an encoding operation;
[0180] Perform compression based on the encoded quality evaluation sequences in the file to be compressed.
[0181] In an alternative embodiment, when the processor 71 performs an encoding operation after filtering the quality evaluation sequence based on the target evaluation value, it can be specifically used for:
[0182] Encode the quality evaluation sequence into a fourth marker array, where the elements in the fourth marker array correspond one-to-one to the elements in the quality evaluation sequence. Among them, the elements in the quality evaluation sequence that are different from the target evaluation value are encoded as a third value in the fourth marker array, and the elements that are the same as the target evaluation value are encoded as a fourth value in the fourth marker array;
[0183] Record the original element values of the elements in the quality evaluation sequence that are different from the target evaluation value in the first evaluation value array pointed to by the third value in order.
[0184] In an alternative embodiment, the processor 71 may further be used for:
[0185] If the file to be compressed contains the identification information corresponding to each sequencing sequence, delete the identification information from the file to be compressed;
[0186] Merge the compression results obtained by performing compression on the quality evaluation sequence and the sequencing sequence respectively to obtain the compressed file corresponding to the file to be compressed;
[0187] After decompressing the compressed file, regenerate the identification information for each decompressed sequencing sequence to distinguish different sequencing sequences.
[0188] Further, as Figure 7 shown, the computing device further includes: a power supply component 73 and other components. Figure 7 Only some components are schematically shown in Figure 7 and it does not mean that the computing device only includes
[0189] It should be noted that for the technical details in the above embodiments of the computing device, reference can be made to the relevant descriptions in the foregoing method embodiments. To save space, they will not be repeated here, but this should not cause loss of the protection scope of this application.
[0190] Correspondingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed, it can implement each step that can be executed by the computing device in the above method embodiments.
[0191] The above Figure 7The memory therein is used to store computer programs and can be configured to store various other data to support operations on the computing platform. Examples of such data include instructions for any application or method operating on the computing platform, contact data, phone book data, messages, pictures, videos, etc. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disks or optical discs.
[0192] The above Figure 7 The communication component therein is configured to facilitate communication in a wired or wireless manner between the device where the communication component is located and other devices. The device where the communication component is located can access a wireless network based on communication standards, such as WiFi, 2G, 3G, 4G / LTE, 5G and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.
[0193] The above Figure 7 The power supply component therein provides power for various components of the device where the power supply component is located. The power supply component can include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device where the power supply component is located.
[0194] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0195] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram can be implemented by computer program instructions, as well as the combination of flows and / or blocks in the flowchart and / or block diagram. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0196] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0197] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more of the flows Figure 1 one or more of the flows and / or blocks Figure 1 or means for implementing the functions specified in one or more of the blocks.
[0198] It should also be noted that the term "including", "comprising", or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, commodity, or device including a series of elements includes not only those elements but also other elements not explicitly listed, or also includes elements inherent to such process, method, commodity, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, commodity, or device including the said element.
[0199] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0200] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.
Claims
1. A compression method for sequencing data files, characterized in that, Including: For any sequencing sequence in the file to be compressed, using at least two matching components to collaboratively perform a matching operation on a reference sequence to determine a reference fragment for the sequencing sequence; Constructing alignment information between the sequencing sequence and the reference fragment; Performing compression based on the alignment information generated for each sequencing sequence in the file to be compressed.
2. The method according to claim 1, wherein For any sequencing sequence in the file to be compressed, using at least two matching components to collaboratively perform a matching operation on a reference sequence to determine a reference fragment for the sequencing sequence, including: Using the at least two matching components to respectively match similar fragments for the sequencing sequence on the reference sequence based on subsequences extracted from the sequencing sequence, and the subsequences used by different matching components do not overlap with each other; If the at least two matching components match the same similar fragment and the similar fragment meets the preset similarity requirement, then determining the similar fragment as the reference fragment corresponding to the sequencing sequence.
3. The method according to claim 2, wherein Also including: If it is detected that the at least two matching components match similar fragments at different positions, and / or, the at least two matching components match the same similar fragment but the similarity between the similar fragment and the sequencing sequence does not meet the preset similarity requirement, then determining the sequencing sequence as a variant sequence; Adopting a local alignment technique to determine a reference fragment for each variant sequence included in the file to be compressed on the reference sequence; Wherein, in the case of determining that the sequencing sequence is a variant sequence, the at least two matching components no longer extract the remaining subsequences on the sequencing sequence.
4. The method according to claim 2, wherein A sliding window tool is respectively set in the at least two matching components, and the method further includes: Allocating non-overlapping sequence fragments for different sliding window tools on the sequencing sequence; Controlling each sliding window tool to slide on the sequence fragment allocated to itself to extract a subsequence that meets the target requirements; Wherein, the target requirements include that the sequence element at the end of the subsequence is of a preset type, and / or the sequence element at the end of the subsequence is different from its adjacent sequence element.
5. The method according to claim 1, wherein The alignment information corresponding to a single sequencing sequence contains multiple types of information content. Performing compression based on the alignment information generated for each sequencing sequence in the file to be compressed, including: Classifying the information content in the alignment information generated for each sequencing sequence included in the file to be compressed according to the content type; Encoding the information content under different content types respectively using corresponding encoding rules according to the correspondence between the content type and the encoding rule; Respectively compressing the results generated after encoding under different content types.
6. The method according to claim 5, wherein The content type includes a position type, and the position type information content is used to describe the position of the reference fragment corresponding to the sequencing sequence. When classifying the information content in the alignment information generated for each sequencing sequence included in the file to be compressed according to the content type, including: Reordering the alignment information for each sequencing sequence in the file to be compressed according to the reference fragment position; Under the target content type, in the order of the re-sorted comparison information, query the information content belonging to the target content type from the comparison information corresponding to each sequencing sequence in turn to determine the information content under the target content type; Among them, the target content type is any content type.
7. The method according to claim 6, wherein If the file to be compressed is a sequencing data file generated based on single-end sequencing technology, the process of encoding the location type information content includes: Encoding each location type information content into a first marker array, where the elements in the first marker array correspond to each sequencing sequence one by one; Successively determine for each sequencing sequence whether the distance difference between the reference fragment position and the reference fragment position of the adjacent previous sequencing sequence exceeds a preset distance threshold; Record the location type information content corresponding to the sequencing sequences determined to exceed the preset distance threshold in the first position array in turn, and mark the corresponding elements in the first marker array as values for pointing to the first position array; Record the distance differences corresponding to the sequencing sequences determined not to exceed the preset distance threshold in the second position array in turn, and mark the corresponding elements in the first marker array as values for pointing to the second position array.
8. The method according to claim 5 or 6, characterized in that The content type includes strand attribute types. The strand attribute type information content is used to describe whether the sequencing sequence matches the positive strand or the negative strand in the reference sequence. If the file to be compressed is a sequencing data file generated based on single-end sequencing technology, the process of encoding the strand attribute type information content includes: Encoding each strand attribute type information content into a second marker array, where the elements in the second marker array correspond to each sequencing sequence one by one; Mark the corresponding elements of the sequencing sequences that match the positive strand as the first value in the second marker array; Mark the corresponding elements of the sequencing sequences that match the negative strand as the second value in the second marker array.
9. The method according to claim 5 or 6, characterized in that, The content type includes result types. The result type information content includes an array for marking the consistency of sequence elements and each sequence element marked as inconsistent. The process of encoding the result type information content includes: In the result type information content corresponding to a single sequencing sequence in the file to be compressed, store each sequence element marked as inconsistent into the result array in order.
10. The method according to claim 1, wherein It also includes: For any quality evaluation sequence in the file to be compressed, check whether there is a target evaluation value whose proportion exceeds a preset proportion threshold; If so, perform an encoding operation after filtering the quality evaluation sequence based on the target evaluation value; Perform compression based on the encoded quality evaluation sequences in the file to be compressed.
11. The method according to claim 10, wherein Performing an encoding operation after filtering the quality evaluation sequence based on the target evaluation value includes: Encoding the quality evaluation sequence into a fourth marker array, where the elements in the fourth marker array correspond to each element in the quality evaluation sequence. Among them, each element in the quality evaluation sequence that is different from the target evaluation value is encoded as the third value in the fourth marker array, and each element that is the same as the target evaluation value is encoded as the fourth value in the fourth marker array; Record the original element values of the elements in the quality evaluation sequence that are different from the target evaluation value in the first evaluation value array pointed to by the third value in order.
12. The method according to claim 10, wherein Further included are: If the to-be-compressed file contains the identification information corresponding to each sequencing sequence, delete the identification information from the to-be-compressed file; Merge the compression results obtained by performing compression on the quality evaluation sequence and the sequencing sequence respectively to obtain the compressed file corresponding to the to-be-compressed file; After decompressing the compressed file, regenerate the identification information for each decompressed sequencing sequence to distinguish different sequencing sequences.
13. A compression method for sequencing data files, characterized in that, Included are: Based on the reference segments matched by each sequencing sequence contained in the to-be-compressed file on the reference sequence, construct the alignment information corresponding to each sequencing sequence, and the alignment information contains various types of information content; Classify the information content in the alignment information generated under each sequencing sequence contained in the to-be-compressed file according to the content type; According to the correspondence between the content type and the coding rule, encode the information content under different content types respectively using the corresponding coding rule; Compress the results generated after encoding under different content types respectively to obtain the compression result of the sequencing sequence in the to-be-compressed file.
14. A computing device, characterized in that, Including a memory and a processor; The memory is used to store one or more computer instructions; The processor is coupled to the memory and is used to execute the one or more computer instructions to execute the compression method of the sequencing data file according to any one of claims 1-13.
15. A computer-readable storage medium storing computer instructions, characterized in that, When the computer instructions are executed by one or more processors, cause the one or more processors to execute the compression method of the sequencing data file according to any one of claims 1-13.
16. A computer program product, characterized in that, Including a computer program / instructions, wherein when the computer program is executed by a processor, cause the processor to implement the compression method of the sequencing data file according to any one of claims 1-13.