Compression methods for sequencing data file, device and storage medium

By using multiple matching components on the reference sequence to determine the reference fragment and perform content classification encoding and compression, the problem of low compression rate of Fastq files is solved, and more efficient data compression is achieved.

WO2025158223A1PCT designated stage expired Publication Date: 2025-07-31CLOUD INTELLIGENCE ASSETS HOLDING (SINGAPORE) PTE LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/IB2025/050212
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-22
Filing Date
2025-01-09
Publication Date
2025-07-31

AI Technical Summary

Technical Problem

The existing Fastq file compression scheme has low compression rate, resulting in high storage and transmission costs.

Method used

At least two matching components are used to perform matching operations on the reference sequence, determine the reference fragment for the sequencing sequence, construct the alignment information between the sequencing sequence and the reference fragment, and encode and compress according to the content type using different encoding rules.

Benefits of technology

Improves the compression rate of sequencing data files and reduces storage and transmission costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure IB2025050212_31072025_PF_FP_ABST
    Figure IB2025050212_31072025_PF_FP_ABST
Patent Text Reader

Abstract

Provided in the embodiments of the present application are compression methods for a sequencing data file, a device and a storage medium. Provided is a compression solution for a sequencing data file, comprising: for each sequencing sequence contained in a sequencing data file, proposing to use at least two matching components to cooperatively determine a reference fragment for the same sequencing sequence on a reference sequence; and, on the basis of the reference fragment determined accordingly, constructing comparison information between the sequencing sequence and the reference fragment, and compressing each piece of comparison information constructed under a file to be compressed. Thus, by means of cooperative operation of at least two matching components, the efficiency of searching reference fragments for sequencing sequences can be improved, and matching results can be mutually verified among matching components, thus effectively improving the matching accuracy of reference fragments.
Need to check novelty before this filing date? Find Prior Art

Description

Compression Method, Device, and Storage Medium for Sequencing Data Files - Technical Field

[0001] This application relates to the technical field of data processing, and particularly to a compression method, device, and storage medium for sequencing data files. Background Art

[0002] The Fastq format is a commonly used format for storing raw sequencing data in the bioinformatics field. A Fastq file stores biological sequences (usually nucleic acid sequences) and corresponding quality evaluation value sequences.

[0003] As the number of Fastq files generated in the bioinformatics field increases, new challenges are brought to the storage cost and transmission cost of Fastq files. Compressing Fastq files can greatly reduce the file size, thereby saving storage costs and transmission costs. However, the compression ratios that can be achieved by current compression schemes for Fastq files are generally not high, resulting in poor compression effects for Fastq files.

[0004] Therefore, there is an urgent need for a solution that can provide a higher compression ratio for Fastq files. Summary of the Invention

[0005] Multiple aspects of this application provide a compression method, device, and storage medium for sequencing data files to improve the compression ratio of sequencing data files.

[0006] An embodiment of this application provides a compression method for sequencing data files, including: for any sequencing sequence in the file to be compressed, using at least two matching components to perform a matching operation on a reference sequence in cooperation to determine a reference fragment for the sequencing sequence; constructing alignment information between the sequencing sequence and the reference fragment; and performing compression based on the alignment information generated for each sequencing sequence in the file to be compressed.

[0007] Another embodiment of this application provides a compression method for sequencing data files, including: respectively constructing alignment information corresponding to each sequencing sequence based on the reference fragments matched by each sequencing sequence included in the file to be compressed on the reference sequence, where the alignment information includes various types of information content; classifying the information content in the alignment information generated for each sequencing sequence included in the file to be compressed according to the content type; encoding the information content under different content types respectively using corresponding encoding rules according to the correspondence between the content type and the encoding rule; and respectively compressing the results generated after encoding under different content types to obtain the compression result of the sequencing sequence in the file to be compressed.

[0008] An embodiment of the present application further provides a computing device, including a memory and a processor; the memory is used to store one or more computer instructions; The processor is coupled to the memory and is configured to execute the one or more computer instructions to execute the foregoing compression method for sequencing data files.

[0009] An embodiment of the present application further provides a computer-readable storage medium storing computer instructions, which, when executed by one or more processors, cause the one or more processors to execute the foregoing compression method for sequencing data files.

[0010] An embodiment of the present application further provides a computer program product, including a computer program / instructions, wherein, when the computer program is executed by a processor, the processor is caused to implement the foregoing compression method for sequencing data files

[0011] In an embodiment of the present application, a compression scheme for sequencing data files is proposed. For each sequencing sequence included in the sequencing data file, it is proposed to use at least two matching components to cooperate to determine a reference fragment for the same sequencing sequence on the reference sequence; based on the determined reference fragment, alignment information between the sequencing sequence and the reference fragment can be constructed, and compression is performed on each piece of alignment information constructed under the file to be compressed. In this way, by the cooperation of at least two matching components, not only can the efficiency of searching for reference fragments for sequencing sequences be improved, but also the matching results between the matching components can be mutually verified, which can effectively improve the matching accuracy of the reference fragments, thereby improving the compression ratio for sequencing sequences, and further improving the compression ratio of the sequencing data file. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0013] FIG. 1 is a schematic flowchart of a compression method for a sequencing data file provided by an exemplary embodiment of the present application;

[0014] FIG. 2a is a schematic internal structure diagram of a Fastq file provided by an exemplary embodiment of the present application;

[0015] FIG. 2b is a schematic logical diagram of a content splitting method for a Fastq file provided by an exemplary embodiment of the present application;

[0016] FIG. 3 is a schematic logical diagram of an exemplary scheme for extracting subsequences provided by an exemplary embodiment of the present application;

[0017] FIG. 4 is a schematic flowchart of another method for compressing a sequencing data file provided by an exemplary embodiment of the present application;

[0018] FIG. 5 is a schematic flowchart of yet another method for compressing a sequencing data file provided by an exemplary embodiment of the present application;

[0019] FIG. 6 is a schematic flowchart of a method for compressing a sequencing data file provided by another exemplary embodiment of the present application;

[0020] FIG. 7 is a schematic structural diagram of a computing device provided by yet another exemplary embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0021] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions of the present application will be clearly and completely described below in conjunction with the specific embodiments of the present application and the corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0022] Before starting to elaborate on the technical solutions provided by the embodiments of the present application, several technical concepts involved in the present application are briefly explained as follows.

[0023] A sequencing data file can be understood as a sequencing result file output after using a sequencing instrument to perform DNA sequencing on a sample. The sequencing data file at least contains the sequencing sequences obtained by sequencing (which can also be called biological sequences, usually base sequences, or can also be amino acid sequences).

[0024] A Fastq file is a sequencing data file in a typical format, used to store sequencing sequences and corresponding quality evaluation sequences. The Fastq file contains multiple reads. Each read generally contains four lines. The first line starts with @ and is followed by the description information of the biological sequence. The second line is the biological sequence. The third line starts with + and can also be followed by the description information of the biological sequence. The fourth line is the quality evaluation corresponding to the biological sequence in the second line (quality values, the quality evaluation of sequencing). The number of elements in the quality evaluation sequence in the fourth line is the same as the number of elements in the biological sequence in the second line.

[0025] A reference sequence refers to a reference genomic sequence used for comparison and analysis. It is a known genomic sequence obtained through sequencing and assembly, and is usually a representative sample of the genome of a certain species. Reference sequences generally come from public databases, which provide a large number of known genomic sequences for researchers to use in gene sequencing and bioinformatics research. At the same time, with the development of technology and the implementation of new sequencing projects, reference sequences are constantly updated and improved to meet the research needs of more species and greater precision.

[0026] As mentioned in the background art, with the continuous progress of gene sequencing technology, the types of sequencing data files generated are increasing, bringing new challenges to storage and transmission. The inventors found in the research process that there are currently some solutions for compressing sequencing data files in the field, but the compression ratios achieved by these solutions are generally not high, resulting in little cost savings in storage and transmission of the compressed files.

[0027] Therefore, in this embodiment, a compression method for sequencing data files is proposed to improve the compression ratio of sequencing data files.

[0028] The following will detail the technical solutions provided by each embodiment of the present application with reference to the accompanying drawings.

[0029] FIG. 1 is a schematic flowchart of a compression method for sequencing data files provided by an exemplary embodiment of the present application. This method can be executed by a file compression device, which can be implemented as software, hardware, or a combination of software and hardware, and can be integrated in a computing device. Referring to FIG. 1, this method may include steps 100 to 102.

[0030] Step 100: For any sequencing sequence in the file to be compressed, use at least two matching components to perform a matching operation on the reference sequence in cooperation to determine a reference fragment for the sequencing sequence.

[0031] Step 101: Construct alignment information between the sequencing sequence and the reference fragment.

[0032] Step 102: Perform compression based on the alignment information generated for each sequencing sequence in the file to be compressed.

[0033] In this embodiment, the sequencing data file to be compressed is described as a file to be compressed. It should be noted that in this embodiment, the file format of the file to be compressed is not limited. For example, the file to be compressed can be the Fastq file mentioned above, or other format files that can be used to store sequencing sequences currently or in the future. No more examples of file formats are given here. In this embodiment, the file to be compressed may at least contain sequencing sequences, and usually there are multiple sequencing sequences in the file to be compressed. In this embodiment, other file contents included in the file to be compressed are not limited. For different file formats, the other file contents included in the file to be compressed may not be exactly the same. For example, in the Fastq file mentioned above, in addition to containing sequencing sequences, it also contains quality evaluation sequences and identification information of the sequencing sequences and other file contents. No more examples of other file contents included in other format files are given either.

[0034] The compression method for the sequencing data file provided in this embodiment can perform content splitting on the file to be compressed to obtain each sequencing sequence included in the file to be compressed.

[0035] FIG. 2a is a schematic diagram of the internal structure of a Fastq file provided by an exemplary embodiment of the present application. FIG. 2b is a schematic diagram of the logic of a content splitting method for a Fastq file provided by an exemplary embodiment of the present application. Referring to FIG. 2a, in this case, the file to be compressed contains multiple reads. What is shown in FIG. 2a is the 4 lines of content included in one read. Among them, the second line of the read is the sequencing sequence, the fourth line of the read is the quality evaluation sequence, and the first line and the third line of the read are the identification information of the sequencing sequence. Based on this, referring to FIG. 2b, in this embodiment, it is proposed that the file to be compressed can be split into three parts: the sequencing sequence part, the quality evaluation sequence part, and the identification part. Among them, the sequencing sequence part contains the sequencing sequences included in each read in the file to be compressed, the quality evaluation sequence part contains the quality evaluation sequences included in each read in the file to be compressed, and the identification part contains the identification information included in each read in the file to be compressed.

[0036] On this basis, in this embodiment, an improved solution is provided for the compression of the sequencing sequences in the file to be compressed. Based on the improved solution provided in this embodiment, the sequencing sequence part cut out from the file to be compressed can be compressed separately to improve the compression rate of the sequencing sequence part, and thus improve the compression rate of the file to be compressed. In this embodiment, it is further proposed that an improved solution can also be provided for the compression of other parts cut out from the file to be compressed, which will be described in detail later.

[0037] Referring to FIG. 2b, in this embodiment, different file parts cut from the file to be compressed can be compressed separately. By improving the compression ratio of each individual file part, the overall compression ratio of the file to be compressed can be increased. Among them, after the compression results generated by separately compressing each file part are combined, a compressed file corresponding to the file to be compressed can be generated. The compressed file will be the object for storage and transmission.

[0038] The following will elaborate on each step in the file method provided in this embodiment in detail. In this embodiment, the compression scheme for the sequencing sequence part can at least include two links: the first is a preprocessing link that takes a single sequencing sequence as the object, and the second is an execution compression link that takes each sequencing sequence as the object. The following will elaborate on the two links separately in detail.

[0039] Regarding the first link, considering that the preprocessing logic implemented for each sequencing sequence is the same, for the convenience of description, any sequencing sequence in the file to be compressed will be taken as an example to elaborate on the first link in detail.

[0040] Referring to FIG. 1, in step 100, for any sequencing sequence in the file to be compressed, at least two matching components are used to perform a matching operation on the reference sequence in cooperation to determine a reference fragment for the sequencing sequence. Among them, the matching component can be understood as a logical component with the function of matching similar fragments for the sequencing sequence on the reference sequence. In step 100, at least two matching components can be used to work together to determine a reference fragment for the same sequencing sequence on the reference sequence. The cooperation in this embodiment can be understood as reasonably dividing labor, cooperating with each other, and corroborating with each other among at least two matching components, so as to determine a reference fragment for the same sequencing sequence by synthesizing the matching results of each matching component. As mentioned above, each matching component has the ability to match similar fragments for the sequencing sequence on the reference sequence. Therefore, in step 100, the matching results obtained by each matching component for the same sequencing sequence can be corroborated with each other to finally determine the reference fragment corresponding to the sequencing sequence.

[0041] In a preferred implementation manner: at least two matching components can be used to match similar fragments for the sequencing sequence on the reference sequence respectively based on the subsequences extracted from the sequencing sequence. The exemplary logic for matching similar fragments can be: using the subsequence used as a seed to search on the reference sequence to see if there is a fragment identical to the subsequence; if there is If so, map the relative positional relationship between the subsequence and the sequencing sequence to the reference sequence to locate the similar segment for the sequencing sequence. It should be understood that this is only exemplary, and the matching logic in the matcher component in this embodiment is not limited to this, and no more examples are given here.

[0042] The subsequence in this embodiment can be understood as a subset of the sequencing sequence. In this embodiment, it is proposed that the subsequences used in different matching components do not repeat each other, which can effectively avoid the problem of repeatedly performing matching operations on the subsequence, so as to avoid affecting the matching efficiency. This embodiment provides an exemplary subsequence extraction scheme: sliding window tools can be respectively set in at least two matching components. Based on this, non-overlapping sequence segments can be allocated to different sliding window tools on the sequencing sequence; control each sliding window tool to slide on the sequence segment allocated to itself to extract the subsequence that meets the target requirements.

[0043] FIG. 3 is a schematic logical diagram of an exemplary scheme for extracting a subsequence provided by an exemplary embodiment of the present application. Referring to FIG. 3, for example, two matching components can be used to match similar segments for the sequencing sequence, and the sequencing sequence can be divided into two sequence segments and allocated to the sliding window tools in the two matching components (the two squares shown in FIG. 3 are the sliding windows). The starting positions and sliding directions of the two sliding window tools can also be configured. Based on this, the two sliding window tools can be controlled to slide on their respective sequence segments. It should be understood that the width of the sliding window should be the same as the length of the subsequence. Referring to FIG. 3, the two sliding window tools can slide from both ends of the sequencing sequence to the middle, and the subsequence is the one hit by the sliding window.

[0044] In this exemplary subsequence extraction scheme, the aforementioned target requirements can be flexibly set according to actual needs. Exemplarily, the target requirements may include but are not limited to that the sequence elements at the ends of the subsequence are of a preset type, and / or the sequence elements at the ends of the subsequence are different from their adjacent sequence elements. Here, the sequence elements can be bases or amino acids, etc. In this way, the two matching components can independently extract the subsequences they use, and the two matching components will not extract the same subsequence.

[0045] In addition, in step 100, preferably, at least two matching components can run in parallel. From the perspective of a single matching component, after matching a similar segment for the sequencing sequence on the reference sequence based on a certain extracted subsequence, it can pause to wait for the matching results in other matching components.

[0046] During the research process, the inventors found that sequencing sequences usually fall into two categories: variant sequences and non-variant sequences. Among them, variant sequences can be understood as those that have undergone insertions or deletions of sequence elements relative to a certain fragment on the reference sequence. That is to say, variant sequences are usually due to Insertion and Deletion (InDel) events, and InDel events are usually caused by processes such as gene mutations, chromosomal rearrangements, and gene recombinations. Correspondingly, non-variant sequences can be understood as sequencing sequences that have not undergone InDel events.

[0047] In this embodiment, by using at least two matching components to match similar fragments for the same sequencing sequence on the reference sequence, it is possible to efficiently identify whether the sequencing sequence belongs to a variant sequence or a non-variant sequence.

[0048] Continuing to refer to Figure 1, it is proposed in step 101 that if at least two matching components match the same similar fragment and the similar fragment meets the preset similarity requirement, the similar fragment is determined as the reference fragment corresponding to the sequencing sequence. The inventors found during the research process that for non-variant sequences, there are fragments on the reference sequence that can be aligned with the non-variant sequence and have a high enough similarity, and such fragments can be used as the reference fragments for non-variant sequences. Among them, the reference fragment can be understood as the fragment on the reference sequence used as the compression reference for the sequencing sequence. Therefore, in this embodiment, for a sequencing sequence, if at least two matching components match the same similar fragment and the similar fragment meets the preset similarity requirement, it can be determined that the sequencing sequence is a non-variant sequence. And the matching similar fragment can be used as the reference fragment corresponding to the sequencing sequence.

[0049] The inventors also found that in practical applications, since most subsequences in non-variant sequences can match the final reference fragment on the reference sequence, in this embodiment, at least two matching components generally only need to perform a small number of matching operations and then can match the similar fragment and pause. In this way, in this embodiment, at least two matching components can be used to efficiently match the reference fragment for non-variant sequences.

[0050] Moreover, since at least two matching components are used to perform matching operations for the same sequencing sequence in this embodiment, the matching results obtained by at least two matching components can be mutually verified. As mentioned before, at least two matching components need to match the same similar fragment and the similar fragment needs to meet the preset similarity requirement before the similar fragment is determined as the reference fragment. This can effectively improve the matching accuracy of the reference fragment and avoid matching an inappropriate reference sequence for variant sequences.

[0051] In other words, in this embodiment, if it is identified by at least two matching components that the sequencing sequence belongs to a non-variant sequence, the reference segment corresponding to the sequencing sequence will be determined based on the matching results of the at least two matching components. If it is identified by at least two matching components that the sequencing sequence belongs to a variant sequence, then the method will switch to other reference segment determination methods more suitable for variant sequences to more accurately determine the reference segment for the variant sequence.

[0052] In this embodiment, it is proposed that if it is detected that the at least two matching components match similar segments at different positions, and / or the at least two matching components match the same similar segment but the similarity between the similar segment and the sequencing sequence does not meet the preset similarity requirement, then it is determined that the sequencing sequence is a variant sequence.

[0053] As mentioned above, for a variant sequence, insertions or deletions of sequence elements have occurred relative to the reference sequence. Therefore, there may be multiple subsequences on the variant sequence that can match similar segments on the reference sequence. However, the positions of these similar segments are usually different. For this reason, in this embodiment, when at least two matching components match different similar segments for the same sequencing sequence, it can be determined that the sequencing sequence is a variant sequence. In addition, considering that the subsequences used in at least two matching components may be located in the same region as the inserted or deleted sequence elements in the variant sequence, in this case, if at least two matching components may match the same similar segment for the variant sequence, for this reason, in this embodiment, when the at least two matching components match the same similar segment but the similarity between the similar segment and the sequencing sequence does not meet the preset similarity requirement, it can also be determined that the sequencing sequence is a variant sequence.

[0054] Furthermore, in this embodiment, it is also proposed that after it is determined based on the above logic that the sequencing sequence is a variant sequence, the at least two matching components will no longer extract the remaining subsequences on the sequencing sequence. In this way, variant sequences can be identified more efficiently, and the at least two matching components do not need to traverse all the subsequences on the variant sequence.

[0055] Of course, in this embodiment, if it is detected that any one of the matching components has not matched a similar segment for the sequencing sequence on the reference sequence after traversing all the subsequences it has extracted, then it can also be determined that the sequencing sequence is a variant sequence. The other matching components can no longer extract the remaining subsequences on the sequencing sequence.

[0056] In this way, in this embodiment, the sequencing sequences that may have mutated can be identified more comprehensively and efficiently. Thus, the determination method of the reference fragment can be switched in a timely manner for such sequencing sequences, and further, the accuracy of the reference fragment of such sequencing sequences can be ensured.

[0057] Exemplarily, this embodiment proposes that local alignment technology can be used to determine reference fragments for each mutated sequence included in the file to be compressed on the reference sequence. Local alignment technology: It is to find a suitable matching region in two sequences for alignment to find a suitable matching scheme. Local alignment technology is usually used to align sequences with relatively low similarity. In this embodiment, the type of algorithm used in the local alignment technology is not limited. For example, the BWA (Burrows-Wheeler Aligner) algorithm or the Wave Front Alignment (WFA) algorithm, etc. can be used. Here, the principle of the local alignment technology will not be elaborated in detail. Any algorithm in the local alignment technology can be used in this embodiment to determine the reference fragment for the mutated sequence.

[0058] Continuing to refer to FIG. 1, after determining the reference fragment for the sequencing sequence, the alignment information between the sequencing sequence and the reference fragment can be constructed. Among them, the alignment information is used to describe the differences between the sequencing sequence and the reference fragment.

[0059] In this way, by implementing the above-mentioned preprocessing logic for each sequencing sequence included in the file to be compressed, the alignment information generated under each sequencing sequence in the file to be compressed can be obtained. As mentioned above, in this embodiment, the reference fragment can be determined more accurately for the sequencing sequence, which can effectively reduce the amount of information of the alignment information generated by each sequencing sequence.

[0060] For the second link in the compression scheme for the sequencing sequence part, referring to FIG. 1, in step 103, compression can be performed based on the alignment information generated under each sequencing sequence in the file to be compressed. That is, each sequencing sequence in the file to be compressed will be converted into the corresponding alignment information. By performing compression on these alignment information, the compression of the sequencing sequence part cut out from the file to be compressed can be achieved. Since the reference fragment can be determined more accurately in this embodiment, the amount of information of the alignment information can be effectively reduced, and thus an excellent compression ratio can be obtained in the second link.

[0061] In summary, in this embodiment, a compression scheme for sequencing data files is proposed. For each sequencing sequence included in the sequencing data file, it is proposed to use at least two matching components to cooperate to determine a reference fragment for the same sequencing sequence on the reference sequence. Based on the determined reference fragment, the alignment information between the sequencing sequence and the reference fragment can be constructed, and the alignment information constructed under the file to be compressed is compressed. In this way, through the cooperation of at least two matching components, not only can the efficiency of searching for reference fragments for sequencing sequences be improved, but also the matching results can be mutually verified between the matching components, which can effectively improve the matching accuracy of the reference fragments, thereby improving the compression rate for sequencing sequences and further improving the compression rate of the sequencing data file.

[0062] FIG. 4 is a schematic flowchart of another compression method for a sequencing data file provided by an exemplary embodiment of the present application. Referring to FIG. 4, the method may include steps 401 to 405.

[0063] Step 401: For any sequencing sequence in the file to be compressed, use at least two matching components to perform a matching operation on the reference sequence in cooperation to determine a reference fragment for the sequencing sequence.

[0064] Step 402: Construct the alignment information between the sequencing sequence and the reference fragment.

[0065] Step 403: Classify the information content in the alignment information generated under each sequencing sequence included in the file to be compressed according to the content type.

[0066] Step 404: According to the correspondence between the content type and the encoding rule, encode the information content under different content types respectively using the corresponding encoding rules.

[0067] Step 405: Compress the results generated after encoding under different content types respectively.

[0068] Among them, steps 401 to 402 can refer to the relevant descriptions in the foregoing embodiments and will not be repeated here. In this embodiment, an optional implementation manner for compressing the alignment information generated in the file to be compressed is provided based on steps 403 to 405. This optional implementation manner can be combined with other implementation manners in the above or following embodiments to generate a variety of technical solutions.

[0069] Referring to FIG. 4, in this embodiment, the alignment information corresponding to a single sequencing sequence may include various types of information content. The content types in this embodiment may include but are not limited to position types, strand attribute types, result types, etc. Among them, the position The location class information content can be used to describe the position of the reference segment corresponding to the sequencing sequence; the strand attribute class information content can be used to describe whether the sequencing sequence matches the positive strand or the negative strand in the reference sequence; the result class information content can include an array for marking the consistency of sequence elements and each sequence element marked as inconsistent.

[0070] In step 403, according to the content type, the information content in the alignment information generated for each sequencing sequence included in the file to be compressed can be classified. It should be understood that since the file to be compressed contains multiple sequencing sequences, therefore, in this embodiment, the alignment information generated under the file to be compressed will also be multiple. Moreover, since there is usually an order relationship between the sequencing sequences in the file to be compressed, therefore, in this embodiment, it is proposed that the alignment information generated under the file to be compressed can inherit this order relationship. Based on this, in step 403, the information content in each alignment information can be divided into the corresponding content type, so that a group of information content arranged in order can be obtained under a single content type.

[0071] In a preferred solution, the alignment information for each sequencing sequence in the file to be compressed can be re - sorted according to the reference segment position; under the target content type, according to the order between the re - sorted alignment information, the information content belonging to the target content type is sequentially queried from the alignment information corresponding to each sequencing sequence to determine the information content under the target content type; where the target content type is any one of the content types. That is, in this exemplary solution, the sequencing sequences in the file to be compressed are re - sorted, so that the order relationship generated between the re - sorted alignment information will be inherited under each content type. Based on this preferred solution, the reference segment positions in each alignment information can be adjacent to each other in the reference sequence, so that the result generated by performing encoding in the information class will be more conducive to compression.

[0072] Of course, this is only a preference. In this embodiment, re - sorting may not be performed either. Whether re - sorting is performed or not, the order relationship between the alignment information will be inherited under each content type according to the actual arrangement state.

[0073] Referring to FIG. 4, after the classification of the information content is completed, in step 404, it is proposed to provide custom encoding rules for different content types to obtain a better encoding effect under each content type. Based on this, according to the correspondence relationship between the content type and the encoding rule, the information content under different content types is encoded using the corresponding encoding rules respectively.

[0074] In the process of research, the inventor found that the commonly used sequencing technologies in the current field include single-end sequencing technology and paired-end sequencing technology. Among them, the single-end sequencing technology can be understood as sequencing unidirectionally from one end to the other end of the gene fragment fragments. If sequencing is performed in two directions from both ends to the other end respectively, it is called paired-end sequencing. The single-end sequencing technology generates one sequencing sequence for each gene fragment, while the paired-end sequencing generates two sequencing sequences with opposite directions for each gene fragment. From the perspective of the sequencing data file, the single-end sequencing technology will generate one sequencing data file for the same sample, while the paired-end sequencing technology will generate a pair of sequencing data files for the same sample. For the convenience of description, in this embodiment, the two sequencing sequences generated for a fragment in the paired-end sequencing file are collectively referred to as the sequencing data segment. It can be understood that a single sequencing data segment contains a pair of sequencing sequences, which are stored one-to-one in the aforementioned pair of sequencing data files. In addition, it is worth noting that the sequencing sequences between the pair of sequencing data files generated under the paired-end sequencing technology are aligned. Under the aforementioned reordering scheme, it can be understood that the sequencing data segments are reordered, and the sequencing sequences in the pair of sequencing data files will change their positions synchronously with the change of the position of the sequencing data segment.

[0075] Therefore, in this embodiment, it is proposed that for the files to be compressed generated by the single-end sequencing technology and the files to be compressed generated by the paired-end sequencing technology, different encoding rules can be adopted for information content encoding under some content types.

[0076] The following will separately expand and explain the exemplary encoding processes for different content types.

[0077] Position type information content.

[0078] If the file to be compressed is a sequencing data file generated based on single-end sequencing technology, an exemplary encoding scheme for positional information content can be as follows: encode each piece of positional information content into a first marker array, where the elements in the first marker array correspond to each sequencing sequence; sequentially determine for each sequencing sequence whether the distance difference between the reference fragment position and the reference fragment position of the adjacent previous sequencing sequence exceeds a preset distance threshold; record the positional information content corresponding to each sequencing sequence determined to exceed the preset distance threshold in the first position array in sequence, and mark the corresponding element in the first marker array with a value used to point to the first position array; record the distance difference corresponding to each sequencing sequence determined not to exceed the preset distance threshold in the second position array in sequence, and mark the corresponding element in the first marker array with a value used to point to the second position array. Among them, the element values in the first marker array can be 1 or 0.

[0079] For example, if the file to be compressed contains 5 sequencing sequences, after reordering according to the reference fragment position mentioned above, the 5 pieces of positional information content are: coordinate A, coordinate B, coordinate C, coordinate D, and coordinate E. In this embodiment, instead of directly compressing these coordinates, they are encoded. For example, the distance difference between coordinate B and coordinate A is 3, the distance difference between coordinate C and coordinate B is 7, the distance difference between coordinate D and coordinate C is 2, and the distance difference between coordinate E and coordinate D is 1. If the preset distance threshold is 5, the encoded first marker array will be [1, 0, 1, 0, 0], where the value 1 in the first marker array points to the first position array, and the value 0 points to the second position array. And the encoded first position array will be [coordinate A, coordinate C], and the encoded second position array will be [3, 2, 1]. It can be seen that after encoding, only two original pieces of positional information content - coordinate A and coordinate C - are left, and the other positional information has been encoded into an array that is more conducive to compression.

[0080] If the file to be compressed is one of the two sequencing data files generated for the same sample based on the paired-end sequencing technology, and the other sequencing data file has completed the encoding of the location-related information content according to the encoding rules set for the single-end sequencing technology, then for the file to be compressed, an exemplary encoding scheme for the location-related information content can be as follows: after encoding each location-related information content into a first marker array, successively determine for each sequencing sequence whether the position difference value between the reference fragment position and the reference fragment position of the corresponding sequencing sequence in the other sequencing data file exceeds a preset difference threshold; record the location-related information content corresponding to each sequencing sequence determined to exceed the preset difference threshold successively in a first position array, and mark the corresponding element in the first marker array as the value pointing to the first position array; record the position difference values corresponding to each sequencing sequence determined not to exceed the preset difference threshold successively in a second position array, and mark the corresponding element in the first marker array as the value pointing to the second position array.

[0081] That is to say, for one of the paired-end sequencing-generated sequencing data files, the encoding scheme provided for the single-end testing technology can be independently used to complete the encoding of the location-related information content. For the other sequencing data file, the location-related information content is encoded according to the exemplary encoding scheme provided here. In this case, first, each location-related information content is encoded into a first marker array. Then, instead of comparing the reference fragment positions between the sequencing sequences within the file to be compressed, the reference fragment positions of each sequencing sequence in the file to be compressed are respectively compared with the corresponding sequencing sequences in the other paired sequencing data file, and then the first position array and the second position array are encoded according to the comparison results and the element values are determined in the first marker array.

[0082] Continuing with the above example, if the file to be compressed is paired with sequencing data file 1, then the coordinates A, B, C, D, and E under the file to be compressed will respectively correspond to the reference Coordinates of the fragments: Coordinate A, Coordinate B, Coordinate C', Coordinate D, and Coordinate E, are compared. The position difference value between Coordinate A and Coordinate A, is 1, the position difference value between Coordinate B and Coordinate B, is 3, the position difference value between Coordinate C and Coordinate C, is 2, the position difference value between Coordinate D and Coordinate D, is 0, and the position difference value between Coordinate E and Coordinate E, is 1. If the preset difference threshold is 4, the first marker array encoded will be [0, 0, 0, 0, 01, where the value 1 in the first marker array points to the first position array, and the value 0 points to the second position array. The encoded first position array will be empty, and the encoded second position array will be [1, 3, 2, 0, 1]. It can be seen that after encoding, no original position type information content is left, but all have been encoded into arrays that are more conducive to compression.

[0083] It should be understood that the above encoding scheme for position type information content provided under the paired-end sequencing technology is exemplary. This embodiment is not limited to this. Under the paired-end sequencing technology, it is also possible not to perform reference fragment position comparison between paired sequencing data files, but to perform independent comparison within a single sequencing data file without interference.

[0084] In this way, in this embodiment, for the file to be compressed, the position type information content will be encoded into 3 arrays: the first marker array, the first position array, and the second position array. These arrays are more conducive to compression than the position type information content itself and can obtain a higher compression rate.

[0085] Chain attribute type information content.

[0086] If the file to be compressed is a sequencing data file generated based on the single-end sequencing technology, an exemplary encoding method for the chain attribute type information content can be: encoding each chain attribute type information content into a second marker array, where the elements in the second marker array correspond one-to-one to each sequencing sequence; marking the corresponding element in the second marker array with the first value for the sequencing sequence matching the positive strand; and marking the corresponding element in the second marker array with the second value for the sequencing sequence matching the negative strand. Among them, the element values in the second marker array can be 1 or 0. Here, it is only required that the first value and the second value are different, and it is not limited whether the first value is 1 or 0.

[0087] For example, if the file to be compressed contains 5 sequencing sequences, and the 5 strand attribute information contents are in sequence: forward strand, forward strand, forward strand, reverse strand, reverse strand. In this embodiment, these strand attribute type information contents are not directly compressed, but are encoded into a second marker array which will be [1, 1, 1, 0, 0], where the element value 1 in the second marker array represents the forward strand, and the element value 0 represents the reverse strand.

[0088] If the file to be compressed is any one of the two sequencing data files generated for the same sample based on the paired-end sequencing technology, then an exemplary encoding method for the strand attribute type information content can be: configure a third marker array under the strand attribute type, and the elements in the third marker array correspond one by one to each sequencing data segment; for the first type of sequencing data segment containing sequencing sequences with different strand attributes, perform an exchange operation on the strand attribute type information content as needed, so as to concentrate the same strand attribute type information content under the same sequencing data file, and mark the elements corresponding to each first type of sequencing data segment in the third marker array as the target value; after the exchange operation is completed, perform the operation of encoding each strand attribute type information content into a second marker array and element marking for the file to be compressed.

[0089] That is to say, in the case of the paired-end sequencing technology, the sequencing sequence exchange operation can be first performed between the paired sequencing data files to try to concentrate the sequencing sequences with the same strand attribute into the same sequencing data file. After the exchange operation is completed, the two sequencing data files can then independently implement the foregoing encoding process for the second marker array.

[0090] Continuing with the above example, the 5 strand attribute information contents under the file to be compressed are in sequence: forward strand, forward strand, forward strand, reverse strand, reverse strand. If the 5 strand attribute information contents under the paired sequencing data file 1 of the file to be compressed are in sequence: reverse strand, reverse strand, forward strand, forward strand, forward strand, then the target strand attribute can be first determined for the file to be compressed. For example, the target strand attribute Set as the positive strand. After that, it can be detected that the 4th and 5th sequencing data segments need to perform a swap operation. After performing the swap operation accordingly, the 5 chain attribute information contents under the file to be compressed will be in turn: positive strand, positive strand, positive strand, positive strand, positive strand; and the 5 chain attribute information contents under the sequencing data file 1 will be in turn: negative strand, negative strand, positive strand, negative strand, negative strand. Correspondingly, the second marker array encoded for the file to be compressed will be [1,1,1,1,1], and the second marker array encoded for the sequencing data file 1 will be [0,0,1,0,0]. And the third marker array configured under the chain attribute class can be encoded as [0,0,0,1,1], where the element value of 1 in the third array indicates that an interaction operation has occurred, and the element value of 0 indicates that no swap operation has occurred.

[0091] It can be understood that after the foregoing interaction operation, the same chain attribute class information contents are concentrated under the same sequencing data file. This makes it such that in the second marker arrays encoded under the two paired sequencing data files, the element values will basically be the same. And the more elements with the same value in a single array, the higher the compression rate that can be obtained. Therefore, this exemplary encoding scheme can effectively improve the compression rate of the chain attribute class information contents.

[0092] In this way, in this embodiment, for the file to be compressed, the chain attribute class information content will be encoded into 1 array: the second marker array. Under the paired-end sequencing technology, a third marker array can also be provided for the two paired sequencing data files. These arrays are more conducive to compression compared to the chain attribute class information content itself and can obtain a higher compression rate.

[0093] Result class information content.

[0094] As mentioned above, the result class information content in this embodiment may include an array for marking the consistency of sequence elements and each sequence element marked as inconsistent. Among them, the array for marking the consistency of sequence elements is already in the form of an array and can be retained in this embodiment. And for each sequence element marked as inconsistent, this embodiment proposes that in the result class information content corresponding to a single sequencing sequence in the file to be compressed, each sequence element marked as inconsistent can be stored in the result array in sequence. That is, a result array is used to carry each sequence element marked as inconsistent.

[0095] So far, in this embodiment, the corresponding information content can be encoded into an array under different content types, and through the designed encoding rules, the encoded array can obtain a higher compression rate.

[0096] On this basis, continue to refer to FIG. 4. In step 405, the results generated after encoding for different content types can be compressed separately, which can further improve the compression ratio.

[0097] In summary, in this embodiment, it is proposed to classify the information content in each comparison information generated under the file to be compressed according to the content type, so as to separate the information content of different types. Moreover, a clever encoding rule is proposed for different content types, so that the information content under different content types can be encoded into an array that is more conducive to compression. In this way, the results generated after encoding for different content types can be compressed separately, thereby effectively improving the compression ratio of the comparison information.

[0098] FIG. 5 is a schematic flowchart of another method for compressing a sequencing data file provided by an exemplary embodiment of the present application. Referring to FIG. 5, the method may include steps 500 to 502:

[0099] Step 500: For any quality evaluation sequence in the file to be compressed, check whether there is a target evaluation value whose proportion exceeds a preset proportion threshold.

[0100] Step 501: If it exists, after filtering the quality evaluation sequence based on the target evaluation value, perform an encoding operation.

[0101] Step 502: Perform compression based on the encoded quality evaluation sequences in the file to be compressed.

[0102] In this embodiment, an improved scheme for compressing the quality evaluation sequence part in the file to be compressed is provided. This improved scheme can be combined with the improved scheme for compressing the sequencing sequence part in the foregoing embodiment to further improve the compression ratio of the file to be compressed.

[0103] Referring to FIG. 5, in step 500, for each quality evaluation sequence included in the file to be compressed, check whether there is a target evaluation value whose proportion exceeds a preset proportion threshold. Among them, the preset proportion threshold can be flexibly set as needed, for example, it can be set to 50% etc., and no limitation is made here.

[0104] For any quality evaluation sequence, if there is a target evaluation value therein, then in step 501, after filtering the quality evaluation sequence based on the target evaluation value, perform an encoding operation. In this way, the repeated content in the quality evaluation sequence can be removed, thereby reducing the amount of data to be compressed.

[0105] —An exemplary coding scheme can be: encoding the quality evaluation sequence into a fourth marker array, where the elements in the fourth marker array correspond one-to-one with the elements in the quality evaluation sequence. Among them, the elements in the quality evaluation sequence that are different from the target evaluation value are encoded as the third value in the fourth marker array, and the elements that are the same as the target evaluation value are encoded as the fourth value in the fourth marker array; the original element values of the elements in the quality evaluation sequence that are different from the target evaluation value are recorded in the first evaluation value array pointed to by the third value in order.

[0106] Of course, the target evaluation value can also be recorded in the second evaluation value array. Obviously, the second evaluation value array can only contain 1 element, and the value of the element is the target evaluation value.

[0107] In this exemplary coding scheme, it is proposed to encode each quality evaluation sequence included in the file to be compressed into a fourth marker array. Moreover, the fourth marker array only contains two element values, which effectively ensures that the fourth marker array can achieve an excellent compression ratio. The other evaluation values in the quality evaluation sequence that are different from the target evaluation value will be stored in the first evaluation value array. The inventor found in the research process that the proportion of the remaining evaluation values in the quality evaluation sequence that are different from the target evaluation value is usually relatively low. Therefore, the number of elements in the first evaluation value array will not be too large, usually much smaller than the number of elements in the quality evaluation sequence. In this way, for the quality evaluation sequence part divided from the file to be compressed, the amount of data to be compressed will be effectively reduced, and the array form is also more conducive to obtaining a higher compression ratio.

[0108] On this basis, continuing to refer to FIG. 5, in step 502, compression can be performed based on the encoded quality evaluation sequences in the file to be compressed. As described above, the encoded result obtained after encoding has a smaller data volume compared to the quality evaluation sequence. Moreover, the exemplary coding scheme proposed above encodes the quality evaluation sequence into an array form, which is also more convenient for compression, thereby improving the compression ratio for the quality evaluation sequence part cut out from the file to be compressed.

[0109] In addition, referring to FIG. 2b, in this embodiment, it is further proposed that if the file to be compressed contains the identification information corresponding to each sequencing sequence, the identification information is deleted from the file to be compressed. That is, there is no need to compress the identification information part cut out from the file to be compressed, which can further reduce the amount of data to be compressed.

[0110] Thus, in this embodiment, referring to FIG. 2b, the compression results obtained by performing compression on the quality evaluation sequence and the sequencing sequence respectively can be merged to obtain the compressed file corresponding to the file to be compressed. As mentioned above, in this embodiment, the compression ratios for the quality evaluation sequence and the sequencing sequence can be effectively improved. Therefore, the improvement in the compression ratios in both aspects, plus the compression ratio saved by deleting the identification information, can significantly improve the overall compression ratio of the file to be compressed.

[0111] Considering that the identification information contained in the file to be compressed is deleted during the compression process in this embodiment, it will result in the inability to restore the identification information after decompressing the compressed file subsequently. For this reason, this embodiment further proposes that after decompressing the compressed file, identification information can be regenerated for each of the decompressed sequencing sequences to distinguish different sequencing sequences. The inventors found during the research process that the identification information in the file to be compressed is usually used to distinguish different sequencing sequences and does not refer to information such as the order between sequencing sequences. Therefore, the method of regenerating identification information will not affect the use of the sequencing data generated after decompression.

[0112] In summary, in this embodiment, a compression improvement scheme for several other exemplary file parts cut out from the file to be compressed is proposed. By deleting and no longer compressing the identification information in the file to be compressed, a certain compression ratio can be improved. Also, by filtering and encoding the quality evaluation sequence part, the data volume of the result generated after encoding can be lower than the data volume of the quality evaluation sequence itself, and the result generated after encoding adopts an array form, which can further improve the compression ratio for the quality evaluation sequence part. In this way, the compression ratio of each file part cut out from the file to be compressed can be improved, thereby significantly improving the overall compression ratio of the file to be compressed.

[0113] FIG. 6 is a flowchart of a method for compressing a sequencing data file provided by another exemplary embodiment of the present application. Referring to FIG. 6, the method may include steps 600 to 603.

[0114] Step 600: Based on the reference segments matched by each sequencing sequence contained in the file to be compressed on the reference sequence, alignment information corresponding to each sequencing sequence is respectively constructed, and the alignment information contains various types of information content.

[0115] Step 601: Classify the information content in the alignment information generated for each sequencing sequence contained in the file to be compressed according to the content type.

[0116] Step 602: According to the correspondence between the content type and the encoding rule, encode the information content under different content types respectively using the corresponding encoding rules.

[0117] Step 603: Compress the results generated after encoding under different content types respectively to obtain the compression results of the sequencing sequences in the file to be compressed.

[0118] Among them, in this embodiment, the implementation method used to construct the alignment information corresponding to each sequencing sequence is not limited. An exemplary method may be: for any sequencing sequence in the file to be compressed, use at least two matching components to respectively match similar segments for the sequencing sequence on the reference sequence based on the subsequences extracted from the sequencing sequence, and the subsequences used by different matching components do not repeat each other; if at least two matching components match the same similar segment and the similar segment meets the preset similarity requirement, then determine the similar segment as the reference segment corresponding to the sequencing sequence; construct the alignment information between the sequencing sequence and the reference segment. Regarding this implementation method, reference can be made to the relevant descriptions in the foregoing embodiments and will not be elaborated here.

[0119] In this embodiment, other implementation methods can be supported to construct the alignment information corresponding to each sequencing sequence.

[0120] Regardless of the implementation method used to construct the alignment information corresponding to each sequencing sequence, in this embodiment, an improved scheme for performing the compression process on the alignment information generated under each sequencing sequence in the file to be compressed can be provided to improve the compression rate of the file to be compressed. Regarding the technical details of this improved scheme, reference can also be made to the relevant descriptions in the foregoing embodiments and will not be elaborated here.

[0121] In this embodiment, by improving the compression process of the alignment information generated under each sequencing sequence in the file to be compressed, the compression rate of the alignment information can be effectively improved, and thus the compression rate of the file to be compressed can be improved.

[0122] It should be noted that in some of the processes described in the above embodiments and the accompanying drawings, a plurality of operations appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different arrays, values, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types.

[0123] FIG. 7 is a schematic structural diagram of a computing device provided by another exemplary embodiment of the present application. As shown in FIG. 7, the computing device includes: a memory 70, a processor 71, and a communication component 72.

[0124] The processor 71 is coupled to the memory 70 and the communication component 72 and is configured to execute a computer program in the memory 70 for: for any sequencing sequence in the file to be compressed, using at least two matching components to perform a matching operation on the reference sequence in cooperation to determine a reference fragment for the sequencing sequence; constructing alignment information between the sequencing sequence and the reference fragment; and performing compression based on the alignment information generated for each sequencing sequence in the file to be compressed.

[0125] In an optional embodiment, when the processor 71 uses at least two matching components to perform a matching operation on the reference sequence in cooperation to determine a reference fragment for any sequencing sequence in the file to be compressed, it may specifically be configured to: use the at least two matching components to respectively match similar fragments for the sequencing sequence on the reference sequence based on subsequences extracted from the sequencing sequence, and the subsequences used by different matching components are non-repetitive; if the at least two matching components match the same similar fragment and the similar fragment meets the preset similarity requirement, then determine the similar fragment as the reference fragment corresponding to the sequencing sequence.

[0126] In an alternative embodiment, the processor 71 may further be configured to: if it is detected that the at least two matching components match similar segments at different positions, and / or the at least two matching components match the same similar segment but the similarity between the similar segment and the sequencing sequence does not meet the preset similarity requirement, then determine that the sequencing sequence is a variant sequence; and use local alignment technology to determine a reference segment on the reference sequence for each variant sequence included in the file to be compressed.

[0127] In an alternative embodiment, when it is determined that the sequencing sequence is a variant sequence, the at least two matching components no longer extract the remaining subsequences on the sequencing sequence.

[0128] In an alternative embodiment, sliding window tools are respectively arranged in the at least two matching components, and the processor 71 may further be configured to: allocate non-overlapping sequence segments to different sliding window tools on the sequencing sequence; control each sliding window tool to slide on the sequence segment allocated to itself to extract a subsequence that meets the target requirements; wherein the target requirements include that the sequence elements at the ends of the subsequence are of a preset type, and / or the sequence elements at the ends of the subsequence are different from their adjacent sequence elements.

[0129] In an alternative embodiment, the comparison information corresponding to a single sequencing sequence includes various types of information content. When the processor 71 performs compression based on the comparison information generated under each sequencing sequence in the file to be compressed, it may specifically be configured to: classify the information content in the comparison information generated under each sequencing sequence included in the file to be compressed according to the content type; encode the information content under different content types respectively using corresponding encoding rules according to the correspondence between the content type and the encoding rule; and compress the results generated after encoding under different content types respectively.

[0130] In an alternative embodiment, the content type includes a position type, and the position type information content is used to describe the position of the reference segment corresponding to the sequencing sequence. When the processor 71 classifies the information content in the comparison information generated under each sequencing sequence included in the file to be compressed according to the content type, it may specifically be configured to: reorder the comparison information under each sequencing sequence in the file to be compressed according to the reference segment position; in the target content type, query the information content belonging to the target content type from the comparison information corresponding to each sequencing sequence in turn according to the order between the reordered comparison information, so as to determine the information content under the target content type; wherein the target content type is any one of the content types.

[0131] In an alternative embodiment, if the file to be compressed is a sequencing data file generated based on single-end sequencing technology, when the processor 71 encodes the positional information content, it may specifically be configured to: encode each piece of positional information content into a first marker array, where the elements in the first marker array correspond one-to-one to the respective sequencing sequences; sequentially determine, for each sequencing sequence, whether the distance difference between the reference fragment position and the reference fragment position of the adjacent preceding sequencing sequence exceeds a preset distance threshold; sequentially record the positional information content corresponding to the sequencing sequences determined to exceed the preset distance threshold in a first position array, and mark the corresponding elements in the first marker array with values for pointing to the first position array; sequentially record the distance differences corresponding to the sequencing sequences determined not to exceed the preset distance threshold in a second position array, and mark the corresponding elements in the first marker array with values for pointing to the second position array.

[0132] In an alternative embodiment, if the file to be compressed is one of two sequencing data files generated for the same sample based on paired-end sequencing technology, a single sequencing data segment generated under paired-end sequencing technology contains a pair of sequencing sequences, which are stored one-to-one in the file to be compressed and the paired other sequencing data file, and the other sequencing data file has completed the encoding of the positional information content according to the encoding rules set for single-end sequencing technology, then when the processor 71 encodes the positional information content for the file to be compressed, it may specifically be configured to: after encoding each piece of positional information content into the first marker array, sequentially determine, for each sequencing sequence, the position difference value between the reference fragment position and the reference fragment position of the corresponding sequencing sequence in the other sequencing data file, whether it exceeds a preset difference threshold; sequentially record the positional information content corresponding to the sequencing sequences determined to exceed the preset difference threshold in the first position array, and mark the corresponding elements in the first marker array with values for pointing to the first position array; sequentially record the position difference values corresponding to the sequencing sequences determined not to exceed the preset difference threshold in the second position array, and mark the corresponding elements in the first marker array with values for pointing to the second position array.

[0133] In an alternative embodiment, the content type includes a chain attribute class. The information content of the chain attribute class is used to describe whether the sequencing sequence matches the positive strand or the negative strand in the reference sequence. If the file to be compressed is a sequencing data file generated based on single-end sequencing technology, then when the processor 71 encodes the information content of the chain attribute class, it can be specifically used for: encoding each piece of information content of the chain attribute class into a second marker array, where the elements in the second marker array correspond one by one to each sequencing sequence; marking the corresponding element of the sequencing sequence that matches the positive strand as a first value in the second marker array; and marking the corresponding element of the sequencing sequence that matches the negative strand as a second value in the second marker array.

[0134] In an alternative embodiment, if the file to be compressed is any one of two sequencing data files generated for the same sample based on paired-end sequencing technology, and a single sequencing data segment generated by paired-end sequencing technology contains a pair of sequencing sequences, which are stored one-to-one in the file to be compressed and the paired other sequencing data file, then when the processor 71 encodes the information content of the chain attribute class, it can be specifically used for: configuring a third marker array under the chain attribute class, where the elements in the third marker array correspond one by one to each sequencing data segment; for the first type of sequencing data segment containing sequencing sequences with different chain attributes, performing an exchange operation on the information content of the chain attribute class as needed to concentrate the same information content of the chain attribute class under the same sequencing data file, and marking the corresponding element of each first type of sequencing data segment as a target value in the third marker array; after completing the exchange operation, encoding each piece of information content of the chain attribute class into the second marker array and performing an element marking operation for the file to be compressed.

[0135] In an alternative embodiment, the content type includes a result class. The information content of the result class includes an array for marking the consistency of sequence elements and each sequence element marked as inconsistent. When the processor 71 encodes the information content of the result class, it can be specifically used for: sequentially storing each sequence element marked as inconsistent into a result array in the information content of the result class corresponding to a single sequencing sequence in the file to be compressed.

[0136] In an alternative embodiment, the processor 71 can also be used for: for any quality evaluation sequence in the file to be compressed, checking whether there is a target evaluation value whose proportion exceeds a preset proportion threshold; if so, filtering the quality evaluation sequence based on the target evaluation value and then performing an encoding operation; Performing compression based on the encoded quality evaluation sequences in the file to be compressed.

[0137] In an alternative embodiment, when the processor 71 performs an encoding operation after filtering the quality evaluation sequence based on the target evaluation value, it may specifically be used to: encode the quality evaluation sequence into a fourth marker array, where the elements in the fourth marker array correspond one-to-one to the elements in the quality evaluation sequence. Among them, the elements in the quality evaluation sequence that are different from the target evaluation value are encoded as a third value in the fourth marker array, and the elements that are the same as the target evaluation value are encoded as a fourth value in the fourth marker array; record the original element values of the elements in the quality evaluation sequence that are different from the target evaluation value in the first evaluation value array pointed to by the third value in order.

[0138] In an alternative embodiment, the processor 71 may also be used to: if the to-be-compressed file contains identification information corresponding to each sequencing sequence, delete the identification information from the to-be-compressed file; merge the compression results obtained by performing compression on the quality evaluation sequence and the sequencing sequence respectively to obtain the compressed file corresponding to the to-be-compressed file; after decompressing the compressed file, regenerate identification information for each decompressed sequencing sequence to distinguish different sequencing sequences.

[0139] Further, as shown in FIG. 7, the computing device further includes: a power supply component 73 and other components. Only some components are schematically shown in FIG. 7, which does not mean that the computing device only includes the components shown in FIG. 7.

[0140] It should be noted that for the technical details in the above embodiments of the computing device, reference may be made to the relevant descriptions in the foregoing method embodiments. For the sake of brevity, they will not be repeated here, but this should not cause loss of the protection scope of this application.

[0141] Correspondingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, and when the computer program is executed, it can implement each step that can be executed by the computing device in the above method embodiments.

[0142] The memory in FIG. 7 above is used to store computer programs and can be configured to store various other data to support operations on the computing platform. Examples of such data include instructions for any application or method operating on the computing platform, contact data, phone book data, messages, pictures, videos, etc. The memory can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0143] The communication component in FIG. 7 above is configured to facilitate communication, either wired or wireless, between the device in which the communication component is located and other devices. The device in which the communication component is located can access a wireless network based on a communication standard, such as WiFi, 2G, 3G, 4G / LTE, 5G and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0144] The power supply component in FIG. 7 above provides power to various components of the device in which the power supply component is located. The power supply component can include a power management system, one or more power supplies, and other components associated with generating, managing and distributing power for the device in which the power supply component is located.

[0145] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program code.

[0146] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, so that the instructions executed by the processors of the computer or other programmable data processing devices generate a device for implementing the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.

[0147] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.

[0148] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more flows in the flowchart and / or one or more blocks in the block diagram.

[0149] It should also be noted that the term "including", "comprising", or any other variation thereof is intended to cover non-exclusive inclusion, so that a process, method, commodity, or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or elements inherent to such process, method, commodity, or device. Without further limitation, an element defined by the statement "including a....." does not exclude the existence of additional identical elements in the process, method, commodity, or device including the said element. It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.

[0151] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included within the protection scope of the present application.

Claims

Claims 1. A method for compressing a sequencing data file, wherein, Including: For any sequencing sequence in the file to be compressed, at least two matching components are used to collaboratively perform a matching operation on a reference sequence to determine a reference fragment for the sequencing sequence; Construct alignment information between the sequencing sequence and the reference fragment; based on the alignment information generated for each sequencing sequence in the file to be compressed, perform compression.

2. The method according to claim 1, wherein, For any sequencing sequence in the file to be compressed, using at least two matching components to collaboratively perform a matching operation on a reference sequence to determine a reference fragment for the sequencing sequence, includes: using the at least two matching components to respectively match similar fragments for the sequencing sequence on the reference sequence based on subsequences extracted from the sequencing sequence, and the subsequences used by different matching components do not overlap; if the at least two matching components match the same similar fragment and the similar fragment meets the preset similarity requirement, then determine the similar fragment as the reference fragment corresponding to the sequencing sequence.

3. The method according to claim 2, wherein It further includes: if it is detected that the at least two matching components match similar fragments at different positions, and / or the at least two matching components match the same similar fragment but the similarity between the similar fragment and the sequencing sequence does not meet the preset similarity requirement, then determine the sequencing sequence as a variant type sequence; adopt a local alignment technique to determine a reference fragment for each variant type sequence included in the file to be compressed on the reference sequence; wherein, in the case of determining the sequencing sequence as a variant type sequence, the at least two matching components no longer extract the remaining subsequences on the sequencing sequence.

4. The method according to claim 2, wherein Sliding window tools are respectively set in the at least two matching components, and the method further includes: allocating non-overlapping sequence fragments for different sliding window tools on the sequencing sequence; controlling each sliding window tool to slide on the sequence fragment allocated to itself to extract subsequences that meet the target requirements; wherein, the target requirements include that the sequence elements at the ends of the subsequences are of a preset type, and / or the sequence elements at the ends of the subsequences are different from their adjacent sequence elements.

5. The method according to claim 1, wherein, The alignment information corresponding to a single sequencing sequence contains multiple types of information content. Based on the alignment information generated for each sequencing sequence in the file to be compressed, performing compression includes: classifying the information content in the alignment information generated for each sequencing sequence included in the file to be compressed according to the content type; according to the correspondence between the content type and the coding rule, respectively adopting the corresponding coding rule to code the information content under different content types; respectively compressing the results generated after coding under different content types.

6. The method according to claim 5, wherein The content type includes location types, and the location type information content is used to describe the reference fragment positions corresponding to the sequencing sequences; when classifying the information content in the alignment information generated for each sequencing sequence included in the file to be compressed according to the content type, it includes: reordering the alignment information for each sequencing sequence in the file to be compressed according to the reference fragment positions; under the target content type, querying the information content belonging to the target content type from the alignment information corresponding to each sequencing sequence in turn according to the order between the reordered alignment information, so as to determine the information content under the target content type; where the target content type is any one of the content types. If the file to be compressed is a sequencing data file generated based on single-end sequencing technology, the process of encoding the location type information content includes: encoding each location type information content into a first marker array, where the elements in the first marker array correspond one by one to the sequencing sequences; successively determining for each sequencing sequence whether the distance difference between the reference fragment position and the reference fragment position of the adjacent previous sequencing sequence exceeds a preset distance threshold; successively recording the location type information content corresponding to the sequencing sequences determined to exceed the preset distance threshold in a first position array, and marking the corresponding elements in the first marker array with values used to point to the first position array; successively recording the distance differences corresponding to the sequencing sequences determined not to exceed the preset distance threshold in a second position array, and marking the corresponding elements in the first marker array with values used to point to the second position array.

7. The method according to claim 6, wherein, The content type includes strand attribute types, and the strand attribute type information content is used to describe whether the sequencing sequence matches the positive strand or the negative strand in the reference sequence. If the file to be compressed is a sequencing data file generated based on single-end sequencing technology, the process of encoding the strand attribute type information content includes: encoding each strand attribute type information content into a second marker array, where the elements in the second marker array correspond one by one to the sequencing sequences; marking the corresponding elements of the sequencing sequences that match the positive strand with a first value in the second marker array; marking the corresponding elements of the sequencing sequences that match the negative strand with a second value in the second marker array.

8. The method according to claim 5 or 6, wherein The content type includes result types, and the result type information content includes an array for marking the consistency of sequence elements and each sequence element marked as inconsistent. The process of encoding the result type information content includes: sequentially storing each sequence element marked as inconsistent into a result array in the result type information content corresponding to a single sequencing sequence in the file to be compressed.

9. The method according to claim 5 or 6, wherein ​ 10. The method according to claim 1, wherein, It further includes: for any quality evaluation sequence in the file to be compressed, searching for whether there is a target evaluation value whose proportion exceeds a preset proportion threshold; if there is, then after filtering the quality evaluation sequence based on the target evaluation value, performing an encoding operation; based on each encoded quality evaluation sequence in the file to be compressed, performing compression.

11. The method according to claim 10, wherein After filtering the quality evaluation sequence based on the target evaluation value and then performing an encoding operation, it includes: encoding the quality evaluation sequence into a fourth marker array, where the elements in the fourth marker array correspond one-to-one to the elements in the quality evaluation sequence. Among them, each element in the quality evaluation sequence that is different from the target evaluation value is encoded as a third value in the fourth marker array, and each element that is the same as the target evaluation value is encoded as a fourth value in the fourth marker array; the original element values of each element in the quality evaluation sequence that is different from the target evaluation value are sequentially recorded in the first evaluation value array pointed to by the third value.

12. The method according to claim 10, wherein It further includes: if the file to be compressed contains identification information corresponding to each sequencing sequence, then deleting the identification information from the file to be compressed; combining the compression results obtained by performing compression on the quality evaluation sequence and the sequencing sequence respectively to obtain the compressed file corresponding to the file to be compressed; after decompressing the compressed file, regenerating identification information for each decompressed sequencing sequence to distinguish different sequencing sequences.

13. A method for compressing a sequencing data file, wherein, It includes: Based on the reference fragments matched by each sequencing sequence included in the file to be compressed on the reference sequence, constructing the alignment information corresponding to each sequencing sequence, where the alignment information contains various types of information content; classifying the information content in the alignment information generated for each sequencing sequence included in the file to be compressed according to the content type; according to the correspondence between the content type and the encoding rule, respectively adopting the corresponding encoding rule to encode the information content under different content types; respectively compressing the results generated after encoding under different content types to obtain the compression result of the sequencing sequence in the file to be compressed.

14. A computing device, wherein, It includes a memory and a processor; the memory is used to store one or more computer instructions; the processor is coupled to the memory and is used to execute the one or more computer instructions to execute the compression method of the sequencing data file according to any one of claims 1-13.

15. A computer-readable storage medium storing computer instructions, wherein, When the computer instructions are executed by one or more processors, it causes the one or more processors to execute the compression method of the sequencing data file according to any one of claims 1-13.

16. A computer program product, wherein, It includes a computer program / instructions, where when the computer program is executed by a processor, it causes the processor to implement the compression method of the sequencing data file according to any one of claims 1-13.

Citation Information

Patent Citations

  • Base sequence coding method and system in FASTQ file compression

    CN112102883A

  • Batch distributed compression method based on FASTQ gene big data

    CN113268459A

  • Apparatus and methods for genomic sequencing

    CN114787930A

  • Gene sequence processing method and system

    CN116312799A

  • Systems and Methods for Compressing Genetic Sequencing Data and Uses Thereof

    US20200058379A1