Field compression method and device, processor and electronic equipment
By dividing and separating the target fields of gene sequencing data, the problems of low field compression efficiency and lack of scalability in the prior art are solved, and a more efficient field compression effect is achieved.
Patent Information
- Application Number
- CN202411983423.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-02
AI Technical Summary
In the prior art, the field compression efficiency of gene sequencing data is low and does not have scalability, resulting in the compression being limited by stand-alone resources.
By obtaining the target field in the target file and dividing it into three categories: the first target field representing the quality metric of gene data sequence alignment, the second target field associated with the sequence alignment, and the third target field of other fields. Each type of field is compressed separately to obtain the respective compression results.
By separating and compressing different target fields, the scalability of the compressed target files is ensured, and the problem of field compression is limited by stand-alone resources is avoided, thereby improving the efficiency of field compression.
Smart Images

Figure CN119920330A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of data processing, and in particular to a field compression method, device, processor and electronic device. Background Art
[0002] As the amount of gene sequencing data grows exponentially, it is necessary to quickly compress the gene sequencing data. However, in related technologies, thread pools are often introduced to quickly compress gene sequencing data. This compression method is limited by single-machine resources and lacks scalability, resulting in low efficiency of field compression.
[0003] Currently, no effective solution has been proposed to the technical problem of low efficiency of the above-mentioned field compression. Summary of the invention
[0004] The embodiments of the present invention provide a field compression method, device, processor and electronic device to at least solve the technical problem of low efficiency of field compression.
[0005] According to one aspect of an embodiment of the present invention, a field compression method is provided, the method comprising: obtaining a target file, wherein the target file comprises a target field associated with genetic data; dividing the target field in the target file to obtain a first target field, a second target field and a third target field, wherein the first target field is used to represent a quality indicator of a sequence alignment of the genetic data, the second target field is used to represent a field associated with the sequence alignment, and the third target field is used to represent a field in the target field other than the first target field and the second target field; compressing the first target field to obtain a first compression result of the first target field, compressing the second target field to obtain a second compression result of the second target field, and compressing the third target field to obtain a third compression result of the third target field.
[0006] Optionally, compressing the first target field to obtain a first compression result of the first target field includes: respectively determining the scores of multiple first target subfields in the first target field; sorting the multiple first target subfields based on the score of each first target subfield; and compressing the sorted multiple first target subfields to obtain a first compression result.
[0007] Optionally, based on the score of each first target subfield, multiple first target subfields are sorted, including: in response to the existence of the same score in the scores of the multiple first target subfields, constructing the multiple first target subfields corresponding to the same score into corresponding graphs; determining the similarity between the multiple graphs corresponding to the multiple first target subfields; and sorting the multiple first target subfields based on the similarity.
[0008] Optionally, sorting the multiple first target subfields based on the similarity includes: sorting the similarity; and sorting the multiple first target subfields according to the sorted similarity.
[0009] Optionally, compressing the second target field to obtain a second compression result of the second target field includes: constructing a target sequence corresponding to the second target field; and compressing the second target field using the target sequence to obtain a second compression result.
[0010] Optionally, compressing the third target field to obtain a third compression result of the third target field includes: segmenting the third target field; in response to the segmented third target field including a target element, determining a change state of the third target field; determining a coding strategy that matches the change state, wherein the coding strategy is used to control encoding of the third target field; encoding the third target field according to the coding strategy, and compressing the encoded third target field to obtain a third compression result.
[0011] According to one aspect of an embodiment of the present invention, a field compression device is provided, which may include: an acquisition unit, used to acquire a target file, wherein the target file includes a target field associated with genetic data; a division unit, used to divide the target field in the target file to obtain a first target field, a second target field and a third target field, wherein the first target field is used to represent a quality indicator of a sequence alignment of the genetic data, the second target field is used to represent a field associated with the sequence alignment, and the third target field is used to represent a field in the target field other than the first target field and the second target field; a compression unit, used to compress the first target field to obtain a first compression result of the first target field, compress the second target field to obtain a second compression result of the second target field, and compress the third target field to obtain a third compression result of the third target field.
[0012] According to another aspect of an embodiment of the present invention, a processor is further provided, wherein the processor is used to run a program, wherein when the program is run by the processor, the field compression method in the embodiment of the present invention is executed.
[0013] According to another aspect of an embodiment of the present invention, there is further provided an electronic device, comprising: a memory storing an executable program; and a processor for running the program, wherein the field compression method in each embodiment of the present invention is executed when the program is running.
[0014] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the field compression method in the present invention.
[0015] According to another aspect of the embodiments of the present invention, a computer program product is provided. The computer program product includes a computer program. When the computer program is executed by a processor, the field compression method in the embodiments of the present invention is implemented.
[0016] According to another aspect of an embodiment of the present invention, a computer program product is also provided, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the field compression method in the embodiment of the present invention is implemented.
[0017] According to another aspect of an embodiment of the present invention, an embodiment of the present application further provides a computer program, which implements the field compression method in the above-mentioned embodiment of the present invention when the computer program is executed by a processor.
[0018] In an embodiment of the present invention, when compressing a field, a target file can be obtained, and the target field in the obtained target file can be divided to obtain a first target field, a second target field, and a third target field. By compressing the first target field obtained by division, a first compression result of the first target field can be obtained, by compressing the second target field obtained by division, a second compression result of the second target field can be obtained, and by compressing the third target field obtained by division, a third compression result of the third target field can be obtained. Since the different target fields obtained by division are compressed separately, the scalability of the compressed target file can be guaranteed, thereby achieving the purpose of avoiding the field compression being limited by the resources of a single machine, thereby solving the technical problem of low efficiency of field compression, and further achieving the technical effect of improving the efficiency of field compression. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:
[0020] Figure 1 is a flow chart of a field compression method according to an embodiment of the present invention;
[0021] FIG. 2( a ) is a flow chart of a method for compressing a QUAL field according to an embodiment of the present invention;
[0022] FIG2( b ) is a flow chart of a method for compressing a SEQ field according to an embodiment of the present invention;
[0023] FIG2( c ) is a schematic diagram of a load balancing scheduling and monitoring architecture for massive SAM compression tasks according to an embodiment of the present invention;
[0024] FIG2( d ) is a schematic diagram of a Flink-based SAM streaming parallel compression architecture according to an embodiment of the present invention;
[0025] Figure 3 It is a schematic diagram of a field compression device according to an embodiment of the present invention. DETAILED DESCRIPTION
[0026] In order to enable those skilled in the art to better understand the scheme of the present invention, the technical scheme in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only embodiments of a part of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0027] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units that are clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.
[0028] According to an embodiment of the present invention, a method for compressing a field is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.
[0029] Figure 1 is a flowchart of a field compression method according to an embodiment of the present invention, the method may include the following steps:
[0030] Step S101, obtaining a target file.
[0031] In the technical solution provided in step S101 of the present invention, the target file may include a target field associated with the gene data. For example, the target file may be a file storing the target field associated with the gene data in a sequence alignment and mapping (SAM) format.
[0032] In this embodiment, the above-mentioned gene data may be generated by a gene sequencing machine, and the above-mentioned target field associated with the gene data may be used to store base information, reference sequence name, and difference information between comparison fragments.
[0033] In this embodiment, the target file is obtained. Optionally, this embodiment uses a gene sequencing machine to generate gene data. The generated gene data is stored in the initial file in SAM format to obtain the target file, thereby achieving the purpose of obtaining the target file.
[0034] Step S102, dividing the target field in the target file to obtain a first target field, a second target field and a third target field.
[0035] In the technical solution provided in the above step S102 of the present invention, the above first target field can be used to represent the quality index of the sequence alignment of the genetic data, the above second target field can be used to represent the field associated with the sequence alignment, and the above third target field can be used to represent the fields in the target field except the first target field and the second target field.
[0036] In this embodiment, after acquiring the target file, the target field in the target file is divided to obtain the first target field, the second target field and the third target field. Optionally, based on the acquisition of the target file, this embodiment divides the target field in the target file according to the types of different target fields to obtain the first target field, the second target field and the third target field. That is to say, if the type of a certain target field is a quality indicator type, the target field is divided into the first target field, if the type of a certain target field is a field type, the target field is divided into the second target field, and if the type of a certain target field is neither a field type nor a quality indicator type, the target field is divided into the third target field.
[0037] Step S103, compressing the first target field to obtain a first compression result of the first target field, compressing the second target field to obtain a second compression result of the second target field, and compressing the third target field to obtain a third compression result of the third target field.
[0038] In the technical solution provided in step S103 of the present invention, the prediction result can be used to indicate the similarity between the current behavior and the target behavior associated with the target feature. The target behavior can be a pre-set behavior, for example, the pre-set behavior can be a smoking behavior, which is only used as an example and is not specifically limited.
[0039] In this embodiment, after dividing the target field in the target file to obtain the first target field, the second target field and the third target field, the first target field is compressed to obtain a first compression result of the first target field. Optionally, on the basis of obtaining the first target field, the second target field and the third target field, this embodiment reorders the first target field obtained by division, and compresses the reordered first target field to obtain a first compression result of the first target field. The reordering operation may include a reordering operation based on frequency score and a reordering operation based on graph similarity.
[0040] In this embodiment, after dividing the target field in the target file to obtain the first target field, the second target field and the third target field, the second target field is compressed to obtain a second compression result of the second target field. Optionally, on the basis of obtaining the first target field, the second target field and the third target field, this embodiment encodes the second target field obtained by division, and compresses the encoded second target field to obtain a second compression result of the second target field. The encoding operation may include a difference encoding operation, a run-length encoding operation, a differential encoding operation and a matching encoding operation based on a Bloom filter index.
[0041] In this embodiment, after dividing the target field in the target file to obtain the first target field, the second target field and the third target field, the third target field is compressed to obtain the third compression result of the third target field. Optionally, on the basis of obtaining the first target field, the second target field and the third target field, the embodiment can determine the characteristics of each field in the third target field, wherein the characteristics of each field can be used to represent the format and characteristics of each field. According to the characteristics of each field, a compression algorithm matching the characteristics of each field can be selected. Using the selected compression algorithm, each field in the third target field is compressed respectively, and the compression results of each field in the third target field can be obtained, and the compression results of each field in the third target field are used as the third compression results of the third target field.
[0042] In the above steps S101 to S103 of the present application, when compressing fields, a target file can be obtained, and the target fields in the obtained target file can be divided to obtain a first target field, a second target field and a third target field. By compressing the above first target field obtained by division, a first compression result of the first target field can be obtained, by compressing the above second target field obtained by division, a second compression result of the second target field can be obtained, and by compressing the above third target field obtained by division, a third compression result of the third target field can be obtained. Since the different target fields obtained by division are compressed separately, the scalability of the compressed target file can be guaranteed, thereby achieving the purpose of avoiding the field compression being limited by the resources of a single machine, thereby solving the technical problem of low efficiency of field compression, and further achieving the technical effect of improving the efficiency of field compression.
[0043] The above method of this embodiment is further introduced below.
[0044] As an optional implementation method, step S103 compresses the first target field to obtain a first compression result of the first target field, including: respectively determining the scores of multiple first target subfields in the first target field; sorting the multiple first target subfields based on the score of each first target subfield; and compressing the sorted multiple first target subfields to obtain a first compression result.
[0045] In this embodiment, the first target field may be a quality (QUAL) field, wherein the first target subfield may be used to represent each row of fields in the QUAL field, that is, each row of QUAL sequences; and the score may be used to represent the frequency score of each row of QUAL sequences.
[0046] In this embodiment, after dividing the target field in the target file to obtain the first target field, the second target field and the third target field, the scores of the multiple first target subfields in the first target field are determined respectively. Optionally, based on the first target field, the second target field and the third target field, the embodiment can determine the frequency of occurrence of characters in the multiple first target subfields in the first target field, and score the frequency of occurrence of characters in the multiple first target subfields obtained to obtain the scores of the multiple first target subfields.
[0047] In this embodiment, after determining the scores of the plurality of first target subfields in the first target field, the plurality of first target subfields are sorted based on the score of each first target subfield. Optionally, based on determining the scores of the plurality of first target subfields, this embodiment sorts the plurality of first target subfields in ascending order according to the determined score of each first target subfield.
[0048] In this embodiment, after sorting the plurality of first target subfields based on the score of each first target subfield, the sorted plurality of first target subfields are compressed to obtain a first compression result. Optionally, this embodiment performs field compression on the sorted plurality of first target subfields based on sorting the first target subfields to obtain a first compression result of the first target field.
[0049] As an optional implementation method, multiple first target subfields are sorted based on the score of each first target subfield, including: in response to the existence of the same score in the scores of the multiple first target subfields, the multiple first target subfields corresponding to the same score are constructed into corresponding graphs; the similarity between the multiple graphs corresponding to the multiple first target subfields is determined; and based on the similarity, the multiple first target subfields are sorted.
[0050] In this embodiment, the above-mentioned graph may include a directed graph and an undirected graph.
[0051] In this embodiment, after determining the scores of the plurality of first target subfields in the first target field, in response to the existence of the same score in the scores of the plurality of first target subfields, the plurality of first target subfields corresponding to the same score are constructed into a corresponding graph. Optionally, based on determining the scores of the plurality of first target subfields, this embodiment determines whether the same score exists in the scores of the plurality of first target subfields determined, and if it is determined that the same score exists in the scores of the plurality of first target subfields determined, the plurality of first target subfields corresponding to the same score are modeled, thereby constructing a corresponding graph.
[0052] In this embodiment, in response to the existence of the same score in the scores of the plurality of first target subfields, after the plurality of first target subfields corresponding to the same score are constructed into corresponding graphs, the similarity between the plurality of graphs corresponding to the plurality of first target subfields is determined; and the plurality of first target subfields are sorted based on the similarity. Optionally, on the basis of constructing the plurality of first target subfields corresponding to the same score into corresponding graphs, the embodiment can determine the structural similarity of the plurality of graphs corresponding to the plurality of first target subfields. Based on the determined similarity, the plurality of first target subfields are sorted in ascending order, thereby achieving the purpose of sorting the first target subfields.
[0053] As an optional implementation manner, sorting the multiple first target subfields based on similarity includes: sorting the similarity; and sorting the multiple first target subfields according to the sorted similarity.
[0054] In this embodiment, after determining the similarities between the multiple graphs corresponding to the multiple first target subfields, the similarities are sorted; and the multiple first target subfields are sorted according to the sorted similarities. Optionally, this embodiment, based on determining the similarities between the multiple graphs corresponding to the multiple first target subfields, sorts the determined similarities, and sorts the multiple first target subfields in the first target field in ascending order according to the sorted similarities. For example, in the ascending order, if the similarities between two or several graphs are similar, the two or several graphs are arranged together.
[0055] As an optional implementation method, step S103, compressing the second target field to obtain a second compression result of the second target field, includes: constructing a target sequence corresponding to the second target field; using the target sequence, compressing the second target field to obtain a second compression result.
[0056] In this embodiment, the second target field may include a sequence (SEQ) field, a name (RNAME) field, a position (POS) field, and a compact (Compact Idiosyncratic Gapped Alignment Report, referred to as CIGAR field) field, etc. The target sequence may be a hypothetical reference sequence corresponding to the second target field.
[0057] In this embodiment, after dividing the target field in the target file to obtain the first target field, the second target field and the third target field, a target sequence corresponding to the second target field is constructed; the second target field is compressed using the target sequence to obtain a second compression result. Optionally, on the basis of obtaining the first target field, the second target field and the third target field, this embodiment can construct a target sequence corresponding to the second target field, encode the second target field using the constructed target sequence, and compress the encoded second target field to obtain a second compression result for the second target field.
[0058] Optionally, in the second target field, based on constructing a hypothetical reference sequence, a CIGAR field is introduced for the second target field available to the RNAME field, and differential encoding is performed on the second target field available to the RNAME field. The second target field available to the RNAME field after differential encoding is compressed, and a first initial compression result of the second target field available to the RNAME field can be obtained.
[0059] At the same time, for the second target field where the RNAME field is unavailable, a matching coding algorithm based on a Bloom filter index is used to match code the second target field where the RNAME field is unavailable, and the second target field where the RNAME field is unavailable after matching coding is compressed, so as to obtain a second initial compression result of the second target field where the RNAME field is unavailable; and, the RNAME field, the CIGAR field, and the POS field in the second target field are run-length coded or differentially coded, and the RNAME field, the CIGAR field, and the POS field after run-length coding or differential coding are compressed, so as to obtain a third initial compression result of the RNAME field, a fourth initial compression result of the CIGAR field, and a fifth initial compression result of the POS field. In addition, the second target field where the RNAME field is unavailable after matching coding is compressed, so as to obtain a sixth initial compression result of the second target field where the RNAME field is unavailable.
[0060] Furthermore, the obtained first initial compression result, second initial compression result, third initial compression result, fourth initial compression result, fifth initial compression result and sixth initial compression result are used as the second compression result of the second target field.
[0061] As an optional implementation method, step S103, compressing the third target field to obtain a third compression result of the third target field, includes: segmenting the third target field; in response to the segmented third target field including the target element, determining a change state of the third target field; determining a coding strategy that matches the change state; encoding the third target field according to the coding strategy, compressing the encoded third target field, and obtaining a third compression result.
[0062] In this embodiment, the third target field may be the remaining fields in the target field except the QUAL field and the SEQ field. The target element may be a number. The change state may be a state where no change occurs or a state where a change occurs.
[0063] In this embodiment, after dividing the target field in the target file to obtain the first target field, the second target field and the third target field, the third target field is segmented; in response to the segmented third target field including the target element, the change state of the third target field is determined. Optionally, on the basis of obtaining the first target field, the second target field and the third target field, this embodiment segments the obtained third target field, determines the relationship between the segmented third target field and the target element, and if it is determined that the segmented third target field includes the target element, the change state of the third target field can be determined to be a changed state.
[0064] Optionally, if it is determined that the segmented third target field does not include the target element, it may be determined that the change state of the third target field is a state of no change.
[0065] In this embodiment, the above encoding strategy may be used to control encoding of the third target field.
[0066] In this embodiment, in response to the third target field after segmentation including the target element, after determining the change state of the third target field, a coding strategy matching the change state is determined; according to the coding strategy, the third target field is encoded, and the encoded third target field is compressed to obtain a third compression result. Optionally, based on the determination of the change state of the third target field, if the change state of the third target field is a changed state, the coding strategy matching the change state may be a differential coding strategy, and according to the differential coding strategy, the third target field is differentially encoded, and the differentially encoded third target field is compressed to obtain a third compression result. Alternatively, if the change state of the third target field is a state that does not change, the coding strategy matching the change state may be a run-length coding strategy, and according to the run-length coding strategy, the third target field is run-length encoded, and the run-length encoded third target field is compressed to obtain a third compression result.
[0067] In an embodiment of the present invention, when compressing a field, a target file can be obtained, and the target field in the obtained target file can be divided to obtain a first target field, a second target field, and a third target field. By compressing the first target field obtained by division, a first compression result of the first target field can be obtained, by compressing the second target field obtained by division, a second compression result of the second target field can be obtained, and by compressing the third target field obtained by division, a third compression result of the third target field can be obtained. Since the different target fields obtained by division are compressed separately, the scalability of the compressed target file can be guaranteed, thereby achieving the purpose of avoiding the field compression being limited by the resources of a single machine, thereby solving the technical problem of low efficiency of field compression, and further achieving the technical effect of improving the efficiency of field compression.
[0068] The technical solution of the embodiment of the present invention is illustrated below in conjunction with preferred implementation modes.
[0069] As the amount of gene sequencing data grows exponentially, it is necessary to quickly compress the gene sequencing data. However, in related technologies, thread pools are often introduced to quickly compress gene sequencing data. This compression method is limited by single-machine resources and lacks scalability, resulting in low efficiency of field compression.
[0070] In order to solve the above technical problems, an embodiment of the present invention proposes a method for compressing a field. When compressing a field, a target file can be obtained, and the target field in the obtained target file can be divided to obtain a first target field, a second target field, and a third target field. By compressing the first target field obtained by division, a first compression result of the first target field can be obtained, by compressing the second target field obtained by division, a second compression result of the second target field can be obtained, and by compressing the third target field obtained by division, a third compression result of the third target field can be obtained, thereby achieving the purpose of avoiding the field compression being limited by the resources of a single machine, thereby solving the technical problem of low efficiency of field compression, and further achieving the technical effect of improving the efficiency of field compression.
[0071] In this embodiment, the compression method of the QUAL field is executed to compress the QUAL field to obtain a compression result of the QUAL field. For example, FIG2(a) is a flowchart of a compression method of the QUAL field according to an embodiment of the present invention. As shown in FIG2(a), the method may include the following steps:
[0072] Step S201, obtaining the QUAL field.
[0073] In the technical solution provided in the above step S201 of the present invention, the fields of the SAM file are split to obtain the QUAL field, the SEQ field and the remaining fields.
[0074] After the QUAL field is obtained, the process proceeds to step S202 to record the original order of the QUAL field.
[0075] In the technical solution provided in the above step S202 of the present invention, since reordering is introduced, the original order of the QUAL field needs to be recorded before the ordering, thereby ensuring that the original file can be completely restored during the decompression process.
[0076] After recording the original order of the QUAL fields, the process proceeds to step S203 to reorder the QUAL fields based on the frequency scores.
[0077] In the technical solution provided in the above step S203 of the present invention, in the actual sorting process, the QUAL field is sorted in ascending order according to the frequency score of each row of the QUAL sequence.
[0078] That is to say, for each row of QUAL sequence, a score representing the frequency of the characters in the QUAL sequence can be generated, and then the QUAL field is sorted in ascending order according to the frequency score. Among them, rows with the same or similar character frequencies will be gathered together as much as possible. When calculating the frequency score of each row, first divide the interval [33, 104] evenly into n equal parts (n = 4 by default), generating n small intervals, denoted as [S i ,E i ],i∈[1,n]. For each line of QUAL sequence, count the frequency of all characters in each small interval to get n frequency values, denoted as P i ,i∈[1,n]. Assuming the total number of times is 100, the frequency value P i Multiply by 100 and round down to a two-digit frequency value L i According to the following formula (1) and formula (2), the final frequency score C is an integer string with a length of 2n:
[0079] C i =L i %10 (1)
[0080] C n+i =L i / 10 (2)
[0081] Among them, C i It can be used to represent the remainder obtained by taking the modulus of the frequency value, C n+i Can be used to express relative multiples of frequency values.
[0082] It should be noted that the advantage of the above frequency scoring algorithm is its fast operation speed. However, in order to avoid the length of the frequency scoring result being too long, the number of intervals n is usually less than 10, and a large number of identical frequency scoring results may appear in the entire QUAL sequence. In order to improve the local redundancy of the sorted files, when the frequency scores are the same, the QUAL fields are reordered according to the similarity of the graphs corresponding to the sequences.
[0083] Alternatively, after recording the original order of the QUAL fields, the process proceeds to step S204 to reorder the QUAL fields based on the frequency scores.
[0084] In the technical solution provided in the above step S204 of the present invention, the QUAL fields are sorted in ascending order according to the structural similarities between the graphs.
[0085] In this embodiment, the above-mentioned graph can be a set of vertices and edges, usually saved in the form of an adjacency matrix or an adjacency list. Graph similarity can be used to indicate the degree of structural similarity between two graphs. For any QUAL sequence, a graph is constructed according to a specified method, and its character distribution is modeled so that the QUAL sequence is converted from a one-dimensional to a directed graph. A row of QUAL sequence can correspond to a similarity, which is obtained by calculating the similarity of the graph corresponding to the high-frequency sequence and the current sequence. Finally, according to the similarity of the QUAL sequence, the QUAL field is sorted in ascending order, thereby achieving the purpose of arranging QUAL sequences with similar distributions together.
[0086] After the QUAL fields are reordered, the process proceeds to step S205 and step S206 to record the final order of the reordered QUAL fields and to adaptively encode the reordered QUAL fields.
[0087] In the technical solution provided in the above step S206 of the present invention, adjacent quality scores can be encoded into one character by adaptively encoding the reordered QUAL field.
[0088] After recording the final order of the reordered QUAL fields, the process proceeds to step S207 to perform differential encoding on the reordered QUAL fields.
[0089] In the technical solution provided in step S207 of the present invention, the differential coding can be used in the form of A(i+1)-A(i) to save the differences between adjacent elements of the sequence, so as to be suitable for processing digital data with small element differences. In most cases, the adjacent sequence numbers generated by reordering have the same number of digits, and the length of the numbers can be significantly reduced after differential coding. For example, in the optimal case, only one digit is needed to save the next sequence. The result of differential coding will eventually be compressed using a universal compressor, thereby achieving the purpose of reducing data redundancy.
[0090] After adaptively encoding the reordered QUAL field and differentially encoding the reordered QUAL field, the process proceeds to step S208 to compress the encoded QUAL field to obtain a first compression result.
[0091] In this embodiment, the compression method of the SEQ field is executed to compress the SEQ field to obtain a compression result of the SEQ field. For example, FIG2(b) is a flowchart of a compression method of the SEQ field according to an embodiment of the present invention. As shown in FIG2(b), the method may include the following steps:
[0092] Step S211, obtaining the SEQ field, RNAME field, POS field, CIGAR field, etc.
[0093] After obtaining the SEQ field, the RNAME field, the POS field, the CIGAR field, etc., the process proceeds to step S212 to construct a hypothetical reference sequence.
[0094] After constructing the hypothetical reference sequence, the process proceeds to step S213, where the CIGAR field is introduced to perform differential encoding on the SEQ field available in the RNAME field.
[0095] In the technical solution provided by the above step S213 of the present invention, all the information such as SEQ field, RNAME field, POS field and CIGAR field are read out from the original SAM gene sequence, and the association relationship between each domain is maintained. In order to ensure that the order of the SEQ sequence after decompression is completely consistent with the original file, before constructing the hypothetical reference sequence, it is first necessary to mark the original order of each row. The 4 associated fields are sorted according to the RNAME field and the POS field. If the RNAME field is "*", it is arranged to the front. Adjacent SEQ sequences with overlapping fragments form a hypothetical reference sequence according to the principle of minority obeys majority, and the offset (Offset) address of the SEQ sequence relative to the hypothetical reference sequence and the final arrangement order information of each row are recorded.
[0096] It should be noted that the CIGAR field can include the difference information of the SEQ sequence relative to the original reference sequence. Since the constructed hypothetical reference sequence is sorted according to the RNAME field and the POS field and follows the principle of minority obeys majority, the hypothetical reference sequence and the original reference sequence will have some similar fragments. When the SEQ sequence can be restored by the hypothetical reference sequence, the Offset information and the CIGAR field, the difference information is actually represented in the CIGAR field, and there is no need to record additional difference information.
[0097] In this embodiment, the amount of data required for encoding is reduced by introducing the CIGAR field, and each SEQ sequence encoding result is merged into one line using a unified format. When differential encoding is performed, the SEQ field is mapped to the assumed reference sequence according to the offset address of the record, and the SEQ field is attempted to be restored in combination with the CIGAR field and the assumed reference sequence. Traverse the CIGAR field, and for the "M" and "=" characters representing the match, copy the bases of the same length in the assumed reference sequence. The "D" character can be used to indicate deletion, and the "N" character can be used to indicate skipping. The fragments represented by the above two marks exist in the reference sequence, but do not exist in the SEQ field, so the bases of the same length in the assumed reference sequence are directly skipped. The fragments marked by "H" and "P" themselves are not in the assumed reference sequence and SEQ sequence, so they can be directly ignored. "I" can be used to indicate insertion, and "S" can be used to indicate soft cutting. The fragments represented by the above two marks do not exist in the reference sequence, but exist in the SEQ field, so they are directly returned and cannot be restored. "X" indicates a complete mismatch, so it is directly returned and cannot be restored.
[0098] Optionally, if the recovery is successful, a carriage return is added to the output stream buffer, indicating that the complete difference information can be represented by the CIGAR field without adding additional information. Since the "M" type match may include mismatched characters, the inserted characters are not included in the reference sequence, etc., there may be situations where full recovery is not possible. If the recovery fails, the two sequences are traversed through the double pointer, and when different characters are encountered,<len,str> The length of the same segment and the mismatched characters are saved in the form of . In order to further reduce the amount of data, when there are continuous mismatches, run-length encoding is used to process str. After the difference encoding, the encoding result and the hypothetical reference sequence are compressed by the Prediction by Partial Matching algorithm (referred to as PPMD).
[0099] After constructing the hypothetical reference sequence, the process proceeds to step S214 to perform run-length encoding or differential encoding on the RNAME field, the POS field, and the CIGAR field.
[0100] In the technical solution provided in step S214 of the present invention, after the RNAME field is sorted, the same RNAME is grouped together, so it is suitable to be processed using run-length encoding. Finally, the RNAME field is encoded into<name,len> In the form of. Among them, name can be used to indicate the newly appeared RNAME field, and len can be used to indicate the number of consecutive appearances of this RNAME. The POS field can be used to indicate the starting address of the SEQ field relative to the reference sequence when generating the SAM file. After sorting, the POS of the same RNAME field are grouped together and show an increasing trend, so it is suitable to be processed using differential coding. However, it should be noted that differential coding should be performed when the RNAME fields are the same. When a new RNAME is encountered, the POS corresponding to the new RNAME field is selected as the first digit, and the differential operation is restarted. The CIGAR field can be composed of consecutive numbers and letters, and can be compressed directly using PPMD, and differential coding can be used to compress the final arrangement order.
[0101] It should be noted that the hypothetical reference sequence constructed in the previous text is used as the reference sequence, the SEQ fields where the RNAME field is unavailable are spliced together in order as the sequence to be compressed, and the sequence to be compressed is matched with the hypothetical reference sequence to reduce the amount of data.
[0102] After constructing the hypothetical reference sequence, the process proceeds to step S215, where a matching coding algorithm based on a Bloom filter index is used to perform matching coding on the SEQ field where the RNAME field is unavailable.
[0103] After the encoding is completed, enter step S216, step S217 and step S218, compress the SEQ field available to the RNAME field after difference encoding, and obtain the first initial compression result of the SEQ field available to the RNAME field; compress the RNAME field after run-length encoding or differential encoding, and obtain the third initial compression result of the RNAME field; compress the CIGAR field after run-length encoding or differential encoding, and obtain the fourth initial compression result of the CIGAR field; and compress the POS field after run-length encoding or differential encoding to obtain the fifth initial compression result of the POS field; and compress the SEQ field that is unavailable to the RNAME field after matching encoding, and obtain the sixth initial compression result of the SEQ field that is unavailable to the RNAME field.
[0104] After obtaining the first initial compression result, the second initial compression result, the third initial compression result, the fourth initial compression result, the fifth initial compression result and the sixth initial compression result, the process proceeds to step S219 to output the second compression result.
[0105] In the technical solution provided in the above step S217 of the present invention, the second compression result is output, that is, the compression result of the SEQ field, the RNAME field, the POS field and the CIGAR field is output.
[0106] In this embodiment, the execution status of the compression method of the field in this application can be monitored through the massive SAM compression task load balancing scheduling and monitoring architecture. For example, Figure 2(c) is a schematic diagram of a massive SAM compression task load balancing scheduling and monitoring architecture according to an embodiment of the present invention. As shown in Figure 2(c), the architecture may include: master nodes 2211 to 2213, working nodes 2221 to 2214, Zookeeper cluster 223, Kafka cluster 224, database server 225 and network server 226. The master nodes 2211 to 2213 can be used to control the working nodes 2221 to 2214. The working nodes 2221 to 2214 can be used to upload the running status of the compression task to the database server 225 and the Kafka cluster 224. The database server 225 can control the master nodes 2211 to 2213 to select the node to which the message needs to be sent from the working nodes 2221 to 2214 according to resource utilization and file location. The network server 226 can be used to send decompression tasks and / or compression tasks to the Kafka cluster 224. The Kafka cluster can be used to interact with the master nodes 2211 to 2213 for data.
[0107] In this embodiment, the master nodes (TSMaster) 2211 to 2213 may be used to receive compression tasks and distribute the received compression tasks to the worker nodes (TSWorker) 2221 to 2214. The master nodes may be generated by a distributed lock algorithm.
[0108] In this embodiment, the above-mentioned working nodes 2221 to 2214 can be used to trigger the execution of the compression task and monitor the execution status of the compression task.
[0109] In this embodiment, when the database server 225 in the SAM compression system is started, a temporary ordered child node is first created under the directory (Zookeeper's / ZSAM_Sytem / Task / Master). After the child node is successfully created, the information of all child nodes under the directory is obtained. If the child node created by the database server 225 is the first three nodes with the smallest sequence number, it means that a distributed lock has been obtained, and the current server will create TSMaster nodes 2211 to 2213, otherwise it will register and listen to / ZSAM_Sytem / Task / Master. When a child node is deleted due to TSMaster downtime, an event notification will be triggered, and the database server 225 will recheck the node information and try to obtain a distributed lock. If successful, the TSMaster process will be started, otherwise it will be tried again at the next notification.
[0110] In this embodiment, the Kafka message queue is used to implement asynchronous submission of user tasks. The tasks submitted by users on the front-end page are sent to a topic named task-submit in Kafka through a web server, and TSMaster is used to pull messages from task-submit.
[0111] In this embodiment, a Web server is used to partition messages using a polling strategy, thereby evenly sending task messages to the nine partitions of task-submit, thereby achieving the purpose of ensuring load balancing. TSMaster is used to partition messages using a collaborative stickiness strategy, and when a faulty TSMaster goes offline and a newly elected TSMaster comes online, the message pulling behavior of other normal TSMasters will not be stopped, thereby avoiding system response delays.
[0112] In this embodiment, after TSMaster obtains the task, it reads the resource utilization of the server and the storage location of the file to be compressed from the database. The system will select the TSWorker on the server based on the streaming computing engine (Flink) with the lowest resource utilization to send the message. After Flink receives the command, it will perform distributed streaming compression on the SAM gene sequence stored on the distributed file system (Hadoop Distributed File System, referred to as HDFS).
[0113] In this embodiment, HDFS uses TSWorker to monitor distributed compression tasks. If an error occurs during the execution of a compression or decompression task, TSWorker sends a task failure message to a Topic named task-status in Kafka and modifies the relevant records in the database, where TSMaster pulls messages from task-status and sends email notifications to users. Through the above operations, users can obtain the task execution status in real time and control the life cycle of the task. It should be noted that the submission monitoring method for decompression tasks and compression tasks is the same.
[0114] Optionally, by compressing the first target field obtained by the division, a first compression result of the first target field can be obtained by compressing the second target field obtained by the division, a second compression result of the second target field can be obtained, and a third compression result of the third target field can be obtained by compressing the third target field obtained by the division. For example, FIG2(d) is a schematic diagram of a SAM streaming parallel compression architecture based on Flink according to an embodiment of the present invention. As shown in FIG2(d), the architecture may include: a consumer component (FlinkKafkaConsumer) 231, output components (hdfsSink) 2321 to 2323, and a distributed file system 233.
[0115] In this embodiment, the above-mentioned consumption component 231 can be used to receive a SAM target file, and filter each line of data in the target file through a filtering operator to separate the QUAL field, SEQ field and remaining fields from the target file.
[0116] In this embodiment, the separated QUAL field is reordered, the reordered QUAL field is adaptively encoded, and the adaptively encoded QUAL field is compressed; at the same time, the separated SEQ field is reference compressed, and the separated remaining field is segmented, and the segmented remaining field is compressed through PPMD.
[0117] Optionally, corresponding compression algorithms are designed for different fields of SAM. For example, for the QUAL field, the same samples are first arranged in a centralized manner through reordering, and then the QUAL field is mapped to an enlarged interval that is easy to encode through adaptive encoding, and finally PPMD compression is performed to obtain the compression result of the QUAL field; for the SEQ field, the corresponding mapping fragments are read from the historical data (Redis), and then the differences between the fragments are matched and encoded; for the remaining fields, they are first split according to the tab character (TAB), and the number of substrings in each line is recorded, and the obtained substrings are split again according to ":", and then the label names, label types and label values of multiple lines are compressed by PPMD respectively, where the number of substrings can be differentially encoded first and then compressed by PPMD.
[0118] In this embodiment, the output components 2321 to 2323 may be used to write the compression results obtained by compression into the distributed file system 233. For example, the output component 2321 may be used to write the compression results of the QUAL field obtained by compression into area 1 of the distributed file system 233, the output component 2322 may be used to write the compression results of the SEQ field obtained by compression into area 2 of the distributed file system 233, and the output component 2323 may be used to write the compression results of the remaining fields obtained by compression into area n of the distributed file system 233.
[0119] In this embodiment, when compressing a field, a target file can be obtained, and the target field in the obtained target file can be divided to obtain a first target field, a second target field, and a third target field. By compressing the first target field obtained by division, a first compression result of the first target field can be obtained, by compressing the second target field obtained by division, a second compression result of the second target field can be obtained, and by compressing the third target field obtained by division, a third compression result of the third target field can be obtained. Since the different target fields obtained by division are compressed separately, the scalability of the compressed target file can be guaranteed, thereby achieving the purpose of avoiding the field compression being limited by the resources of a single machine, thereby solving the technical problem of low efficiency of field compression, and further achieving the technical effect of improving the efficiency of field compression.
[0120] According to an embodiment of the present invention, a field compression device is also provided. It should be noted that the field compression device can be used to execute a field compression method in the embodiment.
[0121] Figure 3 FIG. 1 is a schematic diagram of a field compression device according to an embodiment of the present invention. Figure 3As shown, the compression device 300 of the field may include: an acquisition unit 301, a division unit 302 and a compression unit 303.
[0122] The acquisition unit 301 is used to acquire a target file, wherein the target file includes a target field associated with gene data.
[0123] The division unit 302 is used to divide the target field in the target file to obtain a first target field, a second target field and a third target field, wherein the first target field is used to represent the quality index of the sequence alignment of the genetic data, the second target field is used to represent the field associated with the sequence alignment, and the third target field is used to represent the field in the target field except the first target field and the second target field.
[0124] The compression unit 303 is used to compress the first target field to obtain a first compression result of the first target field, compress the second target field to obtain a second compression result of the second target field, and compress the third target field to obtain a third compression result of the third target field.
[0125] Optionally, the compression unit 303 may include: a first determination module, used to respectively determine the scores of multiple first target subfields in the first target field; a sorting module, used to sort the multiple first target subfields based on the score of each first target subfield; and a first compression module, used to compress the sorted multiple first target subfields to obtain a first compression result.
[0126] Optionally, the sorting module may include: a construction submodule, used to construct multiple first target subfields corresponding to the same scores into corresponding graphs in response to the existence of the same score in the scores of multiple first target subfields; a determination submodule, used to determine the similarity between multiple graphs corresponding to the multiple first target subfields; and a sorting submodule, used to sort the multiple first target subfields based on the similarity.
[0127] Optionally, the sorting submodule may sort the multiple first target subfields based on the similarity by the following steps: sorting the similarity; and sorting the multiple first target subfields according to the sorted similarity.
[0128] Optionally, the compression unit 303 may include: a construction module, used to construct a target sequence corresponding to the second target field; and a second compression module, used to compress the second target field using the target sequence to obtain a second compression result.
[0129] Optionally, the compression unit 303 may include: a segmentation module, used to segment the third target field; a second determination module, used to determine the change state of the third target field in response to the segmented third target field including the target element; a third determination module, used to determine a coding strategy that matches the change state, wherein the coding strategy is used to control the encoding of the third target field; a third compression module, used to encode the third target field according to the coding strategy, compress the encoded third target field, and obtain a third compression result.
[0130] In this embodiment, an acquisition unit is used to acquire a target file, wherein the target file includes a target field associated with genetic data; a division unit is used to divide the target field in the target file to obtain a first target field, a second target field and a third target field, wherein the first target field is used to represent a quality indicator of a sequence alignment of the genetic data, the second target field is used to represent a field associated with the sequence alignment, and the third target field is used to represent a field in the target field other than the first target field and the second target field; a compression unit is used to compress the first target field to obtain a first compression result of the first target field, compress the second target field to obtain a second compression result of the second target field, and compress the third target field to obtain a third compression result of the third target field, thereby achieving the purpose of avoiding field compression being limited by single-machine resources, thereby solving the technical problem of low efficiency of field compression, and further achieving the technical effect of improving the efficiency of field compression.
[0131] According to an embodiment of the present invention, a processor is further provided. The processor is used to run a program, wherein the program executes the field compression method in the embodiment when the program is run by the processor.
[0132] According to an embodiment of the present invention, there is further provided an electronic device, comprising: a memory storing an executable program; and a processor for running the program, wherein the field compression method in the embodiment is executed when the program is running.
[0133] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium includes a stored program, wherein when the program is executed, the device where the computer-readable storage medium is located is controlled to execute the field compression method in the embodiment.
[0134] According to an embodiment of the present invention, a computer program product is further provided. The computer program product includes a computer program. When the computer program is executed by a processor, the field compression method in the embodiment is implemented.
[0135] According to an embodiment of the present invention, a computer program product is also provided, including a non-volatile computer-readable storage medium, wherein the non-volatile computer-readable storage medium is used to store a computer program, and when the computer program is executed by a processor, the field compression method in the embodiment is implemented.
[0136] According to an embodiment of the present invention, a computer program is also provided. When the computer program is executed by a processor, the field compression method in the embodiment is implemented.
[0137] The serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0138] In the above embodiments of the present invention, the description of each embodiment has its own emphasis. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0139] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only schematic. For example, the division of units can be a logical function division. There may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, which can be electrical or other forms.
[0140] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed over multiple units. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.
[0141] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.
[0142] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the relevant technology or the whole or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions for a computer device (which can be a personal computer, a server or a network device, etc.) to perform all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, read-only memory (ROM, referred to as Read-Only Memory), random access memory (RAM, referred to as Random Access Memory), mobile hard disk, magnetic disk or optical disk and other media that can store program codes.
[0143] The above are only preferred embodiments of the present invention. It should be pointed out that, for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.
Claims
1. A method for compressing a field, characterized in that: include: Acquire a target file, wherein the target file includes a target field associated with the gene data; Dividing the target field in the target file to obtain a first target field, a second target field and a third target field, wherein the first target field is used to represent a quality indicator of a sequence alignment of the gene data, the second target field is used to represent a field associated with the sequence alignment, and the third target field is used to represent fields in the target field other than the first target field and the second target field; The first target field is compressed to obtain a first compression result of the first target field, the second target field is compressed to obtain a second compression result of the second target field, and the third target field is compressed to obtain a third compression result of the third target field.
2. The method according to claim 1, characterized in that Compressing the first target field to obtain a first compression result of the first target field includes: respectively determining scores of a plurality of first target subfields in the first target field; sorting the plurality of first target subfields based on the score of each of the first target subfields; The sorted plurality of the first target subfields are compressed to obtain the first compression result.
3. The method according to claim 2, characterized in that Sorting the plurality of first target subfields based on the score of each of the first target subfields comprises: In response to the existence of an identical score among the scores of the plurality of first target subfields, constructing the plurality of first target subfields corresponding to the identical scores into corresponding graphs; Determining similarities between a plurality of the graphs corresponding to a plurality of the first target subfields; Based on the similarity, the first target subfields are sorted.
4. The method according to claim 3, characterized in that Sorting the plurality of first target subfields based on the similarity includes: sorting the similarities; The plurality of first target subfields are sorted according to the sorted similarities.
5. The method according to claim 1, characterized in that Compressing the second target field to obtain a second compression result of the second target field includes: constructing a target sequence corresponding to the second target field; The second target field is compressed using the target sequence to obtain the second compression result.
6. The method according to claim 1, characterized in that Compressing the third target field to obtain a third compression result of the third target field includes: Segmenting the third target field; In response to the segmented third target field including the target element, determining a change state of the third target field; Determining a coding strategy that matches the change state, wherein the coding strategy is used to control encoding of the third target field; According to the compression strategy, the third target field is encoded, and the encoded third target field is compressed to obtain the third compression result.
7. A field compression device, characterized in that: include: An acquisition unit, configured to acquire a target file, wherein the target file includes a target field associated with the gene data; a dividing unit, configured to divide the target field in the target file to obtain a first target field, a second target field, and a third target field, wherein the first target field is used to represent a quality indicator of a sequence alignment of the gene data, the second target field is used to represent a field associated with the sequence alignment, and the third target field is used to represent a field in the target field other than the first target field and the second target field; A compression unit is used to compress the first target field to obtain a first compression result of the first target field, to compress the second target field to obtain a second compression result of the second target field, and to compress the third target field to obtain a third compression result of the third target field.
8. A processor, characterized in that: The processor is used to run a program, wherein the program, when run by the processor, executes the field compression method described in any one of claims 1 to 6.
9. An electronic device, characterized in that: include: A memory storing an executable program; A processor, used to run the program, wherein the program executes the field compression method described in any one of claims 1 to 6 when running.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored executable program, wherein when the executable program is run, the device where the storage medium is located is controlled to execute the field compression method described in any one of claims 1 to 6.