A method, device and computer-readable storage medium for compressing genotype information

The index array is generated by chunking, rearrangement and sparse matrix coding of genotype bitmatrix, combined with fast-lzma2 algorithm compression, which solves the problems of low compression rate and poor compatibility of existing tools, and realizes efficient genotype information compression and fast query.

CN115691683BActive Publication Date: 2025-07-29SHENZHEN UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211384583.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-07
Publication Date
2025-07-29
Estimated Expiration
2042-11-07

AI Technical Summary

Technical Problem

Existing genotype information compression tools such as GenoType Compressor (GTC) are insufficient in compression rate and compatibility, and do not support efficient random retrieval, resulting in low genotype information compression efficiency.

Method used

The genotype bit matrix is used to block, rearrange and operation to generate a sparse matrix, and the index array is generated through sparse matrix encoding, and the index array is compressed using preset compression algorithms such as fast-lzma2 to form an efficient compressed file.

Benefits of technology

It improves the compression efficiency of genotype information, supports efficient random retrieval and query, and has better compression performance than existing tools.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115691683B_ABST
    Figure CN115691683B_ABST
Patent Text Reader

Abstract

The present application relates to a method, apparatus, and computer-readable storage medium for compressing genotype information. The method includes: partitioning a genotype bit matrix according to a preset rule to obtain genotype bit matrix blocks; wherein the genotype bit matrix is a 01 matrix; rearranging and operating on the genotype bit matrix blocks to obtain a sparse matrix; encoding the sparse matrix to generate an index array; and compressing the index array based on a preset compression algorithm to obtain a compressed file. Through the implementation of the solution of the present application, the genotype bit matrix is partitioned, then the genotype bit matrix blocks are rearranged and operated on to obtain a sparse matrix, then the sparse matrix is encoded to obtain an index array, and finally the index array is compressed, thereby effectively improving the efficiency of compressing genotype information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of bioinformatics technology, and in particular, to a genotype information compression method, device, and computer-readable storage medium. Background Art

[0002] Currently, sequencing technology is becoming increasingly mature, and the accompanying sequencing costs are constantly decreasing, reaching a situation where almost everyone can afford to have their genes sequenced. As a result, gene sequencing data has shown exponential explosive growth. Although the sequencing cost has decreased, the storage and transmission costs and pressures are still a major problem. Currently, the compression effect of general compression tools on gene sequencing data cannot meet the actual needs, and there is still an urgent need to develop more effective data exploration and compression strategies. One of the main purposes of gene sequencing is to mine and analyze variant information. Currently, variant information is still stored in the file format of Variant Call Format (VCF). This file format is an important file format for downstream analysis and has become the standard for genomic variant research.

[0003] As the number of people being sequenced continues to grow, large-scale sequencing projects are also increasing. The VCF files generated in these projects can reach hundreds of gigabytes or even larger. Therefore, there is still a need for an efficient compression method.

[0004] The VCF file contains a list of variants in a genome, as well as information on the reference / non-reference alleles present at each specific variant position in each genome. Since they are intensively searched, the applied compression scheme should support various types of fast queries. So when the query is about a single variant or their range, VCFtools (Danecek et al., 2011) or BCFtools can be used to effectively query indexed and gzip-compressed VCF files. However, retrieving sample data means time-consuming decompression and processing of the entire file.

[0005] Some existing tools for compressing VCF files, such as compression tools like PBWT, BGT, GTC, GTShark, VCFShark, genozip, etc. However, most of these tools do not support random retrieval, such as GTShark and VCFShark, etc. While BGT and GTC support random sample indexing, their compression ratios are not high enough. Currently, the most mainstream tool for genotype compression and random retrieval of VCF files is the GenoType Compressor (GTC) proposed by Agnieszka Danek. The core idea of this tool's compression is to block genotypes and perform the nearest neighbor greedy sorting based on Hamming distance on the sample columns between the blocked data, so that similar sample columns can be better concentrated together. Then, perform run-based tuple encoding on it, and finally perform Huffman encoding on the tuples to complete further compression. Since this encoding method makes the coupling between data not high, this method supports compression and random indexing. However, this compression tool has deficiencies in compression ratio because it needs to record the order of column sorting during sorting and the subsequent encoding is not the optimal encoding. Secondly, due to its excessive dependence on the htslib library, its compatibility is also not good. Summary of the Invention

[0006] The embodiments of the present application provide a genotype information compression method, device, and computer-readable storage medium, which can at least solve the problem of low efficiency in compressing genotype information in VCF files in related technologies.

[0007] The first aspect of the embodiments of the present application provides a genotype information compression method, including:

[0008] Block the genotype bit matrix according to a preset rule to obtain genotype bit matrix blocks; wherein, the genotype bit matrix is a 01 matrix;

[0009] Rearrange and operate on the genotype bit matrix blocks to obtain a sparse matrix;

[0010] Encode the sparse matrix to generate an index array;

[0011] Compress the index array based on a preset compression algorithm to obtain a compressed file.

[0012] The second aspect of the embodiments of the present application provides an electronic device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, it implements each step in the genotype information compression method provided in the first aspect of the embodiments of the present application.

[0013] In the third aspect of the embodiments of the present application, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, each step in the genotype information compression method provided in the first aspect of the embodiments of the present application is implemented.

[0014] As can be seen from the above, according to the genotype information compression method, device, and computer-readable storage medium provided by the solution of the present application, the genotype bit matrix is partitioned according to a preset rule to obtain genotype bit matrix blocks; wherein, the genotype bit matrix is a 01 matrix; the genotype bit matrix blocks are rearranged and operated to obtain a sparse matrix; the sparse matrix is encoded to generate an index array; and the index array is compressed based on a preset compression algorithm to obtain a compressed file. By implementing the solution of the present application, the genotype bit matrix is partitioned, then the genotype bit matrix blocks are rearranged and operated to obtain a sparse matrix, then the sparse matrix is encoded to obtain an index array, and finally the index array is compressed, thereby effectively improving the efficiency of compressing genotype information. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] Figure 1 is a schematic flowchart of the basic process of the genotype information compression method provided in the first embodiment of the present application;

[0016] Figure 2 is a schematic diagram of the result of partitioning a genotype bit matrix provided in the first embodiment of the present application;

[0017] Figure 3 is a schematic diagram of the result of preprocessing genotype data provided in the first embodiment of the present application;

[0018] Figure 4 is a schematic diagram of the result of rearranging and performing exclusive OR operations on genotype bit matrix blocks provided in the first embodiment of the present application;

[0019] Figure 5 is a schematic diagram of the result of hiding the storage arrangement order values provided in the first embodiment of the present application;

[0020] Figure 6 is a schematic flowchart of the detailed process of the genotype information compression method provided in the second embodiment of the present application;

[0021] Figure 7 is a schematic diagram of the modules of a genotype information compression device provided in the third embodiment of the present application;

[0022] Figure 8 is a schematic diagram of the structure of an electronic device provided in the fourth embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] In order to make the invention object, technical solution and advantages of the present application more obvious and understandable, the following will describe the technical solutions in the embodiments of the present application clearly and completely in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work belong to the protection scope of the present application.

[0024] To solve the problem of low efficiency in compressing genotype information in a VCF file in related technologies, the first embodiment of the present application provides a genotype information compression method, which is implemented by (Genotype SparseCompressor) GSC. When compressing a VCF file, GSC first divides the VCF file into three parts: header information, variant description information, and genotype information. Among them, the header information and the variant description information are regarded as a whole and compressed using the fast-lzma2 algorithm. The compression tool GSC mainly processes the genotype information part. As Figure 1 This is the basic flowchart of the genotype information compression method provided in this embodiment. The genotype information compression method includes the following steps:

[0025] Step 101: Divide the genotype bit matrix according to a preset rule to obtain genotype bit matrix blocks.

[0026] Specifically, in this embodiment, the genotype bit matrix is a 01 matrix. GSC divides the genotype bit matrix into fixed blocks according to the rule that the number of variants = the number of samples * the number of haplotypes to obtain genotype bit matrix blocks. Dividing the genotype bit matrix according to this rule is to make the columns of the array perm[] storing the sample sorting order correspond one by one with the variant field array pos in the variant description information during subsequent processing. The form of the divided 01 bit matrix block (with a row size of block_size and a column size of samples * the number of haplotypes) is as Figure 2 shown.

[0027] In some embodiments of this embodiment, before the step of dividing the genotype bit matrix according to a preset rule to obtain genotype bit matrix blocks, it further includes: obtaining the mutation types of samples in the genotype data; where the mutation types include: no mutation, the first type of mutation, other mutations, unknown mutations; encoding the genotype data according to the mutation types to obtain genotype bit vectors corresponding to the mutation types, and generating a genotype bit matrix.

[0028] Specifically, in this embodiment, before dividing the genotype bit matrix, the genotype data is first processed. First, obtain the mutation types of samples in the genotype data, such asFigure 3 The types of mutations shown can be no mutation, i.e., the reference allele "0", the first type of mutation, i.e., "1", other mutations, i.e., "2", and unknown mutations, i.e., ".". The GSC in this embodiment uses the same genotype encoding process as the BGT tool, that is, two bit vectors are used to encode different mutation types. The genotype bit vector of "00" is obtained by encoding the reference allele "0", the genotype bit vector of "01" is obtained by encoding the first type of mutation "1", the genotype bit vector of "11" is obtained by encoding other mutations "2", and the genotype bit vector of "10" is obtained by encoding unknown mutations ".". Through encoding, each haplotype can be represented by two bits, and a genotype bit matrix is obtained based on these genotype bit vectors.

[0029] In some other embodiments of this embodiment, after the step of obtaining the genotype bit matrix blocks by partitioning the genotype bit matrix according to a preset rule, it further includes: when there are identical rows in the genotype bit matrix blocks, obtaining the positions of the identical rows and marking them; where the identical rows include: all-zero rows, and identical non-all-zero rows; when the identical rows are non-all-zero identical rows, storing the replicated rows of the original rows corresponding to the non-all-zero identical rows in the replicated row marking array to obtain the genotype bit matrix blocks without replicated rows.

[0030] Specifically, in this embodiment, after obtaining the genotype bit matrix blocks, the GSC will also mark each row of the genotype bit matrix blocks to find rows that are all zero or completely identical rows other than all zero in the matrix block, such as rows that are all one or identical rows containing zero and one. In this embodiment, two bit vectors of size block_size are used to store the marked rows, with 0 representing rows other than all-zero rows or non-all-zero identical rows, and 1 representing all-zero rows or non-all-zero identical rows. If there are non-all-zero identical rows, after marking the identical rows, the row numbers of the replicated rows corresponding to the identical rows are stored in the pre-set row number array of the replicated rows and saved to the compressed file obtained by compressing the header information and variant description information, and the replicated rows of the original rows corresponding to the identical rows are stored in the marking array to obtain the genotype bit matrix blocks without identical non-zero rows. Through simple processing, the number of rows in the original matrix block is greatly reduced.

[0031] Step 102: Rearrange and operate on the genotype bit matrix blocks to obtain a sparse matrix.

[0032] Specifically, in this embodiment, after obtaining the genotype bit matrix blocks, the matrix blocks are rearranged and operated on to obtain a sparse matrix.

[0033] In some embodiments of this embodiment, the steps of rearranging and operating on the genotype bit matrix block to obtain a sparse matrix include: rearranging each column in the genotype bit matrix block without duplicate rows according to the nearest neighbor heuristic algorithm; performing an exclusive OR operation on the preset columns in the rearranged genotype bit matrix block to obtain a sparse matrix.

[0034] Specifically, in this embodiment, the rearrangement of the genotype bit matrix block is performed by using the nearest neighbor heuristic algorithm and based on the Hamming distance. After obtaining the rearranged matrix block, an exclusive OR operation is performed on adjacent columns in the matrix block. It can be an exclusive OR operation every eight columns or an exclusive OR operation between adjacent two columns, so as to reduce the number of 1s in the matrix block, which can make the 01 matrix become sparser. As Figure 4 shown in the processing result of rearranging and performing an exclusive OR operation on the genotype bit matrix block, where 0 is represented by a white grid and 1 is represented by a gray grid. In this embodiment, it is optionally adopted to perform an exclusive OR operation every eight columns, and the exclusive OR operation is performed based on the first column of every eight columns. Since each grid represents one bit of data, the advantage of performing the exclusive OR operation is that every eight columns of data are within one byte, reducing the coupling between data. During the decompression process, only the corresponding byte needs to be processed, which can meet the requirement of random retrieval.

[0035] Step 103: Encode the sparse matrix to generate an index array.

[0036] Specifically, in this embodiment, after obtaining the sparse matrix, the sparse matrix is further encoded to obtain an index array.

[0037] In some embodiments of this embodiment, the steps of encoding the sparse matrix to generate an index array include: encoding each row in the sparse matrix with a fixed number of bytes to obtain an index; where the index is used to record the position information of the preset bytes; performing differential encoding on the index; generating an index array based on the index obtained by the differential encoding.

[0038] Specifically, in this embodiment, GSC encodes each row of the sparse matrix after XOR operation using a fixed - byte (default two - byte) encoding to obtain an index. The index is used to store the position information of preset bytes. For example, it records the position information where "1" exists in each row, and uses 0 as the delimiter for each row. If the positions of 2, 6, 7, 9 in the first row are "1", then the index [2, 6, 7, 9, 0] can be obtained. If the positions of 2 and 11 in the second row are "1", then the index [2, 11, 0] can be obtained. And differential encoding is performed on each index in each row. For example, differential encoding of the index [2, 6, 7, 9, 0] and the index [2, 11, 0] gives [2, 4, 1, 2, 0] and [2, 9, 0]. All the indexes obtained after differential encoding are stored in an array samples_indexes[], resulting in an index array. If the last matrix block and the previous matrix blocks are all square, they are rearranged, operated on, and encoded in the same way. However, usually, in the case of the last block, the size of the array perm[] storing the sample permutation order (i.e., the corresponding number of samples * ploidy) is greater than the size of pos, so the method of hiding perm[] into the columns of pos cannot be used. Therefore, GSC is not suitable for rearranging, operating on, and encoding it. Currently, considering that the size of the last matrix block is very small in the whole data, the XOR operation method is not used for this matrix block, but it is directly subjected to sparse encoding, differential encoding, and compression using the fast - lzma2 algorithm.

[0039] Step 104: Compress the index array based on a preset compression algorithm to obtain a compressed file.

[0040] Specifically, in this embodiment, GSC compresses the index array according to a preset compression algorithm such as the LZMA algorithm, the fast - lzma2 algorithm, or the zstd algorithm. In this embodiment, the fast - lzma2 algorithm is optionally used to compress the index array to obtain a compressed file.

[0041] In some embodiments of this embodiment, the step of compressing the index array based on a preset compression algorithm to obtain a compressed file includes: merging all index arrays into a total index array; compressing the total index array based on a preset compression algorithm to obtain a compressed file.

[0042] Specifically, in this embodiment, GSC compresses the total index array all_samples_index[] according to the preset fast - lzma2 algorithm to obtain a compressed file, where the total index array is the sum of the index arrays in each matrix block.

[0043] In some other embodiments of this embodiment, after the step of compressing the index array based on a preset compression algorithm to obtain a compressed file, the following steps are further included: decompressing the compressed file according to the instruction type of the query instruction to obtain query content corresponding to the instruction type; wherein, the instruction types include: query by mutation site and query by sample name.

[0044] Specifically, in this embodiment, GSC supports various types of queries in the same way as GTC, including queries by mutation site and by sample name, and decompresses the compressed file according to different query instructions to obtain corresponding query results.

[0045] Further, in some embodiments of this embodiment, the step of decompressing the compressed file according to the instruction type to obtain query content corresponding to the instruction type includes: when the instruction type is query by mutation site, obtaining a first target genotype bit matrix block corresponding to the mutation site information and mutation information in the compressed file according to the mutation site information; obtaining the columns of the first mutation field array, and rearranging the columns of the first mutation field array to obtain the order of the columns in the first target genotype bit matrix block; wherein, the first mutation field array is used to hide the array storing the arrangement order of the first target genotype bit matrix block according to a preset transformation rule; according to the order of the columns in the first target genotype bit matrix block, decompressing the rows corresponding to the mutation information in the first target genotype bit matrix block, and performing an exclusive OR operation on the decompressed rows to obtain query content corresponding to the mutation site information.

[0046] Specifically, in this embodiment, when the instruction type is query by mutation site, the step of obtaining a first target genotype bit matrix block corresponding to the mutation site information and mutation information in the compressed file according to the mutation site information includes: obtaining a first target genotype bit matrix block corresponding to the mutation site information according to the mutation site information input by the user for query, and obtaining the index corresponding to the mutation site information, and locating the mutation information row_id corresponding to the index according to the index. The step of decompressing the rows corresponding to the mutation information in the first target genotype bit matrix block according to the order of the columns in the first target genotype bit matrix block further includes: obtaining the columns of the first mutation field array, and rearranging the columns according to the merge sort method to obtain the order of the columns in the first target genotype bit matrix block, wherein the first mutation field array POS[] has the same size as the array Perm[] storing the arrangement order corresponding to the samples in the first target genotype bit matrix block, and is used to hide the array storing the arrangement order to the first mutation field array according to a preset transformation rule. For example, according to the rule Perm[i]=m, new POS[m]=POS[i], the array storing the arrangement order is hidden to the mutation field array, such as Figure 5The hidden result is shown. Since the Perm[] array corresponding to the sample is ordered, and the columns of POS[] in the variant description information happen to be ordered as well, and when partitioning, it is ensured that the sizes of both are equal. Thus, when sorting, the columns of the corresponding POS[] are transformed equally. When restoring the order of the array storing the permutation order, only the variant field array needs to be sorted, and the same operation is performed on the array storing the permutation order together.

[0047] In this embodiment, the step of decompressing the rows corresponding to the mutation information in the first target genotype bit matrix block according to the order of the columns in the first target genotype bit matrix block includes: according to the order of the columns in the first target genotype bit matrix block, obtaining the all-0 rows and non-all-0 identical rows in the first target genotype bit matrix block. When the mutation information row_id is marked as the all-0 row, decompress the all-0 row corresponding to the mutation information; when the mutation information row_id is marked as a non-all-0 identical row, obtain the row number of the copy row of the original row corresponding to the non-all-0 identical row, and decompress the copy row according to the row number; when the mutation information row_id is not marked as an all-0 row or a non-all-0 identical row, decompress all the bytes of the row corresponding to the mutation information. And perform an exclusive OR operation on the decompressed rows, and restore the original bit vector of the row according to the order of the columns in the first target genotype bit matrix block to obtain the query content corresponding to the mutation site information, such as the variant row.

[0048] Furthermore, in some embodiments of this embodiment, the step of decompressing the compressed file according to the instruction type to obtain the query content corresponding to the instruction type includes: when the instruction type is query by sample name, obtaining the target position of the target sample corresponding to the sample name in the compressed file; according to the target position, obtaining the second target genotype bit matrix block corresponding to the target sample in the compressed file; obtaining the columns of the second variant field array, and rearranging the columns of the second variant field array to obtain the order of the columns in the second target genotype bit matrix block; where the second variant field array is used to hide the array storing the permutation order of the second target genotype bit matrix block according to a preset transformation rule; according to the order of the columns in the second target genotype bit matrix block, decompressing the bytes corresponding to the target sample in the second target genotype bit matrix block, and performing an exclusive OR operation on the decompressed bytes to obtain the query content corresponding to the sample name.

[0049] Specifically, in this embodiment, when the instruction type is query by sample name, the step of obtaining the target position of the target sample corresponding to the sample name in the compressed file includes: obtaining the input sample name, obtaining the original position of the target sample corresponding to the sample name according to the sample name, and obtaining the target position of the target sample in the compressed file according to the original position, the rearranged position and index of the second variant field array POS[]. According to the target position, obtain the second target genotype bit matrix block corresponding to the target sample in the compressed file; obtain the columns of the second variant field array, and rearrange the columns of the second variant field array to obtain the order of the columns in the second target genotype bit matrix block; wherein, the second variant field array is used to hide the array storing the arrangement order of the second target genotype bit matrix block according to a preset transformation rule; wherein, the second variant field array POS[] is equal in size to the array Perm[] storing the arrangement order corresponding to the samples in the second target genotype bit matrix block, and is used to hide the array storing the arrangement order into the second variant field array according to a preset transformation rule. For example, according to the rule Perm[i]=m, new POS[m]=POS[i], the array storing the arrangement order is hidden into the second variant field array. Since the Perm[] array corresponding to the samples is ordered, and the columns of POS[] in the variant description information are also ordered, and the sizes of both are ensured to be equal when partitioning, so that the corresponding columns of POS[] perform the same transformation during sorting. When restoring the order of the array storing the arrangement order, only need to sort the second variant field array, and the array storing the arrangement order performs the same operation together; according to the order of the columns in the second target genotype bit matrix block, decompress the bytes corresponding to the target sample in the second target genotype bit matrix block, and perform an exclusive OR operation on the decompressed bytes to obtain the query content corresponding to the sample name, such as the sample to be queried. It can be found that there is no longer a need to decompress an entire row of genotype bit vectors as when querying by mutation site before. Only need to index the "1" position information in the sparse coding, then decompress the corresponding bytes, and then perform an exclusive OR operation inside the bytes to query the corresponding result.

[0050] As can be seen from the above, GSC does not need to decompress the entire data when facing these two types of queries. When facing queries by mutation sites, it only needs to decompress the blocks spanned by these mutation sites, and when querying by sample, it only needs to decompress its corresponding bytes. This greatly improves the query speed. However, for both types of queries, it is necessary to decompress all its variant description information to obtain the columns of its mutation field array, obtain the order of the columns of the matrix block according to the columns of the mutation field array, and then decompress the corresponding matrix block according to the column order. By combining the methods of column sorting order hiding, column XOR operation (XOR operation between genotype data columns, similar to differential coding), sparse matrix coding, and the fast-lzma2 algorithm, the compression performance of the GSC tool will be higher than that of GTC. At the same time, due to the selectivity of the column XOR operation, either the XOR operation can be performed between every eight columns of genotype data within each genotype matrix block, or the XOR operation can be performed between each column of data within each genotype matrix block, so as to achieve the effect of higher compression ratio and query speed. No matter which XOR operation scheme is used, the compression performance is better than that of GTC.

[0051] Based on the above technical solution of the embodiment of the present application, the genotype bit matrix is partitioned according to a preset rule to obtain genotype bit matrix blocks; wherein, the genotype bit matrix is a 01 matrix; the genotype bit matrix blocks are rearranged and operated to obtain a sparse matrix; the sparse matrix is encoded to generate an index array; and the index array is compressed based on a preset compression algorithm to obtain a compressed file. Through the implementation of the solution of the present application, the genotype bit matrix is partitioned, then the genotype bit matrix blocks are rearranged and operated to obtain a sparse matrix, then the sparse matrix is encoded to obtain an index array, and finally the index array is compressed, thereby effectively improving the efficiency of compressing genotype information.

[0052] Figure 6 The method in [ ] provides a refined genotype information compression method for the second embodiment of the present application. The genotype information compression method includes:

[0053] Step 601, encode the genotype data according to the mutation type to obtain a genotype bit vector corresponding to the mutation type, and generate a genotype bit matrix.

[0054] Specifically, the GSC in this embodiment uses the same genotype encoding process as the BGT tool, that is, two-bit vectors are used to encode different mutation types. The genotype bit vector of "00" is obtained by encoding the reference allele "0", the genotype bit vector of "01" is obtained by encoding the first mutation "1", the genotype bit vector of "11" is obtained by encoding other mutations "2", and the genotype bit vector of "10" is obtained by encoding the unknown mutation ".". Through the encoding process, each haplotype can be represented by two bits, and a genotype bit matrix is obtained based on these genotype bit vectors.

[0055] Step 602: Divide the genotype bit matrix according to a preset rule to obtain genotype bit matrix blocks.

[0056] Specifically, in this embodiment, the genotype bit matrix is a 01 matrix. GSC divides the genotype bit matrix into fixed blocks according to the rule that the number of variants = the number of samples * the number of haplotypes to obtain genotype bit matrix blocks. Dividing the genotype bit matrix according to this rule is to make the columns of the array perm[] storing the sample sorting order and the mutation field array pos in the variant description information correspond one by one during subsequent processing.

[0057] Step 603: When there are identical rows in the genotype bit matrix block, obtain the positions of the identical rows and mark them.

[0058] Step 604: When the identical rows are non-all-zero identical rows, store the copied rows of the original rows corresponding to the non-all-zero identical rows into the copied row marker array to obtain a genotype bit matrix block without copied rows.

[0059] Specifically, in this embodiment, after obtaining the genotype bit matrix block, GSC will also mark each row of the genotype bit matrix block to find rows that are all 0 or exactly the same rows except all 0 in the matrix block, such as rows that are all 1 or identical rows containing 0 and 1. In this embodiment, two bit vectors of size block_size are used to store the marked rows. 0 represents rows other than all-zero rows or non-all-zero identical rows, and 1 represents all-zero rows or non-all-zero identical rows. If there are non-all-zero identical rows, after marking the identical rows, the row numbers of the copied rows corresponding to the identical rows are stored in the pre-set row number array of the copied rows and saved to the compressed file obtained by compressing the header information and variant description information, and the copied rows of the original rows corresponding to the identical rows are stored in the marker array to obtain a genotype bit matrix block without identical non-zero rows. Through simple processing, the number of rows in the original matrix block is greatly reduced.

[0060] Step 605: Rearrange each column in the genotype bit matrix block without duplicate rows according to the nearest neighbor heuristic algorithm.

[0061] Step 606: Perform an exclusive OR operation on the preset columns in the rearranged genotype bit matrix block to obtain a sparse matrix.

[0062] Specifically, in this embodiment, the rearrangement of the genotype bit matrix block is performed by using the nearest neighbor heuristic algorithm based on the Hamming distance. After obtaining the rearranged matrix block, an exclusive OR operation is performed on adjacent columns in the matrix block. It can be an exclusive OR operation every eight columns or an exclusive OR operation between adjacent columns, so as to reduce the number of 1s in the matrix block, making the 01 matrix sparser. In this embodiment, it is optionally an exclusive OR operation every eight columns, with the first column of every eight columns as the reference for the exclusive OR operation.

[0063] Step 607: Encode the sparse matrix to generate an index array.

[0064] Specifically, in this embodiment, after obtaining the sparse matrix, the sparse matrix is also encoded to obtain an index array.

[0065] Step 608: Compress the index array based on a preset compression algorithm to obtain a compressed file.

[0066] Specifically, in this embodiment, GSC compresses the index array according to a preset compression algorithm such as the LZMA algorithm, the fast-lzma2 algorithm, or the zstd algorithm. In this embodiment, it is optionally the fast-lzma2 algorithm that is used to compress the index array to obtain a compressed file.

[0067] Based on the above technical solution of the embodiment of the present application, the genotype bit matrix is partitioned according to a preset rule to obtain a genotype bit matrix block; wherein, the genotype bit matrix is a 01 matrix; the genotype bit matrix block is rearranged and operated to obtain a sparse matrix; the sparse matrix is encoded to generate an index array; the index array is compressed based on a preset compression algorithm to obtain a compressed file. Through the implementation of the solution of the present application, the genotype bit matrix is partitioned, then the genotype bit matrix block is rearranged and operated to obtain a sparse matrix, then the sparse matrix is encoded to obtain an index array, and finally the index array is compressed, thereby effectively improving the efficiency of compressing genotype information.

[0068] Figure 7 This is a genotype information compression device provided in the third embodiment of the present application. This genotype information compression device can be used to implement the genotype information compression method in the foregoing embodiment. As Figure 7 shown, this genotype information compression device mainly includes:

[0069] The chunking module 701 is used to chunk the genotype bit matrix according to a preset rule to obtain genotype bit matrix chunks; wherein, the genotype bit matrix is a 01 matrix.

[0070] The operation module 702 is used to rearrange and operate on the genotype bit matrix chunks to obtain a sparse matrix.

[0071] The encoding module 703 is used to encode the sparse matrix to generate an index array. <o

[0072] The compression module 704 is used to compress the index array based on a preset compression algorithm to obtain a compressed file.

[0073] In some embodiments of this embodiment, the genotype information compression device further includes: a generation module, which is used to obtain the mutation types of samples in the genotype data; wherein, the mutation types include: no mutation, the first mutation, other mutations, unknown mutations; according to the mutation types, encode and process the genotype data to obtain genotype bit vectors corresponding to the mutation types, and generate a genotype bit matrix.

[0074] In some embodiments of this embodiment, the genotype information compression device further includes: a storage module, which is used to obtain the positions of identical rows and mark them when there are identical rows in the genotype bit matrix chunks; wherein, the identical rows include: all-zero rows, non-all-zero identical rows; when the identical rows are non-all-zero identical rows, store the replicated rows of the original rows corresponding to the non-all-zero identical rows into a replicated row marker array to obtain a genotype bit matrix chunk without replicated rows.

[0075] In some embodiments of this embodiment, the operation module 702 is specifically used to: rearrange each column in the genotype bit matrix chunk without replicated rows according to the nearest neighbor heuristic algorithm; perform exclusive OR operations on preset columns in the rearranged genotype bit matrix chunk to obtain a sparse matrix.

[0076] In some embodiments of this embodiment, the encoding module 703 is specifically used to: encode each row in the sparse matrix with a fixed number of bytes to obtain an index; wherein, the index is used to record the position information of preset bytes; perform differential encoding on the index; generate an index array based on the index obtained by differential encoding.

[0077] In some embodiments of this embodiment, the compression module 704 is specifically used to: merge all index arrays into a total index array; compress the total index array based on a preset compression algorithm to obtain a compressed file.

[0078] It should be noted that there is a small error in the original text where "<o " should probably be " ", and this has been corrected in the translation.Further, in some embodiments of the present embodiment, the genotype information compression device further includes: a decompression module, configured to decompress the compressed file according to the instruction type of the query instruction to obtain the query content corresponding to the instruction type; wherein, the instruction types include: query by mutation site and query by sample name.

[0079] Still further, in some embodiments of the present embodiment, the decompression module is specifically configured to: when the instruction type is query by mutation site, obtain the first target genotype bit matrix block corresponding to the mutation site information and the mutation information in the compressed file according to the mutation site information; obtain the columns of the first mutation field array, and rearrange the columns of the first mutation field array to obtain the order of the columns in the first target genotype bit matrix block; wherein, the first mutation field array is used to hide the array for storing and arranging the order of the first target genotype bit matrix block according to a preset transformation rule; according to the order of the columns in the first target genotype bit matrix block, decompress the rows corresponding to the mutation information in the first target genotype bit matrix block, and perform an exclusive OR operation on the decompressed rows to obtain the query content corresponding to the mutation site information.

[0080] Still further, in some other embodiments of the present embodiment, the decompression module is specifically configured to: when the instruction type is query by sample name, obtain the target position of the target sample corresponding to the sample name in the compressed file; according to the target position, obtain the second target genotype bit matrix block corresponding to the target sample in the compressed file; obtain the columns of the second mutation field array, and rearrange the columns of the second mutation field array to obtain the order of the columns in the second target genotype bit matrix block; wherein, the second mutation field array is used to hide the array for storing and arranging the order of the second target genotype bit matrix block according to a preset transformation rule; according to the order of the columns in the second target genotype bit matrix block, decompress the bytes corresponding to the target sample in the second target genotype bit matrix block, and perform an exclusive OR operation on the decompressed bytes to obtain the query content corresponding to the sample name.

[0081] It should be noted that the genotype information compression methods in the first and second embodiments can both be implemented based on the genotype information compression device provided in the present embodiment. Those of ordinary skill in the art can clearly understand that for the convenience and brevity of description, the specific working process of the genotype information compression device described in the present embodiment can refer to the corresponding process in the foregoing method embodiments and will not be elaborated herein.

[0082] The genotype information compression device provided according to this embodiment divides a genotype bit matrix into blocks according to a preset rule to obtain genotype bit matrix blocks; wherein, the genotype bit matrix is a 01 matrix; rearranges and operates on the genotype bit matrix blocks to obtain a sparse matrix; encodes the sparse matrix to generate an index array; and compresses the index array based on a preset compression algorithm to obtain a compressed file. Through the implementation of the solution of this application, the genotype bit matrix is divided into blocks, then the genotype bit matrix blocks are rearranged and operated to obtain a sparse matrix, then the sparse matrix is encoded to obtain an index array, and finally the index array is compressed, thereby effectively improving the efficiency of compressing genotype information.

[0083] Please refer to Figure 8 , Figure 8 which is an electronic device provided in the fourth embodiment of this application. This electronic device can be used to implement the genotype information compression method in the foregoing embodiments. As Figure 8 shown, this electronic device mainly includes:

[0084] A memory 801, a processor 802, and a computer program 803 stored on the memory 801 and executable on the processor 802. The memory 801 and the processor 802 are connected through a bus. When the processor 802 executes the computer program 803, the genotype information compression method in the foregoing embodiments is implemented. Among them, the number of processors can be one or more.

[0085] The memory 801 can be a high-speed random access memory (RAM, Random Access Memory) or a non-volatile memory, such as a disk memory. The memory 801 is used to store executable program codes, and the processor 802 is coupled to the memory 801.

[0086] Furthermore, the embodiment of this application also provides a computer-readable storage medium, which can be set in the electronic device in the foregoing embodiments. The computer-readable storage medium can be the memory in the foregoing Figure 8 shown embodiments.

[0087] A computer program is stored on this computer-readable storage medium, and when the program is executed by a processor, the genotype information compression method in the foregoing embodiments is implemented. Furthermore, the computer-readable storage medium can also be various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a RAM, a magnetic disk, or an optical disc that can store program codes.

[0088] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is only a logical function division. In actual implementation, there can be other division methods. For example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling, direct coupling, or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or modules can be in electrical, mechanical, or other forms.

[0089] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0090] In addition, in each embodiment of the present application, the functional modules can be integrated in a processing module, or each module can exist physically alone, or two or more modules can be integrated in one module. The above-mentioned integrated modules can be implemented in the form of hardware or in the form of software functional modules.

[0091] If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. And the aforementioned readable storage medium includes: USB flash drives, mobile hard disks, ROM, RAM, magnetic disks, or optical discs and other various media that can store program codes.

[0092] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily all essential to the present application.

[0093] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not elaborated in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0094] The above is the description of the genotype information compression method, device, and computer-readable storage medium provided by the present application. For those skilled in the art, according to the idea of the embodiments of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation on the present application.

Claims

1. A genotype information compression method, characterized in that Including: Partition the genotype bit matrix according to a preset rule to obtain genotype bit matrix blocks; wherein, the genotype bit matrix is a 01 matrix; Rearrange and operate on the genotype bit matrix blocks to obtain a sparse matrix; Encode the sparse matrix to generate an index array; Compress the index array based on a preset compression algorithm to obtain a compressed file; Decompress the compressed file according to the instruction type of the query instruction to obtain query content corresponding to the instruction type; wherein, the instruction type includes: query by mutation site, query by sample name; The step of decompressing the compressed file according to the instruction type of the query instruction to obtain query content corresponding to the instruction type includes: When the instruction type is the query by mutation site, according to the mutation site information, obtain the first target genotype bit matrix block and mutation information corresponding to the mutation site information in the compressed file; Obtain the columns of the first mutation field array, and rearrange the columns of the first mutation field array to obtain the order of the columns in the first target genotype bit matrix block; wherein, the first mutation field array is used to hide the array storing the arrangement order of the first target genotype bit matrix block according to a preset transformation rule; According to the order of the columns in the first target genotype bit matrix block, decompress the rows corresponding to the mutation information in the first target genotype bit matrix block, and perform an exclusive OR operation on the decompressed rows to obtain query content corresponding to the mutation site information.

2. The genotype information compression method according to claim 1, wherein Before the step of partitioning the genotype bit matrix according to a preset rule to obtain genotype bit matrix blocks, it further includes: Obtain the mutation types of the samples in the genotype data; wherein, the mutation types include: no mutation, first type of mutation, other mutations, unknown mutations; Encode and process the genotype data according to the mutation type to obtain genotype bit vectors corresponding to the mutation type, and generate the genotype bit matrix.

3. The genotype information compression method according to claim 2, wherein After the step of partitioning the genotype bit matrix according to a preset rule to obtain genotype bit matrix blocks, it further includes: When there are identical rows in the genotype bit matrix blocks, obtain the positions of the identical rows and mark them; wherein, the identical rows include: all-0 rows, identical non-all-0 rows; When the identical rows are the non-all-0 identical rows, store the copied rows of the original rows corresponding to the non-all-0 identical rows into a copied row marker array to obtain genotype bit matrix blocks without the copied rows.

4. The genotype information compression method according to claim 3, wherein The step of rearranging and operating on the genotype bit matrix blocks to obtain a sparse matrix includes: Rearrange each column in the genotype bit matrix blocks without the copied rows according to the nearest neighbor heuristic algorithm; Perform an exclusive OR operation on the preset columns in the rearranged genotype bit matrix blocks to obtain a sparse matrix.

5. The genotype information compression method according to claim 1, characterized in that, The step of encoding the sparse matrix to generate an index array includes: Encode each row in the sparse matrix with a fixed number of bytes to obtain an index; wherein, the index is used to record the position information of the preset bytes; Differentially encode the index; Generate an index array based on the index obtained from the differential encoding; The step of compressing the index array based on a preset compression algorithm to obtain a compressed file includes: Merge all the index arrays into a total index array; Compress the total index array based on a preset compression algorithm to obtain the compressed file.

6. The genotype information compression method according to claim 1, characterized in that The step of decompressing the compressed file according to the instruction type of the query instruction to obtain the query content corresponding to the instruction type includes: When the instruction type is query by sample name, obtain the target position of the target sample corresponding to the sample name in the compressed file; According to the target position, obtain the second target genotype bit matrix block corresponding to the target sample in the compressed file; Obtain the columns of the second variant field array, and rearrange the columns of the second variant field array to obtain the order of the columns in the second target genotype bit matrix block; wherein, the second variant field array is used to hide the array for storing the arrangement order of the second target genotype bit matrix block according to a preset transformation rule; According to the order of the columns in the second target genotype bit matrix block, decompress the bytes corresponding to the target sample in the second target genotype bit matrix block, and perform exclusive OR operation on the decompressed bytes to obtain the query content corresponding to the sample name.

7. An electronic device, characterized in that, Includes: A memory and a processor, wherein: The processor is used to execute the computer program stored on the memory; When the processor executes the computer program, the steps in the method according to any one of claims 1 to 6 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, the steps in the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Speech signal compression storage and reconstruction method based on superposition sequence

    CN108962265A

  • Gene sequencing quality row data compression preprocessing, decompression and reduction method and system

    CN110428868A