Compression program, compression method and compression device

By compressing protein database information using a reversible compression function and storing it on high-speed devices, the solution addresses the slow read speeds in large protein databases, enhancing search efficiency and speed.

JP2025091801APending Publication Date: 2025-06-19FUJITSU LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023207264
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-07
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

Large protein databases, such as the Big Fantastic Database (BFD), are stored in low-speed network file systems, leading to decreased read speeds during similarity searches.

Method used

A compression program and method that compresses protein information using a reversible compression function based on a defined compression policy, allowing for faster read speeds during similarity searches by utilizing high-speed storage devices for compressed data.

Benefits of technology

The proposed solution significantly improves read speeds during similarity searches by reducing data size and utilizing high-speed storage, while maintaining the ability to decompress and search the data efficiently.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025091801000001_ABST
    Figure 2025091801000001_ABST
Patent Text Reader

Abstract

To improve a reading speed in the case of performing similarity search.SOLUTION: A compression device acquires information of proteins expressed by encoded amino acid sequences. The compression device compresses the information of proteins on the basis of a compression policy that defines a relation between a condition of compressibility of the information of proteins or a condition of a time required in developing the information of proteins and a reversible compression function. The compression device registers the compressed information of proteins in a storage device. The compression device develops the compressed information of proteins registered in the storage device in the case of receiving access to the storage device.SELECTED DRAWING: Figure 2
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a compression program and the like.

Background Art

[0002] A protein, which is a kind of polymer compound, is chain-formed by 20 types of amino acids. For example, as a format for describing the amino acid sequence of a protein as a character string, the FASTA format is known. In the following description, the data of the protein in the FASTA format is appropriately referred to as "FASTA data".

[0003] FIG. 14 is a diagram showing an example of FASTA data. For example, the protein "PROTEIN A" composed of three amino acids, alanine ("A"), aspartate ("B"), and cystine ("C"), is represented by FASTA data 5.

[0004] For example, as technologies for performing a similarity search of an input sequence from a protein database storing FASTA data, there are HMMER, HH-suite3, and the like. FIG. 15 is a diagram for explaining an example of the similarity search of the prior art. In the example shown in FIG. 15, the FASTA data of the proteins "PROTEIN A", "PROTEIN B", and "PROTEIN C" are stored in the protein database 6.

[0005] Let the amino acid sequence of the protein "PROTEIN A" be "ABC". Let the amino acid sequence of the protein "PROTEIN B" be "DEF". Let the amino acid sequence of the protein "PROTEIN C" be "GHI".

[0006] For example, when "AAA" is specified as the input sequence 7, the protein "PROTEIN A" in which the first character of the amino acid sequence matches is searched.

[0007] The actual protein database 6 is, for example, BFD (Big Fantastic Database), which stores FASTA data of 2.5 billion proteins. BFD stores FASTA data in an uncompressed state, and the data size is 1.7 TiB.

Prior Art Documents

Patent Documents

[0008]

Patent Document 1

Patent Document 2

Patent Document 3

Patent Document 4

Summary of the Invention

Problems to be Solved by the Invention

[0009] Huge protein databases such as the above-mentioned BFD are stored in a large-capacity but low-speed network file system, etc., so the read speed decreases when performing a similarity search. Therefore, there is room for improvement in improving the read speed when performing a similarity search.

[0010] In one aspect, an object of the present invention is to provide a compression program, a compression method, and a compression device that can improve the read speed when performing a similarity search.

Means for Solving the Problems

[0011] In the first proposal, the computer executes the following processes. The computer acquires information on a protein represented by an encoded amino acid sequence. The computer compresses the protein information based on a compression policy that defines the relationship between the condition of the compression rate of the protein information or the time required to decompress the compressed protein information and a reversible compression function. The computer registers the compressed protein information in a storage device. When the computer receives access to the storage device, it decompresses the compressed protein information registered in the storage device.

Advantages of the Invention

[0012] The reading speed in the case of performing a similarity search can be improved.

Brief Description of the Drawings

[0013]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

BEST MODE FOR CARRYING OUT THE INVENTION

[0014] Hereinafter, embodiments of the compression program, compression method, and compression device disclosed in the present application will be described in detail with reference to the drawings. Note that the present invention is not limited by this embodiment.

EXAMPLE

[0015] An example of the processing of the compression device according to this embodiment will be described. The compression device reversibly compresses protein data (such as FASTA data) represented by an encoded amino acid sequence based on a preset compression policy, thereby reducing the data size of the protein database. This makes it possible to store the protein database in a storage device that has a small capacity but can access the data at high speed.

[0016] Storage devices that have a small capacity but can be accessed at high speed include a CPU (Central Processing Unit) memory, an SSD (Solid State Drive), etc. Note that the reversibly compressed protein database shall be decompressed when a similarity search is received.

[0017] In the following description, a storage device that has a large capacity but slow access to data is referred to as the "first storage device". The first storage device is, for example, an HDD (Hard Disk Drive), a network file system, or the like. On the other hand, storage devices such as the above-mentioned CPU memory and SSD are referred to as the "second storage device". Also, reversible compression is simply referred to as "compression". A compressed protein database is referred to as a "compressed database".

[0018] When the compression device receives an access (a request for a similar search) to the compressed database stored in the second storage device, it quickly reads out the compressed database in the second storage device, decompresses it, and then performs a similar search. By doing so, the read speed when performing a similar search can be improved.

[0019] FIG. 1 is a diagram for explaining the breakdown of the time required for a similar search. For example, in the conventional method, when reading out a slow DB (protein database) from the first storage device and performing a search, the time required for the similar search is t1 hours. In contrast, in the compression device of the present invention, when reading out a high-speed DB (compressed database) from the second storage device, decompressing it, and performing a search, the time required for the similar search is t2 hours. Even considering the time for pre-compression that is executed only for the first time, the time required for the similar search by the compression device is shorter than the time required for the similar search by the conventional method.

[0020] That is, according to the compression device according to the present embodiment, although an overhead due to decompression of the compressed data occurs, by reducing the data size of the entire protein database, the second storage device can be used. As a result, the read speed when performing a similar search can be improved.

[0021] Subsequently, the preconditions of this embodiment will be described. The preconditions are as shown in 1 to 6 below. 1. Let the set consisting of byte sequences be B, and let the number of bytes of the byte sequence b ∈ B be |B| ∈ N (Natural number). 2. Let the set consisting of the amino acid sequences of proteins be Bp Let it be ⊂B. In the following description, the amino acid sequence of a protein is denoted as "sequence". For example, the sequence "ABC" means "ABC" ∈ B p becomes 3. The set S of reversible compression functions applied to each sequence during pre-compression comp is defined as shown in Equation (1).

Equation

Equation

Equation

Number

Number

[0022] In the above description, a function that performs Deflate compression is shown as the reversible compression function, but it is not limited to this, and functions such as LZ77 / LZ78 compression, Bzip2, LZF, and Snappy may also be used. Also, although the compression policy is set as Equation (3), it is not limited to this. The compression device may use any of the other compression policies 1 to 4 described below.

[0023] Another compression policy 1 will be described. Another compression policy 1 performs compression by the reversible compression function f only when the compression rate is n ≧ 1 times or more, and is shown in Equation (6).

Number

[0024] Another compression policy 2 will be described. Another compression policy 2 uses the one with the higher compression rate among the reversible compression functions f1 and f2, and is shown in Equation (7).

Number

[0025] Another compression policy 3 will be described. For another compression policy 3, with the target compression ratio being n≥1, a plurality of reversible compression functions with different compression efficiencies are used. The plurality of reversible compression functions with different compression efficiencies are represented by Equation (8). Another compression policy 3 is represented by Equation (9) using a plurality of reversible compression functions with different compression efficiencies. However, it is assumed that a reversible compression function with a higher compression ratio generally has a longer expansion time. When the size (sizeA) of the database before compression and the size (sizeB) of the storage to be used are known, the compression device may determine the target compression ratio n by the ratio "sizeA / sizeB".

Number

Number

[0026] Another compression policy 4 will be described. For another compression policy 4, among the reversible compression functions f1 and f2, the reversible compression function with the shorter time required to expand the compressed array p is used. For example, the time required to expand the array p compressed by the reversible compression function f1 is represented by Equation (10). The time required to expand the array p compressed by the reversible compression function f2 is represented by Equation (11). Another compression policy 4 is represented as shown in Equation (12) using Equation (10) and Equation (11).

Number

Number

Number

[0027] The above has described the other compression policies 1 to 4. For example, when the protein database consists of a combination of a plurality of different database files, a separate compression policy may be applied to each database file. Also, a compression policy may be applied to some of the database files. Note that the compression policy can be said to be information that defines the relationship between the condition of the compression rate of the array or the condition of the time to decompress the compressed array and the reversible compression function.

[0028] Subsequently, the information after compression for several arrays p is shown in FIG. 2. FIG. 2 is a diagram for explaining the information after compression for several arrays p. However, the conditions for compressing the array p are as shown in 1 to 4 below. 1. f def is a function that performs the default Deflate compression implemented by zlib v1.2.11. 2. The set of reversible compression functions is the set shown in Equation (2). 3. The compression policy is the compression policy shown in Equation (13). The compression policy shown in Equation (13) corresponds to setting the value of n in the compression policy shown in the above Equation (6) to 1.5. That is, the compression device compresses when the compression rate is 1.5 times or more. [Number] 4. Set the sign of id to "0" and the sign of f def to "1".

[0029] Explain the first row of Table 10 in FIG. 2. The array p is "AAAAAAAAAAAAAAAAAAAA". When such an array p is compressed by f def , f def (p) becomes "78 9C 73 74 C4 04 00 (omitted hereinafter)". |p| becomes "20", and |f def (p)| becomes "11", so the compression rate becomes "1.82". The compression device selects the reversible compression function f def based on the compression policy and replaces the array p with (c(p), s(p)). c(p) is fdef It is the same as (p). "1" is set for s(p).

[0030] Describe the second row of Table 10 in Figure 2. The array p is "ABCDEABCDEABCDEABCDE". Compress such an array p with f def Then f def (p) becomes "78 9C 73 74 72 76 71 (omitted hereinafter)". |p| is "20", and |f def (p)| is "15", so the compression ratio is "1.33". The compression device selects id based on the compression policy and replaces the array p with (c(p), s(p)). c(p) is the same as the array p. "0" is set for s(p).

[0031] Describe the third row of Table 10 in Figure 2. The array p is "ABCDEFGHIJKLMNOPQRSTUVWXYZ". Compress such an array p with f def Then f def (p) becomes "78 9C 73 74 72 76 71 (omitted hereinafter)". |p| is "26", and |f def (p)| is "34", so the compression ratio is "0.76". The compression device selects id based on the compression policy and replaces the array p with (c(p), s(p)). c(p) is the same as the array p. "0" is set for s(p).

[0032] Describe the fourth row of Table 10 in Figure 2. The array p is "MLEADDQGCIEEQGVEDSAN (omitted hereinafter)". Compress such an array p with f def Then f def (p) becomes "78 9C 0D 8F CB 0D 44 (omitted hereinafter)". |p| is "305", and |f def (p)| is "195", so the compression ratio is "1.56". The compression device selects the reversible compression function f def based on the compression policy and replaces the array p with (c(p), s(p)). c(p) is the same as f def (p). "1" is set for s(p).

[0033] Explanation will be given for the fifth row of Table 10 in FIG. 2. The array p is "MGDGGEGEDEVQFLRTDDEV (hereinafter omitted)". When such an array p is compressed by f def , f def (p) becomes "78 9C ED 56 51 CE 34 (hereinafter omitted)". |p| becomes "5037", and |f def (p)| becomes "1016", so the compression ratio becomes "4.96". The compression device selects a reversible compression function f def based on the compression policy and replaces the array p with (c(p), s(p)). c(p) is the same as f def (p). "1" is set for s(p).

[0034] Next, a configuration example of the compression device that executes the above-described processing will be described. FIG. 3 is a functional block diagram showing the configuration of the compression device according to the present embodiment. As shown in FIG. 3, this compression device 100 includes a communication unit 110, an input unit 120, a display unit 130, a first storage unit 140, a second storage unit 150, and a control unit 160.

[0035] The communication unit 110 executes data communication with an external device or the like via a network. The communication unit 110 is a NIC (Network Interface Card) or the like. For example, the compression device 100 may acquire data to be stored in the protein database 141 from an external device.

[0036] The input unit 120 is an input device that inputs various types of information to the control unit 160 of the compression device 100. For example, the input unit 120 corresponds to a keyboard, a mouse, a touch panel, or the like. When performing a similarity search, the user operates the input unit 120 to specify an input array.

[0037] The display unit 130 is a display device that displays information output from the control unit 160. For example, the display unit 130 displays the search results of a similarity search.

[0038] The first storage unit 140 has a protein database 141. The first storage unit 140 is a storage device with a large capacity but slow access to data. The first storage unit 140 is an HDD, a network file system, or the like.

[0039] The protein database 141 is a database that stores a plurality of sequences.

[0040] The second storage unit 150 has a compressed database 151 and a similar sequence database 152. The second storage unit 150 is a storage device with a small capacity but fast access. The second storage unit 150 is a CPU memory, an SSD, or the like.

[0041] The compressed database 151 is a database that stores compressed sequences (compressed sequences).

[0042] The similar sequence database 152 is a database that stores sequences similar to the input sequence. Note that the similar sequence database 152 does not necessarily have to be stored in the second storage unit 150 and may be stored in the first storage unit 140.

[0043] The control unit 160 has a compression processing unit 161 and a similarity search unit 162. The control unit 160 is a CPU, a GPU (Graphics Processing Unit), or the like.

[0044] The compression processing unit 161 acquires a sequence from the protein database 141 and selects a reversible compression function based on a pre-set compression policy. The compression processing unit 161 compresses the sequence with the selected reversible compression function and stores the compressed sequence in the compressed database 151. For example, the compression processing unit 161 performs compression that replaces the sequence p with the compressed sequence (c(p), s(p)).

[0045] Here, an example of the processing procedure of the compression processing unit 161 will be described. FIG. 4 is a flowchart showing the processing procedure of the compression processing unit according to the present embodiment. As shown in FIG. 4, the compression processing unit 161 of the compression device 100 initializes the compression database 151 (step S101).

[0046] The compression processing unit 161 sets 1 to i (step S102). When N (a preset numerical value) is set to i in the compression processing unit 161 (step S103, Yes), the processing ends. On the other hand, when N is not set to i in the compression processing unit 161 (step S103, No), the process proceeds to step S104.

[0047] The compression processing unit 161 reads the i-th array p from the protein database 141 (step S104). The compression processing unit 161 executes compression processing on the array p (step S105). The compression processing unit 161 writes the compressed array p' to the compression database 151 (step S106).

[0048] The compression processing unit 161 increments i by 1 (step S107) and proceeds to step S103.

[0049] Subsequently, the processing procedure of the compression processing shown in step S105 of FIG. 4 will be described. FIG. 5 is a flowchart showing the processing procedure of the compression processing. As shown in FIG. 5, the compression processing unit 161 calculates q by compressing the array p by f def (step S201).

[0050] When |p| / |q|≥n in the compression processing unit 161 (step S202, Yes), it sets "01" to s(p) (step S203), sets q to c(p) (step S204), and proceeds to step S207.

[0051] On the other hand, when |p| / |q|≧n does not hold (step S202, No), the compression processing unit 161 sets "00" in s(p) (step S205), sets p in c(p) (step S206), and proceeds to step S207.

[0052] The compression processing unit 161 concatenates s(p) and c(p) to obtain a compressed array p' (step S207).

[0053] Return to the description of FIG. 3. When the similarity search unit 162 receives a specification of an input array from the input unit 120 or the like, it expands the compressed arrays in the compressed database 151 and compares the expanded arrays with the input array. The similarity search unit 162 stores, in the similarity array database 152, arrays in the expanded arrays that are similar to the input array. The similarity search unit 162 outputs the information in the similarity array database 152 to the display unit 130 for display as a similarity search result.

[0054] The similarity search unit 162 identifies the reversible compression function used for array compression based on the numerical value set in s(p) of the compressed array, and restores the array by inverse transformation of the identified reversible compression function.

[0055] The similarity search unit 162 searches for arrays in the expanded arrays that are similar to the input array using search algorithms implemented in existing software such as HMMER and HH-suite3. For example, the similarity search unit 162 evaluates the degree of coincidence between the target array and the input array, and determines that the target array is similar to the input array when the degree of coincidence is equal to or greater than a threshold value.

[0056] Here, an example of the processing procedure of the similarity search unit 162 will be described. FIG. 6 is a flowchart showing the processing procedure of the similarity search unit according to the present embodiment. As shown in FIG. 6, the similarity search unit 162 of the compression device 100 receives an input array from the input unit 120 (step S301).

[0057] The similarity search unit 162 initializes the similarity array database (step S302). The similarity search unit 162 sets i to 1 (step S303). When i is set to N (a preset numerical value) (step S304, Yes), the similarity search unit 162 outputs the similarity array database to the display unit 130 for display (step S311).

[0058] On the other hand, when i is not set to N (step S304, No), the similarity search unit 162 proceeds to step S305. The similarity search unit 162 reads the i-th compressed array p' from the compressed database 151 (step S305).

[0059] The similarity search unit 162 executes the decompression process (step S306). The similarity search unit 162 calculates the degree of match between the array p and the input array (step S307). When the degree of match is equal to or greater than the threshold value (step S308, Yes), the similarity search unit 162 writes the array p to the end of the similarity array database (step S309) and proceeds to step S310.

[0060] On the other hand, when the degree of match is not equal to or greater than the threshold value (step S308, No), the similarity search unit 162 proceeds to step S310.

[0061] The similarity search unit 162 increments i by 1 (step S310) and proceeds to step S304.

[0062] Subsequently, the processing procedure of the decompression process shown in step S306 of FIG. 6 will be described. FIG. 7 is a flowchart showing the processing procedure of the decompression process. As shown in FIG. 7, the similarity search unit 162 of the compression device 100 extracts s(p) and c(p) included in the compressed array p' (step S401).

[0063] When the value of s(p) is "01" (step S402, Yes), the similarity search unit 162 restores the array p by inverse transformation of the reversible compression function that performs Deflate compression (step S403) and outputs the array p (step S405).

[0064] On the other hand, when the value of s(p) is not "01" (step S402, No), the similarity search unit 162 sets c(p) as the array p (step S404) and proceeds to step S405.

[0065] As described above, with reference to FIGS. 3 to 7, an example of the configuration example and processing procedure of the compression device 100 according to the present embodiment has been described.

[0066] Next, a case where the above compression device 100 is applied to a database in the FFindex format will be described. The FFindex format is a format used in an actual database.

[0067] FFindex is a format for storing one or more variable-length named binary data, and is composed of a.FFdata file and a.FFindex file. The.FFdata file is a file that stores binary data concatenated with '\0' delimiters. The.FFindex file is a tab-separated file that lists the name, offset (including '\0'), and protein length (including '\0') for each row. The name is the name of the corresponding protein. The offset is the position information of the array (binary data) of the corresponding protein, indicating the offset from the beginning. The length is the length of the array of the corresponding protein.

[0068] FIG. 8 is a diagram showing an example of a.FFdata file and a.FFindex file. The data in the first row of the.FFdata file 10a shown in FIG. 8 is data obtained by concatenating the amino acid sequence "ABC" of the protein "PROTEIN A" as binary data. "41", "42", "43" are the hexadecimal notations of A, B, and C. "00" is the hexadecimal notation of '\0'.

[0069] The data on the first line of the.FFindex file 10b has the name "PROTEIN A", offset "0", and length "4" set. That is, it is shown that the array at offset "0" and length "4" in the.FFdata file 10a is "PROTEIN A".

[0070] The data on the second line of the.FFdata file 10a is the data obtained by concatenating the amino acid sequence "DEF" of the protein "PROTEIN B" as binary data. "44", "45", "46" are the hexadecimal notations of D, E, and F. "00" is the hexadecimal notation of '\0'.

[0071] The data on the second line of the.FFindex file 10b has the name "PROTEIN B", offset "4", and length "4" set. That is, it is shown that the array at offset "4" and length "4" in the.FFdata file 10a is "PROTEIN B".

[0072] The data on the third line of the.FFdata file 10a is the data obtained by concatenating the amino acid sequence "GHI" of the protein "PROTEIN C" as binary data. "47", "48", "49" are the hexadecimal notations of G, H, and I. "00" is the hexadecimal notation of '\0'.

[0073] The data on the third line of the.FFindex file 10b has the name "PROTEIN C", offset "8", and length "4" set. That is, it is shown that the array at offset "8" and length "4" in the.FFdata file 10a is "PROTEIN C".

[0074] As shown in FIG. 8, since PROTEIN A, PROTEIN B, and PROTEIN C are all 3 bytes, they are concatenated with a length of 4 bytes including '\0'.

[0075] For example, assume that the conditions when applying the above compression device 100 to a database in the FFindex format are shown as follows from 1 to 4. 1. Let the set S of reversible compression functions be the set shown in Equation (2). comp 2. Let the compression policy be the policy shown in Equation (6) and set n = 1.5. 3. As the code to be set for s(p), set the code of id to "00" and the code of f to "01". def 4. Append the size of the data before compression (excluding '\0') to the end of each line of the.FFindex file. By adding the size of the data before compression, it is possible to grasp the memory size required after decompression before decompression and reduce the number of memory copy operations. Figure 9 is a diagram showing an example when a compression device is applied to a database in the FFindex format. As shown in Figure 9, let the.FFdata file before compression be the.FFdata file 20a, and the.FFindex file before compression be the.FFindex file 20b. On the other hand, let the.FFdata file after compression be the.FFdata file 21a, and the.FFindex file after compression be the.FFindex file 21b.

[0076] For example, the binary data "41×20" included in the.FFdata file 20a indicates 20 repetitions of the amino acid code "A" (hereinafter, A20). In the.FFindex file 20b, as the index related to A20, the name "A20", the offset "0", and the length "20" are set.

[0077] Since the compression ratio of the compression device 100 for A20 is 1.5 or more, as a reversible compression function, "f

[0078] " is selected. The compression device 100 compresses A20 with f def def ​By compressing, c(p) = "78 9C 73 74 C4 04 00 35 66 05 15" is generated. The compression device 100 sets the pair of s(p) = 01 (code 21a - 1) and c(p) = "78 9C 73 74 C4 04 00 35 66 05 15" in the.FFdata file 21a. Also, the compression device 100 sets the delimiter "00" after c(p) = "78 9C 73 74 C4 04 00 35 66 05 15".

[0079] Based on the compression result regarding A20, the compression device 100 registers the data in the first line of the.FFindex file 21b. For example, the compression device 100 sets the name "A20", offset "0", length "13", and the data size before compression "20".

[0080] The binary data "41 42 43 44 45×5" included in the.FFdata file 20a indicates 4 repetitions of the amino acid symbol sequence "ABCDE" (hereinafter, ABCDE). In the.FFindex file 20b, as the index regarding ABCDE, the name "ABCDE", offset "21", and length "21" are set.

[0081] Since the compression ratio of ABCDE is less than 1.5, the compression device 100 selects "id" as the reversible compression function. The compression device 100 compresses ABCDE with id (identity mapping) to generate c(p) = "41 42 43 44 45×4". The compression device 100 sets the pair of s(p) = 00 (code 21a - 2) and c(p) = "41 42 43 44 45×4" in the.FFdata file 21a. Also, the compression device 100 sets the delimiter "00" after c(p) = "41 42 43 44 45×4".

[0082] Based on the compression result regarding ABCDE, the compression device 100 registers the data in the second line of the.FFindex file 21b. For example, the compression device 100 sets the name "ABCDE", offset "22", length "22", and the data size before compression "20".

[0083] As described above, by compressing the compression device 100 from the.FFdata file 20a to the.FFdata file 21a, the size is reduced from 42 bytes to 35 bytes.

[0084] Note that when the value of n is changed to 1, 1.5, 2 or more, the size of the.FFdata file 20a shown in FIG. 9 becomes as shown in FIG. 10. FIG. 10 is a diagram for explaining the size of the.FFdata file according to n.

[0085] The record in the first row of Table 30 in FIG. 10 will be described. Before compression, the number of bytes of A20 is "21", the number of bytes of ABCDE is "21", and the file size of the.FFdata file 20a is "42".

[0086] The record in the second row of Table 30 in FIG. 10 will be described. When the compression device 100 is set to n ≧ 2, "id" is selected as the reversible compression function for A20 and ABCDE, and A20 and ABCDE are not compressed (identity mapping). As a result, the number of bytes of A20 becomes "22", the number of bytes of ABCDE becomes "22", and the file size of the compressed.FFdata file 21a becomes "44".

[0087] The record in the third row of Table 30 in FIG. 10 will be described. When the compression device 100 is set to n = 1.5, "f def " is selected as the reversible compression function for A20, and "id" is selected as the reversible compression function for ABCDE. The compression device 100 compresses A20 by f def , and the number of bytes of A20 becomes "13". ABCDE is not compressed (identity mapped), and the number of bytes becomes "22". As a result, the file size of the compressed.FFdata file 21a becomes "35".

[0088] The record in the fourth row of Table 30 in FIG. 10 will be described. When the compression device 100 is set to n = 1, as the reversible compression function for A20 and ABCDE, "f def " is selected. The compression device 100 compresses A20 by f def , and the number of bytes of A20 becomes "13". The compression device 100 compresses ABCDE by f def , and the number of bytes of ABCDE becomes "17". As a result, the file size of the compressed.FFdata file 21a becomes "30".

[0089] When the value of n is changed to 1, 1.5, or 2 in the compression device 100, the relationship between the value of n and the file size of the.FFdata file 21a may be output to and displayed on the display unit 130. Further, the compression device 100 may also display the capacity of the second storage unit 150 together.

[0090] For example, assume that the capacity of the second storage unit 150 is 35 bytes. As described with reference to FIG. 10, before compression and when n ≥ 2, the file sizes of the.FFdata file 21a are "42" and "44" respectively, so they cannot be stored in the second storage unit 150.

[0091] When n = 1.5, since the file size of the.FFdata file 21a is "35", it can be stored in the second storage unit 150.

[0092] When n = 1, since the file size of the.FFdata file 21a is "30", it can be stored in the second storage unit 150. However, when n = 1, compared with the case of n = 1.5, ABCDE is compressed, so the time cost of expansion during similarity search increases.

[0093] Therefore, when the reading of the second storage unit 150 is fast enough, it can be said that n = 1.5 is the optimal setting. The compression device 100 may highlight the maximum value of n among the values of n for which the file size of the.FFdata file 21a fits within the capacity of the second storage unit 150. For example, when the capacity of the second storage unit 150 is 35 bytes, the compression device 100 may highlight n = 1.5.

[0094] Next, the effects of the compression device 100 according to this embodiment will be described. The compression device 100 acquires protein information and compresses the protein information based on a compression policy. The compression device 100 registers the compressed protein information in the second storage unit 150. When the compression device 100 receives an access related to a similarity search request, it decompresses the compressed protein information registered in the second storage unit 150. As a result, protein information such as a protein database can be stored in the second storage unit 150, which has a small capacity but allows for fast access to the data, and the reading speed when performing a similarity search can be improved.

[0095] For example, the compression policy defines a plurality of reversible compression functions with different compression ratios. The compression device 100 selects a reversible compression function that maximizes the compression ratio when compressing protein information, and compresses the protein information based on the selected reversible compression function. Thereby, the compression ratio of the protein can be increased.

[0096] For example, the compression policy is defined such that Delfate compression is executed when the compression ratio is equal to or greater than a threshold value, and an identity mapping is executed when the compression ratio is less than the threshold value. The compression device 100 selects Delfate compression or an identity mapping to compress the protein information. As a result, information with a compression ratio equal to or greater than the threshold value can be Delfate compressed, and information with a compression ratio less than the threshold value can be subjected to an identity mapping. In addition, since information with a low compression ratio is subjected to an identity mapping, the processing time at the time of decompression can be reduced.

[0097] For example, the compression policy defines a plurality of reversible compression functions that require different times to expand the information of the compressed protein. The compression device 100 selects a reversible compression function that requires the minimum time to expand the information of the compressed protein, and compresses the protein information based on the selected reversible compression function. As a result, the time required for similarity search can be further shortened.

[0098] Subsequently, the compression method in the FFindex format described in this embodiment is applied to the a3m and hmm of the Uniclust30 database, and the results of evaluating the effect when performing similarity search using hhblits (Hidden Markov Model HMM-based BLAST) will be described. FIG. 11 is a diagram showing the results of applying a3m and hmm of the Uniclust30 database.

[0099] As shown in the first row of Table 35 in FIG. 11, before compression, the size of cs219.ffdata is "3.6 GiB", the size of a3m.ffdata is "64.7 GiB", the size of hhm.ffdata is "13.2 GiB", and the total is "81.5 GiB".

[0100] As shown in the second row of Table 35 in FIG. 11, when n = 4, the size of cs219.ffdata is "3.6 GiB", the size of a3m.ffdata is "35.8 GiB", the size of hhm.ffdata is "11.6 GiB", and the total is "50.9 GiB".

[0101] As shown in the third row of Table 35 in FIG. 11, when n = 2, the size of cs219.ffdata is "3.6 GiB", the size of a3m.ffdata is "16.9 GiB", the size of hhm.ffdata is "4.2 GiB", and the total is "24.6 GiB".

[0102] As shown in FIG. 11, as n decreases, the total size decreases. When n = 2, the total size is reduced to about 0.3 times that before compression.

[0103] In this experiment, a similarity search (executing hhblits) was performed on 6R83, which is the protein with the largest number of hits targeting the BFD·Uniclust30 database among the proteins in the PDB database included in the OpenProteinSet dataset.

[0104] Figure 12 is a diagram showing an example of the execution time when hhblits is executed. In graph 50 of Figure 12, the vertical axis represents time (execution time), and the horizontal axis represents conditions.

[0105] "NFS" indicates the case when all files of Uniclust30 are placed on the Lustre file system and executed, and the time is "688.0 s". Note that error bar 51 indicates the minimum and maximum values of the execution time. Although the illustration is omitted, the maximum value corresponding to NFS is "1962.1 s".

[0106] "SSD" indicates the case when all files of Uniclust30 are copied from the Lustre file system to the SSD of the compute node (compressor 100) and then executed, and the time is "167.5 s". Note that the time required for copying is "62.4 s".

[0107] "Memory" indicates the case when all files of Uniclust30 are copied to the memory ( / dev / shm) of the compute node and then executed, and the time is "167.4 s". Note that the time required for copying is "63.4 s".

[0108] "Memory + the present invention (n = 4)" indicates the case when all pre-compressed (n = 4) files of Uniclust30 are copied to the memory ( / dev / shm) of the compute node and then executed, and the time is "158.5 s". Note that the time required for copying is "41.4 s".

[0109] "Memory + the present invention (n = 2)" shows the case where the pre-compressed (n = 2) entire files of Uniclust30 are copied to the memory ( / dev / shm) of the computing node and then executed, and the time is "157.7 s". The time required for copying is "24.6 s".

[0110] From the above experimental results, the following can be said. Only for "NFS", the execution time of hhblits is extremely long. This is because the database reading is via the network, which has become a performance bottleneck.

[0111] The hhblits execution time of "Memory + the present invention" is almost equivalent to that of "SSD" and "Memory". This is because by using high-speed storage, parts other than the reading of a3m and hhm (and in the case of using the present invention, the expansion in addition) have become bottlenecks.

[0112] Since the Uniclust30 database has a smaller data size when n is smaller in the present invention, the copy time is shortened.

[0113] Assuming that the upper limit of the SSD and memory sizes available for holding the database is 64 GiB, "SSD" and "Memory" cannot be executed, and without the present invention, slow "NFS" has to be used. On the other hand, when using the present invention, it is possible to copy to the memory in both cases of n = 2 and 4, and the execution of hhblits is significantly accelerated. Specifically, "Memory + the present invention (n = 2)" is 4.36 times faster (only the hhblits execution time) or 3.77 times faster (overall) compared to "NFS".

[0114] Next, an example of the hardware configuration of a computer that realizes the same functions as the compression device 100 described above will be described. FIG. 13 is a diagram showing an example of the hardware configuration of a computer that realizes the same functions as the compression device of the embodiment.

[0115] As shown in FIG. 13, the computer 200 includes a CPU 201 that executes various arithmetic processes, an input device 202 that receives input of data from a user, and a display 203. The computer 200 also includes a communication device 204 that exchanges data with external devices or the like via a wired or wireless network, and an interface device 205. The computer 200 further includes a RAM 206 that temporarily stores various information, a hard disk device 207, and an SSD 208 that can be accessed at high speed. Each of the devices 201 to 208 is connected to a bus 209.

[0116] The hard disk device 207 has a compression processing program 207a and a similar search program 207b. The CPU 201 reads out each of the programs 207a and 207b and expands them in the RAM 206.

[0117] The compression processing program 207a functions as a compression processing process 206a. The similar search program 207b functions as a similar search process 206b.

[0118] The process of the compression processing process 206a corresponds to the process of the compression processing unit 161. The process of the similar search process 206b corresponds to the process of the similar search unit 162.

[0119] Note that each of the programs 207a and 207b does not necessarily have to be stored in the hard disk device 207 from the beginning. For example, each program is stored in a "portable physical medium" such as a flexible disk (FD), CD-ROM, DVD, magneto-optical disk, or IC card inserted into the computer 200. Then, the computer 200 may read out and execute each of the programs 207a and 207b.

[0120] Regarding the embodiments including the above embodiments, the following additional remarks are disclosed.

[0121] (Additional Remark 1) Obtaining information on a protein expressed by an encoded amino acid sequence, Based on a compression policy that defines the relationship between the conditions of the compression rate of the protein information or the time required to decompress the compressed protein information and a reversible compression function, compress the protein information, Register the compressed protein information in a storage device, When receiving access to the storage device, decompress the compressed protein information registered in the storage device A compression program characterized by causing a computer to execute the process.

[0122] (Appendix 2) The compression policy defines a plurality of reversible compression functions with different compression rates, and the process of compressing the protein information selects a reversible compression function with the maximum compression rate when compressing the protein information, and based on the selected reversible compression function, The compression program according to Appendix 1, characterized in that the protein information is compressed.

[0123] (Appendix 3) The compression policy is defined to execute Delfate compression when the compression rate is equal to or higher than a threshold value, and to execute an identity mapping when the compression rate is lower than the threshold value. The compression process is to select the Delfate compression or the identity mapping, The compression program according to Appendix 2, characterized in that the protein information is compressed.

[0124] (Appendix 4) The compression policy defines a plurality of reversible compression functions with different times required to decompress the compressed protein information, and the process of compressing the protein information selects a reversible compression function with the minimum time required to decompress the compressed protein information, and based on the selected reversible compression function, The compression program according to Appendix 1, characterized in that the protein information is compressed.

[0125] (Appendix 5) Obtain protein information represented by an encoded amino acid sequence, Based on a compression policy that defines the relationship between the condition of the compression rate of the protein information or the time required to decompress the compressed protein information and a reversible compression function, the protein information is compressed, the compressed protein information is registered in a storage device, when receiving access to the storage device, the computer executes a process of decompressing the compressed protein information registered in the storage device A compression method, characterized in that a computer executes the process.

[0126] (Appendix 6) The compression policy defines a plurality of reversible compression functions with different compression rates. The process of compressing the protein information selects a reversible compression function with the maximum compression rate when compressing the protein information, and compresses the protein information based on the selected reversible compression function. The compression method according to Appendix 5, characterized by this.

[0127] (Appendix 7) The compression policy is defined to execute Delfate compression when the compression rate is equal to or higher than a threshold value, and to execute an identity mapping when the compression rate is less than the threshold value. The compression process selects the Delfate compression or the identity mapping to compress the protein information. The compression method according to Appendix 6, characterized by this.

[0128] (Appendix 8) The compression policy defines a plurality of reversible compression functions with different times required to decompress the compressed protein information. The process of compressing the protein information selects a reversible compression function with the minimum time required to decompress the compressed protein information, and compresses the protein information based on the selected reversible compression function. The compression method according to Appendix 5, characterized by this.

[0129] (Appendix 9) Obtain protein information represented by an encoded amino acid sequence, Based on a compression policy that defines the relationship between the conditions of the compression rate of the protein information or the time required to decompress the compressed protein information and a reversible compression function, the protein information is compressed, the compressed protein information is registered in a storage device, when access to the storage device is received, the compressed protein information registered in the storage device is decompressed A compression device having a control unit that executes processing.

[0130] (Appendix 10) The compression policy defines a plurality of reversible compression functions with different compression rates, and the process of compressing the protein information selects a reversible compression function with the maximum compression rate when compressing the protein information, and based on the selected reversible compression function, the compression device according to Appendix 9, characterized in that the protein information is compressed.

[0131] (Appendix 11) The compression policy is defined to execute Delfate compression when the compression rate is equal to or higher than a threshold value, and to execute an identity mapping when the compression rate is less than the threshold value, and the compression process selects the Delfate compression or the identity mapping to compress the protein information, and the compression device according to Appendix 10, characterized in that the protein information is compressed.

[0132] (Appendix 12) The compression policy defines a plurality of reversible compression functions with different times required to decompress the compressed protein information, and the process of compressing the protein information selects a reversible compression function with the minimum time required to decompress the compressed protein information, and based on the selected reversible compression function, the compression device according to Appendix 9, characterized in that the protein information is compressed.

Explanation of Reference Numerals

[0133] 100 Compression device 110 Communication unit 120 Input unit 130 Display unit 140 First storage unit 141 Protein database 150 Second memory unit 151 Compressed database 152 Similar sequence database 160 Control unit 161 Compression processing unit 162 Similarity search unit

Claims

1. Obtain information on a protein represented by an encoded amino acid sequence, Based on a compression policy that defines the relationship between the condition of the compression rate of the protein information or the time required to decompress the compressed protein information and a reversible compression function, compress the protein information, Register the compressed protein information in a storage device, When receiving access to the storage device, decompress the compressed protein information registered in the storage device A compression program, characterized in that the computer is caused to execute the processing.

2. The compression policy defines a plurality of reversible compression functions with different compression rates, and the process of compressing the protein information selects a reversible compression function with the maximum compression rate when compressing the protein information, and based on the selected reversible compression function, The compression program according to claim 1, wherein the protein information is compressed.

3. The compression policy is defined to execute Delfate compression when the compression rate is equal to or higher than a threshold value, and to execute an identity mapping when the compression rate is lower than the threshold value. The compression process selects the Delfate compression or the identity mapping to compress the protein information. The compression program according to claim 2.

4. The compression policy defines a plurality of reversible compression functions with different times required to decompress the compressed protein information. The process of compressing the protein information selects a reversible compression function with the minimum time required to decompress the compressed protein information, and based on the selected reversible compression function, The compression program according to claim 1, wherein the protein information is compressed.

5. Obtain information on a protein represented by an encoded amino acid sequence, Based on a compression policy that defines the relationship between the condition of the compression rate of the protein information or the time required to decompress the compressed protein information and a reversible compression function, the protein information is compressed, the compressed protein information is registered in a storage device, when access to the storage device is received, the compressed protein information registered in the storage device is decompressed A compression method characterized in that a computer executes the processing.

6. obtain protein information represented by an encoded amino acid sequence, Based on a compression policy that defines the relationship between the condition of the compression rate of the protein information or the time required to decompress the compressed protein information and a reversible compression function, the protein information is compressed, the compressed protein information is registered in a storage device, when access to the storage device is received, the compressed protein information registered in the storage device is decompressed A compression device having a control unit that executes the processing.

Citation Information

Patent Citations

  • DNA sequence encoder and method

    JP2004240975A

  • Genome analysis program, recording medium with this program recorded, genome analysis device and genome analysis method

    JP2007193708A

  • Genome compression and decompression

    US20160306919A1

  • Gene sequencing data compression method and decompression method, system and computer-readable medium

    US20200294629A1