Data compression method and related device
Patent Information
- Application Number
- CN202380069055.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-09-29
- Filing Date
- 2023-08-08
- Publication Date
- 2025-05-06
AI Technical Summary
Existing compression tools have unsatisfactory compression effects on target files, resulting in larger software package sizes and affecting user experience.
By determining the constant strings and special string information in the software package, replacing them with identifiers, and combining file similarity grouping, customized compression algorithms and bit stream allocation strategies are used to compress target files and non-target files in the software package.
Effectively reduce the size of software packages, improve user experience, shorten compression time, improve compression efficiency, and reduce computing resource consumption.
Smart Images

Figure CN119948471A_ABST
Abstract
Description
Method and related device for compressing data
[0001] This application claims priority to the Russian Federation patent application filed with the Russian Federal Patent Office on September 29, 2022, with application number 2022125457 and application name “Method and related device for compressing data”, the entire content of which is incorporated herein by reference. Technical Field
[0002] The embodiments of the present application relate to the field of information technology, and more specifically, to a method and related apparatus for compressing data. Background Art
[0003] A software package is a collection of files and directories required for a software product. Software packages are typically designed and generated by application developers after the application code is developed. Software products need to be organized into one or more packages so they can be easily distributed and installed.
[0004] Object files are an important part of a software package. They contain the object code. Object code is the code generated by a compiler or assembler after processing the source code. Object code typically consists of machine code or code close to machine language.
[0005] Currently, the compression effect of commonly used compression tools on target files is not ideal. Therefore, how to better compress target files is a problem that needs to be solved in the industry.
[0006] Summary of the Invention
[0007] The embodiments of the present application provide a method and related apparatus for compressing data, which can reduce the size of a software package and improve user experience.
[0008] In a first aspect, an embodiment of the present application provides a method for compressing data, comprising: determining N target files included in a software package, where N is a positive integer greater than or equal to 1; determining constant string information, where the constant string information is used to indicate at least one constant string and a constant string identifier corresponding to each constant string in the at least one constant string, and each target file in the N target files includes the at least one constant string; determining N special string information, where the N special string information corresponds one-to-one with the N target files, where first special string information is used to indicate at least one special string in the first target file and a special string identifier corresponding to each special string in the at least one special string, where the first special string information is any one of the N special string information, and the first target file is the target file corresponding to the first special string; replacing the constant string and special string of each of the N target files with the corresponding identifier to obtain N replaced target files; and compressing the software package according to first information to be compressed, where the first information to be compressed includes the constant string information, the N special string information, and the N replaced target files.
[0009] The size of a software package determines the user experience. Larger packages take longer to download; smaller packages take less time. The above technical solution reduces the size of the software package by compressing the target files within it, allowing users to download / transfer the software package faster, thereby improving the user experience.
[0010] Optionally, the software package may be a software package within an integrated development environment (IDE). For example, the software package may be an installation package for the IDE's main program or an installation package for an IDE extension. The IDE may be a traditional IDE running on a local computer or a cloud IDE (e.g., an online integrated development environment or a web IDE).
[0011] When using an IDE for cross-environment development (e.g., remote development or development using a cloud IDE), software packages frequently need to be transferred between multiple physical environments (e.g., between multiple computers or systems). The above technical solution can identify and compress target files within a software package, reducing its size, facilitating transfer, and improving the user experience.
[0012] Optionally, the software package may be a compressed file. In this case, the software package may be decompressed to obtain a non-compressed file, and then the target file in the software package may be determined.
[0013] In combination with the first aspect, in a possible implementation of the first aspect, the method further includes: determining M non-target files in the software package, where M is a positive integer greater than or equal to 1; grouping the non-target files to obtain third information to be compressed, where the third information to be compressed includes at least one file set, where files belonging to the same file set have the same characteristics; and compressing the software package based on the first information to be compressed includes: compressing the first information to be compressed and the second information to be compressed to obtain a compressed software package.
[0014] Compression algorithms typically predict Y bits after the current bit based on the X bits preceding it. If this fails, they try X-1 bits until they succeed. Therefore, compressing similar content (files) together improves overall prediction accuracy, thereby shortening compression time and increasing the compression ratio. Furthermore, grouping files based on similarity and applying customized compression methods to specific files can further improve the compression ratio.
[0015] In combination with the first aspect, in a possible implementation of the first aspect, the multiple file sets include a first file set and at least one small file, where the small file is a non-target file among the M non-target files whose size is less than or equal to a file size threshold.
[0016] From the perspective of saving computing resources and improving compression efficiency, not compressing small files allows for faster package compression without significantly impacting the final package size. The above technical solution groups small files that are not part of the target file into a group. In subsequent processing, the small files are not compressed, but directly combined with the compressed files to form a compressed package.
[0017] In combination with the first aspect, in a possible implementation of the first aspect, the multiple file sets include at least one second file set, wherein multiple non-target files belonging to the same file set have the same extension, the same encoding method, and / or the same file type.
[0018] In combination with the first aspect, in a possible implementation of the first aspect, before compressing the first information to be compressed and the second information to be compressed to obtain a compressed software package, the method further includes: determining K compression workloads, the K compression workloads corresponding one-to-one to K bit streams, each of the K bit streams including part or all of files from the same object to be compressed, wherein the object to be compressed includes the constant string information, the special string information, the replaced target file, and the file set, and K is a positive integer greater than or equal to 2; and allocating the K bit streams to P operation units for compression based on the K compression workloads, wherein the difference between the first workload and the second workload is less than a workload threshold, wherein the first workload is the sum of the workloads of the bit streams allocated to the first operation unit, and the second workload is the sum of the workloads of the bit streams allocated to the second operation unit, the first operation unit and the second operation unit are any two operation units among the P operation units, and P is a positive integer greater than or equal to 2.
[0019] The above technical solution allocates different bit streams to different operation units for compression based on the compression workload of the bit stream, which can further shorten the compression time.
[0020] In combination with the first aspect, in a possible implementation of the first aspect, determining K compression workloads includes: determining the compressibility score of the i-th bit stream among the K bit streams, i = 1,…, K; determining the i-th compression workload among the K compression workloads based on the compressibility score of the i-th bit stream and the size of the i-th bit stream.
[0021] In combination with the first aspect, in a possible implementation of the first aspect, the compressibility score of the i-th bit stream is the distance between an amplitude histogram of information included in the i-th bit stream and an amplitude histogram of Gaussian white noise.
[0022] The greater the randomness of the data, the more it conforms to a Gaussian distribution. Therefore, using a Gaussian white noise histogram can easily determine the randomness of a bitstream. The smaller the distance between the amplitude histogram and the Gaussian white noise histogram, the greater the randomness of the bitstream, the greater the compression, and the corresponding workload. Conversely, the greater the distance between the amplitude histogram and the Gaussian white noise histogram, the less random the bitstream, the less compression, and the corresponding workload.
[0023] In combination with the first aspect, in a possible implementation of the first aspect, determining the i-th compression workload among the K compression workloads according to the compressibility score of the i-th bitstream and the size of the i-th bitstream includes: determining the i-th compression workload according to the following formula: i =(1-Grd i )×Sizei , among which, Comp i is the i-th compression workload, Grd i is the compressibility score of the i-th bitstream, Size i is the size of the i-th bit stream.
[0024] In a second aspect, an embodiment of the present application provides a computer device, which includes a unit for implementing the first aspect or any possible implementation manner of the first aspect.
[0025] In a third aspect, an embodiment of the present application provides a computer device, which includes a processor, wherein the processor is used to couple with a memory, read and execute instructions and / or program codes in the memory to execute the first aspect or any possible implementation of the first aspect.
[0026] In a fourth aspect, an embodiment of the present application provides a chip system, which includes a logic circuit, which is used to couple with an input / output interface and transmit data through the input / output interface to execute the first aspect or any possible implementation method of the first aspect.
[0027] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores program code. When the computer storage medium runs on a computer, it enables the computer to execute the first aspect or any possible implementation of the first aspect.
[0028] In a sixth aspect, an embodiment of the present application provides a computer program product, which includes: a computer program code, which, when running on a computer, enables the computer to execute the first aspect or any possible implementation of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] FIG1 is a schematic flow chart of a method for compressing data provided by an embodiment of the present application.
[0030] FIG2 is a schematic flow chart of a method for compressing data provided in an embodiment of the present application.
[0031] FIG3 is a schematic structural block diagram of a computer device provided according to an embodiment of the present application. DETAILED DESCRIPTION
[0032] The technical solutions in the embodiments of the present application will be described below with reference to the accompanying drawings.
[0033] The computer device referred to in the embodiments of the present application may be a desktop computer, a laptop computer, a tablet computer, a server, or other computer device.
[0034] Figure 1 is a schematic flow chart of a method for compressing data provided by an embodiment of the present application. The method shown in Figure 1 can be executed by a computer device or a component in a computer device (such as a chip or system chip, etc.). For ease of description.
[0035] 101. Determine whether the software package is compressed. If the software package is compressed, execute steps 102 and 103; if the software package is not compressed, execute step 103 directly.
[0036] The embodiments of the present application do not limit the compression format of the software package. For example, the compression format of the software package can be jar format, zip format, rar format, etc.
[0037] In some embodiments, the software package may be a software package in an integrated development environment (IDE), for example, the software package may be an IDE main program installation package, or an IDE extension program installation package.
[0038] 102. Decompress the compressed software package to obtain a decompressed software package.
[0039] Depending on whether the file is compressed, the file can be divided into a compressed file and an uncompressed file. The embodiment of the present application does not limit the format of the compressed file. For example, the format of the compressed file can be a jar format, a zip format, a rar format, etc. An uncompressed file may include a target file, and may also include any one or more files other than the compressed formats such as jar, zip or rar. For example, an uncompressed file may include any one or more types of files: an executable file (such as a file with an extension of .exe), a library file (such as a file with an extension of .lib, .dll, .a or .so, etc.), a text file (such as a file with an extension of .txt, .doc, etc.), a sound file (such as a file with an extension of .mp3, .wav, .flac, etc.), a video file (such as a file with an extension of .mp4, .mkv, .avi, or .rmvb, etc.), or a picture file (such as a file with an extension of .jpg, .gif, .bmp, etc.), etc.
[0040] For ease of description, uncompressed files are referred to as program files in the embodiments of the present application.
[0041] 103. Determine whether the software package includes a compressed file.
[0042] In some embodiments, if the software package includes a compressed file, step 104 may be executed; if the software package does not include a compressed file, step 105 may be executed.
[0043] In other embodiments, if the software package includes compressed files and uncompressed files, step 105 may be performed on the uncompressed files first, and then step 105 may be performed after the compressed files are decompressed to obtain the uncompressed files (ie, step 104 is performed first).
[0044] 104. Decompress the compressed file.
[0045] The compressed files referred to in the embodiments of the present application may be compressed files obtained through single compression or files obtained through nested compression. If a compressed file is obtained through single compression, then decompressing the compressed file will result in an uncompressed file. If a compressed file is compressed through nested compression, then decompressing the compressed file will also result in a compressed file. The embodiments of the present application do not limit the number of nested layers. For example, the number of nested layers may be one, two, or more than two.
[0046] If all that is obtained after the compressed file is decompressed are uncompressed files (program files), then step 105 can be executed; if there are still compressed files after the compressed file is decompressed, then continue to decompress the compressed files until there are no more compressed files after decompression.
[0047] 105, determine the target file in the software package.
[0048] The target file can be identified by the file extension. For example, common target file extensions include .obj, .o, .class, etc.
[0049] Program files can be divided into target files and non-target files based on whether they are target files or not. For ease of description, assume that a software package contains N target files and M non-target files, where N and M are both positive integers greater than or equal to 1.
[0050] The non-target file may be any type of file, such as a text file, a video file, an audio file, an executable file, etc.
[0051] 106 , decompose and replace the N target files to obtain constant string information, N special string information, and N replaced target files.
[0052] Target files are generated by compiling source files. Because target files are closely related to the compilation system, their metadata contains a large number of standard strings defined by the compilation system. These strings include constant strings and special strings. Constant strings appear in every target file, while special strings appear in a specific target file.
[0053] It can be seen from this that N target files can have only one constant string information, and the constant string information includes a constant string that appears in each target file and a constant string identifier used to distinguish the constant string. N target files have N special string information, and the N special string information corresponds one-to-one to the N target files. Assume that the first target file is any target file among the N target files, and the first special string information is the special string information corresponding to the first target file. Then, the first special string information can include the special string contained in the first target file and the special string identifier used to distinguish the special string contained in the first target file. Each special string information in the N special string information can include an identity identifier, and the identity identifier is used to indicate in which target file the special string contained in the special string information appears.
[0054] As described above, each constant string has a corresponding constant string identifier, and each special string has a corresponding special string identifier. The constant string identifier can be X bits long, and the special string identifier can be Y bits long. Both X and Y are positive integers greater than or equal to 1. X and Y can be the same value.
[0055] In some embodiments, the values of X and Y may be predetermined. For example, X (or Y) may be equal to 8, 12, 16, 24, etc.
[0056] In some embodiments, the values of X and Y may be determined based on the number of constant strings. For example, assuming the number of constant strings is Num C , then X must be greater than or equal to Num C The number of bits in the corresponding binary number. For example, suppose the constant string has a length of 80. The binary representation of 80 is 1010000, which has seven bits. Therefore, the value of X can be a positive integer greater than or equal to 7. For example, X can be 7 or 8.
[0057] In some embodiments, the specific positions of the constant string identifier and the special string identifier can be used to distinguish whether the identifier is a constant string identifier or a special string identifier. For example, assuming that the values of X and Y are both 16, then the first 8 bits can be used to distinguish between the constant string identifier and the special string identifier. For example, the first 8 bits of the constant string identifier can be 00000000, and the first 8 bits of the special string identifier can be 11111111. In this way, it can be determined whether an identifier is a constant string identifier or a special string identifier based on the first 8 bits.
[0058] The constant and special strings can be determined from the metadata of the target files. The metadata of the target files is stored in the target file repository. Therefore, the metadata of the N target files can be queried from the target file repository to obtain the constant string and the special string for each target file.
[0059] For example, the metadata of the first target file contains the following information:
[0060] The metadata of the first target file mentioned above indicates that the first target file contains two standard string nodes, "References" and "Namepool". For the convenience of description, the metadata of the first target file only shows the constant string and special string of the node "References". As shown above, the node "References" contains 5 constant strings such as Methodref, corresponding to identifiers: 1-5; special strings are matched by regular expressions. The above regular expression "Regex:(?!#SN-)([AZ]+[0-9]+)\b" means that special strings include any string starting with SN and including letters A to Z and numbers 0 to 9. In addition, special strings need to filter any results containing the special characters '@', '.', and ' / '. That is to say, even if a string starts with SN and contains letters A to Z and numbers 0 to 9, if the string contains '@', '.', and ' / ', then the string is not a special string.
[0061] The constant string identifier can be determined based on the identifier of the constant string in the metadata. For example, if the identifier of the constant string Methodref in the metadata of the first target file is 1, then the constant string identifier of the constant string Methodref can be 0000 0000 0000 0001, where the first 8 bits are flag bits used to distinguish constant strings from special strings, and the last 8 bits are the identifier of the constant string Methodref in the metadata.
[0062] In some embodiments, the special string identifier can be determined based on the position at which the special string appears in the metadata. For example, the decimal number corresponding to the last 8 digits of the special string identifier of the first appearing special string can be 1, the decimal number corresponding to the last 8 digits of the special string identifier of the second appearing special string can be 2, and so on. For example, the special string identifier of the eighth appearing special string can be 1111 1111 0000 1000, where the first 8 digits are a flag used to distinguish between a constant string and a special string, and the last 8 digits are the position at which the special string appears in the metadata (i.e., the eighth appearance).
[0063] In other embodiments, the special string identifier can also be determined based on the identifier of the constant string in the metadata and the position where the special string appears in the metadata. For example, if the maximum value of the identifier of the constant string in the first target file is 5, then the decimal number corresponding to the last 8 digits of the special string identifier of the first special string that appears can be 6, the decimal number corresponding to the last 8 digits of the special string identifier of the second special string that appears can be 7, and so on. For example, the special string identifier of the eighth special string that appears can be 1111 1111 0000 1000, where the first 8 digits are a flag used to distinguish between a constant string and a special string, and the last 8 digits are the position where the special string appears in the metadata (i.e., the eighth appearance).
[0064] In some embodiments, the constant string information may include each constant string and a constant string identifier of each constant string. const constant string, then the constant string information can include N const constant strings and N const A constant string identifier. N const A constant string with N const There is a one-to-one correspondence between the constant string identifiers.
[0065] In other embodiments, the constant string information may include each constant string, the order in which each constant string appears, and the constant string identifier of the first constant string that appears. The constant string identifier of each constant string may be determined based on the constant string identifier and the order in which the constant string appears. For example, assuming that N const constant string, then the constant string information can include N const A constant string, N const The constant string identifier of the first constant string that occurs among the constant strings, and N const The order in which the constant strings appear. N const The nth occurrence of a constant string among constant strings (where n is greater than or equal to 2 and less than or equal to N const The constant string identifier of the first constant string can satisfy the following relationship:
[0066] ID n =ID1+n×α, (Formula 1)
[0067] Where ID nis the constant string identifier of the nth constant string occurrence, ID1 is the constant string identifier of the first constant string occurrence, and α is a positive integer greater than or equal to 1. Thus, based on the identifier of the first constant string occurrence and the order in which the constant strings occur, the constant string identifier of the i-th constant string occurrence can be determined.
[0068] Similarly, in some embodiments, the special string information may include each special string and the special string identifier of each special string. In other embodiments, the special string information may include each special string, the order of occurrence of each special string and the special string identifier of the first special string that occurs.
[0069] Replace constant strings in N target files with constant string identifiers. For example, if the constant string identifier of the constant string Methodref is 0000 0000 0000 0001, then replace all occurrences of the constant string Methodref in each of the N target files with 0000 0000 0000 0001. Similarly, replace all special strings in N target files with special string identifiers. For example, if the special string identifier of the special string SNAZ14389 in the first target file is 1111 1111 0000 1000, then replace all occurrences of the special string SNAZ14389 in the first target file with 1111 1111 0000 1000. The target file that has completed the replacement of the constant string identifier and the special string identifier can be called a replaced target file.
[0070] Replacing constant and special strings in the target file with their respective identifiers can reduce the target file size. Furthermore, constant and special string identifiers contain multiple recurring numbers, such as a flag used to distinguish constant and special strings. These recurring numbers can result in a higher compression ratio.
[0071] For the convenience of description, the constant character string information, the N special character string information, and the N replaced target files may be collectively referred to as first compression information.
[0072] 107 , group the M non-target files to obtain second compressed information, where the second compressed information includes multiple file sets.
[0073] In some embodiments, a file set can be first determined based on the size of the non-target files. The file set includes all non-target files whose sizes are less than or equal to a file size threshold among the M non-target files. For ease of description, reference files whose file sizes are less than or equal to the non-target file size threshold are referred to as small files, and non-target files whose file sizes are greater than the file size threshold are referred to as large files. The file set including all small files can be referred to as file set 1. It is understood that if the file sizes of the M non-target files are all greater than the file size threshold, then the second compression information may not include file set 1.
[0074] The file size threshold may be a system default or a setting. For example, in some embodiments, the file size threshold may be less than or equal to 1024 bytes (byte, B). For example, the file size threshold may be 1024B, 1000B, 512B, 300B, 256B, 200B, 128B, or 100B. For another example, in some embodiments, the file size threshold may be less than or equal to 512B. For example, the file size threshold may be 512B, 500B, 300B, 256B, 200B, 128B, or 100B. For another example, in some embodiments, the file size threshold may be less than or equal to 256B. For example, the file size threshold may be 256B, 200B, 128B, or 100B.
[0075] It is understandable that the files included in the file set 1 are only selected based on file size. Therefore, the file set 1 may include files of various formats. For example, the file set 1 may include one or more of text files, library files, and image files.
[0076] Files in the same file set have the same characteristics. For example, the characteristics of all files in file set 1 are that they are less than or equal to the file size threshold.
[0077] In some embodiments, the plurality of file sets may further include a plurality of file sets 2. Each file set in the plurality of file sets 2 includes files that are non-target files whose file size is greater than a file size threshold. Files belonging to the same file set 5 have the same characteristics and can be compressed using the same compression algorithm.
[0078] As described above, non-target files may include files of different types. For example, non-target files may include text files, video files, audio files, executable files, etc.
[0079] In some embodiments, non-target files can be grouped according to file type, with files of the same type belonging to the same file set. Different file sets may include different types of non-target files. For example, file set 2-1 may include text files; file set 2-2 may include executable files; file set 2-3 may include audio files; and file set 2-4 may include image files.
[0080] Different types of files have different extensions, and files of the same type may also have different extensions. Therefore, in some embodiments, non-target files can be grouped according to their extensions. For example, file set 2-1 includes all files with a .dll extension; file set 2-2 includes all files with a .exe extension; and file set 2-3 includes all files with a .txt extension.
[0081] Furthermore, the following situations may occur: 1. Files with different extensions can achieve good compression using the same compression algorithm, even though these files may be of different file types; 2. Files of the same type with different extensions can achieve better compression using different compression algorithms; 3. Files with the same extension but different encoding methods can achieve better compression using different compression algorithms. Therefore, in some embodiments, grouping information can be pre-set. In this way, non-target files can be grouped directly based on this grouping information.
[0082] For example, the Lempel-Ziv-Markov chain algorithm 2 (LZMA2) has a good compression effect on files with the extension .exe and files with the extension .dll. Therefore, all files with the extension .exe and files with the extension .dll can belong to the same file set 2.
[0083] For example, Huffman coding has a good compression effect on text files encoded in Unicode. Therefore, all text files encoded in Unicode can be classified into the same file set. For text files encoded in ASCII (American Standard Code for Information Interchange), the Lempel-Ziv-Welch (LZW) algorithm has a good compression effect. Therefore, although they are both text files, due to the different encoding methods, text files encoded in Unicode and text files encoded in ASCII belong to two different file sets.
[0084] According to the above relationship, the grouping information shown in Table 1 can be obtained.
[0085] Table 1
[0086] As shown in Table 1, *.dll indicates all files with the extension dll; *.exe indicates all files with the extension exe; txt file (ASCII) indicates a text file encoded in the American Standard Code for Information Interchange (ASCII); and txt file (unicode) indicates a text file encoded in Unicode.
[0087] According to the grouping information shown in Table 1, three file sets 5 can be determined, wherein file set 2-1 includes files with extensions dll and exe; file set 2-2 includes text files encoded in ASCII; and file set 2-3 includes text files encoded in unicode.
[0088] 108. Determine multiple bit streams.
[0089] The information to be compressed in the software package includes the first information to be compressed and the second information to be compressed determined in the above steps. For the convenience of description, the concept of object to be compressed is introduced. The object to be compressed can be the constant string information, the special string information, the replaced target file and the file set. As mentioned above, the first information to be compressed includes a constant string information, N special string information and N replaced target files. If it is assumed that the second information to be compressed includes a file set 1 and N S file set 2. Then the software package contains 1+N+N+1+N S =N TB objects to be compressed. N TB One of the objects to be compressed is a constant string information, one of N special string information, one of N replaced target files, or a file set 1 or N S One of 2 file sets.
[0090] Each bit stream in the multiple bit streams includes part or all of the files in the same object to be compressed.
[0091] In some embodiments, the multiple bitstreams correspond one-to-one to the objects to be compressed determined in step 107 , and each bitstream includes all files in the corresponding object to be compressed.
[0092] In other embodiments, if the size of the object to be compressed exceeds a bitstream threshold, the object to be compressed may be divided into multiple bitstreams, each bitstream having a size not exceeding the bitstream threshold. In this case, a single object to be compressed may correspond to multiple bitstreams, each containing only a portion of the file in the corresponding object to be compressed.
[0093] 109 , determining the compression workload of the bit stream that needs to be compressed.
[0094] In some embodiments, the bitstream corresponding to the small file may not be compressed. The main reason is that the compression efficiency of small files is not good. For example, the compression ratio of small files is not high; or, although the compression ratio of small files is relatively high, the computing resources occupied are not worth the compression ratio of the small files. For example, a small file of 100B may only be 30B after compression, but the same computing resources can compress a file of 100 megabytes (MB) to 40MB. Therefore, the same computing resources can only save 70B of capacity when compressing a small file. Compared with a software package, the saved capacity is very small. Therefore, from the perspective of saving computing resources and improving compression efficiency, not compressing small files can complete the compression of the software package more quickly and will not have a substantial impact on the final size of the software package.
[0095] Of course, in other embodiments, small files may also be compressed.
[0096] For any bitstream, the compression workload of the bitstream can be determined based on the compressibility score of the bitstream and the size of the bitstream. The compressibility score of the bitstream is the distance between the amplitude histogram of the information included in the bitstream and the Gaussian white noise histogram. The bitstream is digitally mapped (each 8 bits or 1 byte is mapped to the interval 0-255) to obtain the gradient histogram of the bitstream, and then the gradient amplitude histogram is further calculated. The distance between the amplitude histogram and the Gaussian white noise histogram can be the Euclidean distance, standard Euclidean distance, or Mahalanobis distance between the amplitude histogram and the Gaussian white noise histogram.
[0097] The smaller the distance between the amplitude histogram and the Gaussian white noise histogram, the greater the randomness of the bit stream, the greater the compression amount, and the corresponding workload; conversely, the larger the distance between the amplitude histogram and the Gaussian white noise histogram, the smaller the randomness of the bit stream, the smaller the compression amount, and the corresponding workload.
[0098] In some embodiments, the compression workload of the bitstream can be determined according to the following formula: i =(1-Grd i )×Size i ,
[0099] Among them, Comp i is the compression workload of the i-th bit stream, Grd i is the compressibility score of the i-th bitstream, Size i is the size of the i-th bitstream.
[0100] In some embodiments, the compression workload of the i-th bitstream may reflect the proportion of the compression time of the i-th bitstream in the total compression time of the K bitstreams. For example, if the compression workload of the i-th bitstream is 10, it means that the compression time of the i-th bitstream accounts for 10% of the total compression time of the K bitstreams.
[0101] In other embodiments, the compression workload of the bitstream can be determined according to a predetermined correspondence relationship. For example, Table 2 shows the correspondence relationship between the compression workload, the bitstream size, and the compressibility score.
[0102] Table 2
[0103] As shown in Table 2, if the compressibility score of the bitstream is greater than or equal to S1 and less than S2 and the bitstream size of the bitstream is less than 1000KB, then the compression workload of the bitstream is 10; if the compressibility score of the bitstream is greater than or equal to S1 and less than S2 and the bitstream size of the bitstream is greater than 10MB, then the compression workload of the bitstream is 30.
[0104] 110, according to the compression workload of the bit stream, allocate an operation unit to each bit stream.
[0105] Assuming that the compression workload for K bitstreams is determined and there are P arithmetic units available for compression, the K bitstreams can be evenly distributed among the P arithmetic units so that the sum of the compression workloads for the bitstreams allocated to different arithmetic units is the same or approximately the same.
[0106] For example, if the workload for bitstream 1 is 10, the workload for bitstream 2 is 20, the workload for bitstream 3 is 30, and the workload for bitstream 4 is 40, and there are two arithmetic units, arithmetic unit 1 and arithmetic unit 2, that can perform compression work, then bitstream 1 and bitstream 4 can be assigned to arithmetic unit 1, and bitstream 2 and bitstream 3 can be assigned to arithmetic unit 2. In this way, the workload assigned to each of the two arithmetic units is 50.
[0107] In some embodiments, an arithmetic unit may be a processor or a component (e.g., a core) within a processor. For example, in some embodiments, a computer device may include multiple processors, each of which may be an arithmetic unit. For another example, in other embodiments, a computer device may include a processor comprising multiple cores. An arithmetic unit is a core within the processor.
[0108] 111. The operation unit compresses the allocated bit stream to obtain a compressed bit stream.
[0109] 112, assembling the compressed bit stream to obtain a compressed software package.
[0110] It is understandable that if the bit stream corresponding to the file set 1 is not compressed, the bit stream to be assembled in step 112 also includes the bit stream corresponding to the file set 1.
[0111] Figure 2 is a schematic flow chart of a method for compressing data provided by an embodiment of the present application. The method shown in Figure 2 can be executed by a computer device or a component in a computer device (such as a chip or system chip, etc.). For ease of description.
[0112] 201, determine the target file in the software package.
[0113] The target file can be identified by the file extension. For example, common target file extensions include .obj, .o, .class, etc.
[0114] Program files can be divided into target files and non-target files based on whether they are target files or not. For ease of description, assume that a software package contains N target files and M non-target files, where N and M are both positive integers greater than or equal to 1.
[0115] The non-target file may be any type of file, such as a text file, a video file, an audio file, an executable file, etc.
[0116] Optionally, in some embodiments, the software package is a compressed file or includes one or more compressed files. In this case, the compressed file can be first decompressed to obtain an uncompressed file, and then the target file in the uncompressed file can be determined. If the compressed files are nested, the nested compressed files can be decompressed after the compressed file is decompressed until no compressed files are left after decompression.
[0117] 202, determine the constant string information.
[0118] The constant string information is used to indicate at least one constant string and a constant string identifier corresponding to each constant string in the at least one constant string. Each target file in the N target files includes the at least one constant string.
[0119] 203. Determine N special character string information.
[0120] The N special string information corresponds one-to-one to the N target files. Assuming that the first string information is any one of the N special string information, the first target file is the target file corresponding to the first string information among the N target files. The first string information can be used to indicate at least one special string in the first target file and a string identifier corresponding to each of the at least one special string.
[0121] The method for determining the constant character string information and the special character string information may refer to the embodiment shown in FIG1 , and for the sake of brevity, it will not be described in detail here.
[0122] 204 , replacing the constant character string and the special character string of each target file in the N target files with the corresponding identifier to obtain N replaced target files.
[0123] The replacement method of the constant character string and the special character string can refer to the embodiment shown in FIG1 , and for the sake of brevity, it will not be described here in detail.
[0124] 205. Compress the software package according to the first compression information. The first compression information may include the constant character string, the N special character strings, and the N replaced target files.
[0125] In the above technical solution, constant and special strings in the target file are replaced with corresponding identifiers before the replaced target file is compressed. This can reduce the space occupied by special and constant strings in the target file, thereby improving the compression efficiency of the target file.
[0126] In some embodiments, two bitstreams may be determined: bitstream 1 and bitstream 2. Bitstream 1 includes the first compression information, and bitstream 2 includes M non-target files in the software package. Bitstream 1 and bitstream 2 are compressed separately to obtain compression results for bitstream 1 and bitstream 2. The compression results for bitstream 1 and bitstream 2 are then combined to obtain a compressed software package.
[0127] In other embodiments, non-target files can be first classified to obtain two file sets: the first file set includes non-target files whose file sizes are less than or equal to a file size threshold; the second file set includes non-target files whose file sizes are greater than the file size threshold. In this case, three bitstreams can be determined: bitstream 1, bitstream 2, and bitstream 3. Bitstream 1 includes the constant string information, the N special string information, and the N replaced target files. Bitstream 2 can include files from the first file set, and bitstream 3 can include files from the second file set. Bitstream 1 and bitstream 3 are compressed separately to obtain the compression results of bitstream 1 and bitstream 3. The compression results of bitstream 2, bitstream 1, and bitstream 3 are combined to obtain a compressed software package. In other words, in this embodiment, non-target files whose file sizes are less than or equal to the file size threshold are not compressed. Instead, only the decomposition and replacement results of the target file (i.e., the constant string, the N special string information, and the N replaced target files) and non-target files whose file sizes are greater than the file size threshold are compressed. This can conserve computing resources of the computer device and more quickly obtain a compressed software package.
[0128] In some embodiments, the sizes of all non-target files are less than or equal to the file size threshold. In this case, only the constant string, the N special strings, and the N replaced target files can be compressed, and then the compression result and the non-target files are combined to obtain a compressed software package.
[0129] In other embodiments, all non-target file sizes are greater than the file size threshold. In this case, the non-target files, the constant string, the N special strings, and the N replaced target files can be compressed, and then the compression results are combined to obtain a compressed software package.
[0130] In the above embodiment, the decomposition and replacement results of the target file (i.e., the constant string, the N special strings, and the N replaced target files) are all compressed in one bitstream. In other embodiments, the constant string, the N special strings, and the N replaced target files may belong to different bitstreams, respectively. For example, in some embodiments, four bitstreams may be determined, bitstream 1 to bitstream 4. Bitstream 1 includes the constant string. Bitstream 2 includes the N special strings. Bitstream 3 includes the N replaced target files. Bitstream 4 includes M non-target files. In this case, bitstreams 1 to 4 may be compressed separately to obtain the compression results of bitstreams 1 to 4. Then, the compression results of bitstreams 1 to 4 are combined to obtain a compressed software package. For another example, in some embodiments, five bitstreams may be determined, bitstream 1 to bitstream 5. Bitstream 1 includes the constant string. Bitstream 2 includes the N special strings. Bitstream 3 includes the N replaced target files. Bitstream 4 includes the first file set. Bitstream 5 includes the second file set. Bitstreams 1 to 3, and bitstream 5 are then compressed separately to obtain compression results for bitstream 1, bitstream 2, bitstream 3, and bitstream 5. Bitstream 4, the compression results for bitstream 1, bitstream 2, bitstream 3, and bitstream 5 are combined to obtain a compressed software package.
[0131] Similar to the embodiment shown in FIG1 , if multiple bitstreams require compression, the compression workload for each bitstream can be determined separately, and then a computing unit can be assigned to each bitstream based on the compression workload. The method for determining the compression workload and assigning computing units can be referenced to the embodiment shown in FIG1 , and for the sake of brevity, will not be further described here.
[0132] FIG3 is a schematic block diagram of a computer device according to an embodiment of the present application. The computer device shown in FIG3 includes a processing unit 301 and a compression unit 302 .
[0133] The processing unit 301 is configured to determine N target files included in the software package, where N is a positive integer greater than or equal to 1.
[0134] The processing unit 301 is further used to determine constant string information, where the constant string information is used to indicate at least one constant string and a constant string identifier corresponding to each constant string in the at least one constant string, and each target file in the N target files includes the at least one constant string.
[0135] The processing unit 301 is further used to determine N special string information, where the N special string information corresponds one-to-one to the N target files, the first special string information is used to indicate at least one special string in the first target file and a special string identifier corresponding to each special string in the at least one special string, the first special string is any one of the N special strings, and the first target file is the target file corresponding to the first special string.
[0136] The processing unit 301 is further configured to replace the constant character string and the special character string of each of the N target files with a corresponding identifier to obtain N replaced target files.
[0137] The compression unit 302 is configured to compress the software package according to first information to be compressed, where the first information to be compressed includes the constant character string information, the N special character string information, and the N replaced target files.
[0138] The specific functions and beneficial effects of the processing unit 301 and the compression unit 302 can be found in the description of the above embodiments, and for the sake of brevity, they will not be repeated here.
[0139] The processing unit 301 and the compression unit 302 may be implemented by a processor.
[0140] The present application also provides a computer device including a processor and a memory. The processor is coupled to the memory to read and execute instructions and / or program codes in the memory to perform the steps of the above method embodiment.
[0141] It should be understood that the processor may be a chip. For example, the processor may be a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), a graphics processing unit (GPU), a system on chip (SoC), a central processor unit (CPU), a network processor (NP), a digital signal processor (DSP), a microcontroller unit (MCU), a programmable logic device (PLD), other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or other integrated chips.
[0142] During implementation, each step of the above method can be completed by an integrated logic circuit of the hardware in the processor or by instructions in the form of software. The steps of the method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The software module can be located in a storage medium mature in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in conjunction with its hardware. To avoid repetition, it will not be described in detail here.
[0143] It should be noted that the processor in the embodiment of the present application can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method embodiment can be completed by an integrated logic circuit of the hardware in the processor or by instructions in the form of software. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the embodiment of the present application can be directly embodied as being executed by a hardware decoding processor, or can be executed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory, and the processor reads the information in the memory and completes the steps of the above method in combination with its hardware.
[0144] It is understood that the memory in the embodiments of the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus RAM (DR RAM). It should be noted that the memory of the systems and methods described herein is intended to include, but is not limited to, these and any other suitable types of memory.
[0145] According to the method provided in the embodiments of the present application, the present application also provides a computer program product, which includes: computer program code, which, when running on a computer, enables the computer to execute each step in the above embodiments.
[0146] According to the method provided in the embodiment of the present application, the present application also provides a computer-readable medium, which stores program code. When the program code is run on a computer, the computer executes the various steps in the above embodiments.
[0147] According to the method provided in an embodiment of the present application, an embodiment of the present application provides a chip system, which includes a logic circuit, which is used to couple with an input / output interface and transmit data through the input / output interface to execute the various steps in the above embodiments.
[0148] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0149] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0150] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0151] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0152] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0153] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.
[0154] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A method for compressing data, characterized in that: include: Determine N target files included in the software package, where N is a positive integer greater than or equal to 1; Determining constant string information, where the constant string information is used to indicate at least one constant string and a constant string identifier corresponding to each constant string in the at least one constant string, each target file in the N target files including the at least one constant string; Determining N special string information, the N special string information corresponding one-to-one to the N target files, wherein first special string information is used to indicate at least one special string in a first target file and a special string identifier corresponding to each special string in the at least one special string, the first special string information is any one of the N special string information, and the first target file is a target file corresponding to the first special string; Replace the constant character string and the special character string of each target file of the N target files with the corresponding identifier to obtain N replaced target files; The software package is compressed according to first information to be compressed, where the first information to be compressed includes the constant character string information, the N special character string information, and the N replaced target files.
2. The method according to claim 1, characterized in that The method further includes: determining M non-target files in the software package, where M is a positive integer greater than or equal to 1; Grouping the non-target files to obtain third information to be compressed, wherein the third information to be compressed includes at least one file set, wherein files belonging to the same file set have the same characteristics; The step of compressing the software package according to the first information to be compressed includes: The first information to be compressed and the second information to be compressed are compressed to obtain a compressed software package.
3. The method according to claim 2, characterized in that The multiple file sets include a first file set and at least one small file, where the small file is a non-target file of the M non-target files whose size is less than or equal to a file size threshold.
4. The method according to claim 2 or 3, characterized in that The multiple file sets include at least one second file set, wherein the multiple non-target files belonging to the same file set have the same extension, the same encoding method, and / or the same file type.
5. The method according to any one of claims 2 to 4, characterized in that Before compressing the first information to be compressed and the second information to be compressed to obtain a compressed software package, the method further includes: Determine K compression workloads, where the K compression workloads correspond one-to-one to K bitstreams, each of the K bitstreams includes part or all of files from the same object to be compressed, where the object to be compressed includes the constant string information, the special string information, the replaced target file, and the file set, and K is a positive integer greater than or equal to 2; According to the K compression workloads, the K bit streams are allocated to P operation units for compression, wherein the difference between the first workload and the second workload is less than a workload threshold, wherein the first workload is the sum of the workloads of the bit streams allocated to the first operation unit, and the second workload is the sum of the workloads of the bit streams allocated to the second operation unit, and the first operation unit and the second operation unit are any two operation units among the P operation units, and P is a positive integer greater than or equal to 2.
6. The method according to claim 5, characterized in that The determining of K compression workloads includes: Determining a compressibility score of an i-th bitstream among the K bitstreams, where i=1, ..., K; An i-th compression workload among the K compression workloads is determined according to the compressibility score of the i-th bitstream and the size of the i-th bitstream.
7. The method according to claim 6, characterized in that The compressibility score of the i-th bit stream is a distance between an amplitude histogram of information included in the i-th bit stream and an amplitude histogram of Gaussian white noise.
8. The method according to claim 6 or 7, characterized in that The determining, according to the compressibility score of the i-th bitstream and the size of the i-th bitstream, the i-th compression workload among the K compression workloads includes: The i-th compression workload is determined according to the following formula: Comp i =(1-Grd i )×Size i , Among them, Comp i is the i-th compression workload, Grd i is the compressibility score of the i-th bitstream, Size i It is the The size of the i bitstreams.
9. A computer device, characterized in that: include: a processing unit, configured to determine N target files included in the software package, where N is a positive integer greater than or equal to 1; The processing unit is further configured to determine constant string information, the constant string information being configured to indicate at least one constant string and a constant string identifier corresponding to each constant string in the at least one constant string, each target file in the N target files including the at least one constant string; The processing unit is further configured to determine N special string information, wherein the N special string information corresponds one-to-one to the N target files, wherein the first special string information is used to indicate at least one special string in the first target file and a special string identifier corresponding to each special string in the at least one special string, the first special string is any one of the N special strings, and the first target file is a target file corresponding to the first special string; The processing unit is further configured to replace the constant character string and the special character string of each of the N target files with a corresponding identifier to obtain N replaced target files; The compression unit is configured to compress the software package according to first information to be compressed, where the first information to be compressed includes the constant character string information, the N special character string information, and the N replaced target files.
10. The computer device according to claim 9, wherein: The processing unit is further configured to determine M non-target files in the software package, group the non-target files, and obtain third information to be compressed, wherein the third information to be compressed includes at least one file set, where M is a positive integer greater than or equal to 1, and files belonging to the same file set have the same characteristics; The compression unit is specifically configured to compress the first information to be compressed and the second information to be compressed to obtain a compressed software package.
11. The computer device according to claim 10, wherein: The multiple file sets include a first file set and at least one small file, where the small file is a non-target file of the M non-target files whose size is less than or equal to a file size threshold.
12. The computer device according to claim 10 or 11, characterized in that The multiple file sets include at least one second file set, wherein the multiple non-target files belonging to the same file set have the same extension, the same encoding method, and / or the same file type.
13. The computer device according to any one of claims 10 to 12, characterized in that The processing unit is further configured to determine K compression workloads, wherein the K compression workloads correspond one-to-one to K bit streams, each of the K bit streams includes part or all of files from the same object to be compressed, wherein the object to be compressed includes the constant string information, the special string information, the replaced target file, and the file set, and K is a positive integer greater than or equal to 2; According to the K compression workloads, the K bit streams are allocated to P operation units for compression, wherein the difference between the first workload and the second workload is less than a workload threshold, wherein the first workload is the sum of the workloads of the bit streams allocated to the first operation unit, and the second workload is the sum of the workloads of the bit streams allocated to the second operation unit, and the first operation unit and the second operation unit are any two operation units among the P operation units, and P is a positive integer greater than or equal to 2.
14. The computer device according to claim 13, wherein: The processing unit is specifically configured to determine a compressibility score of an i-th bitstream among the K bitstreams, where i=1, ..., K; An i-th compression workload among the K compression workloads is determined according to the compressibility score of the i-th bitstream and the size of the i-th bitstream.
15. The computer device according to claim 14, wherein: The compressibility score of the i-th bit stream is a distance between an amplitude histogram of information included in the i-th bit stream and an amplitude histogram of Gaussian white noise.
16. The computer device according to claim 14 or 15, characterized in that The processing unit is specifically configured to determine the i-th compression workload according to the following formula: i =(1-Grd i )×Size i , Among them, Comp i is the i-th compression workload, Grd i is the compressibility score of the i-th bitstream, Size i is the size of the i-th bitstream.
17. A computer device, characterized in that: include: A processor, wherein the processor is coupled to a memory, and is configured to read and execute instructions and / or program codes in the memory to perform the method according to any one of claims 1 to 8.
18. A chip system, characterized in that: include: A logic circuit, wherein the logic circuit is configured to be coupled to an input / output interface and transmit data through the input / output interface to execute the method according to any one of claims 1 to 8.
19. A computer-readable medium, characterized in that The computer-readable medium stores a program code, and when the computer program code is run on a computer, the computer is caused to perform the method according to any one of claims 1 to 8.