Data compression method and device, electronic equipment, storage medium and program product
By dividing the data into data blocks and selecting the appropriate compression algorithm based on feature similarity, the data blocks are compressed, and the problem of selecting the appropriate compression algorithm is solved, achieving a more efficient data compression effect.
Patent Information
- Application Number
- CN202510518010.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-23
- Publication Date
- 2025-08-08
AI Technical Summary
In the prior art, how to choose a suitable compression algorithm to improve the compression effect of different types of data has become an urgent problem.
The to be processed data is divided into multiple to be processed data blocks, and the target compression algorithm is determined by obtaining the characteristic similarity between the to be compressed feature data of each to be processed data block and the algorithm feature data preset by each compression algorithm, and compressing the data blocks based on the characteristic similarity.
Improves compression effect, improves data compression rate and reduces data compression or decompression time.
Smart Images

Figure CN120447832A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computers, and in particular to a data compression method, device, electronic device, storage medium, and program product. Background Art
[0002] Compressing data using different compression algorithms can save storage space and improve storage efficiency. Different data compression algorithms are used in different application scenarios, and choosing the right one depends on the data type, performance requirements, and available resources. For example, for compressing large data sets, algorithms that reduce the data size are often chosen to save storage space or transmission bandwidth. Similarly, for medical imaging or data that requires exact replication, lossless compression algorithms that maintain the original data quality are often used.
[0003] Different compression algorithms have different compression effects. How to select appropriate compression algorithms to improve compression effects for different types of data has become an urgent problem to be solved. Summary of the Invention
[0004] The present application provides a data compression method, device, electronic device, storage medium and program product for determining a compression algorithm to improve the compression effect.
[0005] In a first aspect, the present application provides a data compression method, comprising:
[0006] Dividing the data to be processed into multiple data blocks to be processed;
[0007] Obtaining feature similarity between the to-be-compressed feature data of each to-be-processed data block and the algorithm feature data preset by each compression algorithm, where different compression algorithms correspond to preset algorithm feature data;
[0008] The target compression algorithm for the corresponding data block to be processed is determined based on the feature similarity between each feature data to be compressed and the algorithm feature data, and data compression is performed on the corresponding data block to be processed based on the target compression algorithm for each data block to be processed.
[0009] In one possible implementation, determining a target compression algorithm for a data block to be processed based on feature similarity of each data block to be processed includes:
[0010] Determine the algorithm feature data whose feature similarity between each feature data to be compressed and the algorithm feature data is greater than a first preset threshold as the target algorithm feature data of the corresponding data block to be processed;
[0011] The compression algorithm corresponding to the target algorithm characteristic data of each data block to be processed is determined as the target compression algorithm of the corresponding data block to be processed.
[0012] In one implementable manner, the step of compressing the corresponding data blocks to be processed based on the target compression algorithm of each data block to be processed includes:
[0013] Determining a target data block to be processed having the same target compression algorithm among a plurality of data blocks to be processed;
[0014] Merging the target to-be-processed data blocks into merged to-be-processed data;
[0015] The merged data to be processed is compressed based on a target compression algorithm corresponding to the merged data to be processed.
[0016] In one implementable manner, merging the target to-be-processed data blocks into merged to-be-processed data includes:
[0017] If the number of the target data blocks to be processed is less than or equal to a second preset threshold, merging the target data blocks to be processed into one merged data block to be processed;
[0018] If the number of the target data blocks to be processed is greater than the second preset threshold, the target data blocks to be processed are merged to obtain multiple merged data to be processed, and the number of target data blocks to be processed included in each merged data to be processed in the multiple merged data to be processed is less than or equal to the second preset threshold.
[0019] In one possible implementation, before obtaining the feature similarity between the feature data to be compressed of each data block to be processed and the algorithm feature data preset by each compression algorithm, the method further includes:
[0020] Compressing the sample data using different compression algorithms to obtain sample compressed data of the sample data under different compression algorithms;
[0021] Determining target sample compressed data of the sample data from the sample compressed data under different compression algorithms based on compression performance parameters corresponding to the sample compressed data;
[0022] The characteristic data of the sample data is determined as the algorithm characteristic data of the compression algorithm corresponding to the target sample compression data.
[0023] In one implementable manner, the compression performance parameter includes at least one compression performance index; and determining the target sample compression data of the sample data from the sample compression data under different compression algorithms based on the compression performance parameter corresponding to each sample compression data includes:
[0024] assigning a first compression weight value to the compression performance indicator;
[0025] Determine a first index score value of each sample compressed data in the compression performance index, and perform a weighted sum calculation based on the first index score value of each sample compressed data and the corresponding first compression weight value to determine a first compression performance score value of the corresponding sample compressed data;
[0026] The sample compressed data whose first compression performance score value is greater than a third preset threshold is determined as the target sample compressed data of the sample data.
[0027] In one possible implementation, before obtaining the feature similarity between the feature data to be compressed of each data block to be processed and the algorithm feature data preset by each compression algorithm, the method further includes:
[0028] Extract at least part of the data to be processed to obtain feature value data;
[0029] Compressing the eigenvalue data using different compression algorithms to obtain eigenvalue compressed data of the eigenvalue data under different compression algorithms;
[0030] Determining target eigenvalue compressed data of the eigenvalue data from the eigenvalue compressed data under different compression algorithms based on compression performance parameters corresponding to the eigenvalue compressed data;
[0031] The characteristic data of the characteristic value data is determined as the algorithm characteristic data of the compression algorithm corresponding to the target characteristic value compression data.
[0032] In one implementable manner, the compression performance parameter includes at least one compression performance index; and determining, based on the compression performance parameter corresponding to each eigenvalue compressed data, target eigenvalue compressed data of the eigenvalue data from the eigenvalue compressed data under different compression algorithms includes:
[0033] assigning a second compression weight value to the compression performance indicator;
[0034] Determine a second index score value of each eigenvalue compressed data in the compression performance index, and perform a weighted sum calculation based on the second index score value of each eigenvalue compressed data and the corresponding second compression weight value to determine a second compression performance score value of the corresponding eigenvalue compressed data;
[0035] The eigenvalue compressed data whose second compression performance score value is greater than a third preset threshold is determined as the target eigenvalue compressed data of the eigenvalue compressed data.
[0036] In a second aspect, the present application provides a data compression device, comprising: a data splitting module, configured to divide data to be processed into a plurality of data blocks to be processed;
[0037] A feature calculation module is used to obtain feature similarity between the feature data to be compressed of each data block to be processed and the algorithm feature data preset by each compression algorithm. Different compression algorithms correspond to preset algorithm feature data.
[0038] The compression module is used to determine the target compression algorithm corresponding to the data block to be processed based on the feature similarity between each feature data to be compressed and the algorithm feature data, and compress the corresponding data block to be processed based on the target compression algorithm of each data block to be processed.
[0039] In a third aspect, the present application provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor;
[0040] The memory stores computer-executable instructions;
[0041] The processor executes the computer-executable instructions stored in the memory to implement the method according to the first aspect.
[0042] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions are used to implement the method of the first aspect when executed by a processor.
[0043] In a fifth aspect, the present application provides a computer program product, comprising a computer program, which implements the method described in the first aspect when executed by a processor.
[0044] The data compression method, device, electronic device, storage medium and program product provided by the present application, when compressing the data to be processed, first divide the data to be processed into small-volume data blocks to be processed, and then select the corresponding target compression algorithm for the data blocks to be processed. Since the data volume of the data blocks to be processed is small, the matched target compression algorithm has higher accuracy.
[0045] At the same time, feature similarity is calculated based on the feature data of the data block to be processed and the feature data of the compression algorithm, and a target compression algorithm that is more suitable for the data block to be processed is determined based on the feature similarity. The data block to be processed is compressed using the target compression algorithm, which can improve the compression effect. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.
[0047] Figure 1 is an implementation scenario diagram shown in an exemplary embodiment;
[0048] Figure 2is a flow chart of a data compression method shown in an exemplary embodiment;
[0049] Figure 3 is a flow chart of a data compression method shown in another exemplary embodiment;
[0050] Figure 4 is a flow chart of a data compression method shown in another exemplary embodiment;
[0051] Figure 5 is a flow chart of a data compression method shown in another exemplary embodiment;
[0052] Figure 6 is a flow chart of a data compression method shown in another exemplary embodiment;
[0053] Figure 7 is a structural diagram of a data compression device shown in an exemplary embodiment;
[0054] Figure 8 It is a block diagram of an electronic device shown in an exemplary embodiment.
[0055] The above drawings illustrate specific embodiments of the present application, which will be described in more detail below. These drawings and the textual description are not intended to limit the scope of the present application in any way, but rather to illustrate the concepts of the present application to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION
[0056] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all embodiments consistent with the present application. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0057] Data compression can significantly reduce the size of data, thereby saving storage resources. It can be applied to hard drive storage, cloud storage, and database storage to reduce storage costs. It can also improve data transmission efficiency, increase data security, and save resources.
[0058] Different data compression algorithms are used in different application scenarios. How to choose the appropriate compression algorithm depends on the data type, performance requirements and available resources. Different types of data, such as text, images, audio and video, require different compression algorithms to achieve better compression effects.
[0059] Common compression algorithms include:
[0060] Differential compression and merge compression are commonly used for large datasets to reduce the size of the data in order to save storage space or transmission bandwidth. Both involve identifying and processing redundant information in the data to reduce its redundancy. Differential compression focuses on changes or differences in the data, compressing the data by comparing the differences between two datasets. It is mainly used in scenarios such as historical records, version control, and incremental backups, where different versions of data need to be compared and processed. Merge compression identifies and eliminates redundant information in the data, merging duplicate data into a single instance to reduce data storage and transmission costs. It is commonly used for general-purpose compression of most data types, such as text, images, audio, and video.
[0061] The LZ77 compression algorithm is a dictionary-based compression algorithm suitable for continuously repeated data. It has a high compression ratio but high matching complexity. If it is used to compress small data, the cost-effectiveness will be very low.
[0062] The Huffman compression algorithm is a compression algorithm based on character frequency distribution code length. It has a high compression rate, but the encoding and decoding process has a large overhead and is only suitable for data with uneven character distribution.
[0063] The Deflate compression algorithm combines the LZ77 compression algorithm and the Huffman compression algorithm. It has high compression ratio and fast decompression process and is suitable for various data types, but the compression process is relatively time-consuming.
[0064] The LZW dynamic dictionary compression algorithm is simple to implement and has a high compression ratio. However, its compression efficiency is affected when the probability of string repetition is low. It is suitable for bytes and strings that appear continuously and repeatedly in data streams.
[0065] The Deflate compression algorithm, the LZ77 compression algorithm, and the Huffman compression algorithm are suitable for reducing redundant information in text files, such as repeated words or sentences.
[0066] Image data compression typically uses lossless and lossy compression algorithms. Lossless compression algorithms, such as LZW and PNG (a compression algorithm), are suitable for scenarios where the original image quality must be maintained, such as medical imaging or images that require accurate reproduction. Lossy compression algorithms, such as JPEG (a compression algorithm), are suitable for scenarios where image quality is less critical, such as web images or printouts. They reduce file size by discarding information that is not sensitive to the human visual system.
[0067] Audio and video data compression typically uses specialized encoding technologies, such as the MP3 compression algorithm (for audio) and the H.264 compression algorithm (for video). These technologies use complex algorithms and mathematical models to reduce data size while maintaining audio or video quality. These algorithms typically require specialized decoders to restore the original data.
[0068] Different compression algorithms have different compression effects on different data. Selecting appropriate compression algorithms for different types of data can effectively improve the compression rate of data compression and reduce the time spent on data compression or decompression. How to choose a compression algorithm for data compression has become an urgent problem to be solved.
[0069] The data compression method, device, electronic device, storage medium and program product provided in this application are intended to solve the above technical problems in the prior art.
[0070] The following specific embodiments describe in detail the technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.
[0071] The service deployment method provided in this application can be implemented on any electronic device with data processing capabilities, or it can be a service deployment system. It should be noted that the service deployment system can be deployed individually on an electronic device in any environment (for example, individually on an edge server in an edge environment), or it can be deployed entirely in a cloud environment, or it can be deployed in a distributed manner in different environments.
[0072] For example, the service deployment system can be logically divided into multiple parts, each with different functions. The various parts in the service deployment system can be deployed in any two or three of the electronic device (located on the user side, such as the client), the edge environment, and the cloud environment. The edge environment is an environment that includes a collection of edge electronic devices that are close to the electronic device. The edge electronic devices include: edge servers, edge stations with computing power, etc. The various parts of the container startup system deployed in different environments or devices work together to realize the functions of the data processing platform.
[0073] It should be understood that this application does not restrict the deployment of which parts of the service deployment system are deployed in what specific environment. In actual application, adaptive deployment can be carried out according to the computing power of the electronic device, the resource availability of the edge environment and the cloud environment, or the specific application requirements.
[0074] like Figure 11 is an implementation scenario diagram of an exemplary embodiment. The data compression system corresponding to the implementation scenario includes a client 10 and a server 11. The client 10 and the server 11 are connected via network communication.
[0075] The execution subject of the method of the embodiment of the present application is the server 11, and the client 10 can receive the data to be processed. The client 10 sends the data to be processed to the server 11 through the network. The server 11 compresses the data and can feed back the data compression results to the client 10 through the network. The client 10 has a graphic user interface, and the data compression results can be displayed on the graphic interface.
[0076] In some embodiments, after receiving the data to be processed, the server 11 divides the data to be processed into multiple data blocks to be processed; obtains the feature similarity between the feature data to be compressed of each data block to be processed and the algorithm feature data preset by each compression algorithm, and different compression algorithms correspond to preset algorithm feature data; determines the target compression algorithm corresponding to the data block to be processed based on the feature similarity between each feature data to be compressed and the algorithm feature data, and compresses the corresponding data block to be processed based on the target compression algorithm of each data block to be processed.
[0077] It is understandable that the data compression device can be set at Figure 1 In the server 11, but the example shown in this embodiment Figure 1 The implementation environment shown is only exemplary. In other embodiments, the data compression method may also be applied to other implementation environments, and the data compression device may also be set in other structures in other implementation environments. No specific limitations are made here.
[0078] In this embodiment, the client 10 is an electronic device on the user side, which can be a wired terminal with a visual structure or a wireless terminal. In other embodiments, the terminal can be an electronic device with a visual structure such as a mobile phone, computer, tablet, and vehicle-mounted equipment.
[0079] The server 11 can be an edge environment and a cloud environment, such as a physical server, a server cluster, and a cloud server, etc., and no specific restrictions are made here.
[0080] Figure 2 This is a flowchart of a data compression method shown in an exemplary embodiment, which is applied to Figure 1 The server 11 in Figure 2 As shown, the method includes steps S201 to S203, which are described in detail as follows:
[0081] S201: Divide the data to be processed into a plurality of data blocks to be processed.
[0082] In some embodiments, the data to be processed may be text, image, video, audio, or other data.
[0083] In some embodiments, the data to be processed is divided into small-volume data blocks to be processed.
[0084] For example, if the data to be processed is divided into data blocks to be processed of the same size, the size of a data block to be processed may be 8k (Kilobyte), 16k, 32k, etc., and no specific limitation is given here.
[0085] The data to be processed of the assigned tasks can also be divided into data blocks of different sizes, but the size of the data blocks to be processed is set to be smaller than a certain value, such as the size of the data blocks to be processed is smaller than 20k, 26k, 30k, etc., and no specific restrictions are made here.
[0086] In an embodiment of the present application, the data to be processed is divided into multiple data blocks to be processed, and the volume of the data blocks to be processed is relatively small. Therefore, the target compression algorithm obtained for the data blocks to be processed can be more suitable for the data blocks to be processed, and the corresponding target compression algorithm is more accurate.
[0087] S202: Acquire feature similarities between the to-be-compressed feature data of each to-be-processed data block and the preset algorithm feature data of each compression algorithm. Different compression algorithms correspond to preset algorithm feature data.
[0088] In some embodiments, for each data block to be processed, feature data to be compressed of the data block to be processed is obtained. The feature data to be compressed may include the hash value, time series features, statistical features, scalar features, etc. of the data block to be processed, and no specific restrictions are made here.
[0089] The feature data to be compressed may be data obtained by inputting the data block to be processed into a feature extraction model, and the feature extraction module may be a neural network model, which is not specifically limited here.
[0090] In some embodiments, the algorithm characteristic data corresponding to the compression algorithm is the characteristic data of the data suitable for the compression algorithm. For example, when a data is compressed, it is determined that compression algorithm A is suitable for compressing the data, then the characteristic data of the data is the algorithm characteristic data of the compression algorithm.
[0091] In some embodiments, the feature similarity between the feature data to be compressed and the algorithm feature data can indicate the compatibility between the corresponding data block to be processed and the corresponding compression algorithm, or indicate the compression effect of the compression algorithm corresponding to the algorithm feature data on the data block to be processed.
[0092] It can be understood that the greater the feature similarity between the feature data to be compressed of each data block to be processed and the algorithm feature data preset by each compression algorithm, the higher the adaptability between the data block to be processed and the corresponding compression algorithm, and the better the compression effect of the compression algorithm on the data block to be processed; the smaller the feature similarity between the feature data to be compressed of each data block to be processed and the algorithm feature data preset by each compression algorithm, the lower the adaptability between the data block to be processed and the corresponding compression algorithm, and the worse the compression effect of the compression algorithm on the data block to be processed.
[0093] S203 , determining a target compression algorithm for a corresponding data block to be processed based on the feature similarity between each feature data to be compressed and the algorithm feature data, and performing data compression on the corresponding data block to be processed based on the target compression algorithm for each data block to be processed.
[0094] In some embodiments, a target compression algorithm corresponding to the data block to be processed is determined based on the feature similarity between each feature data to be compressed and the algorithm feature data. The target compression algorithm can be regarded as a compression algorithm with good compression effect for the data block to be processed.
[0095] In some embodiments, the algorithm feature data whose feature similarity between each feature data to be compressed and the algorithm feature data is greater than a first preset threshold is determined as the target algorithm feature data of the corresponding data block to be processed; the compression algorithm corresponding to the target algorithm feature data of each data block to be processed is determined as the target compression algorithm of the corresponding data block to be processed.
[0096] In some embodiments, after the target compression algorithm of the data block to be processed is determined, the corresponding data block to be processed can be directly compressed using the target compression algorithm of each data block to be processed.
[0097] In some embodiments, data blocks to be processed with the same target compression algorithm may be merged and then compressed using the same target compression algorithm to improve compression efficiency.
[0098] In an embodiment of the present application, when compressing the data to be processed, the data to be processed is first divided into small-volume data blocks to be processed, and then a corresponding target compression algorithm is selected for the data blocks to be processed. Since the data volume of the data blocks to be processed is small, the matched target compression algorithm has higher accuracy.
[0099] At the same time, feature similarity is calculated through the feature data of the data block to be processed and the feature data of the compression algorithm, and a target compression algorithm that is more suitable for the data block to be processed is determined through the feature similarity. The data block to be processed is compressed using the target compression algorithm, which can improve the compression effect, such as increasing the compression rate of data compression, reducing the time consumed in data compression or decompression, etc.
[0100] Figure 3 is a flowchart of a data compression method shown in another exemplary embodiment, which is applied to Figure 1 The server 11 in Figure 3 As shown, the method includes steps S301 to S302. The method proposes a method for determining a target compression algorithm, which is described in detail as follows:
[0101] S301: Determine the algorithm feature data whose feature similarity between each feature data to be compressed and the algorithm feature data is greater than a first preset threshold as target algorithm feature data corresponding to the data block to be processed.
[0102] In some embodiments, the first preset threshold value may be set by an empirical parameter. In some embodiments, there is at least one algorithm feature data having a feature similarity greater than the first preset threshold value.
[0103] In some embodiments, there is only one algorithm feature data with feature similarity greater than a first preset threshold, that is, the algorithm feature data with the greatest feature similarity between each feature data to be compressed and the algorithm feature data is determined as the target algorithm feature data of the corresponding data block to be processed.
[0104] For example, in one embodiment, there is feature data A to be compressed of a data block to be processed, and the algorithm feature data includes algorithm feature data a of compression algorithm a, algorithm feature data b of compression algorithm b, and algorithm feature data c of compression algorithm c.
[0105] The feature similarity between the feature data to be compressed A and the algorithm feature data a, the feature similarity between the feature data to be compressed A and the algorithm feature data b, and the feature similarity between the feature data to be compressed A and the algorithm feature data c are calculated respectively.
[0106] Compare the feature similarity between the feature data A to be compressed and the algorithm feature data a, the feature similarity between the feature data A to be compressed and the algorithm feature data b, and the feature similarity between the feature data A to be compressed and the algorithm feature data c. If the feature similarity between the feature data A to be compressed and the algorithm feature data a is the largest, or the feature similarity between the feature data A to be compressed and the algorithm feature data a is greater than the first preset threshold, then the algorithm feature data a is the target algorithm feature data of the database to be processed corresponding to the feature data A to be compressed.
[0107] S302: Determine the compression algorithm corresponding to the target algorithm characteristic data of each data block to be processed as the target compression algorithm of the corresponding data block to be processed.
[0108] In some embodiments, the target algorithm characteristic data of each data block to be processed can be determined. Since the algorithm characteristic data corresponds one-to-one to the compression algorithm, the compression algorithm corresponding to the target algorithm characteristic data of each data block to be processed can be determined as the target compression algorithm of the corresponding data block to be processed through the target algorithm characteristic data of each data block to be processed.
[0109] In the embodiment of the present application, feature similarity is used to refer to the degree of adaptation between the data block to be processed and the compression algorithm. Thus, the target compression algorithm for each data block to be processed is determined by feature similarity, and the data block to be processed is compressed by the target compression algorithm to improve the compression effect.
[0110] Figure 4 is a flowchart of a data compression method shown in another exemplary embodiment, which is applied to Figure 1 The server 11 in Figure 4 As shown, the method includes steps S401 to S403. The method proposes a compression method, which is described in detail as follows:
[0111] S401: Determine a target data block to be processed having the same target compression algorithm among multiple data blocks to be processed.
[0112] In some embodiments, among multiple data blocks to be processed, there may be two or more data blocks to be processed with the same target compression algorithm. For a target compression algorithm, the data to be processed using the target compression algorithm is determined as the target data block to be processed of the target compression algorithm.
[0113] If there are data blocks to be processed 1, data blocks to be processed 2 and data blocks to be processed 3, the target compression algorithms of data blocks to be processed 1 and data blocks to be processed 2 are both compression algorithm a, and the target compression algorithm of data block to be processed 3 is compression algorithm b, then data blocks to be processed 1 and data blocks to be processed 2 are the target data blocks to be processed of compression algorithm a.
[0114] S402: Merge the target to-be-processed data blocks into merged to-be-processed data.
[0115] In some embodiments, since the target compression algorithms of the target data blocks to be processed are the same, the target data blocks to be processed can be directly merged into merged data to be processed, and the merged data to be processed can be compressed using the corresponding target compression algorithm.
[0116] In some embodiments, a method for obtaining merged data to be processed is also proposed. If the number of target data blocks to be processed is less than or equal to a second preset threshold, the target data blocks to be processed are merged into one merged data to be processed; if the number of target data blocks to be processed is greater than the second preset threshold, the target data blocks to be processed are merged to obtain multiple merged data to be processed, and the number of target data blocks to be processed included in each merged data to be processed in the multiple merged data to be processed is less than or equal to the second preset threshold.
[0117] In some embodiments, the second preset threshold can be set according to the performance of the target compression algorithm. For example, the second preset threshold can be a value such as 50, 80, 100, or 130, and no specific limitation is given here.
[0118] In some embodiments, for a target compression algorithm, if the number of target data blocks to be processed is less than or equal to a second preset threshold, they can be directly merged into a merged data to be processed, and the merged data to be processed can be compressed by the corresponding target compression algorithm, thereby reducing the number of compressions and improving compression efficiency.
[0119] In other embodiments, for a target compression algorithm, if the number of its target data blocks to be processed is greater than a second preset threshold, it means that the number of target data blocks to be processed is too large. At this time, the target data blocks to be processed are merged into one merged data to be processed. The volume of the merged data to be processed is too large, the compression time is long, and after the compression is completed, due to the large volume of the target data blocks to be processed before compression, subsequent reading of the target data blocks to be processed is difficult. Therefore, the target data blocks to be processed can be merged into multiple merged data to be processed, and the number of target data blocks to be processed included in each merged data to be processed in the multiple merged data to be processed is less than or equal to the second preset threshold, thereby improving the compression efficiency while facilitating the subsequent reading of data.
[0120] S403: compress the merged data to be processed based on a target compression algorithm corresponding to the merged data to be processed.
[0121] In some embodiments, after the merged data to be processed is obtained, since the target compression algorithms of the target data blocks to be processed in the merged data to be processed are the same, the merged data to be processed can be directly compressed using the target compression algorithm.
[0122] In an embodiment of the present application, for data blocks to be processed with the same target compression algorithm, the data blocks to be processed with the same target compression algorithm are merged and the merged data blocks to be processed are compressed, which can reduce the number of compression times and improve compression efficiency.
[0123] Figure 5 is a flowchart of a data compression method shown in another exemplary embodiment, which is applied to Figure 1The server 11 in Figure 5 As shown, the method includes steps S501 to S503. The method proposes a method for obtaining algorithm feature data, which is described in detail as follows:
[0124] S501 : Compress sample data using different compression algorithms to obtain sample compressed data obtained using different compression algorithms.
[0125] In some embodiments, the sample data may be text, image, video, audio, or other data.
[0126] In some embodiments, the sample data is directly compressed using different compression algorithms to obtain sample compressed data using different compression algorithms.
[0127] S502: Determine target sample compressed data of the sample data from the sample compressed data under different compression algorithms based on compression performance parameters corresponding to each sample compressed data.
[0128] In some embodiments, the compression performance parameter includes at least one compression performance indicator.
[0129] The compression performance indicator may include one or more of compression ratio, compression time, decompression time, etc., which are not specifically limited here.
[0130] In some embodiments, the compression performance parameter indicates that the sample compressed data has better performance, and the sample compressed data with better performance is the target sample compressed data of the sample data.
[0131] For example, in one embodiment, the compression performance indicator of the compression performance parameter is the compression ratio. A higher compression ratio indicates better compression performance. In this case, the sample compressed data with the highest compression ratio is determined as the target sample compressed data.
[0132] In some embodiments, a method for determining target sample compression data is also proposed, which assigns a first compression weight value to the compression performance index; determines the first indicator score value of each sample compression data in the compression performance index, and performs a weighted sum calculation based on the first indicator score value of each sample compression data and the corresponding first compression weight value to determine the first compression performance score value of the corresponding sample compression data; and determines the sample compression data whose first compression performance score value is greater than a third preset threshold as the target sample compression data of the sample data.
[0133] In some embodiments, a first compression weight value may be assigned to the compression performance indicator, the sum of the weight values of different first compression performance indicators is 1, and the number of compression performance indicators is one or more.
[0134] The first compression weight values of the different compression performance indicators can be set according to empirical parameters. For example, if the requirement for a certain compression performance indicator is relatively high, the first compression weight value of the compression performance indicator is also relatively high.
[0135] In some embodiments, a first indicator score value under different compression performance indicators can be determined based on the sample compressed data. The first indicator score value can be a value obtained by normalizing the sample compressed data under the compression performance indicator.
[0136] In this way, through the above method, the first indicator score value and the first compression weight value of each sample compressed data under different compression performance indicators can be obtained. By weighted summing the first indicator score value of the compression performance indicator and the corresponding first compression weight value, the first compression performance score value of each sample compressed data can be obtained.
[0137] The third preset threshold can be set by empirical parameters. It can be understood that the number of first compression performance score values greater than the third preset threshold is one, that is, the sample compression data with the largest first compression performance score value is determined as the target sample compression data of the sample data.
[0138] S503: Determine the characteristic data of the sample data as the algorithm characteristic data of the compression algorithm corresponding to the target sample compression data.
[0139] In some embodiments, after determining the target sample compressed data, the compression algorithm for obtaining the target sample compressed data may be determined, thereby determining the characteristic data of the sample data as the algorithm characteristic data of the compression algorithm corresponding to the target sample compressed data.
[0140] For example, in one embodiment, the sample data m is compressed respectively by compression algorithm a, compression algorithm b, and compression algorithm c to obtain sample compressed data a1, sample compressed data b1, and sample compressed data c1 respectively, and the first compression performance score values of the sample compressed data a1, the sample compressed data b1, and the sample compressed data c1 are calculated. The sample compressed data with the largest first compression performance score value among the sample compressed data a1, the sample compressed data b1, and the sample compressed data c1 is determined as the target sample compressed data, or the sample compressed data with the first compression performance score value greater than the third preset threshold among the sample compressed data a1, the sample compressed data b1, and the sample compressed data c1 is determined as the target sample compressed data.
[0141] If the sample compressed data b1 is the target sample compressed data, and the compression algorithm corresponding to the target sample compressed data b1 is compression algorithm b, then the feature data of the sample data m is the algorithm feature data of the compression algorithm b.
[0142] In an embodiment of the present application, sample data is compressed by a compression algorithm, and a compression algorithm with a better compression effect on the compressed sample data is determined based on the performance of the compressed sample data. The feature data of the sample data is used as the algorithm feature data of the compression algorithm with a better compression effect on the sample data. Subsequently, the compression effect of the corresponding compression algorithm on the data block to be processed can be indicated based on the feature similarity between the algorithm feature data and the feature data of the data block to be processed.
[0143] Figure 6 is a flowchart of a data compression method shown in another exemplary embodiment, which is applied to Figure 1 The server 11 in Figure 6 As shown, the method includes steps S601 to S604. The method proposes a compression method, which is described in detail as follows:
[0144] S601: Extract at least part of the data to be processed to obtain eigenvalue data.
[0145] In some embodiments, in order to make the algorithm feature data more consistent with the features of the data to be processed, at least part of the data to be processed is extracted to obtain feature value data, and the algorithm feature data is extracted using the feature value data.
[0146] In some embodiments, data of a size such as 20%, 30%, or 40% is extracted from the data to be processed to obtain feature value data.
[0147] S602 : compress the eigenvalue data using different compression algorithms to obtain eigenvalue compressed data of the eigenvalue data using different compression algorithms.
[0148] In the embodiment of the present application, the eigenvalue data is compressed using different compression algorithms to obtain eigenvalue compressed data of the eigenvalue data under different compression algorithms.
[0149] S603 : Determine target eigenvalue compressed data from the eigenvalue compressed data under different compression algorithms based on compression performance parameters corresponding to each eigenvalue compressed data.
[0150] In some embodiments, the compression performance parameter includes at least one compression performance indicator.
[0151] The compression performance indicator may include one or more of compression ratio, compression time, decompression time, etc., which are not specifically limited here.
[0152] In some embodiments, the compression performance parameter indicates that the sample compressed data has better performance, and the sample compressed data with better performance is the target sample compressed data of the sample data.
[0153] In some embodiments, another method for determining target sample compression data is proposed, which assigns a second compression weight value to the compression performance index; determines the second indicator score value of each eigenvalue compression data in the compression performance index, and performs a weighted summation calculation based on the second indicator score value of each eigenvalue compression data and the corresponding second compression weight value to determine the second compression performance score value of the corresponding eigenvalue compression data; determines the eigenvalue compression data whose second compression performance score value is greater than a third preset threshold as the target eigenvalue compression data of the eigenvalue compression data.
[0154] In some embodiments, a second compression weight value may be assigned to the compression performance indicator, the sum of the second weight values of different compression performance indicators is 1, and the number of compression performance indicators is one or more.
[0155] The first compression weight values of the same compression performance indicator may be the same or different, and there is no specific limitation here.
[0156] The second compression weight value of the different compression performance indicators can be set according to empirical parameters, and the setting method of the first compression weight value can be referred to, which will not be described here.
[0157] In some embodiments, a second indicator score value under different compression performance indicators can be determined based on the eigenvalue compressed data. The second indicator score value can be a normalized value of the eigenvalue compressed data under the compression performance indicator.
[0158] In this way, through the above method, the second indicator score value and the second compression weight value of each eigenvalue compressed data under different compression performance indicators can be obtained, and the second compression performance score value of each eigenvalue compressed data can be obtained by weighted summation of the second indicator score value of the compression performance indicator and the corresponding second compression weight value.
[0159] The eigenvalue compressed data whose second compression performance score is greater than the third preset threshold is determined as the target eigenvalue compressed data, that is, the eigenvalue compressed data with the largest second compression performance score is determined as the target eigenvalue compressed data of the sample data.
[0160] S604: Determine the characteristic data of the characteristic value data as the algorithm characteristic data of the compression algorithm corresponding to the target characteristic value compression data.
[0161] In some embodiments, after determining the target eigenvalue compressed data, a compression algorithm for obtaining the target eigenvalue compressed data may be determined, thereby determining the characteristic data of the eigenvalue data as the algorithm characteristic data of the compression algorithm corresponding to the target eigenvalue compressed data.
[0162] In an embodiment of the present application, algorithm feature data is generated by using part of the data to be processed. The generated algorithm feature data can better reflect the adaptability of different parts of the data to be processed to different compression algorithms. When the target compression algorithm is subsequently determined by the similarity between the algorithm feature data and the data block to be processed, a target compression algorithm with better compression effect for the data to be processed can be obtained more accurately.
[0163] Figure 7 FIG. 7 is a structural diagram of a data compression device according to an exemplary embodiment. The data compression device 700 is applied to a server of a data compression system and may include:
[0164] The data splitting module 710 is used to divide the data to be processed into multiple data blocks to be processed;
[0165] The feature calculation module 730 is used to obtain the feature similarity between the feature data to be compressed of each data block to be processed and the algorithm feature data preset by each compression algorithm. Different compression algorithms correspond to preset algorithm feature data.
[0166] The compression module 750 is used to determine the target compression algorithm for the corresponding data block to be processed based on the feature similarity between each feature data to be compressed and the algorithm feature data, and to compress the corresponding data block to be processed based on the target compression algorithm.
[0167] In one possible implementation, the compression module includes:
[0168] a target feature data determining unit, configured to determine the algorithm feature data whose feature similarity between each feature data to be compressed and the algorithm feature data is greater than a first preset threshold as target algorithm feature data for the corresponding data block to be processed;
[0169] The target compression algorithm determining unit is used to determine the compression algorithm corresponding to the target algorithm characteristic data of each data block to be processed as the target compression algorithm of the corresponding data block to be processed.
[0170] In one possible implementation, the compression module includes:
[0171] a target data to be processed determining unit, configured to determine a target data block to be processed having the same target compression algorithm among a plurality of data blocks to be processed;
[0172] A merging unit, configured to merge target to-be-processed data blocks into merged to-be-processed data;
[0173] The compression unit is used to compress the merged data to be processed based on a target compression algorithm corresponding to the merged data to be processed.
[0174] In one possible implementation, the merging unit includes:
[0175] A first merging section is used to merge the target data blocks to be processed into one merged data block to be processed if the number of the target data blocks to be processed is less than or equal to a second preset threshold;
[0176] The second merging section is used to merge the target data blocks to be processed if the number of target data blocks to be processed is greater than the second preset threshold value, so as to obtain multiple merged data to be processed, and the number of target data blocks to be processed included in each merged data to be processed in the multiple merged data to be processed is less than or equal to the second preset threshold value.
[0177] In one possible implementation, the data compression device further includes:
[0178] The sample data compression module is used to compress the sample data using different compression algorithms to obtain sample compressed data of the sample data under different compression algorithms;
[0179] A first performance calculation module is configured to determine target sample compressed data of the sample data from the sample compressed data under different compression algorithms based on the compression performance parameters corresponding to the sample compressed data;
[0180] The first algorithm characteristic data acquisition module is used to determine the characteristic data of the sample data as the algorithm characteristic data of the compression algorithm corresponding to the target sample compression data.
[0181] In one implementation, the compression performance parameter includes at least one compression performance index; and the first performance calculation module includes:
[0182] A first weight allocation unit, configured to allocate a first compression weight value to the compression performance indicator;
[0183] a first performance calculation unit, configured to determine a first index score value of the compression performance index for each sample compressed data, and to determine a first compression performance score value for the corresponding sample compressed data by performing a weighted sum calculation based on the first index score value of each sample compressed data and a corresponding first compression weight value;
[0184] The first performance processing unit is configured to determine the sample compressed data having a first compression performance score greater than a third preset threshold as the target sample compressed data of the sample data.
[0185] In one possible implementation, the data compression method further includes:
[0186] A data extraction module is used to extract at least part of the data to be processed to obtain characteristic value data;
[0187] The eigenvalue data compression module is used to compress the eigenvalue data using different compression algorithms to obtain eigenvalue compressed data of the eigenvalue data under different compression algorithms;
[0188] A second performance calculation module is used to determine target eigenvalue compressed data of the eigenvalue data from the eigenvalue compressed data under different compression algorithms based on the compression performance parameters corresponding to the eigenvalue compressed data;
[0189] The second algorithm characteristic data acquisition module is used to determine the characteristic data of the characteristic value data as the algorithm characteristic data of the compression algorithm corresponding to the target characteristic value compression data.
[0190] In one implementation, the compression performance parameter includes at least one compression performance index; and the second performance calculation module includes:
[0191] A second weight allocation unit, configured to allocate a second compression weight value to the compression performance indicator;
[0192] a second performance calculation unit, configured to determine a second index score value of the compression performance index for each eigenvalue compressed data, and to determine a second compression performance score value for the corresponding eigenvalue compressed data by performing a weighted sum calculation based on the second index score value of each eigenvalue compressed data and a corresponding second compression weight value;
[0193] The second performance processing unit is configured to determine the eigenvalue compressed data whose second compression performance score is greater than a third preset threshold as target eigenvalue compressed data of the eigenvalue compressed data.
[0194] The data compression device provided in this embodiment can be used to execute the above-mentioned data compression method. Its implementation principle and technical effects are similar and will not be described in detail in this embodiment.
[0195] Figure 8 This is a block diagram of an electronic device shown in an exemplary embodiment. Figure 8 The electronic device 800 may include: a processor 81 and a memory 82, wherein the processor 81 and the memory 82 can communicate; illustratively, the processor 81 and the memory 82 communicate via a communication bus 83, the memory 82 is used to store computer-executable instructions, and the processor 81 is used to call the computer-executable instructions in the memory to execute the data compression method shown in any of the above method embodiments.
[0196] The processor may be a central processing unit (CPU), or other general-purpose processor, a digital signal processor (DSP), or an application-specific integrated circuit (ASIC). The general-purpose processor may be a microprocessor or any conventional processor. The steps of the method disclosed in this application may be directly implemented as being executed by a hardware processor, or may be implemented by a combination of hardware and software modules in the processor.
[0197] The present application provides a computer-readable storage medium having computer-executable instructions stored thereon; when the computer-executable instructions are executed by a processor, they are used to implement the data compression method as described in any of the above embodiments.
[0198] An embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the above-mentioned data compression method is implemented.
[0199] It should be noted that for the aforementioned method embodiments, for the sake of simplicity, they are all expressed as a series of action combinations, but those skilled in the art should be aware that this application is not limited by the order of the actions described, because according to this application, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all optional embodiments, and the actions and modules involved are not necessarily required by this application.
[0200] It should be further noted that, although the various steps in the flowchart are shown in sequence as indicated by the arrows, these steps are not necessarily performed in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps may be performed in other orders. Moreover, at least a portion of the steps in the flowchart may include multiple sub-steps or multiple stages, and these sub-steps or stages are not necessarily performed at the same time, but may be performed at different times. The execution order of these sub-steps or stages is not necessarily to be performed in sequence, but may be performed in turn or alternately with other steps or at least a portion of the sub-steps or stages of other steps.
[0201] It should be understood that the above-described device embodiments are merely illustrative, and the device of the present application may also be implemented in other ways. For example, the division of units / modules in the above-described embodiments is merely a logical functional division, and actual implementations may employ other division methods. For example, multiple units, modules, or components may be combined or integrated into another system, or some features may be omitted or not implemented.
[0202] In addition, unless otherwise specified, the functional units / modules in the various embodiments of the present application may be integrated into a single unit / module, each unit / module may exist physically separately, or two or more units / modules may be integrated together. The aforementioned integrated units / modules may be implemented in the form of hardware or software program modules.
[0203] If the integrated unit / module is implemented in hardware, the hardware may be digital circuits, analog circuits, etc. The physical implementation of the hardware structure includes, but is not limited to, transistors, memristors, etc. Unless otherwise specified, the processor may be any appropriate hardware processor, such as a CPU, GPU, FPGA, DSP, and ASIC. Unless otherwise specified, the storage unit may be any appropriate magnetic storage medium or magneto-optical storage medium, such as resistive random access memory (RRAM), dynamic random access memory (DRAM), static random access memory (SRAM), enhanced dynamic random access memory (EDRAM), high-bandwidth memory (HBM), hybrid memory cube (HMC), etc.
[0204] If the integrated unit / module is implemented in the form of a software program module and sold or used as an independent product, it can be stored in a computer-readable memory. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a memory and includes a number of instructions for enabling a computer device (which can be a personal computer, server or network device, etc.) to execute all or part of the steps of the various embodiments of the present application. The aforementioned memory includes various media that can store program codes, such as a USB flash drive, a read-only memory (ROM), a random access memory (RAM), a mobile hard disk, a magnetic disk or an optical disk.
[0205] In the above embodiments, the description of each embodiment has its own emphasis. For parts not described in detail in a particular embodiment, please refer to the relevant description of other embodiments. The technical features of the above embodiments can be combined in any way. To keep the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0206] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, and the true scope and spirit of the present application are indicated by the following claims.
[0207] It should be understood that the present application is not limited to the exact structure described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.
Claims
1. A data compression method, characterized in that: include: Dividing the data to be processed into multiple data blocks to be processed; Obtaining feature similarity between the to-be-compressed feature data of each to-be-processed data block and the algorithm feature data preset by each compression algorithm, where different compression algorithms correspond to preset algorithm feature data; The target compression algorithm for the corresponding data block to be processed is determined based on the feature similarity between each feature data to be compressed and the algorithm feature data, and data compression is performed on the corresponding data block to be processed based on the target compression algorithm for each data block to be processed.
2. The method according to claim 1, characterized in that The step of determining a target compression algorithm for a corresponding data block to be processed based on the feature similarity of each data block to be processed includes: Determine the algorithm feature data whose feature similarity between each feature data to be compressed and the algorithm feature data is greater than a first preset threshold as the target algorithm feature data of the corresponding data block to be processed; The compression algorithm corresponding to the target algorithm characteristic data of each data block to be processed is determined as the target compression algorithm of the corresponding data block to be processed.
3. The method according to claim 1, characterized in that The method of compressing the corresponding data blocks to be processed based on the target compression algorithm of each data block to be processed includes: Determining a target data block to be processed having the same target compression algorithm among a plurality of data blocks to be processed; Merging the target to-be-processed data blocks into merged to-be-processed data; The merged data to be processed is compressed based on a target compression algorithm corresponding to the merged data to be processed.
4. The method according to claim 3, characterized in that The step of merging the target to-be-processed data blocks into merged to-be-processed data includes: If the number of the target data blocks to be processed is less than or equal to a second preset threshold, merging the target data blocks to be processed into one merged data block to be processed; If the number of the target data blocks to be processed is greater than the second preset threshold, the target data blocks to be processed are merged to obtain multiple merged data to be processed, and the number of target data blocks to be processed included in each merged data to be processed in the multiple merged data to be processed is less than or equal to the second preset threshold.
5. The method according to any one of claims 1 to 4, characterized in that Before obtaining the feature similarity between the feature data to be compressed of each data block to be processed and the algorithm feature data preset by each compression algorithm, the method further includes: Compressing the sample data using different compression algorithms to obtain sample compressed data of the sample data under different compression algorithms; Determining target sample compressed data of the sample data from the sample compressed data under different compression algorithms based on compression performance parameters corresponding to the sample compressed data; The characteristic data of the sample data is determined as the algorithm characteristic data of the compression algorithm corresponding to the target sample compression data.
6. The method according to claim 5, characterized in that The compression performance parameter includes at least one compression performance index; and determining the target sample compression data of the sample data from the sample compression data under different compression algorithms based on the compression performance parameter corresponding to each sample compression data includes: assigning a first compression weight value to the compression performance indicator; Determine a first index score value of each sample compressed data in the compression performance index, and perform a weighted sum calculation based on the first index score value of each sample compressed data and the corresponding first compression weight value to determine a first compression performance score value of the corresponding sample compressed data; The sample compressed data whose first compression performance score value is greater than a third preset threshold is determined as the target sample compressed data of the sample data.
7. The method according to any one of claims 1 to 4, characterized in that Before obtaining the feature similarity between the feature data to be compressed of each data block to be processed and the algorithm feature data preset by each compression algorithm, the method further includes: Extract at least part of the data to be processed to obtain feature value data; Compressing the eigenvalue data using different compression algorithms to obtain eigenvalue compressed data of the eigenvalue data under different compression algorithms; Determining target eigenvalue compressed data of the eigenvalue data from the eigenvalue compressed data under different compression algorithms based on compression performance parameters corresponding to the eigenvalue compressed data; The characteristic data of the characteristic value data is determined as the algorithm characteristic data of the compression algorithm corresponding to the target characteristic value compression data.
8. The method according to claim 7, characterized in that The compression performance parameter includes at least one compression performance index; and determining target eigenvalue compressed data of the eigenvalue data from the eigenvalue compressed data under different compression algorithms based on the compression performance parameter corresponding to each eigenvalue compressed data includes: assigning a second compression weight value to the compression performance indicator; Determine a second index score value of each eigenvalue compressed data in the compression performance index, and perform a weighted sum calculation based on the second index score value of each eigenvalue compressed data and the corresponding second compression weight value to determine a second compression performance score value of the corresponding eigenvalue compressed data; The eigenvalue compressed data whose second compression performance score value is greater than a third preset threshold is determined as the target eigenvalue compressed data of the eigenvalue compressed data.
9. A data compression device, characterized in that: include: A data splitting module is used to divide the data to be processed into multiple data blocks to be processed; A feature calculation module is used to obtain feature similarity between the feature data to be compressed of each data block to be processed and the algorithm feature data preset by each compression algorithm. Different compression algorithms correspond to preset algorithm feature data. The compression module is used to determine the target compression algorithm corresponding to the data block to be processed based on the feature similarity between each feature data to be compressed and the algorithm feature data, and compress the corresponding data block to be processed based on the target compression algorithm of each data block to be processed.
10. An electronic device, characterized in that: include: a processor, and a memory communicatively connected to the processor; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory to implement the method according to any one of claims 1 to 8.
11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer-executable instructions, which are used to implement the method according to any one of claims 1 to 8 when executed by a processor.
12. A computer program product, characterized in that The invention comprises a computer program, which implements the method according to any one of claims 1 to 8 when the computer program is executed by a processor.
Citation Information
Patent Citations
Data compression method and device, electronic equipment and computer readable medium
CN112994701A
Data compression method and device
CN115145884A
Data compression method and related device
CN116932493A