Data compression method and device, computing equipment and computer readable storage medium

By training the data sequence sample data, a coding mode suitable for the Simple encoding scheme is generated, which solves the problem that a single encoding scheme in the prior art is difficult to ensure the compression rate, and achieves efficient data compression.

CN119921780AActive Publication Date: 2025-05-02HUAWEI TECH CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202311429721.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-10-30
Publication Date
2025-05-02
Estimated Expiration
2043-10-30

AI Technical Summary

Technical Problem

When using a single encoding scheme, existing Simple algorithms are difficult to ensure the compression rate of data and cannot adapt to the actual size distribution of data sequences.

Method used

By obtaining sample data of the data sequence, the encoding mode of the Simple encoding scheme is trained to obtain the encoding mode suitable for the data sequence, and thus compressed encoding.

Benefits of technology

It ensures the compression rate of the data sequence, is suitable for the actual size distribution of the data sequence, and improves the efficiency of compression coding.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119921780A_ABST
    Figure CN119921780A_ABST
Patent Text Reader

Abstract

The invention discloses a data compression method and device, computing equipment and a computer readable storage medium, and belongs to the technical field of compression coding. According to the method, the sample data is obtained based on the to-be-compressed data sequence, so that the sample data can reflect the actual size distribution characteristics of the to-be-compressed data sequence, and then the coding mode of the Simple coding scheme is trained based on the sample data, so that the coding unit in the trained coding mode can meet the actual size distribution of the data sequence; the trained coding mode can be suitable for the data sequence, so that the data sequence is compressed and coded by using the trained coding mode, and the compression ratio of the data sequence can be ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of compression coding technology, and in particular to a data compression method, apparatus, computing device, and computer-readable storage medium. Background Art

[0002] Simple algorithms are a family of algorithms widely used to compress information such as time series data and integer data. Time series data includes data in multiple time series databases such as InfluxDB and TimescaleDB, and integer data includes inverted index tables of major search engines. Simple algorithms compress and encode data to be compressed through the encoding mode in the encoding scheme. The encoding scheme of Simple algorithms is also called Simple encoding scheme.

[0003] There are many existing Simple algorithms, such as Simple9, Simple16, Simple8b and other Simple algorithms. Different Simple algorithms have different Simple encoding schemes. Correspondingly, different Simple encoding schemes have different encoding modes. Once the Simple algorithm to be used is determined, the Simple encoding scheme to be used is also determined. If the encoding modes in the Simple encoding scheme do not cover the actual size distribution of the data to be compressed, continuing to compress and encode the data to be compressed according to the encoding mode in the Simple encoding scheme will make it difficult to guarantee the compression rate of the data to be compressed. Summary of the invention

[0004] The embodiments of the present application provide a data compression method, apparatus, computing device, and computer-readable storage medium, which can solve the problem that the compression rate is difficult to guarantee when using a single Simple encoding scheme. The technical solution is as follows:

[0005] In a first aspect, a data compression method is provided, the method comprising the following steps: for any data sequence, first obtaining multiple sample data based on the data sequence; then training a first encoding mode based on the multiple sample data to obtain a second encoding mode, and then compressing and encoding the data sequence based on the second encoding mode, wherein the first encoding mode is used to indicate a division method of encoding units in a payload encoded based on a Simple coding scheme.

[0006] The method obtains sample data based on the data sequence so that the sample data can reflect the actual size distribution characteristics of the data sequence to be compressed, and then trains the encoding mode of the Simple encoding scheme based on the sample data, so that the encoding units under the trained encoding mode can meet the actual size distribution of the data sequence, and the trained encoding mode can be applicable to the data sequence, so that the trained encoding mode is used to compress and encode the data sequence, which can ensure the compression rate of the data sequence.

[0007] In a possible implementation, the above-mentioned process of training the first coding mode based on multiple sample data to obtain the second coding mode includes: first training the first coding mode based on the compression bit width of the multiple sample data and the bit width of the payload to obtain multiple candidate coding modes, and then determining the second coding mode from the multiple candidate coding modes, wherein each candidate coding mode corresponds to at least one sample data, and the compression bit width is the bit width of the compressed data of the sample data.

[0008] Based on the above possible implementation manner, the determined second encoding mode can be applicable to the data sequence, so that the compression rate of the data sequence can be further ensured when the determined second encoding mode is subsequently used on the data sequence.

[0009] In a possible implementation, the process of training the first coding mode based on the compressed bit widths of multiple sample data and the bit width of the payload to obtain multiple candidate coding modes includes: based on the bit width of the payload, dividing the compressed bit widths of the multiple sample data into multiple bit width groups, each bit width group includes the compressed bit width of at least one sample data, the total bit width of the bit width group is less than or equal to the bit width of the payload, and the total bit width is the sum of the compressed bit widths in the bit width group; for any bit width group, based on the total bit width of the bit width group and the bit width of the payload, determining the candidate coding mode corresponding to the bit width group, wherein each coding unit in the candidate coding mode corresponds to a compressed bit width in the bit width group, the bit width of each coding unit in the candidate coding mode is less than or equal to the corresponding compressed bit width, and the total bit width of each coding unit in the candidate coding mode is equal to the bit width of the payload.

[0010] Based on the above possible implementation methods, based on the bit width of the payload, the compressed bit width of multiple sample data is divided into multiple bit width groups, and then based on the total bit width of each bit width group and the bit width of the payload, each candidate encoding mode is determined, and the second encoding mode is determined from the candidate encoding modes. There is no need to use dynamic programming to determine the second encoding mode, which reduces the amount of calculation. Therefore, even under low computing power settings, large-scale sequence data can be efficiently compressed and encoded.

[0011] In a possible implementation, the process of determining the second coding mode from the multiple candidate coding modes includes: determining at least one coding mode with the highest occurrence frequency among the multiple candidate coding modes as the second coding mode.

[0012] Based on the above possible implementation methods, when the second encoding mode is subsequently used to compress and encode the data sequence, it can be more suitable for the data size distribution of the data sequence and can further improve the compression rate of the data sequence.

[0013] In a possible implementation, the method also includes the following steps: if the bit width of the coding unit under each second coding mode is smaller than the maximum compression bit width of the data sequence, determine a third coding mode based on the maximum compression bit width; replace the second coding mode with the lowest frequency with the third coding mode, wherein the maximum compression bit width is the compression bit width of the largest data in the data sequence, and the bit width of each coding unit under the third coding mode is greater than or equal to the maximum compression bit width.

[0014] Based on the above possible implementation manner, the compressed data of the maximum data in the data sequence can be stored by the encoding unit in the third encoding mode to avoid the situation where the maximum data cannot be compressed and encoded.

[0015] In a second aspect, a data compression device is provided, comprising a functional module for executing the data compression method provided in the first aspect or any optional manner of the first aspect.

[0016] According to a third aspect, a computing device is provided, the computing device comprising a processor, wherein the processor is configured to execute program code so that the computing device performs the data compression method provided in the first aspect or any optional method of the first aspect.

[0017] According to a fourth aspect, a computer-readable storage medium is provided, wherein at least one program code is stored in the storage medium, and the program code is read by a processor to enable a computing device to execute a data compression method as provided in the first aspect or any optional method of the first aspect.

[0018] In a fifth aspect, a computer program product or a computer program is provided, which includes a program code, and the program code is stored in a computer-readable storage medium. A processor of a computing device reads the program code from the computer-readable storage medium, and the processor executes the program code, so that the computer device executes the method provided in the above-mentioned first aspect or various optional implementations of the first aspect.

[0019] Based on the implementations provided in the above aspects, this application can also be further combined to provide more implementations. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] Figure 1 It is a schematic diagram of an implementation environment of an application data compression method provided in an embodiment of the present application;

[0021] Figure 2 is a schematic diagram of an inverted index provided in an embodiment of the present application;

[0022] Figure 3 It is a storage space distribution diagram of a word provided by an embodiment of the present application;

[0023] Figure 4 is a flow chart of a data compression method provided in an embodiment of the present application;

[0024] Figure 5 It is a structural diagram of a character provided in an embodiment of the present application;

[0025] Figure 6 is a structural schematic diagram of a data compression device provided in an embodiment of the present application;

[0026] Figure 7 It is a structural diagram of a computing device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0027] In order to facilitate the understanding of this application, some of the terms involved in this application are introduced as follows:

[0028] Binary area: is an area used to store binary data.

[0029] Word: It is the basic output unit obtained by encoding with Simple algorithms. One or more data to be compressed can be compressed and encoded into a word. A word is a binary area. Each word consists of a selector and a payload. The bit width of the word is the sum of the bit width of the selector and the bit width of the payload in the word. For the encoding scheme and encoding mode of Simple algorithms, the bit width of the selector, the bit width of the payload, and the bit width of the word are all set values. For example, the bit width of the selector of Simple8b is 4 bits, and the bit width of the payload is 60 bits. The 4-bit selector and the 60-bit payload form a 64-bit word.

[0030] Payload: is a binary area used to store compressed data of consecutive elements in an input sequence. The input sequence is a data sequence to be compressed, which includes multiple data to be compressed. Each data to be compressed in the data sequence is an element.

[0031] Unit: A payload is divided into one or more units, each unit is used to store the compressed data of an element in the input sequence, and in this application, the unit is also called a coding unit.

[0032] Pattern: A division method of dividing a payload into coding units is called a pattern. In this application, the pattern is also called a coding mode. The coding units divided according to a coding mode are called coding units under the coding mode. The number of coding units under different coding modes is different, and / or the bit width of coding units under different coding modes is different. The division method indicated by each coding mode is represented by the division information of the coding mode. For example, the division information includes the number of coding units under the coding mode and the bit width of each coding unit under the coding mode. According to the division information, the payload can be divided into a specified number of coding units and coding units of a specified bit width. In another possible implementation, the division information includes the unit bit width of each coding unit under the coding mode, but does not include the number of coding units. The number of coding units is represented by the number of coding unit bit widths in the division information. In another possible implementation, if the bit widths of each coding unit under the coding mode are the same, the division information includes the number of coding units and the bit width of one coding unit, but does not include the bit widths of all coding units under the coding mode. In another possible implementation, if the sum of the bit widths of each coding unit in the encoding mode is less than the bit width of the payload, the division information also includes a residual bit width, which is the bit width of the binary area in the payload except the coding unit. The residual bit width is not used to store compressed data, and the residual bit width is also the wasted payload bit width in the payload.

[0033] Selector: It is a binary area used to record a coding mode. For example, different coding modes are identified by different mode identifiers. The selector is used to store a mode identifier of a coding mode. The mode identifier consists of at least one bit of binary data, such as 0 or 1.

[0034] Coding scheme: used to indicate the correspondence between mode identifiers and coding modes. A coding scheme is represented by scheme information, which includes division information of at least one coding mode and a mode identifier of each coding mode in the at least one coding mode. The division information of each coding mode corresponds to its own mode identifier. In this application, this coding scheme is also called Simple coding scheme.

[0035] Next, combine the attached Figure 1 , an exemplary description is given of the implementation environment involved in the embodiments of the present application.

[0036] Figure 1It is a schematic diagram of an implementation environment of an application data compression method provided in an embodiment of the present application. The implementation environment includes an encoder 101 and a decoder 102. The encoder 101 is a component in the implementation environment for executing the data compression method provided in the present application, and the decoder 102 is a component in the implementation environment for decoding the compression result of the encoder 101.

[0037] For example, each time a data sequence is input into the encoder 101, the encoder 101 executes the data compression method of the present application on the data sequence to compress and encode the data sequence into at least one word, and sends the at least one word to the decoder 102, which decodes the at least one word to restore the data sequence.

[0038] Any one of the encoder 101 and the decoder 102 may be implemented by software or hardware. For example, the encoder 101 is taken as an example to introduce the implementation of the encoder 101. The implementation of the decoder 102 may refer to the implementation of the encoder 101.

[0039] Among them, the encoder 101 is taken as an example of a software functional unit, and the encoder 101 may include code running on a computing instance. Among them, the computing instance may include at least one of a physical host (computing device), a virtual machine, and a container. Further, the above-mentioned computing instance may be one or more. For example, the encoder 101 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed in the same region (region) or in different regions. Furthermore, the multiple hosts / virtual machines / containers used to run the code may be distributed in the same availability zone (AZ) or in different AZs, each AZ including a data center or multiple data centers with close geographical locations. Among them, usually a region may include multiple AZs.

[0040] Similarly, multiple hosts / virtual machines / containers used to run the code can be distributed in the same virtual private cloud (VPC) or in multiple VPCs. Usually, a VPC is set up in a region. For cross-region communication between two VPCs in the same region and between VPCs in different regions, a communication gateway needs to be set up in each VPC to achieve interconnection between VPCs through the communication gateway.

[0041] The encoder 101 is taken as an example of a hardware functional unit, and the encoder 101 may include at least one computing device, such as a server, etc. Alternatively, the encoder 101 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0042] The multiple computing devices included in the encoder 101 can be distributed in the same region or in different regions. The multiple computing devices included in the encoder 101 can be distributed in the same AZ or in different AZs. Similarly, the multiple computing devices included in the encoder 101 can be distributed in the same VPC or in multiple VPCs. The multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.

[0043] The encoder 101 and the decoder 102 may be located in the same computing device, for example, the encoder 101 and the decoder 102 are both located in a server. Alternatively, the encoder 101 and the decoder 102 are located in different computing devices, for example, the encoder 101 is located in one server and the decoder 102 is located in another server.

[0044] Before the encoder (such as the encoder 101 described above) executes the data compression method provided by the present application, the following two configurations may be performed for the encoder:

[0045] 1. Configure the bit width of the payload and the bit width of the selector.

[0046] For ease of description, the bit width of the load is recorded as P, and the bit width of the selector is recorded as S, where both P and S are integers greater than 0. Since a load and a selector form a word, the sum of the configured P and S is equal to the bit width of the word.

[0047] Exemplarily, the user can use the word length of the computing device where the encoder is located as the bit width of the word to be encoded by the encoder, wherein, if the encoder is software, the computing device where the encoder is located is the device running the encoder, if the encoder is a hardware module in a certain device, the computing device where the encoder is located is the certain device, and if the encoder is an independent computing device, the computing device where the encoder is located is the encoder itself. The word length of the computing device refers to the length of the operation word inside the central processing unit (CPU) of the computing device. Taking a computing device with a 64-bit word length as an example, the user can use 64 bits as the bit width of the word to be encoded by the encoder. After determining the bit width of the word to be encoded, the user configures P and S with the bit width of the word as a constraint condition so that the sum of P and S is equal to the bit width of the word.

[0048] In some embodiments, the user can also obtain the maximum element width of the data to be compressed from the data source for the data to be compressed, and configure P with the maximum element width as a constraint, so that the configured P is greater than or equal to the maximum element width, and the sum of P and S is equal to the bit width of the word. Among them, the data type of the data to be compressed is an integer, there are multiple data to be compressed, and the data source is a device or software for providing the data to be compressed. The data source can be a time series database, a search engine with an inverted index, or other devices, device clusters, or software for providing data to be compressed. Here, the embodiments of the present application do not limit the data source. Taking the time series database providing the data to be compressed as an example, the time series data in the time series database is an integer, and the time series data can be used as the data to be compressed. Taking the search engine providing the data to be compressed as an example, the inverted index is an important data structure in the search engine. Figure 2 The schematic diagram of the inverted index shown, the document set provided by the search engine includes multiple documents, and the document identification (identity, ID) of each document is an integer, such as ID is an integer such as 1, 2, 3, 4. For each word appearing in each document, the document ID sequence of each word is composed of the ID of each document where each word is located, and the document ID sequence of all words forms an inverted index, and the ID in each document ID sequence in the inverted index can be used as data to be compressed. In another possible implementation, the data source may also provide integer data of other application types other than time series data or inverted index as data to be compressed, for example, taking the database of the traffic log of the data communication device as an example, a column of data in the traffic log is the number of terminal connections, and the number of terminal connections is an integer, then this column of terminal connection number can be used as data to be compressed. Here, the application embodiment of the present application does not limit the application type of the compressed data.

[0049] The maximum element bit width is the bit width of the compressed data of the maximum data to be compressed provided by the data source. The user sets P to be greater than or equal to the maximum element bit width to avoid the compressed data of the maximum data to be compressed being unable to use the P-bit payload for storage.

[0050] In some embodiments, in view of the characteristics of modern computer hardware, the user can also configure both P and S to be powers of 2 so that the computer device where the encoder is located or the device receiving the word processes the word. Exemplarily, the configured P and S are both powers of 2, and P is greater than or equal to the maximum element bit width, and the sum of P and S is equal to the bit width of the word.

[0051] In some embodiments, if the data size distribution of the data to be compressed is more diverse, more types of coding modes are required, and S can be configured to be larger. If the data size distribution of the data to be compressed is less diverse, the types of coding modes required are relatively fewer, and S can be configured to be relatively smaller to meet the needs of different data size distributions of the data to be compressed. For example, if the data size distribution of the data to be compressed is diverse, S is configured to 4 or 8. If the data size distribution of the data to be compressed is single, S is configured to 2.

[0052] 2. Configure the storage space for the payload and the storage space for the selector.

[0053] The storage space of the payload is used to store the payload of each word encoded by the encoder, and the storage space of the selector is used to store the selector of each word encoded by the encoder.

[0054] The user can configure the storage space of the payload and the storage space of the selector to be the same storage space or different storage spaces based on the data size distribution of the data to be compressed. For example, if the data size distribution of the data to be compressed is diverse, the storage space of the payload and the storage space of the selector can be configured as the same storage space; if the data size distribution of the data to be compressed is single, the storage space of the payload and the storage space of the selector can be configured as different storage spaces. Figure 3 As shown in the storage space distribution diagram of the characters, if the storage space of the payload and the storage space of the selector are configured as the same storage space, the subsequent encoder will store the payload and the selector of each character in the same storage space in sequence when storing the encoded characters, so as to achieve encoding the payload and the selector in the same storage space. If the storage space of the payload and the storage space of the selector are configured as different storage spaces, the subsequent encoder will store the payload of each encoded character in the storage space of the payload, and store the selector of each character in the storage space of the selector, so as to achieve encoding the payload and the selector in different storage spaces.

[0055] In some embodiments, since the storage space is generally a power of 2, if the bit width of the data stored in the storage space is also a power of 2, it is convenient to align the data with the storage address of the storage space, which can improve the utilization of the storage space. Based on this, if the storage space of the load and the storage space of the selector are configured as the same storage space, the user also configures P and S with the sum of P and S being a power of 2 as a constraint, so that the bit width of the word encoded by the subsequent encoder is a power of 2, so that the encoded word can be aligned with the storage address of the storage space to improve the utilization of the storage space. For example, when the maximum element bit width of the data to be compressed is 62 and the data size distribution of the data to be compressed is diverse, P can be configured to 64 and S can be configured to 4, and the storage space of the load and the storage space of the selector can be configured as different storage spaces. When the maximum element bit width of the data to be compressed is 32 and the data size distribution of the data to be compressed is single, P can be configured to 30 and S can be configured to 2, and the storage space of the load and the storage space of the selector can be configured as the same storage space.

[0056] If the storage space of the payload and the storage space of the selector are configured as different storage spaces, the storage space of the payload and the storage space of the selector may be provided by different storage devices or by the same storage device, and the storage device and the encoder may be located in the same computing device or in different computing devices.

[0057] After the encoder is configured, the encoder can compress and encode the compressed data by executing the data compression method provided by the present application. Next, Figure 4 The data compression method provided by this application is introduced in detail.

[0058] Figure 4 This is a flowchart of a data compression method provided in an embodiment of the present application. The method is executed by an encoder and includes the following steps.

[0059] 401. The encoder obtains a plurality of sample data based on the data sequence.

[0060] The data sequence is any data sequence to be compressed provided by the data source, and the data sequence includes multiple data to be compressed, and the data type of each data to be compressed is an integer. The data sequence can be composed of all the data to be compressed from the data source, or it can be composed of part of the data to be compressed from the data source. For example, if the data source has 10,000 data to be compressed, these 10,000 data to be compressed can form a data sequence, or these 10,000 data to be compressed can also be 10 data sequences, and the number of data to be compressed in each data sequence can be the same or different, so that Figure 2The inverted index shown is an example of all the data to be compressed of the data source, and the document ID sequence of each word can be a data sequence. Here, the embodiment of the present application does not limit the number of data to be compressed in the data sequence.

[0061] When any data sequence to be compressed of the data source is input into the encoder, the encoder samples the data to be compressed in the data sequence to obtain n sample data, so that the data size distribution of the n sample data is similar to the data size distribution of the data sequence. At this time, each sample data is a data to be compressed in the data sequence, and n is greater than 1 and less than the number of data to be compressed in the data sequence. For example, the encoder can use the first n data to be compressed in the data sequence as sample data, or use the last n data to be compressed in the data sequence as sample data, or the encoder collects a sample data from the data sequence every m data to be compressed, m is greater than or equal to 1, and less than the number of data to be compressed in the data sequence. Here, the embodiment of the present application does not limit the sampling method of the compressed data.

[0062] After obtaining n sample data, the encoder groups the n sample data into a sample sequence, and uses the sample sequence as a training sample to train the coding mode of the Simple coding scheme (such as step 402 below). Taking the first n to-be-compressed data in the data sequence as sample data as an example, assuming that the data sequence is: [4, 2, 4, 130, 12, 6, 8, 191, 0, 2, 1, 127, 3, 1, 5, 198, 0, 2, ...], assuming that n = 9, the encoder groups the first 9 to-be-compressed data in the data sequence into a sample sequence [4, 2, 4, 130, 12, 6, 8, 191, 0].

[0063] In another possible implementation, different data sources in the encoder are configured with different sample sequences, and the data size distribution of the sample data in each sample sequence is similar to the data size distribution of the data to be compressed in the corresponding data source. For any data sequence to be compressed, the encoder uses the sample sequence of the data source to which the data sequence belongs as the sample sequence of the data sequence, without obtaining the sample sequence from the data sequence in a sampling manner.

[0064] 402. The encoder trains a first coding mode based on the multiple sample data to obtain a trained second coding mode, where the first coding mode is used to indicate a division method of coding units in a payload encoded based on a Simple coding scheme.

[0065] The first encoding mode is the encoding mode in the Simple encoding scheme, and the second encoding mode is the trained encoding mode suitable for the data sequence, that is, the second encoding mode is obtained by training the encoding mode in the Simple encoding scheme. There is at least one second encoding mode, and there can be at most 2 second encoding modes. S kind.

[0066] In a possible implementation, during the training process, the encoder trains the encoding mode based on the compression bit width of each sample data in the sample sequence, the bit width P of the configured payload, and the bit width S of the selector to obtain the second encoding mode. The training process is introduced as follows with the following steps 4021 to 4022.

[0067] Step 4021: The encoder trains the first coding mode based on the compression bit width of the multiple sample data and the bit width of the payload to obtain multiple candidate coding modes, each candidate coding mode corresponding to at least one sample data.

[0068] Among them, the candidate coding modes are candidate modes of the second coding mode, each candidate coding mode is obtained through training based on corresponding sample data, and the bit width of the payload is P configured above.

[0069] The compression bit width of any sample data is the bit width of the compressed data of the sample data. In this application, the compressed data of any data (such as sample data or the maximum data to be compressed provided by the data source) refers to the binary data of the sample data after removing the leading zeros from its binary representation. For example, if a sample data is the number 5 stored in Int8, the binary representation of the number 5 is 00000101, after removing the leading zeros, the remaining 101 is the compressed data of the number 5. If a sample data is the number 0 stored in Int8, for example, if a sample data is the number 0 stored in Int8, the binary representation of the number 0 is 00000000, and the remaining 1-bit 0 is the compressed data of the number 0.

[0070] In a possible implementation, the encoder uses the bit width of the payload as a constraint condition and customizes the encoding mode for the continuous sample data based on the compressed bit width of each sample data in the sample sequence, so that each encoding unit in the customized encoding mode corresponds to a sample data in the continuous sample data, and each encoding unit can store compressed data of the corresponding sample data, wherein the customized encoding mode is the candidate encoding mode, and the continuous sample data refers to a plurality of adjacent sample data in the sample sequence. Still taking the sample sequence [4, 2, 4, 130, 12, 6, 8, 191, 0] as an example, 4, 2, 4, 130 in the sample sequence are continuous sample data. Exemplarily, the process is described in detail through the following steps A1 to A2.

[0071] Step A1: The encoder divides the compressed bit widths of the multiple sample data into multiple bit width groups based on the bit width of the payload, each bit width group includes the compressed bit width of at least one sample data, the total bit width of the bit width group is less than or equal to the bit width of the payload, and the total bit width is the sum of the compressed bit widths in the bit width group.

[0072] The bit width of the payload is the bit width P of the payload configured above.

[0073] The encoder first creates an empty bit width group, starts scanning from the first sample data in the sample sequence, and scans each sample data in the sample sequence in sequence. Each time a sample data is scanned, the scanned sample data is used as the current sample data, the encoder obtains the compressed bit width of the current sample data, adds the compressed bit width of the current sample data to the bit width group, and calculates the total bit width of the bit width group based on each compressed bit width in the bit width group. If the total bit width is less than P and the current sample data is not the last sample data in the sample sequence, continue to scan the next sample data of the current sample data in the sample sequence, and so on, until the total bit width of the bit width group is greater than or equal to P, or the last sample data is scanned.

[0074] In the case where the total bit width of the bit width group is greater than or equal to P, if the total bit width of the bit width group is equal to P, and the current sample data is not the last sample data in the sample sequence, the encoder creates another empty bit width group as the next bit width group of the bit width group, continues to scan the next sample data of the current sample data in the sample sequence, and adds the compressed bit width of the newly scanned sample data to the next bit width group in the manner of adding the compressed bit width to the bit width group. If the total bit width of the bit width group is equal to P, and the current sample data is the last sample data in the sample sequence, the sample sequence scanning ends.

[0075] In the case where the total bit width of the bit width group is greater than or equal to P, if the total bit width of the bit width group is greater than P, the encoder creates another empty bit width group as the next bit width group of the bit width group, and transfers the compressed bit width of the current sample data in the bit width group (i.e., the last compressed bit width in the bit width group) to the next bit width group, so that the total bit width of the bit width group is less than P. If the current sample data is not the last sample data in the sample sequence, the encoder continues to scan the next sample data of the current sample data in the sample sequence, and adds the compressed bit width of the newly scanned sample data to the next bit width group in the manner of adding the compressed bit width in the bit width group. This is repeated in this way until the last sample data in the sample sequence is scanned and the last sample data is added to the bit width group. In this way, the encoder can divide the compressed bit widths of the multiple sample data into multiple bit width groups, so that each bit width group is composed of the compressed bit widths of the continuous compressed data in the sample sequence.

[0076] Next, taking the sample sequence [4, 2, 4, 130, 12, 6, 8, 191, 0] as an example, the method of adding compressed bit width to the bit width group is introduced as follows.

[0077] Assume that P = 32. The encoder scans each sample data in the sample sequence in turn. After scanning the sample data "4", the compressed bit width "3" of the sample data "4" is added to an empty bit width group [] to obtain the bit width group [3]. The total bit width of the bit width group [3] is 3, which is less than 32. Continue to scan the next sample data "2" after the sample data "4"; add the compressed bit width "2" of the sample data "2" to the bit width group [3] to obtain the bit width group [3, 2]. The total bit width of the bit width group [3, 2] is 5, which is less than 32. Continue to scan the next sample data "4" after the sample data "2"; add the compressed bit width "3" of the sample data "4" to the bit width group [3, 2] to obtain the bit width group [3, 2, 3]. The bit width group [3, 2,3] has a total bit width of 8, which is less than 32. Continue to scan the next sample data "130" after the sample data "4"; add the compressed bit width "8" of the sample data "130" to the bit width group [3,2,3] to obtain the bit width group [3,2,3,8]. The total bit width of the bit width group [3,2,3,8] is 16, which is less than 32. Continue to scan the next sample data "12" after the sample data "130"; add the compressed bit width "4" of the sample data "12" to the bit width group [3,2,3,8] to obtain the bit width group [3,2,3,8,4]. The total bit width of the bit width group [3,2,3,8,4] is 20, which is less than 32. Continue to scan the next sample data "6" after the sample data "12"; add the sample data The compressed bit width "3" of "6" is added to the bit width group [3, 2, 3, 8, 4] to obtain the bit width group [3, 2, 3, 8, 4, 3]. The total bit width of the bit width group [3, 2, 3, 8, 4, 3] is 23, which is less than 32. Continue to scan the next sample data "8" after the sample data "6"; add the compressed bit width "4" of the sample data "8" to the bit width group [3, 2, 3, 8, 4, 3] to obtain the bit width group [3, 2, 3, 8, 4, 3, 4]. The total bit width of the bit width group [3, 2, 3, 8, 4, 3, 4] is 27, which is less than 32. Continue to scan the next sample data "191" after the sample data "8"; add the compressed bit width "8" of the sample data "191" to the bit width group [3, 2, 3, 8 , 4,3,4], and get the bit width group [3,2,3,8,4,3,4,8]. The total bit width of the bit width group [3,2,3,8,4,3,4,8] is 35, which is greater than 32. The last compressed bit width "8" in the bit width group [3,2,3,8,4,3,4,8] is transferred to the next empty bit width group [] to get the bit width group [3,2,3,8,4,3,4] and the next bit width group [8]. Continue to scan the next sample data "0" after the sample data "191"; add the compressed bit width "1" of the sample data "0" to the bit width group [8] to get the bit width group [8,1]. The sample sequence scan is completed, and finally the bit width groups [3,2,3,8,4,3,4] and the bit width group [8,1] are obtained.

[0078] In the above, the compressed bit width is first added to the bit width group, and then the total bit width of the bit width group is calculated. Then, according to whether the total bit width of the bit width group is less than P, it is decided whether to continue to add bit width to the bit width group. In another possible implementation, the encoder may also calculate the sum of the total bit width of the current bit width group and the compressed bit width of the current sample data after obtaining the compressed bit width of the current sample data. If the sum is less than or equal to P, the compressed bit width is added to the current bit width group. If the sum is greater than P, the encoder creates an empty bit width group, takes the newly created bit width group as the current bit width group, and continues to add compressed bit width to the newly created bit width group.

[0079] The above is introduced by taking the scanning of sample data of a sample sequence, and obtaining the compressed bit width of the sample data and adding the compressed bit width of the sample data to the bit width group as an example each time a sample data is scanned. In another possible implementation, the encoder first obtains the compressed bit width of each sample data in the sample sequence, and compresses the compressed bit widths of each sample data into a compressed bit width sequence according to the order of the sample data in the sample sequence, wherein the order of the compressed bit widths in the compressed bit width sequence is the same as the order of the sample data in the sample sequence, for example, a certain sample data is the first sample data in the sample sequence, and the compressed bit width of the sample data is the first compressed bit width in the compressed bit width sequence. After obtaining the compressed bit width sequence, the encoder sequentially scans each compressed bit width in the compressed bit width sequence, and each time a compressed bit width is scanned, the scanned compressed bit width is added to the bit width group. The method of adding to the bit width group can refer to the method of adding the compressed bit width to the bit width group when scanning the sample sequence above, and will not be repeated here.

[0080] For at least one determined bit width group, the at least one bit width group corresponds to a first coding mode respectively. For any bit width group, the encoder changes the compressed bit width in the bit width group to the bit width of each coding unit under the corresponding first coding mode. The sum of each compressed bit width in the bit width group may be the same as or different from the bit width of the payload. Based on the bit width of the payload, the first coding mode is trained by adjusting the compressed bit width in the bit width group. The training process is such as the following step A2.

[0081] Step A2: For any bit width group, the encoder determines a candidate coding mode corresponding to the bit width group based on the total bit width of the bit width group and the bit width of the payload, wherein each coding unit under the candidate coding mode corresponds to a compressed bit width in the bit width group, and the bit width of each coding unit under the candidate coding mode is less than or equal to the corresponding compressed bit width, and the total bit width of each coding unit under the candidate coding mode is equal to the bit width of the payload.

[0082] If the total bit width of any bit width group is equal to the bit width of the payload, the encoder determines the bit width group as the bit width group corresponding to the candidate coding mode, and uses each compressed bit width in the bit width group as the bit width of each coding unit under the candidate coding, so as to determine a candidate coding mode.

[0083] If the total bit width of any bit width group is smaller than the bit width of the payload, the encoder adjusts the value of the compressed bit width in the bit width group so that the total bit width of the bit width group is equal to the bit width of the payload. After adjusting the total bit width of the bit width group to the bit width of the payload, the bit width group after the total bit width adjustment is determined as the bit width group corresponding to the candidate coding mode, and each compressed bit width in the bit width group after the total bit width adjustment is used as the bit width of each coding unit under the candidate coding, so as to determine a candidate coding mode.

[0084] Among them, the adjustment method of the value of the compressed bit width in the bit width group is, for example, that the encoder adds 1 to a minimum compressed bit width in the bit width group, so that the total bit width of the bit width group increases by 1. If the total bit width is still smaller than the bit width of the payload after adding 1, for the bit width group after adding 1, the encoder again executes the step of adding 1 to a minimum compressed bit width in the bit width group until the total bit width of the bit width group is equal to the bit width of the payload.

[0085] Assuming that the bit width P of the payload is still 32, taking the above bit width group [3, 2, 3, 8, 4, 3, 4] as an example, the total bit width 27 of the bit width group [3, 2, 3, 8, 4, 3, 4] is less than 32. After adding 1 to the minimum compressed bit width "2", the bit width group [3, 3, 3, 8, 4, 3, 4] is obtained; the total bit width 28 of the bit width group [3, 3, 3, 8, 4, 3, 4] is less than 32. Since there are 4 minimum compressed bit widths "3" in the bit width group [3, 3, 3, 8, 4, 3, 4], any minimum compressed bit width "3" is selected and added by 1. For example, the first minimum compressed bit width "3" is selected and added by 1 to obtain the bit width group [4, 3, 3, 8, 4, 3, 4]; the total bit width 29 of the bit width group [4, 3, 3, 8, 4, 3, 4] If the bit width is less than 32, select the first minimum compression bit width "3" and add 1 to get the bit width group [4, 4, 3, 8, 4, 3, 4]; the total bit width 30 of the bit width group [4, 4, 3, 8, 4, 3, 4] is less than 32, select the first minimum compression bit width "3" and add 1 to get the bit width group [4, 4, 4, 8, 4, 3, 4]; the total bit width 31 of the bit width group [4, 4, 4, 8, 4, 3, 4] is less than 32, select the first minimum compression bit width "3" and add 1 to get the bit width group [4, 4, 4, 8, 4, 4, 4], the total bit width of the bit width group [4, 4, 4, 8, 4, 4, 4] is equal to 32, end adjusting the compression bit width, and take the bit width group [4, 4, 4, 8, 4, 4, 4] as a candidate coding mode bit width group. Since each compressed bit width in the bit width group will eventually be used as the bit width of a coding unit in a candidate coding mode, the bit widths of subsequent coding units in the candidate coding mode can be balanced by continuously adding 1 to the minimum compressed bit width in the bit width group.

[0086] In another possible implementation, if the total bit width of any bit width group is smaller than the bit width of the payload, the encoder obtains the difference between the total bit width and the bit width of the payload, increases the minimum compressed bit width in the bit width group by the difference, so that the total bit width of the bit width group is equal to the bit width of the payload, and then determines the bit width group as a bit width group of a candidate coding mode, and uses each compressed bit width in the bit width group as the bit width of each coding unit under the candidate coding, so as to determine a candidate coding mode.

[0087] Based on the above step A2, since each coding unit under the candidate coding mode corresponds to a compressed bit width in the bit width group respectively, the bit width of each coding unit under the candidate coding mode is less than or equal to the corresponding compressed bit width, so that the coding unit under the candidate coding mode can store the compressed data belonging to the corresponding compressed bit width, and the total bit width of each coding unit under the candidate coding mode is equal to the bit width of the payload, so that the bit width of the payload under the candidate coding mode can be allocated to each coding unit. If the candidate coding mode is subsequently used to encode the data to be compressed in the data sequence, the residual bit width in the payload can be avoided, thereby improving the utilization rate of the payload.

[0088] In another possible implementation, if the total bit width of any bit width group is smaller than the bit width of the payload, the encoder does not adjust the value of the compressed bit width in the bit width group, but instead uses the difference between the bit width of the payload and the total bit width of the bit width group as the remaining bit width of the candidate coding mode, and the encoder determines the candidate coding mode by using each compressed bit width in the bit width group as the bit width of a coding unit under the candidate coding mode.

[0089] For each bit width group in the multiple bit width groups, the encoder can determine a candidate coding mode by executing step A2, so that the encoder can finally determine multiple candidate coding modes, each candidate coding mode corresponds to a bit width group. The multiple candidate coding modes may be the same candidate coding mode or different candidate coding modes. If the multiple candidate coding modes are the same candidate coding mode, the candidate coding mode is used as the second coding mode. If there are multiple coding modes in the multiple candidate coding modes, the encoder determines the second coding mode by executing the following step 4022.

[0090] Step 4022: The encoder determines a second coding mode from the multiple candidate coding modes.

[0091] The encoder determines the second coding mode based on the bit width S of the selector and / or the frequency of occurrence of each candidate coding mode in the multiple candidate coding modes, and the number of types of the second coding mode is less than or equal to 2 S .

[0092] For example, the encoder determines 2 from the multiple candidate coding modes based on the bit width S of the selector. S For example, the encoder randomly selects 2 second encoding modes from the multiple candidate encoding modes. S A candidate coding mode is used as the second coding mode.

[0093] Alternatively, the encoder determines at least one coding mode with the highest frequency among the multiple candidate coding modes as the second coding mode based on the frequency of occurrence of each candidate coding mode among the multiple candidate coding modes. For example, the K coding modes with the highest frequency among the multiple candidate coding modes are determined as the second coding mode, where K is less than or equal to 2. S .

[0094] Alternatively, the encoder determines the second coding mode based on the bit width S of the selector and the frequency of occurrence of each candidate coding mode in the multiple candidate coding modes. For example, when there are multiple coding modes in the multiple candidate coding modes, if the number of types of coding modes in the multiple candidate coding modes is less than or equal to 2 S , the encoder determines the multiple coding modes as the second coding mode. If the number of coding modes in the multiple candidate coding modes is greater than 2 S , the encoder counts the frequency of occurrence of each coding mode in the multiple candidate coding modes, and selects the 2 most frequently occurring coding modes in the multiple candidate coding modes. S The encoding mode is determined as the second encoding mode.

[0095] By selecting less than or equal to 2 from the candidate coding modes S The second coding mode is selected to avoid the selector corresponding to the second coding mode being unable to store the mode identifiers of various second coding modes. S The encoding mode is determined as the second encoding mode, so that when the second encoding mode is subsequently used to compress and encode the data sequence, it can be more suitable for the data size distribution of the data sequence, and the compression rate of the data sequence can be further improved.

[0096] In another possible implementation, the encoder also determines the maximum compression bit width of the data sequence, which is the compression bit width of the largest data in the data sequence. If the bit width of the coding unit in each second coding mode is smaller than the maximum compression bit width, the encoder determines a third coding mode based on the maximum compression bit width. The third coding mode is a coding mode in which the sum of the bit widths of each coding unit in the third coding mode is equal to the bit width of the payload, and the bit width of each coding unit is greater than or equal to the maximum compression bit width, so that the third coding mode can compress and encode the largest data in the data sequence, avoiding the situation where the maximum data cannot be compressed and encoded using the coding mode. Therefore, the third coding mode is also called a secure coding mode.

[0097] The method for determining the third coding mode is, for example, recording the maximum compression bit width as w, recording the bit width of the payload as P, if P is divisible by w, the encoder groups multiple w into a candidate bit width group, at this time, the candidate bit width group is [w, w, ..., w], and the sum of each w in the candidate bit width group is equal to P; if P is not divisible by w, the encoder groups multiple r into a candidate bit width group, the candidate bit width group is [r, r, ..., r], and the sum of each r in the candidate bit width group is equal to P, r is greater than w and less than or equal to P, and r is divisible by P. After determining the candidate bit width group, the encoder generates the division information of the third coding mode based on the candidate bit width group, and the division information includes the candidate bit width group to indicate the division of the coding unit according to the candidate bit width group, and each bit width in the candidate bit width group is the bit width of each coding unit under the third coding mode, so as to determine the third coding mode.

[0098] After the third coding mode is determined, if the number of types of the second coding mode is less than 2 S , the encoder determines the third encoding mode as a new second encoding mode to supplement the determined second encoding mode. If the number of types of the second encoding mode is equal to 2 S , the encoder replaces the second encoding mode with the lowest frequency with the third encoding mode.

[0099] Since the bit width of the coding unit in the third coding mode is greater than or equal to the maximum compression bit width of the data sequence, the third coding mode is determined as the second coding mode so that the compressed data of the maximum data in the data sequence can be stored through the coding unit in the third coding mode, thereby avoiding the situation where the maximum data cannot be compressed and encoded.

[0100] Some Simple algorithms in related technologies have good adaptability and robust performance, but the calculations are complex and not suitable for scenarios with large amounts of data and low computing power, such as network equipment. For example, Simple algorithms for splicing words, such as Relative10, Carryover12, and Slide, can only work when overly wide elements (such as compressed data) happen to be at the edge of the word, and they rely on serial decompression and are difficult to accelerate in parallel. Among them, Relative10 and Carryover12 use the wasted load bit width of adjacent words to store memory. Slide splices words together to avoid wasting space because the element corresponding to the last coding unit is too wide, resulting in only being able to select a coding mode with fewer coding units. For example, Simple algorithms that adaptively optimize the payload width, such as vector of splits encoding (VSEncoding) and memory inverted list compression (MLIC), must first scan the data sequence completely before they can run. In addition, their adaptive process relies on dynamic programming, and the overall computational overhead is too large. Among them, VSEncoding uses dynamic programming to adaptively optimize the bit width of the payload when the bit width of the encoding unit is determined. MILC is a single instruction multiple data (SIMD) acceleration of the optimized prefix frame of reference (Opt-PFor) algorithm. The Opt-PFor algorithm is a Simple algorithm. The Opt-PFor algorithm takes out the 10% elements with the largest bit width in each word and encodes them separately.

[0101] In the process shown in the above steps 4021 and 4022, during the training coding mode, the compressed bit width of multiple sample data is divided into multiple bit width groups based on the bit width of the payload, and then each candidate coding mode is determined based on the total bit width of each bit width group and the bit width of the payload, and the second coding mode is determined from the candidate coding modes. There is no need to use dynamic programming to determine the second coding mode, which reduces the amount of calculation. Therefore, the data compression method provided by the present application has low requirements on computing power, and thus, under low computing power settings, the encoder can also efficiently compress and encode large-scale sequence data, and can automatically adapt and match different data distributions to ensure the compression rate.

[0102] 403. The encoder performs compression encoding on the data sequence based on the second encoding mode.

[0103] The encoder determines a second coding scheme based on each second coding mode determined in the above step 402. The second coding scheme is a Simple coding scheme suitable for the data sequence trained based on the sample data. For example, the encoder generates scheme information of the second coding scheme based on each second coding mode. The scheme information includes a bit width group of each second coding mode and a mode identifier of each second coding mode, and the bit width group of each second coding mode corresponds to its own mode identifier, such as the scheme information shown in Table 1 below.

[0104] Table 1

[0105] Mode flag of the second encoding mode Bit width group of the second coding mode 00 [4,4,4,8,4,4,4] 01 [2,3,8,4,3,4,8] 10 [3,8,4,3,4,8,2] 11 [8,8,8,8]

[0106] In another possible implementation, if any second coding mode still has residual bit width, the scheme information of the second coding scheme also includes the residual bit width of the second coding mode, and the residual bit width corresponds to the mode identifier and bit width group of the second coding mode.

[0107] After determining the second coding scheme, the encoder compresses and encodes the data sequence based on the second coding scheme to obtain at least one word. Next, based on the following steps 4031 to 4032, the process is described in detail.

[0108] Step 4031: The encoder obtains compressed data of each data to be compressed in the data sequence to obtain a compressed data sequence.

[0109] The compressed data sequence includes compressed data of each data to be compressed in the data sequence, and the ordering method of the compressed data in the compressed data sequence is the same as the ordering method of the data to be compressed in the data sequence.

[0110] For example, for any data to be compressed in the data sequence, the encoder removes the leading zero from the binary representation of the data to be compressed as the compressed data of the data to be compressed, and the compressed data of each data to be compressed form the compressed data sequence. Taking the data sequence [2, 1, 127, 3, 1, 5, 198, 0, 2, ...] as an example, the data compression sequence of the data sequence is [10, 1, 1111111, 11, 1, 101, 11000110, 0, 10, ...].

[0111] Step 4032: Based on each second encoding mode in the second encoding scheme, encapsulate the compressed data in the data compression sequence into at least one word, each encoding unit in the payload of each word is used to store a compressed data in the data compression sequence, and the selector in each word is used to store the mode identifier of the second encoding mode to which the encoding unit in the payload belongs.

[0112] The compressed data stored in the encoding unit in each word is continuous in the compressed data sequence.

[0113] The encoder starts with the first compressed data in the data compression sequence, and tries the second encoding mode with the most encoding units in the second encoding scheme to see whether the second encoding mode with the most encoding units can store multiple consecutive compressed data starting from the first compressed data in the data compression sequence. If so, the encoder determines the second encoding mode with the most encoding units as the fourth encoding mode. Otherwise, the encoder tries whether the second encoding mode with the second most encoding units can store the multiple consecutive compressed data until the fourth encoding mode is determined, wherein the fourth encoding mode is the second encoding mode to be used, i.e., the target encoding mode, and the encoding units under the fourth encoding mode can store multiple consecutive compressed data in the data sequence, and the fourth encoding mode is applicable to the multiple consecutive compressed data.

[0114] When trying to see whether any second encoding mode can store multiple continuous compressed data, the encoder can first obtain the bit width of the multiple continuous compressed data, and compare the bit width of the multiple continuous compressed data with the bit width of the encoding unit under the second encoding mode in turn to determine whether the encoding unit under the second encoding mode can store multiple continuous compressed data; if the encoding unit under the second encoding mode can store multiple continuous compressed data, the second encoding mode is determined as the fourth encoding mode; otherwise, an attempt is made to see whether another second encoding mode can store multiple continuous compressed data.

[0115] Taking the compressed sequence [10, 1, 1111111, 11, 1, 101, 11000110, 0, 10, …] as an example, assuming that the second coding scheme is as shown in Table 1, for the convenience of description, the second coding modes in Table 1 are referred to as coding modes. Starting from the first compressed data "10" in the compressed sequence, try the coding mode with the most coding units in the second coding scheme. Since there are 3 coding modes with the most coding units in the second coding scheme shown in Table 1, the encoder can try from any coding mode with the most coding units, such as starting from the first coding mode "00" with the most coding units: the bit width of the compressed data "10" is smaller than the bit width of the first coding mode "00" under the coding mode "00". The bit width of a coding unit is "4", which means that the first coding unit can store the compressed data "10"; the bit width of the next compressed data "1" of the compressed data "10" is smaller than the bit width "4" of the second coding unit under the coding mode "00", which means that the second coding unit can store the compressed data "1"; the bit width of the next compressed data "1111111" of the compressed data "1" is larger than the bit width "4" of the third coding unit under the coding mode "00", which means that the third coding unit cannot store the compressed data "1111111", and the coding unit under the coding mode "00" cannot store continuous compressed data starting from the first compressed data, so the attempt of coding mode "00" fails.

[0116] The encoder continues to try the coding mode "01" with the second most coding units in the second coding scheme: the bit width of the compressed data "10" is equal to the bit width "2" of the first coding unit under the coding mode "01", indicating that the first coding unit can store the compressed data "10"; the bit width of the next compressed data "1" of the compressed data "01" is less than the bit width "3" of the second coding unit under the coding mode "01", indicating that the second coding unit can store the compressed data "1"; the bit width of the next compressed data "1111111" of the compressed data "1" is less than the bit width "8" of the third coding unit under the coding mode "01", indicating that the third coding unit can store the compressed data "1111111"; the bit width of the next compressed data "11" of the compressed data "1111111" is less than the bit width "4" of the fourth coding unit under the coding mode "01", indicating that the fourth coding unit It can store the compressed data "11"; the bit width of the next compressed data "1" of the compressed data "11" is smaller than the bit width "3" of the fifth coding unit under the coding mode "01", which means that the fifth coding unit can store the compressed data "1"; the bit width of the next compressed data "101" of the compressed data "1" is smaller than the bit width "4" of the sixth coding unit under the coding mode "01", which means that the sixth coding unit can store the compressed data "101"; the bit width of the next compressed data "11000110" of the compressed data "101" is equal to the bit width "8" of the seventh coding unit under the coding mode "01", which means that the seventh coding unit can store the compressed data "11000110". It can be seen that the 7 coding units under the coding mode "01" can store the first 7 consecutive compressed data in the data sequence in sequence, and the encoder determines the coding mode "01" as the fourth coding mode.

[0117] After determining the fourth coding mode, the encoder encodes multiple continuous compressed data in the data sequence applicable to the fourth coding mode based on the fourth coding mode to obtain a word. For example, the encoder encodes the multiple continuous compressed data into each coding unit in the fourth coding mode in sequence, each coding unit is used to store a compressed data, encodes the mode identifier of the first coding mode to the selector, and the selector and the coding unit storing the compressed data in the fourth coding mode form a word. If there is a residual bit width in the load under the first coding mode, the encoder forms the selector, the coding unit storing the compressed data in the fourth coding mode, and the number of 0s of the residual bit width into a word.

[0118] Still taking the encoding mode "01" in Table 1 as the fourth encoding mode, and the plurality of continuous compressed data being the first 7 compressed data of the above compression sequence as an example, Figure 5As shown in the structural diagram of the word, the encoder encodes the mode identifier "01" of the encoding mode "01" to the selector, and encodes 7 compressed data such as "10", "1", "1111111", "11", "1", "101" and "11000110" to the encoding unit 1 to the encoding unit 7 under the encoding mode "01" in sequence. Figure 5 As shown, when encoding compressed data in a coding unit, the encoder may first encode the compressed data from the low bits of the coding unit, and after the compressed data is encoded, if there are still high bits remaining in the coding unit, the remaining high bits may be filled with 0. In other embodiments, the encoder may also first encode the compressed data from the high bits of the coding unit, and after the compressed data is encoded, if there are still low bits remaining in the coding unit, the remaining low bits may be filled with 0.

[0119] For the compressed data that has not been encoded in the data sequence, the encoder starts from the first compressed data in the compressed data that has not been encoded, and tries again from the second encoding mode with the most encoding units in the second encoding scheme to see whether the second encoding mode with the most encoding units can store multiple continuous compressed data starting from the first compressed data in the data compression sequence. If it can, the encoder determines the second encoding mode with the most encoding units as the new fourth encoding mode. Otherwise, the encoder tries whether the second encoding mode with the second most encoding units can store the multiple continuous compressed data until the new fourth encoding mode is determined, and encodes the multiple continuous compressed data based on the new fourth encoding mode to obtain a new word. This process is repeated until the last compressed data in the data compression sequence is encoded, and the data sequence is compressed and encoded into at least one word.

[0120] After compressing the data sequence into at least one word, the encoder stores the payload of each word into the storage space of the payload, and stores the selector of each word into the storage space of the selector, so that the at least one word can be subsequently restored based on the stored payload and the stored selector, and the at least one word can be sent to the decoder for decoding by the decoder. Alternatively, after compressing the data sequence into at least one word, the payload of each word is not stored into the storage space of the payload, and the selector of each word is not stored into the storage space of the selector, but the at least one word is sent to the decoder, and the decoder receives the at least one word, and restores the data in the data sequence based on the at least one word. For example, each time the decoder receives a word, based on the mode identifier stored in the selector of the word, the decoder determines the second encoding mode used by the word, based on the division method indicated by the second encoding mode, determines each encoding unit in the payload of the word, obtains a compressed data from each encoding unit, converts the compressed data into a data to be compressed, and the restored data to be compressed form the data sequence.

[0121] Among them, the decoder can be the above Figure 1 The decoder 102 in the decoder, in addition, before the decoder decodes the word, the decoder also obtains the scheme information of the second coding scheme from the encoder, so as to decode the mode identifier stored in the selector based on the word, and queries the second coding mode indicated by the mode identifier and the bit width group of the second coding mode from the scheme information, and determines each coding unit in the payload of the word by dividing the second coding mode with the bit width group.

[0122] based on Figure 4 The method flow shown, by obtaining sample data based on the data sequence to be compressed, so that the sample data can reflect the actual size distribution characteristics of the data sequence to be compressed, and then training the encoding mode of the Simple encoding scheme based on the sample data, so that the encoding unit under the trained encoding mode can meet the actual size distribution of the data sequence, so that the trained encoding mode can be applicable to the data sequence, so that the data sequence is compressed and encoded using the trained encoding mode, which can ensure the compression rate of the data sequence. For each data sequence to be compressed, the encoder can customize an applicable encoding mode for each data sequence according to the data compression method provided by the present application. Even if the data distribution of different data sequences is different, based on the encoding modes customized for different data sequences, different data sequences are compressed and encoded, and the compression rate of each data sequence can be guaranteed, avoiding the situation where the compression rate deteriorates when facing different data sequences with different data size distributions. By using sampled data to train the encoding mode of the Simple encoding scheme to update the Simple encoding scheme, the updating and configuration of the encoding scheme can be automatically completed without manual intervention. In the process of training the coding mode, based on the bit width of the payload, the compressed bit width of multiple sample data is divided into multiple bit width groups, and then based on the total bit width of each bit width group and the bit width of the payload, each candidate coding mode is determined, and the second coding mode is determined from the candidate coding modes. There is no need to use dynamic programming to determine the second coding mode, which reduces the amount of calculation. Therefore, the data compression method provided by the present application has low requirements on computing power, and even under low computing power settings, the encoder can also efficiently compress and encode large-scale sequence data, and can automatically adapt and match different data distributions to ensure the compression rate.

[0123] The above is introduced using the example of the data to be compressed in the data sequence being integer data. In another possible implementation, the data to be compressed in the data sequence is not integer data but non-integer data. In this case, the encoder may first convert each data to be compressed in the data to be compressed into integer data, and execute the above steps 401 to 403 on the converted data sequence.

[0124] In another possible implementation, the encoder supports at least one of an automatic encoding function and a fixed encoding function, wherein the automatic encoding function refers to using the data compression method provided in the present application to compress and encode the data sequence to be compressed, and the fixed encoding function refers to that the encoder uses a coding scheme to compress and encode the data sequence to be compressed. When the user turns on the automatic encoding function of the encoder, the encoder uses the data compression method provided in the present application to compress and encode the data sequence each time it receives a data sequence to be compressed. When the user turns on the fixed encoding function of the encoder, the encoder also provides a configuration interface for the fixed encoding scheme, which displays the scheme information of the Simple encoding scheme to be used. The user can modify the bit width of the encoding unit under each encoding mode in the scheme information in the configuration interface to update the configuration of the Simple encoding scheme. Each time the encoder receives a data sequence to be compressed, the Simple encoding scheme indicated by the scheme information in the configuration interface is used to compress and encode the data sequence.

[0125] In another possible implementation, after determining the second coding scheme, the encoder does not perform the step of compressing and encoding the data sequence based on the second coding scheme to obtain at least one word. Instead, other devices provide the second coding scheme, and other devices compress and encode the data sequence based on the second coding scheme provided by the encoder to obtain at least one word. In this implementation, the encoder is responsible for customizing the coding scheme for the data sequence to be compressed, but is not responsible for compressing and encoding the data sequence. At this time, the encoder may not be called an encoder, but may be called a Simple coding scheme determination device. Step 401, step 402, and the method flow consisting of the step of determining the second coding scheme based on the above step 402 are called the flow of the Simple coding scheme determination method.

[0126] The above describes the method of the embodiment of the present application, and the following describes the device of the embodiment of the present application. It should be understood that the device described below has any function of the encoder in the above method. Figures 1 to 5 The transmission compression method according to the embodiment of the present application is described in detail. Based on the same inventive concept, the following will be combined with Figure 6 to Figure 7 The device according to the embodiment of the present application is described. It should be understood that the technical features described in the method embodiment are also applicable to the following device embodiment.

[0127] Figure 6 is a structural schematic diagram of a data compression device provided in an embodiment of the present application, Figure 6 The device 600 shown may be an encoder or a part of an encoder in the above embodiments, and is used to execute the data compression method executed by the encoder. Figure 6 As shown, the device 600 includes:

[0128] An acquisition module 601 is used to acquire a plurality of sample data based on a data sequence;

[0129] A training module 602 is used to train a first coding mode based on the multiple sample data to obtain a second coding mode, where the first coding mode is used to indicate a division method of coding units in a payload encoded based on a simple coding scheme;

[0130] The encoding module 603 is used to compress and encode the data sequence based on the second encoding mode.

[0131] In a possible implementation, the training module 602 includes:

[0132] A training unit, configured to train the first coding mode based on the compression bit widths of the plurality of sample data and the bit width of the payload, to obtain a plurality of candidate coding modes, each of the candidate coding modes corresponding to at least one sample data, and the compression bit width being the bit width of compressed data of the sample data;

[0133] A determination unit is used to determine the second coding mode from the multiple candidate coding modes.

[0134] In a possible implementation, the training unit is used to:

[0135] Based on the bit width of the payload, the compressed bit widths of the plurality of sample data are divided into a plurality of bit width groups, each of the bit width groups includes the compressed bit width of at least one sample data, the total bit width of the bit width groups is less than or equal to the bit width of the payload, and the total bit width is the sum of the compressed bit widths in the bit width groups;

[0136] For any of the bit width groups, based on the total bit width of the bit width group and the bit width of the payload, a candidate encoding mode corresponding to the bit width group is determined, each encoding unit under the candidate encoding mode corresponds to a compressed bit width in the bit width group, the bit width of each encoding unit under the candidate encoding mode is less than or equal to the corresponding compressed bit width, and the total bit width of each encoding unit under the candidate encoding mode is equal to the bit width of the payload.

[0137] In a possible implementation manner, the determining unit is configured to:

[0138] At least one coding mode with the highest occurrence frequency among the multiple candidate coding modes is determined as the second coding mode.

[0139] In a possible implementation, the apparatus 600 further includes:

[0140] a determination module, configured to determine, if the bit width of the coding unit in each of the second coding modes is less than the maximum compression bit width of the data sequence, a third coding mode based on the maximum compression bit width, wherein the maximum compression bit width is the compression bit width of the largest data in the data sequence, and the bit width of each coding unit in the third coding mode is greater than or equal to the maximum compression bit width;

[0141] A replacement module is used to replace the second encoding mode with the lowest frequency with the third encoding mode.

[0142] It should be understood that the device 600 corresponds to the encoder in the above-mentioned method embodiment, and the modules in the device 600 and the above-mentioned other operations and / or functions are respectively various steps and methods implemented by the encoder in the method embodiment to implement the specific details. Please refer to the above-mentioned method embodiment. For the sake of brevity, they will not be repeated here.

[0143] It should be understood that when the device 600 compresses and encodes the data sequence, only the division of the above-mentioned functional modules is used as an example. In practical applications, the above-mentioned functions can be assigned to different functional modules as needed, that is, the internal structure of the device 600 is divided into different functional modules to complete all or part of the functions described above. In addition, the device 600 provided in the above embodiment and the above method embodiment belong to the same concept, and the specific implementation process thereof is detailed in the above method embodiment, which will not be repeated here.

[0144] Figure 7 is a schematic diagram of a computing device provided in an embodiment of the present application, such as Figure 7 As shown, computing device 700 includes: bus 702, processor 704, memory 706 and communication interface 708. Processor 704, memory 706 and communication interface 708 communicate through bus 702. Computing device 700 can be an electronic device such as a server or terminal device. It should be understood that the present application does not limit the number of processors and memories in computing device 700.

[0145] The bus 702 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 7 The bus 702 may include a path for transmitting information between various components of the computing device 700 (eg, the memory 706, the processor 704, and the communication interface 708).

[0146] The processor 704 may include any one or more processors such as a CPU, a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0147] The memory 706 may include a volatile memory, such as a random access memory (RAM). The processor 704 may also include a non-volatile memory, such as a read-only memory (ROM), a flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0148] The memory 706 stores executable program codes, and the processor 704 executes the executable program codes to respectively implement the functions of the acquisition module 601, the training module 602, and the encoding module 603, thereby implementing the data compression method provided by the present application. That is, the memory 706 stores instructions for executing the data compression method.

[0149] The communication interface 708 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 700 and other devices or communication networks.

[0150] In another possible implementation, the computing device 700 is a chip, and the chip is implemented by any combination of ASIC, PLD, CPLD, FPGA, GAL, etc.

[0151] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including a program code, and the program code can be executed by a processor in a computing device to perform the data compression method in the above embodiment. For example, the computer-readable storage medium is a non-temporary computer-readable storage medium, such as a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), a magnetic tape, a floppy disk, and an optical data storage device.

[0152] An embodiment of the present application also provides a computer program product or a computer program, which includes a program code. The computer instructions are stored in a computer-readable storage medium. The processor of the computing device reads the program code from the computer-readable storage medium, and the processor executes the program code, so that the computing device executes the above-mentioned data compression method.

[0153] In addition, an embodiment of the present application also provides a device, which can specifically be a chip, component or module, and the device may include a connected processor and memory; wherein the memory is used to store computer execution instructions, and when the device is running, the processor can execute the computer execution instructions stored in the memory so that the chip executes the data compression method in the above-mentioned method embodiments.

[0154] Among them, the device, equipment, computer-readable storage medium, computer program product or chip provided in this embodiment is used to execute the corresponding method provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects in the corresponding method provided above, and will not be repeated here.

[0155] Through the description of the above implementation mode, those skilled in the art can understand that, for the convenience and simplicity of description, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above. In addition, the data compression method embodiment provided in the above embodiment belongs to the same concept, and its specific implementation process is detailed in the method embodiment, which will not be repeated here.

[0156] In the several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0157] The units described as separate components may or may not be physically separated, and the components shown as units may be one physical unit or multiple physical units, that is, they may be located in one place or distributed in multiple different places. Some or all of the units may be selected according to actual needs to achieve the purpose of the present embodiment.

[0158] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0159] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium, including several instructions to enable a device (which can be a single-chip microcomputer, chip, etc.) or a processor (processor) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard drives, ROM, RAM, magnetic disks or optical disks.

[0160] In the description of this application, unless otherwise specified, " / " means "or", for example, A / B can mean A or B. "And / or" in this article is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone. In addition, "at least one" means one or more, and "plurality" means two or more. The words "first", "second", etc. do not limit the quantity and execution order, and the words "first", "second", etc. do not limit them to be different.

[0161] In this application, the words "exemplary" or "for example" are used to indicate examples, illustrations or descriptions. Any embodiment or design described as "exemplary" or "for example" in this application should not be interpreted as being more preferred or more advantageous than other embodiments or designs. Specifically, the use of the words "exemplary" or "for example" is intended to present the related concepts in a concrete way.

[0162] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions. For example, the data sequences involved in this application are all obtained with full authorization.

[0163] All the above optional technical solutions can be arbitrarily combined to form optional embodiments of the present disclosure, and will not be described in detail here.

[0164] The above description is only an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present application shall be included in the protection scope of the present application.

Claims

1. A data compression method, characterized in that: The method comprises: Acquire multiple sample data based on the data sequence; Training a first coding mode based on the plurality of sample data to obtain a second coding mode, wherein the first coding mode is used to indicate a division method of coding units in a payload encoded based on a simple coding scheme; The data sequence is compression-encoded based on the second encoding mode.

2. The method according to claim 1, characterized in that The training of the first encoding mode based on the plurality of sample data to obtain the second encoding mode comprises: Based on the compressed bit widths of the multiple sample data and the bit width of the payload, the first coding mode is trained to obtain multiple candidate coding modes, each of the candidate coding modes corresponds to at least one sample data, and the compressed bit width is the bit width of compressed data of the sample data; The second encoding mode is determined from the multiple candidate encoding modes.

3. The method according to claim 2, characterized in that The training of the first coding mode based on the compression bit width of the plurality of sample data and the bit width of the payload to obtain a plurality of candidate coding modes includes: Based on the bit width of the payload, the compressed bit widths of the plurality of sample data are divided into a plurality of bit width groups, each of the bit width groups includes the compressed bit width of at least one sample data, the total bit width of the bit width groups is less than or equal to the bit width of the payload, and the total bit width is the sum of the compressed bit widths in the bit width groups; For any of the bit width groups, based on the total bit width of the bit width group and the bit width of the payload, a candidate encoding mode corresponding to the bit width group is determined, each encoding unit under the candidate encoding mode corresponds to a compressed bit width in the bit width group, the bit width of each encoding unit under the candidate encoding mode is less than or equal to the corresponding compressed bit width, and the total bit width of each encoding unit under the candidate encoding mode is equal to the bit width of the payload.

4. The method according to claim 2 or 3, characterized in that: The determining the second coding mode from the multiple candidate coding modes comprises: At least one coding mode with the highest occurrence frequency among the multiple candidate coding modes is determined as the second coding mode.

5. The method according to claim 4, characterized in that The method further comprises: If the bit width of the coding unit in each of the second coding modes is smaller than the maximum compression bit width of the data sequence, a third coding mode is determined based on the maximum compression bit width, the maximum compression bit width is the compression bit width of the largest data in the data sequence, and the bit width of each coding unit in the third coding mode is greater than or equal to the maximum compression bit width; The second encoding mode with the lowest occurrence frequency is replaced by the third encoding mode.

6. A data compression device, characterized in that: The device comprises: An acquisition module, used for acquiring multiple sample data based on a data sequence; A training module, used for training a first coding mode based on the plurality of sample data to obtain a second coding mode, wherein the first coding mode is used to indicate a division method of coding units in a payload encoded based on a simple coding scheme; An encoding module is used to compress and encode the data sequence based on the second encoding mode.

7. The device according to claim 6, characterized in that The training module includes: A training unit, configured to train the first coding mode based on the compression bit widths of the plurality of sample data and the bit width of the payload, to obtain a plurality of candidate coding modes, each of the candidate coding modes corresponding to at least one sample data, and the compression bit width being the bit width of compressed data of the sample data; A determination unit is used to determine the second coding mode from the multiple candidate coding modes.

8. The device according to claim 7, characterized in that The training unit is used for: Based on the bit width of the payload, the compressed bit widths of the plurality of sample data are divided into a plurality of bit width groups, each of the bit width groups includes the compressed bit width of at least one sample data, the total bit width of the bit width groups is less than or equal to the bit width of the payload, and the total bit width is the sum of the compressed bit widths in the bit width groups; For any of the bit width groups, based on the total bit width of the bit width group and the bit width of the payload, a candidate encoding mode corresponding to the bit width group is determined, each encoding unit under the candidate encoding mode corresponds to a compressed bit width in the bit width group, the bit width of each encoding unit under the candidate encoding mode is less than or equal to the corresponding compressed bit width, and the total bit width of each encoding unit under the candidate encoding mode is equal to the bit width of the payload.

9. The device according to claim 7 or 8, characterized in that The determining unit is used for: At least one coding mode with the highest occurrence frequency among the multiple candidate coding modes is determined as the second coding mode.

10. The device according to claim 9, characterized in that The device also includes: a determination module, configured to determine, if the bit width of the coding unit in each of the second coding modes is less than the maximum compression bit width of the data sequence, a third coding mode based on the maximum compression bit width, wherein the maximum compression bit width is the compression bit width of the largest data in the data sequence, and the bit width of each coding unit in the third coding mode is greater than or equal to the maximum compression bit width; A replacement module is used to replace the second encoding mode with the lowest frequency with the third encoding mode.

11. A computing device, characterized in that: The computing device comprises a processor, and the processor is configured to execute program code so that the computing device performs the method according to any one of claims 1 to 5.

12. A computer-readable storage medium, characterized in that: At least one program code is stored in the storage medium, and the at least one program code is read by the processor to enable the computing device to execute the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Image encoding device, image encoding method, and image encoding system

    CN101816182A

  • Video decoding method and apparatus, video coding method and apparatus, and storage medium

    CN110933409A

  • Data compression method and computing equipment

    CN113055017A

  • Lossless compression method and system for waveform data and medium

    CN115940960A

  • Data processing method and system based on ANS coding, storage medium and equipment

    CN116248130A