Data compression method, apparatus, computing device, and computer-readable storage medium
Patent Information
- Application Number
- CN202311429721.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-10-30
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2043-10-30
AI Technical Summary
[0004]本申请实施例提供了一种数据压缩方法、装置、计算设备以及计算机可读存储介质,能够解决使用单一Simple编码方案所导致的压缩率难以保障保证的问题
[0004] This application provides a data compression method, apparatus, computing device, and computer-readable storage medium, which can solve the problem that the compression ratio is difficult to guarantee when using a single Simple encoding scheme. The technical solution is as follows:
Smart Images

Figure CN119921780B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of compression coding technology, and in particular to a data compression method, apparatus, computing device, and computer-readable storage medium. Background Technology
[0002] Simple algorithms are a family of algorithms widely used for compressing time-series data, integer data, and other types of information. Time-series data includes data from various time-series databases such as InfluxDB and TimescaleDB, while integer data includes inverted index tables from major search engines. Simple algorithms compress and encode the data using an encoding scheme; this encoding scheme is also known as the Simple encoding scheme.
[0003] There are currently several Simple-class algorithms, such as Simple9, Simple16, and Simple8b. Different Simple-class algorithms have different Simple encoding schemes, and correspondingly, different Simple encoding schemes have different encoding modes. Once the Simple-class algorithm is determined, the Simple encoding scheme is also determined. If the encoding modes in the Simple encoding scheme do not cover the actual size distribution of the data to be compressed, continuing to compress the data according to the encoding modes in the Simple encoding scheme will make it difficult to guarantee the compression rate of the data. Summary of the Invention
[0004] This application provides a data compression method, apparatus, computing device, and computer-readable storage medium, which can solve the problem that the compression ratio is difficult to guarantee when using a single Simple encoding scheme. The technical solution is as follows: Firstly, a data compression method is provided, which includes the following steps: for any data sequence, firstly, multiple sample data are obtained based on the data sequence; then, a first encoding mode is trained based on the multiple sample data to obtain a second encoding mode; then, the data sequence is compressed and encoded based on the second encoding mode, wherein the first encoding mode is used to indicate the division method of encoding units in the payload encoded based on the Simple encoding scheme.
[0005] This method obtains sample data based on the data sequence, so that the sample data can reflect the actual size distribution characteristics of the data sequence to be compressed. Then, it trains the encoding mode of the Simple encoding scheme based on the sample data, so that the encoding units in the trained encoding mode can meet the actual size distribution of the data sequence, and the trained encoding mode can be applied to the data sequence. Thus, the data sequence can be compressed and encoded using the trained encoding mode, which can ensure the compression rate of the data sequence.
[0006] In one possible implementation, the process of training the first coding mode based on multiple sample data to obtain the second coding mode includes: firstly, training the first coding mode based on the compressed bit width and the bit width of the payload of multiple sample data to obtain multiple candidate coding modes, and then determining the second coding mode from the multiple candidate coding modes, wherein each candidate coding mode corresponds to at least one sample data, and the compressed bit width is the bit width of the compressed data of the sample data.
[0007] Based on the above possible implementation methods, the determined second encoding mode can be applied to the data sequence, so that the compression rate of the data sequence can be further ensured when the determined second encoding mode is used to process the data sequence in the future.
[0008] In one possible implementation, the process of training the first coding mode based on the compressed bit width of multiple sample data and the bit width of the payload to obtain multiple candidate coding modes includes: dividing the compressed bit width of multiple sample data into multiple bit width groups based on the bit width of the payload, each bit width group including the compressed bit width of at least one sample data, the total bit width of the bit width group being less than or equal to the bit width of the payload, and the total bit width being the sum of the compressed bit widths in the bit width group; for any bit width group, determining the candidate coding mode corresponding to the bit width group based on the total bit width of the bit width group and the bit width of the payload, wherein each coding unit in the candidate coding mode corresponds to a compressed bit width in the bit width group, the bit width of each coding unit in the candidate coding mode is greater than or equal to the corresponding compressed bit width, and the total bit width of each coding unit in the candidate coding mode is equal to the bit width of the payload.
[0009] Based on the above possible implementation methods, the compressed bit width of multiple sample data is divided into multiple bit width groups based on the bit width of the payload. Then, based on the total bit width of each bit width group and the bit width of the payload, each candidate coding mode is determined. The second coding mode is determined from the candidate coding modes. There is no need to use dynamic programming to determine the second coding mode, which reduces the amount of computation. Therefore, even with low computing power settings, it is possible to perform efficient compression coding on large-scale sequence data.
[0010] In one possible implementation, the process of determining the second coding mode from multiple candidate coding modes includes: determining at least one coding mode that appears most frequently among the multiple candidate coding modes as the second coding mode.
[0011] Based on the above possible implementation methods, when the second encoding mode is used to compress and encode the data sequence, it can be more suitable for the data size distribution of the data sequence, and can further improve the compression rate of the data sequence.
[0012] In one possible implementation, the method further includes the following steps: if the bit width of the coding unit in each second coding mode is less than the maximum compressed bit width of the data sequence, a third coding mode is determined based on the maximum compressed bit width; the second coding mode with the lowest frequency of occurrence is replaced with the third coding mode, wherein the maximum compressed bit width is the compressed bit width of the largest data in the data sequence, and the bit width of each coding unit in the third coding mode is greater than or equal to the maximum compressed bit width.
[0013] Based on the above possible implementation methods, the compressed data of the largest data in the data sequence can be stored through the encoding unit in the third encoding mode, so as to avoid the situation where the largest data cannot be compressed and encoded.
[0014] In a second aspect, a data compression apparatus is provided, including a functional module for performing the data compression method provided in the first aspect or any alternative manner of the first aspect.
[0015] Thirdly, a computing device is provided, the computing device including a processor for executing program code, causing the computing device to perform a data compression method as provided in the first aspect or any alternative method of the first aspect.
[0016] Fourthly, a computer-readable storage medium is provided, wherein at least one piece of program code is stored therein, which is read by a processor to cause a computing device to perform a data compression method as provided in the first aspect or any alternative method of the first aspect.
[0017] Fifthly, a computer program product or computer program is provided, the computer program product or computer program including program code stored in a computer-readable storage medium, a processor of a computing device reading the program code from the computer-readable storage medium, the processor executing the program code, causing the computer device to perform the method provided in the first aspect or various optional implementations of the first aspect.
[0018] Based on the implementation methods provided in the above aspects, this application can be further combined to provide more implementation methods. Attached Figure Description
[0019] Figure 1 This is a schematic diagram of the implementation environment of an application data compression method provided in an embodiment of this application; Figure 2 This is a schematic diagram of an inverted index provided in an embodiment of this application; Figure 3 This is a storage space distribution diagram of a word provided in an embodiment of this application; Figure 4 This is a flowchart of a data compression method provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of a character provided in an embodiment of this application; Figure 6 This is a schematic diagram of the structure of a data compression device provided in an embodiment of this application; Figure 7 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application. Detailed Implementation
[0020] To facilitate understanding of this application, some of the terms used in this application will be introduced below: Binary area: This is the area used to store binary data.
[0021] A word is the basic output unit obtained by Simple-class algorithms. One or more pieces of data to be compressed can be compressed into a word. A word is a binary region, and each word consists of a selector and a payload. The bit width of the word is the sum of the bit width of the selector and the bit width of the payload. For Simple-class algorithms, the bit width of the selector, the bit width of the payload, and the bit width of the word are all set values. For example, in Simple8b, the selector has a bit width of 4 bits, and the payload has a bit width of 60 bits. A 4-bit selector and a 60-bit payload together form a 64-bit word.
[0022] Payload: A binary region used to store compressed data of consecutive elements in the input sequence. The input sequence is a data sequence to be compressed, which includes multiple data items to be compressed, with each data item to be compressed being an element.
[0023] Unit: A payload is divided into one or more units, each of which is used to store compressed data of one element in the input sequence. In this application, the unit is also called an encoding unit.
[0024] Pattern: A method of dividing the payload into coding units is called a pattern. In this application, the pattern is also called a coding pattern. The coding units divided according to a coding pattern are called coding units under that coding pattern. The number of coding units is different under different coding patterns, and / or the bit width of the coding units is different under different coding patterns. The division method indicated by each coding pattern is represented by the division information of that coding pattern. For example, the division information includes the number of coding units under that coding pattern and the bit width of each coding unit under that coding pattern. According to the division information, the payload can be divided into a specified number of coding units and coding units of a specified bit width. In another possible implementation, the division information includes the unit bit width of each coding unit under that coding pattern, but does not include the number of coding units. The number of coding units is represented by the number of coding unit bit widths in the division information. In another possible implementation, if the bit width of each coding unit under that coding pattern is the same, the division information includes the number of coding units and the bit width of one coding unit, but does not include the bit width of all coding units under that coding pattern. In another possible implementation, if the sum of the bit widths of each coding unit in the coding mode is less than the bit width of the payload, the partitioning information also includes the remaining bit width. The remaining bit width is the bit width of the binary region in the payload other than the coding unit. The remaining bit width is not used to store compressed data, and the remaining bit width is also the wasted payload bit width in the payload.
[0025] Selector: A binary area used to record an encoding mode. For example, different encoding modes are identified by different mode identifiers. The selector is used to store the mode identifier of an encoding mode. The mode identifier consists of at least one bit of binary data, such as 0 or 1.
[0026] Encoding scheme: Used to indicate the correspondence between pattern identifiers and encoding patterns. An encoding scheme is represented by scheme information, which includes the partitioning information of at least one encoding pattern and the pattern identifier of each encoding pattern among the at least one encoding pattern. The partitioning information of each encoding pattern corresponds to its respective pattern identifier. In this application, the encoding scheme is also referred to as the Simple encoding scheme.
[0027] Next, in conjunction with the appendix Figure 1 The implementation environment involved in the embodiments of this application will be described by way of example.
[0028] Figure 1 This is a schematic diagram of an implementation environment for an application data compression method provided in this application embodiment. The implementation environment includes an encoder 101 and a decoder 102. The encoder 101 is a component in the implementation environment used to execute the data compression method provided in this application, and the decoder 102 is a component in the implementation environment used to decode the compression result of the encoder 101.
[0029] For example, for each input data sequence in encoder 101, encoder 101 performs the data compression method of this application on the data sequence to compress and encode the data sequence into at least one word, and sends the at least one word to decoder 102, which decodes the at least one word to recover the data sequence.
[0030] Either encoder 101 or decoder 102 can be implemented in software or in hardware. For example, encoder 101 will be used as an example to illustrate its implementation. The implementation of decoder 102 can be found in the implementation of encoder 101.
[0031] The encoder 101, as an example of a software functional unit, may include code running on a computing instance. The computing instance may include at least one of a physical host (computing device), a virtual machine, or a container. Furthermore, the aforementioned computing instance may be one or more. For example, the encoder 101 may include code running on multiple hosts / virtual machines / containers. It should be noted that the multiple hosts / virtual machines / containers used to run the code may be distributed within the same region or in different regions. Further, the multiple hosts / virtual machines / containers used to run the code may be distributed within the same availability zone (AZ) or in different AZs, each AZ including one or more geographically proximate data centers. Typically, a region may include multiple AZs.
[0032] Similarly, multiple hosts / virtual machines / containers used to run this code can be distributed within the same Virtual Private Cloud (VPC) or across multiple VPCs. Typically, a VPC is set up within a region. Communication between two VPCs within the same region, as well as between VPCs in different regions, requires a communication gateway to be set up within each VPC to enable interconnection between VPCs.
[0033] Encoder 101, as an example of a hardware functional unit, may include at least one computing device, such as a server. Alternatively, encoder 101 may also be a device implemented using an application-specific integrated circuit (ASIC) or a programmable logic device (PLD). The PLD may be implemented using a complex programmable logical device (CPLD), a field-programmable gate array (FPGA), generic array logic (GAL), or any combination thereof.
[0034] The encoder 101 includes multiple computing devices that can be distributed within the same region or in different regions. Similarly, the encoder 101 includes multiple computing devices that can be distributed within the same Availability Zone (AZ) or in different AZs. Likewise, the encoder 101 includes multiple computing devices that can be distributed within the same Virtual Private Cloud (VPC) or in multiple VPCs. These multiple computing devices can be any combination of computing devices such as servers, ASICs, PLDs, CPLDs, FPGAs, and GALs.
[0035] The encoder 101 and decoder 102 can be located in the same computing device, for example, both encoder 101 and decoder 102 can be located on a server. Alternatively, the encoder 101 and decoder 102 can be located in different computing devices, for example, encoder 101 can be located on one server and decoder 102 can be located on another server.
[0036] Before the encoder (such as encoder 101 mentioned above) executes the data compression method provided in this application, the encoder can be configured in the following two ways: 1. Configure the bit width of the load and the bit width of the selector.
[0037] For ease of description, the bit width of the load is denoted as P, and the bit width of the selector is denoted as S, where both P and S are integers greater than 0. Since a load and a selector constitute a word, the sum of the configured P and S equals the word bit width.
[0038] For example, a user can use the word length of the computing device where the encoder resides as the bit width of the word to be encoded by the encoder. If the encoder is software, the computing device where the encoder resides is the device running the encoder; if the encoder is a hardware module within a device, the computing device where the encoder resides is that specific device; if the encoder is an independent computing device, then the computing device where the encoder resides is the encoder itself. The word length of the computing device refers to the length of the operation word within the central processing unit (CPU). Taking a 64-bit computing device as an example, the user can use 64 bits as the bit width of the word to be encoded by the encoder. After determining the bit width of the word to be encoded, the user configures P and S as constraints, such that the sum of P and S equals the bit width of the word.
[0039] In some embodiments, the user can also obtain the maximum element bit width of the data to be compressed from the data source used for the data to be compressed, and configure P with the maximum element bit width as a constraint. The configured P is greater than or equal to the maximum element bit width, and the sum of P and S is equal to the bit width of the word. The data type of the data to be compressed is integer, there are multiple data sets to be compressed, and the data source is a device or software used to provide the data to be compressed. The data source can be a time-series database, a search engine with an inverted index, or other devices, device clusters, or software used to provide the data to be compressed. This application embodiment does not limit the data source. Taking a time-series database providing the data to be compressed as an example, the time-series data in the time-series database is integer, and the time-series data can be used as the data to be compressed. Taking a search engine providing the data to be compressed as another example, the inverted index is an important data structure in a search engine. Figure 2 The diagram illustrates an inverted index. The document collection provided by the search engine includes multiple documents, each with a document identifier (ID) that is an integer, such as 1, 2, 3, 4, etc. For each word appearing in each document, a document ID sequence is formed using the IDs of the documents containing each word. All document ID sequences form an inverted index, and the IDs in each document ID sequence within the inverted index can be used as data to be compressed. In another possible implementation, the data source may also provide integer data of other application types besides time-series data or inverted indexes as data to be compressed. For example, taking a database storing traffic logs of communication devices as an example, a column in the traffic logs contains the number of terminal connections, which is an integer; this column of terminal connection counts can then be used as data to be compressed. Here, this embodiment does not limit the application type of the data to be compressed.
[0040] The maximum element bit width is the bit width of the compressed data for the largest amount of data to be compressed provided by the data source. Users set P to be greater than or equal to the maximum element bit width to avoid the compressed data for the largest amount of data to be compressed being unable to use the payload of P bits.
[0041] In some embodiments, given the characteristics of modern computer hardware, the user may also configure both P and S to powers of 2, so that the computer device where the encoder is located or the device receiving the word processes the word. For example, both P and S are configured to powers of 2, and P is greater than or equal to the maximum element bit width, and the sum of P and S equals the bit width of the word.
[0042] In some embodiments, if the data size distribution of the data to be compressed is more diverse, and more types of encoding modes are required, then S can be configured to be larger. Conversely, if the data size distribution of the data to be compressed is less diverse, and fewer types of encoding modes are required, S can be configured to be relatively smaller to meet the needs of different data size distributions. For example, if the data size distribution of the data to be compressed is diverse, S can be configured to 4 or 8; if the data size distribution of the data to be compressed is uniform, S can be configured to 2.
[0043] 2. Configure the storage space for the payload and the storage space for the selector.
[0044] The payload storage space is used to store the payloads of each word encoded by the encoder, and the selector storage space is used to store the selectors of each word encoded by the encoder.
[0045] Users can configure the payload's storage space and the selector's storage space to be the same or different storage spaces based on the distribution of the data size to be compressed. For example, if the data size distribution is diverse, the payload's storage space and the selector's storage space can be configured to be the same; if the data size distribution is uniform, the payload's storage space and the selector's storage space can be configured to be different. Figure 3 The diagram shows the storage space distribution of the characters. If the storage space for the payload and the storage space for the selector are configured to be the same, the subsequent encoder will store the payload and selector of each character sequentially in the same storage space when storing the encoded characters, thus achieving the encoding of the payload and selector into the same storage space. If the storage space for the payload and the storage space for the selector are configured to be different, the subsequent encoder will store the payload of each encoded character in the payload storage space and the selector of each character in the selector storage space, thus achieving the encoding of the payload and selector into different storage spaces.
[0046] In some embodiments, since storage space is generally a power of 2, if the bit width of the data stored in that storage space is also a power of 2, it facilitates data alignment with the storage address of the storage space, thereby improving storage space utilization. Based on this, if the storage space of the payload and the storage space of the selector are configured as the same storage space, the user further configures P and S with the constraint that the sum of P and S is a power of 2, so that the bit width of the words encoded by the subsequent encoder is a power of 2, enabling the encoded words to align with the storage address of that storage space, thus improving the utilization of that storage space. For example, when the maximum element bit width of the data to be compressed is 62 and the data size distribution of the data to be compressed is diverse, P can be configured as 64 and S as 4, configuring the storage space of the payload and the storage space of the selector as different storage spaces. When the maximum element bit width of the data to be compressed is 32 and the data size distribution of the data to be compressed is uniform, P can be configured as 30 and S as 2, configuring the storage space of the payload and the storage space of the selector as the same storage space.
[0047] If the storage space for the load and the storage space for the selector are configured as different storage spaces, the storage space for the load and the storage space for the selector can be provided by different storage devices or by the same storage device. The storage device and the encoder can be located on the same computing device or on different computing devices.
[0048] After the encoder is configured, it can compress and encode the data to be compressed by executing the data compression method provided in this application. Next, combined with... Figure 4 The data compression method provided in this application is described in detail.
[0049] Figure 4 This is a flowchart of a data compression method provided in an embodiment of this application. The method is executed by an encoder and includes the following steps.
[0050] 401. The encoder acquires multiple sample data based on the data sequence.
[0051] Here, the data sequence is any data sequence to be compressed provided by the data source. This data sequence includes multiple data items to be compressed, each of which is an integer. The data sequence can consist of all the data items to be compressed from the data source, or it can consist of a portion of the data items to be compressed from the data source. For example, if the data source has 10,000 data items to be compressed, these 10,000 data items can form one data sequence, or they can form 10 data sequences. The number of data items to be compressed in each data sequence can be the same or different. Figure 2Taking the inverted index shown as an example of all the data to be compressed from the data source, the document ID sequence for each word can be a data sequence. Here, this embodiment does not limit the number of data to be compressed in the data sequence.
[0052] When any data sequence to be compressed from the data source is input into the encoder, the encoder samples the data to be compressed in the data sequence to obtain n sample data, so that the data size distribution of the n sample data is similar to the data size distribution of the data sequence. At this time, each sample data is a piece of data to be compressed in the data sequence, where n is greater than 1 and less than the number of pieces of data to be compressed in the data sequence. For example, the encoder can use the first n pieces of data to be compressed in the data sequence as sample data, or use the last n pieces of data to be compressed in the data sequence as sample data, or the encoder can collect a sample data every m pieces of data to be compressed in the data sequence, where m is greater than or equal to 1 and less than the number of pieces of data to be compressed in the data sequence. Here, the sampling method of the data to be compressed in this embodiment of the application is not limited.
[0053] After acquiring n sample data, the encoder combines the n sample data into a sample sequence and uses this sample sequence as training samples to train the encoding mode of the Simple encoding scheme (as shown in step 402 below). Taking the first n data to be compressed in the data sequence as sample data as an example, assuming the data sequence is: [4, 2, 4, 130, 12, 6, 8, 191, 0, 2, 1, 127, 3, 1, 5, 198, 0, 2, ...], assuming n=9, the encoder combines the first 9 data to be compressed in the data sequence into a sample sequence [4, 2, 4, 130, 12, 6, 8, 191, 0].
[0054] In another possible implementation, different data sources are configured with different sample sequences in the encoder. The data size distribution of the sample data in each sample sequence is similar to the data size distribution of the data to be compressed in the corresponding data source. For any data sequence to be compressed, the encoder uses the sample sequence of the data source to which the data sequence belongs as the sample sequence of the data sequence, without having to obtain the sample sequence from the data sequence by sampling.
[0055] 402. The encoder trains the first coding mode based on the multiple sample data to obtain the trained second coding mode. The first coding mode is used to indicate the division of coding units in the payload encoded based on the Simple coding scheme.
[0056] The first encoding pattern is a pattern found in the Simple encoding scheme, and the second encoding pattern is a trained encoding pattern suitable for the data sequence; that is, the second encoding pattern is obtained by training the encoding patterns in the Simple encoding scheme. There can be at least one second encoding pattern, and at most two second encoding patterns. S kind.
[0057] In one possible implementation, during training, the encoder trains the encoding mode based on the compression bit width of each sample data in the sample sequence, the bit width P of the configured payload, and the bit width S of the selector to obtain a second encoding mode. The training process is described below with steps 4021 to 4022.
[0058] Step 4021: The encoder trains the first coding mode based on the compressed bit width of the multiple sample data and the bit width of the payload, obtaining multiple candidate coding modes, each corresponding to at least one sample data. The candidate coding modes are candidate modes for the second coding mode, and each candidate coding mode is trained based on its corresponding sample data. The bit width of the payload is P as configured above.
[0059] The compression bit width of any sample data is the bit width of the compressed data of that sample data. In this application, the compressed data of any data (such as sample data or the largest data to be compressed provided by the data source) refers to the binary representation of the sample data after removing the leading zeros. Taking a sample data that is the number 5 stored in Int8 as an example, the binary representation of the number 5 is 00000101. After removing the leading zeros, the remaining 101 is the compressed data of the number 5. Taking a sample data that is the number 0 stored in Int8 as an example, the binary representation of the number 0 is 00000000. The remaining 1 bit 0 is the compressed data of the number 0.
[0060] In one possible implementation, the encoder, constrained by the bit width of the payload, customizes an encoding mode for continuous sample data based on the compressed bit width of each sample data in the sample sequence. Each encoding unit in the customized encoding mode corresponds to one sample data in the continuous sample data, and each encoding unit can store the compressed data of the corresponding sample data. The customized encoding mode is the candidate encoding mode. Continuous sample data refers to multiple adjacent sample data in the sample sequence. Taking the sample sequence [4, 2, 4, 130, 12, 6, 8, 191, 0] as an example, 4, 2, 4, and 130 in this sample sequence are continuous sample data. The process will be described in detail through steps A1 to A2 below.
[0061] Step A1: Based on the bit width of the load, the encoder divides the compressed bit width of the multiple sample data into multiple bit width groups. Each bit width group includes the compressed bit width of at least one sample data. The total bit width of the bit width group is less than or equal to the bit width of the load. The total bit width is the sum of the compressed bit widths in the bit width group.
[0062] The bit width of the load is the bit width P of the load configured above.
[0063] The encoder first creates an empty bit-width group and starts scanning from the first sample data in the sample sequence, scanning each sample data in the sequence sequentially. Upon scanning each sample data, the encoder takes that sample data as the current sample data, obtains the compressed bit width of the current sample data, adds the compressed bit width of the current sample data to the bit-width group, and calculates the total bit width of the bit-width group based on the compressed bit widths in the group. If the total bit width is less than P, and the current sample data is not the last sample data in the sample sequence, the encoder continues scanning the next sample data in the sample sequence, and so on, until the total bit width of the bit-width group is greater than or equal to P, or until the last sample data has been scanned.
[0064] If the total bit width of the current bit width group is greater than or equal to P, and if the total bit width of the current bit width group equals P, and the current sample data is not the last sample data in the sample sequence, the encoder creates an empty bit width group as the next bit width group, and continues to scan the next sample data in the sample sequence. Following the method of adding compressed bit width to the current bit width group, the compressed bit width of the newly scanned sample data is added to the next bit width group. If the total bit width of the current bit width group equals P, and the current sample data is the last sample data in the sample sequence, the sample sequence scanning ends.
[0065] If the total bit width of a bit width group is greater than or equal to P, and if the total bit width of the bit width group is greater than P, the encoder creates an empty bit width group as the next bit width group. The compressed bit width of the current sample data in the current bit width group (i.e., the last compressed bit width in the current bit width group) is transferred to the next bit width group, making the total bit width of the current bit width group less than P. If the current sample data is not the last sample data in the sample sequence, the encoder continues to scan the next sample data in the sample sequence, adding the compressed bit width of the newly scanned sample data to the next bit width group in the same way as adding the compressed bit width in the current bit width group. This process continues until the last sample data in the sample sequence is scanned, at which point the last sample data is added to the bit width group. In this way, the encoder can divide the compressed bit widths of multiple sample data into multiple bit width groups, so that each bit width group consists of the compressed bit widths of consecutive compressed data in the sample sequence.
[0066] Next, taking the sample sequence [4, 2, 4, 130, 12, 6, 8, 191, 0] as an example, we will introduce the method of adding compressed bit width to the bit width group as follows.
[0067] Assuming P=32, the encoder scans each sample data in the sample sequence sequentially. After scanning sample data "4", the compressed bit width "3" of sample data "4" is added to an empty bit width group [], resulting in bit width group [3]. The total bit width of bit width group [3] is 3, which is less than 32. The encoder continues to scan the next sample data "2" after sample data "4". The compressed bit width "2" of sample data "2" is added to bit width group [3], resulting in bit width group [3, 2]. The total bit width of bit width group [3, 2] is 5, which is less than 32. The encoder continues to scan the next sample data "4" after sample data "2". The compressed bit width "3" of sample data "4" is added to bit width group [3, 2], resulting in bit width group [3, 2, 3]. The total bit width of [2, 3] is 8, which is less than 32. Continue scanning the next sample data "130" after sample data "4". Add the compressed bit width "8" of sample data "130" to the bit width group [3, 2, 3] to get the bit width group [3, 2, 3, 8]. The total bit width of the bit width group [3, 2, 3, 8] is 16, which is less than 32. Continue scanning the next sample data "12" after sample data "130". Add the compressed bit width "4" of sample data "12" to the bit width group [3, 2, 3, 8] to get the bit width group [3, 2, 3, 8, 4]. The total bit width of the bit width group [3, 2, 3, 8, 4] is 20, which is less than 32. Continue scanning the next sample data "6" after sample data "12". The compressed bit width "3" of "6" is added to the bit width group [3, 2, 3, 8, 4], resulting in the bit width group [3, 2, 3, 8, 4, 3]. The total bit width of the bit width group [3, 2, 3, 8, 4, 3] is 23, which is less than 32. Continue scanning the next sample data "8" after the sample data "6". Add the compressed bit width "4" of the sample data "8" to the bit width group [3, 2, 3, 8, 4, 3], resulting in the bit width group [3, 2, 3, 8, 4, 3, 4]. The total bit width of the bit width group [3, 2, 3, 8, 4, 3, 4] is 27, which is less than 32. Continue scanning the next sample data "191" after the sample data "8". Add the compressed bit width "8" of the sample data "191" to the bit width group [3, 2, 3, 8, 4, 3, 4]. [3, 2, 3, 8, 4, 3, 4, 8], the total bit width of the bit width group [3, 2, 3, 8, 4, 3, 4, 8] is 35, which is greater than 32. The last compressed bit width "8" in the bit width group [3, 2, 3, 8, 4, 3, 4, 8] is moved to the next empty bit width group [], resulting in the bit width group [3, 2, 3, 8, 4, 3, 4] and the next bit width group [8]. Continue scanning the next sample data "0" after the sample data "191". Add the compressed bit width "1" of the sample data "0" to the bit width group [8], resulting in the bit width group [8, 1]. The sample sequence scanning is completed, and finally the bit width group [3, 2, 3, 8, 4, 3, 4] and the bit width group [8, 1] are obtained.
[0068] The above method involves first adding compressed bit width to the bit width group, then calculating the total bit width of the bit width group, and then deciding whether to continue adding bit width to the bit width group based on whether the total bit width of the bit width group is less than P. In another possible implementation, the encoder can also calculate the sum of the total bit width of the current bit width group and the compressed bit width of the current sample data after obtaining the compressed bit width of the current sample data. If the sum is less than or equal to P, the compressed bit width is added to the current bit width group. If the sum is greater than P, the encoder creates an empty bit width group and uses the newly created bit width group as the current bit width group, and continues to add compressed bit width to the newly created bit width group.
[0069] The above example illustrates how, upon scanning a sample sequence, the encoder acquires the compressed bit width of each sample data point and adds it to the bit width group. In another possible implementation, the encoder first acquires the compressed bit width of each sample data point in the sample sequence. Following the order of the sample data in the sequence, the encoder assembles the compressed bit widths into a compressed bit width sequence. The order of the compressed bit widths in the sequence is the same as the order of the sample data in the sample sequence. For example, if a sample data point is the first sample data point in the sequence, its compressed bit width is the first compressed bit width in the compressed bit width sequence. After acquiring the compressed bit width sequence, the encoder sequentially scans each compressed bit width in the sequence. Upon acquiring a compressed bit width, it adds it to the bit width group. The method of adding compressed bit widths to the bit width group is similar to the method described above for adding compressed bit widths to the bit width group when scanning the sample sequence, and will not be repeated here.
[0070] For at least one bit width group, each bit width group corresponds to a first coding mode. For any bit width group, the encoder compresses the bit width in the bit width group to the bit width of each coding unit in the corresponding first coding mode. The sum of the compressed bit widths in the bit width group may be the same as or different from the bit width of the payload. Based on the bit width of the payload, the first coding mode is trained by adjusting the compressed bit width in the bit width group. The training process is as follows: step A2.
[0071] Step A2: For any bit width group, the encoder determines the candidate coding mode corresponding to the bit width group based on the total bit width of the bit width group and the bit width of the payload. Each coding unit in the candidate coding mode corresponds to a compressed bit width in the bit width group. The bit width of each coding unit in the candidate coding mode is greater than or equal to the corresponding compressed bit width, and the total bit width of each coding unit in the candidate coding mode is equal to the bit width of the payload.
[0072] If the total bit width of any bit width group is equal to the bit width of the payload, the encoder determines the bit width group corresponding to the candidate coding mode, so that each compressed bit width in the bit width group is used as the bit width of each coding unit under the candidate coding mode, thereby determining a candidate coding mode.
[0073] If the total bit width of any bit width group is less than the bit width of the payload, the encoder adjusts the value of the compressed bit width in the bit width group so that the total bit width of the bit width group is equal to the bit width of the payload. After adjusting the total bit width of the bit width group to the bit width of the payload, the bit width group with the adjusted total bit width is determined as the bit width group corresponding to the candidate coding mode. Each compressed bit width in the bit width group with the adjusted total bit width is used as the bit width of each coding unit under the candidate coding mode, so as to determine a candidate coding mode.
[0074] The adjustment method for the value of the compressed bit width in the bit width group is as follows: the encoder adds 1 to the minimum compressed bit width in the bit width group, so that the total bit width of the bit width group increases by 1. If the total bit width is still less than the bit width of the load after adding 1, the encoder performs the step of adding 1 to the minimum compressed bit width in the bit width group again for the bit width group after adding 1, until the total bit width of the bit width group is equal to the bit width of the load.
[0075] Still assuming the load bit width P = 32, taking the bit width group [3, 2, 3, 8, 4, 3, 4] as an example, the total bit width of the bit width group [3, 2, 3, 8, 4, 3, 4] is 27, which is less than 32. After adding 1 to the minimum compression bit width "2", we get the bit width group [3, 3, 3, 8, 4, 3, 4]. The total bit width of the bit width group [3, 3, 3, 8, 4, 3, 4] is 28, which is less than 32. Since there are 4 minimum compression bit widths of "3" in the bit width group [3, 3, 3, 8, 4, 3, 4], we select any minimum compression bit width "3" and add 1, for example, select the first minimum compression bit width "3" and add 1, to get the bit width group [4, 3, 3, 8, 4, 3, 4]. The total bit width of the bit width group [4, 3, 3, 8, 4, 3, 4] is 29. If the value is less than 32, select the first minimum compressed bit width "3" and add 1 to obtain the bit width group [4, 4, 3, 8, 4, 3, 4]. The total bit width of the bit width group [4, 4, 3, 8, 4, 3, 4] is 30, which is less than 32. Select the first minimum compressed bit width "3" and add 1 to obtain the bit width group [4, 4, 4, 8, 4, 3, 4]. The total bit width of the bit width group [4, 4, 4, 8, 4, 3, 4] is 31, which is less than 32. Select the first minimum compressed bit width "3" and add 1 to obtain the bit width group [4, 4, 4, 8, 4, 4, 4]. The total bit width of the bit width group [4, 4, 4, 8, 4, 4, 4] is equal to 32. The adjustment of the compressed bit width ends, and the bit width group [4, 4, 4, 8, 4, 4, 4] is taken as a bit width group for a candidate encoding mode. Since each compressed bit width in this bit width group will eventually be used as the bit width of a coding unit in a candidate coding mode, by continuously increasing the minimum compressed bit width in the bit width group by 1, the bit width of each coding unit in the subsequent candidate coding mode can be balanced.
[0076] In another possible implementation, if the total bit width of any bit width group is less than the bit width of the payload, the encoder obtains the difference between the total bit width and the bit width of the payload, adds the difference to the minimum compressed bit width in the bit width group, so that the total bit width of the bit width group is equal to the bit width of the payload, and then determines the bit width group of the bit width group as a candidate coding mode bit width group, so that each compressed bit width in the bit width group is used as the bit width of each coding unit under the candidate coding mode, thereby determining a candidate coding mode.
[0077] Based on step A2 above, since each coding unit in the candidate coding mode corresponds to a compressed bit width in the corresponding bit width group, and the bit width of each coding unit in the candidate coding mode is greater than or equal to the corresponding compressed bit width, the coding unit in the candidate coding mode can store the compressed data to which the corresponding compressed bit width belongs. Moreover, the total bit width of each coding unit in the candidate coding mode is equal to the bit width of the payload, so that the bit width of the payload can be allocated to each coding unit in the candidate coding mode. If the candidate coding mode is used to encode the data to be compressed in the data sequence, the remaining bit width in the payload can be avoided, thus improving the utilization rate of the payload.
[0078] In another possible implementation, if the total bit width of any bit width group is less than the bit width of the payload, the encoder does not adjust the value of the compressed bit width in that bit width group. Instead, it uses the difference between the bit width of the payload and the total bit width of the bit width group as the remaining bit width of the candidate coding mode. The encoder then uses the compressed bit widths in the bit width group as the bit width of a coding unit in the candidate coding mode to determine the candidate coding mode.
[0079] For each of the multiple bit-width groups, the encoder determines a candidate coding pattern by executing step A2, thus ultimately determining multiple candidate coding patterns, each corresponding to a bit-width group. These multiple candidate coding patterns may be the same or different. If they are the same, that candidate coding pattern is used as the second coding pattern. If multiple candidate coding patterns exist, the encoder determines the second coding pattern by executing step 4022.
[0080] Step 4022: The encoder determines the second encoding mode from the multiple candidate encoding modes.
[0081] The encoder determines a second coding mode based on the selector's bit width S and / or the frequency of occurrence of each candidate coding mode among multiple candidate coding modes, wherein the number of types of the second coding mode is less than or equal to 2. S .
[0082] For example, the encoder determines 2 from the multiple candidate coding patterns based on the bit width S of the selector. S A second coding mode. For example, the encoder randomly selects 2 from these multiple candidate coding modes. S One candidate coding pattern is used as the second coding pattern.
[0083] Alternatively, the encoder may determine the second coding mode based on the frequency of occurrence of each candidate coding mode among the multiple candidate coding modes, selecting the at least one coding mode with the highest frequency among the multiple candidate coding modes. For example, the second coding mode may be determined by selecting the K coding modes with the highest frequency among the multiple candidate coding modes, where K is less than or equal to 2. S .
[0084] Alternatively, the encoder determines the second coding mode based on the selector's bit width S and the frequency of occurrence of each candidate coding mode among the multiple candidate coding modes. For example, if multiple coding modes exist among the multiple candidate coding modes, and the number of coding modes among the multiple candidate coding modes is less than or equal to 2... S If the encoder determines all of these multiple coding modes as the second coding mode, and the number of coding modes among the multiple candidate coding modes is greater than 2, then the encoder will determine the second coding mode. S The encoder then counts the frequency of each coding mode among the multiple candidate coding modes, and selects the 2 most frequent candidate coding modes. S The first encoding mode was determined to be the second encoding mode.
[0085] By selecting less than or equal to 2 from candidate coding patterns S The second encoding mode is selected to avoid the selector corresponding to the second encoding mode being unable to store the mode identifiers of various second encoding modes. This is achieved by selecting the 2 most frequent candidate encoding modes from among multiple candidate encoding modes. S The first encoding mode is determined to be the second encoding mode, so that when the data sequence is compressed and encoded using the second encoding mode in the future, it can be more suitable for the data size distribution of the data sequence and can further improve the compression rate of the data sequence.
[0086] In another possible implementation, the encoder also determines the maximum compression bit width of the data sequence, which is the compression bit width of the largest data in the data sequence. If the bit width of the coding unit in each second coding mode is less than the maximum compression bit width, the encoder determines a third coding mode based on the maximum compression bit width. The third coding mode is a coding mode in which the sum of the bit widths of each coding unit is equal to the bit width of the payload, and the bit width of each coding unit is greater than or equal to the maximum compression bit width. This allows the third coding mode to compress and encode the largest data in the data sequence, avoiding the situation where the largest data cannot be compressed and encoded using the coding mode. Therefore, the third coding mode is also called a secure coding mode.
[0087] The method for determining the third coding mode is as follows: Let the maximum compressed bit width be denoted as *w*, and the bit width of the payload as *P*. If *P* is divisible by *w*, the encoder forms a candidate bit width group from multiple *w*, where the candidate bit width group is [w, w, ..., w], and the sum of each *w* in the candidate bit width group equals *P*. If *P* is not divisible by *w*, the encoder forms a candidate bit width group from multiple *r*, where the candidate bit width group is [r, r, ..., r], and the sum of each *r* in the candidate bit width group equals *P*. *r* is greater than *w* and less than or equal to *P*, and *r* is divisible by *P*. After determining the candidate bit width group, the encoder generates partitioning information for the third coding mode based on this group. This partitioning information includes the candidate bit width group to indicate the division of coding units according to the candidate bit width group. Each bit width in the candidate bit width group represents the bit width of each coding unit under the third coding mode, thus determining the third coding mode.
[0088] After determining the third encoding pattern, if the number of types of the second encoding pattern is less than 2... S The encoder identifies the third coding mode as a new second coding mode to supplement the existing second coding modes. If the number of second coding modes is equal to 2... S The encoder will replace the second encoding mode, which appears least frequently, with the third encoding mode.
[0089] Since the bit width of the encoding unit in the third encoding mode is greater than or equal to the maximum compressed bit width of the data sequence, the third encoding mode is determined as the second encoding mode so that the compressed data of the largest data in the data sequence can be stored through the encoding unit in the third encoding mode, thereby avoiding the situation where the largest data cannot be compressed and encoded.
[0090] Some Simple-type algorithms in related technologies have good adaptability and performance robustness, but they are computationally complex and unsuitable for scenarios with large data volumes and low computing power, such as network devices. For example, Simple-type algorithms that concatenate words, such as Relative10, Carryover12, and Slide, only work when excessively wide elements (such as compressed data) are exactly at the word edge, and they rely on serial decompression, making parallel acceleration difficult. Among them, Relative10 and Carryover12 utilize the wasted payload bits of adjacent words for storage. Slide concatenates words to avoid wasting space by having to choose an encoding mode with fewer encoding units if the element corresponding to the last encoding unit is too wide. For example, simple algorithms like vector of splits encoding (VSEncoding) and memory inverted list compression (MLIC), which adaptively optimize payload width, must first scan the entire data sequence before running. Furthermore, their adaptive process relies on dynamic programming, resulting in excessive overall computational overhead. VSEncoding, for instance, uses dynamic programming to adaptively optimize the payload width given a fixed bit width for the encoding unit. MILC accelerates the optimized prefix frame of reference (Opt-PFor) algorithm using single instruction multiple data (SIMD). The Opt-PFor algorithm is a simple algorithm that extracts the 10% of the elements with the largest bit width within each word and encodes them separately.
[0091] The process described in steps 4021 and 4022 above involves dividing the compressed bit width of multiple sample data into multiple bit width groups based on the bit width of the payload during the training of the coding mode. Then, based on the total bit width of each bit width group and the bit width of the payload, each candidate coding mode is determined, and the second coding mode is determined from the candidate coding modes. This eliminates the need for dynamic programming to determine the second coding mode, reducing the computational load. Therefore, the data compression method provided in this application has low computational requirements. Consequently, even with low computational power settings, the encoder can efficiently compress and encode large-scale sequence data and can automatically adapt and match to different data distributions to ensure compression ratio.
[0092] 403. The encoder compresses and encodes the data sequence based on the second encoding mode.
[0093] The encoder determines a second encoding scheme based on each of the second encoding modes determined in step 402 above. The second encoding scheme is a Simple encoding scheme trained based on sample data and suitable for the data sequence. For example, the encoder generates scheme information for the second encoding scheme based on each of the second encoding modes. The scheme information includes the bit width group of each second encoding mode and the mode identifier of each second encoding mode. The bit width group of each second encoding mode corresponds to its respective mode identifier, such as the scheme information shown in Table 1 below.
[0094] Table 1
[0095] In another possible implementation, if any second encoding mode still has remaining bit width, the scheme information of the second encoding scheme also includes the remaining bit width of the second encoding mode, and the remaining bit width corresponds to the mode identifier and bit width group of the second encoding mode.
[0096] After determining the second encoding scheme, the encoder compresses and encodes the data sequence based on the second encoding scheme to obtain at least one character. Next, the process will be described in detail based on the following steps 4031 to 4032.
[0097] Step 4031: The encoder obtains the compressed data of each data to be compressed in the data sequence to obtain the compressed data sequence.
[0098] The compressed data sequence includes compressed data of each data to be compressed in the data sequence, and the sorting method of the compressed data in the compressed data sequence is the same as the sorting method of the data to be compressed in the data sequence.
[0099] For example, for any data to be compressed in the data sequence, the encoder uses the binary representation of the data to be compressed (with leading zeros removed) as the compressed data, and combines the compressed data of each data to be compressed into the compressed data sequence. Taking the data sequence [2, 1, 127, 3, 1, 5, 198, 0, 2, ...] as an example, the compressed data sequence of this data sequence is [10, 1, 1111111, 11, 1, 101, 11000110, 0, 10, ...].
[0100] Step 4032: Based on each of the second encoding modes in the second encoding scheme, encapsulate the compressed data in the data compression sequence into at least one word. Each encoding unit in the payload of each word is used to store one compressed data in the data compression sequence. The selector in each word is used to store the mode identifier of the second encoding mode to which the encoding unit in the payload belongs.
[0101] In this context, the compressed data stored in the encoding unit of each word is continuous in the compressed data sequence.
[0102] The encoder starts with the first compressed data in the data compression sequence and tries the second encoding mode with the most encoding units in the second encoding scheme. It checks whether the second encoding mode with the most encoding units can store multiple consecutive compressed data starting from the first compressed data in the data compression sequence. If it can, the encoder determines the second encoding mode with the most encoding units as the fourth encoding mode. Otherwise, the encoder tries the second encoding mode with the second most encoding units to see if it can store the multiple consecutive compressed data until the fourth encoding mode is determined. The fourth encoding mode is the second encoding mode to be used, i.e., the target encoding mode. The encoding units in the fourth encoding mode can store multiple consecutive compressed data in the data sequence, and the fourth encoding mode is suitable for the multiple consecutive compressed data.
[0103] When attempting to determine whether any second encoding mode can store multiple consecutive compressed data, the encoder can first obtain the bit width of multiple consecutive compressed data, and then compare the bit width of the multiple consecutive compressed data with the bit width of the encoding unit in the second encoding mode to determine whether the encoding unit in the second encoding mode can store multiple consecutive compressed data. If the encoding unit in the second encoding mode can store multiple consecutive compressed data, then the second encoding mode is determined as the fourth encoding mode; otherwise, an attempt is made to determine whether another second encoding mode can store multiple consecutive compressed data.
[0104] Taking the compressed sequence [10, 1, 1111111, 11, 1, 101, 11000110, 0, 10, ...] as an example, assuming the second encoding scheme is as shown in Table 1, for ease of description, each of the second encoding modes in Table 1 will be simply referred to as an encoding mode. Starting from the first compressed data "10" in the compressed sequence, we will begin trying the encoding mode with the most encoding units in the second encoding scheme. Since there are 3 encoding modes with the most encoding units in the second encoding scheme shown in Table 1, the encoder can start trying from any encoding mode with the most encoding units, for example, starting from the first encoding mode with the most encoding units, "00": the bit width of the compressed data "10" is less than the bit width of the first encoding mode "00". A bit width of "4" in a coding unit indicates that the first coding unit can store compressed data "10". The bit width of the next compressed data "1" after compressed data "10" is less than the bit width of the second coding unit "4" under coding mode "00", indicating that the second coding unit can store compressed data "1". The bit width of the next compressed data "1111111" after compressed data "1" is greater than the bit width of the third coding unit "4" under coding mode "00", indicating that the third coding unit cannot store compressed data "1111111". Coding units under coding mode "00" cannot store consecutive compressed data starting from the first compressed data, so the attempt of coding mode "00" fails.
[0105] The encoder continues to attempt the second-most encoding mode "01" in the second encoding scheme: the bit width of compressed data "10" is equal to the bit width "2" of the first encoding unit under encoding mode "01", indicating that the first encoding unit can store compressed data "10"; the bit width of the next compressed data "1" after compressed data "01" is less than the bit width "3" of the second encoding unit under encoding mode "01", indicating that the second encoding unit can store compressed data "1"; the bit width of the next compressed data "1111111" after compressed data "1" is less than the bit width "8" of the third encoding unit under encoding mode "01", indicating that the third encoding unit can store compressed data "1111111"; the bit width of the next compressed data "11" after compressed data "1111111" is less than the bit width "4" of the fourth encoding unit under encoding mode "01", indicating that the fourth encoding unit... The encoder can store compressed data "11". The bit width of the next compressed data "1" after compressed data "11" is less than the bit width "3" of the fifth coding unit under coding mode "01", which means that the fifth coding unit can store compressed data "1". The bit width of the next compressed data "101" after compressed data "1" is less than the bit width "4" of the sixth coding unit under coding mode "01", which means that the sixth coding unit can store compressed data "101". The bit width of the next compressed data "11000110" after compressed data "101" is equal to the bit width "8" of the seventh coding unit under coding mode "01", which means that the seventh coding unit can store compressed data "11000110". Therefore, it can be seen that the 7 coding units under coding mode "01" can store the first 7 consecutive compressed data in the data sequence in sequence. Thus, the encoder determines coding mode "01" as the fourth coding mode.
[0106] After determining the fourth encoding mode, the encoder encodes multiple consecutive compressed data in the data sequence that are applicable to the fourth encoding mode, based on the fourth encoding mode, to obtain a word. For example, the encoder sequentially encodes the multiple consecutive compressed data into each encoding unit of the fourth encoding mode, with each encoding unit used to store one compressed data. The mode identifier of the first encoding mode is encoded into the selector, and the selector and the encoding unit storing the compressed data in the fourth encoding mode are combined to form a word. If there is remaining bit width in the payload of the first encoding mode, the encoder combines the selector, the encoding unit storing the compressed data in the fourth encoding mode, and the number of 0s of the remaining bit width into a word.
[0107] Taking the encoding pattern "01" in Table 1 as the fourth encoding pattern, and using the first 7 compressed data points of the above compressed sequence as an example, for instance... Figure 5The diagram shows the structure of the character. The encoder encodes the mode identifier "01" of encoding mode "01" into the selector, and sequentially encodes the seven compressed data characters "10", "1", "1111111", "11", "101", and "11000110" into encoding units 1 to 7 under encoding mode "01". For example, Figure 5 As shown, when encoding compressed data in the encoding unit, the encoder can first encode the compressed data from the low-order bits of the encoding unit. After the compressed data is encoded, if there are any remaining high-order bits in the encoding unit, the remaining high-order bits can be padded with 0. In some other embodiments, the encoder can also first encode the compressed data from the high-order bits of the encoding unit. After the compressed data is encoded, if there are any remaining low-order bits in the encoding unit, the remaining low-order bits can be padded with 0.
[0108] For unencoded compressed data in the data sequence, the encoder starts with the first unencoded compressed data and tries the second encoding scheme with the most encoding units. It checks if the second encoding scheme with the most encoding units can store multiple consecutive compressed data starting from the first compressed data. If it can, the encoder determines this second encoding scheme with the most encoding units as the new fourth encoding scheme. Otherwise, the encoder tries the second encoding scheme with the second most encoding units to see if it can store the multiple consecutive compressed data, until a new fourth encoding scheme is determined. Based on the new fourth encoding scheme, the multiple consecutive compressed data are encoded to obtain a new word. This process continues until the last compressed data in the data sequence is encoded, thus compressing the data sequence into at least one word.
[0109] After compressing the data sequence into at least one word, the encoder stores the payload of each word in the payload storage space and the selector of each word in the selector storage space. This allows for subsequent reconstruction of the at least one word based on the stored payload and selector, and then sends the at least one word to the decoder for decoding. Alternatively, after compressing the data sequence into at least one word, instead of storing the payload of each word in the payload storage space and the selector of each word in the selector storage space, the encoder sends the at least one word to the decoder. The decoder receives the at least one word and reconstructs the data in the data sequence based on it. For example, each time the decoder receives a word, it determines the second encoding mode used by the word based on the mode identifier stored in the word selector. Based on the partitioning method indicated by the second encoding mode, it determines each encoding unit in the payload of the word, obtains compressed data from each encoding unit, converts the compressed data into uncompressed data, and assembles the reconstructed uncompressed data into the data sequence.
[0110] The decoder can be one of the above. Figure 1 In the decoder 102, before the decoder decodes the word, the decoder also obtains the scheme information of the second encoding scheme from the encoder so as to decode the mode identifier stored in the word selector, query the second encoding mode indicated by the mode identifier and the bit width group of the second encoding mode from the scheme information, and use the bit width group as the division method of the second encoding mode to determine each encoding unit in the payload of the word.
[0111] based on Figure 4 The method described herein obtains sample data based on the data sequence to be compressed, ensuring that the sample data reflects the actual size distribution characteristics of the data sequence. Then, based on the sample data, a Simple encoding scheme encoding mode is trained, so that the encoding units in the trained encoding mode can satisfy the actual size distribution of the data sequence. This makes the trained encoding mode applicable to the data sequence, and the data sequence is compressed using the trained encoding mode, ensuring a high compression ratio. For each data sequence to be compressed, the encoder, according to the data compression method provided in this application, can customize an applicable encoding mode for each data sequence. Even if different data sequences have different data distributions, compression ratios can be guaranteed for each data sequence based on the customized encoding modes, avoiding degradation in compression ratio when dealing with different data sequences with different size distributions. By training the Simple encoding scheme encoding mode using sampled data, the Simple encoding scheme can be updated automatically without manual intervention. During the training of the encoding mode, based on the bit width of the payload, the compression bit width of multiple sample data is divided into multiple bit width groups. Then, based on the total bit width of each bit width group and the bit width of the payload, each candidate encoding mode is determined. The second encoding mode is determined from the candidate encoding modes. There is no need to use dynamic programming to determine the second encoding mode, which reduces the amount of computation. Therefore, the data compression method provided in this application has low computational requirements. Thus, even with low computational power settings, the encoder can efficiently compress and encode large-scale sequence data, and can automatically adapt and match to different data distributions to ensure compression ratio.
[0112] The above example uses integer data as the data to be compressed in the data sequence. In another possible implementation, the data to be compressed in the data sequence is not integer data, but non-integer data. In this case, the encoder can first convert each data to be compressed in the data to be compressed into integer data, and then perform the above steps 401 to 403 on the converted data sequence.
[0113] In another possible implementation, the encoder supports at least one of the following encoding functions: automatic encoding and fixed encoding. Automatic encoding refers to using the data compression method provided in this application to compress and encode the data sequence to be compressed. Fixed encoding refers to the encoder using an encoding scheme to compress and encode the data sequence to be compressed. When the user enables the encoder's automatic encoding function, the encoder compresses and encodes each received data sequence using the data compression method provided in this application. When the user enables the encoder's fixed encoding function, the encoder also provides a configuration interface for the fixed encoding scheme. This configuration interface displays scheme information for the Simple encoding scheme to be used. The user can modify the bit width of the encoding unit in each encoding mode in the scheme information in this configuration interface to update the configuration of the Simple encoding scheme. Each time the encoder receives a data sequence to be compressed, it compresses and encodes the data sequence using the Simple encoding scheme indicated in the scheme information in the configuration interface.
[0114] In another possible implementation, after the second encoding scheme is determined, the encoder does not perform the step of compressing and encoding the data sequence based on the second encoding scheme to obtain at least one word. Instead, another device provides the second encoding scheme, and the other device compresses and encodes the data sequence based on the second encoding scheme provided by the encoder to obtain at least one word. In this implementation, the encoder is responsible for customizing the encoding scheme for the data sequence to be compressed, but not for compressing and encoding the data sequence. In this case, the encoder may not be called an encoder, but a Simple encoding scheme determination device. The method flow consisting of steps 401, 402, and the step of determining the second encoding scheme based on step 402 is called the flow of the Simple encoding scheme determination method.
[0115] The methods of the embodiments of this application have been described above; the apparatus of the embodiments of this application will be described below. It should be understood that the apparatus described below has any of the encoder functions described in the above methods. (The above is in conjunction with...) Figures 1 to 5 The transmission compression method according to embodiments of this application is described in detail. Based on the same inventive concept, the following will be combined with... Figures 6 to 7 The apparatus according to embodiments of this application is described. It should be understood that the technical features described in the method embodiments are also applicable to the following apparatus embodiments.
[0116] Figure 6 This application provides a schematic diagram of the structure of a data compression device. Figure 6 The device 600 shown can be an encoder or a portion of an encoder as described in the preceding embodiments, used to execute the data compression method performed by the encoder. For example... Figure 6 As shown, the device 600 includes: Acquisition module 601 is used to acquire multiple sample data based on a data sequence; Training module 602 is used to train a first encoding mode based on the plurality of sample data to obtain a second encoding mode. The first encoding mode is used to indicate the division of encoding units in the payload encoded based on the Simple encoding scheme. The encoding module 603 is used to compress and encode the data sequence based on the second encoding mode.
[0117] In one possible implementation, training module 602 includes: The training unit is used to train the first coding mode based on the compressed bit width of the plurality of sample data and the bit width of the payload to obtain a plurality of candidate coding modes, each of the candidate coding modes corresponding to at least one sample data, and the compressed bit width is the bit width of the compressed data of the sample data; A determining unit is configured to determine the second encoding pattern from the plurality of candidate encoding patterns.
[0118] In one possible implementation, the training unit is used for: Based on the bit width of the payload, the compressed bit width of the plurality of sample data is divided into a plurality of bit width groups. Each bit width group includes the compressed bit width of at least one sample data. The total bit width of the bit width group is less than or equal to the bit width of the payload. The total bit width is the sum of the compressed bit widths in the bit width group. For any bit width group, a candidate coding mode corresponding to the bit width group is determined based on the total bit width of the bit width group and the bit width of the payload. Each coding unit in the candidate coding mode corresponds to a compressed bit width in the bit width group. The bit width of each coding unit in the candidate coding mode is greater than or equal to the corresponding compressed bit width, and the total bit width of each coding unit in the candidate coding mode is equal to the bit width of the payload.
[0119] In one possible implementation, the determining unit is used for: The most frequently occurring encoding pattern among the multiple candidate encoding patterns is determined as the second encoding pattern.
[0120] In one possible implementation, device 600 further includes: The determining module is configured to determine a third encoding mode based on the maximum compressed bit width if the bit width of the encoding unit in each of the second encoding modes is less than the maximum compressed bit width of the data sequence, wherein the maximum compressed bit width is the compressed bit width of the largest data in the data sequence, and the bit width of each encoding unit in the third encoding mode is greater than or equal to the maximum compressed bit width. The replacement module is used to replace the second encoding pattern, which appears least frequently, with the third encoding pattern.
[0121] It should be understood that device 600 corresponds to the encoder in the above method embodiments. The modules in device 600 and the other operations and / or functions described above are respectively for implementing various steps and methods of the encoder in the method embodiments. For specific details, please refer to the above method embodiments. For the sake of brevity, they will not be repeated here.
[0122] It should be understood that when device 600 compresses and encodes data sequences, the division of the above-described functional modules is only used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of device 600 can be divided into different functional modules to complete all or part of the functions described above. In addition, the device 600 provided in the above embodiments and the above method embodiments belong to the same concept, and its specific implementation process is detailed in the above method embodiments, which will not be repeated here.
[0123] Figure 7 This is a schematic diagram of the structure of a computing device provided in an embodiment of this application, such as... Figure 7 As shown, the computing device 700 includes a bus 702, a processor 704, a memory 706, and a communication interface 708. The processor 704, the memory 706, and the communication interface 708 communicate with each other via the bus 702. The computing device 700 can be an electronic device such as a server or a terminal device. It should be understood that this application does not limit the number of processors and memories in the computing device 700.
[0124] The 702 bus can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, Figure 7 The bus 702 may be represented by a single line, but this does not mean that there is only one bus or one type of bus. The bus 702 may include a path for transmitting information between various components of the computing device 700 (e.g., memory 706, processor 704, communication interface 708).
[0125] Processor 704 may include any one or more processors such as CPU, graphics processing unit (GPU), microprocessor (MP), or digital signal processor (DSP).
[0126] The memory 706 may include volatile memory, such as random access memory (RAM). The processor 704 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).
[0127] The memory 706 stores executable program code, and the processor 704 executes the executable program code to implement the functions of the aforementioned acquisition module 601, training module 602, and encoding module 603, thereby implementing the data compression method provided in this application. That is, the memory 706 stores instructions for executing the data compression method.
[0128] The communication interface 708 uses transceiver modules, such as, but not limited to, network interface cards and transceivers, to enable communication between the computing device 700 and other devices or communication networks.
[0129] In another possible implementation, the computing device 700 is a chip, which is implemented by any combination of ASIC, PLD, CPLD, FPGA and GAL.
[0130] In an exemplary embodiment, a computer-readable storage medium is also provided, such as a memory including program code that can be executed by a processor in a computing device to perform the data compression method described above. For example, the computer-readable storage medium is a non-transitory computer-readable storage medium, such as read-only memory (ROM), random access memory (RAM), compact disc read-only memory (CD-ROM), magnetic tape, floppy disk, and optical data storage devices.
[0131] This application also provides a computer program product or computer program, which includes program code. The computer instructions are stored in a computer-readable storage medium. The processor of the computing device reads the program code from the computer-readable storage medium and executes the program code, causing the computing device to perform the above-described data compression method.
[0132] In addition, embodiments of this application also provide an apparatus, which may specifically be a chip, component or module. The apparatus may include a connected processor and a memory; wherein the memory is used to store computer execution instructions, and when the apparatus is running, the processor may execute the computer execution instructions stored in the memory to cause the chip to execute the data compression methods in the above-described method embodiments.
[0133] In this embodiment, the apparatus, device, computer-readable storage medium, computer program product or chip are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding methods provided above, and will not be repeated here.
[0134] Through the above description of the embodiments, those skilled in the art will understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. Furthermore, the data compression method embodiments provided above belong to the same concept, and their specific implementation processes are detailed in the method embodiments, and will not be repeated here.
[0135] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another device, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0136] The units described as separate components may or may not be physically separate. A component shown as a unit can be one or more physical units; that is, it can be located in one place or distributed in multiple different locations. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0137] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0138] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solutions of the embodiments of this application, essentially, or the parts that contribute to the prior art, or all or part of the technical solutions, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, ROM, RAM, magnetic disks, or optical disks.
[0139] In the description of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. The "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" means one or more, and "multiple" means two or more. The terms "first," "second," etc., do not limit the quantity or order of execution, and "first," "second," etc., do not necessarily imply differences.
[0140] In this application, the terms "exemplary" or "for example" are used to indicate that something is an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or designs. Specifically, the use of the terms "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.
[0141] It should be noted that all information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the data sequences involved in this application were all obtained under full authorization.
[0142] All of the above-mentioned optional technical solutions can be combined in any way to form optional embodiments of this disclosure, and will not be described in detail here.
[0143] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A data compression method, characterized in that, The method includes: A sample sequence is obtained based on a data sequence, wherein the sample sequence includes multiple sample data. Using the bit width of the payload encoded based on the Simple coding scheme as a constraint, and based on the compressed bit width of each sample data in the sample sequence, candidate coding modes are customized for multiple adjacent sample data in the sample sequence to obtain multiple candidate coding modes. The compressed bit width is the bit width of the compressed data of the sample data, and the candidate coding mode is the first coding mode after training. The first coding mode is used to indicate the division method of the coding unit in the payload. From the plurality of candidate encoding patterns, a second encoding pattern is determined; The data sequence is compressed and encoded based on the second encoding mode.
2. The method according to claim 1, characterized in that, The method uses the bit width of the payload encoded based on the Simple coding scheme as a constraint, and based on the compressed bit width of each sample data in the sample sequence, customizes candidate coding modes for multiple adjacent sample data in the sample sequence, resulting in multiple candidate coding modes including: Based on the bit width of the payload, the compressed bit width of multiple sample data in the sample sequence is divided into multiple bit width groups. Each bit width group includes the compressed bit width of multiple adjacent sample data in the sample sequence. The total bit width of the bit width group is less than or equal to the bit width of the payload. The total bit width is the sum of the compressed bit widths in the bit width group. For any bit width group, a candidate coding mode corresponding to the bit width group is determined based on the total bit width of the bit width group and the bit width of the payload. Each coding unit in the candidate coding mode corresponds to a compressed bit width in the bit width group. The bit width of each coding unit in the candidate coding mode is greater than or equal to the corresponding compressed bit width, and the total bit width of each coding unit in the candidate coding mode is equal to the bit width of the payload.
3. The method according to claim 1 or 2, characterized in that, Determining the second encoding pattern from the plurality of candidate encoding patterns includes: The most frequently occurring encoding pattern among the multiple candidate encoding patterns is determined as the second encoding pattern.
4. The method according to claim 3, characterized in that, The method further includes: If the bit width of the coding unit in each of the second coding modes is less than the maximum compressed bit width of the data sequence, a third coding mode is determined based on the maximum compressed bit width, wherein the maximum compressed bit width is the compressed bit width of the largest data in the data sequence, and the bit width of each coding unit in the third coding mode is greater than or equal to the maximum compressed bit width. The second encoding pattern, which appears least frequently, is replaced with the third encoding pattern.
5. A data compression device, characterized in that, The device includes an acquisition module, a training module, and an encoding module, wherein the training module includes a training unit and a determination unit. The acquisition module is used to acquire a sample sequence based on a data sequence, wherein the sample sequence includes multiple sample data. The training unit is used to customize candidate coding patterns for multiple adjacent sample data in the sample sequence based on the compressed bit width of each sample data in the sample sequence, with the bit width of the payload encoded based on the Simple coding scheme as a constraint, thereby obtaining multiple candidate coding patterns. The compressed bit width is the bit width of the compressed data of the sample data, and the candidate coding pattern is the first coding pattern after training. The first coding pattern is used to indicate the division method of the coding units in the payload. The determining unit is configured to determine a second coding pattern from the plurality of candidate coding patterns; The encoding module is used to compress and encode the data sequence based on the second encoding mode.
6. The apparatus according to claim 5, characterized in that, The training unit is used for: Based on the bit width of the payload, the compressed bit width of multiple sample data in the sample sequence is divided into multiple bit width groups. Each bit width group includes the compressed bit width of multiple adjacent sample data in the sample sequence. The total bit width of the bit width group is less than or equal to the bit width of the payload. The total bit width is the sum of the compressed bit widths in the bit width group. For any bit width group, a candidate coding mode corresponding to the bit width group is determined based on the total bit width of the bit width group and the bit width of the payload. Each coding unit in the candidate coding mode corresponds to a compressed bit width in the bit width group. The bit width of each coding unit in the candidate coding mode is greater than or equal to the corresponding compressed bit width, and the total bit width of each coding unit in the candidate coding mode is equal to the bit width of the payload.
7. The apparatus according to claim 5 or 6, characterized in that, The determining unit is used for: The most frequently occurring encoding pattern among the multiple candidate encoding patterns is determined as the second encoding pattern.
8. The apparatus according to claim 7, characterized in that, The device further includes: The determining module is configured to determine a third encoding mode based on the maximum compressed bit width if the bit width of the encoding unit in each of the second encoding modes is less than the maximum compressed bit width of the data sequence, wherein the maximum compressed bit width is the compressed bit width of the largest data in the data sequence, and the bit width of each encoding unit in the third encoding mode is greater than or equal to the maximum compressed bit width. The replacement module is used to replace the second encoding pattern, which appears least frequently, with the third encoding pattern.
9. A computing device, characterized in that, The computing device includes a processor for executing program code that causes the computing device to perform the method as described in any one of claims 1 to 4.
10. A computer-readable storage medium, characterized in that, The storage medium stores at least one piece of program code, which is read by a processor to cause a computing device to perform the method as described in any one of claims 1 to 4.
Citation Information
Patent Citations
Video decoding method and apparatus, video coding method and apparatus, and storage medium
CN110933409A
Data compression method and computing equipment
CN113055017A