Floating point number compression method and decompression method, and computing device
By entropy encoding and reserved short encoding processing of the floating-point exponential part in the large language model, the problem of insufficient communication bandwidth in multi-device training is solved, and efficient parameter transmission and training is achieved.
Patent Information
- Application Number
- PCT/CN2024/142217
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-27
- Filing Date
- 2024-12-25
- Publication Date
- 2025-07-03
AI Technical Summary
During the training of large language models, the communication bandwidth between multiple devices is insufficient, making it difficult to meet the transmission requirements of parameter quantities, resulting in low training efficiency.
Entropy encoding is used to compress the exponential part of the floating point number losslessly. By differentiating the exponential and predicted value of the floating point number into a normal distribution, and a short encoding is reserved to process zero-value floating point numbers, lossless compression is achieved and communication bandwidth occupation is reduced.
It realizes efficient transmission of large language model parameters under limited communication bandwidth, reducing bandwidth usage during communication without additional computing resources.
Smart Images

Figure CN2024142217_03072025_PF_FP_ABST
Abstract
Description
Floating point number compression method, decompression method and computing device
[0001] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of China on December 27, 2023, with application number 202311826087.3 and application name “A compression method, decompression method and computing device for floating-point numbers”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The present invention relates to the field of artificial intelligence technology, and in particular to a floating-point number compression method, a decompression method, and a computing device. Background Art
[0003] With the development of large language models such as GPT and Chat AI, the number of parameters in these models has increased, reaching hundreds of billions. Due to the large number of parameters in large language models, training them requires the participation of multiple devices. When training parameters distributed across multiple devices, these devices must communicate with each other to transfer parameters from one device to another. Typically, the bandwidth required for inter-device communication is very limited, making it difficult to meet the communication requirements of multiple devices training large language models. Therefore, given the limited bandwidth required for inter-device communication, how to enable multiple devices to train large language models is a pressing issue. Summary of the Invention
[0004] To address the aforementioned issues, an embodiment of the present application provides a floating-point number compression method that performs entropy coding on the exponential portion of a floating-point number to achieve lossless compression, and reserves a short code as the result of the lossless compression of the floating-point number with a value of zero, thereby achieving lossless compression of the transmitted data and reducing the communication bandwidth occupied during transmission. Furthermore, the present application also provides a floating-point number compression device corresponding to the floating-point number compression method, a floating-point number decompression method, a floating-point number decompression device corresponding to the floating-point number decompression method, and a computing device.
[0005] To this end, the following technical solutions are adopted in the embodiments of the present application:
[0006] In a first aspect, an embodiment of the present application provides a method for compressing floating-point numbers, which is executed by a computing device and includes: obtaining a target floating-point number; the target floating-point number includes an original sign part, an original exponent part, and an original mantissa part; subtracting the original exponent part from a predicted value to obtain an exponent difference of the target floating-point number; the predicted value is the average value of a set number of floating-point numbers adjacent to the target floating-point number in a floating-point number queue, and the floating-point number queue is a queue that sorts floating-point numbers in chronological order; entropy encoding the exponent difference of the target floating-point number to obtain an entropy-encoded exponent part; the original sign part, the original mantissa part, the entropy-encoded exponent part, and the predicted value constitute a data message.
[0007] In this embodiment, after receiving the target floating-point number, the computing device subtracts the exponent of the target floating-point number from the predicted value, thereby converting the target floating-point number from a skewed distribution to a normal distribution, so that the exponent of the target floating-point number is losslessly compressed using entropy coding, thereby enabling the computing device to losslessly compress the sent data, reduce the communication bandwidth occupied during the transmission process, and do not require additional computing resources of other modules.
[0008] In one embodiment, before obtaining the target floating-point number, the method further includes: utilizing a data sliding window to slide on the floating-point number queue to obtain a plurality of floating-point numbers; and averaging the values of the plurality of floating-point numbers to obtain a predicted value of the current data sliding window.
[0009] In this embodiment, the computing device can call a data sliding window to slide on the floating-point number queue, and calculate the average value of the exponents of several floating-point numbers within the data sliding window as a temporary prediction value for the entire batch of floating-point numbers. The computing device does not need to calculate the prediction value for the entire batch of floating-point numbers, thereby reducing the computing workload of the computing device.
[0010] In one embodiment, before averaging the values of the multiple floating-point numbers to obtain the predicted value of the current data sliding window, the method further includes: detecting whether there is a floating-point number with a value of zero among the multiple floating-point numbers; if there is a floating-point number with a value of zero among the multiple floating-point numbers, replacing the value of the floating-point number with a value of zero with the predicted value of the data sliding window of the previous slide.
[0011] In this embodiment, since the exponential part of a floating-point number with a value of zero is greatly different from the exponential part of other floating-point numbers, the computing device can use the predicted value of the data sliding window of the previous jump as the value of the floating-point number with a value of zero, which can avoid the influence of the floating-point number with a value of zero on the normal prediction value, resulting in a particularly large difference error in the exponent of the calculated floating-point number.
[0012] In one embodiment, the floating-point number of the current data sliding window includes the first floating-point number in the floating-point number queue. Before averaging the values of the multiple floating-point numbers to obtain the predicted value of the current data sliding window, the method also includes: detecting whether there is a floating-point number with a value of zero among the multiple floating-point numbers; if there is a floating-point number with a value of zero among the multiple floating-point numbers, eliminating the floating-point number with a value of zero.
[0013] In this embodiment, if the data sliding window contains the first floating-point number in the floating-point number queue, it indicates that the data sliding window is the first to slide, and the floating-point number with a value of zero cannot be replaced. Therefore, when calculating the predicted value, the computing device can eliminate the floating-point number with a value of zero in the data sliding window to avoid the influence of the floating-point number with a value of zero on the normal predicted value, which may cause a particularly large difference error in the exponent of the calculated floating-point number.
[0014] In one embodiment, before subtracting the original exponential part from the predicted value to obtain the exponential difference of the target floating-point number, the method further includes: detecting whether the value of the target floating-point number is zero; if the value of the target floating-point number is zero, converting the target floating-point number into a specific code; the specific code instructs other computing devices to generate a floating-point number with a value of zero.
[0015] In this embodiment, since the sign part, mantissa part and exponent part of the zero-valued floating-point number are all fixed values, the zero-valued floating-point number can be converted into a special code, allowing the decompressor to directly recognize the special code to output the zero-valued floating-point number, thereby further reducing the communication bandwidth for transmitting the zero-valued floating-point number and reducing the computational complexity of processing the zero-valued floating-point number.
[0016] In a second aspect, an embodiment of the present application provides a floating-point number decompression method, characterized in that the method is executed by a computing device, and the method includes: receiving a data message; the data message includes a sign part, a mantissa part, an exponential part after entropy coding, and a predicted value, the predicted value being the average value of a set number of floating-point numbers adjacent to a floating-point number in a floating-point number queue, the floating-point number queue being a queue that sorts floating-point numbers in chronological order; decoding the exponential part after entropy coding to obtain an exponential difference of the floating-point number;
[0017] The exponent difference of the floating point number is added to the predicted value to obtain the original exponent part; and a floating point number is constructed according to the sign part, the mantissa part and the original exponent part.
[0018] In this embodiment, after receiving a data message, the computing device parses the data message to obtain the sign portion, mantissa portion, entropy-coded exponent portion, and predicted value. The computing device can decode the entropy-coded exponent portion, add the decoded difference value to the predicted value, and restore the original exponent portion, thereby restoring the original target floating-point number. This allows the computing device to reduce the communication bandwidth occupied during transmission without requiring additional computing resources from other modules.
[0019] In one embodiment, the data message includes a specific code, and the method further includes: generating a floating point number with a value of zero according to the specific code.
[0020] In this embodiment, since the sign, mantissa, and exponent of a zero-valued floating-point number are all fixed values, the zero-valued floating-point number can be converted into a special code. After receiving the special code, the computing device can directly recognize the special code and output the zero-valued floating-point number, thereby reducing the communication bandwidth occupied by the computing device during transmission and reducing the computational effort required to process the zero-valued floating-point number.
[0021] In a third aspect, an embodiment of the present application provides a floating-point number compression device, characterized in that it includes: a first processing unit for obtaining a target floating-point number; the target floating-point number includes an original sign part, an original exponent part and an original mantissa part; a second processing unit for subtracting the original exponent part from a predicted value to obtain an exponent difference of the target floating-point number; the predicted value is the average value of a set number of floating-point numbers adjacent to the target floating-point number in a floating-point number queue, and the floating-point number queue is a queue that sorts floating-point numbers in chronological order; a third processing unit for entropy encoding the exponent difference of the target floating-point number to obtain the exponent part after entropy encoding; a fourth processing unit for forming a data message with the original sign part, the original mantissa part, the exponent part after entropy encoding and the predicted value.
[0022] In one embodiment, the first processing unit is further configured to utilize a data sliding window to slide on the floating-point number queue to obtain a plurality of floating-point numbers; and average the values of the plurality of floating-point numbers to obtain a predicted value of the current data sliding window.
[0023] In one embodiment, the first processing unit is further used to detect whether there is a floating-point number with a value of zero among the multiple floating-point numbers; if there is a floating-point number with a value of zero among the multiple floating-point numbers, the value of the floating-point number with a value of zero is replaced with the predicted value of the data sliding window of the last sliding.
[0024] In one embodiment, the first processing unit is further used to detect whether there is a floating-point number with a value of zero among the multiple floating-point numbers when the floating-point number of the current data sliding window includes the first floating-point number in the floating-point number queue; and if there is a floating-point number with a value of zero among the multiple floating-point numbers, eliminate the floating-point number with a value of zero.
[0025] In one embodiment, the second processing unit is further used to detect whether the value of the target floating-point number is zero; if the value of the target floating-point number is zero, convert the target floating-point number into a specific code; the specific code instructs other computing devices to generate a floating-point number with a value of zero.
[0026] In a fourth aspect, an embodiment of the present application provides a floating-point number decompression device, comprising: a first processing unit for receiving a data message; the data message comprises a sign part, a mantissa part, an exponential part after entropy coding, and a predicted value, the predicted value being the average value of a set number of floating-point numbers adjacent to a floating-point number in a floating-point number queue, the floating-point number queue being a queue that sorts floating-point numbers in chronological order; a second processing unit for decoding the exponential part after entropy coding to obtain an exponential difference of the floating-point number; a third processing unit for adding the exponential difference of the floating-point number to the predicted value to obtain the original exponential part; and a fourth processing unit for constructing a floating-point number based on the sign part, the mantissa part, and the original exponential part.
[0027] In one embodiment, the fourth processing unit is further configured to generate a floating point number with a value of zero according to the specific code when the data message includes the specific code.
[0028] In a fifth aspect, an embodiment of the present application provides a computing device, comprising: at least one memory; and at least one processor, the processor being configured to execute instructions stored in the memory, so that the computing device executes each possible implementation embodiment of the first aspect, and / or executes each possible implementation embodiment of the second aspect.
[0029] In a sixth aspect, an embodiment of the present application provides a computer-readable storage medium comprising computer program instructions. When the computer program instructions are executed by a network device, the computing device executes each possible implementation embodiment of the first aspect and / or executes each possible implementation embodiment of the second aspect.
[0030] In the seventh aspect, an embodiment of the present application provides a computer program product comprising instructions, characterized in that the computer program product stores instructions, and when the instructions are executed by a network device, the network device implements the various possible implementation embodiments of the first aspect, and / or executes the various possible implementation embodiments of the second aspect. BRIEF DESCRIPTION OF THE DRAWINGS
[0031] The following is a brief introduction to the drawings required for describing the embodiments or prior art.
[0032] FIG1 is a schematic diagram of bits when storing a floating-point number;
[0033] FIG2 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application;
[0034] FIG3 is a schematic diagram of a processing process of a compressor provided in an embodiment of the present application;
[0035] FIG4 is a schematic diagram of a processing process of a decompressor provided in an embodiment of the present application;
[0036] FIG5 is a schematic flow chart of a floating-point number compression method provided in an embodiment of the present application;
[0037] FIG6 is a flow chart of a method for decompressing a floating-point number provided in an embodiment of the present application;
[0038] FIG7 is a schematic structural diagram of a floating-point number compression device provided in an embodiment of the present application;
[0039] FIG8 is a schematic structural diagram of a floating-point number decompression device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0040] The technical solutions in the embodiments of the present application will be described below in conjunction with the drawings in the embodiments of the present application.
[0041] The term "and / or" as used herein describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. The symbol " / " as used herein indicates that the related objects are in an "or" relationship, for example, A / B means either A or B.
[0042] The terms "first" and "second" in this specification and claims are used to distinguish different objects rather than to describe a specific order of objects. For example, "first response message" and "second response message" are used to distinguish different response messages rather than to describe a specific order of response messages.
[0043] In the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or descriptions. Any embodiment or design described as "exemplary" or "for example" in the embodiments of this application should not be interpreted as being preferred or advantageous over other embodiments or designs. Rather, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0044] In the description of the embodiments of the present application, unless otherwise specified, "multiple" means two or more, for example, multiple processing units means two or more processing units, etc.; multiple elements means two or more elements, etc.
[0045] In order to solve the problem of insufficient communication bandwidth for data transmission between multiple devices, related technologies can perform offline lossless compression on the sent data and online decompression on the received data when multiple devices are transmitting data. The data compression process is specifically as follows:
[0046] Take the neural network processor as an example. The neural network processor uses run-length coding to encode continuous zero values into some special values, and then uses Golomb-Rice coding (GRC) to further compress the index value and run-length value. Given a parameter M, GRC performs fixed-length binary coding on the last M bits of the original data and unary coding on the remaining bits. It is suitable for scenarios where the frequency of smaller values is much higher than the frequency of larger values. Because GRC contains a variable-length unary code, the neural network processor needs to serially count the number of consecutive bits 1 during the decoding process, which makes it difficult for the neural network processor to parallelize the decoding process and has poor scalability. It is not possible to improve the decoding efficiency by doubling the hardware. Moreover, it is difficult to achieve hardware-in-line compression during the compression process of the neural network processor, and software statistical parameter distribution is required for offline compression.
[0047] The parameters of large language models are usually represented and stored using floating-point numbers. Floating-point numbers are a data type used to represent real numbers. According to the definition of the Institute of Electronics Engineers (IEEE), floating-point numbers are composed of a sign, an exponent, and a fraction. The sign part indicates whether the floating-point number is positive or negative. The exponent part is used to indicate the order of magnitude (i.e., power) of the floating-point number, usually in the form of an optional counting method. The fraction part is also called a significant digit, which represents the main part of the floating-point number and is used to store the actual value. Compared with integers, floating-point numbers can better represent non-integer, extremely large or extremely small numbers.
[0048] When the memory stores floating-point numbers in single-precision mode, a floating-point number occupies 32 bits. The 32nd bit represents the sign, with "0" representing a positive value and "1" representing a negative value. Bits 31-24 represent the exponent. Bits 1-23 represent the mantissa. Taking the real number "5.0" as an example, the memory converts "5.0" into a binary number, "101.0". The memory then moves the decimal point in "101.0" forward two places, resulting in the floating-point number: 5.0 = 1.010 × 2 2
[0049] Among them, the sign part is a positive number, the mantissa part is "010", and the exponent part is "2".
[0050] The sign portion stored in memory is "0." When the exponent portion is stored in memory in single-precision, the single-precision offset value 127 is added to the exponent, resulting in an exponent portion of "100 000 01." The mantissa portion stored in memory is "010 000 000 000 000 000 000 000 00." The bits of the real number "5.0" ultimately stored in memory are shown in Figure 1.
[0051] In large language models, parameter distribution has limitations. For example, the weights, feature maps, gradients, and other data within a single copy exhibit normal or skewed distributions. Because the parameters within a large language model are relatively concentrated, the exponent values of the floating-point numbers corresponding to the same parameters are also relatively concentrated, creating potential for entropy coding. The mantissa of a floating-point number is typically uniformly and randomly distributed, making it suitable only for lossy compression.
[0052] To address the shortcomings of related technologies, embodiments of the present application provide a floating-point number compression method, decompression method, and computing device. The computing device can perform entropy coding on the exponential portion of the floating-point number to achieve lossless compression, and reserve a short code as the result of the lossless compression of the floating-point number with a value of zero, thereby achieving lossless compression of the transmitted data, reducing the communication bandwidth occupied during transmission, and eliminating the need to occupy additional computing resources of other modules.
[0053] FIG2 is a schematic diagram of the structure of a computing device provided in an embodiment of the present application. As shown in FIG2 , the computing device 200 includes at least one processor 210, a cache unit 220, a storage unit 230, a compressor 240, a decompressor 250, and a communication unit 260. The at least one processor 210 is connected to the cache unit 220. The cache unit 220 can be connected to the storage unit 230. The storage unit 230 is respectively connected to the compressor 240 and the decompressor 250. The compressor 240 and the decompressor 250 are both connected to the communication unit 260.
[0054] The computing device 200 may be a network interface card (NIC), a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), or other devices, or may be an electronic device that integrates a network card, a CPU, a GPU, an NPU, and other devices, such as a computer, a server, a desktop computer, and the like.
[0055] The processor 210 can be the core of a device such as a network card, CPU, GPU, NPU, etc., or it can be a device such as a network card, CPU, GPU, NPU, etc., used to process the parameter quantity of a large language model and output floating-point numbers.
[0056] The cache unit 220 includes private L2 cache (PL2), shared L2 cache (SL2), CPU cache, hard disk cache, network card cache, database cache, etc. The user caches the cache data of each processor 210, such as floating point numbers.
[0057] The storage unit 230 may be a random access memory (RAM), a hard disk drive (HDD), a solid state drive (SSD), etc., and is used to persistently store data such as floating-point numbers.
[0058] After receiving the floating-point number, the compressor 240 can use entropy coding to losslessly compress the exponential part of any floating-point number, and then form a data message with the compressed exponential part, the original sign part, and the original mantissa part, and send it to the communication unit 260. In the embodiment of the present application, as shown in Figure 3, after the compressor 240 receives the floating-point number, it can sort the floating-point numbers according to the time when the floating-point number is received to form a floating-point number queue. The compressor 240 can maintain a data sliding window and allow the data sliding window to slide on the floating-point number queue. The unit length of the data sliding window is one exponent, and the length of the data sliding window is greater than or equal to two exponents.
[0059] The compressor 240 can call a predictor to calculate the average value of the exponents of several floating-point numbers within the data sliding window as a temporary predicted value for the entire batch of floating-point numbers. The predictor does not need to calculate the predicted value for the entire batch of floating-point numbers, thereby reducing the amount of calculation of the compressor 240. In the embodiment of the present application, the predictor obtains the floating-point numbers within the data sliding window and extracts the exponential part of each floating-point number. The predictor averages the exponents of each floating-point number within the data sliding window and uses the average value as the predicted value of the exponents of several floating-point numbers within the data sliding window. The calculation formula of the predicted value is:
[0060] Among them, emean t represents the smoothed average value at time t, α is a smoothing parameter between 0 and 1, N is the window size, and e i is the exponential value of the i-th floating point number in the window.
[0061] When a floating-point number is zero, the exponent part of the floating-point number is usually the minimum value. According to the IEEE 754 floating-point number standard, the exponent range of a floating-point number is [-126, 127]. The exponent part when a floating-point number is zero is -127, which is a value that differs greatly from the exponent parts of other floating-point numbers. If there is a zero-value floating-point number in the data sliding window, the predicted value calculated by the predictor deviates greatly from the normal predicted value, resulting in a particularly large error in the difference of the exponent calculated for the floating-point number. In an embodiment of the present application, the predictor can detect whether the value of the floating-point number in the data sliding window is zero before calculating the predicted value. In one case, the predictor detects that there is a zero-value floating-point number in the current data sliding window, and the predicted value of the data sliding window of the previous jump can be used as the value of the zero-value floating-point number in the current data sliding window.
[0062] In another case, the predictor determines that the floating-point numbers in the current data sliding window include the first floating-point number in the floating-point number queue and there are floating-point numbers with a value of zero. The predictor can eliminate the floating-point numbers with a value of zero and use the exponents of the remaining floating-point numbers to calculate the predicted value of the exponent of the floating-point numbers in the current sliding window.
[0063] After compressor 240 calculates the predicted value, it can sequentially split each floating-point number within the data sliding window into three components: sign, exponent, and mantissa. Compressor 240 can subtract the exponent of the floating-point number from the predicted value to obtain the exponent difference of the floating-point number, thereby converting the floating-point number from a skewed distribution to a normal distribution, thereby enabling lossless compression using entropy coding.
[0064] After the compressor 240 obtains the exponential difference of the floating-point number, it performs static entropy coding on the exponential difference of the floating-point number to obtain the exponential part after entropy coding. The algorithm used by the compressor 240 for static entropy coding can be GRC, static Huffman coding, etc. In one embodiment, the specific process of the compressor 240 performing static entropy coding on the exponential difference of the floating-point number is as follows:
[0065] Compressor 240 converts the exponential portion into a binary number. It then counts the number of "1" occurrences in the binary number and uses a shorter code to represent exponential portions with a higher number of "1" occurrences, while using a longer code to represent exponential portions with a lower number of "1" occurrences. Finally, compressor 240 reassembles the encoded exponential portions into a binary integer and converts the integer back into the exponential portion of a floating-point number. Entropy encoding the exponential portion of the floating-point number by compressor 240 effectively improves data compression efficiency and storage space utilization, and is particularly valuable in large-scale computations and processing of floating-point numbers.
[0066] In the IEEE 754 floating-point number standard, the sign part of a floating-point number when it is zero is an integer, and the mantissa part is 0, so the sign part, mantissa part, and exponent part of a floating-point number when it is zero are all fixed values. In the embodiment of the present application, before the compressor 240 splits the floating-point number, it detects the value of the floating-point number. If the floating-point number is zero, it is not necessary to split the floating-point number, and the floating-point number with zero value can be directly converted into a special code. The special code is a specific code that allows the decompressor 250 to directly identify that the current floating-point number is zero. The format and bit position of the special code can be the same as the exponential part after entropy coding.
[0067] Compressor 240 calculates the entropy-coded exponential portion and can form a data message with the original sign portion, the original mantissa portion, the entropy-coded exponential portion, and the predicted value, and transmit the message to communication unit 260. If the floating-point number is zero, compressor 240 can form a data message with a special code. After compressor 240 completes the conversion of all floating-point numbers in the data sliding window into data messages, it can instruct the data sliding window to slide once so that the floating-point numbers in the next data sliding window can be processed.
[0068] After receiving the data message, the decompressor 250 can perform operations such as unpacking, decoding, and reassembling the data message to restore the original floating-point number. In the embodiment of the present application, after receiving the data message, the decompressor 250 splits the data message to obtain the original sign portion, the original mantissa portion, the exponential portion after entropy coding, and the predicted value. If the decompressor 250 cannot unpack or only splits out a special code, the decompressor 250 can directly output a floating-point number with a value of zero.
[0069] After obtaining the entropy-coded exponential portion, decompressor 250 decodes the entropy-coded exponential portion to restore the exponential difference of the floating-point number. Decompressor 250 adds the exponential difference of the floating-point number to the predicted value to obtain the original exponential portion. Decompressor 250 assembles the original exponential portion, the original sign portion, and the original mantissa portion to restore the original floating-point number.
[0070] The communication unit 260 can be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, etc., which is used to transmit data with other computing devices, send data packets to other computing devices, and receive data packets from other computing devices and send them to the decompressor 250.
[0071] In the embodiment of the present application, the compressor 240 and the decompressor 250 can be arranged not only between the storage unit 230 and the communication unit 260, but also between multiple processors 210 and the cache unit 220. If the compressor 240 and the decompressor 250 are arranged between multiple processors 210 and the cache unit 220, the compressor 240 and the decoder 250 can relieve the bandwidth pressure between the processor 210 and the cache unit 220 and between the cache unit 220 and the storage unit 230 while compressing or decompressing the communication data, and improve the hit rate of the cache unit 220.
[0072] FIG5 is a flow chart of a method for compressing floating-point numbers provided in an embodiment of the present application. As shown in FIG5 , the method for compressing floating-point numbers can be executed by the compressor 240 described above, or by other processors, cores, or other computing devices of a computing device. Taking the compressor 240 as an example, the processing process is as follows:
[0073] S501: The compressor 240 calculates a predicted value of the current data sliding window.
[0074] After receiving the floating-point numbers, compressor 240 can sort them according to the time they were received, forming a floating-point number queue. Compressor 240 can maintain a data sliding window that slides across the floating-point number queue. Compressor 240 obtains the floating-point numbers within the data sliding window and extracts the exponents of each floating-point number. Compressor 240 averages the exponents of the floating-point numbers within the data sliding window and uses this average as a predicted value for the exponents of the floating-point numbers within the data sliding window.
[0075] Before calculating the predicted value, the compressor 240 may detect whether the value of the floating-point number in the data sliding window is zero. In one case, the compressor 240 detects that there is a floating-point number with a value of zero in the current data sliding window, and the predicted value of the data sliding window of the previous jump may be used as the value of the floating-point number with a value of zero in the current data sliding window. In another case, the compressor 240 determines that the floating-point number in the current data sliding window is the floating-point number of the first jump of the floating-point number queue and there is a floating-point number with a value of zero. The compressor 240 may eliminate the floating-point number with a value of zero and use the exponents of the remaining floating-point numbers to calculate the predicted value of the exponent of the floating-point number in the current sliding window.
[0076] S502: The compressor 240 obtains a target floating-point number.
[0077] S503, determine whether the target floating point number is zero. In one case, when the target floating point number is zero, execute step S506. In another case, when the target floating point number is not zero, execute step S504.
[0078] S504 , the compressor 240 separates the target floating-point number into the original sign part, the original exponent part, and the original mantissa part.
[0079] S505 , the compressor 240 subtracts the original exponential part from the predicted value to obtain an exponential difference of the target floating-point number.
[0080] S506 , the compressor 240 performs static entropy coding on the exponent difference of the target floating-point number to obtain an entropy-coded exponential part.
[0081] When the target floating-point number is not zero, the compressor 240 may separate the target floating-point number obtained into three parts: a sign, an exponent, and a mantissa. The compressor 240 may subtract the exponent of the target floating-point number from the predicted value to obtain an exponent difference of the target floating-point number. After obtaining the exponent difference of the target floating-point number, the compressor 240 performs static entropy coding on the exponent difference of the target floating-point number to obtain an entropy-coded exponent portion.
[0082] When the target floating-point number is zero, since the sign part, mantissa part and exponent part of the zero-valued floating-point number are all fixed values, the compressor 240 does not need to split the target floating-point number, and can directly convert the zero-valued floating-point number into a special code, allowing the decompressor 250 to directly recognize that the target floating-point number is zero.
[0083] S507: The compressor 240 constructs a data message.
[0084] The compressor 240 calculates the exponential portion after entropy coding and can form a data message with the original sign portion, the original mantissa portion, the exponential portion after entropy coding and the predicted value. If the value of the target floating point number is zero, the compressor 240 can form a data message with special coding.
[0085] At step S508, the compressor 240 detects whether all floating-point numbers in the current data sliding window are constructed into data packets. In one case, if the compressor 240 detects that all floating-point numbers in the current data sliding window are constructed into data packets, step S509 is executed. In another case, if the compressor 240 detects that not all floating-point numbers in the current data sliding window are constructed into data packets, step S502 is executed.
[0086] S509 , the compressor 240 instructs the data sliding window to slide once on the floating point number queue.
[0087] After the compressor 240 completes converting all floating-point numbers in the data sliding window into data messages, it may instruct the data sliding window to slide once so as to process the floating-point numbers in the next data sliding window.
[0088] In this embodiment of the present application, after receiving the target floating-point number, compressor 240 subtracts the exponent of the target floating-point number from the predicted value, thereby converting the target floating-point number from a skewed distribution to a normal distribution. This allows the exponent of the target floating-point number to be losslessly compressed using entropy coding. This enables compressor 240 to perform lossless compression on the transmitted data, reducing the communication bandwidth occupied during transmission without requiring additional computing resources from other modules.
[0089] FIG6 is a flow chart of a method for decompressing a floating-point number provided in an embodiment of the present application. As shown in FIG6 , the method for decompressing a floating-point number can be executed by the above-mentioned decompressor 250, or by other processors, cores, or other computing devices of a computing device. Taking the decompressor 250 as an example, the processing process is as follows:
[0090] S601: After receiving a data message, the decompressor 250 splits the data message.
[0091] After receiving the data packet, the decompressor 250 decomposes the data packet to obtain the original sign portion, the original mantissa portion, the exponential portion after entropy coding, and the predicted value. If the decompressor 250 cannot decompress the packet or only decomposes a special code, the decompressor 250 can directly output a floating point number with a value of zero.
[0092] S602: The decompressor 250 decodes the exponential part after entropy coding to restore the exponential difference of the floating point number.
[0093] S603 , the decompressor 250 adds the exponential difference of the floating point number to the predicted value to obtain the original exponential part.
[0094] S604: The decompressor 250 assembles the original exponent part, the original sign part, and the original mantissa part to obtain an original floating-point number.
[0095] In the embodiment of the present application, after receiving a data message, the decompressor 250 parses the data message to obtain the sign portion, the mantissa portion, the entropy-coded exponent portion, and the predicted value. The decompressor 250 can decode the entropy-coded exponent portion, add the decoded difference value to the predicted value, and restore the original exponent portion, thereby restoring the original target floating-point number. This reduces the communication bandwidth occupied during transmission and does not require additional computing resources from other modules.
[0096] FIG7 is a schematic diagram of the structure of a floating-point number compression device provided in an embodiment of the present application. As shown in FIG7 , the floating-point number compression device 700 can be divided into a first processing unit 710, a second processing unit 720, a third processing unit 730, and a fourth processing unit 740 according to the execution function. The specific implementation process of the floating-point number compression device 700 is as follows:
[0097] The first processing unit 710 is used to obtain a target floating-point number. The target floating-point number includes an original sign portion, an original exponent portion, and an original mantissa portion. The second processing unit 720 is used to subtract the original exponent portion from the predicted value to obtain the exponent difference of the target floating-point number. The predicted value is the average value of a set number of floating-point numbers adjacent to the target floating-point number in the floating-point number queue. The floating-point number queue is a queue that sorts floating-point numbers in chronological order. The third processing unit 730 is used to perform entropy coding on the exponent difference of the target floating-point number to obtain the entropy-coded exponent portion. The fourth processing unit 740 is used to form a data message with the original sign portion, the original mantissa portion, the entropy-coded exponent portion, and the predicted value.
[0098] In one embodiment, the first processing unit 710 is further configured to use a data sliding window to slide on the floating point number queue to obtain multiple floating point numbers. The first processing unit 710 is further configured to average the values of the multiple floating point numbers to obtain a predicted value of the current data sliding window.
[0099] In one embodiment, the first processing unit 710 is further configured to detect whether there is a floating-point number with a value of zero among the multiple floating-point numbers. The first processing unit 710 is further configured to replace the value of the floating-point number with a predicted value of the last sliding data sliding window if there is a floating-point number with a value of zero among the multiple floating-point numbers.
[0100] In one embodiment, the first processing unit 710 is further configured to, when the floating-point number in the current data sliding window includes the first floating-point number in the floating-point number queue, detect whether there is a floating-point number with a value of zero among the multiple floating-point numbers. The first processing unit 710 is further configured to, when there is a floating-point number with a value of zero among the multiple floating-point numbers, eliminate the floating-point number with a value of zero.
[0101] In one embodiment, the second processing unit 720 is further configured to detect whether the value of the target floating-point number is zero. The second processing unit 720 is further configured to convert the target floating-point number into a specific encoding when the value of the target floating-point number is zero. The specific encoding instructs other computing devices to generate a floating-point number with a value of zero.
[0102] FIG8 is a schematic diagram of the structure of a floating-point number decompression device provided in an embodiment of the present application. As shown in FIG8 , the floating-point number decompression device 800 can be divided into a first processing unit 810, a second processing unit 820, a third processing unit 830, and a fourth processing unit 840 according to the execution function. The specific implementation process of the floating-point number compression device 800 is as follows:
[0103] The first processing unit 810 is used to receive a data message. The data message includes a sign portion, a mantissa portion, an entropy-coded exponential portion, and a predicted value. The predicted value is the average of a set number of floating-point numbers adjacent to a floating-point number in a floating-point number queue. The floating-point number queue is a queue that sorts floating-point numbers in chronological order. The second processing unit 820 is used to decode the entropy-coded exponential portion to obtain the exponential difference of the floating-point number. The third processing unit 830 is used to sum the exponential difference of the floating-point number with the predicted value to obtain the original exponential portion. The fourth processing unit 840 is used to construct a floating-point number based on the sign portion, the mantissa portion, and the original exponential portion.
[0104] In one embodiment, the fourth processing unit 840 is further configured to generate a floating point number with a value of zero according to the specific code when the data message includes the specific code.
[0105] A computing device is also provided in an embodiment of the present application. The computing device includes at least one memory and at least one processor. The at least one main processor can execute the technical solutions shown in Figures 1-6 and the above-mentioned corresponding protections, so that the computing device has the technical effects of the above-mentioned technical solutions.
[0106] In an embodiment of the present application, a computer-readable storage medium is further provided, on which a computer program is stored. When the computer program is executed in a computer, the computer is caused to execute the floating-point number compression method or the floating-point number decompression method described in any one of the above Figures 1 to 6 and the corresponding description content.
[0107] A computer program product is also provided in an embodiment of the present application. The computer program product stores instructions, and when the instructions are executed by a computer, the computer implements the floating-point number compression method or floating-point number decompression method described in any one of Figures 1 to 6 and the corresponding description content.
[0108] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the embodiments of this application.
[0109] In addition, various aspects or features of the embodiments of the present application can be implemented as methods, devices or products using standard programming and / or engineering techniques. The term "product" used in this application covers computer programs that can be accessed from any computer-readable device, carrier or medium. For example, computer-readable media may include, but are not limited to: magnetic storage devices (e.g., hard disks, floppy disks or tapes, etc.), optical disks (e.g., compact discs (CDs), digital versatile discs (DVDs), etc.), smart cards and flash memory devices (e.g., erasable programmable read-only memories (EPROMs), cards, sticks or key drives, etc.). In addition, the various storage media described herein may represent one or more devices and / or other machine-readable media for storing information. The term "machine-readable medium" may include, but is not limited to, wireless channels and various other media capable of storing, containing and / or carrying instructions and / or data.
[0110] In the above embodiment, the floating-point compression device 700 described in FIG7 and the floating-point decompression device 800 described in FIG8 can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) mode. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium may be a magnetic medium (eg, a floppy disk, a hard disk, a magnetic tape), an optical medium (eg, a high-density digital video disc (DVD)), or a semiconductor medium (eg, a solid state disk (SSD)).
[0111] It should be understood that in various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.
[0112] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0113] In the several embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0114] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.
[0115] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiment of the present application is essentially or the part that contributes to the prior art or the part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for enabling a computer device (which can be a personal computer, a server, or an access network device, etc.) to execute all or part of the steps of the method described in each embodiment of the embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0116] The above is only a specific implementation of the embodiment of the present application, but the protection scope of the embodiment of the present application is not limited to this. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed in the embodiment of the present application, and they should all be covered by the protection scope of the embodiment of the present application.
Claims
1. A compression method for floating-point numbers, characterized in that, The method is executed by a computing device, and the method includes: Obtain a target floating-point number; the target floating-point number includes an original sign part, an original exponent part, and an original mantissa part; Subtract the original exponent part from a predicted value to obtain an exponent difference of the target floating-point number; the predicted value is an average value of a set number of floating-point numbers adjacent to the target floating-point number in a floating-point number queue, and the floating-point number queue is a queue in which floating-point numbers are sorted in chronological order; Perform entropy coding on the exponent difference of the target floating-point number to obtain an entropy-coded exponent part; Construct a data packet from the original sign part, the original mantissa part, the entropy-coded exponent part, and the predicted value.
2. The method according to claim 1, characterized in that Before obtaining the target floating-point number, the method further includes: Use a data sliding window to slide on the floating-point number queue to obtain a plurality of floating-point numbers; Calculate the average value of the values of the plurality of floating-point numbers to obtain a predicted value of the current data sliding window.
3. The method according to claim 2, characterized in that, Before calculating the average value of the values of the plurality of floating-point numbers to obtain a predicted value of the current data sliding window, the method further includes: Detect whether there is a floating-point number with a value of zero among the plurality of floating-point numbers; In the case where there is a floating-point number with a value of zero among the plurality of floating-point numbers, replace the value of the floating-point number with a value of zero with the predicted value of the previous sliding data sliding window.
4. The method according to claim 2, wherein The floating-point numbers in the current data sliding window include the first floating-point number in the floating-point number queue, Before calculating the average value of the values of the plurality of floating-point numbers to obtain a predicted value of the current data sliding window, the method further includes: Detect whether there is a floating-point number with a value of zero among the plurality of floating-point numbers; In the case where there is a floating-point number with a value of zero among the plurality of floating-point numbers, remove the floating-point number with a value of zero.
5. The method according to any one of claims 1-4, characterized in that Before subtracting the original exponent part from the predicted value to obtain the exponent difference of the target floating-point number, the method further includes: Detect whether the value of the target floating-point number is zero; In the case where the value of the target floating-point number is zero, convert the target floating-point number into a specific code; the specific code instructs other computing devices to generate a floating-point number with a value of zero.
6. A method for decompressing a floating-point number, characterized in that, The method is executed by a computing device, and the method includes: Receive a data packet; the data packet includes a sign part, a mantissa part, an entropy-coded exponent part, and a predicted value, the predicted value is an average value of a set number of floating-point numbers adjacent to a floating-point number in a floating-point number queue, and the floating-point number queue is a queue in which floating-point numbers are sorted in chronological order; Decode the entropy-coded exponent part to obtain an exponent difference of the floating-point number; Add the exponent difference of the floating-point number to the predicted value to obtain the original exponent part; Construct a floating-point number according to the sign part, the mantissa part, and the original exponent part.
7. The method according to claim 6, wherein The data packet includes a specific code, and the method further includes: Generate a floating-point number with a value of zero according to the specific code.
8. A compression device for floating-point numbers, characterized in that, Includes: A first processing unit for obtaining a target floating-point number; The target floating-point number includes an original sign part, an original exponent part, and an original mantissa part; A second processing unit, configured to subtract the predicted value from the original exponent part to obtain an exponent difference of the target floating-point number; The predicted value is an average value of a set number of floating-point numbers adjacent to the target floating-point number in a floating-point number queue, and the floating-point number queue is a queue in which floating-point numbers are sorted in chronological order; A third processing unit, configured to perform entropy encoding on the exponent difference of the target floating-point number to obtain an entropy-encoded exponent part; A fourth processing unit, configured to form a data packet with the original sign part, the original mantissa part, the entropy-encoded exponent part, and the predicted value; 9. The apparatus according to claim 8, wherein: The first processing unit is further configured to slide a data sliding window on the floating-point number queue to obtain a plurality of floating-point numbers; Calculate an average value of the values of the plurality of floating-point numbers to obtain a predicted value of the current data sliding window.
10. The apparatus according to claim 9, wherein: The first processing unit is further configured to detect whether there is a floating-point number with a value of zero among the plurality of floating-point numbers; In the case where there is a floating-point number with a value of zero among the plurality of floating-point numbers, replace the value of the floating-point number with a value of zero with the predicted value of the previous sliding data sliding window.
11. The apparatus according to claim 9, wherein: The first processing unit is further configured to detect whether there is a floating-point number with a value of zero among the plurality of floating-point numbers when the floating-point numbers in the current data sliding window include the first floating-point number in the floating-point number queue; In the case where there is a floating-point number with a value of zero among the plurality of floating-point numbers, remove the floating-point number with a value of zero.
12. The apparatus according to any one of claims 8-11, wherein: The second processing unit is further configured to detect whether the value of the target floating-point number is zero; In the case where the value of the target floating-point number is zero, convert the target floating-point number into a specific code; the specific code instructs other computing devices to generate a floating-point number with a value of zero.
13. A floating-point number decompression device, characterized in that, Comprising: A first processing unit, configured to receive a data packet; The data packet includes a sign part, a mantissa part, an entropy-encoded exponent part, and a predicted value, and the predicted value is an average value of a set number of floating-point numbers adjacent to a floating-point number in a floating-point number queue, and the floating-point number queue is a queue in which floating-point numbers are sorted in chronological order; A second processing unit, configured to decode the entropy-encoded exponent part to obtain an exponent difference of the floating-point number; A third processing unit, configured to add the exponent difference of the floating-point number and the predicted value to obtain an original exponent part; A fourth processing unit, configured to form a floating-point number according to the sign part, the mantissa part, and the original exponent part; 14. The apparatus according to claim 13, wherein: The fourth processing unit is further configured to generate a floating-point number with a value of zero according to the specific code when the data packet includes the specific code.
15. A computing device, characterized in that, Comprising: At least one memory; At least one processor, the processor being configured to execute instructions stored in a memory to cause the computing device to perform the method according to any one of claims 1-5, and / or to perform the method according to any one of claims 6-7.
16. A computer-readable storage medium, characterized in that, Comprising computer program instructions which, when executed by a computing device, cause the computing device to perform the method according to any one of claims 1-5, and / or to perform the method according to any one of claims 6-7.
17. A computer program product comprising instructions, characterized in that, The computer program product stores instructions which, when executed by a computing device, cause the computing device to implement the method according to any one of claims 1-5, and / or to perform the method according to any one of claims 6-7.
Citation Information
Patent Citations
Floating-point number processing device, floating-point number addition device, and floating-point number processing method
CN112230882A
Method and device for encoding and decoding floating-point number
CN115113924A
Data processing method and device, electronic equipment and computer readable storage medium
CN117193707A
Data processing method and device and related equipment
CN117217296A
Content-aware compression of data with reduced number of class codes to be encoded
US10122379B1