Floating-point number compression method, floating-point number decompression method and computing equipment
Through the lossless compression method of entropy encoding of the exponential part of floating-point numbers in training large language model, the problem of insufficient communication bandwidth in multi-device training is solved, and efficient data transmission and optimization of computing resources are achieved.
Patent Information
- Application Number
- CN202311826087.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-27
- Publication Date
- 2025-06-27
AI Technical Summary
When multi-device training large language models, the communication bandwidth between devices is insufficient, making it difficult to meet training needs.
The compression method of floating-point numbers is adopted, and the exponential part of the floating-point number is entropy-encoding for lossless compression, and a short encoding is reserved as the zero-value floating-point number lossless compression is reduced, thereby reducing the communication bandwidth.
Lossless compression of the transmitted data is realized, reducing the communication bandwidth occupied during transmission, and no additional computing resources of other modules are required.
Smart Images

Figure CN120215873A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular, to a method for compressing floating-point numbers, a method for decompressing floating-point numbers, and a computing device. Background Art
[0002] With the development of large language models such as GPT and chat AI, the number of parameters of large language models is increasing and reaching a scale of hundreds of billions. Since the number of parameters of large language models is relatively large, the training process of large language models requires multiple devices to participate. When training the parameters distributed to multiple devices, communication is required between multiple devices so that the parameters of one device can be transmitted to another device. Usually, the communication bandwidth between devices is very small and it is difficult to meet the communication requirements for training large language models by multiple devices. Therefore, in the case of limited communication bandwidth between devices, how to implement the training of large language models by multiple devices is an urgent problem to be solved currently. Summary of the Invention
[0003] To solve the above problems, an embodiment of the present application provides a method for compressing floating-point numbers, which performs entropy encoding on the exponent part of the floating-point number to achieve lossless compression, and reserves a short code as the result of lossless compression of the floating-point number with a zero value, thereby realizing lossless compression of the data to be sent and reducing the communication bandwidth occupied during the transmission process. In addition, the present application also provides a floating-point number compression device corresponding to the floating-point number compression method, a floating-point number decompression method, a floating-point number decompression device corresponding to the floating-point number decompression method, and a computing device.
[0004] Therefore, the following technical solutions are adopted in the embodiments of the present application:
[0005] In a first aspect, an embodiment of the present application provides a method for compressing floating-point numbers, which is executed by a computing device. The method includes: obtaining a target floating-point number; the target floating-point number includes an original sign part, an original exponent part, and an original mantissa part; subtracting the prediction value from the original exponent part to obtain an exponent difference of the target floating-point number; the prediction value is the average value of a set number of floating-point numbers adjacent to the target floating-point number in a floating-point number queue, and the floating-point number queue is a queue in which floating-point numbers are sorted in chronological order; performing entropy encoding on the exponent difference of the target floating-point number to obtain an entropy-encoded exponent part; and forming a data packet with the original sign part, the original mantissa part, the entropy-encoded exponent part, and the prediction value.
[0006] In this embodiment, after the computing device receives the target floating-point number, it subtracts the exponent of the target floating-point number from the predicted value, thereby converting the target floating-point number from a skewed distribution to a normal distribution, so that the exponent of the target floating-point number can be losslessly compressed using entropy coding, enabling the computing device to perform lossless compression on the transmitted data, reducing the communication bandwidth occupied during transmission, and not requiring additional computing resources of other modules.
[0007] In one embodiment, before obtaining the target floating-point number, the method further includes: sliding a data sliding window over the floating-point number queue to obtain a plurality of floating-point numbers; and calculating the average value of the values of the plurality of floating-point numbers to obtain the predicted value of the current data sliding window.
[0008] In this embodiment, the computing device can call the data sliding window to slide over the floating-point number queue, and use the average value of the exponents of several floating-point numbers within the data sliding window as the predicted value of the temporary batch of floating-point numbers. The computing device does not need to calculate the predicted value for the entire batch of floating-point numbers, thereby reducing the computational amount of the computing device.
[0009] In one embodiment, before calculating the average value of the values of the plurality of floating-point numbers to obtain the predicted value of the current data sliding window, the method further includes: detecting whether there is a floating-point number with a value of zero among the plurality of floating-point numbers; and in the case where there is a floating-point number with a value of zero among the plurality of floating-point numbers, replacing the value of the floating-point number with a value of zero with the predicted value of the previous data sliding window.
[0010] In this embodiment, since the exponent part of the floating-point number with a value of zero is significantly different in value from the exponent parts of other floating-point numbers, the computing device can use the predicted value of the previous data sliding window as the value of the floating-point number with a value of zero, which can avoid the influence of the floating-point number with a value of zero on the normal predicted value and cause a particularly large difference error in the calculated exponent of the floating-point number.
[0011] In one embodiment, the floating-point numbers in the current data sliding window include the first floating-point number in the floating-point number queue. Before calculating the average value of the values of the plurality of floating-point numbers to obtain the predicted value of the current data sliding window, the method further includes: detecting whether there is a floating-point number with a value of zero among the plurality of floating-point numbers; and in the case where there is a floating-point number with a value of zero among the plurality of floating-point numbers, removing the floating-point number with a value of zero.
[0012] In this embodiment, if the first floating-point number in the floating-point number queue is included in the data sliding window, it indicates that the data sliding window slides for the first time, and the floating-point number with a value of zero cannot be replaced. Therefore, during the process of calculating the predicted value by the computing device, the floating-point number with a value of zero in the data sliding window can be excluded to avoid the influence of the floating-point number with a value of zero on the normal predicted value, resulting in a particularly large difference error in the exponent of the calculated floating-point number.
[0013] In one embodiment, before subtracting the predicted value from the original exponent part to obtain the exponent difference of the target floating-point number, the method further includes: detecting whether the value of the target floating-point number is zero; in the case where the value of the target floating-point number is zero, converting the target floating-point number into a specific code; the specific code instructs other computing devices to generate a floating-point number with a value of zero.
[0014] In this embodiment, since the sign part, mantissa part, and exponent part of the floating-point number with a value of zero are all fixed values, the floating-point number with a value of zero can be converted into a special code, and the decompressor can directly recognize the special code to output the floating-point number with a value of zero, thereby further reducing the communication bandwidth for transmitting the floating-point number with a value of zero and reducing the computational amount for processing the floating-point number with a value of zero.
[0015] In a second aspect, an embodiment of the present application provides a method for decompressing a floating-point number, characterized in that the method is executed by a computing device, and the method includes: receiving a data packet; the data packet includes a sign part, a mantissa part, an entropy-encoded exponent part, and a predicted value, the predicted value is the average value of a set number of floating-point numbers adjacent to a floating-point number in a floating-point number queue, and the floating-point number queue is a queue in which floating-point numbers are sorted in chronological order; decoding the entropy-encoded exponent part to obtain the exponent difference of the floating-point number;
[0016] Adding the exponent difference of the floating-point number to the predicted value to obtain the original exponent part; constructing a floating-point number according to the sign part, the mantissa part, and the original exponent part.
[0017] In this embodiment, after receiving the data packet, the computing device parses the sign part, mantissa part, entropy-encoded exponent part, and predicted value of the data packet. The computing device can decode the entropy-encoded exponent part, add the decoded difference to the predicted value, and restore the original exponent part, thereby restoring the original target floating-point number, realizing that the computing device reduces the communication bandwidth occupied during the transmission process and does not require additional occupation of the computing resources of other modules.
[0018] In one embodiment, the data packet includes a specific code, and the method further includes: generating a floating-point number with a value of zero according to the specific code.
[0019] In this embodiment, since the sign part, mantissa part, and exponent part of a floating-point number with a zero value are all fixed values, the floating-point number with a zero value can be converted into a special code. After receiving the special code, the computing device can directly recognize the special code to output the floating-point number with a zero value, so that the computing device reduces the communication bandwidth occupied during the transmission process and reduces the computational amount for processing the floating-point number with a zero value.
[0020] In a third aspect, an embodiment of the present application provides a floating-point number compression device, which is characterized by including: a first processing unit for obtaining a target floating-point number; the target floating-point number includes an original sign part, an original exponent part, and an original mantissa part; a second processing unit for subtracting a predicted value from the original exponent part to obtain an exponent difference of the target floating-point number; the predicted value is an average value of a set number of floating-point numbers adjacent to the target floating-point number in a floating-point number queue, and the floating-point number queue is a queue in which floating-point numbers are sorted in chronological order; a third processing unit for performing entropy coding on the exponent difference of the target floating-point number to obtain an entropy-coded exponent part; a fourth processing unit for forming a data packet with the original sign part, the original mantissa part, the entropy-coded exponent part, and the predicted value.
[0021] In one embodiment, the first processing unit is further configured to slide a data sliding window on the floating-point number queue to obtain a plurality of floating-point numbers; calculate an average value of the values of the plurality of floating-point numbers to obtain a predicted value of the current data sliding window.
[0022] In one embodiment, the first processing unit is further configured to detect whether there is a floating-point number with a zero value among the plurality of floating-point numbers; in the case where there is a floating-point number with a zero value among the plurality of floating-point numbers, replace the value of the floating-point number with a zero value with the predicted value of the previous sliding data sliding window.
[0023] In one embodiment, when the floating-point numbers in the current data sliding window include the first floating-point number in the floating-point number queue, the first processing unit is further configured to detect whether there is a floating-point number with a zero value among the plurality of floating-point numbers; in the case where there is a floating-point number with a zero value among the plurality of floating-point numbers, remove the floating-point number with a zero value.
[0024] In one embodiment, the second processing unit is further configured to detect whether the value of the target floating-point number is a zero value; in the case where the value of the target floating-point number is a zero value, convert the target floating-point number into a specific code; the specific code instructs other computing devices to generate a floating-point number with a zero value.
[0025] Fourthly, an apparatus for decompressing floating-point numbers according to an embodiment of the present application includes: a first processing unit configured to receive a data packet; the data packet includes a sign part, a mantissa part, an entropy-coded exponent part, and a predicted value, where the predicted value is an average value of a set number of floating-point numbers adjacent to a floating-point number in a floating-point number queue, and the floating-point number queue is a queue in which floating-point numbers are sorted in chronological order; a second processing unit configured to decode the entropy-coded exponent part to obtain an exponent difference of a floating-point number; a third processing unit configured to add the exponent difference of the floating-point number to the predicted value to obtain an original exponent part; and a fourth processing unit configured to form a floating-point number according to the sign part, the mantissa part, and the original exponent part.
[0026] In one implementation, the fourth processing unit is further configured to generate a floating-point number with a value of zero according to the specific coding when the data packet includes the specific coding.
[0027] Fifthly, a computing device according to an embodiment of the present application includes: at least one memory; at least one processor, where the processor is configured to execute instructions stored in the memory so that the computing device executes each possible implementation embodiment of the first aspect and / or executes each possible implementation embodiment of the second aspect.
[0028] Sixthly, a computer-readable storage medium according to an embodiment of the present application includes computer program instructions. When the computer program instructions are executed by a network device, the computing device executes each possible implementation embodiment of the first aspect and / or executes each possible implementation embodiment of the second aspect.
[0029] Seventhly, a computer program product including instructions according to an embodiment of the present application is characterized in that the computer program product stores instructions, and when the instructions are executed by a network device, the network device implements each possible implementation embodiment of the first aspect and / or executes each possible implementation embodiment of the second aspect. Description of the Drawings
[0030] The following briefly introduces the drawings required for description in the embodiments or the prior art.
[0031] Figure 1 It is a schematic diagram of bit positions during storage of a floating-point number;
[0032] Figure 2 It is a schematic structural diagram of a computing device provided in an embodiment of the present application;
[0033] Figure 3 It is a schematic diagram of a processing process of a compressor provided in an embodiment of the present application;
[0034] Figure 4 It is a schematic diagram of the processing process of a decompressor provided in the embodiment of the present application;
[0035] Figure 5 It is a schematic flowchart of a method for compressing floating-point numbers provided in the embodiment of the present application;
[0036] Figure 6 It is a schematic flowchart of a method for decompressing floating-point numbers provided in the embodiment of the present application;
[0037] Figure 7 It is a schematic structural diagram of a device for compressing floating-point numbers provided in the embodiment of the present application;
[0038] Figure 8 It is a schematic structural diagram of a device for decompressing floating-point numbers provided in the embodiment of the present application. Detailed implementation manners
[0039] Next, the technical solutions in the embodiments of the present application will be described with reference to the accompanying drawings in the embodiments of the present application.
[0040] The term "and / or" in this article is an association relationship describing associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The symbol " / " in this article represents an "or" relationship between associated objects. For example, A / B represents A or B.
[0041] The terms "first", "second", etc. in the description and claims of this article are used to distinguish different objects, rather than to describe a specific order of objects. For example, the first response message and the second response message are used to distinguish different response messages, rather than to describe the specific order of response messages.
[0042] In the embodiments of the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0043] In the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality" refers to two or more. For example, a plurality of processing units refers to two or more processing units, etc.; a plurality of elements refers to two or more elements, etc.
[0044] To solve the problem of insufficient communication bandwidth for data transmission between multiple devices, in related technologies, when multiple devices perform data transmission, the data to be sent can be losslessly compressed offline, and the received data can be decompressed online. Among them, the process of data compression is specifically as follows:
[0045] Taking a neural network processor as an example. The neural network processor uses run-length encoding to encode consecutive zero values into some special values, and then further compresses the index values and run lengths using Golomb-Rice coding (GRC). Given a parameter M, GRC performs fixed-length binary encoding on the last M bits of the original data and unary encoding on the remaining bits, which is suitable for scenarios where the frequency of smaller values is much higher than that of larger values. Since GRC contains a variable-length unary encoding, the neural network processor needs to serially count the number of consecutive bit 1s during the decoding process, resulting in difficulty in parallel processing during decoding by the neural network processor and poor scalability, and it is impossible to improve the decoding efficiency by doubling the hardware. Moreover, it is difficult to achieve hardware in-line during the compression process of the neural network processor, and software needs to be used to statistically analyze the parameter distribution for offline compression.
[0046] The number of parameters of large language models is usually represented and stored using floating-point numbers. A floating-point number is a data type used to represent real numbers. According to the definition of the Institute of Electrical and Electronics Engineers (IEEE), a floating-point number consists of a sign, an exponent, and a fraction. Among them, the sign part indicates whether the floating-point number is positive or negative. The exponent part is used to represent the order of magnitude (i.e., power) of the floating-point number, usually in the form of an optional counting method. The fraction part, also known as the significant digits, represents the main part of the floating-point number and is used to store the actual value. Compared with integers, floating-point numbers can better represent non-integer, extremely large, or extremely small numbers.
[0047] When the memory stores floating-point numbers in single-precision format, one floating-point number occupies 32 bits. Among them, the 32nd bit represents the sign part, where "0" indicates a positive value and "1" indicates a negative value. The 31st - 24th bits represent the exponent part. The 1st - 23rd bits represent the fraction part. Taking the real number "5.0" as an example, the memory converts "5.0" into a binary number, which is "101.0". Then, the memory moves the decimal point in "101.0" two places forward, and the resulting floating-point number is:
[0048] 5.0 = 1.010 × 2 2
[0049] Among them, the sign part is positive, the fraction part is "010", and the exponent part is "2".
[0050] The symbol part stored in the memory is "0". When the memory stores the exponent part in single-precision format, it is necessary to add the single-precision offset value 127 to the exponent, resulting in the exponent part being "100 000 01". The mantissa part stored in the memory is "010 000000 000 000 000 000 00". The bit positions of the real number "5.0" finally stored in the memory are as Figure 1 shown.
[0051] In large language models, there are limitations in the distribution of the number of parameters. For example, data such as weights, feature maps, and gradients of the same type show a normal or skewed distribution. Since the number of parameters of the same type in large language models is relatively concentrated, the values of the exponent parts of the floating-point numbers corresponding to the same number of parameters are relatively concentrated, having the potential for entropy coding. The mantissa parts of floating-point numbers are usually uniformly randomly distributed, so they are only suitable for lossy compression.
[0052] To solve the deficiencies in the related art, embodiments of the present application provide a method for compressing floating-point numbers, a method for decompressing them, and a computing device. The computing device can perform entropy coding on the exponent part of the floating-point number to achieve lossless compression, and reserve a short code as the result of the lossless compression of the floating-point number with a zero value, thereby achieving lossless compression of the transmitted data, reducing the communication bandwidth occupied during the transmission process, and not requiring additional computing resources of other modules.
[0053] Figure 2 This is a schematic structural diagram of a computing device provided in an embodiment of the present application. As Figure 2 shown, the computing device 200 includes at least one processor 210, a cache unit 220, a storage unit 230, a compressor 240, a decompressor 250, and a communication unit 260. At least one processor 210 is connected to the cache unit 220. The cache unit 220 can be connected to the storage unit 230. The storage unit 230 is respectively connected to the compressor 240 and the decompressor 250. The compressor 240 and the decompressor 250 are both connected to the communication unit 260.
[0054] The computing device 200 can be devices such as a network interface card (NIC), a central processing unit (CPU), a graphics processing unit (GPU), a neural processing unit (NPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), etc., or can be an electronic device integrated with devices such as a NIC, a CPU, a GPU, an NPU, etc., such as a computer, a server, a desktop computer, etc.
[0055] The processor 210 can be the core of devices such as a NIC, a CPU, a GPU, an NPU, etc., or can be devices such as a NIC, a CPU, a GPU, an NPU, etc., and is used to process the number of parameters of the large language model and output floating-point numbers.
[0056] The cache unit 220 is a private level 2 cache (PL2), a shared level 2 cache (SL2), a CPU cache, a hard disk cache, a NIC cache, a database cache, etc., and the user caches the cache data of each processor 210, such as floating-point numbers, etc.
[0057] The storage unit 230 can be a random access memory (RAM), a hard disk drive (HDD), a solid state drive (SSD), etc., and is used for persistent storage of data such as floating-point numbers.
[0058] After receiving the floating-point number, the compressor 240 can use entropy coding to losslessly compress the exponent part of any floating-point number, and then form a data packet with the compressed exponent part, the original sign part, and the original mantissa part, and send it to the communication unit 260. In the embodiments of the present application, as Figure 3 shown, after receiving the floating-point number, the compressor 240 can sort each floating-point number according to the time when the floating-point number is received to form a floating-point number queue. The compressor 240 can maintain a data sliding window and let the data sliding window slide on the floating-point number queue. The unit length of the data sliding window is one exponent, and the length of the data sliding window is greater than or equal to two exponents.
[0059] The compressor 240 can call a predictor to calculate the average value of the exponents of several floating-point numbers within the data sliding window as the predicted value of the temporary batch of floating-point numbers. The predictor can calculate the predicted value for the batch of floating-point numbers without having to do so, thereby reducing the computational amount of the compressor 240. In the embodiments of the present application, the predictor obtains the floating-point numbers within the data sliding window and extracts the exponent parts of the respective floating-point numbers. The predictor calculates the average value of the exponents of the respective floating-point numbers within the data sliding window and uses this average value as the predicted value of the exponents of several floating-point numbers within the data sliding window. The calculation formula for the predicted value is:
[0060]
[0061] where emean t represents the smoothed average value at time t, α is a smoothing parameter between 0 and 1, N is the window size, and e i is the exponent value of the i-th floating-point number within the window.
[0062] When the floating-point number is a zero value, the exponent part of the floating-point number is usually the minimum value. According to the IEEE 754 floating-point standard, the exponent range of the floating-point number is [-126, 127]. When the floating-point number is a zero value, the exponent part is -127, and this value differs greatly from the values of the exponent parts of other floating-point numbers. If there are zero-valued floating-point numbers within the data sliding window, the predicted value calculated by the predictor deviates greatly from the normal predicted value, resulting in a particularly large difference error in the calculated exponents of the floating-point numbers. In the embodiments of the present application, before calculating the predicted value, the predictor can detect whether the value of the floating-point number within the data sliding window is a zero value. In one case, when the predictor detects that there is a zero-valued floating-point number within the current data sliding window, it can use the predicted value of the previous-hop data sliding window as the value of the zero-valued floating-point number within the current data sliding window.
[0063] In another case, when the predictor determines that the floating-point numbers within the current data sliding window include the first floating-point number of the floating-point number queue and there are zero-valued floating-point numbers, the predictor can eliminate the zero-valued floating-point numbers and calculate the predicted value of the exponents of the floating-point numbers within the current sliding window using the exponents of the remaining floating-point numbers.
[0064] After the compressor 240 calculates the predicted value, it can successively split each floating-point number within the data sliding window into three parts: sign, exponent, and mantissa. The compressor 240 can subtract the predicted value from the exponent of the floating-point number to obtain the exponent difference of the floating-point number, thereby converting the floating-point number from a skewed distribution to a normal distribution for lossless compression using entropy coding.
[0065] After the compressor 240 obtains the exponent difference of the floating-point number, it performs static entropy encoding on the exponent difference of the floating-point number to obtain the exponent part after entropy encoding. The algorithms for the compressor 240 to perform static entropy encoding can be algorithms such as GRC and static Huffman encoding. In one embodiment, the specific process of the compressor 240 performing static entropy encoding on the exponent difference of the floating-point number is as follows:
[0066] The compressor 240 converts the exponent part into a binary number. Then, the compressor 240 counts the number of times "1" appears in the binary number, and represents the exponent part with a higher number of "1" occurrences using a shorter code, and represents the exponent part with a lower number of "1" occurrences using a longer code. Finally, the compressor 240 recombines the encoded exponent part into a binary integer and converts the integer back to the exponent part of the floating-point number. The compressor 240 performing entropy encoding on the exponent part of the floating-point number can effectively improve the data compression efficiency and the utilization rate of the storage space, and has important application value especially when performing large-scale calculations and processing on floating-point numbers.
[0067] In the IEEE 754 floating-point number standard, when the floating-point number is a zero value, the sign part is an integer, the mantissa part is 0, so the sign part, mantissa part, and exponent part of the floating-point number when it is a zero value are all fixed values. In the embodiment of the present application, before the compressor 240 splits the floating-point number, it detects the value of the floating-point number. If the floating-point number is a zero value, there is no need to split the floating-point number, and the floating-point number with a zero value can be directly converted into a special code. The special code is a specific code that can allow the decompressor 250 to directly recognize that the current floating-point number is a zero value. The format and bit positions of the special code can be the same as the exponent part after entropy encoding.
[0068] The compressor 240 calculates the exponent part after entropy encoding, and can form a data packet with the original sign part, the original mantissa part, the exponent part after entropy encoding, and the predicted value, and send it to the communication unit 260. If the value of the floating-point number is a zero value, the compressor 240 can form a data packet with the special code. After the compressor 240 completes the conversion of all floating-point numbers in the data sliding window into data packets, it can instruct the data sliding window to slide once to process the floating-point numbers in the next data sliding window.
[0069] After receiving the data packet, the decompressor 250 can perform operations such as unpacking, decoding, and assembling on the data packet to restore the original floating-point number. In the embodiment of the present application, after the decompressor 250 receives the data packet, it splits the data packet to obtain the original sign part, the original mantissa part, the exponent part after entropy encoding, and the predicted value. If the decompressor 250 cannot unpack or only splits out a special code, the decompressor 250 can directly output a floating-point number with a zero value.
[0070] After the decompressor 250 obtains the entropy-encoded exponent part, it performs a decoding operation on the entropy-encoded exponent part to restore the exponent difference of the floating-point number. The decompressor 250 adds the exponent difference of the floating-point number to the predicted value to obtain the original exponent part. The decompressor 250 assembles the original exponent part, the original sign part, and the original mantissa part to restore the original floating-point number.
[0071] The communication unit 260 can be a peripheral component interconnect (PCI) bus, an extended industry standard architecture (EISA) bus, etc., and is used for data transmission with other computing devices, so as to send data packets to other computing devices, receive data packets from other computing devices, and send them to the decompressor 250.
[0072] In the embodiments of the present application, the compressor 240 and the decompressor 250 can be not only arranged between the storage unit 230 and the communication unit 260, but also arranged between multiple processors 210 and the cache unit 220. If the compressor 240 and the decompressor 250 are arranged between multiple processors 210 and the cache unit 220, while compressing or decompressing communication data, the compressor 240 and the decoder 250 can also relieve the bandwidth pressure between the processor 210 and the cache unit 220, between the cache unit 220 and the storage unit 230, and improve the hit rate of the cache unit 220.
[0073] Figure 5 It is a schematic flowchart of a method for compressing a floating-point number provided in the embodiments of the present application. As Figure 5 shown, the method for compressing a floating-point number can be executed by the above-mentioned compressor 240, or other processors, cores, etc. of the computing device with computing functions. Taking the compressor 240 as an example below, the processing process is as follows:
[0074] S501, the compressor 240 calculates the predicted value of the current data sliding window.
[0075] After receiving the floating-point numbers, the compressor 240 can sort the floating-point numbers according to the time when the floating-point numbers are received to form a floating-point number queue. The compressor 240 can maintain a data sliding window and let the data sliding window slide on the floating-point number queue. The compressor 240 obtains the floating-point numbers in the data sliding window and extracts the exponent parts of the respective floating-point numbers. The compressor 240 calculates the average value of the exponents of the respective floating-point numbers in the data sliding window and uses this average value as the predicted value of the exponents of several floating-point numbers in the data sliding window.
[0076] Before calculating the predicted value, the compressor 240 can detect whether the value of the floating-point number in the data sliding window is zero. In one case, when the compressor 240 detects that there is a floating-point number with a zero value in the current data sliding window, it can use the predicted value of the previous-hop data sliding window as the value of the floating-point number with a zero value in the current data sliding window. In another case, when the compressor 240 determines that the floating-point number in the current data sliding window is the first-hop floating-point number of the floating-point number queue and there is a floating-point number with a zero value, the compressor 240 can eliminate the floating-point number with a zero value and calculate the predicted value of the exponent of the floating-point number in the current sliding window using the exponents of the remaining floating-point numbers.
[0077] S502, the compressor 240 obtains the target floating-point number.
[0078] S503, determine whether the target floating-point number is a zero value. In one case, when the target floating-point number is a zero value, step S506 is executed. In another case, when the target floating-point number is not a zero value, step S504 is executed.
[0079] S504, the compressor 240 splits the original sign part, original exponent part, and original mantissa part in the target floating-point number.
[0080] S505, the compressor 240 subtracts the predicted value from the original exponent part to obtain the exponent difference of the target floating-point number.
[0081] S506, the compressor 240 performs static entropy encoding on the exponent difference of the target floating-point number to obtain the entropy-encoded exponent part.
[0082] When the target floating-point number is not a zero value, the compressor 240 can split a target floating-point number it obtains into three parts: sign, exponent, and mantissa. The compressor 240 can subtract the predicted value from the exponent of the target floating-point number to obtain the exponent difference of the target floating-point number. After the compressor 240 obtains the exponent difference of the target floating-point number, it performs static entropy encoding on the exponent difference of the target floating-point number to obtain the entropy-encoded exponent part.
[0083] When the target floating-point number is a zero value, since the sign part, mantissa part, and exponent part of the floating-point number with a zero value are all fixed values, the compressor 240 does not need to split the target floating-point number and can directly convert the floating-point number with a zero value into a special encoding to allow the decompressor 250 to directly recognize that the target floating-point number is a zero value.
[0084] S507, the compressor 240 constructs a data packet.
[0085] The compressor 240 calculates the entropy-encoded exponent part, and can form a data packet with the original symbol part, the original mantissa part, the entropy-encoded exponent part, and the predicted value. When the value of the target floating-point number is zero, the compressor 240 can form a data packet with a special encoding.
[0086] S508, the compressor 240 detects whether all the floating-point numbers in the current data sliding window have been formed into data packets. In one case, when the compressor 240 detects that all the floating-point numbers in the current data sliding window have been formed into data packets, step S509 is executed. In another case, when the compressor 240 detects that not all the floating-point numbers in the current data sliding window have been formed into data packets, step S502 is executed.
[0087] S509, the compressor 240 instructs the data sliding window to slide once on the floating-point number queue.
[0088] After the compressor 240 has completed converting all the floating-point numbers in the data sliding window into data packets, it can instruct the data sliding window to slide once to process the floating-point numbers in the next data sliding window.
[0089] In the embodiment of the present application, after receiving the target floating-point number, the compressor 240 subtracts the exponent of the target floating-point number from the predicted value, thereby converting the target floating-point number from a skewed distribution to a normal distribution, so that the exponent of the target floating-point number can be losslessly compressed using entropy coding, realizing lossless compression of the data sent by the compressor 240, reducing the communication bandwidth occupied during the transmission process, and not requiring additional computing resources of other modules.
[0090] Figure 6 It is a schematic flowchart of a method for decompressing a floating-point number provided in the embodiment of the present application. As Figure 6 shown, the method for decompressing a floating-point number can be executed by the above decompressor 250, or by other processors, cores, etc. of the computing device with computing functions. Taking the decompressor 250 as an example below, the processing process is as follows:
[0091] S601, after receiving the data packet, the decompressor 250 splits the data packet.
[0092] After receiving the data packet, the decompressor 250 splits the data packet to obtain the original symbol part, the original mantissa part, the entropy-encoded exponent part, and the predicted value. If the decompressor 250 cannot unpack or only splits out a special encoding, the decompressor 250 can directly output a floating-point number with a value of zero.
[0093] S602, the decompressor 250 performs a decoding operation on the entropy-encoded exponent part to restore the exponent difference of the floating-point number.
[0094] S603, the decompressor 250 adds the exponent difference of the floating-point number and the predicted value to obtain the original exponent part.
[0095] S604, the decompressor 250 assembles the original exponent part, the original sign part, and the original mantissa part to obtain the original floating-point number.
[0096] In the embodiment of the present application, after receiving the data packet, the decompressor 250 parses the sign part, the mantissa part, the entropy-coded exponent part, and the predicted value of the data packet. The decompressor 250 can decode the entropy-coded exponent part, add the decoded difference and the predicted value to restore the original exponent part, thereby restoring the original target floating-point number, reducing the communication bandwidth occupied during the transmission process, and not requiring additional computing resources of other modules.
[0097] Figure 7 It is a schematic structural diagram of a floating-point number compression device provided in the embodiment of the present application. As Figure 7 shown, the floating-point number compression device 700 can be divided into a first processing unit 710, a second processing unit 720, a third processing unit 730, and a fourth processing unit 740 according to the executed functions. The specific implementation process of the floating-point number compression device 700 is as follows:
[0098] The first processing unit 710 is used to obtain the target floating-point number. The target floating-point number includes the original sign part, the original exponent part, and the original mantissa part. The second processing unit 720 is used to subtract the predicted value from the original exponent part to obtain the exponent difference of the target floating-point number. The predicted value is the average value of a set number of floating-point numbers adjacent to the target floating-point number in the floating-point number queue. The floating-point number queue is a queue in which floating-point numbers are sorted in chronological order. The third processing unit 730 is used to perform entropy coding on the exponent difference of the target floating-point number to obtain the entropy-coded exponent part. The fourth processing unit 740 is used to form a data packet with the original sign part, the original mantissa part, the entropy-coded exponent part, and the predicted value.
[0099] In one implementation, the first processing unit 710 is further used to slide a data sliding window on the floating-point number queue to obtain multiple floating-point numbers. The first processing unit 710 is further used to calculate the average value of the values of the multiple floating-point numbers to obtain the predicted value of the current data sliding window.
[0100] In one implementation, the first processing unit 710 is further used to detect whether there is a floating-point number with a value of zero among the multiple floating-point numbers. The first processing unit 710 is further used to replace the value of the floating-point number with a value of zero with the predicted value of the previous sliding data sliding window when there is a floating-point number with a value of zero among the multiple floating-point numbers.
[0101] In one embodiment, the first processing unit 710 is further configured to detect whether there is a floating-point number with a value of zero among multiple floating-point numbers when the floating-point numbers in the current data sliding window include the first floating-point number in the floating-point number queue. The first processing unit 710 is further configured to remove the floating-point number with a value of zero in the case where there is a floating-point number with a value of zero among the multiple floating-point numbers.
[0102] In one embodiment, the second processing unit 720 is further configured to detect whether the value of the target floating-point number is zero. The second processing unit 720 is further configured to convert the target floating-point number into a specific encoding in the case where the value of the target floating-point number is zero. The specific encoding instructs other computing devices to generate a floating-point number with a value of zero.
[0103] Figure 8 FIG. is a schematic structural diagram of a floating-point number decompression device provided in an embodiment of the present application. As Figure 8 shown, the floating-point number decompression device 800 can be divided into a first processing unit 810, a second processing unit 820, a third processing unit 830, and a fourth processing unit 840 according to the executed functions. The specific implementation process of the floating-point number compression device 800 is as follows:
[0104] The first processing unit 810 is configured to receive a data packet. The data packet includes a sign part, a mantissa part, an entropy-encoded exponent part, and a predicted value. The predicted value is the average value of a set number of floating-point numbers adjacent to a floating-point number in the floating-point number queue. The floating-point number queue is a queue in which floating-point numbers are sorted in chronological order. The second processing unit 820 is configured to decode the entropy-encoded exponent part to obtain the exponent difference of the floating-point number. The third processing unit 830 is configured to add the exponent difference of the floating-point number to the predicted value to obtain the original exponent part. The fourth processing unit 840 is configured to form a floating-point number according to the sign part, the mantissa part, and the original exponent part.
[0105] In one embodiment, the fourth processing unit 840 is further configured to generate a floating-point number with a value of zero according to the specific encoding when the data packet includes the specific encoding.
[0106] An embodiment of the present application further provides a computing device, which includes at least one memory and at least one processor. The at least one main processor can execute as Figures 1 - 6 and the corresponding protected technical solutions described above, so that the computing device has the technical effects of the protected technical solutions described above.
[0107] An embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed on a computer, the computer is made to execute the above Figures 1 - 6The floating-point number compression method or floating-point number decompression method described in any one of the corresponding description contents.
[0108] In an embodiment of the present application, a computer program product is further provided. The computer program product stores instructions that, when executed by a computer, cause the computer to implement the above Figures 1 - 6 The floating-point number compression method or floating-point number decompression method described in any one of the corresponding description contents.
[0109] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the embodiments of the present application.
[0110] In addition, various aspects or features of the embodiments of the present application can be implemented as a method, apparatus, or article of manufacture using standard programming and / or engineering techniques. The term "article of manufacture" used in the present application covers a computer program accessible from any computer-readable device, carrier, or medium. For example, the computer-readable medium can include, but is not limited to: magnetic storage devices (such as hard disks, floppy disks, or magnetic tapes, etc.), optical discs (such as compact discs (CDs), digital versatile discs (DVDs), etc.), smart cards, and flash memory devices (such as erasable programmable read-only memories (EPROMs), cards, sticks, or key drives, etc.). Additionally, the various storage media described herein can represent one or more devices and / or other machine-readable media for storing information. The term "machine-readable medium" can include, but is not limited to, wireless channels and various other media capable of storing, containing, and / or carrying instructions and / or data.
[0111] In the above embodiments, Figure 7 The floating-point number compression device 700 in Figure 8The floating-point number decompression device 800 described above can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that the computer can access or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a high-density digital video disc (DVD)), or a semiconductor medium (such as a solid-state disk (SSD)), etc.
[0112] It should be understood that in various embodiments of the embodiments of the present application, the magnitudes of the sequence numbers of the above processes do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0113] Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be described herein again.
[0114] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.
[0115] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0116] If the above functions are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the embodiments of this application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or an access network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.
[0117] The above is only the specific implementation manner of the embodiments of this application, but the protection scope of the embodiments of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in the embodiments of this application, and all should be covered by the protection scope of the embodiments of this application.
Claims
1. A compression method for floating-point numbers, characterized in that, The method is executed by a computing device, and the method includes: Obtain a target floating-point number; the target floating-point number includes an original sign part, an original exponent part, and an original mantissa part; Subtract the original exponent part from a predicted value to obtain an exponent difference of the target floating-point number; the predicted value is an average value of a set number of floating-point numbers adjacent to the target floating-point number in a floating-point number queue, and the floating-point number queue is a queue in which floating-point numbers are sorted in chronological order; Perform entropy encoding on the exponent difference of the target floating-point number to obtain an entropy-encoded exponent part; Construct a data packet from the original sign part, the original mantissa part, the entropy-encoded exponent part, and the predicted value.
2. The method according to claim 1, characterized in that Before obtaining the target floating-point number, the method further includes: Use a data sliding window to slide on the floating-point number queue to obtain a plurality of floating-point numbers; Calculate an average value of the values of the plurality of floating-point numbers to obtain a predicted value of the current data sliding window.
3. The method according to claim 2, wherein Before calculating an average value of the values of the plurality of floating-point numbers to obtain a predicted value of the current data sliding window, the method further includes: Detect whether there is a floating-point number with a value of zero among the plurality of floating-point numbers; In the case where there is a floating-point number with a value of zero among the plurality of floating-point numbers, replace the value of the floating-point number with a value of zero with the predicted value of the previous sliding data sliding window.
4. The method according to claim 2, wherein The floating-point numbers in the current data sliding window include the first floating-point number in the floating-point number queue, Before calculating an average value of the values of the plurality of floating-point numbers to obtain a predicted value of the current data sliding window, the method further includes: Detect whether there is a floating-point number with a value of zero among the plurality of floating-point numbers; In the case where there is a floating-point number with a value of zero among the plurality of floating-point numbers, remove the floating-point number with a value of zero.
5. The method according to any one of claims 1-4, characterized in that, Before subtracting the original exponent part from the predicted value to obtain an exponent difference of the target floating-point number, the method further includes: Detect whether the value of the target floating-point number is zero; In the case where the value of the target floating-point number is zero, convert the target floating-point number into a specific code; the specific code instructs other computing devices to generate a floating-point number with a value of zero.
6. A method for decompressing a floating-point number, characterized in that, The method is executed by a computing device, and the method includes: Receive a data packet; the data packet includes a sign part, a mantissa part, an entropy-encoded exponent part, and a predicted value, the predicted value is an average value of a set number of floating-point numbers adjacent to a floating-point number in a floating-point number queue, and the floating-point number queue is a queue in which floating-point numbers are sorted in chronological order; Decode the entropy-encoded exponent part to obtain an exponent difference of the floating-point number; Add the exponent difference of the floating-point number to the predicted value to obtain the original exponent part; Construct a floating-point number according to the sign part, the mantissa part, and the original exponent part.
7. The method according to claim 6, wherein The data packet includes a specific code, and the method further includes: Generate a floating-point number with a value of zero according to the specific code.
8. A compression device for floating-point numbers, characterized in that, Includes: A first processing unit for obtaining a target floating-point number; The target floating-point number includes an original sign part, an original exponent part, and an original mantissa part; A second processing unit, configured to subtract the predicted value from the original exponent part to obtain an exponent difference of the target floating-point number; The predicted value is an average value of a set number of floating-point numbers adjacent to the target floating-point number in a floating-point number queue, and the floating-point number queue is a queue in which floating-point numbers are sorted in chronological order; A third processing unit, configured to perform entropy encoding on the exponent difference of the target floating-point number to obtain an entropy-encoded exponent part; A fourth processing unit, configured to form a data packet with the original sign part, the original mantissa part, the entropy-encoded exponent part, and the predicted value.
9. The apparatus according to claim 8, wherein The first processing unit is further configured to slide a data sliding window on the floating-point number queue to obtain a plurality of floating-point numbers; Calculate an average value of the values of the plurality of floating-point numbers to obtain a predicted value of the current data sliding window.
10. The apparatus according to claim 9, wherein The first processing unit is further configured to detect whether there is a floating-point number with a value of zero among the plurality of floating-point numbers; In the case where there is a floating-point number with a value of zero among the plurality of floating-point numbers, replace the value of the floating-point number with a value of zero with the predicted value of the previous sliding data sliding window.
11. The apparatus according to claim 9, wherein The first processing unit is further configured to detect whether there is a floating-point number with a value of zero among the plurality of floating-point numbers when the floating-point numbers in the current data sliding window include the first floating-point number in the floating-point number queue; In the case where there is a floating-point number with a value of zero among the plurality of floating-point numbers, remove the floating-point number with a value of zero.
12. The apparatus according to any one of claims 8-11, wherein The second processing unit is further configured to detect whether the value of the target floating-point number is zero; In the case where the value of the target floating-point number is zero, convert the target floating-point number into a specific code; the specific code instructs other computing devices to generate a floating-point number with a value of zero.
13. A decompression device for floating-point numbers, characterized in that, Comprising: A first processing unit, configured to receive a data packet; The data packet includes a sign part, a mantissa part, an entropy-encoded exponent part, and a predicted value, and the predicted value is an average value of a set number of floating-point numbers adjacent to a floating-point number in a floating-point number queue, and the floating-point number queue is a queue in which floating-point numbers are sorted in chronological order; A second processing unit, configured to decode the entropy-encoded exponent part to obtain an exponent difference of the floating-point number; A third processing unit, configured to add the exponent difference of the floating-point number to the predicted value to obtain an original exponent part; A fourth processing unit, configured to form a floating-point number according to the sign part, the mantissa part, and the original exponent part.
14. The apparatus according to claim 13, wherein The fourth processing unit is further configured to generate a floating-point number with a value of zero according to the specific code when the data packet includes the specific code.
15. A computing device, characterized in that, Comprising: At least one memory; At least one processor, the processor being configured to execute instructions stored in a memory to cause the computing device to perform the method according to any one of claims 1-5, and / or to perform the method according to any one of claims 6-7.
16. A computer-readable storage medium, characterized in that, Comprising computer program instructions which, when executed by a computing device, cause the computing device to perform the method according to any one of claims 1-5, and / or to perform the method according to any one of claims 6-7.
17. A computer program product containing instructions, characterized in that, The computer program product stores instructions which, when executed by a computing device, cause the computing device to implement the method according to any one of claims 1-5, and / or to perform the method according to any one of claims 6-7.