A communication data compression method and system based on quantization technology
By using quantization technology to compress and recover transmission vectors in the distributed GMRES algorithm, the problem of data communication accounting for a large proportion in the distributed GMRES algorithm is solved, and the data transmission volume is reduced and iterative convergence is improved.
Patent Information
- Application Number
- CN202411125553.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-16
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2044-08-16
AI Technical Summary
During the large-scale sparse linear system solution, the data communication between nodes accounts for a large proportion of the total running time of the iterative algorithm, becoming a performance bottleneck.
The communication data compression method based on quantization technology is adopted, and the transmission vector is compressed in the distributed iteration process through quantization and inverse quantization strategies, and the quantization module and inverse quantization module are used for quantization compression and accuracy recovery respectively.
It significantly reduces the amount of data transmission, improves the utilization efficiency of bit width, reduces the impact on iterative convergence, and supports compression of arbitrary bit widths.
Smart Images

Figure CN118944678B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of communications, and particularly relates to a communication data compression method and system based on quantization technology. Background Art
[0002] The Generalized Minimal Residual method (GMRES) is an iterative algorithm for solving the sparse linear system Ax = b, where A is a sparse matrix, and x and b are dense vectors. There are usually two methods for solving sparse linear systems, namely the direct method and the iterative method. The direct method has the advantages of good versatility and high accuracy of calculation results, but it faces problems such as high storage requirements and large computational amounts when solving large-scale sparse linear systems. The iterative method generates a series of approximate solutions approaching the exact solution through multiple iterations, and has the advantages of small computational amount, low memory requirements, and can make full use of the sparsity of the matrix. It is the mainstream method for solving large-scale sparse linear systems and is widely used in scientific computing and simulation fields such as computational electromagnetics and computational fluid dynamics.
[0003] Since a single computing node cannot meet the memory and computing resource requirements for solving large-scale sparse linear systems, distributed iterative solvers are more commonly used in scientific and engineering computing. Although distributed systems have more computing and storage resources available for problem solving, the communication delay between nodes will increase significantly compared to the communication delay within a single node.
[0004] The distributed computing of GMRES is an important means for solving large-scale sparse linear systems. Its distributed implementation distributes the main computational kernels such as sparse matrix-vector multiplication (SpMV), vector addition, and vector inner product to each node for calculation. To ensure the normal operation of the iteration, each process needs to merge and reduce the local calculation results of typical vector operators such as vector addition, vector inner product, and two-norm. These calculation results are usually scalars and have relatively small communication overheads. For the SpMV operator y = Ap involving sparse matrix operations, each process needs to obtain the complete multiplying vector p through inter-process communication before calling the SpMV kernel in each iteration, and the data for inter-process communication is a vector. The analysis of the experimental results of distributed GMRES on 29 sparse linear systems shows that the average ratio of the vector communication time before the execution of the SpMV operator to the total running time of the iterative algorithm exceeds 80%. This means that the data communication between distributed nodes is the main performance bottleneck of distributed GMRES, and how to reduce its communication overhead is an important direction for optimizing the overall efficiency of distributed GMRES. Summary of the Invention
[0005] To solve the above technical problems, the present invention provides a communication data compression method and system based on quantization technology. Among them, a communication data compression method based on quantization technology includes:
[0006] Perform distributed iteration based on quantization and dequantization strategies. After each process obtains its own local calculation result, quantize and compress the transmission vector based on the numerical characteristics of the local calculation result.
[0007] Each process sends the quantized result to other processes through the network, and each process obtains the complete quantized result.
[0008] Based on the complete quantized result, each process dequantizes the data it receives to obtain a vector represented in the original precision, and then performs local SpMV operations and subsequent operations on the vector represented in the original precision.
[0009] Preferably, the process of performing distributed iteration based on quantization and dequantization strategies includes:
[0010] In each iteration process, each process first quantizes the data represented in high precision that it wants to transmit to other processes to a fixed-point representation with w-bit width through the quantization function FP2INTw, then sends it to other processes, and receives the quantized data from other processes to obtain the complete quantized result v w ; then call INTw2FP to dequantize the complete quantized result v w to a high-precision floating-point representation v out , and then continue with subsequent SpMV and other operations.
[0011] Preferably, the process of quantizing and compressing the transmission vector based on the numerical characteristics of the local calculation result includes:
[0012] Map all floating-point numbers in the interval [a, b] to an integer interval with a fixed step size based on the communication compression method of uniform quantization. Given a floating-point vector x in and an unsigned integer grid, perform quantization operations to map the real number vector to the integer grid with a lower bit width to achieve data compression.
[0013] Preferably, the quantization process is expressed as:
[0014]
[0015] Among them, means rounding the input data to the nearest integer, and the definition of the clamp function is:
[0016]
[0017] Preferably, according to different choices of zeros, the uniform quantization in the communication compression method based on uniform quantization includes symmetric quantization and asymmetric quantization;
[0018] The parameter composition of the asymmetric quantization includes a scaling factor S, a zero point z, and a bit width w; among them, the scaling factor S is a floating-point number used to determine the step size of the quantizer.
[0019] Preferably, the process of quantizing and compressing the transmission vector based on the numerical characteristics of the local calculation result includes:
[0020] For the communication compression method based on non-uniform quantization, a quantization technique based on a piecewise linear function is used to introduce a segmentation point in the quantization range, splitting the quantization interval into two non-overlapping regions to obtain a central dense region and an edge discrete region;
[0021] Based on the quantization technique of the piecewise linear function, the data distributed in the two quantization intervals are uniformly quantized respectively, and the quantization loss of the edge discrete region and the central dense region is regulated by controlling the segmentation point.
[0022] Preferably, the process of using the quantization technique based on the piecewise linear function to introduce a segmentation point in the quantization range and splitting the quantization interval into two non-overlapping regions includes:
[0023] Introduce a parameter p to represent the proportion of elements included in the central dense region, and by adjusting the parameter p, the segmentation of the central dense region and the edge discrete region is achieved, and the segmentation point setting for data distribution adaptation is completed.
[0024] Preferably, the process of uniformly quantizing the data distributed in the two quantization intervals respectively based on the quantization technique of the piecewise linear function includes:
[0025] Through the hybrid compression of quantization mapping and floating-point representation, the elements in the edge discrete region are not quantized and compressed, and the original FP64 or FP32 representation is still used;
[0026] Define the proportion of elements included in the central dense region as α, then the proportion of elements included in the edge discrete region is 1 - α. Assume that 64-bit integers are used to represent the edge discrete region, and the central dense region is compressed by uniform quantization with a quantization bit width of w. Then the communication compression ratio η achieved by quantization is:
[0027]
[0028] Preferably, based on the complete quantization result, the process of each process performing inverse quantization on the data it receives includes:
[0029] After inter-process communication of the compressed data, each process performs inverse quantization on the received data, and the floating-point vector obtained by inverse quantization is approximately equal to the initial floating-point vector;
[0030] Among them, the process of the inverse quantization is expressed as:
[0031] x out = s·(x w - z)
[0032] The quantization factor s is determined by the quantization range [a, b] and the quantization bit width w in the following manner:
[0033]
[0034] The present invention also provides a communication data compression system based on quantization technology, including:
[0035] A quantization module, configured to, after each process obtains its own local calculation result, perform quantization compression on the transmission vector based on the numerical characteristics of the local calculation result; each process sends the quantized result to other processes through the network, and each process obtains the complete quantized result;
[0036] An inverse quantization module, connected to the quantization module, configured to, based on the complete quantized result, each process perform inverse quantization on the data it receives to obtain a vector represented in the original precision, and then perform local SpMV operations and subsequent operations on the vector represented in the original precision.
[0037] Compared with the prior art, the present invention has the following advantages and technical effects:
[0038] By quantizing floating-point numbers and then transmitting them, the present invention can, on the one hand, significantly reduce the data transmission volume, and on the other hand, can effectively utilize the bit width, thereby reducing the impact on the iterative convergence.
[0039] The communication compression method based on quantization technology proposed by the present invention maps a given set of floating-point numbers to a fixed-point number interval within a specific range according to the numerical characteristics of the limited floating-point data. Compared with the floating-point format representation that divides the exponent and mantissa, the quantization representation can improve the utilization efficiency of the limited bit width, thereby improving the representation precision of the floating-point numbers. More importantly, the data compressed based on quantization technology is only used for communication and not for calculation, so compression with any bit width is supported. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] The drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:
[0041] Figure 1 It is the numerical distribution diagram of the communication vector in the iterative algorithm of the embodiment of the present invention; among them, (a) is the exponent distribution diagram; (b) is the mantissa distribution diagram; (c) is the distribution diagram of the change of the information entropy of the value;
[0042] Figure 2 Element ratio distribution diagram of the same index for the embodiments of the present invention; wherein, (a) is the ratio distribution diagram of nd12k; (b) is the ratio distribution diagram of nd24k; (c) is the ratio distribution diagram of bmwcra_1;
[0043] Figure 3 Schematic diagram of the distributed iterative process based on quantization technology for the embodiments of the present invention;
[0044] Figure 4 Data distribution diagram of the communication floating-point vector for the embodiments of the present invention; wherein, (a) is the data distribution diagram of bmwcra_1; (b) is the data distribution diagram of nd12k. Detailed implementation manners
[0045] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine with the embodiments to detail this application.
[0046] It should be noted that the steps shown in the flowchart of the drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0047] Embodiment 1
[0048] This embodiment provides a communication data compression method based on quantization technology, including:
[0049] Performing distributed iteration based on quantization and inverse quantization strategies. When each process obtains its own local calculation result, quantize and compress the transmission vector based on the numerical characteristics of the local calculation result;
[0050] Each process sends the quantized result to other processes through the network, and each process obtains the complete quantized result;
[0051] Based on the complete quantized result, each process performs inverse quantization on the data it receives to obtain a vector represented in the original precision, and then performs local SpMV operations and subsequent operations on the vector represented in the original precision.
[0052] Further, the process of performing distributed iteration based on quantization and inverse quantization strategies includes:
[0053] In each iteration process, each process first quantizes the data represented in high precision that it wants to transmit to other processes to a fixed-point representation with w-bit width through the quantization function FP2INTw, then sends it to other processes, and receives the quantized data from other processes to obtain the complete quantized result v w; Then call INTw2FP to convert the complete quantization result v w to a high-precision floating-point representation v out , and then continue with subsequent SpMV and other operations.
[0054] Furthermore, the process of quantizing and compressing the transmission vector based on the numerical characteristics of the local calculation results includes:
[0055] The communication compression method based on uniform quantization maps all floating-point numbers in the interval [a, b] to an integer interval with a fixed step size. Given a floating-point vector x in and an unsigned integer grid, perform quantization operations to map the real number vector to the lower-width integer grid to achieve data compression.
[0056] Furthermore, the quantization process is expressed as:
[0057]
[0058] where represents rounding the input data to the nearest integer, and the definition of the clamp function is:
[0059]
[0060] Furthermore, according to different choices of the zero point, the uniform quantization in the communication compression method based on uniform quantization includes symmetric quantization and asymmetric quantization;
[0061] The parameter composition of asymmetric quantization includes a scaling factor S, a zero point z, and a bit width w; among them, the scaling factor S is a floating-point number used to determine the step size of the quantizer.
[0062] Furthermore, the process of quantizing and compressing the transmission vector based on the numerical characteristics of the local calculation results includes:
[0063] The communication compression method based on non-uniform quantization uses a quantization technique based on a piecewise linear function to introduce a breakpoint in the quantization range, splitting the quantization interval into two non-overlapping regions to obtain a central dense region and an edge discrete region;
[0064] The quantization technique based on the piecewise linear function uniformly quantizes the data distributed in the two quantization intervals respectively, and controls the quantization loss of the edge discrete region and the central dense region by adjusting the breakpoint.
[0065] Furthermore, the process of using the quantization technique based on the piecewise linear function to introduce a breakpoint in the quantization range and splitting the quantization interval into two non-overlapping regions includes:
[0066] Introduce a parameter p to represent the proportion of elements contained in the central dense region. By adjusting the parameter p, segment the central dense region and the edge discrete region to complete the setting of the segmentation point for adaptive data distribution.
[0067] Furthermore, the process of uniformly quantizing the data distributed in the two quantization intervals based on the quantization technique of the piecewise linear function includes:
[0068] Through the hybrid compression of quantization mapping and floating-point representation, the elements in the edge discrete region are not quantized and compressed, and the original FP64 or FP32 representation is still used;
[0069] Define the proportion of elements contained in the central dense region as α, then the proportion of elements contained in the edge discrete region is 1 - α. Assume that 64-bit integers are used to represent the edge discrete region, and the central dense region is compressed by uniform quantization with a quantization bit width of w. Then the communication compression ratio η achieved by quantization is:
[0070]
[0071] Furthermore, based on the complete quantization result, the process of each process dequantizing the data it receives includes:
[0072] After inter-process communication of the compressed data, each process dequantizes the received data, and the dequantized floating-point vector is approximately equal to the initial floating-point vector;
[0073] Among them, the process of dequantization is expressed as:
[0074] x out =s·(x w -z)
[0075] The quantization factor s is determined by the quantization range [a, b] and the quantization bit width w in the following way:
[0076]
[0077] Embodiment 2
[0078] This embodiment also provides a communication data compression system based on quantization technology, including:
[0079] A quantization module, which is used to, after each process obtains its own local calculation result, quantize and compress the transmission vector based on the numerical characteristics of the local calculation result; each process sends the quantized result to other processes through the network, and each process obtains the complete quantization result;
[0080] The dequantization module, connected to the quantization module, is used to dequantize the data received by each process based on the complete quantization result to obtain a vector represented in the original precision, and then perform local SpMV operations and subsequent operations on the vector represented in the original precision.
[0081] Embodiment III
[0082] The prior art hides data communication or reduces communication overhead through methods such as sparse matrix rearrangement, sparse matrix partitioning based on graph partitioning algorithms, computation and communication overlap, and selective transmission of communication vectors. Among them, sparse matrix rearrangement and sparse matrix partitioning based on graph partitioning minimize communication requirements at the source, and computation and communication overlap hide part of the communication process through asynchronous execution of tasks. Selective transmission means communicating only the data that each process actually needs. These methods have significantly improved the computational efficiency of the distributed GMRES algorithm. However, through statistical analysis of the optimized algorithm, it is found that the communication overhead still accounts for a large proportion.
[0083] Table 1
[0084]
[0085] Currently, in distributed iterative algorithms, communication data uses IEEE 754-like floating-point formats such as FP64, FP32, BF16, and FP16. Such floating-point formats have significant storage redundancy for representing communication data in iterative algorithms. As shown in Table 1, taking the GMRES algorithm as an example, Algorithm 1 shows its iterative process. Through the inner iteration of Algorithm 1 (lines 3 - 14), it can be found that in the SpMV operation Av j (line 4), the vector v to be communicated j After normalization, its elements are all floating-point numbers with absolute values less than 1. The exponent bits of the four traditional mainstream floating-point formats are 11, 7, 7, and 5 respectively, and their representation ranges far exceed the value ranges of the vector elements during the iteration process. This means that when representing vector elements using IEEE 754, the exponent part has very limited values, and the limited floating-point exponent bits are not fully utilized in the scenario of iterative solution algorithms. Taking the sparse linear equations with coefficient matrices bmwcra_1, nd12k, and nd24k from the SuiteSparse matrix collection as examples, the inventor further analyzed their iterative solution processes. As Figure 1 shown, it shows the changes in the information entropy of the exponent, mantissa, and value of the transmitted vector v during 155 iterations. It can be observed that during the solution process of the same linear system, the information entropy of the mantissa and value of the transmitted floating-point vector elements is the same, which is much larger than the information entropy of the exponent. In the three linear systems, the information entropy of the exponent is basically distributed between 2 and 3. Further, the proportion of elements with the same exponent is statistically analyzed. As Figure 2As shown, it can be observed that in most iterations, 1, 2, 4, and 8 shared exponents can cover approximately 30%, 50%, 80%, and 96% of the non-zero elements respectively.
[0086] The above analysis shows that problems such as ineffective utilization of bit positions will occur in actual data representation with floating-point representation based on the IEEE74 standard or similar representation methods using exponent + mantissa.
[0087] To solve the above technical problems, this embodiment provides a communication data compression method based on quantization technology, as Figure 3 shown, which demonstrates a distributed iterative process using quantization and dequantization strategies. After each process obtains its own local calculation result, it quantizes and compresses the transmission vector based on the numerical characteristics of the calculation result. Then each process sends the quantized result to other processes through the network, so that each process has a complete quantized result. Finally, each process dequantizes the data it receives to obtain a vector represented in the original precision, and then performs local SpMV operations and subsequent operations.
[0088] As shown in Table 2, Algorithm 2 demonstrates the core steps of a distributed iterative algorithm using quantization optimization. In each iteration, each process first quantizes the high-precision data it wants to transmit to other processes to a fixed-point representation with a width of w through the quantization function FP2INTw, then sends it to other processes, and receives the quantized data from other processes to obtain a complete v w . Then it calls INTw2FP to dequantize it to a high-precision floating-point representation v out , and then continues with subsequent SpMV and other operations. Although the quantization before transmission and the dequantization after transmission introduce additional data conversion overheads, both the quantization and dequantization processes are performed element by element, the elements are independent of each other, and the data being operated on is all in the GPU memory. Therefore, it can be quickly completed using the thousands of computing cores of the GPU. Compared with the benefits in terms of communication, the overheads of quantization and dequantization are almost negligible.
[0089] Table 2
[0090]
[0091]
[0092] The existing floating-point format defined by the IEEE 754 standard adopts the same representation rule for all floating-point numbers and can meet the general computing requirements. In practical applications, however, the data of each application usually has unique numerical distribution characteristics, such as the number of data and the data value range. However, the existing IEEE 754-like floating-point representation is difficult to match the application numerical distribution characteristics, resulting in significant waste of bit widths. Quantization technology maps a given set of floating-point numbers into a fixed-point number interval within a specific range according to the numerical characteristics of the finite floating-point data. Compared with the floating-point format representation that divides the exponent and mantissa, the quantization representation can improve the utilization efficiency of the finite bit width, thereby enhancing the representation precision of floating-point numbers. However, the existing quantization technology is designed for artificial intelligence applications, and the representation precision of data in model training and inference is not as sensitive as that in scientific and engineering calculations. Therefore, the quantization technology cannot be directly applied to the communication optimization of distributed iterative algorithms. This embodiment considers how to quantize the communication vectors of iterative algorithms from two cases: uniform quantization and non-uniform quantization.
[0093] (1) Communication compression method based on uniform quantization;
[0094] Uniform quantization maps all floating-point numbers in the interval [a, b] into an integer interval with a fixed step size. According to the different choices of the zero point, uniform quantization can be further divided into symmetric and non-symmetric quantization. Symmetric quantization can be regarded as a special non-symmetric quantization. Non-symmetric quantization consists of three parameters: scaling factor S, zero point z, and bit width w. The scaling factor S is usually a floating-point number that determines the step size of the quantizer. Given a floating-point vector x in and an unsigned integer grid, the quantization process can be expressed as:
[0095]
[0096] where, denotes rounding the input data to the nearest integer, and the definition of the clamp function is:
[0097]
[0098] The above quantization operation can map the real number vector into an integer grid with a lower bit width, thereby achieving data compression. After inter-process communication of the compressed data, each process also needs to dequantize the received data in the following manner, and the dequantized floating-point vector is approximately equal to the initial floating-point vector.
[0099] x out = s·(x w - z)(3)
[0100] The quantization factor s is determined by the quantization range [a, b] and the quantization bit width w in the following way:
[0101]
[0102] In some quantization methods in the field of artificial intelligence, the quantization range is not necessarily an interval consisting of the maximum and minimum values of the current tensor. In this case, if the quantized result exceeds the interval [min, max], it will be directly clamped to the min or max value. This is applicable to the compression of certain data in artificial intelligence applications, because errors in a small part of the data do not necessarily affect the global performance of the model. However, for the scientific computing problems discussed in the present invention, in the iterative solution process, large errors on some components may have a greater impact on the convergence of the iterative algorithm. Therefore, the value of the quantization range of this embodiment is statistically derived from the current vector to be compressed. Specifically, before quantization, the communication vector to be compressed is statistically analyzed to obtain its maximum and minimum values, which is very efficient using multi-threaded parallel technology on mainstream accelerator GPUs.
[0103] In order to perform efficient calculations directly based on existing instruction sets on quantized data, the quantization bit width used in artificial intelligence applications is generally 4 or 8. However, in the problem considered in the present invention, it is not necessary to perform iterative calculations directly based on quantized data, but only for data synchronization and transmission between different computing nodes, so there is no need to be limited to 4 or 8 bit widths, but can be quantized to any bit width as needed. Of course, for iterative algorithms, the bit width below 8 bits indicates that the precision is too low, and the data compression effect is very limited when it exceeds 24 bits, so it is more desirable to adjust as needed near 16 bits. When the bit width is an integer multiple of 8 (i.e., byte alignment), the quantization from floating point numbers to integers is more intuitive, but when the bit width is not an integer multiple of 8 (e.g., 12 and 20), it is necessary to further fill it with bit operations during data encoding and decoding. Bit operations have a lower execution overhead in computers, so that arbitrary bit width quantization can be supported.
[0104] (2) Hybrid compression method based on non-uniform quantization;
[0105] From the uniform quantization formula, we can see that the effect of quantization is very dependent on the maximum and minimum values in the input vector. When the bit width is fixed at w, the larger the interval between the maximum and minimum values, the lower the precision loss caused by quantization. Figure 4 The data distribution of the communication floating-point vector in a certain iteration of the iterative solution of two linear equations is shown. It can be observed that the values of the elements in the vector are very unevenly distributed, with only a few elements far from 0, while the vast majority of elements are near 0. It is precisely because of the very few elements far from 0 that the maximum and minimum values are determined.
[0106] The present invention uses a quantization method based on a piecewise linear function to introduce a segmentation point in the quantization range, splitting the quantization interval into two non-overlapping regions: a dense central region and a sparse edge region. The quantization technique based on the piecewise linear function uniformly quantizes the data distributed in the two quantization intervals. By controlling the segmentation point, the quantization loss of the discrete edge region and the central dense region can be adjusted. In artificial intelligence applications, a segmentation point that increases the representation accuracy of the central dense region is usually selected. However, in the problems discussed in the present invention, the values of the central dense region and the discrete edge region are equally important, and a higher accuracy loss in either party will affect the final convergence of the iteration. For this reason, the present invention proposes two improvement points:
[0107] 1) Adaptive segmentation point setting for data distribution. In linear equations with different sparse matrices as coefficient matrices, the numerical characteristics of the communication vectors are different. Therefore, it is necessary to use different segmentation point settings. The present invention introduces a parameter p to represent the proportion of elements contained in the central dense region, and adjusts p to segment the central dense region and the edge discrete region.
[0108] 2) Hybrid compression of quantization mapping and floating-point representation. The edge discrete region contains fewer elements, and this part of the data can be not quantized and compressed, and the original FP64 or FP32 representation is still used. If the proportion of elements contained in the central dense region is α, then the proportion of elements contained in the edge discrete region is 1 - α. Assuming that 64-bit integers are used to represent the edge discrete region and the central dense region is compressed by uniform quantization with a quantization bit width of w, the communication compression ratio η achieved by quantization at this time is:
[0109]
[0110] When α = 0.9 and w = 16, the hybrid storage can still achieve a compression ratio of about 3.08 times. However, it should be noted that as α decreases, the accuracy loss caused by quantization becomes smaller, but the compression ratio will also decrease accordingly. Therefore, a trade-off needs to be considered according to the requirements.
[0111] (3) Hybrid compression of internal and external iterations;
[0112] Quantization optimization essentially belongs to a data lossy compression technique, so it will have a certain impact on the convergence of the distributed iterative algorithm. By observing Algorithm 2, it can be found that the GMRES algorithm contains two levels of iteration, both of which involve SpMV operations (lines 1 and 4), so the latest multiplication vectors need to be obtained through inter-process communication. The communication of the v vector in the inner iteration is mainly used to construct the Krylov subspace, while the SpMV calculation in the outer iteration is used to obtain the residual caused by the current numerical solution, and its multiplication vector is the approximate solution obtained in the current iteration. Using quantization methods to compress this approximate solution will cause significant accuracy loss. Therefore, in this embodiment, quantization optimization is not used for the communication vectors in the outer iteration, and quantization compression is only used for the communication vectors in the inner iteration. In fact, the number of executions of the inner iteration is much higher than that of the outer iteration, that is, the communication ratio of the inner iteration is significantly higher than that of the outer iteration. Therefore, this hybrid optimization method for the inner and outer iterations can not only significantly reduce the communication overhead, but also not cause high accuracy loss.
[0113] The existing floating-point format defined by the IEEE754 standard uses the same representation rule for all floating-point numbers and can meet the general computing requirements. In practical applications, however, the data of each application usually has unique numerical distribution characteristics. However, the existing IEEE754-like floating-point representation is difficult to match the application numerical distribution characteristics, resulting in significant waste of bit widths. The communication compression method based on quantization technology proposed in the present invention maps a given set of floating-point numbers into a fixed-point number interval within a specific range according to the numerical characteristics of the finite floating-point data. Compared with the floating-point format representation that divides the exponent and mantissa, the quantization representation can improve the utilization efficiency of the limited bit width, thereby improving the representation accuracy of the floating-point numbers. More importantly, the compressed data based on quantization technology is only used for communication and not for calculation, so it supports compression of any bit width.
[0114] The above is only the preferred specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A communication data compression method based on quantization technology, characterized in that: include: Distributed iteration is performed based on quantization and dequantization strategies. When each process obtains its own local calculation result, the transmission vector is quantized and compressed based on the numerical characteristics of the local calculation result. Each process sends the quantized result to other processes through the network, and each process obtains the complete quantized result; Based on the complete quantization result, each process dequantizes the data it receives to obtain a vector represented by the original precision, and then performs a local SpMV operation and subsequent operations on the vector represented by the original precision; The process of distributed iteration based on quantization and anti-quantization strategies includes: In each iteration, each process first transmits the high-precision representation of the data to be transmitted to other processes. Quantize it into a fixed-point representation of w bits through the quantization function FP2INTw, then send it to other processes, and receive the quantized data from other processes to obtain the complete quantization result. ; Then call INTw2FP to get the complete quantization result Dequantize to high-precision floating point representation , and then continue with subsequent SpMV and other operations; The process of quantizing and compressing the transmission vector based on the numerical characteristics of the local calculation result includes: The communication compression method based on uniform quantization maps all floating-point numbers in the interval [a, b] to an integer interval with a fixed step size. Given a floating-point vector and an unsigned integer grid, and performs quantization operations to map real vectors to low-bit-width integer grids to achieve data compression; The process of quantizing and compressing the transmission vector based on the numerical characteristics of the local calculation result includes: The communication compression method based on non-uniform quantization introduces a segmentation point in the quantization range by using the quantization technology based on piecewise linear function, splits the quantization interval into two non-overlapping areas, and obtains the central dense area and the edge discrete area; The quantization technology based on the piecewise linear function uniformly quantizes the data distributed in the two quantization intervals respectively, and the quantization loss of the edge discrete area and the center dense area is regulated by controlling the segmentation point; Based on the complete quantization result, the process of each process dequantizing the data it receives includes: After the compressed data is communicated between processes, each process dequantizes the received data, and the floating-point vector obtained by dequantization is approximately equal to the initial floating-point vector; The dequantization process is expressed as follows: The quantization factor s is determined by the quantization range [a, b] and the quantization bit width w as follows: 。 2. The communication data compression method based on quantization technology according to claim 1, characterized in that: The quantization process is expressed as: in, Indicates that the input data is rounded to the nearest integer. The definition of the clamp function is: 。 3. The communication data compression method based on quantization technology according to claim 1, characterized in that: The communication compression method based on uniform quantization has different zero point selections. Uniform quantization includes symmetric quantization and asymmetric quantization. The asymmetric quantization parameter composition includes a scaling factor S, a zero point z, and a bit width w; wherein the scaling factor S is a floating point number and is used to determine the step size of the quantizer.
4. The communication data compression method based on quantization technology according to claim 1, characterized in that: The process of introducing a segmentation point in the quantization range by using the quantization technology based on piecewise linear function and splitting the quantization interval into two non-overlapping areas includes: The parameter p is introduced to represent the proportion of elements contained in the central dense area. By adjusting the parameter p, the central dense area and the edge discrete area are segmented, and the segmentation point setting for data distribution adaptation is completed.
5. The communication data compression method based on quantization technology according to claim 1, characterized in that: The process of uniformly quantizing the data distributed in two quantization intervals based on the piecewise linear function quantization technology includes: Through the hybrid compression of quantization mapping and floating point representation, the elements in the edge discrete area are not quantized and compressed, and the original FP64 or FP32 representation is still used; Define the proportion of elements contained in the central dense area as α, then the proportion of elements contained in the edge discrete area is 1-α, assuming that a 64-bit integer is used to represent the edge discrete area, uniform quantization is used to compress the central dense area, and the quantization bit width is w, then the communication compression ratio achieved by quantization is for: = 。 6. A communication data compression system based on quantization technology, characterized in that: The method for implementing any one of claims 1 to 5 comprises: A quantization module is used to quantize and compress the transmission vector based on the numerical characteristics of the local calculation results after each process obtains its own local calculation results; each process sends the quantized results to other processes through the network, and each process obtains a complete quantization result; The dequantization module is connected to the quantization module and is used for each process to dequantize the data it receives based on the complete quantization result to obtain a vector represented by the original precision, and then perform local SpMV calculations and subsequent operations on the vector represented by the original precision.
Citation Information
Patent Citations
Multi-level quantization and adaptive adjustment method
CN115103031A