Gradient quantification method and device, equipment, storage medium and product

By dynamically allocating the number of bits of the neural network layer and allocating the number of bits to the target layer according to the quantization error, the resource waste caused by the fixed bit quantization strategy is solved, and the training efficiency of the deep learning model is improved.

CN120218136APending Publication Date: 2025-06-27PENG CHENG LAB
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510251006.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing fixed bit quantization strategy leads to wasted compression resources in deep learning and cannot effectively adapt to the differences in sensitivity of different neural network layers to compression distortions.

Method used

By obtaining the quantization errors and current bit counts of each layer in the neural network, dynamically allocate the number of bits, determine the target layer that needs bit allocation based on the quantization error, and allocate the number of bits to it within the preset bit range.

Benefits of technology

Under the overall bit budget constraint, the computational overhead of model training is significantly reduced, the efficiency of distributed training is improved, and resource waste in traditional fixed bit allocation methods is avoided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120218136A_ABST
    Figure CN120218136A_ABST
Patent Text Reader

Abstract

The invention discloses a gradient quantization method and device, equipment, a storage medium and a product, and relates to the technical field of deep learning, and the gradient quantization method comprises the steps: obtaining a quantization error and a current bit number of each neural network layer in a neural network; determining a target neural network layer needing bit allocation according to the quantization error; and allocating a bit number to the target neural network layer based on a preset bit range, the current bit number and the bit budget. The bit number is dynamically allocated according to the quantization error of each neural network layer, and the overall quantization error is minimized under the constraint of meeting the overall bit budget, so that compared with an existing fixed bit allocation mode, the mode can remarkably reduce the model training calculation overhead, and the model training efficiency is improved. And dynamic bit allocation becomes feasible in actual distributed training.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of deep learning technology, and in particular, to a gradient quantization method, apparatus, device, storage medium, and product. Background Art

[0002] With the continuous growth of the scale of deep learning models and datasets, a single computing node can no longer easily handle the entire training data or model. To solve this problem, data parallel distributed training has become the mainstream solution. In this paradigm, the data is partitioned across multiple computing nodes, and each node maintains a complete copy of the model parameters and processes different data batches. Although this method effectively improves the computing speed, during the gradient synchronization process, since each node needs to exchange gradients to maintain the consistency of the model, a large amount of communication overhead is generated.

[0003] To alleviate the communication bottleneck, researchers have proposed various gradient compression techniques. Among them, the most concerned is the gradient quantization method, which reduces the communication cost by reducing the number of bits used to represent gradient elements. Early works such as 1-bit SGD proposed using a single bit to represent gradients, SignSGD only transmits the gradient signs, QSGD introduces random quantization and provides theoretical guarantees for convergence, and TernGrad restricts the gradient values to the ternary levels {-1, 0, 1}. Although these methods significantly reduce the communication bandwidth requirements, they all have a common limitation: they adopt a fixed-bit quantization strategy, that is, the same number of bits is used for quantization for all layers of the neural network.

[0004] This unified compression strategy ignores the differences in the sensitivity of different layers of the neural network to compression distortion. It will result in layers that are sensitive to compression obtaining fewer bits, which is not enough to retain important information; at the same time, layers that are not sensitive to compression are allocated too many bits, causing waste of compression resources. These problems seriously restrict the training efficiency of distributed deep learning systems. Summary of the Invention

[0005] The main purpose of this application is to provide a gradient quantization method, apparatus, device, storage medium, and product, aiming to solve the technical problem of waste of compression resources caused by the existing fixed-bit quantization strategy.

[0006] To achieve the above object, this application proposes a gradient quantization method, and the gradient quantization method includes:

[0007] Obtain the quantization error and the current number of bits of each neural network layer in the neural network;

[0008] Determine the target neural network layer that needs to perform bit allocation according to the quantization error;

[0009] Allocate bits for the target neural network layer based on a preset bit range, the current number of bits, and the bit budget.

[0010] Optionally, the step of allocating bits for the target neural network layer based on a preset bit range, the current number of bits, and the bit budget includes:

[0011] Determine the currently occupied number of bits according to the current number of bits;

[0012] When the currently occupied number of bits is less than the bit budget, allocate bits for the target neural network layer based on the preset bit range.

[0013] Optionally, before the step of allocating bits for the target neural network layer based on a preset bit range, the current number of bits, and the bit budget, it further includes:

[0014] Obtain the historical bit data of the target neural network layer;

[0015] Determine the bit mean and bit standard deviation according to the historical bit data;

[0016] Determine the preset bit range based on the bit mean and the bit standard deviation.

[0017] Optionally, after the step of determining the target neural network layer that needs bit allocation according to the quantization error, it further includes:

[0018] Obtain the historical bit data of the target neural network layer;

[0019] Determine the coefficient of variation according to the historical bit data;

[0020] When the coefficient of variation is less than the preset variation threshold, determine the bit mode according to the historical bit data;

[0021] Allocate bits for the target neural network layer based on the bit mode.

[0022] Optionally, after the step of determining the coefficient of variation according to the historical bit data, it further includes:

[0023] When the coefficient of variation is not less than the preset variation threshold, allocate bits for the target neural network layer based on a preset bit range, the current number of bits, and the bit budget.

[0024] Optionally, the step of obtaining the quantization error of each neural network layer in the neural network further includes:

[0025] Calculate the quantization error of each neural network layer in the neural network according to the following formula:

[0026]

[0027] Among them, error(l, b l ) is used to represent the quantization error of the l-th layer, and g l is used to represent gradient data, is used to represent the reconstructed gradient, and b l is used to represent the currently allocated number of bits, and |g l | is used to represent the number of parameters of the current layer.

[0028] In addition, to achieve the above object, the present application also proposes a gradient quantization device, and the gradient quantization device includes:

[0029] An acquisition module, configured to acquire the quantization error and the current number of bits of each neural network layer in the neural network;

[0030] A determination module, configured to determine a target neural network layer that needs to perform bit allocation according to the quantization error;

[0031] A bit number allocation module, configured to allocate a bit number for the target neural network layer based on a preset bit range, the current number of bits, and a bit budget.

[0032] In addition, to achieve the above object, the present application also proposes a gradient quantization device, and the device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the computer program is configured to implement the steps of the gradient quantization method as described above.

[0033] In addition, to achieve the above object, the present application also proposes a storage medium, the storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the gradient quantization method as described above are implemented.

[0034] In addition, to achieve the above object, the present application also provides a computer program product, the computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the gradient quantization method as described above are implemented.

[0035] The present application acquires the quantization error and the current number of bits of each neural network layer in the neural network; determines a target neural network layer that needs to perform bit allocation according to the quantization error; and allocates a bit number for the target neural network layer based on a preset bit range, the current number of bits, and a bit budget. Since the present application dynamically allocates bit numbers according to the quantization error of each neural network layer and minimizes the overall quantization error under the constraint of meeting the overall bit budget, compared with the existing fixed bit allocation method, the above method of the present application can significantly reduce the computational overhead of model training and make dynamic bit allocation feasible in actual distributed training. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.

[0037] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0038] Figure 1 It is a schematic flowchart provided for the first embodiment of the gradient quantization method of the present application;

[0039] Figure 2 It is a schematic flowchart of the bit allocation provided for the first embodiment of the gradient quantization method of the present application;

[0040] Figure 3 It is a schematic flowchart provided for the second embodiment of the gradient quantization method of the present application;

[0041] Figure 4 It is a schematic diagram of the module structure of the gradient quantization device in the embodiment of the present application;

[0042] Figure 5 It is a schematic diagram of the device structure of the hardware operating environment involved in the gradient quantization method in the embodiment of the present application.

[0043] The implementation, functional features, and advantages of the present application will be further described with reference to the embodiments and the drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0045] In order to better understand the technical solutions of the present application, the following will be described in detail in combination with the drawings of the specification and the specific implementation manners.

[0046] The main solution of the embodiments of this application is: obtaining the quantization errors and the current number of bits of each neural network layer in a neural network; determining a target neural network layer that needs bit allocation according to the quantization errors; allocating the number of bits for the target neural network layer based on a preset bit range, the current number of bits, and a bit budget. Since this application dynamically allocates the number of bits according to the quantization errors of each neural network layer and minimizes the overall quantization error under the constraint of meeting the overall bit budget, compared with the existing fixed-bit allocation method, the above method of this application can significantly reduce the computational overhead of model training and make dynamic bit allocation feasible in actual distributed training.

[0047] It should be noted that the execution subject of this embodiment can be a computing service device with data processing, network communication, and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device, a gradient quantization device, etc. that can implement the above functions. Hereinafter, taking the gradient quantization device as an example, this embodiment and the following embodiments will be described.

[0048] Based on this, the embodiments of this application provide a gradient quantization method, referring to Figure 1 , Figure 1 which is a schematic flowchart provided for Embodiment 1 of the gradient quantization method of this application.

[0049] In this embodiment, the gradient quantization method includes the following steps:

[0050] Step S10, obtaining the quantization errors and the current number of bits of each neural network layer in a neural network;

[0051] It should be noted that the quantization error (Quantization Error) refers to the error introduced when converting data from a high-precision representation (such as 32-bit floating-point number) to a low-precision representation (such as 8-bit integer). The quantization error is a core issue in quantization technology and directly affects the performance and convergence of the quantized model. The current number of bits can be the number of bits currently allocated to the neural network layer.

[0052] Furthermore, in order to improve the training efficiency of the model, the step of obtaining the quantization errors of each neural network layer in the neural network further includes:

[0053] Calculating the quantization errors of each neural network layer in the neural network according to the following formula:

[0054]

[0055] where error(l,b l ) is used to represent the quantization error of the l-th layer, g l is used to represent the gradient data, is used to represent the reconstructed gradient, bl used to characterize the currently allocated number of bits, |g l | used to characterize the number of parameters of the current layer.

[0056] Step S20, determine the target neural network layer that needs to perform bit allocation according to the quantization error;

[0057] It should be noted that the target neural network layer that needs to perform bit allocation according to the quantization error may be the neural network layer with the largest quantization error, and it is judged whether the current number of bits of this neural network layer is less than the preset maximum number of bits. If the current number of bits is less than the preset maximum number of bits, then this neural network layer is used as the target neural network layer that needs to perform bit allocation.

[0058] Step S30, allocate the number of bits for the target neural network layer based on the preset bit range, the current number of bits, and the bit budget.

[0059] It should be noted that the preset bit range includes the preset minimum number of bits and the maximum number of bits. For example, the preset bit range can be 1 - 8. The bit budget can be the total number of bits that can be allocated to each neural network layer. Its calculation method can be:

[0060]

[0061] where B total used to characterize the bit budget, B C used to characterize the average bit width constraint, L is used to characterize the total number of layers of the neural network, |g l | used to characterize the number of parameters of this layer.

[0062] It should be noted that the allocation of the number of bits for the target neural network layer based on the preset bit range, the current number of bits, and the bit budget may be that when the current number of bits of the target neural network layer is within the preset bit range, and the bit budget minus the number of bits occupied by the current layers is not less than 1, allocate one more bit for the target neural network layer.

[0063] Further, step S30 may include: determining the currently occupied number of bits according to the current number of bits;

[0064] When the currently occupied number of bits is less than the bit budget, allocate the number of bits for the target neural network layer based on the preset bit range.

[0065] It should be noted that the currently occupied number of bits can be the total number of bits occupied by each neural network layer. Allocating the number of bits for the target neural network layer based on the preset bit range can be to allocate one or a number of bits not greater than the maximum value in the preset bit range when the current number of bits of the target neural network layer is less than the maximum value in the preset bit range.

[0066] In specific implementation, reference can be made to Figure 2 , Figure 2 which is a schematic diagram of the bit allocation process provided in the first embodiment of the gradient quantization method of the present application. In this embodiment, during the distributed training process, for the l-th layer in the neural network, its gradient data is represented as g l . The quantization error module first quantizes the original gradient according to the currently allocated number of bits bl to obtain the reconstructed gradient Then calculate the normalized quantization error: where |g l | represents the number of parameters of this layer. This normalization process ensures the comparability of errors between layers of different scales. The adaptive bit allocation process includes:

[0067] First, calculate the overall bit budget B total , where B C is the average bit width constraint. In the initialization stage, allocate the minimum number of bits for all layers, which can be 1, and record the used number of bits B used . In the main loop, calculate the quantization errors of all neural network layers, find the layer with the largest error and not reaching the maximum number of bits B (which can be 8), and allocate one (or multiple) available bits for it. Each allocation should ensure that the overall bit budget constraint is satisfied, and the quantization error of this layer is updated in real time. This process continues until the bit budget is reached or all layers reach the maximum number of bits. Figure 2 The b next in the formula is the number of bits after allocation, and br is the number of bits of the current neural network layer.

[0068] This embodiment obtains the quantization errors and the current number of bits of each neural network layer in the neural network; determines the target neural network layer that needs bit allocation according to the quantization error; allocates the number of bits for the target neural network layer based on the preset bit range, the current number of bits, and the bit budget. Since this embodiment dynamically allocates the number of bits according to the quantization errors of each neural network layer, and minimizes the overall quantization error under the constraint of the overall bit budget, compared with the existing fixed bit allocation method, the above method of this embodiment can significantly reduce the computational overhead of model training and make dynamic bit allocation feasible in actual distributed training.

[0069] This embodiment proposes a hierarchical adaptive gradient quantization method for gradient compression in distributed deep learning. By minimizing the overall quantization error of all layers under a given overall bit budget constraint, dynamic bit allocation is achieved. Specifically, through the standardized calculation of the quantization error, the deviation between layers of different scales is avoided, and the bit allocation is dynamically adjusted according to the quantization sensitivity of each layer, minimizing the compression error while satisfying the overall bit budget constraint. This adaptive allocation mechanism breaks through the limitations of traditional fixed bit allocation schemes and can better adapt to the differences in the sensitivity of different neural network layers to compression.

[0070] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar content as in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 3 , Figure 3 which is a schematic flowchart provided for the second embodiment of the gradient quantization method of the present application. Before the step S30, the following steps are further included:

[0071] Step S201: Obtain the historical bit data of the target neural network layer;

[0072] It should be noted that the historical bit data may be the number of bits of the target neural network layer in each previous iteration process.

[0073] Step S202: Determine the bit mean and bit standard deviation according to the historical bit data;

[0074] It should be noted that determining the bit mean and bit standard deviation according to the historical bit data may be calculating the mean and standard deviation of the number of bits of the target neural network layer in historical iterations.

[0075] Step S203: Determine a preset bit range based on the bit mean and the bit standard deviation.

[0076] It should be noted that determining the preset bit range based on the bit mean and the bit standard deviation may be when the bit standard deviation σ is less than the threshold thr1 (default value 0.5, or other values), the search range is limited to between; when σ is greater than or equal to thr1, the search range is expanded to μ is the bit mean, and this mechanism significantly reduces the search space and improves the algorithm efficiency.

[0077] In a specific implementation, this embodiment introduces a dynamic search range optimization mechanism. This mechanism sets a sliding window with a length of n (default value 50) to record the bit allocation results of the nearest n iterations for each layer. By calculating the mean μ and standard deviation σ of these historical data, the bit search range can be dynamically determined. When σ is less than the threshold thr1 (default value 0.5), the search range is restricted between ; when σ is greater than or equal to thr1, the search range is expanded to This mechanism significantly reduces the search space and improves the algorithm efficiency.

[0078] Furthermore, in order to further improve the efficiency of deep learning, after the step S20, the following steps are also included:

[0079] Obtain the historical bit data of the target neural network layer;

[0080] Determine the coefficient of variation according to the historical bit data;

[0081] In the case where the coefficient of variation is less than the preset variation threshold, determine the bit mode according to the historical bit data;

[0082] Allocate the number of bits for the target neural network layer based on the bit mode.

[0083] It should be noted that the coefficient of variation can be calculated based on the bit mean μ and bit standard deviation σ. Specifically, the coefficient of variation The coefficient of variation CV is used to evaluate the stability of bit allocation. The preset variation threshold can be a pre-set threshold, which can be a value such as 0.2. Determining the bit mode according to the historical bit data can be to determine the mode of the number of bits of the neural network layer in the historical bit data. Using the bit mode as the number of bits of the target neural network layer. In the case where the coefficient of variation is greater than or equal to the preset variation threshold, the number of bits for the target neural network layer is still allocated based on the preset bit range, the current number of bits, and the bit budget.

[0084] In a specific implementation, this embodiment designs a bit pre-allocation mechanism. By calculating the coefficient of variation to evaluate the stability of bit allocation. This mechanism also sets a sliding window with a length of n (default value 50) to record the bit allocation results of the nearest n iterations for each layer. When CV is less than the threshold thr2 (default value 0.2), directly use the mode of the historical data within the sliding window as the predicted value, that is, the number of bits allocated for the target neural network layer, and completely skip the search process. When CV is greater than or equal to thr2, then start the bit allocation process in Embodiment 1 or the search process optimized based on the dynamic search range optimization mechanism. This prediction mechanism greatly reduces the computational overhead.

[0085] Furthermore, in order to minimize the additional time overhead in this embodiment, a parallel processing mechanism is also designed. This mechanism decouples the bit allocation process from the model training process, achieving efficient cooperation between the CPU and the GPU. Specifically, when the training iteration t is in progress, the GPU is responsible for regular model inference and gradient calculation, while the CPU concurrently executes the bit allocation calculation for part of iteration t + 1 (including the bit allocation process in Embodiment 1, the search process optimized based on the dynamic search range optimization mechanism, and the bit pre-allocation mechanism). This pipeline design ensures that most of the bit allocation calculations overlap with the training process, minimizing the additional time overhead.

[0086] This embodiment obtains the historical bit data of the target neural network layer; determines the bit mean and bit standard deviation according to the historical bit data; and determines the preset bit range based on the bit mean and the bit standard deviation. This significantly reduces the search space and improves the algorithm efficiency.

[0087] This embodiment also proposes two key techniques for accelerating bit allocation: bit search range optimization and bit pre-allocation mechanism. The bit search range optimization technique first calculates the average value and variance of bit allocation for each layer within a set historical window, and optimizes the bit search range for each layer in each iteration through these two statistics. The bit pre-allocation mechanism evaluates the stability of bit allocation by calculating the coefficient of variation CV. When CV < thr2, it directly uses the mode of the historical data as the predicted value, completely skipping the search process. The combination of these two techniques significantly reduces the computational complexity of the algorithm, making dynamic bit allocation feasible in actual training.

[0088] This embodiment also designs an efficient CPU-GPU parallel processing mechanism, realizing the decoupling of bit allocation calculation and model training. The core idea of this mechanism is to use the predicted bit allocation scheme in the current training iteration, while using the CPU to concurrently calculate the bit search for the next iteration. Through this pipeline design, most of the computational overhead of bit allocation overlaps with the training process of deep learning, thereby minimizing the additional time overhead.

[0089] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the gradient quantization method of this application. Based on this technical concept, more forms of simple transformation are within the protection scope of this application.

[0090] This application also provides a gradient quantization device. Please refer to Figure 4 , the gradient quantization device includes:

[0091] An acquisition module 10, configured to acquire the quantization error and the current number of bits of each neural network layer in the neural network;

[0092] A determination module 20, configured to determine a target neural network layer for which bit allocation is required according to the quantization error;

[0093] A bit number allocation module 30, configured to allocate bit numbers for the target neural network layer based on a preset bit range, the current bit number, and a bit budget.

[0094] In this embodiment, the quantization errors and current bit numbers of each neural network layer in the neural network are obtained; a target neural network layer for which bit allocation is required is determined according to the quantization error; and bit numbers are allocated for the target neural network layer based on a preset bit range, the current bit number, and a bit budget. Since this embodiment dynamically allocates bit numbers according to the quantization errors of each neural network layer and minimizes the overall quantization error under the constraint of meeting the overall bit budget, compared with the existing fixed bit allocation method, the above method in this embodiment can significantly reduce the computational overhead of model training and make dynamic bit allocation feasible in actual distributed training.

[0095] The gradient quantization device provided in this application adopts the gradient quantization method in the above embodiment and can solve the technical problem of waste of compression resources caused by the existing fixed bit quantization strategy. Compared with the prior art, the beneficial effects of the gradient quantization device provided in this application are the same as those of the gradient quantization method provided in the above embodiment, and other technical features in the gradient quantization device are the same as those disclosed in the method of the above embodiment, and will not be elaborated here.

[0096] This application provides a gradient quantization device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the gradient quantization method in the first embodiment above.

[0097] Next, refer to Figure 5 , which shows a schematic structural diagram of a gradient quantization device suitable for implementing the embodiments of this application. The gradient quantization device in the embodiments of this application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle terminals (such as vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The gradient quantization device shown is only an example and should not impose any limitation on the functions and usage scopes of the embodiments of this application.

[0098] As shown Figure 5 in the figure, the gradient quantization device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which may perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 into a random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the gradient quantization device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems may be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 may allow the gradient quantization device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a gradient quantization device having various systems, it should be understood that it is not required to implement or have all the shown systems. More or fewer systems may be alternatively implemented or had.

[0099] Specifically, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program codes for performing the methods shown in the flowcharts. In such an embodiment, the computer program may be downloaded and installed from a network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above functions defined in the methods of the embodiments disclosed in the present application are executed.

[0100] The gradient quantization device provided by the present application adopts the gradient quantization method in the above embodiments, and can solve the technical problem of waste of compression resources caused by the existing fixed-bit quantization strategy. Compared with the prior art, the beneficial effects of the gradient quantization device provided by the present application are the same as those of the gradient quantization method provided by the above embodiments, and other technical features in the gradient quantization device are the same as those disclosed in the method of the previous embodiment, and will not be elaborated here.

[0101] It should be understood that each part disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.

[0102] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art within the technical scope disclosed in this application can easily think of changes or substitutions, which should all be covered within the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.

[0103] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the gradient quantization method in the above embodiments.

[0104] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0105] The above computer-readable storage medium can be included in the gradient quantization device; it can also exist separately without being assembled into the gradient quantization device.

[0106] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by connecting through the Internet using an Internet service provider).

[0107] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0108] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation to the unit itself in some cases.

[0109] The readable storage medium provided by this application is a computer-readable storage medium. The computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned gradient quantization method, which can solve the technical problem of waste of compression resources caused by the existing fixed-bit quantization strategy. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the gradient quantization method provided by the above embodiments, and will not be elaborated here.

[0110] The present application also provides a computer program product, including a computer program, which implements the steps of the gradient quantization method as described above when executed by a processor.

[0111] The computer program product provided by the present application can solve the technical problem of waste of compression resources caused by the existing fixed-bit quantization strategy. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the gradient quantization method provided by the above embodiments, and will not be elaborated here.

[0112] The above are only partial embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.

Claims

1. A gradient quantization method, characterized in that: The gradient quantization method comprises the following steps: Get the quantization error and current number of bits of each neural network layer in the neural network; Determining a target neural network layer that needs bit allocation according to the quantization error; Allocate bits to the target neural network layer based on a preset bit range, the current bit number and a bit budget.

2. The gradient quantization method according to claim 1, characterized in that: The step of allocating bits to the target neural network layer based on the preset bit range, the current bit number and the bit budget comprises: Determine the currently occupied number of bits according to the current number of bits; When the currently occupied number of bits is less than the bit budget, bits are allocated to the target neural network layer based on the preset bit range.

3. The gradient quantization method according to claim 1, characterized in that: Before the step of allocating the number of bits for the target neural network layer based on the preset bit range, the current number of bits and the bit budget, the step further includes: Get the historical bit data of the target neural network layer; Determine a bit mean and a bit standard deviation based on the historical bit data; A preset bit range is determined based on the bit mean and the bit standard deviation.

4. The gradient quantization method according to claim 1, characterized in that: After the step of determining the target neural network layer that needs to be bit allocated according to the quantization error, the method further includes: Get the historical bit data of the target neural network layer; determining a coefficient of variation based on the historical bit data; When the coefficient of variation is less than a preset variation threshold, determining a bit mode according to the historical bit data; A number of bits is allocated to the target neural network layer based on the bit mode.

5. The gradient quantization method according to claim 4, characterized in that: After the step of determining the coefficient of variation according to the historical bit data, the method further includes: When the coefficient of variation is not less than a preset variation threshold, the number of bits is allocated to the target neural network layer based on a preset bit range, the current number of bits and a bit budget.

6. The gradient quantization method according to any one of claims 1 to 5, characterized in that: The step of obtaining the quantization error of each neural network layer in the neural network also includes: The quantization error of each neural network layer in the neural network is calculated according to the following formula: Among them, error(l,b1) is used to represent the quantization error of the lth layer, g1 is used to represent the gradient data, It is used to characterize the reconstruction gradient, b1 is used to characterize the number of bits currently allocated, and |g1| is used to characterize the number of parameters of the current layer.

7. A gradient quantization device, characterized in that: The gradient quantization device comprises: An acquisition module, used to obtain the quantization error and current bit number of each neural network layer in the neural network; A determination module, used to determine a target neural network layer that needs to be bit allocated according to the quantization error; A bit number allocation module is used to allocate bits to the target neural network layer based on a preset bit range, the current bit number and a bit budget.

8. A gradient quantization device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the gradient quantization method according to any one of claims 1 to 6.

9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the gradient quantization method according to any one of claims 1 to 6 are implemented.

10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the gradient quantization method according to any one of claims 1 to 6 are implemented.