Neural network quantification method, electronic equipment and storage medium
By building an adversarial generation network, using preset loss function to train the generation network, generate training data for neural network quantization, and quantify the pretrained neural network, solving the neural network quantization problem in data-free scenarios, and achieving efficient neural network quantization effect.
Patent Information
- Application Number
- CN202411930365.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-25
- Publication Date
- 2025-05-16
AI Technical Summary
In the data-free scenario, existing neural network quantization methods are difficult to effectively solve the problem of neural network quantization.
A adversarial generation network is constructed, including a generator and a discriminator, where the generator is an auxiliary classification generation generative network, and the discriminator includes a pre-trained neural network and a corresponding initialized quantized neural network. The generated network is trained by a preset loss function, training data for neural network quantization is generated, and the pre-trained neural network is quantified.
In the data-free scenario, through adversarial generation network generation, neural network quantization can be effectively quantified, and the instructions and diversity of generated data can be improved, thereby improving the effect of neural network quantization.
Smart Images

Figure CN120012839A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of neural network quantization, and more specifically, to a neural network quantization method, an electronic device, and a storage medium. Background Art
[0002] Most of the existing neural network quantization methods are based on data dependence. However, in actual application scenarios, data may not be available due to security, user privacy, cost, etc. Therefore, neural network quantization based on data-free scenarios is a problem that needs to be solved in current actual application scenarios. Summary of the invention
[0003] One purpose of the present disclosure is to provide a neural network quantization method to solve the neural network quantization problem in a data-free scenario.
[0004] According to a first aspect of the present disclosure, a method for quantizing a neural network is provided, comprising:
[0005] Constructing a generative adversarial network, the generative adversarial network comprising a generator and a discriminator, wherein the generator comprises a generative network in an auxiliary classification generative adversarial network, and the discriminator comprises a pre-trained neural network and an initialized quantized neural network corresponding to the pre-trained neural network;
[0006] Training the generative adversarial network by using a preset loss function, and generating training data for neural network quantization based on the trained generative adversarial network;
[0007] The pre-trained neural network is quantized using the training data.
[0008] Optionally, the preset loss function includes a first loss function, a second loss function, a third loss function and a fourth loss function;
[0009] The first loss function is used to make the distribution of the training data generated by the generating network close to the distribution of the original training data;
[0010] The second loss function is used to improve the adaptability of generated data;
[0011] The third loss function is used to limit the adaptability of the generated data to avoid the situation where the adaptability is too high or too low;
[0012] The fourth loss function is used to enable the generation network to generate data that is more beneficial for reconstruction based on block-level output.
[0013] Optionally, the first loss function L BNS It is expressed by the following formula:
[0014]
[0015] Among them, L BN is the set of all batch normalization layers in the pre-trained neural network, and They represent the mean and standard deviation of the output dynamically calculated before the i-th batch normalization layer when the training data generated by the generator network is forward propagated in the pre-trained neural network. and They respectively represent the mean and standard deviation of the statistics of the i-th batch normalization layer during the training process.
[0016] Optionally, the second loss function L SG It is expressed by the following formula:
[0017] L SG =L as +L ds
[0018] Among them, L as and L ds They represent inconsistency loss and consistency loss respectively, and their formulas are as follows:
[0019]
[0020] in, is the Gaussian noise input to the generator network, and onehot(y) is the generator network based on The one-hot expression of the label y corresponding to the generated training data, p ds Vector and p as The vector can represent the probability distribution of inconsistency and consistency in all categories, and its formula is as follows:
[0021] p ds =Softmax(y M -y Q )
[0022] p as =Softmax(y M +y Q )
[0023] Among them, y M and Q Respectively based on The generated training data is used to train the neural network and quantize the prediction results of the neural network inference.
[0024] Optionally, the third loss function L margin It is expressed by the following formula:
[0025]
[0026] Among them, λ i , u represents the lower and upper bounds of the normalized information entropy, and 0≤λ i <λ u ≤1,H ′ (p ds ) and H ′ (p ds ) represent the normalized information entropy of the probability distribution of inconsistency and consistency respectively;
[0027] Among them, the normalized information entropy of the probability distribution of inconsistency and consistency is determined by the following formula:
[0028]
[0029] Among them, category c belongs to set C, p(c) represents the probability corresponding to category c in probability distribution p, H info (p) represents the information entropy of probability distribution p.
[0030] Optionally, the fourth loss function L adv It is expressed by the following formula:
[0031]
[0032] Among them, B M,i and B Q,i Respectively represent the corresponding blocks in the pre-trained neural network and the initialized quantized neural network, and Respectively represent the training data generated by the generating network in the pre-trained neural network and the initialized quantized neural network forward propagation through B M,i and B Q,i Output.
[0033] Optionally, the pre-trained neural network is quantized based on the following quantization function formula:
[0034]
[0035] Among them, x q is the parameter of the neural network after quantization, x is the parameter of the full-precision neural network before quantization, s represents the scaling factor, z represents the zero point, b represents the quantization bit width, V is a learnable parameter used to guide the rounding direction of the parameters in the block, and σ is a continuous monotonically increasing function with a value range between [0,1].
[0036] Optionally, in the process of quantizing the pre-trained neural network, the neural network is quantized by an enhanced loss function, wherein the enhanced loss function is used to enhance local information of the training data, and the enhanced loss function includes an enhanced first loss function, the second loss function, the third loss function and an enhanced fourth loss function;
[0037] The enhanced first loss function is expressed by the following formula:
[0038]
[0039] The enhanced fourth loss function is expressed by the following formula:
[0040]
[0041] Among them, B Qj and B Mj Respectively represent the block B currently being output for reconstruction Qj and the corresponding block B in the pre-trained neural network Mj , and Respectively represent that the training data is forward propagated through B in the pre-trained neural network and the initialized quantized neural network M,i and B Q,i The output of , α is the weight parameter used to enhance the loss, where α>1.
[0042] According to a second aspect of the present disclosure, an electronic device is provided, including a processor and a memory, wherein the memory stores computer instructions, and when the computer instructions are executed by the processor, the steps of any one of the methods described in the first aspect are implemented.
[0043] According to a third aspect of the present disclosure, there is provided a storage medium on which computer instructions are stored, and when the computer instructions are executed by a processor, the steps of any one of the methods described in the first aspect are implemented.
[0044] One technical effect of the present disclosure is that a neural network quantization method is provided, which can first construct an adversarial generative network, including a generator and a discriminator, wherein the generator is a generative network in an auxiliary classification generative adversarial network, and the discriminator includes a pre-trained neural network and an initialized quantized neural network corresponding to the pre-trained neural network. The generative network is trained based on a preset loss function, and training data is generated based on the trained generative network. The pre-trained neural network is then quantized using the training data. In this way, the present application can train the adversarial generative network in a data-free scenario. The discriminator of the adversarial generative network in this embodiment can determine whether the generated data belongs to the correct category based on the prior classification label. In this way, the instructions and diversity of the generated data of the adversarial network are improved to improve the effect of neural network quantization.
[0045] Other features and advantages of the embodiments of the present disclosure will become apparent from the following detailed description of exemplary embodiments of the present disclosure with reference to the accompanying drawings. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] The accompanying drawings, which constitute a part of the specification, illustrate embodiments of the present disclosure and, together with the description, serve to explain the principles of the embodiments of the present disclosure.
[0047] Figure 1 is a flow chart of a neural network quantization method according to one embodiment;
[0048] Figure 2 is a schematic structural diagram of an electronic device according to an embodiment; DETAILED DESCRIPTION
[0049] Various exemplary embodiments of the present disclosure will now be described in detail with reference to the accompanying drawings. It should be noted that the relative arrangement of components and steps, numerical expressions and numerical values set forth in these embodiments do not limit the scope of the present invention unless otherwise specifically stated.
[0050] The following description of at least one exemplary embodiment is merely illustrative in nature and is in no way intended to limit the invention, its application, or uses.
[0051] Techniques and equipment known to ordinary technicians in the relevant art may not be discussed in detail, but where appropriate, the techniques and equipment should be considered part of the specification.
[0052] In all examples shown and discussed herein, any specific values should be interpreted as merely exemplary and not limiting. Therefore, other examples of the exemplary embodiments may have different values.
[0053] It should be noted that like reference numerals and letters refer to similar items in the following figures, and therefore, once an item is defined in one figure, it need not be further discussed in subsequent figures.
[0054] The present application embodiment discloses a method for neural network quantization, such as Figure 1 As shown, it includes step S11 to step S12.
[0055] Step S11, constructing a generative adversarial network, the generative adversarial network includes a generator and a discriminator, wherein the generator includes a generative network in an auxiliary classification generative adversarial network, and the discriminator includes a pre-trained neural network and an initialized quantized neural network corresponding to the pre-trained neural network.
[0056] The adversarial generative network usually has at least one discriminator and a generator, which are used to optimize themselves by adversarial means. The generator strives to generate more realistic fake data to deceive the discriminator, while the discriminator strives to improve its discrimination ability. In this embodiment, the discriminator uses a pre-trained neural network and its corresponding initialized quantized neural network. The pre-trained neural network is the neural network model to be quantized. The initialized quantized neural network corresponds to the pre-trained neural network, and the blocks in the two neural network models are also in a one-to-one correspondence.
[0057] The discriminant network is used to determine whether the generated data has the same characteristics as the real data, while the pre-trained neural network M that replaces the discriminant network and the initialized quantized neural network can determine whether the generated forged data belongs to the correct category based on the prior classification label, thereby realizing the function of the discriminant network.
[0058] In one example of this embodiment, the generation network in the optimized auxiliary classification generation adversarial network can be used as the generation network of the present invention. The Gaussian noise input to the generation network The training data is obtained through a fully connected mapping, two "interpolation upsampling-convolution-activation" modules and convolution operations.
[0059] In one example, the batch normalization layer of the generator network is replaced with the conditional classification batch normalization layer (Categorical Conditional Batch Normalization, CCBN) used in the Spectral Normalization GAN (SNGAN).
[0060] Step S12: training the generative adversarial network based on a preset loss function, and generating training data for neural network quantization based on the trained generative adversarial network.
[0061] In an example of this embodiment, the preset loss function includes a first loss function, a second loss function, a third loss function and a fourth loss function; the first loss function is used to make the distribution of the training data generated by the generation network close to the distribution of the original training data; the second loss function is used to improve the adaptability of the generated data; the third loss function is used to limit the adaptability of the generated data to avoid the situation where the adaptability is too high or too low; the fourth loss function is used to make the generation network generate data that is more beneficial to reconstruction based on block-level output.
[0062] In this example, a loss function composed of batch normalization information, adaptive perception, adversarial exploration learning and other losses is pre-constructed. The preset loss function T G , the formula is as follows:
[0063] L G =L BNS +L SG +L margin +L adv
[0064] Among them, the first loss function L BNS The purpose is to make the distribution of the fake data generated by the generative network G closer to the distribution of the original training data. The second loss function L SG The purpose is to balance the data generation learning based on adaptability perception, so that the generated data has more ideal adaptability. The third loss function L margin It is a hinge loss for the adaptability boundary, which aims to limit the adaptability of the generated data to avoid situations where the adaptability is too high or too low. adv The block-based adversarial learning is the target loss for exploration, which aims to train the generative network G to generate data that is more beneficial for reconstruction based on block-level output.
[0065] In one example of this embodiment, the first loss function L BNS It is expressed by the following formula:
[0066]
[0067] Among them, L BN is the set of all batch normalization layers in the pre-trained neural network, and They represent the mean and standard deviation of the output dynamically calculated before the i-th batch normalization layer when the training data generated by the generator network is forward propagated in the pre-trained neural network. and They respectively represent the mean and standard deviation of the statistics of the i-th batch normalization layer during the training process.
[0068] In this example, optimizing the loss during the training process of the generative network can gradually reduce the difference between the input of each batch normalization layer of the pre-trained neural network and the first-order and second-order distribution information obtained by statistics of the training set data when the network was previously trained during the forward propagation of the data generated by the generative network. This allows the generative network to extract information about the distribution of the original training set from the pre-trained neural network and transfer it to the generated fake data, thereby improving the performance of the generative network.
[0069] In one example of this embodiment, the second loss function L SG It is expressed by the following formula:
[0070] L SG =L as +L ds
[0071] Among them, L as and L ds They represent inconsistency loss and consistency loss respectively, and their formulas are as follows:
[0072]
[0073] in, is the Gaussian noise of the input generation network, and onehot(y) is the generation network based on The one-hot expression of the label y corresponding to the generated training data, p ds Vector and p as The vector can represent the probability distribution of inconsistency and consistency in all categories, and its formula is as follows:
[0074] p ds =Softmax(y M -y Q )
[0075] p as =Softmax(y M +y Q )
[0076] Among them, y M and Q Respectively based on The generated training data is used to train the neural network and quantize the prediction results of the neural network inference.
[0077] In this case, L SGThe purpose of is to balance the data generation learning based on adaptability perception, so that the generated data has a more ideal adaptability. It consists of two parts: consistency loss and inconsistency loss. onehot(y) is the one-hot expression of the generated training data corresponding to the label y, that is, the probability of one item is close to 1 and the probability of other items is close to 0. y M and Q Respectively based on The generated training data is used to obtain prediction results through training neural network and quantized neural network reasoning, and their corresponding probability distribution can be calculated through the softmax function.
[0078] In one example of this embodiment, the third loss function L margin It is expressed by the following formula:
[0079]
[0080] Among them, λ i , u represents the lower and upper bounds of the normalized information entropy, and 0≤λ i <λ u ≤1,H ′ (p ds ) and H ′ (p ds ) represent the normalized information entropy of the probability distribution of inconsistency and consistency respectively;
[0081] Among them, the normalized information entropy of the probability distribution of inconsistency and consistency is determined by the following formula:
[0082]
[0083] Among them, category c belongs to set C, p(c) represents the probability corresponding to category c in probability distribution p, H info (p) represents the information entropy of probability distribution p.
[0084] L margin It is a hinge loss for the adaptability boundary. Its purpose is to limit the adaptability of the generated data to avoid situations where the adaptability is too high or too low. The upper and lower bounds of the normalized information entropy constrain the upper and lower bounds of the adaptability of the training data during the learning process of the generative network, thereby ensuring that the adaptability remains within a stable range.
[0085] In this example, p(c) represents the probability corresponding to the category c∈C in the probability distribution p, H info (p) represents the information entropy of the probability distribution p. The normalized information entropy is positively correlated with the adaptability of the generated data to the quantized neural network and can be used as the target loss function for optimizing the generated network.
[0086] In one example of this embodiment, the fourth loss function L adv It is expressed by the following formula:
[0087]
[0088] Among them, B M,i and B Q,i They represent the corresponding blocks in the pre-trained neural network and the initialized quantized neural network, respectively. and They represent the training data generated by the generative network in the pre-trained neural network and the initialized quantized neural network forward propagation through B M,i and B Q,i Output.
[0089] L adv The target loss for exploring block-level adversarial learning is to train the generative network to generate data that is more beneficial for reconstructing block-level outputs.
[0090] In an example of this embodiment, the quantization of the pre-trained neural network is performed based on the following quantization function formula:
[0091]
[0092] Among them, x q is the parameter of the neural network after quantization, x is the parameter of the full-precision neural network before quantization, s represents the scaling factor, z represents the zero point, b represents the quantization bit width, V is a learnable parameter used to guide the rounding direction of the parameters in the block, and σ is a continuous monotonically increasing function with a value range between [0,1].
[0093] Unlike the traditional linear uniform quantization function, in order to take into account the correlation between weight parameters between different levels in the block-level granularity, the post-training quantization method based on block-granularity output reconstruction modifies the original level local loss and designs it as the mean square error of the output before and after the block-granularity quantization. At the same time, the post-training quantization method based on block-granularity output reconstruction introduces a set of learnable parameters V to guide the rounding direction of the parameters within the block. Among them, clamp represents the clamping operation, and floor represents the flooring operation. By using the mean square error of the output before and after the block-granularity quantization as the objective function to optimize the learnable parameter V, the weight parameters can learn the adaptive rounding direction during the calibration process, thereby achieving better quantized neural network performance.
[0094] In an example of this embodiment, in the process of quantizing the pre-trained neural network, the neural network is quantized by an enhanced loss function, wherein the enhanced loss function is used to enhance local information of the training data, and the enhanced loss function includes an enhanced first loss function, a second loss function, a third loss function, and an enhanced fourth loss function;
[0095] The enhanced first loss function is expressed as follows:
[0096]
[0097] The enhanced fourth loss function is expressed by the following formula:
[0098]
[0099] Among them, B Qj and B Mj Respectively represent the block B currently being output for reconstruction Qj and its corresponding block B in the pre-trained neural network Mj , and Respectively represent the forward propagation of training data in the pre-trained neural network and the initialized quantized neural network through B M,i and B Q,i The output of , α is the weight parameter used to enhance the loss, where α>1.
[0100] In this embodiment, after the generated network has been trained and optimized according to the algorithm, the set of modules of all block-level granularity in the pre-trained neural network and the initialized quantized neural network is respectively B M and B Q , the set of all batch normalization layers contained in the pre-trained neural network is L BN , then for the initialized quantized neural network currently outputting the reconstructed block B Qj ∈B Q and the block B at the corresponding position in the pre-trained neural network Mj ∈B M , L defined in this formula G,enhanced The block B that is currently being reconstructed is output Q and the corresponding B M The weight of the loss term in the optimization target loss is increased from 1 to α (α>1), while the weight of the loss term for other modules outside the current block remains unchanged. Under this condition, the generative network can provide more local information from each block and can generate enhanced data samples that achieve better block-level granularity output reconstruction.
[0101] like Figure 2As shown, an embodiment of the present application further provides an electronic device 200, including a processor 201 and a memory 202, wherein the memory 202 stores computer instructions, and when the computer instructions are executed by the processor 201, the steps of any method in the embodiment of the neural network quantization method are implemented.
[0102] The embodiment of the present application also provides a storage medium on which computer instructions are stored. When the computer instructions are executed by a processor, any one of the above-mentioned neural network quantization embodiments is implemented and the same technical effect can be achieved. To avoid repetition, they will not be described here.
[0103] Each embodiment in the present disclosure is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and equipment embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0104] The above describes specific embodiments of the present disclosure. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0105] The embodiments of the present disclosure may be systems, methods and / or computer program products. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the embodiments of the present disclosure.
[0106] A computer-readable storage medium may be a tangible device that can hold and store instructions used by an instruction execution device. A computer-readable storage medium may be, for example, but not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples of computer-readable storage media (a non-exhaustive list) include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disk read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, such as a punch card or a raised structure in a groove on which instructions are stored, and any suitable combination of the foregoing. As used herein, a computer-readable storage medium is not to be interpreted as a transient signal per se, such as a radio wave or other freely propagating electromagnetic wave, an electromagnetic wave propagating through a waveguide or other transmission medium (e.g., a light pulse through a fiber optic cable), or an electrical signal transmitted through a wire.
[0107] The computer-readable program instructions described herein can be downloaded from a computer-readable storage medium to each computing / processing device, or downloaded to an external computer or external storage device via a network, such as the Internet, a local area network, a wide area network, and / or a wireless network. The network can include copper transmission cables, optical fiber transmissions, wireless transmissions, routers, firewalls, switches, gateway computers, and / or edge servers. The network adapter card or network interface in each computing / processing device receives the computer-readable program instructions from the network and forwards the computer-readable program instructions for storage in the computer-readable storage medium in each computing / processing device.
[0108] The computer program instructions for performing the operation of the embodiments of the present disclosure may be assembly instructions, instruction set architecture (ISA) instructions, machine instructions, machine-related instructions, microcode, firmware instructions, state setting data, or source code or object code written in any combination of one or more programming languages, including object-oriented programming languages such as Smalltalk, C++, etc., and conventional procedural programming languages such as "C" language or similar programming languages. Computer-readable program instructions may be executed completely on a user's computer, partially on a user's computer, as an independent software package, partially on a user's computer, partially on a remote computer, or completely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer via any type of network including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., using an Internet service provider to connect via the Internet). In some embodiments, an electronic circuit, such as a programmable logic circuit, a field programmable gate array (FPGA), or a programmable logic array (PLA), may be personalized by utilizing the state information of a computer-readable program instruction, and the electronic circuit may execute a computer-readable program instruction, thereby realizing various aspects of the embodiments of the present disclosure.
[0109] Various aspects of the embodiments of the present disclosure are described herein with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer-readable program instructions.
[0110] These computer-readable program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, thereby producing a machine, so that when these instructions are executed by the processor of the computer or other programmable data processing device, a device that implements the functions / actions specified in one or more boxes in the flowchart and / or block diagram is generated. These computer-readable program instructions can also be stored in a computer-readable storage medium, and these instructions cause the computer, programmable data processing device, and / or other equipment to work in a specific manner, so that the computer-readable medium storing the instructions includes a manufactured product, which includes instructions for implementing various aspects of the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0111] Computer-readable program instructions may also be loaded onto a computer, other programmable data processing apparatus, or other device so that a series of operating steps are performed on the computer, other programmable data processing apparatus, or other device to produce a computer-implemented process, thereby causing the instructions executed on the computer, other programmable data processing apparatus, or other device to implement the functions / actions specified in one or more boxes in the flowchart and / or block diagram.
[0112] The flowcharts and block diagrams in the accompanying drawings show the possible architecture, functions and operations of the systems, methods and computer program products according to multiple embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of an instruction, and a part of a module, a program segment or an instruction contains one or more executable instructions for realizing the specified logical function. In some alternative implementations, the functions marked in the box can also occur in a different order from the order marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or the flowchart, and the combination of the boxes in the block diagram and / or the flowchart can be implemented by a dedicated hardware-based system that performs the specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions. It is well known to those skilled in the art that it is equivalent to implement it by hardware, implement it by software, and implement it by combining software and hardware.
[0113] The embodiments of the present disclosure have been described above, and the above description is exemplary, not exhaustive, and is not limited to the disclosed embodiments. Many modifications and changes will be apparent to those of ordinary skill in the art without departing from the scope of the described embodiments. The selection of terms used herein is intended to best explain the principles of the embodiments, practical applications, or improvements to the technology in the market, or to enable other persons of ordinary skill in the art to understand the embodiments disclosed herein.
Claims
1. A neural network quantization method, characterized in that: include: Constructing a generative adversarial network, the generative adversarial network comprising a generator and a discriminator, wherein the generator comprises a generative network in an auxiliary classification generative adversarial network, and the discriminator comprises a pre-trained neural network and an initialized quantized neural network corresponding to the pre-trained neural network; Training the generative adversarial network by using a preset loss function, and generating training data for neural network quantization based on the trained generative adversarial network; The pre-trained neural network is quantized using the training data.
2. The method according to claim 1, characterized in that The preset loss function includes a first loss function, a second loss function, a third loss function and a fourth loss function; The first loss function is used to make the distribution of the training data generated by the generating network close to the distribution of the original training data; The second loss function is used to improve the adaptability of generated data; The third loss function is used to limit the adaptability of the generated data to avoid the situation where the adaptability is too high or too low; The fourth loss function is used to enable the generation network to generate data that is more beneficial for reconstruction based on block-level output.
3. The method according to claim 2, characterized in that The first loss function L BNS It is expressed by the following formula: Among them, L BN is the set of all batch normalization layers in the pre-trained neural network, and They represent the mean and standard deviation of the output dynamically calculated before the i-th batch normalization layer when the training data generated by the generator network is forward propagated in the pre-trained neural network. and They respectively represent the mean and standard deviation of the statistics of the i-th batch normalization layer during the training process.
4. The method according to claim 3, characterized in that The second loss function L SG It is expressed by the following formula: L SG =L as +L ds Among them, L as and L ds They represent inconsistency loss and consistency loss respectively, and their formulas are as follows: in, is the Gaussian noise input to the generator network, and onehot(y) is the generator network based on The one-hot expression of the label y corresponding to the generated training data, p ds Vector and p as The vector can represent the probability distribution of inconsistency and consistency in all categories, and its formula is as follows: p ds =Softmax(and M -and Q ) p as =Softmax(and M +y Q ) Among them, y M and Q Respectively based on The generated training data is used to train the neural network and quantize the prediction results of the neural network inference.
5. The method according to claim 4, characterized in that The third loss function L margin It is expressed by the following formula: Among them, λ i , u represents the lower and upper bounds of the normalized information entropy, and 0≤λ i <λ u ≤1, H′(p ds ) and H′(p ds ) represent the normalized information entropy of the probability distribution of inconsistency and consistency respectively; Among them, the normalized information entropy of the probability distribution of inconsistency and consistency is determined by the following formula: Among them, category c belongs to set C, p(c) represents the probability corresponding to category c in probability distribution p, H info (p) represents the information entropy of probability distribution p.
6. The method according to claim 5, characterized in that The fourth loss function L adv It is expressed by the following formula: Among them, B M,i and B Q,i Respectively represent the corresponding blocks in the pre-trained neural network and the initialized quantized neural network, and Respectively represent the training data generated by the generating network in the pre-trained neural network and the initialized quantized neural network forward propagation through B M,i and B Q,i Output.
7. The method according to claim 6, characterized in that The pre-trained neural network is quantized based on the following quantization function formula: Among them, x q is the parameter of the neural network after quantization, x is the parameter of the full-precision neural network before quantization, s represents the scaling factor, z represents the zero point, b represents the quantization bit width, V is a learnable parameter used to guide the rounding direction of the parameters in the block, and σ is a continuous monotonically increasing function with a value range between [0, 1].
8. The method according to claim 7, characterized in that In the process of quantizing the pre-trained neural network, the neural network is quantized by an enhanced loss function, wherein the enhanced loss function is used to enhance local information of the training data, and the enhanced loss function includes an enhanced first loss function, the second loss function, the third loss function and an enhanced fourth loss function; The enhanced first loss function is expressed by the following formula: The enhanced fourth loss function is expressed by the following formula: Among them, B Qj and B Mj Respectively represent the block B currently being output for reconstruction Qj and the corresponding block B in the pre-trained neural network Mj , and Respectively represent that the training data is forward propagated through B in the pre-trained neural network and the initialized quantized neural network M,i and B Q,i The output of , α is the weight parameter used to enhance the loss, where α>1.
9. An electronic device, characterized in that: The method comprises a processor and a memory, wherein the memory stores computer instructions, and when the computer instructions are executed by the processor, the steps of the method described in any one of claims 1 to 8 are implemented.
10. A storage medium, characterized in that: Computer instructions are stored thereon, and when the computer instructions are executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.