Neural Network Model Quantization Method and Device

The method addresses the challenge of optimizing mixed precision quantization in neural networks by determining target parameters and adjusting precision based on storage pressure, enhancing efficiency and automation in IoT devices.

CN114580610BActive Publication Date: 2025-07-15ALIBABA (SHENZHEN) TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210114146.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-30
Publication Date
2025-07-15
Estimated Expiration
2042-01-30

AI Technical Summary

Technical Problem

In the scenario of storage restricted, how to automatically improve the accuracy of the neural network model, while optimizing memory usage without exceeding the memory limit, especially in hybrid precision quantization, how to find a better combination to balance performance and accuracy.

Method used

The target parameters of the neural network model are determined through active variable analysis, and the parameter accuracy of the target layer network is adjusted according to the storage pressure. The hybrid precision quantization method is used to automatically adjust the parameter accuracy of different layers of networks to meet memory requirements.

Benefits of technology

It improves the automation quantization efficiency and accuracy of the neural network model, optimizes the use of memory resources, and ensures that the model can still run effectively under storage restricted conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114580610B_ABST
    Figure CN114580610B_ABST
Patent Text Reader

Abstract

The embodiments of this specification provide a neural network model quantization method and apparatus. The neural network model quantization method includes: determining target parameters of multiple layers of a target model according to the parameter attributes of the initial parameters of the target model, determining the storage pressure of multiple layers of the target model according to the target parameters, and adjusting the target parameters of the target layer network of the target model according to the storage pressure of the multiple layers. By determining the storage pressure of multiple layers of the target model through the target parameters, determining the target layer network and target parameters whose precision needs to be adjusted according to the storage pressure of the multiple layers, and then performing different precision adjustments on different target parameters, the automation degree and efficiency are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the technical field of neural networks, and particularly to a method for quantifying a neural network model. Background Art

[0002] Model quantization is an important means for model acceleration, which can effectively reduce the model size, improve the inference speed, and at the same time reduce the power consumption. Especially for low-power IoT devices, model quantization has become an essential key step in the engineering deployment of neural networks. Common quantizations include 8-bit / 16-bit integer quantization, and there are ultra-low-precision 2-bit and 4-bit quantizations on IoT devices. Lower-precision quantization can further improve the inference performance and compress the model memory overhead, but the low precision also leads to a decrease in the accuracy and precision of the model results.

[0003] Mixed-precision quantization allows different layers or even each different operator in the model to use different quantization precisions, which is a relatively flexible quantization method that balances performance and precision. However, how to find a better mixed-precision scheme is a challenging task. Taking the mixed quantization of int8 and int16 types as an example, there are a total of 2 to the power of n different combination ways, where n is the number of layers. Taking resnet50 as an example, n is 50. If the quantization is at the granularity of operators, the number of combination ways is even more. In scenarios with limited storage, the use of memory is the priority consideration. Under the condition of ensuring that the memory usage does not exceed the limit, how to automatically improve the precision has become an urgent problem to be solved. Summary of the Invention

[0004] In view of this, the embodiments of this specification provide a method for quantifying a neural network model. One or more embodiments of this specification also relate to a device for quantifying a neural network model, a computing device, a computer-readable storage medium, and a computer program, so as to solve the technical defects existing in the prior art.

[0005] According to the first aspect of the embodiments of this specification, a method for quantifying a neural network model is provided, including:

[0006] Determining target parameters of the multi-layer network of the target model according to the parameter attributes of the initial parameters of the target model based on a neural network;

[0007] Determining the storage pressure of the multi-layer network of the target model according to the target parameters;

[0008] Adjusting the target parameters of the target layer network of the target model according to the storage pressure of the multi-layer network.

[0009] According to a second aspect of the embodiments of the present specification, there is provided a neural network model quantization device, including:

[0010] A parameter determination module, configured to determine target parameters of multiple layers of the target model according to parameter attributes of initial parameters of a target model based on a neural network;

[0011] A pressure determination module, configured to determine storage pressure of multiple layers of the target model according to the target parameters;

[0012] An adjustment module, configured to adjust target parameters of a target layer network of the target model according to the storage pressure of the multiple layers of the network.

[0013] According to a third aspect of the embodiments of the present specification, there is provided a computing device, including:

[0014] A memory and a processor;

[0015] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above neural network model quantization method are implemented.

[0016] According to a fourth aspect of the embodiments of the present specification, there is provided a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the steps of the above neural network model quantization method are implemented.

[0017] According to a fifth aspect of the embodiments of the present specification, there is provided a computer program, wherein when the computer program is executed on a computer, the computer is made to execute the steps of the above neural network model quantization method.

[0018] The embodiments of the present specification provide a neural network model quantization method. The neural network model quantization method includes: determining target parameters of multiple layers of a target model according to parameter attributes of initial parameters of the target model, determining storage pressure of multiple layers of the target model according to the target parameters, and adjusting target parameters of a target layer network of the target model according to the storage pressure of the multiple layers of the network. By determining the storage pressure of multiple layers of the target model through the target parameters, determining the target layer network and target parameters whose precision needs to be adjusted according to the storage pressure of the multiple layers of the network, and then performing different precision adjustments on different target parameters, the automation degree and efficiency are improved. Description of the Drawings

[0019] Figure 1 is a flowchart of a neural network model quantization method provided by an embodiment of the present specification;

[0020] Figure 2It is a process flow chart of a neural network model quantization method provided by an embodiment of this specification;

[0021] Figure 3 It is a schematic structural diagram of a neural network model quantization device provided by an embodiment of this specification;

[0022] Figure 4 It is a structural block diagram of a computing device provided by an embodiment of this specification. Detailed implementation manners

[0023] Many specific details are set forth in the following description in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.

[0024] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more of the associated listed items.

[0025] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".

[0026] First, the noun terms related to one or more embodiments of this specification are explained.

[0027] Model quantization: It is a technology that converts floating-point calculations into low-bit fixed-point calculations, which can effectively reduce the model calculation intensity, parameter size, and memory consumption.

[0028] Live variable analysis: A typical data flow analysis in compilers that calculates whether each variable is live at the exit of which path.

[0029] In this specification, a neural network model quantization method is provided. This specification also relates to a neural network model quantization device, a computing device, a computer-readable storage medium, and a computer program, which will be described in detail one by one in the following embodiments.

[0030] Figure 1 The flowchart of a neural network model quantization method provided according to an embodiment of this specification is shown, which specifically includes the following steps.

[0031] Step 102: Determine the target parameters of the multi-layer network of the target model according to the parameter attributes of the initial parameters of the target model based on a neural network.

[0032] Among them, the target model can be a neural network model, including but not limited to a fully connected neural network (FCN), a convolutional neural network (CNN), a residual network (ResNet), and a feedback neural network; the initial parameters can be the parameters input to the target model, for example: variable A, variable B, and variable C input to the target model; the parameter attribute can be the attribute of the active period of the variable, for example: the active interval of variable A is in the first-layer network and the second-layer network of the target model, that is to say, variable A is used in both the first-layer network and the second-layer network of the target model; the target parameter can be part of the initial parameters, for example: variable A.

[0033] It should be noted that the multi-layer network can be understood as all the layer networks in the target model, or it can be part of the layer networks in the target model, and the same parameters are included between the layer networks in the part of the layer networks.

[0034] In practical applications, the target model is a convolutional neural network, and the initial precision of all parameters in the target model is 8 bits. Model quantization based on 16 bits is performed on the target model to improve the computing precision of the target model. However, in one possibility, the storage space of the hardware memory cannot support the precision of all parameters in the target model to be changed to 16 bits. That is to say, in the case of changing the precision of all parameters in the target model to 16 bits, the memory space may be insufficient. In such a case, it is necessary to select part of the parameters in the target model, change the precision of the part of the parameters to 16 bits, and at the same time, it will not exceed the size of the memory storage space. Thus, the target parameters of each layer in the target model can be determined first.

[0035] For example, the initial parameters of the target model include variable A, variable B, and variable C. The target model includes an L1 layer network, an L2 layer network, and an L3 layer network. The target parameters of the L1 layer network, L2 layer network, and L3 layer network are determined according to the parameter attributes of variable A, variable B, and variable C. The specific implementation is as follows.

[0036] Determining the target parameters of the multi-layer network of the target model according to the parameter attributes of the initial parameters of the target model based on a neural network includes:

[0037] Determine the initial parameters of the target model and the parameter attributes of the initial parameters;

[0038] According to the parameter attributes of the initial parameters, determine the target parameters of the multi-layer network of the target model.

[0039] Continuing with the above example, the initial parameters of the target model include variable A, variable B, and variable C. The target model includes an L1 layer network, an L2 layer network, and an L3 layer network. Among them, the active interval of variable A is the L1 layer network and the L2 layer network, the active interval of variable B is the L2 layer network and the L3 layer network, and the active interval of variable C is the L1 layer network, the L2 layer network, and the L3 layer network. Then, the target parameters of the L1 layer network include variable A and variable C, the target parameters of the L2 layer network include variable A, variable B, and variable C, and the interval of the L3 layer network is variable B and variable C.

[0040] The embodiments of this specification determine the target parameters according to the parameter attributes of the initial parameters, reducing the system overhead and improving the efficiency.

[0041] Determining the initial parameters of the target model and the parameter attributes of the initial parameters includes:

[0042] Determine the initial parameters of the target model;

[0043] Perform active variable analysis on the initial parameters to obtain the survival period of the initial parameters;

[0044] Correspondingly, determining the target parameters of the multi-layer network of the target model according to the parameter attributes of the initial parameters includes:

[0045] According to the survival period of the initial parameters, determine the target parameters of the multi-layer network of the target model.

[0046] Among them, the survival period can be understood as the period during which variables in the model are used, and can also be referred to as the active interval in the above embodiments.

[0047] Continuing with the above example, the initial parameters of the target model include variable A, variable B, and variable C. The target model includes an L1 layer network, an L2 layer network, and an L3 layer network. Among them, variable A, variable B, and variable C can obtain their live ranges through live variable analysis. Specifically, the identity identifier of the variable can be obtained, and it is queried whether the variable exists in multiple layers of the network in the target model using the identity identifier of the variable. Querying whether variable A exists in the L1 layer network, L2 layer network, and L3 layer network using the identity identifier of variable A, and it is found that variable A exists in the L1 layer network and L2 layer network, then the live range of variable A is the L1 layer network and L2 layer network. Similarly, the live range of variable B is obtained as the L2 layer network and L3 layer network through the same query method, and the live range of variable C is the L1 layer network, L2 layer network, and L3 layer network. Then, the target parameters of the L1 layer network include variable A and variable C, the target parameters of the L2 layer network include variable A, variable B, and variable C, and the target parameters of the L3 layer network are variable B and variable C.

[0048] The embodiments of this specification can accurately determine the target parameters of multiple layers of the network according to the lifespan of the initial parameters, improving the accuracy rate.

[0049] Step 104: Determine the storage pressure of multiple layers of the network of the target model according to the target parameters.

[0050] Among them, the storage pressure can be the pressure of multiple layers of the network in the target model on the memory. For example, if the first layer network of the target model occupies 50 bytes of memory, then the storage pressure of the first layer network of the target model is 50 bytes.

[0051] In practical applications, after determining the target parameters of multiple layers of the network of the target model, the storage pressure of multiple layers of the network can be calculated through the storage pressure of the target parameters of multiple layers of the network.

[0052] Specifically, the determining the storage pressure of multiple layers of the network of the target model according to the target parameters includes:

[0053] Obtain the quantity and precision of the target parameters of multiple layers of the network of the target model;

[0054] Obtain the storage pressure of multiple layers of the network according to the quantity and precision of the target parameters of multiple layers of the network.

[0055] In practical applications, determining the storage pressure of multiple layers of the network is also determining how much memory space the current layer needs to occupy during operation. When determining how much memory space the current layer network needs to occupy during operation, it is determined according to how many parameters exist in this layer network and the precision of these parameters. Then, it is necessary to first determine the quantity of parameters, and then multiply by the corresponding precision to obtain the storage pressure of this layer network.

[0056] For example, the target model includes an L1-layer network and an L2-layer network. Among them, the L1-layer network includes variable A and variable B, and the target parameters of the L2-layer network include variable B and variable C. Calculate the storage pressure of the L1-layer network with an initial precision of 8 bits. Among them, variable A occupies 8 bits of storage pressure, that is, 1 byte of storage pressure, and variable B occupies 8 bits of storage pressure, which is also 1 byte of storage pressure. Then the storage pressure of the L1-layer network is 2 bytes; if the dimension of the L2-layer network is 2D, multiply the storage pressures of variable B and variable C by 2 to obtain the storage pressure of the L2-layer network. Variable B occupies 8 bits of storage pressure, which is 1 byte of storage pressure, and variable C occupies 8 bits of storage pressure, which is 1 byte of storage pressure. Because the dimension of the L2-layer network is 2D, the storage pressure of the L2-layer network is 4 bytes.

[0057] The embodiments of this specification can quickly determine the storage pressure of a multi-layer network according to the precision and quantity of the target parameters of the multi-layer network. The calculation method is simple and the calculation efficiency is improved.

[0058] Step 106: Adjust the target parameters of the target layer network of the target model according to the storage pressure of the multi-layer network.

[0059] Among them, the target layer network can be any layer network in the target model. For example, the target layer network is the L1-layer network and the L2-layer network; the target parameters can be any parameters in the target layer network of the target model. For example: the target parameters are variable A and variable B.

[0060] In one case, the adjusting the target parameters of the target layer network of the target model according to the storage pressure of the multi-layer network includes:

[0061] In the case where the storage pressure of the multi-layer network is greater than the storage pressure of the initial storage space of the target object, reduce the precision of the target parameters of the multi-layer network to a preset precision; where the target model runs on the target object.

[0062] The target object here can be Memory. Correspondingly, the storage space can be the storage space of the memory; the preset precision can be any precision less than the current precision of the target parameters. For example, if the current precision of the target parameters is 8 bits, the preset precision can be any one of 2 bits and 4 bits.

[0063] In practical applications, when the parameters in the target model are run with the initially set precision, it is also possible to exceed the storage pressure of the memory, that is, the memory is insufficient. At this time, the target model cannot be run, and the precision of the parameters in the target model needs to be set to a lower precision so that the target model can run.

[0064] For example, the precision of the parameters in the target model is 8 bits. The target model includes a network layer L1, where the network layer L1 includes variables A, B, C, and D. Calculate the storage pressure of the multi-layer network of the target model. If the dimension of the network layer L1 is 2D, the storage pressure of the network layer L1 is 8 bytes. When the storage space of the memory is 6 bytes, the storage pressure of the memory is insufficient to support the operation of the target model. Then, the precision of the network layer L1 is adjusted to 4 bits, and in this way, the storage pressure of the network layer L1 becomes 4 bytes.

[0065] In the embodiments of this specification, when the storage pressure of the storage space is lower than the storage pressure required by the lowest precision, the precision is uniformly reduced to enable the target model to run, reducing the expenditure on system performance in the process and improving the efficiency.

[0066] In another case, the adjustment of the target parameters of the target layer network of the target model according to the storage pressure of the multi-layer network includes:

[0067] Determine the precision ratio of the multi-layer network according to the target precision of the target parameters of the multi-layer network of the target model, and adjust the target parameters of the target layer network of the target model according to the precision ratio of the multi-layer network and the storage space of the target object.

[0068] Determine the target precision of the target parameters of the multi-layer network of the target model;

[0069] Determine the precision ratio according to the target precision and the precision of the target parameters of the multi-layer network;

[0070] Determine the storage pressure of the proportional storage space of the target object according to the precision ratio;

[0071] Adjust the target parameters of the target layer network of the target model according to the storage pressure of the multi-layer network and the storage pressure of the proportional storage space.

[0072] Among them, the target precision can be understood as the precision to be achieved. For example, the current precision is 8 bits, and the target precision can be 16 bits; the precision ratio can be the ratio of the current precision to the target precision. For example: the precision ratio is 8 bits:16 bits = 0.5; the proportional storage space can be the space calculated according to the precision ratio. For example, if the memory is 100 bytes and the precision ratio is 0.5, then the proportional storage space is 50 bytes.

[0073] In practical applications, the proportional storage space can be determined according to the target precision, and how to adjust the precision of the parameters in the target model can be determined according to the proportional storage space. That is to say, it is judged whether the memory space is sufficient when the precision of all parameters of the target model is set to the target precision. When the memory space is sufficient, the precision of all parameters is set to the target precision; when the memory space is insufficient, the precision of some parameters is set to the target precision.

[0074] Specifically, the adjustment of the target parameters of the target layer network of the target model according to the storage pressure of the multi-layer network and the storage pressure of the proportional storage space includes:

[0075] When the storage pressure of the multi-layer network is less than the storage pressure of the proportional storage space, the precision of the target parameters of the multi-layer network is increased to the target precision.

[0076] For example, the initial precision of the parameters of the target model is 8 bits, and the target precision of the target parameters of the multi-layer network of the target model is set to 16 bits. It can be determined that the precision ratio is 0.5. When the memory is 100 bytes, the proportional storage space can be determined to be 50 bytes, and the storage pressure of the proportional storage space is 50 bytes. When the storage pressure of the multi-layer network of the target model calculated with the initial precision of 8 bits is less than 50 bytes, the target precision of the target parameters of the multi-layer network of the target model is set to 16 bits.

[0077] In the embodiment of this specification, when the storage pressure of the storage space can support the storage pressure required by double precision, the precision is uniformly increased to enable the target model to run, reducing the expenditure on system performance in the process and improving the efficiency.

[0078] In another implementable manner, the adjustment of the target parameters of the target layer network of the target model according to the storage pressure of the multi-layer network and the storage pressure of the proportional storage space includes:

[0079] The network in the target model whose storage pressure is less than or equal to the storage pressure of the initial storage space of the target object and greater than the storage pressure of the proportional storage space is determined as the target layer network;

[0080] The precision of the target parameters of the target layer network is adjusted according to the storage pressure of the target layer network.

[0081] For example, the initial precision of the parameters of the target model is 8 bits. If the target precision of the target parameters of the multi-layer network of the target model is set to 16 bits, the precision ratio can be determined to be 0.5. When the memory is 100 bytes, the proportional storage space can be determined to be 50 bytes, and the storage pressure of the proportional storage space is also 50 bytes. If the target model includes an L1 layer network and an L2 layer network, and the storage pressure of the multi-layer network of the target model is calculated with an initial precision of 8 bits, the storage pressure of the L1 layer network is 40 bytes, and the storage pressure of the L2 layer network is 60 bytes, then the L2 layer network is adjusted.

[0082] In the embodiments of the present specification, different precision adjustments are made to different target parameters. Without exceeding the memory limit while ensuring the precision, the utilization efficiency of memory resources is improved.

[0083] Adjusting the precision of the target parameters of the target layer network according to the storage pressure of the target layer network includes:

[0084] Sorting the target layer networks in descending order according to the storage pressure of the target layer network to obtain a first sorted list;

[0085] Adjusting the precision of the target parameters of the target layer network by a first adjustment method according to the first sorted list.

[0086] Among them, the first sorted list can be a data list including the target layer networks. For example, the first sorted list includes the L1 layer network and the L2 layer network; the first adjustment method can be any adjustment method as long as it can achieve the precision adjustment of the target layer network, and the embodiments of the present specification do not make any limitations.

[0087] For example, the target layer network includes an L1 layer network, an L2 layer network, and an L3 layer network. Among them, the storage pressure of the L1 layer network is 60 bytes, the storage pressure of the L2 layer network is 70 bytes, and the storage pressure of the L3 layer network is 75 bytes. Then, sorting according to the storage pressure, the first sorted list is the L3 layer network, the L2 layer network, and the L1 layer network, and the precision of the target parameters of the L3 layer network, the L2 layer network, and the L1 layer network is adjusted by the first adjustment method.

[0088] In the embodiments of the present specification, the target network layers are sorted in descending order, and the target layer network with the greatest influence is processed first, reducing the system overhead of subsequent adjustment steps and improving the efficiency.

[0089] The specific implementation manner of adjusting the precision of the target parameters of the L1 layer network, the L2 layer network, and the L3 layer network by the first adjustment method is as follows.

[0090] Adjusting the accuracy of the target parameters of the target layer network according to the first sorted list by the first adjustment method includes:

[0091] Set the accuracy of the target layer network in the first sorted list to the target accuracy;

[0092] Determine the first target layer network in the first sorted list;

[0093] Obtain the storage pressure difference according to the storage pressure of the first target layer network and the storage pressure of the proportional storage space;

[0094] Adjust the accuracy of the target parameters in the first target layer network according to the storage pressure difference, and delete the first target layer network from the first sorted list;

[0095] Continue to select the first target layer network in the first sorted list until the first sorted list is empty.

[0096] Among them, the storage pressure difference can be understood as the difference between the storage pressure of the target layer network and the storage pressure of the proportional storage space. For example: the storage pressure of the target layer network is 60 bytes, and the storage pressure of the proportional storage space is 50 bytes, then the storage pressure difference is 10 bytes.

[0097] Continuing with the above example, the memory space is 100 bytes, the target layer network includes the L1 layer network, the L2 layer network, and the L3 layer network, and the accuracy of the L1 layer network, the L2 layer network, and the L3 layer network is all 8 bits. Among them, the storage pressure of the L1 layer network is 60 bytes, the storage pressure of the L2 layer network is 70 bytes, and the storage pressure of the L3 layer network is 75 bytes. Then, sorting by storage pressure, the first sorted list is the L3 layer network, the L2 layer network, and the L1 layer network. Set the accuracy of the L1 layer network, the L2 layer network, and the L3 layer network to 16 bits. Then select the first target layer network in the first sorted list: the L3 layer network. Subtract the storage pressure of the proportional storage space of 50 bytes from the storage pressure of the L3 layer network of 75 bytes to obtain a storage pressure difference of 25 bytes. Adjust the accuracy of the L3 layer network according to the storage pressure difference of 25 bytes, and then delete the L3 layer network from the first sorted list. Then, select the L2 layer network again. Subtract the storage pressure of the proportional storage space of 50 bytes from the storage pressure of 70 bytes of the L2 layer network to obtain a storage pressure difference of 20 bytes. Adjust the accuracy of the L2 layer network according to the storage pressure difference of 20 bytes, and then delete the L2 layer network from the first sorted list. Then, select the L1 layer network again. Subtract the storage pressure of the proportional storage space of 50 bytes from the storage pressure of 60 bytes of the L1 layer network to obtain a storage pressure difference of 10 bytes. Adjust the accuracy of the L1 layer network according to the storage pressure difference of 10 bytes, and then delete the L1 layer network from the first sorted list.

[0098] In the embodiments of this specification, the storage pressure difference is first determined and adjusted according to the storage pressure difference, which can ensure the maximum utilization of the memory space and improve the accuracy to the greatest extent.

[0099] Specifically, the accuracy adjustment of the target parameters in the first target layer network according to the storage pressure difference includes:

[0100] Determine the current accuracy of the target parameters in the first target layer network;

[0101] Adjust the storage pressure difference according to the current accuracy of the target parameters, determine the current storage pressure difference, and determine the target parameter list, where the target parameter list includes that the current accuracy of the target parameters is the target accuracy;

[0102] When the current storage pressure difference is greater than the preset pressure threshold, adjust the accuracy of the target parameters in the parameter list.

[0103] Among them, the current accuracy can be the current accuracy of the target parameters. Because in the above embodiments, the target layer network is adjusted cyclically, and when adjusting other target layer networks, the accuracy of the target parameters in the current target layer network may have been adjusted, so it is necessary to first determine the current accuracy of the target parameters; the current storage pressure difference can be the current storage pressure difference during the cyclic execution; the target parameter list can be the parameters selected from the parameters of the target layer network to be adjusted; the preset pressure threshold can be a set threshold, such as 0 bytes, 10 bytes.

[0104] Continuing with the above example, select the first target layer network in the first row list: the L2 layer network. Subtract the storage pressure of 50 bytes in the proportional storage space from the storage pressure of 70 bytes in the L2 layer network to obtain a storage pressure difference of 20 bytes. Adjust the accuracy of the L2 layer network according to the storage pressure difference of 20 bytes. Among them, the L2 layer network includes variable A, variable B, and variable C. Determine that the current accuracy of variable A is 8 bits, the current accuracy of variable B is 16 bits, and the current accuracy of variable C is 16 bits. Then, determine the current storage pressure difference according to the storage pressure difference of 20 bytes and variable A, and put variable B and variable C into the target parameter list. Then, if the current storage pressure difference is greater than the preset pressure threshold of 0 bytes, adjust the accuracy of variable B and variable C in the parameter list.

[0105] The embodiments of this specification determine the current accuracy of the target parameters to avoid repeated calculations and improve the calculation efficiency.

[0106] Specifically, the adjustment of the storage pressure difference according to the current accuracy of the target parameters to determine the current storage pressure difference includes:

[0107] Determine whether there is a previous target parameter for the target parameter;

[0108] If so, when the current precision of the target parameter is inconsistent with the target precision, subtract the stored pressure of the target parameter from the stored pressure difference to obtain the current stored pressure difference;

[0109] If not, when the current precision of the target parameter is inconsistent with the target precision, subtract the stored pressure of the target parameter from the current stored pressure difference of the previous target parameter to obtain the updated current stored pressure difference.

[0110] Continuing with the above example, the current precision of variable A is 8 bits, the current precision of variable B is 16 bits, and the current precision of variable C is 16 bits. Determine the current stored pressure difference based on the stored pressure difference of 20 bytes and variable A. When the dimension of variable A is 1-dimensional and it is the first variable in the current loop, subtract 1 byte from the stored pressure difference of 20 bytes to obtain the current stored pressure difference of 19 bytes.

[0111] In another case, if there is a next variable D with a precision of 8 bits, subtract the stored pressure of variable D from the current stored pressure difference of 19 bytes to obtain the updated current stored pressure difference.

[0112] The embodiments of this specification cyclically update the current stored pressure difference, improving the accuracy of the current stored pressure difference.

[0113] Next, if the current stored pressure difference is greater than the preset pressure threshold of 0 bytes, the specific implementation of adjusting the precision of variables B and C in the parameter list by the threshold is as follows.

[0114] When the current stored pressure difference is greater than the preset pressure threshold, adjusting the precision of the target parameter in the parameter list includes:

[0115] Sort the target parameters in the parameter list in descending order according to the stored pressure of the target parameters in the parameter list to obtain a second sorted list;

[0116] Adjust the precision of the target parameters in the parameter list according to the second sorted list through a second adjustment method.

[0117] Among them, the second sorted list can be a data list including target parameters. For example, the second sorted list includes variables B and C; the second adjustment method can be any adjustment method as long as it can achieve the precision adjustment of the target parameter, and the embodiments of this specification do not limit it.

[0118] Continuing with the above example, the target parameter list includes variable B and variable C. The storage pressure of variable B is 10 bytes, and the storage pressure of variable C is 16 bytes. Then the second sorted list is: variable B, variable C. Then, the precision of variables B and C in the parameter list is adjusted by the second adjustment method.

[0119] After the current storage pressure difference is determined in the embodiment of this specification, the target layer network that has met the storage pressure requirements is skipped, reducing the calculation process and saving system overhead.

[0120] The specific method for adjusting the precision of variables B and C in the parameter list by the second adjustment method is described as follows.

[0121] Adjusting the precision of the target parameters in the parameter list by the second adjustment method according to the second sorted list includes:

[0122] Determine the first target parameter in the second sorted list;

[0123] Adjust the precision of the first target parameter to the initial precision;

[0124] Determine the storage pressure saved for the first target layer network according to the initial precision;

[0125] If the saved storage pressure is less than the current storage pressure difference, delete the first target parameter in the second sorted list;

[0126] Continue to execute the determination of the first target parameter in the second sorted list until the second sorted list is empty.

[0127] Among them, the initial precision can be the precision when the parameters in the target model have not been adjusted. For example, the initial precision is 8 bits; the saved storage pressure can be the storage pressure saved by adjusting to a lower precision. For example, when the precision is adjusted from 16 bits to 8 bits, half of the storage pressure can be saved.

[0128] If the current storage pressure difference is 10 bytes, continuing with the above example, the second sorted list is: variable B, variable C. The storage pressure of variable B is 10 bytes, and the storage pressure of variable C is 16 bytes. Adjust the precision of variable B from 16 bits to 8 bits and delete variable B from the second sorted list, then the saved storage pressure can be obtained as 5 bytes. Since the saved storage pressure of 5 bytes is less than the current storage pressure difference of 10 bytes, then continue to select variable C, adjust the precision of variable C from 16 bits to 8 bits and delete variable C from the second sorted list, then the saved storage pressure can be obtained as 8 bytes. At this time, the saved storage pressure is 5 bytes plus 8 bytes, and the new saved storage pressure is 13 bytes. Then the precision adjustment of the current target layer network is completed.

[0129] In the embodiments of this specification, the target parameters whose precision has not been adjusted are sorted in descending order, and the target parameters with the greatest impact are adjusted until the storage pressure requirement is met, so as to reduce the calculation process and improve the efficiency.

[0130] See Figure 2 , Figure 2 which shows a process flow chart of a neural network model quantization method provided by an embodiment of this specification, specifically including the following steps.

[0131] Step 202: Determine the initial parameters of the target model, perform active variable analysis on the initial parameters, and obtain the survival period of the initial parameters.

[0132] Among them, the target model can be a neural network model, including but not limited to a fully connected neural network (FCN), a convolutional neural network (CNN), a residual network (ResNet), and a feedback neural network; the initial parameters can be the parameters input to the target model, for example: variable A, variable B, and variable C input to the target model; the survival period can be understood as the period during which the variables in the model are used, and can also be called the active interval.

[0133] In practical applications, the target model is a convolutional neural network, and the initial precision of all parameters in the target model is 8 bits. The target model is quantized based on 16 bits to improve the calculation precision of the target model. However, in a possible situation, the storage space of the hardware memory cannot support the precision of all parameters in the target model to be changed to 16 bits. That is to say, in the case of changing the precision of all parameters in the target model to 16 bits, the memory space may be insufficient. In such a case, it is necessary to select some parameters of the target model, change the precision of some parameters to 16 bits, and at the same time, it will not exceed the size of the memory storage space. Thus, the target parameters of each layer in the target model can be determined first.

[0134] For example, the initial parameters of the target model include variable A, variable B, and variable C, and the target model includes an L1 layer network, an L2 layer network, and an L3 layer network. Among them, the active ranges (live ranges) of variable A, variable B, and variable C can be obtained through active variable analysis. Specifically, the identity identifier of the variable can be obtained, and it is queried whether the variable exists in multiple layers of the network in the target model using the identity identifier of the variable. Query whether variable A exists in the L1 layer network, L2 layer network, and L3 layer network using the identity identifier of variable A. If it is obtained that variable A exists in the L1 layer network and L2 layer network, then the active range of variable A is the L1 layer network and L2 layer network. Similarly, the active range of variable B is obtained as the L2 layer network and L3 layer network, and the active range of variable C is the L1 layer network, L2 layer network, and L3 layer network.

[0135] Step 204: Determine the target parameters of the multiple layers of the network of the target model according to the survival period of the initial parameters.

[0136] Among them, the target parameters can be some of the initial parameters. For example: variable A.

[0137] Continuing with the above example, the active range of variable A is the L1 layer network and L2 layer network, the active range of variable B is the L2 layer network and L3 layer network, and the active range of variable C is the L1 layer network, L2 layer network, and L3 layer network. Then the target parameters of the L1 layer network include variable A and variable C, the target parameters of the L2 layer network include variable A, variable B, and variable C, and the target parameters of the L3 layer network are variable B and variable C.

[0138] Step 206: Determine the storage pressure of the multiple layers of the network of the target model according to the target parameters.

[0139] Among them, the storage pressure can be the pressure on the memory of multiple layers of the network in the target model. For example, if the memory occupied by the first layer network of the target model is 50 bytes, then the storage pressure of the first layer network of the target model is 50 bytes.

[0140] In practical applications, after determining the target parameters of the multiple layers of the network of the target model, the storage pressure of the multiple layers of the network can be calculated through the storage pressure of the target parameters of the multiple layers of the network.

[0141] Continuing with the above example, the target model includes an L1 layer network, an L2 layer network, and an L3 layer network. Among them, the L1 layer network includes variable A and variable B, the target parameters of the L2 layer network include variable A, variable B, and variable C, and the target parameters of the L3 layer network are variable B and variable C. Calculate the storage pressure of the L1 layer network with an initial precision of 8 bits. Among them, the storage pressure of variable A is 30 bytes, the storage pressure of variable B is 25 bytes, and the storage pressure of variable C is 40 bytes. Then the storage pressure of the L1 layer network is 55 bytes; the storage pressure of the L2 layer network is 95 bytes; the storage pressure of the L3 layer network is 65 bytes.

[0142] Step 208: Determine the target precision of the target parameters of the multi-layer network of the target model, determine the precision ratio according to the target precision and the precision of the target parameters of the multi-layer network, and determine the storage pressure of the proportional storage space of the target object according to the precision ratio.

[0143] Among them, the target precision can be understood as the precision that one wants to achieve. For example, the current precision is 8 bits, and the target precision can be 16 bits; the precision ratio can be the ratio of the current precision and the target precision. For example: the precision ratio is 8 bits:16 bits = 0.5, and the proportional storage space can be the space calculated according to the precision ratio. For example, the memory is 100 bytes and the precision ratio is 0.5, then the proportional storage space is 50 bytes; the target parameter can be any parameter in the target layer network of the target model.

[0144] In practical applications, the proportional storage space can be determined according to the target precision, and how to adjust the precision of the parameters in the target model can be determined according to the proportional storage space. That is to say, it is judged whether the memory space is sufficient when the precision of all parameters of the target model is set to the target precision. When the memory space is sufficient, the precision of all parameters is set to the target precision. When the memory space is insufficient, the precision of some parameters is set to the target precision.

[0145] For example, the initial precision of the parameters of the target model is 8 bits, and the target precision of the target parameters of the multi-layer network of the target model is set to 16 bits. The precision ratio can be determined to be 0.5. When the memory is 100 bytes, the proportional storage space can be determined to be 50 bytes, and the storage pressure of the proportional storage space is also 50 bytes.

[0146] Step 210: Determine the network in the target model whose storage pressure is less than or equal to the storage pressure of the initial storage space of the target object and greater than the storage pressure of the proportional storage space as the target layer network.

[0147] Among them, the target layer network can be any layer network in the target model.

[0148] Continuing with the above example, the initial precision of the parameters of the target model is 8 bits. Setting the target precision of the target parameters of the multi-layer network of the target model to 16 bits, the precision ratio can be determined to be 0.5. In the case of 100 bytes of memory, the proportional storage space can be determined to be 50 bytes, and the storage pressure of the proportional storage space is also 50 bytes. The storage pressure of the L1 layer network is 55 bytes, the storage pressure of the L2 layer network is 95 bytes, and the storage pressure of the L3 layer network is 65 bytes. Then, the L1 layer network, the L2 layer network, and the L3 layer network are all target layer networks.

[0149] Step 212: Sort the target layer networks in descending order according to the storage pressure of the target layer networks to obtain the first sorted list, and set the precision of the target layer networks in the first sorted list to the target precision.

[0150] Among them, the first sorted list can be a data list including the target layer networks. For example, the first sorted list includes the L1 layer network and the L2 layer network.

[0151] For example, the target layer networks include the L1 layer network, the L2 layer network, and the L3 layer network. Among them, the storage pressure of the L1 layer network is 55 bytes, the storage pressure of the L2 layer network is 95 bytes, and the storage pressure of the L3 layer network is 65 bytes. Then, sorting according to the storage pressure, the first sorted list is the L2 layer network, the L3 layer network, and the L1 layer network, and the precision of the L1 layer network, the L2 layer network, and the L3 layer network is set to 16 bits.

[0152] Step 214: Determine the first target layer network in the first sorted list, and obtain the storage pressure difference according to the storage pressure of the first target layer network and the storage pressure of the proportional storage space.

[0153] Among them, the storage pressure difference can be understood as the difference between the storage pressure of the target layer network and the storage pressure of the proportional storage space. For example: the storage pressure of the target layer network is 60 bytes, and the storage pressure of the proportional storage space is 50 bytes, then the storage pressure difference is 10 bytes.

[0154] Continuing with the above example, select the first target layer network in the first sorted list: the L2 layer network, subtract the storage pressure of 50 bytes of the proportional storage space from the storage pressure of 95 bytes of the L2 layer network, and obtain a storage pressure difference of 45 bytes.

[0155] Step 216: Adjust the storage pressure difference according to the current precision of the target parameters to determine the current storage pressure difference, and determine the target parameter list, where the target parameter list includes that the current precision of the target parameters is the target precision.

[0156] Among them, the current precision can be the current precision of the target parameter. Since the target layer network is cyclically adjusted in the above embodiments, when adjusting other target layer networks, the precision of the target parameter in the current target layer network may have been adjusted. Therefore, it is necessary to first determine the current precision of the target parameter; the current storage pressure difference can be the current storage pressure difference during the cyclic execution; the target parameter list can be the parameters selected from the parameters of the target layer network to be adjusted.

[0157] Continuing with the above example, select the first target layer network in the first row list: the L2 layer network. Subtract the storage pressure of 50 bytes of the proportional storage space from the storage pressure of 95 bytes of the L2 layer network to obtain a storage pressure difference of 45 bytes. Adjust the precision of the L2 layer network according to the storage pressure difference of 45 bytes. Among them, the L2 layer network includes variable A, variable B, and variable C. Determine that the current precision of variable A is 8 bits, the current precision of variable B is 16 bits, and the current precision of variable C is 16 bits. Then, determine the current storage pressure difference according to the storage pressure difference of 45 bytes and the storage pressure of 30 bytes of variable A: 15 bytes. Put variable B and variable C into the target parameter list.

[0158] Step 218: When the current storage pressure difference is greater than the preset pressure threshold, sort the target parameters in the parameter list in descending order according to the storage pressure of the target parameters to obtain the second row list.

[0159] Among them, the preset pressure threshold can be a set threshold, such as 0 bytes, 10 bytes.

[0160] Continuing with the above example, the current storage pressure difference is greater than the preset pressure threshold of 0 bytes. The target parameter list includes variable B and variable C. The storage pressure of variable B is 25 bytes, and the storage pressure of variable C is 40 bytes. Then the second row list is: variable C, variable B.

[0161] Step 220: Determine the first target parameter in the second row list; adjust the precision of the first target parameter to the initial precision; determine the storage pressure saved by the first target layer network according to the initial precision.

[0162] Among them, the initial precision can be the precision when the parameters in the target model have not been adjusted. For example, the initial precision is 8 bits; the storage pressure saved can be the storage pressure saved by adjusting to a lower precision. For example, when adjusting the precision from 16 bits to 8 bits, half of the storage pressure can be saved.

[0163] Continuing with the above example, if the current storage pressure difference is 15 bytes and the second row sequence list is: variable C, variable B, the storage pressure of variable B is 25 bytes, and the storage pressure of variable C is 40 bytes. Adjust the precision of variable B from 16 bits to 8 bits and delete variable B from the second row sequence list, then the saved storage pressure can be obtained as 12.5 bytes.

[0164] Step 222: When the saved storage pressure is less than the current storage pressure difference, delete the first target parameter in the second row sequence list; continue to determine the first target parameter in the second row sequence list until the second row sequence list is empty, and delete the first target layer network in the first row sequence list.

[0165] Continuing with the above example, the saved storage pressure of 12.5 bytes is less than the current storage pressure difference of 15 bytes, so continue to select variable C. Adjust the precision of variable C from 16 bits to 8 bits and delete variable C from the second row sequence list, then the saved storage pressure can be obtained as 20 bytes. At this time, the new saved storage pressure is 12.5 bytes plus 20 bytes, and the new saved storage pressure is 32.5 bytes. The second row sequence list is empty, so the precision adjustment of the L2 layer network is completed, and the L2 layer network is deleted from the first row sequence list.

[0166] Step 224: Determine whether the first row sequence list is empty.

[0167] Continuing with the above example, when the first row sequence list is not empty, continue to execute steps 216 to 222 to adjust the precision of the L1 layer network and the L3 layer network. When the first row sequence list is empty, the adjustment is completed.

[0168] Corresponding to the above method embodiments, this specification also provides embodiments of a neural network model quantization device. Figure 3 It shows a schematic structural diagram of a neural network model quantization device provided by an embodiment of this specification. As Figure 3 shown, the device includes:

[0169] A parameter determination module 302, configured to determine the target parameters of the multi-layer network of the target model according to the parameter attributes of the initial parameters of the target model based on the neural network;

[0170] A pressure determination module 304, configured to determine the storage pressure of the multi-layer network of the target model according to the target parameters;

[0171] An adjustment module 306, configured to adjust the target parameters of the target layer network of the target model according to the storage pressure of the multi-layer network.

[0172] Optionally, the parameter determination module 302 is further configured to:

[0173] Determine the initial parameters of the target model;

[0174] Conduct active variable analysis on the initial parameters to obtain the survival period of the initial parameters;

[0175] Correspondingly, determining the target parameters of the multi-layer network of the target model according to the parameter attributes of the initial parameters includes:

[0176] Determine the target parameters of the multi-layer network of the target model according to the survival period of the initial parameters.

[0177] Optionally, the pressure determination module 304 is further configured to:

[0178] Obtain the quantity and precision of the target parameters of the multi-layer network of the target model;

[0179] Obtain the storage pressure of the multi-layer network according to the quantity and precision of the target parameters of the multi-layer network.

[0180] Optionally, the adjustment module 306 is further configured to:

[0181] Determine the target precision of the target parameters of the multi-layer network of the target model;

[0182] Determine the precision ratio according to the target precision and the precision of the target parameters of the multi-layer network;

[0183] Determine the storage pressure of the proportional storage space of the target object according to the precision ratio;

[0184] Adjust the target parameters of the target layer network of the target model according to the storage pressure of the multi-layer network and the storage pressure of the proportional storage space.

[0185] Optionally, the adjustment module 306 is further configured to:

[0186] Determine the network in the target model whose storage pressure is less than or equal to the storage pressure of the initial storage space of the target object and greater than the storage pressure of the proportional storage space as the target layer network;

[0187] Adjust the precision of the target parameters of the target layer network according to the storage pressure of the target layer network.

[0188] Optionally, the adjustment module 306 is further configured to:

[0189] Sort the target layer network in descending order according to the storage pressure of the target layer network to obtain the first sorted list;

[0190] Adjust the accuracy of the target parameters of the target layer network through the first adjustment method according to the first sorted list.

[0191] Optionally, the adjustment module 306 is further configured to:

[0192] Set the accuracy of the target layer network in the first sorted list to the target accuracy;

[0193] Determine the first target layer network in the first sorted list;

[0194] Obtain a storage pressure difference according to the storage pressure of the first target layer network and the storage pressure of the proportional storage space;

[0195] Adjust the accuracy of the target parameters in the first target layer network according to the storage pressure difference, and delete the first target layer network in the first sorted list;

[0196] Continue to execute the selection of the first target layer network in the first sorted list until the first sorted list is empty.

[0197] Optionally, the adjustment module 306 is further configured to:

[0198] Determine the current accuracy of the target parameters in the first target layer network;

[0199] Adjust the storage pressure difference according to the current accuracy of the target parameters, determine the current storage pressure difference, and determine a target parameter list, where the target parameter list includes that the current accuracy of the target parameters is the target accuracy;

[0200] When the current storage pressure difference is greater than a preset pressure threshold, adjust the accuracy of the target parameters in the parameter list.

[0201] Optionally, the adjustment module 306 is further configured to:

[0202] Determine whether there is a previous target parameter for the target parameter;

[0203] If so, when the current accuracy of the target parameter is inconsistent with the target accuracy, subtract the storage pressure of the target parameter from the storage pressure difference to obtain the current storage pressure difference;

[0204] If not, when the current accuracy of the target parameter is inconsistent with the target accuracy, subtract the storage pressure of the target parameter from the current storage pressure difference of the previous target parameter to obtain the updated current storage pressure difference.

[0205] Optionally, the adjustment module 306 is further configured to:

[0206] Sort the target parameters in the parameter list in descending order according to the storage pressure of the target parameters in the parameter list to obtain a second sorted list;

[0207] Adjust the precision of the target parameters in the parameter list according to the second sorted list through a second adjustment method.

[0208] Optionally, the adjustment module 306 is further configured to:

[0209] Determine the first target parameter in the second sorted list;

[0210] Adjust the precision of the first target parameter to the initial precision;

[0211] Determine the storage pressure saved for the first target layer network according to the initial precision;

[0212] Delete the first target parameter in the second sorted list if the saved storage pressure is less than the current storage pressure difference;

[0213] Continue to execute determining the first target parameter in the second sorted list until the second sorted list is empty.

[0214] An embodiment of this specification provides a neural network model quantization device. This neural network model quantization device determines the target parameters of the multi-layer network of the target model according to the parameter attributes of the initial parameters of the target model, determines the storage pressure of the multi-layer network of the target model according to the target parameters, and adjusts the target parameters of the target layer network of the target model according to the storage pressure of the multi-layer network. By determining the storage pressure of the multi-layer network of the target model according to the target parameters, determining the target layer network and the target parameters whose precision needs to be adjusted according to the storage pressure of the multi-layer network, and then performing different precision adjustments on different target parameters, the automation degree and efficiency are improved.

[0215] The above is a schematic solution of a neural network model quantization device in this embodiment. It should be noted that the technical solution of this neural network model quantization device and the technical solution of the above neural network model quantization method belong to the same concept. For the details not described in the technical solution of this neural network model quantization device, reference can be made to the description of the technical solution of the above neural network model quantization method.

[0216] Figure 4 The structural block diagram of a computing device 400 provided according to an embodiment of this specification is shown. The components of this computing device 400 include but are not limited to a memory 410 and a processor 420. The processor 420 is connected to the memory 410 through a bus 430, and a database 450 is used to store data.

[0217] The computing device 400 also includes an access device 440, which enables the computing device 400 to communicate via one or more networks 460. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 440 may include one or more of any type of wired or wireless network interfaces (e.g., Network Interface Card (NIC)), such as an IEEE802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.

[0218] In one embodiment of the present specification, the above components of the computing device 400 and Figure 4 other components not shown may also be connected to each other, for example, via a bus. It should be understood that Figure 4 the block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of the present specification. Those skilled in the art can add or replace other components as needed.

[0219] The computing device 400 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or PCs. The computing device 400 can also be a mobile or stationary server.

[0220] Wherein, the processor 420 is used to execute the following computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the above neural network model quantization method are implemented.

[0221] The above is a schematic solution of a computing device in this embodiment. It should be noted that the technical solution of the computing device and the technical solution of the above neural network model quantization method belong to the same concept. For the details not described in the technical solution of the computing device, reference can be made to the description of the technical solution of the above neural network model quantization method.

[0222] One embodiment of the present specification also provides a computer-readable storage medium, which stores computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the above neural network model quantization method are implemented.

[0223] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solution of the above neural network model quantization method belong to the same concept. For the details not described in the technical solution of the storage medium, reference can be made to the description of the technical solution of the above neural network model quantization method.

[0224] An embodiment of this specification also provides a computer program. When the computer program is executed on a computer, the computer is made to execute the steps of the above neural network model quantization method.

[0225] The above is a schematic solution of a computer program according to this embodiment. It should be noted that the technical solution of this computer program and the technical solution of the above neural network model quantization method belong to the same concept. For the details not described in the technical solution of the computer program, reference can be made to the description of the technical solution of the above neural network model quantization method.

[0226] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0227] The computer instructions include computer program code, which may be in the form of source code, object code, an executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a removable hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content included in the computer-readable medium may be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0228] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described action sequence, because according to the embodiments of this specification, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential for the embodiments of this specification.

[0229] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0230] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, according to the content of the embodiments of this specification, many modifications and changes can be made. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can well understand and utilize this specification. This specification is only limited by the claims and their full scope and equivalents.

Claims

1. A method for quantifying a neural network model, comprising: Determining target parameters of multiple layers of a target model according to parameter attributes of initial parameters of the target model of the neural network; Determining storage pressure of multiple layers of the target model according to the target parameters; Adjusting target parameters of a target layer network of the target model according to the storage pressure of the multiple layers of the network, wherein adjusting the target parameters of the target layer network of the target model according to the storage pressure of the multiple layers of the network includes: determining a precision ratio of the multiple layers of the network according to a target precision of the target parameters of the multiple layers of the target model, and adjusting the target parameters of the target layer network of the target model according to the precision ratio, storage space of a target object, and storage pressure of the multiple layers of the network, wherein the target model runs on the target object, and the storage pressure is the storage pressure of the multiple layers of the target model on the target object.

2. The method according to claim 1, wherein determining target parameters of multiple layers of the target model according to parameter attributes of initial parameters of a target model based on a neural network includes: Determining initial parameters of the target model; Performing active variable analysis on the initial parameters to obtain a survival period of the initial parameters; Determining target parameters of multiple layers of the target model according to the survival period of the initial parameters.

3. The method according to claim 1, wherein determining storage pressure of multiple layers of the target model according to the target parameters includes: Obtaining the quantity and precision of target parameters of multiple layers of the target model; Obtaining the storage pressure of the multiple layers of the network according to the quantity and precision of the target parameters of the multiple layers of the network.

4. The method according to claim 1, wherein adjusting target parameters of a target layer network of the target model according to the storage pressure of the multiple layers of the network includes: Determining a target precision of target parameters of multiple layers of the target model; Determining a precision ratio according to the target precision and the precision of the target parameters of the multiple layers of the network; Determining storage pressure of a proportional storage space of the target object according to the precision ratio; Adjusting target parameters of a target layer network of the target model according to the storage pressure of the multiple layers of the network and the storage pressure of the proportional storage space.

5. The method according to claim 4, wherein adjusting target parameters of a target layer network of the target model according to the storage pressure of the multiple layers of the network and the storage pressure of the proportional storage space includes: Determining, as the target layer network, a network in the target model whose storage pressure is less than or equal to the storage pressure of the initial storage space of the target object and greater than the storage pressure of the proportional storage space; Adjusting the precision of target parameters of the target layer network according to the storage pressure of the target layer network.

6. The method according to claim 5, wherein adjusting the precision of target parameters of the target layer network according to the storage pressure of the target layer network includes: Performing descending order sorting on the target layer network according to the storage pressure of the target layer network to obtain a first sorted list; Adjust the accuracy of the target parameters of the target layer network through a first adjustment method according to the first sorted list.

7. The method according to claim 6, wherein the adjusting the accuracy of the target parameters of the target layer network through a first adjustment method according to the first sorted list comprises: Set the accuracy of the target layer network in the first sorted list to the target accuracy; Determine the first target layer network in the first sorted list; Obtain a storage pressure difference according to the storage pressure of the first target layer network and the storage pressure of the proportional storage space; Adjust the accuracy of the target parameters in the first target layer network according to the storage pressure difference, and delete the first target layer network in the first sorted list; Continue to execute the selection of the first target layer network in the first sorted list until the first sorted list is empty.

8. The method according to claim 7, wherein the adjusting the accuracy of the target parameters in the first target layer network according to the storage pressure difference comprises: Determine the current accuracy of the target parameters in the first target layer network; Adjust the storage pressure difference according to the current accuracy of the target parameters, determine the current storage pressure difference, and determine a target parameter list, wherein the target parameter list includes that the current accuracy of the target parameters is the target accuracy; When the current storage pressure difference is greater than a preset pressure threshold, adjust the accuracy of the target parameters in the parameter list.

9. The method according to claim 8, wherein the adjusting the storage pressure difference according to the current accuracy of the target parameters to determine the current storage pressure difference comprises: Determine whether there is a previous target parameter for the target parameter; If not, when the current accuracy of the target parameter is inconsistent with the target accuracy, subtract the storage pressure of the target parameter from the storage pressure difference to obtain the current storage pressure difference; If so, when the current accuracy of the target parameter is inconsistent with the target accuracy, subtract the storage pressure of the target parameter from the current storage pressure difference of the previous target parameter to obtain the updated current storage pressure difference.

10. The method according to claim 8, wherein the adjusting the accuracy of the target parameters in the parameter list when the current storage pressure difference is greater than a preset pressure threshold comprises: Sort the target parameters in the parameter list in descending order according to the storage pressure of the target parameters in the parameter list to obtain a second sorted list; Adjust the accuracy of the target parameters in the parameter list through a second adjustment method according to the second sorted list.

11. The method according to claim 10, wherein the adjusting the accuracy of the target parameters in the parameter list through a second adjustment method according to the second sorted list comprises: Determine the first target parameter in the second sorted list; Adjust the accuracy of the first target parameter to the initial accuracy; Determine the storage pressure saved by the first target layer network according to the initial accuracy; When the storage pressure saving is less than the current storage pressure difference, delete the first target parameter in the second sorted list; Continue to determine the first target parameter in the second sorted list until the second sorted list is empty.

12. A neural network model quantization device, comprising: A parameter determination module, configured to determine target parameters of multiple layers of a target model according to parameter attributes of initial parameters of a target model based on a neural network; A pressure determination module, configured to determine the storage pressure of multiple layers of the target model according to the target parameters, wherein adjusting the target parameters of the target layer network of the target model according to the storage pressure of the multiple layers of the network includes: determining the accuracy ratio of the multiple layers of the network according to the target accuracy of the target parameters of the multiple layers of the target model, and adjusting the target parameters of the target layer network of the target model according to the accuracy ratio of the multiple layers of the network, the storage space of the target object, and the storage pressure of the multiple layers of the network, wherein the target model runs on the target object, and the storage pressure is the storage pressure of the multiple layers of the target model on the target object; An adjustment module, configured to adjust the target parameters of the target layer network of the target model according to the storage pressure of the multiple layers of the network.

13. A computing device, comprising: A memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the neural network model quantization method according to any one of claims 1 to 11 are implemented.

14. A computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the steps of the neural network model quantization method according to any one of claims 1 to 11 are implemented.

Citation Information

Patent Citations

  • Neural network quantification method and device, neural network application method and device and computing equipment

    CN111126557A