Low-power-consumption neural network parameter determination method and device based on zero-order optimization
By adding the energy consumption during neural network inference on the digital memory computing chip to the loss function, and adjusting the model parameters using zero-order optimization and simulated annealing algorithm, the problem of high power consumption in high-precision scenarios in the existing technology is solved, and the parameter determination of low-power neural network is achieved.
Patent Information
- Application Number
- CN202510077319.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-17
- Publication Date
- 2025-06-13
AI Technical Summary
Existing digital in-memory computing technology consumes high power in high-precision scenarios and cannot effectively adapt to resource-constrained edge devices or energy-sensitive scenarios.
Using a zero-order optimization method, the energy consumption of the neural network when performing inference on the digital memory computing chip is added to the loss function. By iteratively adjusting the model parameters of the neural network, combined with a simulated annealing algorithm, the parameters are optimized to minimize the loss function.
It effectively reduces the power consumption of neural networks inference on digital in-memory computing chips, and is suitable for resource-constrained edge devices and energy-sensitive scenarios.
Smart Images

Figure CN120146130A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of neural networks, and in particular, to a low-power neural network parameter determination method and device based on zero-order optimization. Background Art
[0002] Currently, with the rapid development of artificial intelligence, neural networks are increasingly widely used in various computationally intensive tasks, such as image recognition, natural language processing, and speech recognition. However, when computing hardware under the traditional von Neumann architecture (such as CPUs and GPUs) processes neural network tasks, it is restricted by the frequent data transmission between the storage unit and the computing unit, and has the following problems: Limited computing speed: A large amount of data transmission leads to a system bandwidth bottleneck, and the performance of the computing unit cannot be fully utilized; High energy consumption: Frequent memory read and write operations not only take time but also consume a large amount of energy, which often exceeds the energy consumption of the computing unit itself in practical applications.
[0003] To solve the above problems, Computing In Memory (CIM) has been proposed as a new computing architecture. By integrating the storage unit and the computing unit, CIM avoids frequent data movement, thereby significantly improving computing efficiency and reducing energy consumption. In CIM technology, digital in-memory computing is mainly used to process neural network inference problems. However, although existing digital in-memory computing technologies can improve computing efficiency, their power consumption is high in high-precision scenarios and cannot well adapt to resource-constrained edge devices or energy-sensitive scenarios.
[0004] Therefore, there are defects in the prior art and it needs to be improved and developed. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a low-power neural network parameter determination method and device based on zero-order optimization in view of the above-mentioned defects of the prior art, aiming to solve the problem of high power consumption caused by using digital in-memory computing to process neural network inference in the prior art.
[0006] The technical solution adopted by the present invention to solve the technical problem is as follows:
[0007] In a first aspect, an embodiment of the present invention provides a low-power neural network parameter determination method based on zero-order optimization, and the method includes:
[0008] Obtain a training set of a neural network, where the neural network is deployed on a digital in-memory computing chip;
[0009] Input the training set into the neural network batch by batch;
[0010] Calculate the loss function for each round of training, where the loss function includes the accuracy loss and the energy consumption of the in-memory computing chip when performing the inference task.
[0011] Combine the zero-order optimization method and the simulated annealing algorithm, and iteratively adjust the model parameters of the neural network based on the loss function for each round of training until the training termination condition is met, to obtain the neural network parameters that satisfy the constraint conditions and minimize the loss function.
[0012] In one implementation, before obtaining the training set of the neural network, it further includes: constructing an energy consumption model, where the energy consumption model is used to calculate the energy consumption of the neural network when performing the inference task on the in-memory computing chip; the specific steps for constructing the energy consumption model include:
[0013] Establish a mapping relationship between the convolutional layer in the neural network and the in-memory computing macro in the in-memory computing chip.
[0014] Based on the mapping relationship, construct a memory cell energy consumption module and an adder tree energy consumption module. The memory cell energy consumption module is used to calculate the dynamic energy consumption generated by the memory cell during the convolution process, and the adder tree energy consumption module is used to calculate the dynamic energy consumption generated by the adder tree during the convolution operation.
[0015] Combine the memory cell energy consumption module and the adder tree energy consumption module to form an energy consumption model.
[0016] In one implementation, the establishing the mapping relationship between the convolutional layer in the neural network and the in-memory computing macro in the in-memory computing chip includes:
[0017] Divide the in-memory computing macros of the in-memory computing chip into libraries equal in number to the number of convolutional output channels in the neural network. Each library is used to store a K×K×Cin convolutional kernel, where K is the size of the convolutional kernel and Cin is the number of input channels.
[0018] Input the feature maps in the neural network into each library to achieve parallel computing of the convolutional kernel and the feature maps.
[0019] In one implementation, before inputting the training set into the neural network batch by batch, it further includes:
[0020] Insert pseudo-quantization nodes in each layer of the neural network.
[0021] In one implementation, the calculating the loss function for each round of training includes:
[0022] Calculate the accuracy loss based on the difference between the output value of the neural network model for the training data in each round of training and the true label value.
[0023] Calculate the quantization results of the activation values and weights for each layer using the pseudo - quantization nodes, and input the quantization results of the activation values and weights for each layer into the energy consumption model;
[0024] Calculate the energy consumption based on the energy consumption model to obtain the energy consumption of each round of training;
[0025] Perform calculations based on the accuracy loss, the energy consumption, and a preset coefficient to obtain the loss function of each round of training.
[0026] In one implementation, combining the zero - order optimization method and the simulated annealing algorithm, iteratively adjust the model parameters of the neural network based on the loss function of each round of training until the training termination condition is met, to obtain the neural network parameters that satisfy the constraint conditions and minimize the loss function, including:
[0027] Update the model parameters according to the gradient direction estimated by zero - order optimization to obtain the initial model parameters;
[0028] Expand the search range of the initial model parameters through the simulated annealing algorithm to obtain new model parameters;
[0029] Based on the new model parameters, perform constraint condition judgment to update or maintain the new model parameters according to the judgment result;
[0030] Repeat the steps of parameter update, search range expansion, and constraint condition judgment until the training termination condition is met, to obtain the neural network parameters that satisfy the constraint conditions and minimize the loss function.
[0031] In one implementation, the constraint conditions include accuracy constraint and energy consumption constraint; the performing constraint condition judgment based on the new model parameters to update or maintain the new model parameters according to the judgment result includes:
[0032] Judge whether the accuracy constraint is met when using the new model parameters to execute the validation set. If the accuracy constraint is not satisfied, adjust the new model parameters;
[0033] If the accuracy constraint is met and the energy consumption constraint is not met when using the new model parameters to execute the validation set, adjust the new model parameters;
[0034] If the accuracy constraint is met and the energy consumption constraint is met when using the new model parameters to execute the validation set, maintain the current parameters of the model.
[0035] In a second aspect, an embodiment of the present invention further provides a low - power neural network parameter determination device based on zero - order optimization, including:
[0036] A dataset acquisition module, configured to acquire a training set of a neural network, where the neural network is deployed on a digital in - memory computing chip;
[0037] A training set input module for inputting the training set into a neural network batch by batch;
[0038] A loss function calculation module for calculating the loss function of each round of training, where the loss function includes accuracy loss and the energy consumption of the in-memory computing chip when performing inference tasks;
[0039] A parameter determination module for iteratively adjusting the model parameters of the neural network based on the loss function of each round of training by combining the zero-order optimization method and the simulated annealing algorithm until the training termination condition is met, so as to obtain the neural network parameters that satisfy the constraint conditions and minimize the loss function.
[0040] In a third aspect, an embodiment of the present invention further provides a terminal, where the terminal includes: a memory, a processor, and a low-power neural network parameter determination program based on zero-order optimization stored in the memory and executable on the processor. When the low-power neural network parameter determination program based on zero-order optimization is executed by the processor, the steps of the above-mentioned low-power neural network parameter determination method based on zero-order optimization are implemented.
[0041] In a fourth aspect, an embodiment of the present invention further provides a computer-readable storage medium, where the computer-readable storage medium stores a low-power neural network parameter determination program based on zero-order optimization, and the low-power neural network parameter determination program based on zero-order optimization can be executed to implement the steps of the above-mentioned low-power neural network parameter determination method based on zero-order optimization.
[0042] The beneficial effects of the present invention: The present invention obtains the training set of the neural network, and the neural network is deployed on the in-memory computing chip; inputs the training set into the neural network batch by batch; calculates the loss function of each round of training, where the loss function includes accuracy loss and the energy consumption of the in-memory computing chip when performing inference tasks; combines the zero-order optimization method and the simulated annealing algorithm, and iteratively adjusts the model parameters of the neural network based on the loss function of each round of training until the training termination condition is met, so as to obtain the neural network parameters that satisfy the constraint conditions and minimize the loss function. By adding the energy consumption during neural network inference to the loss function for training to determine the parameters of the neural network, the present invention can effectively reduce the power consumption of the neural network during inference on the in-memory computing chip. Description of the Drawings
[0043] Figure 1 is a flowchart of a preferred embodiment of the low-power neural network parameter determination method based on zero-order optimization in the present invention.
[0044] Figure 2 is a schematic diagram of pseudo-quantization node insertion in the present invention.
[0045] Figure 3 It is a schematic diagram of the quantization process mapping in the present invention.
[0046] Figure 4 It is a schematic diagram of the structure of the convolutional layer in the present invention.
[0047] Figure 5 It is a schematic diagram of the principle of the simulated annealing algorithm in the present invention.
[0048] Figure 6 It is a schematic diagram of the flow of the simulated annealing algorithm in the present invention.
[0049] Figure 7 It is a schematic diagram of the model parameter update in the present invention.
[0050] Figure 8 It is a schematic diagram of the structure of the preferred embodiment of the low-power neural network parameter determination device based on zero-order optimization in the present invention.
[0051] Figure 9 It is a schematic block diagram of the terminal principle in the present invention. Detailed implementation manners
[0052] To make the objectives, technical solutions and advantages of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific examples described herein are only used to explain the present invention and are not used to limit the present invention.
[0053] Currently, with the rapid development of artificial intelligence, neural networks are increasingly widely used in various computationally intensive tasks, such as image recognition, natural language processing, and speech recognition. However, when computing hardware under the traditional von Neumann architecture (such as CPUs and GPUs) processes neural network tasks, it is restricted by the frequent data transfer between the storage unit and the computing unit, and there are the following problems: limited computing speed: a large amount of data transfer leads to a system bandwidth bottleneck, and the performance of the computing unit cannot be fully utilized; high energy consumption: frequent memory read and write operations not only take time but also consume a large amount of energy, which often exceeds the energy consumption of the computing unit itself in practical applications.
[0054] To solve the above problems, Computing In Memory (CIM) is proposed as a new computing architecture. By integrating the storage unit and the computing unit, CIM avoids frequent data movement, thereby significantly improving computing efficiency and reducing energy consumption. In CIM technology, digital CIM is mainly used to process neural network inference problems. However, although existing digital CIM technologies can improve computing efficiency, their power consumption is high in high-precision scenarios and cannot well adapt to resource-constrained edge devices or energy-sensitive scenarios.
[0055] In view of the above defects of the prior art, the present invention provides a method and apparatus for determining low-power neural network parameters based on zero-order optimization. The method includes: obtaining a training set of a neural network, where the neural network is deployed on a digital in-memory computing chip; inputting the training set into the neural network batch by batch; calculating the loss function for each round of training, where the loss function includes accuracy loss and the energy consumption of the digital in-memory computing chip when performing an inference task; combining the zero-order optimization method and the simulated annealing algorithm, and iteratively adjusting the model parameters of the neural network based on the loss function for each round of training until the training termination condition is satisfied, so as to obtain the neural network parameters that satisfy the constraint conditions and minimize the loss function. By adding the energy consumption during neural network inference to the loss function for training to determine the parameters of the neural network, the present invention can effectively reduce the power consumption of the neural network during inference on the digital in-memory computing chip.
[0056] Please refer to Figure 1 , the method for determining low-power neural network parameters based on zero-order optimization according to the embodiment of the present invention includes the following steps:
[0057] Step S100: Obtain a training set of a neural network, where the neural network is deployed on a digital in-memory computing chip.
[0058] Specifically, the training set of the present invention can be from a publicly available training set or a training set generated on demand, and its main purpose is to train the neural network, which is not limited herein.
[0059] When a CPU and a GPU with a traditional von Neumann architecture execute compute-intensive tasks, a vast amount of parameters need to be read from memory and then computed. The data transfer between the compute unit and the storage unit greatly limits the computing speed, and frequent read and write operations consume a large amount of energy, often exceeding the energy consumption of the compute unit itself. To solve this problem, in-memory computing has emerged. In-memory computing is divided into digital in-memory computing and analog in-memory computing. Analog in-memory computing mainly implements multiplication-accumulation operations on a memory-compute array based on physical laws (Ohm's law and Kirchhoff's law), and its area and power consumption overhead are small. However, since analog signals are sensitive to factors such as process, temperature, voltage, and noise, calculations based on analog signals are also very susceptible to these factors, thereby affecting the accuracy of calculation results. Therefore, analog in-memory computing is suitable for scenarios with small computing power and low power consumption. In some high-precision application scenarios, analog in-memory computing is not the best choice. Digital in-memory computing has characteristics such as high precision, strong noise immunity, and high flexibility, and has been widely used in scenarios such as large computing power, cloud computing, edge computing, and neural network inference. The neural network of the present invention is precisely deployed on a digital in-memory chip for inference. Digital in-memory computing uses all-digital circuits to perform calculations. Its multiplication and accumulation operations are both implemented through logic gates, and the weight data in memory is also entirely represented by binary data, facilitating the rewriting of stored values (analog in-memory computing requires changing each resistance value, while digital in-memory computing only needs to change the state of each bit, 0 or 1). This method can improve flexibility and computing accuracy.
[0060] Please refer to Figure 1 , the method for determining low-power neural network parameters based on zero-order optimization according to the embodiment of the present invention further includes the following steps:
[0061] Step S200: Input the training set into the neural network batch by batch.
[0062] Specifically, after obtaining the training set, it is input into the neural network batch by batch for training.
[0063] In one implementation, before inputting the training set into the neural network batch by batch, it further includes:
[0064] Insert pseudo-quantization nodes into each layer of the neural network.
[0065] Specifically, in order to simulate low-precision operations during training and reduce the precision loss when the model is deployed from high precision to low precision, pseudo-quantization nodes are inserted into each layer of the neural network before training, so as to perform quantization-aware training (QAT) subsequently. The present invention performs training by inserting pseudo-quantization operations at the positions of weights and activations and taking into account the errors caused by quantization during the training process. Through quantization-aware training, the quantized values of the weights and activation values of each layer can be extracted during the training process. The schematic diagram is as shown in Figure 2 shown. Starting from the lower left corner of the figure, the "activations" (activation values) enter the convolutional layer after passing through the pseudo-quantization operation. The output of the convolutional layer is added to the "biases" (biases), and then passes through the ReLU activation function. The output of the ReLU passes through the pseudo-quantization operation again to become the final "output" (output). The "weights" (weights) also enter the convolutional layer after passing through the pseudo-quantization operation. The oval nodes represent the pseudo-quantization operations. Every time a pseudo-quantization node is passed through, the data will be quantized and de-quantized once.
[0066] During the quantization process, the quantization parameters are usually determined according to the range of the input data. Please refer to Figure 3 . The quantization process is to map the data from the floating-point space to the integer space.
[0067] During the quantization and de-quantization process, quantization is the process of converting high-precision floating-point numbers (such as float32) into low-precision integers (such as int8), the purpose of which is to reduce the storage and calculation overhead while trying to maintain the precision of the model. In the present invention, the specific steps of quantization and de-quantization are as follows: First, according to the input data (such as the weights or activation values in the neural network), find the maximum value r max and the minimum value r min . Then use formula (1) to calculate the quantization scale S, and formula (1) is as follows:
[0068]
[0069] where q max is the maximum value of the predefined data type, and q min is the minimum value of the predefined data type.
[0070] Secondly, use formula (2) to calculate the quantization zero point Z, and formula (2) is as follows:
[0071]
[0072] Thirdly, use formula (3) to calculate the quantized data q, and formula (3) is as follows:
[0073]
[0074] Among them, r float represents the original floating-point number, and round represents the rounding operation.
[0075] Finally, the dequantized data is calculated using formula (4), and formula (4) is as follows:
[0076] r quantized = S(q - Z) (4)
[0077] Among them, r quantized is the floating-point number after dequantization.
[0078] Quantization-aware training introduces the quantization process into the training, so the error caused by quantization can be considered during training. The present invention extracts the quantized activation values and weight parameters of each layer of the neural network through quantization-aware training, which can be used to calculate the energy consumption in the subsequent input energy consumption model.
[0079] Please refer to Figure 1 , the method for determining low-power neural network parameters based on zero-order optimization described in the embodiments of the present invention further includes the following steps:
[0080] Step S300, calculate the loss function of each round of training, and the loss function includes the accuracy loss and the energy consumption of the digital in-memory computing chip when performing the inference task.
[0081] Specifically, in the prior art, when training a neural network model, more attention is often paid to how to train the optimal weight parameters, rather than considering how to reduce the power consumption of the hardware. The present invention adds the energy consumption of the digital in-memory computing chip when performing the inference task to the loss function, that is, takes the energy consumption as an optimization goal. By continuous iterative training, the loss function is minimized, and then model parameters with less loss are obtained. Subsequently, when using the determined model parameters in the digital in-memory computing chip to perform the inference task, the power consumption can be effectively reduced.
[0082] In one implementation, before obtaining the training set of the neural network, it further includes: constructing an energy consumption model, where the energy consumption model is used to calculate the energy consumption of the neural network when performing the inference task on the digital in-memory computing chip; the specific steps of constructing the energy consumption model include:
[0083] Establish a mapping relationship between the convolutional layer in the neural network and the in-memory computing macro in the digital in-memory computing chip;
[0084] Based on the mapping relationship, a storage basic unit energy consumption module and an adder tree energy consumption module are constructed. The storage basic unit energy consumption module is used to calculate the dynamic energy consumption generated by the storage basic unit during the convolution process, and the adder tree energy consumption module is used to calculate the dynamic energy consumption generated by the adder tree during the convolution operation;
[0085] Combining the storage basic unit energy consumption module and the adder tree energy consumption module forms an energy consumption model.
[0086] Specifically, the present invention uses the constructed energy consumption model to calculate the energy consumption of each round of training. The steps of establishing the mapping relationship between the convolutional layer in the neural network and the in-memory computing macro in the digital in-memory computing chip include: dividing the in-memory computing macros of the digital in-memory computing chip into libraries equal in number to the number of convolutional output channels in the neural network. Each library is used to store a K×K×Cin convolutional kernel, where k is the size of the convolutional kernel and Cin is the number of input channels; inputting the feature maps in the neural network into each library to achieve parallel computing of the convolutional kernel and the feature maps. By dividing the in-memory computing macros of the digital in-memory computing chip into libraries equal in number to the number of convolutional output channels in the neural network, each convolutional kernel can have a dedicated storage and computing area. In this way, during large-scale convolution operations, multiple convolutional kernels perform calculations on different feature map regions synchronously, greatly shortening the overall calculation time. In addition, since the data is processed locally in each library of the in-memory computing macro, there is no need for a large amount of high-energy-consuming data transmission between the storage unit and the computing unit as in the traditional architecture, so that the energy consumption is mainly concentrated on the effective calculation process.
[0087] In one implementation, the structural schematic diagram of a single convolutional layer can be as Figure 4 shown. The size of the feature map of the input layer is 100×80×4, the convolutional kernel of the weight parameter convolutional layer is 3×3×4, and the number is 2. The size of the feature map of the output layer is 98×78×2. Among them, Hin is the height of the input feature map, Win is the width of the input feature map, Cin is the number of channels of the input feature map, K is the size of the convolutional kernel, Hout is the height of the output feature map, Wout is the width of the output feature map, and Cout is the number of channels of the output feature map.
[0088] In digital in-memory computing, "toggle" refers to the change in the data state in a digital computing circuit, i.e., from 0 to 1 or from 1 to 0. The toggle rate is the number of data state changes per unit time, and it is a key metric for measuring data activity. Toggles have a significant impact on power consumption because each state change involves the movement of charge, thus consuming energy. The power consumption of digital in-memory computing mainly consists of dynamic power and static power. Dynamic power is generated due to the toggles of the storage cell states (i.e., data read and write operations), and it is directly related to the toggle rate. Each toggle requires charging or discharging the nodes of the digital circuit, which involves the charge and discharge process of capacitors, thus consuming energy. Therefore, the higher the toggle rate, the more times the nodes of the digital circuit are charged and discharged, and the greater the dynamic power. Static power is related to the leakage current of the storage cells, and it also exists when the digital computing circuit maintains the data state unchanged. Therefore, in the present invention, the energy consumption is the dynamic power, i.e., the power consumption generated due to data inversion.
[0089] Since the mapping relationship is determined, the data to be processed by each storage basic unit can be determined. The hardware structure of the storage basic unit includes two parts: a storage array (array), which consists of multiple multiplication units (bitcell) and completes K×K×Cin×N multiplication operations, where Cin is the number of input channels and N is the weight quantization bit number, and an adder tree (adder tree), which consists of full adders (full adder) and is responsible for accumulating the K×K×Cin multiplication results to complete the convolution operation.
[0090] The energy consumption of the storage basic unit in the energy consumption model can calculate the energy consumption based on the input pattern of the computing macro. For example, when the input pattern of the in-memory computing macro is "most significant bit first, serial input", since each storage basic unit can only perform a single-bit multiplication operation, the dynamic energy consumption will be affected by the flipping of the activation value. The flipping of the activation value mainly comes from the following two aspects: One is the flipping during the serial input of the activation value: Since the storage basic unit can only perform single-bit multiplication, when the input of the neural network is quantized to N bits, N storage units need to process each bit separately, and each bit may flip multiple times. For example, assume an 8-bit input with an activation value of "01010101", then it will flip 8 times during the input process. The other is the flipping when the convolutional kernel slides: When the convolutional kernel slides on the feature map, it may cause the least significant bit (LSB) and the most significant bit (MSB) of the activation value to flip. For example, before the convolutional operation, the least significant bit of the activation value at a certain position on the feature map is "0", and after the convolutional kernel slides, the most significant bit of the activation value at this position becomes "1", and this kind of bit flipping will also affect the dynamic energy consumption of the storage unit. After determining the flipping situation of the activation value using the energy consumption module of the storage basic unit, the dynamic energy consumption of the storage basic unit array can be determined using the Energy Look-Up Table of the storage basic unit. This energy look-up table contains the following column information: activation value [t - 1]: the activation value at the previous moment; activation value [t]: the activation value at the current moment; weight: the weight of the neural network; energy: the dynamic energy consumption obtained by looking up the table according to different combinations of activation value flipping and weight. By looking up the table, the dynamic energy consumption of the storage basic unit array in each cycle can be accurately calculated, providing a basis for the energy consumption analysis of the entire neural network.
[0091] The dynamic energy of the adder tree can be calculated using the adder tree energy consumption module in the energy consumption model. Since the adder tree is composed of full adders. First, obtain the energy look-up table of the full adder to determine the energy of the full adder. Here, the energy look-up table of the full adder contains the following column information: ABCin[t - 1], ABCin[t], energy. ABCin[t - 1] represents the input at the previous moment, ABCin[t] represents the input at the current moment, and energy is the energy at the current moment. Then, the adder tree energy consumption module can obtain the energy consumption of the adder tree based on the structure of the adder tree and the energy of the full adder.
[0092] Finally, the energy consumption model adds the energy consumption of the adder tree and the energy consumption of all storage basic units to obtain the energy consumption of the neural network when performing the inference task on the digital in-memory computing chip.
[0093] In one implementation, calculating the loss function for each round of training includes:
[0094] Calculate the precision loss based on the difference between the output value of the training data for each round of training by the neural network model and the true label value.
[0095] Use the pseudo-quantization nodes to calculate the quantization results of the activation values and weights for each layer, and input the quantization results of the activation values and weights for each layer into the energy consumption model.
[0096] Calculate the energy consumption according to the energy consumption model to obtain the energy consumption for each round of training.
[0097] Perform calculations based on the precision loss, the energy consumption, and a preset coefficient to obtain the loss function for each round of training.
[0098] Specifically, when training a traditional neural network model, energy consumption is mostly not considered. In the present invention, energy consumption is added to the loss function. The formula for the loss function is: loss = loss accuracy + λEnergy, where loss accuracy is the precision loss, Energy is the energy consumption, and λ is the coefficient.
[0099] In one implementation, loss accuracy is the binary cross-entropy loss function, and the formula is: where m is the number of samples, y is the true label of the sample, is the probability that the model predicts that the sample belongs to the positive class.
[0100] Please refer to Figure 1 , the method for determining low-power neural network parameters based on zero-order optimization described in the embodiments of the present invention further includes the following steps:
[0101] Step S400: Combine the zero-order optimization method and the simulated annealing algorithm, and iteratively adjust the model parameters of the neural network based on the loss function for each round of training until the training termination condition is satisfied, so as to obtain the neural network parameters that satisfy the constraint conditions and minimize the loss function.
[0102] Specifically, since the loss function cannot perform explicit derivation on the parameters after adding energy consumption, zero-order optimization is used for parameter update. Zero-order optimization (Zero-order Optimization, ZOO) is an optimization method that does not require gradient information. It depends on the value of the objective function in function optimization, rather than its derivative, and is applicable to situations where the objective function is non-differentiable, non-differentiable, or the gradient is difficult to calculate. The present invention effectively reduces the complexity and difficulty of gradient calculation by using zero-order optimization.
[0103] During the neural network training process, the gradient of the parameters is usually obtained through backpropagation, and according to the definition of the derivative: The original function can be slightly perturbed to obtain a new function value. Subtracting the original function value from the new function value and then dividing by the perturbation value can approximately obtain the gradient of the function at that point. The above principle is used in zero-order optimization. It represents the slope of the straight line passing through the two points (x, f(x)) and (x + ε, f(x + ε)). When the perturbation value ε approaches 0, this slope approaches the slope of the tangent line of the function f(x) at the point (x, f(x)), that is, the derivative of f(x) at x.
[0104] There are two common types of zero-order optimization: coordinate gradient estimation and stochastic gradient estimation. Coordinate gradient estimation is to perturb only one parameter while keeping other parameters unchanged when estimating the gradient each time, obtaining an approximation of the gradient of one parameter. This method is relatively accurate, but when the number of parameters is very large, it is necessary to perturb d times (d is the total number of parameters) and then calculate the loss function value d times to obtain all the gradients, with a large and cumbersome computational amount.
[0105] The present invention adopts stochastic gradient estimation, and all the gradient estimation values of the parameters can be obtained in one estimation. The formula is as follows: Among them, Both θ and u are d-dimensional vectors. represents all the gradient estimation values of the parameters, θ represents all the parameter values, u is a randomly generated vector obeying the Gaussian distribution, which can represent a direction vector in the d-dimensional space, μ is a very small scalar, and L represents the objective function.
[0106] When μ is very small, the directional derivative L u ′(θ) can be expressed as follows: Since the directional derivative = direction vector · gradient, so Multiply both sides by u and take the expectation (since u follows the Gaussian distribution, so E(uTu) = 1) to obtain:
[0107] So is an unbiased estimate of the gradient dθ of all the parameters. Therefore, by adding a Gaussian perturbation to the original parameters, the gradient value of the parameters can be estimated. Due to the limitations of a single perturbation, in the present invention, positive and negative perturbations are adopted, that is, the gradient value of the parameters can be estimated. Based on the gradient, update the weight θ = θ - adθ, where a is the learning rate.
[0108] After initially calculating the gradient value of the parameters using zero-order optimization, the simulated annealing algorithm is used to further expand the search range of the parameters. Suppose we want to find such as Figure 5The minimum value of the function f(x) shown. Taking the traditional gradient descent method as an example, if the initialized solution is at point D, according to the gradient descent method, the solution will only be optimized in the direction of the gradient descent and will eventually reach point A, the local optimal solution, instead of point B, the global optimal solution. This method is more suitable for optimizing functions with only one extreme value. In the simulated annealing algorithm, if the initial solution is at D, the Metropolis criterion can be used to accept a new solution that is worse than the current solution with a certain probability, that is Figure 5 point E in. Based on the new solution, the Metropolis criterion is used again to determine whether to accept the next solution. If accepted, the next solution is used as the new solution; if rejected, a new solution is searched for and iterated in turn. The core idea of the simulated annealing algorithm is to allow the algorithm to accept deteriorated solutions to jump out of the local optimal solution.
[0109] The simulated annealing algorithm consists of two parts, namely the Metropolis algorithm and the annealing process, corresponding to the inner loop and the outer loop respectively. The Metropolis algorithm is about how to jump out when trapped in the local optimal solution. In 1953, Metropolis proposed importance sampling, that is, accepting a new state with a probability instead of using a completely deterministic rule (simply accepting or rejecting), which is called the Metropolis criterion and can significantly reduce the computational amount. The specific content of the Metropolis criterion is as follows: Assume that the previous state is x(n), the system is perturbed, and the state becomes x(n + 1). Correspondingly, the system energy changes from E(n) to E(n + 1). Then the probability p of accepting the system changing from x(n) to x(n + 1) is defined as:
[0110]
[0111] The above formula shows that when the state changes, if the energy of the system decreases, this state transition is accepted (p = 1). If the energy increases, it means that the system is farther from the position of the global optimal solution (i.e., the point with the lowest energy), but at this time the algorithm will not immediately discard it, but make a probability judgment: a uniformly distributed random number ε is generated in the interval [0, 1]. If ε < p, this state transition is also accepted, otherwise it is rejected and the next perturbation is performed, and so on. It can be seen from the above formula that when E(n + 1) ≥ E(n), if the energy of the next state E(n + 1) is much larger than the energy of the previous state E(n), then p will be very small and the acceptance probability will be very low; on the contrary, if E(n + 1) is very close to E(n), p will be close to 1 and the acceptance probability will be very high. By setting an iteration number L, when the iteration reaches L times, it stops, and the whole process is called an inner loop. After the end of an inner loop, the solution that best meets the conditions among all the accepted solutions is used as the optimal solution of this inner loop.
[0112] In addition, there is another parameter T that affects the acceptance probability. During one inner loop, T is a fixed value, and what affects parameter T is the outer loop, that is, the annealing process. During the solid annealing process, T refers to temperature, and as time goes by, T will continuously decrease. In the simulated annealing algorithm, an initial temperature T0, a stopping temperature Tf (i.e., when T drops to Tf, the annealing is completed), and an annealing rate (i.e., the way the temperature decreases) will be set. At temperature T0, an initial solution is randomly generated and used as a starting point for one inner loop. After the inner loop ends, T0 is decreased once to obtain T1, and then another inner loop is carried out with the decreased T1 as the parameter, and so on iteratively. When T drops to the stopping temperature Tf, it terminates, and the whole operation is called one outer loop. The optimal solution obtained from each inner loop is taken as the optimal again to obtain the final optimal solution. The simplest way of temperature decrease is exponential decrease, that is, T(n + 1) = αT(n), where α is a positive number less than 1, and generally takes a value between 0.8 and 0.99. When T continuously decreases, the probability of accepting a solution that makes the objective function increase will gradually decrease. When the temperature approaches 0, only solutions that make the objective function decrease can be accepted.
[0113] Figure 6 Shows the basic steps of the entire simulated annealing algorithm. Step 1: Randomly generate an initial solution ω and calculate the objective function f(ω); Step 2: Generate a new solution ω' through perturbation and calculate the objective function f(ω') of the new solution; Step 3: Determine whether to accept the new solution. If f(ω') - f(ω) ≤ 0, then accept the new solution. If f(ω') - f(ω) > 0, then decide whether to accept the new solution according to the Metropolis criterion; Step 4: Determine whether the iteration times are reached. If the iteration times are reached, go to the next step. If the iteration times are not reached, return to Step 2 to continue generating new solutions; Step 5: Determine whether the termination condition is satisfied. If the termination condition is satisfied, the operation ends and the optimal solution is output. If the termination condition is not satisfied, the temperature is slowly decreased, the iteration times are reset, and return to Step 2 to continue iterating.
[0114] In one implementation, the combination of the zero-order optimization method and the simulated annealing algorithm iteratively adjusts the model parameters of the neural network based on the loss function of each round of training until the training termination condition is met, and obtains the neural network parameters that satisfy the constraint conditions and minimize the loss function, including:
[0115] Update the model parameters according to the gradient direction estimated by the zero-order optimization to obtain the initial model parameters;
[0116] Expand the search range of the initial model parameters through the simulated annealing algorithm to obtain new model parameters;
[0117] Based on the new model parameters, perform constraint condition judgment to update or maintain the new model parameters according to the judgment result;
[0118] The steps of parameter updating, search range expansion and constraint condition judgment are repeatedly performed until the training termination conditions are met, and the neural network parameters that meet the constraints and minimize the loss function are obtained.
[0119] Specifically, the zero-order optimization method does not need to calculate the exact gradient, but estimates the gradient direction by sampling the function value, which greatly reduces the computational complexity. Especially when dealing with large-scale neural networks, traditional optimization methods based on second-order derivatives often face huge computational and memory requirements, while zero-order optimization can quickly give a rough gradient direction, allowing the model parameters to quickly take the first step towards the optimization goal, saving a lot of time for subsequent fine-tuning. Then, the introduction of the simulated annealing algorithm further expanded the parameter search space. The simulated annealing algorithm simulates the solid annealing process and allows a certain probability to accept a poor solution in the initial stage, which avoids the optimization process from falling into a local optimal solution too early. After the initial model parameters are obtained through zero-order optimization, the simulated annealing algorithm can explore a wider parameter area around it and tap into potential better solutions.
[0120] In one implementation, the constraint condition includes an accuracy constraint and an energy consumption constraint; and the constraint condition judgment based on the new model parameter to update or maintain the new model parameter according to the judgment result includes:
[0121] Determine whether the accuracy constraint is met when executing the validation set using the new model parameters, and if the accuracy constraint is not met, adjust the new model parameters;
[0122] If the accuracy constraint is satisfied and the energy consumption constraint is not satisfied when executing the validation set using the new model parameters, adjusting the new model parameters;
[0123] If the accuracy constraint and the energy consumption constraint are met when executing the validation set using the new model parameters, the current parameters of the model are maintained.
[0124] Specifically, the accuracy constraint is an accuracy threshold D%. The energy consumption constraint is that the energy consumption of the current round of training of the neural network is less than the energy consumption of the previous round of training.
[0125] In one implementation, in order to improve the training efficiency, the training is continued using the pre-trained model. Figure 7 As shown in the figure, the energy consumption model is constructed first. After the energy consumption model is constructed, the energy consumption is calculated in combination with the energy consumption model when training the neural network. The model parameters are updated through zero-order optimization, and the coverage space is expanded through the simulated annealing algorithm. The parameters and energy consumption of the model are recorded only when the accuracy constraint and the energy consumption constraint are met. When the accuracy constraint is not met or the accuracy constraint is met but the energy consumption constraint is not met, the parameters are adjusted continuously.
[0126] In summary, in the present invention, quantization-aware training is added during the training of the neural network, and the energy consumption of the current neural network is calculated in combination with the energy consumption model. Furthermore, by setting the loss function with the energy consumption as an optimization objective, zero-order optimization and simulated annealing search are performed to find the neural network parameters with the lowest energy consumption under the condition of meeting the accuracy requirements, thereby reducing the power consumption of the neural network on the digital in-memory computing chip.
[0127] In one embodiment, as Figure 8 shown, based on the above method for determining low-power neural network parameters based on zero-order optimization, the present invention also correspondingly provides a device for determining low-power neural network parameters based on zero-order optimization, including:
[0128] A dataset acquisition module 100, configured to acquire a training set of the neural network, where the neural network is deployed on a digital in-memory computing chip;
[0129] A training set input module 200, configured to input the training set into the neural network batch by batch;
[0130] A loss function calculation module 300, configured to calculate the loss function for each round of training, where the loss function includes accuracy loss and the energy consumption of the digital in-memory computing chip when performing an inference task;
[0131] A parameter determination module 400, configured to iteratively adjust the model parameters of the neural network based on the loss function for each round of training by combining the zero-order optimization method and the simulated annealing algorithm until the training termination condition is met, so as to obtain the neural network parameters that meet the constraint conditions and minimize the loss function.
[0132] In one embodiment, before acquiring the training set of the neural network, it further includes: constructing an energy consumption model, where the energy consumption model is used to calculate the energy consumption of the neural network when performing an inference task on the digital in-memory computing chip; the device further includes:
[0133] A mapping relationship establishment unit, configured to establish a mapping relationship between the convolutional layer in the neural network and the in-memory computing macro in the digital in-memory computing chip;
[0134] A first module construction unit, configured to construct a storage basic unit energy consumption module and an adder tree energy consumption module based on the mapping relationship, where the storage basic unit energy consumption module is used to calculate the dynamic energy consumption generated by the storage basic unit during the convolution process, and the adder tree energy consumption module is used to calculate the dynamic energy consumption generated by the adder tree during the convolution operation;
[0135] A model formation unit, configured to form an energy consumption model by combining the storage basic unit energy consumption module and the adder tree energy consumption module.
[0136] In one embodiment, the device further includes:
[0137] A library partitioning unit, configured to partition the in-memory computing macros of the digital in-memory computing chip into libraries equal in number to the number of convolutional output channels in the neural network. Each library is used to store a K×K×Cin convolutional kernel, where K is the size of the convolutional kernel and Cin is the number of input channels;
[0138] A feature map input unit, configured to input the feature maps in the neural network into each library to implement parallel computing of the convolutional kernel and the feature map.
[0139] In one embodiment, the device further includes:
[0140] A node insertion unit, configured to insert pseudo-quantization nodes in each layer of the neural network.
[0141] In one embodiment, the device further includes:
[0142] A first calculation unit, configured to calculate the accuracy loss based on the difference between the output value of the training data for each round of training of the neural network model and the true label value;
[0143] A data input unit, configured to calculate the quantization results of the activation values and weights of each layer by using the pseudo-quantization nodes, and input the quantization results of the activation values and weights of each layer into the energy consumption model;
[0144] An energy consumption calculation unit, configured to calculate the energy consumption according to the energy consumption model to obtain the energy consumption for each round of training;
[0145] A loss function calculation unit, configured to perform calculations based on the accuracy loss, the energy consumption, and a preset coefficient to obtain the loss function for each round of training.
[0146] In one embodiment, the device further includes:
[0147] A first parameter update unit, configured to update the model parameters according to the gradient direction estimated by zero-order optimization to obtain the initial model parameters;
[0148] A second parameter update unit, configured to expand the search range of the initial model parameters through a simulated annealing algorithm to obtain new model parameters;
[0149] A constraint condition judgment unit, configured to perform constraint condition judgment based on the new model parameters to update or maintain the new model parameters according to the judgment result;
[0150] A parameter determination unit, configured to repeatedly execute the steps of parameter update, search range expansion, and constraint condition judgment until the training termination condition is met, to obtain the neural network parameters that satisfy the constraint conditions and minimize the loss function.
[0151] In one embodiment, the constraint conditions include accuracy constraint and energy consumption constraint; the device further includes:
[0152] A first judgment unit, configured to judge whether the accuracy constraint is met when the verification set is executed using the new model parameters, and if the accuracy constraint is not satisfied, adjust the new model parameters;
[0153] A second judgment unit, configured to adjust the new model parameters if the accuracy constraint is met and the energy consumption constraint is not met when the verification set is executed using the new model parameters;
[0154] A third judgment unit, configured to maintain the current parameters of the model if the accuracy constraint is met and the energy consumption constraint is met when the verification set is executed using the new model parameters.
[0155] Based on the above embodiment, the present invention further provides a terminal, and its principle block diagram can be as Figure 9 shown. The above terminal includes a processor, a memory, a network interface, and a display screen connected through a device bus. Among them, the processor of the terminal is used to provide computing and control capabilities. The memory of the terminal includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating device and a low-power neural network parameter determination program based on zero-order optimization. The internal memory provides an environment for the operation of the operating device and the low-power neural network parameter determination program in the non-volatile storage medium. The network interface of the terminal is used to communicate with an external terminal through a network connection. When the low-power neural network parameter determination program based on zero-order optimization is executed by the processor, the steps of any one of the above low-power neural network parameter determination methods based on zero-order optimization are implemented. The display screen of the terminal can be a liquid crystal display screen or an electronic ink display screen.
[0156] Those skilled in the art can understand that Figure 9 the principle block diagram shown in
[0157] is only a block diagram of some structures related to the solution of the present invention, and does not constitute a limitation on the terminal to which the solution of the present invention is applied. The specific terminal may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0158] An embodiment of the present invention further provides a computer-readable storage medium. A low-power neural network parameter determination program based on zero-order optimization is stored on the computer-readable storage medium. When the low-power neural network parameter determination program based on zero-order optimization is executed by a processor, the steps of any one of the low-power neural network parameter determination methods provided by the embodiments of the present invention are implemented.
[0159] It should be understood that the sequence numbers of the steps in the above embodiments do not imply the order of execution. The execution order of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.
[0160] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the division of the above function units and modules is used as an example. In practical applications, the above functions can be allocated to different function units and modules according to needs, that is, the internal structure of the above device can be divided into different function units or modules to complete all or part of the functions described above. The function units and modules in the embodiments can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software function unit. In addition, the specific names of the function units and modules are only for the convenience of mutual distinction and do not limit the protection scope of the present invention. The specific working processes of the units and modules in the above device can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0161] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0162] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or by a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.
[0163] In the embodiments provided by the present invention, it should be understood that the disclosed device / terminal device and method can be implemented in other ways. For example, the device / terminal device embodiments described above are only illustrative. For example, the above-mentioned division of modules or units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed.
[0164] The above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not deviate from the spirit and scope of the technical solutions of the present invention, and should all be included in the protection scope of the present invention.
Claims
1. A method for determining parameters of a low-power neural network based on zero-order optimization, characterized in that: The method comprises: Obtaining a training set of a neural network, wherein the neural network is deployed on a digital in-memory computing chip; Inputting the training set into the neural network in batches; Calculate the loss function for each round of training, where the loss function includes the accuracy loss and the energy consumption of the digital in-memory computing chip when performing the inference task; Combining the zero-order optimization method and the simulated annealing algorithm, the model parameters of the neural network are iteratively adjusted based on the loss function of each round of training until the training termination condition is met, thereby obtaining the neural network parameters that meet the constraints and minimize the loss function.
2. The method for determining low-power neural network parameters based on zero-order optimization according to claim 1, characterized in that: Before obtaining the training set of the neural network, the method further includes: constructing an energy consumption model, wherein the energy consumption model is used to calculate the energy consumption of the neural network when performing reasoning tasks on the digital memory computing chip; the specific steps of constructing the energy consumption model include: Establish the mapping relationship between the convolutional layer in the neural network and the in-memory computing macro in the digital in-memory computing chip; Based on the mapping relationship, a storage basic unit energy consumption module and an addition tree energy consumption module are constructed, wherein the storage basic unit energy consumption module is used to calculate the dynamic energy consumption generated by the storage basic unit during the convolution process, and the addition tree energy consumption module is used to calculate the dynamic energy consumption generated by the addition tree during the convolution operation; The energy consumption model is formed by combining the storage basic unit energy consumption module and the addition tree energy consumption module.
3. The method for determining low-power neural network parameters based on zero-order optimization according to claim 2, characterized in that: The step of establishing a mapping relationship between a convolutional layer in a neural network and an in-memory computing macro in a digital in-memory computing chip includes: The in-memory computing macro of the digital in-memory computing chip is divided into banks equal to the number of convolution output channels in the neural network. Each bank is used to store a convolution kernel of size K×K×Cin, where k is the size of the convolution kernel and cin is the number of input channels. The feature maps in the neural network are input into various libraries to achieve parallel calculation of convolution kernels and feature maps.
4. The method for determining low-power neural network parameters based on zero-order optimization according to claim 2, characterized in that: Before inputting the training set into the neural network in batches, the method further includes: Insert pseudo-quantization nodes in each layer of the neural network.
5. The method for determining low-power neural network parameters based on zero-order optimization according to claim 4, characterized in that: The calculation of the loss function for each round of training includes: Calculate the accuracy loss based on the difference between the output value of the neural network model for the training data in each round of training and the true label value; Calculating the quantization results of the activation values and weights of each layer using the pseudo quantization nodes, and inputting the quantization results of the activation values and weights of each layer into the energy consumption model; Calculating energy consumption according to the energy consumption model to obtain energy consumption for each round of training; The loss function for each round of training is obtained by performing calculations based on the accuracy loss, the energy consumption, and a preset coefficient.
6. The method for determining low-power neural network parameters based on zero-order optimization according to claim 1, characterized in that: The zero-order optimization method and the simulated annealing algorithm are combined to iteratively adjust the model parameters of the neural network based on the loss function of each round of training until the training termination condition is met, and the neural network parameters that meet the constraint conditions and minimize the loss function are obtained, including: Update the model parameters according to the gradient direction estimated by zero-order optimization to obtain the initial model parameters; The search range of the initial model parameters is expanded by the simulated annealing algorithm to obtain new model parameters; Perform constraint condition judgment based on the new model parameters to update or maintain the new model parameters according to the judgment result; The steps of parameter updating, search range expansion and constraint condition judgment are repeatedly performed until the training termination conditions are met, and the neural network parameters that meet the constraints and minimize the loss function are obtained.
7. The method for determining low-power neural network parameters based on zero-order optimization according to claim 6, characterized in that: The constraint conditions include accuracy constraint and energy consumption constraint; the constraint condition judgment based on the new model parameters to update or maintain the new model parameters according to the judgment result includes: Determine whether the accuracy constraint is met when executing the validation set using the new model parameters, and if the accuracy constraint is not met, adjust the new model parameters; If the accuracy constraint is satisfied and the energy consumption constraint is not satisfied when executing the validation set using the new model parameters, adjusting the new model parameters; If the accuracy constraint and the energy consumption constraint are met when executing the validation set using the new model parameters, the current parameters of the model are maintained.
8. A low-power neural network parameter determination device based on zero-order optimization, characterized in that: include: A data set acquisition module, used to acquire a training set of a neural network, wherein the neural network is deployed on a digital in-memory computing chip; A training set input module, used to input the training set into the neural network in batches; A loss function calculation module, used to calculate the loss function of each round of training, wherein the loss function includes the precision loss and the energy consumption of the digital in-memory computing chip when performing the inference task; The parameter determination module is used to combine the zero-order optimization method and the simulated annealing algorithm to iteratively adjust the model parameters of the neural network based on the loss function of each round of training until the training termination condition is met, so as to obtain the neural network parameters that meet the constraint conditions and minimize the loss function.
9. A terminal, characterized in that: The terminal includes a memory, a processor, and a low-power neural network parameter determination program based on zero-order optimization stored in the memory and executable on the processor. When the low-power neural network parameter determination program based on zero-order optimization is executed by the processor, the steps of the low-power neural network parameter determination method based on zero-order optimization as described in any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a low-power neural network parameter determination program based on zero-order optimization. When the low-power neural network parameter determination program based on zero-order optimization is executed by a processor, the steps of the low-power neural network parameter determination method based on zero-order optimization as described in any one of claims 1 to 7 are implemented.
Citation Information
Cited By
Low-bit neural network reasoning method and system for microcontroller
CN120654836A