An optimization method for a neural network and related devices

By introducing a new quantitative model into the neural network, the binarization of each layer's weight matrix is ​​related to the front layer weight matrix, which solves the problem of large quantization error in the existing technology and realizes more efficient neural network training and use.

CN111950700BActive Publication Date: 2025-05-27HUAWEI TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202010650726.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-07-06
Publication Date
2025-05-27
Estimated Expiration
2040-07-06

AI Technical Summary

Technical Problem

The application of existing neural networks on edge devices and end-side devices is limited by storage space and calculation amount. Although binary neural networks have high compression rates and fast computing speed, the binaryization of each layer's weight matrix alone leads to large quantization errors.

Method used

Through a new quantization model, the weight matrix of each layer of the neural network is binarized, so that the value of the adjusted weight matrix of each layer is correlated with the value of the unadjusted weight matrix of the previous layers, thereby reducing the quantization error.

Benefits of technology

This optimization method makes the training and use of neural networks more efficient, reduces quantization errors, and makes the application of the model more suitable on edge devices and end-side devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111950700B_ABST
    Figure CN111950700B_ABST
Patent Text Reader

Abstract

Embodiments of this application disclose an optimization method for a neural network and related devices, which can be applied to the field of computer vision in the field of artificial intelligence (such as, image super-resolution reconstruction), etc. The method includes: binarizing the weight matrix / feature representation (or called feature map, activation value) of the neural network through a new quantization model. Specifically, the first quantization model is used to obtain the second weight matrix of the m-th layer of the neural network according to the m first weight matrices of the 1st layer to the m-th layer of the neural network, and the second quantization model is used to obtain the second feature representation of the m-th layer of the neural network according to the m first feature representations of the 1st layer to the m-th layer. This optimization method makes the values of the weight matrix / feature representation of each layer not only related to itself, but also related to the weight matrix / feature representation of other layers, reduces the quantization error, makes the training and use of the neural network more efficient. At the same time, compared with the existing binary neural network, the accuracy of image information processing is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of machine learning, and in particular, to an optimization method for a neural network and related devices. Background Art

[0002] A neural network is a machine learning technology that simulates the human brain's neural network in order to achieve artificial intelligence similar to humans. It is the basis of deep learning. Currently, general neural networks use floating-point calculations, which require a large amount of storage space and computational power, seriously hindering their applications on edge devices (such as cameras) and end-side devices (such as mobile phones). Binary neural networks have become a popular research direction in deep learning in recent years due to their potential advantages of high model compression rate and fast calculation speed.

[0003] A binary neural network (BNN) is based on a neural network, and each weight in the weight matrix of the neural network is binarized to 1 or -1. Through the binarization operation, the parameters of the model occupy less storage space (each original weight requires 32-bit floating-point storage, and now only one bit can store it, and the memory consumption is theoretically reduced to 1 / 32 of the original). The essence of BNN is to binarize the weight matrix of the original neural network (that is, each weight takes a value of +1 or -1), without changing the network structure, and mainly makes some optimization processes in gradient descent, weight update, etc.

[0004] Currently, the binarization method of a neural network binarizes the weight matrix of a single layer. That is to say, this binarization method only quantifies the weight matrices of each layer of the neural network separately, resulting in a large quantization error. Summary of the Invention

[0005] Embodiments of this application provide an optimization method for a neural network and related devices, which are used to adjust the values of each weight in the weight matrix of each layer of the neural network to +1 or -1. The value of the adjusted weight matrix of each layer (such as the weight matrix of the mth layer) is related to the values of the weight matrices of the previous layers (such as the 1st layer to the m - 1st layer) before adjustment. This optimization method makes the values of each weight in the weight matrices of each layer not only related to itself but also related to the weight matrices of other layers, reducing the quantization error and making the training and use of the neural network more efficient.

[0006] Based on this, embodiments of this application provide the following technical solutions:

[0007] In a first aspect, embodiments of this application first provide an optimization method for a neural network, which can be used in the field of artificial intelligence. The neural network includes a first neural network module, and the first neural network module includes n convolutional layers. Specifically, the method includes:

[0008] First, the training device obtains a first quantization model, which is used to obtain the second weight matrix of the m-th layer of the first neural network module according to the m first weight matrices of the first to m-th layers of the first neural network module in the neural network. Here, the first weight matrix of each layer of the first neural network module refers to the initial weight matrix represented by 32-bit floating-point numbers, and the second weight matrix of each layer of the first neural network module refers to the weight matrix assigned as +1 or -1. After the training device obtains the first quantization model, it can, according to this first quantization model, perform a binarization operation on each first weight matrix corresponding to each layer of the first neural network module to obtain each second weight matrix corresponding to each layer of the first neural network module. After the training device binarizes the first weight matrix of each layer of the first neural network module into a second weight matrix according to the first quantization model through the above steps, it can further train the neural network with the training data in the training set to obtain the trained neural network. Finally, the trained neural network can be deployed on the target device for use. It should be noted that in the embodiments of the present application, the target device may specifically be a mobile device, such as edge devices like cameras and smart homes, or end-side devices such as mobile phones, personal computers, computer workstations, tablets, smart wearable devices (such as smart watches, smart bracelets, smart earphones, etc.), game consoles, set-top boxes, and media consumption devices. Specifically, the type of the target device is not limited here.

[0009] In the above implementation manner of the present application, a new quantization model (i.e., the first quantization model) is used to binarize the weight matrix of the neural network. This first quantization model is used to obtain the second weight matrix of the m-th layer of the neural network according to the m first weight matrices of the first to m-th layers of the neural network. Here, the first weight matrix of each layer is a weight matrix represented by 32-bit floating-point numbers, and the second weight matrix of each layer is a weight matrix assigned as +1 or -1. In this way, the value of the adjusted weight matrix of each layer (such as the weight matrix of the m-th layer) is related to the values of the weight matrices of the previous layers (such as the first to m-1-th layers) before adjustment. This optimization method makes the value of each weight in the weight matrices of each layer not only related to itself but also related to the weight matrices of other layers, reducing the quantization error and making the training and use of the neural network more efficient.

[0010] In a possible implementation manner of the first aspect, when the training device, according to this first quantization model, performs a binarization operation on each first weight matrix corresponding to each layer of the first neural network module to obtain each second weight matrix corresponding to each layer of the first neural network module, specifically, it can be through to obtain the second weight matrix of the m-th layer, where W 1 ,W 2 ,…,W mare the first weight matrices from the first layer to the m-th layer, α 1 , α 2 , …, α m are the first linear combination parameters corresponding to W 1 , W 2 , …, W m . WN(·) is the normalization operation performed on α 1 W 1 + α 2 W 2 + … + α m W m . Sign(·) is the sign function, is the second weight matrix of the m-th layer. It can also be the second weight matrix of the m-th layer obtained through , where W 1 , W 2 , …, W m are the first weight matrices from the first layer to the m-th layer, α 1 , α 2 , …, α m are the first linear combination parameters corresponding to W 1 , W 2 , …, W m . k is a non - negative trainable parameter. WN(·) is the normalization operation performed on α 1 W 1 + α 2 W 2 + … + α m W m . Sign(·) is the sign function, is the second weight matrix of the m-th layer. Specifically, the specific form of the first quantization model is not limited here. As long as the first quantization model makes the second weight matrix of the current layer related to the first weight matrices of at least the previous two layers, it belongs to the first quantization model described in this application.

[0011] In the above - mentioned implementation manner of this application, several different quantization forms of the first weight matrix are provided, that is, several specific expression forms of the first quantization model are given, which have selectivity and flexibility.

[0012] In a possible implementation manner of the first aspect, the weight gain of the second weight matrix of the m-th layer can be further determined, and the second weight matrix of the m-th layer is adjusted according to the weight gain of the second weight matrix of the m-th layer, so that the difference between the adjusted second weight matrix of the m-th layer and the first weight matrix of the m-th layer is less than the difference between the second weight matrix of the m-th layer and the first weight matrix of the m-th layer.

[0013] In the above embodiments of the present application, the advantage of adjusting the second weight matrix using the weight gain is that the adjusted second weight matrix is closer to the initial first weight matrix of 32-bit floating-point numbers. In practical applications, this can better retain the accuracy of image information.

[0014] In a possible implementation manner of the first aspect, since the first linear combination parameters are a set of non-negative parameters and their values are not determined finally in the initialization state, the first linear combination parameters can be set as the network parameters of the neural network, so that during the training of the neural network according to the training data in the training set, the first linear combination parameters are trained.

[0015] In the above embodiments of the present application, a specific implementation manner for optimizing the first linear combination parameters is provided. The advantage of this optimization process is that during the training of the neural network, the optimization of the first linear combination parameters is completed simultaneously, which is simple and convenient.

[0016] In a possible implementation manner of the first aspect, the optimization process of the first linear combination parameters can also be: determining the modulus value of the first weight matrix of the m-th layer and the second weight matrix of the m-th layer as α in the first linear combination parameters m , and performing linear regression on this modulus value to obtain the final value of α m .

[0017] In the above embodiments of the present application, another specific implementation manner for optimizing the first linear combination parameters is provided. By using the method of linear regression to obtain the values of each parameter in the first linear combination parameters, the optimization method of the first linear combination parameters has selectivity.

[0018] In a possible implementation of the first aspect, the training device may also calculate the first feature representation of each layer of the first neural network module in the order of connection of the n convolutional layers of the first neural network module, and obtain a second quantization model, which is used to obtain the second feature representation of the m-th layer of the first neural network module according to the m first feature representations of the first neural network module from the 1st layer to the m-th layer. Among them, the first feature representation of each layer is a feature representation represented by 32-bit floating-point numbers, and the second feature representation of each layer is a feature representation assigned +1 or -1. 1 ≤ m ≤ n. Finally, according to the second quantization model, each first feature representation corresponding to each layer of the first neural network module is binarized to obtain each second feature representation corresponding to each layer of the first neural network module. It should be noted that there is no sequence requirement between the training device calculating the first feature representation of each layer of the first neural network module and obtaining the second quantization model. The training device may first calculate the first feature representation of each layer of the first neural network module and then obtain the second quantization model; the training device may also first obtain the second quantization model and then calculate the first feature representation of each layer of the first neural network module. Specifically, it is not limited here.

[0019] Since only the weight matrix is binarized, the feature representations (which can also be called feature maps, activation values, etc.) of each layer are still represented by 32-bit floating-point numbers. When the weight matrix and the feature representation are operated, it still needs to be carried out through 32-bit floating-point numbers, and the computing overhead cannot be saved. It only partially reduces the space occupied by the storage of the neural network model. Therefore, in the above implementation manner of the present application, the feature representations output by each layer of the first neural network module are further binarized, so that the binarized weight matrix and the binarized feature representation can directly perform bit operations, reducing the computing overhead.

[0020] In a possible implementation of the first aspect, the training device binarizes each first feature representation corresponding to each layer of the first neural network module according to the second quantization model to obtain each second feature representation corresponding to each layer of the first neural network module. Specifically, it can be through to obtain the second feature representation of the m-th layer, where A 1 , A 2 , …, A m are the first feature representations from the 1st layer to the m-th layer of the first neural network module, and β 1 , β 2 , …, β m are the second linear combination parameters corresponding to A 1 , A 2 , …, A m , and BN(·) is to β 1 A 1 + β 2 A 2+…+β m A m The normalization operation performed, Sign(·) is the sign function, is the second feature representation of the m-th layer. It can also be obtained by to obtain the second feature representation of the m-th layer, where A 1 , A 2 , …, A m are the first feature representations from the first layer to the m-th layer, β 1 , β 2 , …, β m are the second linear combination parameters corresponding to A 1 , A 2 , …, A m , l is a non-negative parameter that can be trained, and BN(·) is the normalization operation performed on β 1 A 1 +β 2 A 2 +…+β m A m The normalization operation performed, Sign(·) is the sign function, is the second feature representation of the m-th layer. Specifically, the specific form of the second quantization model is not limited here, as long as the second quantization model makes the second feature representation of the current layer related to the first feature representations of at least the previous two layers, then it belongs to the second quantization model described in this application.

[0021] In the above embodiments of this application, several different quantization forms of the first feature representation are provided, that is, several specific expression forms of the second quantization model are given, which have selectivity and flexibility.

[0022] In a possible implementation manner of the first aspect, the activation gain of the second feature representation of the m-th layer can be further determined, and the second feature representation of the m-th layer can be adjusted according to the activation gain of the second feature representation of the m-th layer, so that the difference between the adjusted second feature representation of the m-th layer and the first feature representation of the m-th layer is less than the difference between the second feature representation of the m-th layer and the first feature representation of the m-th layer.

[0023] In the above embodiments of this application, the advantage of adjusting the second feature representation using the activation gain is that: the adjusted second feature representation is closer to the first feature representation of the initial 32-bit floating-point number. Since the feature representation has a greater impact on the accuracy of the image information, in practical applications, the accuracy of the retained image information is further improved.

[0024] In a possible implementation of the first aspect, the training device calculates the first feature representation of each layer of the first neural network module in the order of connection of the n convolutional layers of the first neural network module. Specifically, it can be: calculating the first feature representation of the m-th layer according to the second weight matrix of the m-th layer, the second feature representation of the (m - 1)-th layer, the weight gain of the second weight matrix of the m-th layer, and the activation gain of the second feature representation of the (m - 1)-th layer. Here, it should be noted that since the feature representations of the neural network are calculated layer by layer, and the normal convolution operation is also calculated layer by layer backward, when the training device calculates the feature representation of the second layer of the first neural network module, the feature representation of the first layer has already been calculated. Therefore, in some embodiments of the present application, the second feature representation of the first layer of the first neural network module is directly obtained by the Sign function on the first feature representation of the first layer. When the training device calculates the second feature representations of the second layer and subsequent layers, it can calculate the first feature representation of the m-th layer according to the second weight matrix of the m-th layer, the second feature representation of the (m - 1)-th layer, the weight gain of the second weight matrix of the m-th layer, and the activation gain of the second feature representation of the (m - 1)-th layer.

[0025] In the above embodiments of the present application, an implementation manner for calculating the first feature representation is provided, which has feasibility.

[0026] In a possible implementation of the first aspect, the implementation manner for calculating the first feature representation can specifically be: performing a convolution operation on the second weight matrix of the m-th layer and the second feature representation of the (m - 1)-th layer to obtain a convolution result. Then, performing a dot product operation on the convolution result and the weight gain of the second weight matrix of the m-th layer to obtain a dot product result. Finally, performing a dot product operation on the dot product result and the activation gain of the second feature representation of the (m - 1)-th layer to obtain the first feature representation of the m-th layer.

[0027] In the above embodiments of the present application, a method for obtaining the first feature representation is specifically described, which has feasibility.

[0028] In a possible implementation of the first aspect, since the second linear combination parameters are a set of non-negative parameters and the final values of these parameters are not determined in the initialization state, the second linear combination parameters can be set as the network parameters of the neural network. In this way, during the process of training the neural network according to the training data in the training set, the second linear combination parameters can be trained simultaneously.

[0029] In the above embodiments of the present application, a specific implementation manner for optimizing the second linear combination parameters is provided. The advantage of this optimization process is that during the process of training the neural network, the optimization of the second linear combination parameters is completed simultaneously, which is simple and convenient.

[0030] In a possible implementation of the first aspect, the optimization process of the second linear combination parameter may also be: determining that the modulus value of the first feature representation and the second feature representation of the m-th layer is β in the second linear combination parameter m , and performing linear regression on the modulus value to obtain the final value of β m .

[0031] In the above embodiments of the present application, another specific implementation of optimizing the second linear combination parameter is provided. By means of linear regression, the values of each parameter in the second linear combination parameter are obtained, making the optimization method of the second linear combination parameter selective.

[0032] In a possible implementation of the first aspect, the neural network further includes a second neural network module and a third neural network module. The second neural network module is used to perform full-precision feature extraction on the input image, and the third neural network module is used to perform image reconstruction on the output of the first neural network module to obtain the output image.

[0033] In the above embodiments of the present application, it is described that in addition to including the first neural network module, the neural network may further include a second neural network module and a third neural network module. Among them, the second neural network module is used to perform full-precision feature extraction on the input image, and the third neural network module is used to perform image reconstruction on the output of the first neural network module to obtain the output image. The purpose of the second neural network module and the third neural network module is to adopt a full-precision convolution process in the feature extraction stage and the image reconstruction stage, so as to ensure the performance of the model and make the accuracy of the final output image higher.

[0034] In a possible implementation of the first aspect, the input image includes one or more low-resolution images, and the output image includes one high-resolution image.

[0035] In the above embodiments of the present application, when the neural network is applied to the scenario of image super-resolution reconstruction, then the input image may be one or more low-resolution images, and the output image will be one high-resolution image.

[0036] In the second aspect of the embodiments of the present application, an image processing method is further provided. The method may specifically include: obtaining an input image, and processing the input image through a trained neural network to obtain an output image. The trained neural network is a neural network optimized by the method of the above first aspect or any possible implementation of the first aspect.

[0037] The third aspect of the embodiments of the present application provides a network structure of a neural network, which may specifically include: a first neural network module, a second neural network module, and a third neural network module. Among them, the first neural network module includes n convolutional layers. The second neural network module is used to perform full-precision feature extraction on the input image to obtain a first target feature representation. The first neural network module is used to perform a non-linear mapping on the first target feature representation to obtain a second target feature representation. Among them, the weight matrix of each layer of the first neural network module is a second weight matrix processed by a first quantization model. The first quantization model is used to obtain the second weight matrix of the m-th layer of the first neural network module according to the m first weight matrices of the first layer to the m-th layer of the first neural network module. The first weight matrix of each layer is a weight matrix represented by 32-bit floating-point numbers, and the second weight matrix of each layer is a weight matrix assigned +1 or -1, where 1 ≤ m ≤ n. The third neural network module is used to perform image reconstruction on the second target feature representation to obtain an output image.

[0038] In the above implementation manner of the present application, a network structure of a neural network is introduced. The difference between this neural network and other neural networks is that the weight matrix of each layer of its first neural network module is binarized by a first quantization model, so that the value of the binarized weight matrix (i.e., the second weight matrix) of each layer of the first neural network module is not only related to itself, but also related to the values of all un-binarized weight matrices (i.e., the first weight matrices) of the previous layers, reducing the quantization error and making the training and use of the neural network more efficient.

[0039] In a possible implementation manner of the third aspect, the first quantization model may be: Among them, W 1 , W 2 , …, W m are the first weight matrices of the first layer to the m-th layer of the first neural network module 701, α 1 , α 2 , …, α m are the first linear combination parameters corresponding to W 1 , W 2 , …, W m . WN(·) is the normalization operation performed on α 1 W 1 + α 2 W 2 + … + α m W m . Sign(·) is the sign function, is the second weight matrix of the m-th layer. The first quantization model may also be: Among them, W1 ,W 2 ,…,W m are the first weight matrices from the first layer to the m-th layer, α 1 ,α 2 ,…,α m are the first linear combination parameters corresponding to W 1 ,W 2 ,…,W m , k is a non - negative trainable parameter, WN(·) is the normalization operation performed on α 1 W 1 +α 2 W 2 +…+α m W m , and Sign(·) is the sign function. is the second weight matrix of the m-th layer. Specifically, the specific form of the first quantization model is not limited here. As long as the first quantization model makes the second weight matrix of the current layer related to the first weight matrices of at least the previous two layers, it belongs to the first quantization model described in this application.

[0040] In the above - mentioned embodiments of this application, several specific forms of the first quantization model are given, which has flexibility.

[0041] In a possible implementation manner of the third aspect, the feature representations of each layer of the first neural network module are the second feature representations processed by the second quantization model. The second quantization model is used to obtain the second feature representation of the m-th layer of the first neural network module according to the m first feature representations of the first neural network module from the first layer to the m-th layer. Among them, the first feature representation of each layer is a feature representation represented by 32 - bit floating - point numbers, and the second feature representation of each layer is a feature representation assigned +1 or -1.

[0042] In the above - mentioned embodiments of this application, not only the weight matrix of the first neural network module is binarized, but also the feature representation of the first neural network module is further binarized through the second quantization model. In this way, the binarized weight matrix and the binarized feature representation can directly perform bit operations, reducing the computational overhead.

[0043] In a possible implementation manner of the third aspect, the second quantization model can be: Among them, A 1 ,A 2 ,…,A m are the first feature representations of the first neural network module 701 from the first layer to the m-th layer, β 1 ,β 2 ,…,β m are the ones corresponding to A1 , A 2 , …, A m The corresponding second linear combination parameter, and BN(·) is the normalization operation performed on β 1 A 1 + β 2 A 2 + … + β m A m The Sign(·) is the sign function is the second feature representation of the m-th layer. The second quantization model can also be: Wherein, A 1 , A 2 , …, A m are the first feature representations from the first layer to the m-th layer, and β 1 , β 2 , …, β m are the second linear combination parameters corresponding to A 1 , A 2 , …, A m l is a non-negative parameter that can be trained, and BN(·) is the normalization operation performed on β 1 A 1 + β 2 A 2 + … + β m A m The Sign(·) is the sign function is the second feature representation of the m-th layer. Specifically, the specific form of the second quantization model is not limited here, as long as the second quantization model makes the second feature representation of the current layer related to the first feature representations of at least two previous layers, then it belongs to the second quantization model described in this application

[0044] In the above embodiments of this application, several specific forms of the second quantization model are given, which have flexibility

[0045] The fourth aspect of the embodiments of this application provides a training device, and this training device has the function of implementing the method in the above first aspect or any one of the possible implementation manners of the first aspect. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions

[0046] The fifth aspect of the embodiments of this application provides an execution device, and this execution device has the function of implementing the method in the above second aspect. This function can be implemented by hardware or by hardware executing corresponding software. The hardware or software includes one or more modules corresponding to the above functions

[0047] A sixth aspect of the embodiments of the present application provides a training device, which may include a memory, a processor, and a bus system. Among them, the memory is used to store a program, and the processor is used to call the program stored in the memory to execute the method according to the first aspect or any possible implementation manner of the first aspect of the embodiments of the present application.

[0048] A seventh aspect of the embodiments of the present application provides an execution device, which may include a memory, a processor, and a bus system. Among them, the memory is used to store a program, and the processor is used to call the program stored in the memory to execute the method according to the second aspect of the present application above.

[0049] An eighth aspect of the present application provides a computer-readable storage medium, in which instructions are stored. When it runs on a computer, it enables the computer to execute the method according to the first aspect or any possible implementation manner of the first aspect, or enables the computer to execute the method according to the second aspect.

[0050] A ninth aspect of the embodiments of the present application provides a computer program. When it runs on a computer, it enables the computer to execute the method according to the first aspect or any possible implementation manner of the first aspect, or enables the computer to execute the method according to the second aspect. Description of the Drawings

[0051] Figure 1 It is a schematic diagram of mislabeling during a binarization operation process;

[0052] Figure 2 It is a schematic structural diagram of an artificial intelligence main framework provided by the embodiments of the present application;

[0053] Figure 3 It is a system architecture diagram of a task processing system provided by the embodiments of the present application;

[0054] Figure 4 It is a schematic flowchart of an optimization method for a neural network provided by the embodiments of the present application;

[0055] Figure 5 It is a schematic diagram of adjusting a second weight matrix through weight gain provided by the embodiments of the present application;

[0056] Figure 6 It is a schematic diagram of the overall process of an optimization method for a neural network provided by the embodiments of the present application;

[0057] Figure 7 It is a schematic diagram of the network structure of a neural network provided by the embodiments of the present application;

[0058] Figure 8A schematic diagram of an application scenario of the neural network trained in the embodiment of the present application in image super-resolution reconstruction;

[0059] Figure 9 A schematic diagram of an application scenario of the neural network trained in the embodiment of the present application for object detection on a terminal mobile phone;

[0060] Figure 10 A schematic diagram of an application scenario of the neural network trained in the present application for autonomous driving scene segmentation on a wheeled mobile device;

[0061] Figure 11 A schematic diagram of an application scenario of the neural network trained in the present application in face recognition applications;

[0062] Figure 12 A schematic diagram of an application scenario of the neural network trained in the present application in speech recognition applications;

[0063] Figure 13 A comparison chart of the visual evaluation of the solution provided in the embodiment of the present application and other existing solutions based on the VDSR model;

[0064] Figure 14 A comparison chart of the visual evaluation of the solution provided in the embodiment of the present application and other existing solutions based on the SRRestNet model;

[0065] Figure 15 A schematic diagram of a training device provided in the embodiment of the present application;

[0066] Figure 16 A schematic diagram of an execution device provided in the embodiment of the present application;

[0067] Figure 17 Another schematic diagram of a training device provided in the embodiment of the present application;

[0068] Figure 18 Another schematic diagram of an execution device provided in the embodiment of the present application;

[0069] Figure 19 A schematic diagram of a structure of a chip provided in the embodiment of the present application. Detailed implementation manners

[0070] Embodiments of the present application provide an optimization of a neural network and related devices, which are used to adjust the values of each weight in the weight matrix of each layer of the neural network to +1 or -1. The value of the adjusted weight matrix of each layer (for example, the weight matrix of the m-th layer) is related to the values of the weight matrices of the previous layers (for example, the weight matrices of the 1st layer to the m-1th layer) before adjustment. This optimization method makes the values of each weight in the weight matrices of each layer not only related to itself, but also related to the weight matrices of other layers, reducing the quantization error and making the training and use of the neural network more efficient.

[0071] Terms such as "first" and "second" in the specification, claims and above-mentioned drawings of the present application are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing embodiments of the present application. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product or device comprising a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products or devices.

[0072] Before introducing the embodiments of the present application, the current technology of neural network binarization (i.e., BNN) and related backgrounds are briefly introduced to facilitate the subsequent understanding of the embodiments of the present application. First, the related background of the proposal of BNN is introduced. In the field of deep learning, neural networks are widely used. The central processing unit (CPU) has gradually been unable to meet the requirements of high concurrency and high computational complexity of various deep neural networks (such as convolutional neural networks (CNNs)). Although the graphics processing unit (GPU) can partially solve the problems of high concurrency and high computational complexity, its large power consumption and high price also limit its application in mobile devices (including edge devices and end devices). Generally, only enterprises or research institutions can purchase high-end GPUs for neural network training, testing and application. At present, some mobile phone chips have integrated neural network processors (NPUs), such as the Kirin 970 chip of Huawei. However, how to achieve a balance between power consumption and performance remains an urgent problem to be solved.

[0073] The two main technical problems restricting the application of deep neural networks on mobile devices are: 1) excessive computational complexity; 2) excessive number of parameters in the neural network. Taking CNN as an example, the computational complexity of the convolution operation is huge. For a convolution kernel with hundreds of thousands of parameters, the number of floating point operations (FLOPs) of the convolution operation can reach tens of millions. The total computational complexity of an ordinary existing CNN with n layers can be as high as several billion FLOPs. A CNN that can be operated in real time on a GPU becomes very slow on a mobile device. When the computing resources on a mobile device are difficult to meet the real-time operation requirements of the existing CNN, it is necessary to consider how to reduce the convolution computational complexity. In addition, in the commonly used CNNs currently, the number of parameters in each convolution layer can often reach tens of thousands, hundreds of thousands or even more. The parameters of the entire n-layer network add up to tens of millions, and each parameter is represented by a 32-bit floating point number. This requires hundreds of megabytes of memory or cache to store these parameters. In mobile devices, the memory and cache resources are very limited. How to reduce the number of parameters in the convolution layer so that the CNN can adapt to the relevant devices of mobile devices is also an urgent problem to be solved. Against this background, BNN came into being.

[0074] Currently, the commonly used BNN is based on the existing neural network, and performs binary quantization on the weights, that is, assigns the values of each weight in the weight matrix of each layer of the original neural network to +1 or -1. BNN does not change the network structure of the original neural network. It mainly makes some optimization treatments on gradient descent, weight update, and convolution operations. There are currently two main methods for binary quantization of the weight matrix of a floating-point neural network. The first method is a deterministic method based on the sign function (also known as the Sign function), and the formula (1) is as follows:

[0075]

[0076] where w is the value of each weight in the weight matrix of each layer of the original neural network, and W and W b represent the weight matrix before quantization and the weight matrix after quantization respectively.

[0077] The second method is a random binary quantization method (which can be called a statistical method), and the formula (2) is as follows:

[0078]

[0079] where that is, each weight in the weight matrix W is randomly binary quantized to +1 or -1 with a certain probability σ(W).

[0080] Theoretically speaking, the second method is more reasonable. However, it is difficult to generate random numbers using hardware in actual operation. Therefore, in actual applications, the second method has not been applied yet, and the first method is adopted, that is, binarization is performed through the Sign function.

[0081] However, this binarization method only binarizes the weight matrices of each layer of the neural network separately, without considering the correlation between the weight matrices of each layer, which will cause two problems:

[0082] (1) Large quantization error

[0083] Because this binarization method only binarizes the current weight matrix separately and cannot effectively retain pixel detail information. In a certain layer of the neural network (for example, the m-th layer), some weights that should be binarized to +1 may be binarized to -1, such as Figure 1 the weights with a dark background in Figure 1 are mislabeled as -1; while some weights that should be binarized to -1 may be binarized to +1, such as

[0084] the weights with a light background in

[0085] are mislabeled as +1.

[0086]

[0087] Therefore, when training the BNN, the above Sign function is not differentiable. In this case, generally, the derivative of the binarized weight matrix is directly used to update the floating-point weight matrix, and a clipping operation is adopted to enhance the stability of training, as shown in formula (4):

[0088]

[0089] Among them, Clip represents the clipping operation, C represents the loss function, α represents the trainable scale factor (scale, a non-negative coefficient), g w represents the weight gradient of each layer of the weight matrix W after clipping, η represents the learning rate of the neural network, and UpdataBinaryParameter represents the iterative process of the weight matrix W.

[0090] However, the method of directly updating the floating-point weight matrix using the derivative of the binary weight matrix reduces the accuracy of gradient conduction and is not conducive to the training of neural networks.

[0091] Based on this, to solve the above problems, an embodiment of the present application provides an optimization method for neural networks, which is used to adjust the values of each weight in the weight matrix of each layer of the neural network to +1 or -1. The value of the adjusted weight matrix of each layer (for example, the weight matrix of the m-th layer) is related to the values of the weight matrices of the previous layers (for example, the 1st layer to the m-1th layer) before adjustment. This optimization method makes the values of each weight in the weight matrix of each layer not only related to itself but also related to the weight matrices of other layers, reducing the quantization error and making the training and use of neural networks more efficient.

[0092] Next, the embodiments of the present application will be described in conjunction with the accompanying drawings. As is known to those of ordinary skill in the art, with the development of technology and the emergence of new scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0093] First, the overall working process of the artificial intelligence system will be described. Please refer to Figure 2 , Figure 2 FIG. shows a schematic structural diagram of an artificial intelligence main framework. The above artificial intelligence main framework will be described from two dimensions: "intelligent information chain" (horizontal axis) and "IT value chain" (vertical axis). Among them, the "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it can be a general process of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution and output. In this process, data undergoes a refining process of "data - information - knowledge - wisdom". The "IT value chain" reflects the value brought by artificial intelligence to the information technology industry from the underlying infrastructure of artificial intelligence, information (provision and processing technology implementation) to the industrial ecosystem of the system.

[0094] (1) Infrastructure

[0095] The infrastructure provides computing power support for the artificial intelligence system, realizes communication with the external world, and is supported through the basic platform. Communicate with the outside through sensors; the computing power is provided by intelligent chips (such as hardware acceleration chips like CPU, NPU, GPU, ASIC, FPGA, etc.); the basic platform includes relevant platform guarantees and supports such as distributed computing frameworks and networks, and can include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the outside to obtain data, and these data are provided to the intelligent chips in the distributed computing system provided by the basic platform for computing.

[0096] (2) Data

[0097] The data at the upper layer of the infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voices, texts, and also involves the Internet of Things data of traditional devices, including the business data of existing systems and the sensed data such as force, displacement, liquid level, temperature, humidity, etc.

[0098] (3) Data processing

[0099] Data processing usually includes data training, machine learning, deep learning, search, inference, decision-making, etc.

[0100] Among them, machine learning and deep learning can perform symbolic and formal intelligent information modeling, extraction, preprocessing, training, etc. on data.

[0101] Inference refers to the process of simulating the intelligent reasoning method of humans in a computer or intelligent system, and using formal information to perform machine thinking and solve problems according to the inference control strategy. The typical function is search and matching.

[0102] Decision-making refers to the process of making decisions after intelligent information is inferred, and usually provides functions such as classification, sorting, prediction, etc.

[0103] (4) General capabilities

[0104] After the data is processed through the above-mentioned data processing, some general capabilities can be further formed based on the results of the data processing. For example, it can be an algorithm or a general system. For example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.

[0105] (5) Intelligent products and industry applications

[0106] Intelligent products and industry applications refer to the products and applications of artificial intelligence systems in various fields, which is the encapsulation of the overall artificial intelligence solution, productizing intelligent information decision-making and realizing landing applications. Its application fields mainly include: intelligent terminals, intelligent manufacturing, intelligent transportation, smart homes, intelligent healthcare, intelligent security, autonomous driving, safe cities, etc.

[0107] The embodiments of the present application can be applied to the optimal design of the network structure of neural networks, and the neural network with the weight matrix optimized by the present application can be specifically applied to various sub-fields in the field of artificial intelligence, such as the field of image processing, the field of computer vision, the field of semantic analysis, etc. Specifically, combined with Figure 2In general, the data in the dataset obtained by the infrastructure in the embodiments of the present application can be multiple different types of data obtained through sensors such as cameras and radars (which can also be referred to as training data, and multiple training data constitute a training set), or multiple image data or multiple video data, as long as the training set meets the function of being used for iterative training of the neural network and can be used to optimize the weight matrix of the neural network of the present application. Specifically, the data type in the training set is not limited here.

[0108] Next, the architecture of the task processing system will be introduced. Please refer to Figure 3 , Figure 3 which is a system architecture diagram of the task processing system provided by the embodiments of the present application. In Figure 3 , the task processing system 200 includes an execution device 210, a training device 220, a database 230, a client device 240, a data storage system 250, and a data acquisition device 260. The execution device 210 includes a computing module 211. Among them, the data acquisition device 260 is used to obtain an open-source large-scale dataset (i.e., a training set) required by the user and store the training set in the database 230. The training device 220 trains the neural network 201 of the present application based on the training set maintained in the database 230, and the trained neural network 201 is then applied on the execution device 210. The execution device 210 can call the data, code, etc. in the data storage system 250, or store data, instructions, etc. in the data storage system 250. The data storage system 250 can be placed in the execution device 210, or the data storage system 250 is an external memory relative to the execution device 210.

[0109] The trained neural network 201 obtained by training through the training device 220 can be applied to different systems or devices (i.e., the execution device 210), specifically, it can be an edge device or an end-side device. For example, mobile phones, tablets, laptop computers, monitoring systems (such as cameras), security systems, and so on. In Figure 3In [the figure], the execution device 210 is configured with an I / O interface 212 to interact with external devices, and a "user" can input data to the I / O interface 212 through the client device 240. For example, the client device 240 can be a camera device of a monitoring system, and the image captured by the camera device is input as input data to the computing module 211 of the execution device 210. After the computing module 211 detects the input image, a detection result is obtained, and then the detection result is output to the camera device or directly displayed on the display interface (if any) of the execution device 210. In addition, in some embodiments of the present application, the client device 240 can also be integrated in the execution device 210. For example, when the execution device 210 is a mobile phone, the target task can be directly obtained through the mobile phone (for example, an image can be captured through the camera of the mobile phone, or the target voice recorded by the recording module of the mobile phone, etc., and the target task is not limited here), or the target task sent by other devices (such as another mobile phone) is received, and then the computing module 211 in the mobile phone detects the target task to obtain a detection result, and directly presents the detection result on the display interface of the mobile phone. The product forms of the execution device 210 and the client device 240 are not limited here.

[0110] It should be noted that Figure 3 is only a schematic diagram of a system architecture provided by an embodiment of the present application. The positional relationships between the devices, components, modules, etc. shown in the figure do not constitute any limitation. For example, in Figure 3 the data storage system 250 is an external memory relative to the execution device 210. In other cases, the data storage system 250 can also be placed in the execution device 210; in Figure 3 the client device 240 is an external device relative to the execution device 210. In other cases, the client device 240 can also be integrated in the execution device 210.

[0111] It should also be noted that the training of the neural network 201 described in the embodiments of the present application can be implemented on the cloud side. For example, the training device 220 on the cloud side (the training device 220 can be set on one or more servers or virtual machines) can obtain a training set, and train the neural network according to multiple groups of training data in the training set to obtain the trained neural network 201. Then, the trained neural network 201 is sent to the execution device 210 for application. For example, it is sent to the execution device 210 for image super-resolution reconstruction. Exemplarily, Figure 3As described in the corresponding system architecture, the neural network is trained by the training device 220, and the trained neural network 201 is then sent to the execution device 210 for use. The training of the neural network 201 described in the above embodiments can also be implemented on the terminal side, that is, the training device 220 can be located on the terminal side. For example, the training set can be obtained by a terminal device (such as a mobile phone, a smart watch, etc.), a wheeled mobile device (such as an autonomous driving vehicle, an assisted driving vehicle, etc.), and the neural network is trained according to multiple groups of training data in the training set to obtain the trained neural network 201. The trained neural network 201 can be directly used on the terminal device or sent by the terminal device to other devices for use. Specifically, the embodiments of the present application do not limit on which device (cloud side or terminal side) the neural network 201 is trained or applied.

[0112] Next, an optimization method for the neural network provided by the embodiments of the present application will be introduced. Please refer to Figure 4 , Figure 4 which is a schematic flowchart of an optimization method for the neural network provided by the embodiments of the present application, and specifically may include:

[0113] 401. Obtain a first quantization model, where the first quantization model is used to obtain a second weight matrix of the m-th layer of the first neural network module according to m first weight matrices of the first to m-th layers of the first neural network module.

[0114] First, the training device obtains a first quantization model, and the first quantization model is used to obtain a second weight matrix of the m-th layer of the first neural network module according to m first weight matrices of the first to m-th layers of the neural network. Among them, the first weight matrix of each layer of the first neural network module refers to the initial weight matrix represented by 32-bit floating-point numbers, and the second weight matrix of each layer of the first neural network module refers to the weight matrix assigned to +1 or -1.

[0115] It should be noted that in the embodiments of the present application, since the existing process of neural network-based image processing (such as image super-resolution reconstruction) generally includes three stages: feature extraction, non-linear mapping, and image reconstruction. Assuming that x is the input low-resolution (LR) image and y is the finally reconstructed high-resolution (HR) image, the general neural network model for performing this image processing can be simplified as formula (5):

[0116]

[0117] where ε corresponds to the feature extraction stage, corresponds to the non-linear mapping stage, corresponds to the image reconstruction stage. Generally speaking, ε and In this stage, only one convolutional layer is used to implement the transformation from an image to depth features and its inverse transformation, and the computational complexity of the neural network model almost entirely depends on the design of the stage module, The stage module can include n convolutional layers, which is specifically determined by the design requirements and complexity.

[0118] The above is only an example using the application scenario of image super-resolution reconstruction in image processing. In all image processing processes, generally, there are the above three stages. It's just that in each stage, the number of layers of the neural network of the model will be a little different, but the essence is the same, so it won't be elaborated here.

[0119] Therefore, in some embodiments of the present application, the above first neural network module may include three stages: feature extraction, non-linear mapping, and image reconstruction. In this case, the first quantization model can be applied to all layers of the neural network, that is, the first weight matrix of all layers of the neural network can be binarized to +1 or -1. The advantage of this method is that: the model parameters of the neural network can occupy the smallest storage space (originally each weight needed to be stored as a 32-bit floating-point number, now only one bit is needed to store it, and the memory consumption is theoretically reduced to 1 / 32 of the original), but the accuracy of image processing may be correspondingly reduced. The above first neural network module may also only include the non-linear mapping stage. In this case, the first quantization model is only applied to the stage of the neural network, that is, the first weight matrix of each layer in the stage can be binarized to +1 or -1. The advantage of this method is that: a full-precision convolutional process is adopted in the ε and stages, which can ensure the performance of the model. Only the weight matrix of each layer in the stage is binarized. On the premise of partially reducing the model size, the accuracy of image processing is also ensured. The above first neural network module may also include two stages: feature extraction and non-linear mapping, or may include two stages: non-linear mapping and image reconstruction. Specifically, it is not limited here which stages the first neural network module includes. The difference is only that the stages included in the first neural network module are different, and the positions and numbers of the neural network layers that can be binarized are different. For ease of understanding, in the following embodiments, it is described by taking the first neural network module including the non-linear mapping stage as an example.

[0120] 402. According to the first quantization model, perform a binarization operation on each first weight matrix corresponding to each layer of the first neural network module to obtain each second weight matrix corresponding to each layer of the first neural network module.

[0121] After the training device obtains the first quantization model, it can perform binarization operations on each first weight matrix corresponding to each layer of the first neural network module according to the first quantization model, and obtain each second weight matrix corresponding to each layer of the first neural network module. Specifically, it includes but is not limited to the following methods: 1) By obtain the second weight matrix of the m-th layer, where W 1 , W 2 , …, W m are the first weight matrices from the first layer to the m-th layer, α 1 , α 2 , …, α m are the first linear combination parameters corresponding to W 1 , W 2 , …, W m , WN(·) is the normalization operation performed on α 1 W 1 + α 2 W 2 + … + α m W m , Sign(·) is the sign function, is the second weight matrix of the m-th layer. 2) By obtain the second weight matrix of the m-th layer, where W 1 , W 2 , …, W m are the first weight matrices from the first layer to the m-th layer, α 1 , α 2 , …, α m are the first linear combination parameters corresponding to W 1 , W 2 , …, W m , k is a non-negative parameter that can be trained, WN(·) is the normalization operation performed on α 1 W 1 + α 2 W 2 + … + α m W m , Sign(·) is the sign function, is the second weight matrix of the m-th layer. Specifically, the specific form of the first quantization model is not limited here, as long as the first quantization model makes the second weight matrix of the current layer related to the first weight matrices of at least two previous layers, then it belongs to the first quantization model described in this application.

[0122] It should be noted that in some embodiments of this application, the first linear combination parameters can be optimized through but not limited to the following two methods:

[0123] a. Set the first linear combination parameter as the network parameter of the neural network, so that during the training of the neural network based on the training data in the training set, the first linear combination parameter can be trained simultaneously.

[0124] b. Determine that the modulus values of the first weight matrix and the second weight matrix of the m-th layer are α in the first linear combination parameter m , and perform linear regression on the modulus value to obtain the final value of α m , as shown in formula (6):

[0125]

[0126] where, W m is the first weight matrix of the m-th layer, is the second weight matrix of the m-th layer.

[0127] For any layer in the first neural network module, it can be obtained by calculating the modulus value and performing linear regression on the modulus value, which will not be elaborated here.

[0128] It should be noted that in some embodiments of the present application, the weight gain of the second weight matrix of the m-th layer can be further determined, and the second weight matrix of the m-th layer is adjusted according to the weight gain of the second weight matrix of the m-th layer, so that the difference between the adjusted second weight matrix of the m-th layer and the first weight matrix of the m-th layer is less than the difference between the second weight matrix of the m-th layer and the first weight matrix of the m-th layer. The advantage of adjusting the second weight matrix using the weight gain is that: the adjusted second weight matrix is closer to the initial 32-bit floating-point first weight matrix, so that in practical applications, the accuracy of image information can be better retained.

[0129] It should be noted that in some embodiments of the present application, the weight gain of the second weight matrix of the m-th layer can be a trainable non-negative coefficient, as shown in formula (7):

[0130]

[0131] where, γ m is a trainable non-negative coefficient of the m-th layer, is the second weight matrix of the m-th layer, is the adjusted second weight matrix of the m-th layer, is obtained by dot multiplication of and γ m .

[0132] For easy understanding, please refer to Figure 5, assume that the weight matrix of the 5th layer of the first neural network module is a 3×3 matrix, and γ obtained through training 5 = 4, then after γ 5 adjustment, the obtained is as Figure 5 shown in the right part. It should be noted that for different m, the obtained values of γ m are different. For example, in Figure 5 , γ of the 5th layer 5 = 4, and the value of γ of the 3rd layer 3 after training may be 2 or the like.

[0133] It should also be noted that in some embodiments of the present application, if then the weight gain of the second weight matrix of the mth layer can be as shown in formula (8):

[0134]

[0135] where c in is the input channel, c 0ut is the output channel, and k×k is the size of the convolutional kernel of the mth layer. Duplicating E(|W m | c in ×k×k times to form a new matrix The obtained new matrix does not actually change in essence, but only changes in form. It still represents the weight gain of the second weight matrix of the mth layer. At this time, the adjusted second weight matrix of the mth layer can be represented by formula (9):

[0136]

[0137] It should be noted that in the above embodiments, only the binarization operation is performed on the weight matrix, and the feature representations (also called feature maps, activation values, etc.) of each layer are still represented by 32-bit floating-point numbers. When the weight matrix and the feature representation are operated, it is still necessary to use 32-bit floating-point numbers, and the calculation overhead cannot be saved. Only part of the storage space occupied by the neural network model is reduced. Therefore, in some embodiments of the present application, the feature representations output by each layer of the first neural network module can be further binarized, so that the binarized weight matrix and the binarized feature representation can directly perform bit operations, reducing the calculation overhead.

[0138] In some embodiments of the present application, the specific process of binarizing the feature representations of each layer of the first neural network module may be achieved by, but not limited to, the following method: First, the training device calculates the first feature representation of each layer of the first neural network module in the order of connection of the n convolutional layers of the first neural network module, and obtains a second quantization model, which is used to obtain the second feature representation of the m-th layer of the first neural network module according to the m first feature representations of the 1st to m-th layers of the first neural network module. Among them, the first feature representation of each layer is a feature representation expressed in 32-bit floating-point numbers, and the second feature representation of each layer is a feature representation assigned +1 or -1, where 1 ≤ m ≤ n. Finally, according to this second quantization model, each first feature representation corresponding to each layer of the first neural network module is subjected to a binarization operation to obtain each second feature representation corresponding to each layer of the first neural network module.

[0139] It should be noted that, in some embodiments of the present application, the training device performs a binarization operation on each first feature representation corresponding to each layer of the first neural network module according to this second quantization model to obtain each second feature representation corresponding to each layer of the first neural network module. Specifically, it includes, but is not limited to, the following methods: 1) By obtain the second feature representation of the m-th layer, where A 1 , A2, …, A m are the first feature representations from the 1st layer to the m-th layer, and β 1 , β 2 , …, β m are the second linear combination parameters corresponding to A 1 , A 2 , …, A m , BN(·) is the normalization operation performed on β 1 A 1 + β 2 A 2 + … + β m A m , and Sign(·) is the sign function. is the second feature representation of the m-th layer. 2) By obtain the second feature representation of the m-th layer, where A 1 , A 2 , …, A m are the first feature representations from the 1st layer to the m-th layer, and β 1 , β 2 , …, β m are the second linear combination parameters corresponding to A 1 , A 2 , …, A mThe corresponding second linear combination parameter, where l is a non - negative trainable parameter, and BN(·) is the normalization operation performed on β 1 A 1 +β 2 A 2 +…+β m A m Sign(·) is the sign function. is the second feature representation of the m - th layer. Specifically, the specific form of the second quantization model is not limited here. As long as the second quantization model makes the second feature representation of the current layer related to the first feature representations of at least two previous layers, it belongs to the second quantization model described in this application.

[0140] It should be noted that in some embodiments of this application, the second linear combination parameter can be optimized in the following two ways, but not limited to:

[0141] a. Set the second linear combination parameter as the network parameter of the neural network. In this way, during the training of the neural network according to the training data in the training set, the second linear combination parameter can be trained simultaneously.

[0142] b. Determine that the modulus value of the first feature representation of the m - th layer and the second feature representation of the m - th layer is β in the second linear combination parameter m and perform linear regression on this modulus value to obtain the final value of β m as shown in formula (10):

[0143]

[0144] where A m is the first feature representation of the m - th layer, is the second feature representation of the m - th layer.

[0145] For any layer in the first neural network module, it can be obtained by calculating the modulus value and performing linear regression on the modulus value, which will not be elaborated here.

[0146] It should be noted that in some embodiments of this application, the activation gain of the second feature representation of the m - th layer can be further determined, and the second feature representation of the m - th layer can be adjusted according to the activation gain of the second feature representation of the m - th layer, so that the difference between the adjusted second feature representation of the m - th layer and the first feature representation of the m - th layer is less than the difference between the second feature representation of the m - th layer and the first feature representation of the m - th layer before adjustment. The advantage of adjusting the second feature representation using the activation gain is that the adjusted second feature representation is closer to the first feature representation of the initial 32 - bit floating - point number. Since the feature representation has a greater impact on the accuracy of image information, in practical applications, the accuracy of image information retention is further improved.

[0147] It should be noted that in some embodiments of the present application, the activation gain represented by the second feature of the m-th layer may be a trainable non-negative coefficient, as shown in formula (11):

[0148]

[0149] where σ m is a trainable non-negative coefficient of the m-th layer, is the second feature representation of the m-th layer, is the adjusted second feature representation of the m-th layer, is obtained by dot-multiplying with σ m . The specific process is similar to Figure 5 and will not be elaborated here.

[0150] It should also be noted that in some embodiments of the present application, if then the activation gain of the second feature representation of the m-th layer can also be as shown in formula (12):

[0151]

[0152] where c in is the input channel, c 0ut is the output channel, k×k is the size of the convolutional kernel of the m-th layer, N is the number of feature representations input at one time, H×W is the size of the feature representation, W m is the first weight matrix of the m-th layer, A m is the first feature representation of the m-th layer. Duplicating E(|A m | c out times to form a new matrix The resulting new matrix does not actually change in essence, but only changes in form, and it still represents the activation gain of the second feature representation of the m-th layer. At this time, the adjusted second feature representation of the m-th layer can be expressed by formula (13):

[0153]

[0154] It should be noted here that since the feature representations of the neural network are calculated layer by layer, and the normal convolution operation is also calculated layer by layer backward, when calculating the feature representation of the second layer of the first neural network module, the feature representation of the first layer has already been calculated. Therefore, in some embodiments of the present application, the second feature representation of the first layer of the first neural network module is directly obtained through the Sign function on the first feature representation of the first layer. When calculating the second feature representation of the second layer and subsequent layers, the first feature representation of the m-th layer can be calculated according to the second weight matrix of the m-th layer, the second feature representation of the (m - 1)-th layer, the weight gain of the second weight matrix of the m-th layer, and the activation gain of the second feature representation of the (m - 1)-th layer. Specifically, the first feature representation of the m-th layer can be obtained through the following formula (14):

[0155]

[0156] where, is the convolution operation, and is A m is the first feature representation of the m-th layer, is the second feature representation of the (m - 1)-th layer, is the second weight matrix of the m-th layer, is the weight gain of the second weight matrix of the m-th layer, is the activation gain of the second feature representation of the (m - 1)-th layer.

[0157] In some embodiments of the present application, the first feature representation of the m-th layer can also be obtained through the following formula (15):

[0158]

[0159] where, is the convolution operation, A m is the first feature representation of the m-th layer, is the second feature representation of the (m - 1)-th layer, is the second weight matrix of the m-th layer, γ m is the weight gain of the second weight matrix of the m-th layer, σ m-1 is the activation gain of the second feature representation of the (m - 1)-th layer.

[0160] In the embodiments of the present application, the specific processes of the above steps 401 to 402 can be referred to Figure 6 , which will not be elaborated here.

[0161] 403. Train the neural network with the training data in the training set to obtain the trained neural network.

[0162] After the training device binarizes the first weight matrix of each layer of the first neural network module into a second weight matrix according to the first quantization model through the above steps, or, after binarizing the first weight matrix of each layer of the first neural network module into a second weight matrix according to the first quantization model and binarizing the first feature representation of each layer of the first neural network module into a second feature representation according to the second quantization model, the neural network can be further trained with the training data in the training set to obtain the trained neural network.

[0163] 404. Deploy the trained neural network on the target device.

[0164] After obtaining the trained neural network, the neural network can be deployed on the target device.

[0165] It should be noted that in the embodiments of the present application, the target device may specifically be a mobile device, such as edge devices like cameras and smart homes, or end-side devices such as mobile phones, personal computers, computer workstations, tablets, smart wearable devices (such as smart watches, smart bracelets, smart earphones, etc.), game consoles, set-top boxes, media consumption devices, etc. Specifically, the type of the target device is not limited here.

[0166] It should also be noted that if the first neural network module only includes a non-linear mapping stage, then in some embodiments of the present application, the neural network may further include a second neural network module and a third neural network module, where the second neural network module, the first neural network module, and the third neural network module are connected in sequence. Among them, the second neural network module (i.e., the module corresponding to the ε stage, generally a convolutional layer) is used to perform full-precision feature extraction on the input image, and the third neural network module (i.e., the stage corresponding module, generally a convolutional layer) is used to reconstruct the image from the output of the first neural network module to obtain the output image. It should be noted here that the purpose of the second neural network module and the third neural network module is to use full-precision convolutional processes in the ε and stages to ensure the performance of the model and make the accuracy of the final output image higher. It should be noted that in some embodiments of the present application, when the neural network is applied to the scenario of image super-resolution reconstruction, the input image can be one or more low-resolution images, and the output image will be a high-resolution image.

[0167] In the above embodiments of the present application, a new quantization model (i.e., the first quantization model) is used to binarize the weight matrix of the neural network. The first quantization model is used to obtain the second weight matrix of the m-th layer of the neural network according to the m first weight matrices of the 1st to m-th layers of the neural network. Among them, the first weight matrix of each layer is a weight matrix represented by 32-bit floating-point numbers, and the second weight matrix of each layer is a weight matrix assigned +1 or -1. In this way, the value of the adjusted weight matrix of each layer (e.g., the weight matrix of the m-th layer) is related to the values of the weight matrices of the previous layers (e.g., the 1st to m-1-th layers) before adjustment. This optimization method makes the value of each weight in the weight matrix of each layer not only related to itself, but also related to the weight matrices of other layers, reducing the quantization error and making the training and use of the neural network more efficient.

[0168] After introducing the optimization method of the neural network of the present application, the following introduces a network structure of a neural network provided by an embodiment of the present application. Please refer to Figure 7 , Figure 7 The application scenario of the illustrated neural network is an image super-resolution reconstruction scenario. Therefore, the input image is a low-resolution image, and the output image after being processed by the neural network is a high-resolution image. For details, please refer to Figure 7 , the network structure of the neural network includes a first neural network module 701, a second neural network module 702, and a third neural network module 703. Among them, the first neural network module 701 includes n convolutional layers. The second neural network module 702 is used to perform full-precision feature extraction on the input image to obtain a first target feature representation; the first neural network module 701 is used to perform a non-linear mapping on the first target feature representation to obtain a second target feature representation; among them, the weight matrix of each layer of the first neural network module 701 is the second weight matrix processed by the first quantization model. The first quantization model is used to obtain the second weight matrix of the m-th layer of the first neural network module 701 according to the m first weight matrices of the 1st to m-th layers of the first neural network module 701. The first weight matrix of each layer is a weight matrix represented by 32-bit floating-point numbers, and the second weight matrix of each layer is a weight matrix assigned +1 or -1, 1 ≤ m ≤ n; the third neural network module 703 is used to perform image reconstruction on the second target feature representation to obtain an output image.

[0169] In the above embodiments of the present application, a network structure of a neural network is introduced. The difference between this neural network and other neural networks is that the weight matrices of each layer of the first neural network module 701 are binarized by the first quantization model, so that the value of the binarized weight matrix (i.e., the second weight matrix) of each layer of the first neural network module 701 is not only related to itself, but also related to the values of all non-binarized weight matrices (i.e., the first weight matrices) of the previous layers, reducing the quantization error and making the training and use of the neural network more efficient.

[0170] It should be noted that, in some embodiments of the present application, the first quantization model may be: Where W 1 , W 2 , …, W m are the first weight matrices of the first layer to the m-th layer of the first neural network module 701, α 1 , α 2 , …, α m are the first linear combination parameters corresponding to W 1 , W 2 , …, W m , WN(·) is the normalization operation performed on α 1 W 1 + α 2 W 2 + … + α m W m , Sign(·) is the sign function, is the second weight matrix of the m-th layer. The first quantization model may also be: Where W 1 , W 2 , …, W m are the first weight matrices of the first layer to the m-th layer, α 1 , α 2 , …, α m are the first linear combination parameters corresponding to W 1 , W 2 , …, W m , k is a non-negative parameter that can be trained, WN(·) is the normalization operation performed on α 1 W 1 + α 2 W 2 + … + α m W m , Sign(·) is the sign function, is the second weight matrix of the m-th layer. Specifically, the specific form of the first quantization model is not limited here. As long as the first quantization model makes the second weight matrix of the current layer related to the first weight matrices of at least the previous two layers, it belongs to the first quantization model described in this application.

[0171] In the above embodiments of this application, several specific forms of the first quantization model are given, which have flexibility.

[0172] It should also be noted that, in some embodiments of this application, the feature representations of each layer of the first neural network module 701 are second feature representations processed by a second quantization model. The second quantization model is used to obtain the second feature representation of the m-th layer of the first neural network module 701 according to the m first feature representations of the first layer to the m-th layer of the first neural network module 701. Among them, the first feature representation of each layer is a feature representation represented by 32-bit floating-point numbers, and the second feature representation of each layer is a feature representation assigned +1 or -1.

[0173] In the above embodiments of this application, not only the weight matrix of the first neural network module 701 is binarized, but also the feature representation of the first neural network module 701 is further binarized through the second quantization model. In this way, the binarized weight matrix and the binarized feature representation can directly perform bit operations, reducing the calculation overhead.

[0174] It should also be noted that, in some embodiments of this application, the second quantization model can be: where A 1 , A 2 , …, A m are the first feature representations of the first layer to the m-th layer of the first neural network module 701, β 1 , β 2 , …, β m are the second linear combination parameters corresponding to A 1 , A 2 , …, A m , BN(·) is the normalization operation performed on β 1 A 1 + β 2 A 2 + … + β m A m , Sign(·) is the sign function, is the second feature representation of the m-th layer. The second quantization model can also be: where A 1 , A 2 , …, A mis the first feature representation from the first layer to the m-th layer, β 1 , β 2 , …, β m is the second linear combination parameter corresponding to A 1 , A 2 , …, A m ; l is a non - negative trainable parameter, BN(·) is the normalization operation performed on β 1 A 1 + β 2 A 2 + … + β m A m , and Sign(·) is the sign function. is the second feature representation of the m-th layer. Specifically, the specific form of the second quantization model is not limited here. As long as the second quantization model makes the second feature representation of the current layer related to the first feature representations of at least the previous two layers, it belongs to the second quantization model described in this application.

[0175] In the above - mentioned embodiments of this application, several specific forms of the second quantization model are given, which have flexibility.

[0176] It should be noted that Figure 7 is only an application scenario of the neural network optimized in the embodiments of this application in the image super - resolution reconstruction scenario. In practical applications, since the trained neural network in the embodiments of this application can be used in fields such as intelligent security, safe city, and intelligent terminals to perform task processing (such as image processing, audio processing, semantic analysis, etc.). For example, the trained neural network of this application can be applied to various scenarios and problems in the field of computer vision, such as some common tasks: face recognition, image classification, object detection, semantic segmentation, image super - resolution reconstruction, etc. Each type of scenario involves many efficient neural network models that can be binarized using this application. Below, multiple application scenarios implemented in products will be introduced.

[0177] (1) Image super - resolution reconstruction

[0178] Image super-resolution reconstruction is an image processing technology that improves the resolution of images and has been widely used in many fields, such as video surveillance, medical imaging, remote sensing image processing, etc. With the continuous development of deep learning, convolutional neural networks have made great progress in the field of image super-resolution. However, the continuously deepening convolutional network brings too high storage costs and computational complexity, severely limiting the application of image super-resolution reconstruction models on embedded mobile devices. Therefore, the image super-resolution reconstruction model needs to effectively reduce the consumption of storage and computing resources to meet the needs of existing resource-limited devices. Therefore, the trained neural network of this application can be used as the neural network model for image super-resolution reconstruction. For details, please refer to Figure 8 , because the binarized weight matrices of each layer of the trained neural network of this application are not only related to itself but also related to the weight matrices of other layers. Therefore, the detailed information of image pixels is effectively retained, greatly improving the accuracy of the output image.

[0179] (2) Object detection

[0180] As an example, the trained neural network of this application can be used for object detection on terminals (such as mobile phones, smart watches, personal computers, etc.). For details, please refer to Figure 9 , taking the terminal as a mobile phone as an example, object detection on the mobile phone side is a target detection problem. When the user takes a photo with the mobile phone, automatically capturing targets such as human faces and animals can help the mobile phone autofocus, beautify, etc. Therefore, the mobile phone needs a neural network model for target detection with a small volume and fast operation. Therefore, the trained neural network of this application can be used as the neural network model for the mobile phone. Since the weight matrix of this trained neural network is binarized, and the binarized weight matrix is not only related to itself but also related to the weight matrices of other layers. On the premise that both its computational amount and the number of parameters of the neural network are greatly reduced compared with the previous neural network, the detailed information of image pixels is effectively retained. This makes the mobile phone more fluent when performing the above-mentioned target detection, and the picture quality is also clearer than that of the existing binarized neural network. And the fluency can bring a better user experience to the user and improve the quality of mobile phone products.

[0181] (3) Autonomous driving scene segmentation

[0182] As another example, the trained neural network of this application can also be used for autonomous driving scene segmentation of wheeled mobile devices (such as autonomous driving vehicles, assisted driving vehicles, etc.). For details, please refer to Figure 10, taking a wheeled mobile device as an example of an autonomous vehicle, autonomous driving scene segmentation is a semantic segmentation problem. The camera of the autonomous vehicle captures the road scene, and it is necessary to segment the scene to distinguish different objects such as the road surface, roadbed, vehicle, and pedestrian, so as to keep the vehicle driving in the correct safe area. For autonomous driving with extremely high safety requirements, it is necessary to understand the scene in real time. Therefore, a convolutional neural network capable of performing semantic segmentation in real time is crucial. Since the number of parameters and the amount of computation of the neural network trained in this application are greatly reduced compared to the previous neural networks, it is smaller in size and faster in operation, and can well meet the above series of requirements for the convolutional neural network of autonomous vehicles. Therefore, the neural network trained in this application can also be used as a neural network model for autonomous driving scene segmentation of wheeled mobile devices.

[0183] It should be noted that the wheeled mobile device described in this application can be a wheeled robot, a wheeled construction equipment, an autonomous vehicle, etc. As long as it is a device with wheeled mobility, it belongs to the wheeled mobile device described in this application. In addition, it should also be noted that the autonomous vehicle described above in this application can be a car, a truck, a motorcycle, a bus, a ship, an airplane, a helicopter, a lawn mower, a recreational vehicle, a playground vehicle, a construction equipment, a tram, a golf cart, a train, and a trolley, etc. The embodiments of this application do not make special limitations.

[0184] (4) Face recognition

[0185] As another example, the neural network trained in this application can also be used for face recognition (such as face verification at the entrance gate). For details, please refer to Figure 11 , face recognition is a problem of image similarity comparison. When passengers perform face authentication at the gates at high-speed rail stations, airports, etc., the camera will capture the face image, extract features using a convolutional neural network, and calculate the similarity with the image features of the identity document stored in the system. If the similarity is high, the verification is successful. Among them, extracting features using a convolutional neural network is the most time-consuming. To quickly perform face verification, an efficient convolutional neural network for feature extraction is required. Since the neural network trained in this application has a small number of parameters and a low amount of computation, it is smaller in size and faster in operation, and can well meet the above series of requirements for the convolutional neural network in the application scenario of face recognition.

[0186] (5) Speech recognition

[0187] As another example, the neural network trained in this application can also be used for speech recognition (such as simultaneous interpretation of a translator). For details, please refer to Figure 12, simultaneous interpretation by a translation machine is a problem of speech recognition and machine translation. In the problems of speech recognition and machine translation, convolutional neural networks are also commonly used recognition models. In scenarios that require simultaneous interpretation, real-time speech recognition and translation must be achieved, which requires the convolutional neural network deployed on the device to have a fast computing speed. The trained neural network of this application has a small number of parameters and a low computational load, making it smaller in size and faster in operation, and can also well meet a series of requirements of the above speech recognition application scenarios for convolutional neural networks.

[0188] It should be noted that the trained neural network described in this application can be applied not only to the Figures 8 to 12 application scenarios described above, but also to various sub-fields in the field of artificial intelligence, such as the field of image processing, the field of computer vision, the field of semantic analysis, etc. As long as it is a field and device that can use neural networks, the trained neural network provided by the embodiments of this application can be applied, and no further examples will be given here.

[0189] In order to have a more intuitive understanding of the beneficial effects brought by the embodiments of this application, the following further compares the technical effects brought by the embodiments of this application. The application scenario is image super-resolution reconstruction. Please refer to Table 1, Table 2, Table 3 and Figure 13 , Figure 14 . Among them, it can be seen from Table 1 (VDSR-BAM is the solution provided by the embodiment of this application) and Table 2 (SRResNet-BAM is the solution provided by the embodiment of this application) that the binarization algorithm provided by the embodiment of this application is significantly better than other algorithms in the objective evaluation indicators: peak signal-to-noise ratio (PSNR), structural similarity (SSIM), and visual evaluation, and has obtained a large performance improvement in the field of image super-resolution reconstruction.

[0190] Table 1: Performance comparison of the binarization algorithm based on VDSR

[0191]

[0192] Table 2: Performance comparison of the binarization algorithm based on SRRestNet

[0193]

[0194] In addition to comparing with existing binarization algorithms, the present application also compares the neural network provided in the embodiments of the present application with other existing neural networks that only binarize the weight matrix. In the experiment, the settings of the model structure are completely consistent with this method (that is, only the weight matrix is quantized, and the activation value is 32-bit floating-point). As can be seen from Table 3, the method provided in the embodiments of the present application can achieve better results.

[0195] Table 3: Performance comparison with other existing neural networks that only binarize the weight matrix

[0196]

[0197] Based on the above embodiments, in order to better implement the above solutions of the embodiments of the present application, the following also provides related devices for implementing the above solutions. Specifically, refer to Figure 15 , Figure 15 which is a schematic diagram of a training device provided in an embodiment of the present application. The training device 1500 may specifically include: an acquisition unit 1501, configured to acquire a first quantization model, where the first quantization model is used to obtain a second weight matrix of the mth layer of the first neural network module according to m first weight matrices of the first neural network module from the first layer to the mth layer, where each first weight matrix of each layer is a weight matrix represented by 32-bit floating-point numbers, and each second weight matrix of each layer is a weight matrix assigned +1 or -1, 1 ≤ m ≤ n; a quantization unit 1502, configured to perform a binarization operation on each first weight matrix corresponding to each layer of the first neural network module according to the first quantization model to obtain each second weight matrix corresponding to each layer of the first neural network module; a training unit 1503, configured to train the neural network through training data in a training set to obtain a trained neural network; and a deployment unit 1504, configured to deploy the trained neural network on a target device.

[0198] In the above implementation manner of the present application, a new quantization model (i.e., the first quantization model) is used to binarize the weight matrix of the neural network. The first quantization model is used to obtain the second weight matrix of the mth layer of the neural network according to m first weight matrices of the neural network from the first layer to the mth layer, where each first weight matrix of each layer is a weight matrix represented by 32-bit floating-point numbers, and each second weight matrix of each layer is a weight matrix assigned +1 or -1. In this way, the value of each adjusted weight matrix (e.g., the weight matrix of the mth layer) is related to the values of the weight matrices before adjustment of the previous layers (e.g., from the first layer to the m - 1th layer). This optimization method makes the values of each weight in each layer of the weight matrix not only related to itself but also related to the weight matrices of other layers, reducing the quantization error and making the training and use of the neural network more efficient.

[0199] In a possible design, the quantization unit 1502 is specifically configured to: obtain the second weight matrix of the m-th layer by where W 1 , W 2 , …, W m are the first weight matrices from the first layer to the m-th layer, α 1 , α2, …, α m are the first linear combination parameters corresponding to W 1 , W 2 , …, W m , WN(·) is the normalization operation performed on α 1 W 1 + α 2 W 2 + … + α m W m , and Sign(·) is the sign function, is the second weight matrix of the m-th layer. It can also obtain the second weight matrix of the m-th layer by where W 1 , W 2 , …, W m are the first weight matrices from the first layer to the m-th layer, α 1 , α 2 , …, α m are the first linear combination parameters corresponding to W 1 , W 2 , …, W m , k is a non-negative parameter that can be trained, WN(·) is the normalization operation performed on α 1 W 1 + α 2 W 2 + … + α m W m , and Sign(·) is the sign function, is the second weight matrix of the m-th layer. There is no limitation here.

[0200] In the above embodiments of the present application, several different quantization forms of the first weight matrix are provided, which have selectivity and flexibility.

[0201] In a possible design, the quantization unit 1502 is further configured to: determine the weight gain of the second weight matrix of the m-th layer, and adjust the second weight matrix of the m-th layer according to the weight gain of the second weight matrix of the m-th layer, so that the difference between the adjusted second weight matrix of the m-th layer and the first weight matrix of the m-th layer is less than the difference between the second weight matrix of the m-th layer and the first weight matrix of the m-th layer.

[0202] In the above embodiments of the present application, the advantage of adjusting the second weight matrix using weight gain is that the adjusted second weight matrix is closer to the initial first weight matrix of 32-bit floating-point numbers. In this way, in practical applications, the accuracy of image information can be better retained.

[0203] In a possible design, the quantization unit 1502 is further configured to: set the first linear combination parameter as the network parameter of the neural network, so that during the training of the neural network according to the training data in the training set, the first linear combination parameter is trained.

[0204] In the above embodiments of the present application, a specific implementation manner for optimizing the first linear combination parameter is provided. The advantage of this optimization process is that during the training of the neural network, the optimization of the first linear combination parameter is completed simultaneously.

[0205] In a possible design, the quantization unit 1502 is further configured to: determine that the modulus value of the first weight matrix and the second weight matrix of the m-th layer is α in the first linear combination parameter m and perform linear regression on the modulus value to obtain the final value of α m .

[0206] In the above embodiments of the present application, another specific implementation manner for optimizing the first linear combination parameter is provided. By using the method of linear regression to obtain the value of each parameter in the first linear combination parameter, the optimization method of the first linear combination parameter has selectivity.

[0207] In a possible design, the obtaining unit 1501 is further configured to calculate the first feature representation of each layer of the first neural network module in sequence according to the connection sequence of the n convolutional layers; the obtaining unit 1501 is further configured to obtain a second quantization model, where the second quantization model is used to obtain the second feature representation of the m-th layer of the first neural network module according to the m first feature representations of the 1st to m-th layers of the first neural network module, where the first feature representation of each layer is a feature representation represented by 32-bit floating-point numbers, and the second feature representation of each layer is a feature representation assigned +1 or -1, 1≤m≤n; the quantization unit 1502 is further configured to perform a binarization operation on each first feature representation corresponding to each layer of the first neural network module according to the second quantization model to obtain each second feature representation corresponding to each layer of the first neural network module. It should be noted that there is no sequence requirement between the obtaining unit 1501 calculating the first feature representation of each layer of the first neural network module and obtaining the second quantization model. The obtaining unit 1501 may first calculate the first feature representation of each layer of the first neural network module and then obtain the second quantization model; the obtaining unit 1501 may first obtain the second quantization model and then calculate the first feature representation of each layer of the first neural network module. Specifically, no limitation is made here.

[0208] Since only the weight matrix is binarized, the feature representations (which can also be referred to as feature maps, activation values, etc.) of each layer are still represented by 32-bit floating-point numbers. When the weight matrix and the feature representations are operated, it is still necessary to perform operations through 32-bit floating-point numbers, and the computing overhead cannot be saved. Only part of the space occupied by the storage of the neural network model is reduced. Therefore, in the above embodiments of the present application, the feature representations output by each layer of the first neural network module are further binarized, so that the binarized weight matrix and the binarized feature representations can directly perform bit operations, reducing the computing overhead.

[0209] In a possible design, the quantization unit 1502 is further configured to: can pass through to obtain the second feature representation of the m-th layer, where A 1 , A 2 , …, A m are the first feature representations from the 1st layer to the m-th layer, β 1 , β 2 , …, β m are the second linear combination parameters corresponding to A 1 , A 2 , …, A m , and BN(·) is to β 1 A 1 +β 2 A 2 + … + βm A m The normalization operation is performed, Sign(·) is the sign function, is the second feature representation of the m-th layer. It can also be obtained by to obtain the second feature representation of the m-th layer, where A 1 , A 2 , …, A m are the first feature representations from the first layer to the m-th layer, β 1 , β 2 , …, β m are the second linear combination parameters corresponding to A 1 , A 2 , …, A m , l is a non-negative trainable parameter, BN(·) is the normalization operation performed on β 1 A 1 + β 2 A 2 + … + β m A m , Sign(·) is the sign function, is the second feature representation of the m-th layer. Specifically, it is not limited here.

[0210] In the above embodiments of the present application, several different quantization forms of the first feature representation are provided, which have selectivity and flexibility.

[0211] In a possible design, the quantization unit 1502 is further configured to: determine the activation gain of the second feature representation of the m-th layer, and adjust the second feature representation of the m-th layer according to the activation gain of the second feature representation of the m-th layer, so that the difference between the adjusted second feature representation of the m-th layer and the first feature representation of the m-th layer is less than the difference between the second feature representation of the m-th layer and the first feature representation of the m-th layer.

[0212] In the above embodiments of the present application, the advantage of adjusting the second feature representation using the activation gain is that: the adjusted second feature representation is closer to the initial 32-bit floating-point first feature representation. Since the feature representation has a greater impact on the accuracy of the image information, in practical applications, the accuracy of the retained image information is further improved.

[0213] In a possible design, the obtaining unit 1501 is further specifically configured to: calculate the first feature representation of the m-th layer according to the second weight matrix of the m-th layer, the second feature representation of the (m - 1)-th layer, the weight gain of the second weight matrix of the m-th layer, and the activation gain of the second feature representation of the (m - 1)-th layer. It should be noted here that since the feature representations of the neural network are calculated layer by layer, and the normal convolution operation is also calculated layer by layer backward, when the obtaining unit 1501 calculates the feature representation of the second layer of the first neural network module, the feature representation of the first layer has already been calculated. Therefore, in some embodiments of the present application, the second feature representation of the first layer of the first neural network module is directly obtained through the Sign function on the first feature representation of the first layer. When the obtaining unit 1501 calculates the second feature representations of the second layer and subsequent layers, it can calculate the first feature representation of the m-th layer according to the second weight matrix of the m-th layer, the second feature representation of the (m - 1)-th layer, the weight gain of the second weight matrix of the m-th layer, and the activation gain of the second feature representation of the (m - 1)-th layer.

[0214] In the above embodiments of the present application, an implementation manner for calculating the first feature representation is provided, which has feasibility.

[0215] In a possible design, the obtaining unit 1501 is further specifically configured to: perform a convolution operation on the second weight matrix of the m-th layer and the second feature representation of the (m - 1)-th layer to obtain a convolution result. Then, perform a dot product operation on the convolution result and the weight gain of the second weight matrix of the m-th layer to obtain a dot product result. Finally, perform a dot product operation on the dot product result and the activation gain of the second feature representation of the (m - 1)-th layer to obtain the first feature representation of the m-th layer.

[0216] In the above embodiments of the present application, a method for obtaining the first feature representation is specifically described, which has feasibility.

[0217] In a possible design, the quantization unit 1502 is further configured to: set the second linear combination parameter as the network parameter of the neural network, so that during the training of the neural network according to the training data in the training set, the second linear combination parameter is trained.

[0218] In the above embodiments of the present application, a specific implementation manner for optimizing the second linear combination parameter is provided. The advantage of this optimization process is that during the training of the neural network, the optimization of the second linear combination parameter is completed simultaneously.

[0219] In a possible design, the quantization unit 1502 is further configured to: determine that the modulus value of the first feature representation of the m-th layer and the second feature representation of the m-th layer is β in the second linear combination parameter, m and perform linear regression on the modulus value to obtain the final value of β. m

[0220] In the above embodiments of the present application, another specific implementation manner for optimizing the second linear combination parameter is provided. The value of each parameter in the second linear combination parameter is obtained by linear regression, so that the optimization method of the second linear combination parameter has selectivity.

[0221] In a possible design, the neural network further includes a second neural network module and a third neural network module. The second neural network module is configured to perform full-precision feature extraction on the input image, and the third neural network module is configured to perform image reconstruction on the output of the first neural network module to obtain an output image.

[0222] In the above embodiments of the present application, it is described that in addition to including the first neural network module, the neural network may further include a second neural network module and a third neural network module. Among them, the second neural network module is configured to perform full-precision feature extraction on the input image, and the third neural network module is configured to perform image reconstruction on the output of the first neural network module to obtain an output image. The purpose of the second neural network module and the third neural network module is to adopt a full-precision convolution process in the feature extraction stage and the image reconstruction stage, so as to ensure the performance of the model and make the accuracy of the final output image higher.

[0223] In a possible design, the input image includes one or more low-resolution images, and the output image includes one high-resolution image.

[0224] In the above embodiments of the present application, when the neural network is applied to the scenario of image super-resolution reconstruction, the input image may be one or more low-resolution images, and the output image will be one high-resolution image.

[0225] It should be noted that the information interaction, execution process, etc. between the modules / units in the training device 1500 are based on the same concept as the corresponding method embodiments in the present application. For specific content, reference may be made to the description in the method embodiments shown above in the present application, which will not be elaborated here. Figure 4

[0226] The embodiments of the present application further provide an execution device. Please refer to Figure 16 Figure 16Schematic diagram of an execution device provided by an embodiment of the present application. The execution device 1600 includes: an acquisition unit 1601 and an execution unit 1602. The acquisition unit 1601 is used to acquire an input image. The execution unit 1602 is used to process the input image through a trained neural network to obtain an output image. The trained neural network is a neural network optimized by the corresponding implementation method in the present application Figure 4 The neural network optimized by the corresponding implementation method in the present application. For specific content, reference can be made to the description in the method embodiments shown above in the present application, which will not be elaborated here.

[0227] It should be noted that for the information interaction, execution process, etc. between the modules / units in the execution device 1600, it can be specifically applied to various application scenarios in the corresponding method embodiments of the present application Figures 8 to 12 The corresponding method embodiments of the present application. For specific content, reference can be made to the above Figures 8 to 12 Shown in the method embodiments of the present application, the description will not be elaborated here.

[0228] Next, another training device provided by an embodiment of the present application will be introduced. Please refer to Figure 17 , Figure 17 Schematic diagram of a structure of a training device provided by an embodiment of the present application. The training device 1700 can be deployed with Figure 15 The training device 1500 described in the corresponding embodiment, used to implement Figure 15 The functions of the training device 1500 in the corresponding embodiment. Specifically, the training device 1700 is implemented by one or more servers. The training device 1700 may vary greatly due to configuration or performance differences, and may include one or more central processing units (CPUs) 1722 and a memory 1732, and one or more storage media 1730 (such as one or more mass storage devices) for storing application programs 1742 or data 1744. Among them, the memory 1732 and the storage media 1730 can be transient storage or persistent storage. The program stored in the storage media 1730 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the training device 1700. Further, the central processing unit 1722 can be set to communicate with the storage media 1730 and execute a series of instruction operations in the storage media 1730 on the training device 1700.

[0229] The training device 1700 may also include one or more power supplies 1726, one or more wired or wireless network interfaces 1750, one or more input and output interfaces 1758, and / or, one or more operating systems 1741, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.

[0230] In the embodiment of the present application, the central processor 1722 is used to execute Figure 4 The optimization method of the neural network executed by the training device in the corresponding embodiment. For example, the central processor 1522 can be used to: first, obtain a first quantization model, and the first quantization model is used to obtain the second weight matrix of the mth layer of the first neural network module according to the m first weight matrices of the 1st layer to the mth layer of the first neural network module in the neural network, wherein the first weight matrix of each layer of the first neural network module refers to the initial weight matrix represented by a 32-bit floating point number, and the second weight matrix of each layer of the first neural network module refers to the weight matrix assigned a value of +1 or -1. After obtaining the first quantization model, each first weight matrix corresponding to each layer of the first neural network module is binarized according to the first quantization model to obtain each second weight matrix corresponding to each layer of the first neural network module, and then the neural network is further trained by the training data in the training set to obtain the trained neural network, and finally the trained neural network is deployed on the target device.

[0231] It should be noted that the specific manner in which the CPU 1722 performs the above steps is different from that in the present application. Figure 4 The corresponding method embodiments are based on the same concept, and the technical effects they bring are also the same as those of the above-mentioned embodiments of the present application. For specific contents, please refer to the description in the method embodiments shown above in the present application, and will not be repeated here.

[0232] Next, an execution device provided by an embodiment of the present application is introduced. Figure 18 , Figure 18 A schematic diagram of the structure of the execution device provided in the embodiment of the present application, the execution device 1800 can be specifically manifested as various terminal devices, such as virtual reality VR devices, mobile phones, tablets, laptops, smart wearable devices, monitoring data processing devices or radar data processing devices, etc., which are not limited here. Among them, the execution device 1800 can be deployed with Figure 16 The execution device 1600 described in the corresponding embodiment is used to implement Figure 16Perform the functions of the execution device 1600 in the corresponding embodiment. Specifically, the execution device 1800 includes: a receiver 1801, a transmitter 1802, a processor 1803, and a memory 1804 (where the number of processors 1803 in the execution device 1800 can be one or more, Figure 18 and one processor is taken as an example herein), where the processor 1803 may include an application processor 18031 and a communication processor 18032. In some embodiments of the present application, the receiver 1801, the transmitter 1802, the processor 1803, and the memory 1804 may be connected through a bus or other means.

[0233] The memory 1804 may include a read-only memory and a random access memory, and provide instructions and data to the processor 1803. A part of the memory 1804 may further include a non-volatile random access memory (NVRAM). The memory 1804 stores processor and operation instructions, executable modules, or data structures, or subsets thereof, or extended sets thereof, where the operation instructions may include various operation instructions for implementing various operations.

[0234] The processor 1803 controls the operations of the execution device 1800. In a specific application, the various components of the execution device 1800 are coupled together through a bus system, where the bus system may further include a power bus, a control bus, a status signal bus, etc. in addition to the data bus. However, for the sake of clear illustration, all kinds of buses are referred to as the bus system in the figure.

[0235] The above Figure 4 The method disclosed in the corresponding embodiment of the present application can be applied to the processor 1803 or implemented by the processor 1803. The processor 1803 may be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method may be completed by the integrated logic circuit in the hardware of the processor 1803 or by instructions in the form of software. The above-mentioned processor 1803 may be a general-purpose processor, a digital signal processor (DSP), a microprocessor, or a microcontroller, and may further include an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate, or transistor logic devices, discrete hardware components. The processor 1803 may implement or execute the present application Figure 4The various methods, steps, and logic block diagrams disclosed in the corresponding embodiments. The general-purpose processor can be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by a hardware decoding processor, or can be executed and completed by a combination of hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as random access memory, flash memory, read-only memory, programmable read-only memory, or electrically erasable programmable memory, registers, etc. This storage medium is located in the memory 1804, and the processor 1803 reads the information in the memory 1804 and combines its hardware to complete the steps of the above method.

[0236] The receiver 1801 can be used to receive input digital or character information, and generate signal inputs related to the relevant settings and function controls of the execution device 1800. The transmitter 1802 can be used to output digital or character information through the first interface; the transmitter 1802 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; the transmitter 1802 can also include a display device such as a display screen.

[0237] In an embodiment of the present application, in one case, the processor 1803 is used to process the input image through a trained neural network to obtain an output image. For example, the application processor 18031 can be used to: obtain the input image, and process the input image through a trained neural network to obtain an output image. The trained neural network can be a neural network obtained through the Figure 4 corresponding optimization method of the present application. For specific content, reference can be made to the description in the method embodiments shown above in the present application, and details will not be elaborated here.

[0238] An embodiment of the present application also provides a computer-readable storage medium, in which a program for signal processing is stored. When it runs on a computer, it causes the computer to execute the steps executed by the training device described in the embodiments shown above Figure 4 、 Figure 15 or causes the computer to execute the steps executed by the execution device described in the embodiments shown above Figure 16 as shown.

[0239] The training device, execution device, etc. provided in the embodiments of the present application can specifically be a chip. The chip includes: a processing unit and a communication unit. The processing unit can be, for example, a processor, and the communication unit can be, for example, an input / output interface, a pin, or a circuit, etc. The processing unit can execute the computer execution instructions stored in the storage unit, so that the chip in the training device executes the steps executed by the training device described in the embodiments shown above Figure 4 、 15 as shown, or causes the chip in the execution device to execute as described aboveFigure 16 Steps performed by the execution device described in the illustrated embodiment.

[0240] Optionally, the storage unit is a storage unit within the chip, such as a register, cache, etc., or the storage unit may also be a storage unit outside the chip within the wireless access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.

[0241] Specifically, please refer to Figure 19 , Figure 19 which is a schematic structural diagram of a chip provided by an embodiment of the present application. The chip can be represented as a neural network processor NPU 200. The NPU 200 is mounted as a coprocessor on a main CPU (Host CPU), and tasks are allocated by the Host CPU. The core part of the NPU is the arithmetic circuit 2003. The arithmetic circuit 2003 extracts matrix data from the memory and performs multiplication operations under the control of the controller 2004.

[0242] In some implementations, the arithmetic circuit 2003 includes multiple processing units (process engine, PE) inside. In some implementations, the arithmetic circuit 2003 is a two-dimensional systolic array. The arithmetic circuit 2003 can also be a one-dimensional systolic array or other electronic circuits that can perform mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 2003 is a general matrix processor.

[0243] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit fetches the corresponding data of matrix B from the weight memory 2002 and caches it on each PE in the arithmetic circuit. The arithmetic circuit fetches the data of matrix A from the input memory 2001 and performs matrix operations with matrix B. The partial results or final results of the obtained matrix are saved in the accumulator 2008.

[0244] The unified memory 2006 is used to store input data and output data. The weight data is directly transported through the direct memory access controller (DMAC) 2005, and the DMAC transports it to the weight memory 2002. The input data is also transported to the unified memory 2006 through the DMAC.

[0245] The bus interface unit 2010 (bus interface unit, hereinafter referred to as BIU) is used for the interaction between the AXI bus and the DMAC and the instruction fetch buffer (IFB) 2009.

[0246] The bus interface unit 2010 is used for the instruction fetch buffer 2009 to obtain instructions from the external memory, and is also used for the storage unit access controller 2005 to obtain the original data of the input matrix A or the weight matrix B from the external memory.

[0247] The DMAC is mainly used to transfer the input data in the external memory DDR to the unified memory 2006, or transfer the weight data to the weight memory 2002, or transfer the input data to the input memory 2001.

[0248] The vector calculation unit 2007 includes multiple arithmetic processing units, which, if necessary, further process the output of the arithmetic circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, size comparison, etc. It is mainly used for non-convolution / full connection layer network calculations in neural networks, such as Batch Normalization, pixel-level summation, upsampling of the feature plane, etc.

[0249] In some implementations, the vector calculation unit 2007 can store the processed output vector into the unified memory 2006. For example, the vector calculation unit 2007 can apply a linear function and / or a non-linear function to the output of the arithmetic circuit 2003, such as linear interpolation of the feature plane extracted by the convolutional layer, or for example, a vector of accumulated values to generate activation values. In some implementations, the vector calculation unit 2007 generates normalized values, pixel-level summation values, or both. In some implementations, the processed output vector can be used as the activation input to the arithmetic circuit 2003, for example, for use in subsequent layers in the neural network.

[0250] The instruction fetch buffer 2009 connected to the controller 2004 is used to store the instructions used by the controller 2004;

[0251] The unified memory 2006, the input memory 2001, the weight memory 2002, and the instruction fetch buffer 2009 are all On-Chip memories. The external memory is private to this NPU hardware architecture.

[0252] Wherein, the processor mentioned anywhere above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the program of the method in the first aspect above.

[0253] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in this application, the connection relationship between modules indicates that they have a communication connection, which can be specifically implemented as one or more communication buses or signal lines.

[0254] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures used to implement the same function can also be various, such as analog circuits, digital circuits or dedicated circuits. However, for this application, in more cases, software program implementation is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, USB flash drive, mobile hard disk, ROM, RAM, magnetic disk or optical disc of a computer, and includes several instructions to enable a computer device (which can be a personal computer, training device, or network device, etc.) to execute the methods described in various embodiments of this application.

[0255] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.

[0256] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are wholly or partially generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center by wire (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wirelessly (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.

Claims

1. An optimization method for a neural network, characterized in that, the neural network includes a first neural network module, the first neural network module includes n convolutional layers, and the method includes: obtaining a first quantization model, the first quantization model is used to obtain a second weight matrix of the m-th layer of the first neural network module according to m first weight matrices of the 1st layer to the m-th layer of the first neural network module, wherein each first weight matrix of each layer is a weight matrix represented by 32-bit floating-point numbers, and each second weight matrix of each layer is a weight matrix assigned +1 or -1, 1 ≤ m ≤ n; performing a binarization operation on each first weight matrix corresponding to each layer of the first neural network module according to the first quantization model to obtain each second weight matrix corresponding to each layer of the first neural network module, wherein the neural network is updated based on each second weight matrix; training the neural network with training data in a training set to obtain a trained neural network, wherein a first target feature representation is non-linearly mapped through the first neural network module to obtain a second target feature representation, so that a third neural network module performs image reconstruction on the second target feature representation to obtain an output image, the first target feature representation is obtained by the second neural network module performing full-precision feature extraction on an input image, and the input image belongs to the training set; deploying the trained neural network on a target device.

2. The method according to claim 1, characterized in that, performing a binarization operation on each first weight matrix corresponding to each layer of the first neural network module according to the first quantization model to obtain each second weight matrix corresponding to each layer of the first neural network module includes: By obtain the second weight matrix of the m-th layer, where W 1 , W 2 , …, W m are the first weight matrices from the first layer to the m-th layer, α 1 , α 2 , …, α m are the first linear combination parameters corresponding to W 1 , W 2 , …, W m , WN(·) is the normalization operation performed on α 1 W 1 + α 2 W 2 + … + α m W m , Sign(·) is the sign function for binarizing WN(α 1 W 1 + α 2 W 2 + … + α m W m ), is the second weight matrix of the m-th layer, is the first quantization model.

3. The method according to claim 1, characterized in that, before training the neural network with training data in a training set to obtain a trained neural network, the method further includes: determining a weight gain of the second weight matrix of the m-th layer; adjusting the second weight matrix of the m-th layer according to the weight gain of the second weight matrix of the m-th layer, so that the difference between the adjusted second weight matrix of the m-th layer and the first weight matrix of the m-th layer is less than the difference between the second weight matrix of the m-th layer and the first weight matrix of the m-th layer.

4. The method according to claim 2, characterized in that, before training the neural network with training data in a training set to obtain a trained neural network, the method further includes: setting the first linear combination parameter as a network parameter of the neural network, so that the first linear combination parameter is trained during the process of training the neural network with training data in the training set.

5. The method according to claim 2, characterized in that, before training the neural network with training data in a training set to obtain a trained neural network, the method further includes: Determine that the modulus values of the first weight matrix and the second weight matrix of the m-th layer are α in the first linear combination parameter m ; Perform a linear regression on the modulus value to obtain the final value of the α m .

6. The method according to any one of claims 1-5, characterized in that, After performing a binarization operation on each first weight matrix corresponding to each layer of the first neural network module according to the first quantization model to obtain each second weight matrix corresponding to each layer of the first neural network module, the method further includes: Calculating the first feature representation of each layer of the first neural network module in sequence according to the sequential connection order of the n convolutional layers; Obtaining a second quantization model, which is used to obtain the second feature representation of the m-th layer of the first neural network module according to the m first feature representations of the first layer to the m-th layer of the first neural network module, where the first feature representation of each layer is a feature representation represented by 32-bit floating-point numbers, and the second feature representation of each layer is a feature representation assigned +1 or -1, 1 ≤ m ≤ n; According to the second quantization model, performing a binarization operation on each first feature representation corresponding to each layer of the first neural network module to obtain each second feature representation corresponding to each layer of the first neural network module.

7. The method according to claim 6, wherein performing a binarization operation on each first feature representation corresponding to each layer of the first neural network module according to the second quantization model to obtain each second feature representation corresponding to each layer of the first neural network module includes: By obtain the second feature representation of the m-th layer, where A 1 , A 2 , …, A m are the first feature representations from the first layer to the m-th layer, and β 1 , β 2 , …, β m are the second linear combination parameters corresponding to A 1 , A 2 , …, A m , BN(·) is the normalization operation performed on β 1 A 1 + β 2 A 2 + … + β m A m , and Sign(·) is the sign function for binarizing BN(β 1 A 1 + β 2 A 2 + … + β m A m ), is the second feature representation of the m-th layer, is the second quantization model.

8. The method according to claim 6, characterized in that, Before training the neural network with the training data in the training set to obtain the trained neural network, the method further includes: Determining the activation gain of the second feature representation of the m-th layer; Adjusting the second feature representation of the m-th layer according to the activation gain of the second feature representation of the m-th layer, so that the difference between the adjusted second feature representation of the m-th layer and the first feature representation of the m-th layer is less than the difference between the second feature representation of the m-th layer and the first feature representation of the m-th layer.

9. The method according to claim 8, characterized in that, The calculating the first feature representation of each layer of the first neural network module in sequence according to the sequential connection order of the n convolutional layers includes: Calculating the first feature representation of the m-th layer according to the second weight matrix of the m-th layer, the second feature representation of the (m - 1)-th layer, the weight gain of the second weight matrix of the m-th layer, and the activation gain of the second feature representation of the (m - 1)-th layer.

10. The method according to claim 9, characterized in that, The calculating the first feature representation of the m-th layer according to the second weight matrix of the m-th layer, the second feature representation of the (m - 1)-th layer, the weight gain of the second weight matrix of the m-th layer, and the activation gain of the second feature representation of the (m - 1)-th layer includes: Performing a convolution operation on the second weight matrix of the m-th layer and the second feature representation of the (m - 1)-th layer to obtain a convolution result; Performing a dot product operation on the convolution result and the weight gain of the second weight matrix of the m-th layer to obtain a dot product result; Performing a dot product operation on the dot product result and the activation gain of the second feature representation of the (m - 1)-th layer to obtain the first feature representation of the m-th layer.

11. The method according to claim 7, characterized in that, Before training the neural network with the training data in the training set to obtain the trained neural network, the method further includes: Setting the second linear combination parameter as the network parameter of the neural network, so that during the training of the neural network according to the training data in the training set, the second linear combination parameter is trained.

12. The method according to claim 7, wherein, Before training the neural network with the training data in the training set to obtain the trained neural network, the method further includes: Determine that the modulus value of the first feature representation and the second feature representation of the m-th layer is β in the second linear combination parameter m ; Perform a linear regression on the modulus value to obtain the final value of the β m .

13. The method according to any one of claims 1-5, wherein, The neural network further includes the second neural network module and the third neural network module, and the second neural network module, the first neural network module, and the third neural network module are connected in sequence; The second neural network module is used for performing full-precision feature extraction on the input image, and the third neural network module is used for performing image reconstruction on the output of the first neural network module to obtain the output image.

14. The method according to claim 13, wherein, The input image includes one or more low-resolution images; The output image includes one high-resolution image.

15. An image processing method, wherein, including: Obtaining an input image; Processing the input image through the trained neural network to obtain an output image, where the trained neural network is a neural network optimized by the method according to any one of claims 1-14.

16. A training device, wherein, including: An acquisition unit, configured to acquire a first quantization model, where the first quantization model is used to obtain a second weight matrix of the mth layer of the first neural network module according to m first weight matrices of the first layer to the mth layer of the first neural network module of the neural network. The neural network includes a first neural network module, the first neural network module includes n convolutional layers, the first weight matrix of each layer is a weight matrix represented by 32-bit floating-point numbers, and the second weight matrix of each layer is a weight matrix assigned +1 or -1, 1≤m≤n; A quantization unit, configured to perform a binarization operation on each first weight matrix corresponding to each layer of the first neural network module according to the first quantization model to obtain each second weight matrix corresponding to each layer of the first neural network module, where the neural network is updated based on each second weight matrix; A training unit, configured to train the neural network with the training data in the training set to obtain the trained neural network. Wherein, a second target feature representation is obtained by performing a non-linear mapping on the first target feature representation through the first neural network module, so that the third neural network module performs image reconstruction on the second target feature representation to obtain the output image. The first target feature representation is obtained by the second neural network module performing full-precision feature extraction on the input image, and the input image belongs to the training set; A deployment unit for deploying the trained neural network on a target device.

17. The device according to claim 16, wherein, the quantization unit is specifically configured to: By obtain the second weight matrix of the m-th layer, where W 1 , W 2 , …, W m are the first weight matrices from the first layer to the m-th layer, α 1 , α 2 , …, α m are the first linear combination parameters corresponding to W 1 , W 2 , …, W m , WN(·) is the normalization operation performed on α 1 W 1 + α 2 W 2 + … + α m W m , Sign(·) is the sign function for binarizing WN(α 1 W 1 + α 2 W 2 + … + α m W m ), is the second weight matrix of the m-th layer, is the first quantization model.

18. The device according to claim 16, wherein, the quantization unit is further configured to: Determine the weight gain of the second weight matrix of the m-th layer; Adjust the second weight matrix of the m-th layer according to the weight gain of the second weight matrix of the m-th layer, so that the difference between the adjusted second weight matrix of the m-th layer and the first weight matrix of the m-th layer is less than the difference between the second weight matrix of the m-th layer and the first weight matrix of the m-th layer.

19. The device according to claim 17, wherein, the quantization unit is further configured to: Set the first linear combination parameter as the network parameter of the neural network, so that during the training of the neural network according to the training data in the training set, the first linear combination parameter is trained.

20. The device according to claim 17, wherein, the quantization unit is further configured to: Determine that the modulus values of the first weight matrix and the second weight matrix of the m-th layer are α in the first linear combination parameter m , and perform linear regression on the modulus value to obtain the final value of α m .

21. The device according to any one of claims 16-20, wherein, the obtaining unit is further configured to sequentially calculate the first feature representation of each layer of the first neural network module according to the sequential connection order of the n convolutional layers; The obtaining unit is further configured to obtain a second quantization model, which is used to obtain the second feature representation of the m-th layer of the first neural network module according to the m first feature representations of the first layer to the m-th layer of the first neural network module, wherein the first feature representation of each layer is a feature representation represented by 32-bit floating-point numbers, and the second feature representation of each layer is a feature representation assigned +1 or -1, 1≤m≤n; The quantization unit is further configured to binarize each first feature representation corresponding to each layer of the first neural network module according to the second quantization model to obtain each second feature representation corresponding to each layer of the first neural network module.

22. For the device according to claim 21, the quantization unit is further configured to: By obtain the second feature representation of the m-th layer, wherein, A 1 , A 2 , …, A m are the first feature representations from the first layer to the m-th layer, β 1 , β 2 , …, β m are the second linear combination parameters corresponding to A 1 , A 2 , …, A m , BN(·) is the normalization operation performed on β 1 A 1 + φ 2 A 2 + … + β m A m , Sign(·) is the sign function for binarizing BN(β 1 A 1 + φ 2 A 2 + … + β m A m ), is the second feature representation of the m-th layer, is the second quantization model.

23. The device according to claim 21, wherein, the quantization unit is further configured to: Determine the activation gain of the second feature representation of the m-th layer; Adjust the second feature representation of the m-th layer according to the activation gain of the second feature representation of the m-th layer, so that the difference between the adjusted second feature representation of the m-th layer and the first feature representation of the m-th layer is less than the difference between the second feature representation of the m-th layer and the first feature representation of the m-th layer.

24. The device according to claim 23, wherein, the obtaining unit is specifically further configured to: Calculate the first feature representation of the m-th layer according to the second weight matrix of the m-th layer, the second feature representation of the m-1-th layer, the weight gain of the second weight matrix of the m-th layer, and the activation gain of the second feature representation of the m-1-th layer.

25. The device according to claim 24, wherein, the obtaining unit is specifically further configured to: Perform a convolution operation on the second weight matrix of the m-th layer and the second feature representation of the (m - 1)-th layer to obtain a convolution result; Perform a dot product operation on the convolution result and the weight gain of the second weight matrix of the m-th layer to obtain a dot product result; Perform a dot product operation on the dot product result and the activation gain of the second feature representation of the (m - 1)-th layer to obtain the first feature representation of the m-th layer.

26. The device according to claim 22, wherein, the quantization unit is further configured to: Set the second linear combination parameter as the network parameter of the neural network, so that during the training of the neural network according to the training data in the training set, the second linear combination parameter is trained.

27. The device according to claim 22, wherein, the quantization unit is further configured to: Determine that the modulus value of the first feature representation and the second feature representation of the m-th layer is β in the second linear combination parameter, and perform linear regression on the modulus value to obtain the final value of β. m , and perform linear regression on the modulus value to obtain the final value of β m .

28. The device according to any one of claims 16 - 20, wherein, the neural network further includes the second neural network module and the third neural network module. The second neural network module, the first neural network module, and the third neural network module are connected in sequence. The second neural network module is configured to perform full-precision feature extraction on the input image, and the third neural network module is configured to perform image reconstruction on the output of the first neural network module to obtain an output image.

29. The device according to claim 28, wherein, the input image includes one or more low-resolution images, and the output image includes one high-resolution image.

30. An execution device, wherein, comprising: an acquisition unit configured to acquire an input image; an execution unit configured to process the input image through a trained neural network to obtain an output image, where the trained neural network is a neural network optimized by the method according to any one of claims 1 - 14.

31. A training device includes a processor and a memory, and the processor is coupled to the memory, wherein, the memory is configured to store a program; the processor is configured to execute the program in the memory, so that the training device executes the method according to any one of claims 1 - 14.

32. An execution device includes a processor and a memory, and the processor is coupled to the memory, wherein, the memory is configured to store a program; the processor is configured to execute the program in the memory, so that the execution device executes the method according to claim 15.

33. A computer-readable storage medium includes a program, which when running on a computer, causes the computer to execute the method according to any one of claims 1 - 14, or causes the computer to execute the method according to claim 15.

34. A computer program product containing instructions, which when running on a computer, causes the computer to execute the method according to any one of claims 1 - 14, or causes the computer to execute the method according to claim 15.

35. A chip, the chip comprising a processor and a data interface, the processor reads instructions stored on a memory through the data interface, executes the method according to any one of claims 1-14, or causes a computer to execute the method according to claim 15.

Citation Information

Patent Citations

  • Compression method based on layer-by-layer network binarization

    CN108765506A

  • Balanced binarization neural network quantification method and system

    CN110472725A