Neural network compression method and related equipment thereof
By dynamically adjusting the quantized bit count and resolution of the neural network, the problem of deep neural networks being large in resources and insufficient processing accuracy on terminal devices is solved, and efficient and accurate image processing on terminal devices is achieved.
Patent Information
- Application Number
- CN202510207151.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2021-02-27
- Publication Date
- 2025-07-25
AI Technical Summary
The existing deep neural networks take up a large resource when running on terminal devices, and the model with a fixed number of quantization bits cannot guarantee the accuracy of image processing.
After acquiring the target image, the quantized bit count and resolution of the neural network are dynamically adjusted to adapt to the difficulty of image processing and achieve image processing with different precisions.
It ensures the accuracy of image processing, reduces the resource usage of neural networks, and adapts to the processing needs of different images.
Smart Images

Figure CN120373367A_ABST
Abstract
Description
[0001] This application is a divisional application of the Chinese patent application with the application number 202110221937.1, titled "A Neural Network Compression Method and Related Devices", which was filed with the Chinese Patent Office on February 27, 2021. Technical Field
[0002] This application relates to the field of artificial intelligence technology, and particularly to a neural network compression method and related devices. Background Art
[0003] In recent years, deep neural networks have made great progress in computer vision tasks such as image classification, object detection, and image segmentation. However, deep neural networks often contain a large number of model parameters, which require a large amount of device resources (such as storage space and computational volume), and it is difficult to operate efficiently on terminal devices. Therefore, it is necessary to compress the neural network to reduce the device resources occupied by the neural network.
[0004] Model quantization technology is an effective method for compressing neural networks. This technology can convert the parameters of the neural network from high-bit values (for example, 32 bits) to low-bit values (for example, 4 bits) for representation, thereby significantly reducing the resources occupied by the parameters of the neural network.
[0005] When performing model quantization, the quantization bit number (that is, the bit number of the parameters of the neural network expected by the user) is usually preset and fixed, resulting in the same processing accuracy for any image when the neural network processes images, and it is impossible to guarantee the accuracy of image processing (for example, if the processing difficulty of a certain image is large and the processing accuracy of the neural network for it is low, the processing result will be inaccurate). Summary of the Invention
[0006] The embodiments of this application provide a neural network compression method and related devices, which can enable the compressed neural network to perform image processing with different accuracies on different images, thereby ensuring the accuracy of image processing.
[0007] The first aspect of the embodiments of this application provides a neural network compression method, and the method includes:
[0008] When it is necessary to perform image processing on a target image, the target image to be processed can be obtained first. Further, a first neural network and a second neural network can also be obtained, where the first neural network is used to compress the second neural network, and the second neural network is used to perform image processing on the target image.
[0009] Then, input the target image into the first neural network to obtain the quantization bit number of the second neural network, where the quantization bit number of the second neural network is positively correlated with the amount of computation required for image processing. For example, if the objects in the target image are relatively easy to identify, the amount of computation required for identifying the target image is less, that is, the processing accuracy of the second neural network for the target image can be relatively low, so the quantization bit number of the second neural network is smaller. If the objects in the target image are relatively difficult to identify, the amount of computation required for identifying the target image is more, that is, the processing accuracy of the second neural network for the target image needs to be higher, so the quantization bit number of the second neural network is larger.
[0010] Finally, according to the quantization bit number of the second neural network, perform quantization processing on the parameters of the second neural network to obtain the quantized second neural network, that is, the compressed second neural network. For example, when the quantization bit number of the second neural network is 4 bits, the parameters of the second neural network are represented by 4-bit values, thereby reducing the device resources occupied by the second neural network.
[0011] It can be seen from the above method that after obtaining the target image, the target image can be input into the first neural network to obtain the quantization bit number for compressing the second neural network. This quantization bit number is positively correlated with the amount of computation required for image processing of the target image. If the amount of computation required for image processing of the target image is more, the quantization bit number is larger; if the amount of computation required for image processing of the target image is less, the quantization bit number is smaller. It can be seen that due to the different processing difficulties of the target image, the quantization bit numbers output by the first neural network are different, resulting in different degrees of quantization processing for the second neural network. In this way, for a target image with a relatively small processing difficulty, the quantized second neural network can perform image processing with a relatively small accuracy on it; for a target image with a relatively large processing difficulty, the quantized second neural network can perform image processing with a relatively large accuracy on it, thus ensuring the accuracy of image processing.
[0012] In a possible implementation, the quantization bits of the second neural network include the quantization bits of M layers of the second neural network. According to the quantization bits of the second neural network, the parameters of the second neural network are quantized to obtain the quantized second neural network, which specifically includes: according to the quantization bits of the i-th layer of the second neural network, the parameters of the i-th layer of the network are quantized to obtain the quantized second neural network, where i = 1, 2,..., M, and M is a positive integer. In the foregoing implementation, the first neural network can output the quantization bits of M layers of the second neural network. For any one of the M layers of the second neural network, the parameters of the layer can be quantized according to the quantization bits of the layer, so as to obtain the quantized second neural network. For example, if the quantization bits of the first layer of the second neural network are 4 bits, the parameters of the first layer of the network are represented by a value of 4 bits. If the quantization bits of the second layer of the second neural network are 3 bits, the parameters of the second layer of the network are represented by a value of 3 bits. In this way, the parameters of each layer of the second neural network can be quantized to obtain the quantized second neural network.
[0013] In a possible implementation, inputting the target image into the first neural network to obtain the quantization bits of the second neural network specifically includes: inputting the target image into the first neural network to obtain the probability of the candidate bits of the i-th layer of the second neural network; according to the magnitude of the probability of this part of the candidate bits, the quantization bits of the i-th layer of the network are selected from the candidate bits of the i-th layer of the network. In the foregoing implementation, after the first neural network processes the target image, the probability of the candidate bits of any one of the M layers of the second neural network can be obtained. For any one of the M layers of the network, the quantization bits of the layer can be accurately selected from the candidate bits of the layer based on the magnitude of the probability of the candidate bits of the layer.
[0014] In a possible implementation, the method further includes: after obtaining the quantized second neural network, the target image can be input into the neural network, so that the quantized second neural network performs image processing on the target image (for example, image classification, object detection, image segmentation, etc.) to obtain the features of the target image.
[0015] The second aspect of the embodiments of the present application provides a neural network compression method, the method includes:
[0016] When image processing needs to be performed on the target image, the target image to be processed can be obtained first. Further, a third neural network and a second neural network can also be obtained, where the third neural network is used to compress the input of the second neural network, and the second neural network is used to perform image processing on the target image.
[0017] Then, input the target image into the third neural network to obtain the target resolution corresponding to the second neural network, where the target resolution corresponding to the second neural network is positively correlated with the amount of computation required for image processing. For example, if the objects in the target image are relatively easy to identify, the amount of computation required for identifying the target image is less, that is, the processing accuracy of the second neural network for the target image can be relatively low, so the target resolution corresponding to the second neural network is relatively small. If the objects in the target image are relatively difficult to identify, the amount of computation required for identifying the target image is more, that is, the processing accuracy of the second neural network for the target image needs to be relatively high, so the target resolution corresponding to the second neural network is relatively large.
[0018] Finally, adjust the resolution of the target image to the target resolution to obtain the target image with the target resolution. For example, when the target resolution corresponding to the second neural network is 168×168 and the original resolution of the target image is 224×224, then adjust the resolution of the target image from 224×224 to 168×168 to obtain the target image with the resolution of 168×168.
[0019] In the related art, adjusting the resolution of an image is also one of the methods for compressing a neural network. However, what the user expects is that the resolution of the adjusted image is usually preset and fixed, resulting in the same processing accuracy for any image when the neural network performs image processing, and it is impossible to ensure the accuracy of image processing. It can be seen from the above method that after obtaining the target image, the target image can be input into the third neural network to obtain the target resolution corresponding to the second neural network, which is used to compress the input of the second neural network, that is, the target image. This target resolution is positively correlated with the amount of computation required for image processing of the target image. If the amount of computation required for image processing of the target image is more, then this target resolution is larger; if the amount of computation required for image processing of the target image is less, then this target resolution is smaller. It can be seen that due to the different processing difficulties of the target image, the target resolution output by the third neural network is different, so that the target resolution of the adjusted target image is also different. In this way, for a target image with a relatively small processing difficulty, its resolution can be made smaller, so that the second neural network performs image processing with a relatively low accuracy on it; for a target image with a relatively large processing difficulty, its resolution can be made larger, so that the second neural network performs image processing with a relatively high accuracy on it, thereby ensuring the accuracy of image processing.
[0020] In a possible implementation, inputting the target image into the third neural network to obtain the target resolution corresponding to the second neural network specifically includes: inputting the target image into the third neural network to obtain the probabilities of the candidate resolutions corresponding to the second neural network; and selecting the target resolution corresponding to the second neural network from the candidate resolutions corresponding to the second neural network according to the magnitudes of these probabilities of the candidate resolutions. In the foregoing implementation, after the third neural network performs image processing on the target image, the probabilities of the candidate resolutions corresponding to the second neural network can be obtained. Then, according to the magnitudes of the probabilities of each candidate resolution, the target resolution corresponding to the second neural network is accurately selected from these candidate resolutions.
[0021] In a possible implementation, after adjusting the resolution of the target image to the target resolution to obtain the target image with the target resolution, the method further includes: inputting the target image with the target resolution (i.e., the target image after adjusting the resolution) into the second neural network, so that the second neural network performs image processing (such as image classification, object detection, and image segmentation, etc.) on the target image with the target resolution to obtain the features of the target image with the target resolution.
[0022] The third aspect of the embodiments of the present application provides a model training method, and the method includes: obtaining the image to be trained; inputting the image to be trained into the first model to be trained to obtain the quantization bit number of the second model to be trained; performing quantization processing on the parameters of the second model to be trained according to the quantization bit number of the second model to be trained to obtain the second model to be trained after quantization processing; inputting the image to be trained into the second model to be trained after quantization processing to obtain the features of the image to be trained; and updating the parameters of the first model to be trained and the parameters of the second model to be trained according to the quantization bit number of the second model to be trained and the features of the image to be trained until the model training conditions are satisfied to obtain the first neural network and the second neural network.
[0023] It can be seen from the above method that: by jointly training the first model to be trained and the second model to be trained, the first neural network and the second neural network can be obtained. The first neural network obtained by this method can accurately obtain the quantization bit number of the second neural network based on the input target image, and the second neural network obtained by this method can perform accurate image processing on the target image.
[0024] In a possible implementation, according to the quantization bits of the second model to be trained and the features of the image to be trained, the parameters of the first model to be trained and the parameters of the second model to be trained are updated until the model training conditions are met, and the first neural network and the second neural network are obtained. Specifically, it includes: obtaining the target loss according to the deviation between the quantization bits of the second model to be trained and the preset bits, and the deviation between the features of the image to be trained and the true features of the image to be trained; updating the parameters of the first model to be trained and the parameters of the second model to be trained according to the target loss until the model training conditions are met, and the first neural network and the second neural network are obtained.
[0025] In a possible implementation, the quantization bits of the second model to be trained include the quantization bits of the M layers of the network in the second model to be trained. According to the quantization bits of the second model to be trained, the parameters of the second model to be trained are quantized to obtain the quantized second model to be trained. Specifically, it includes: according to the quantization bits of the i-th layer of the network in the second model to be trained, the parameters of the i-th layer of the network are quantized to obtain the quantized second model to be trained, where i = 1, 2,..., M, and M is a positive integer.
[0026] In a possible implementation, inputting the image to be trained into the first model to be trained to obtain the quantization bits of the second model to be trained specifically includes: inputting the image to be trained into the first model to be trained to obtain the probability of the candidate bits of the i-th layer of the network in the second model to be trained; selecting the quantization bits of the i-th layer of the network from the candidate bits of the i-th layer of the network according to the magnitude of the probability.
[0027] In a possible implementation, the deviation between the quantization bits of the second model to be trained and the preset bits includes the deviation between the quantization bits of the i-th layer of the network in the second model to be trained and the preset bits.
[0028] The fourth aspect of the embodiments of the present application provides a model training method, and the method includes: obtaining the image to be trained; inputting the image to be trained into the third model to be trained to obtain the target resolution corresponding to the second model to be trained; adjusting the resolution of the image to be trained to the target resolution to obtain the image to be trained with the target resolution; inputting the image to be trained with the target resolution into the second model to be trained to obtain the features of the image to be trained with the target resolution; updating the parameters of the third model to be trained and the parameters of the second model to be trained according to the target resolution and the features of the image to be trained with the target resolution until the model training conditions are met, and the third neural network and the second neural network are obtained.
[0029] As can be seen from the above method: by jointly training the third model to be trained and the second model to be trained, a third neural network and a second neural network can be obtained. The third neural network obtained by this method can accurately obtain the target resolution corresponding to the second neural network based on the input target image, and the second neural network obtained by this method can accurately perform image processing on the target image.
[0030] In a possible implementation manner, according to the target resolution and the features of the image to be trained with the target resolution, updating the parameters of the third model to be trained and the parameters of the second model to be trained until the model training conditions are met, and obtaining the third neural network and the second neural network specifically includes: obtaining a target loss according to the deviation between the expected value corresponding to the target resolution and the preset expected value, and the deviation between the features of the image to be trained with the target resolution and the true features of the image to be trained with the target resolution; updating the parameters of the third model to be trained and the parameters of the second model to be trained according to the target loss until the model training conditions are met, and obtaining the third neural network and the second neural network.
[0031] In a possible implementation manner, inputting the image to be trained into the third model to be trained and obtaining the target resolution corresponding to the second model to be trained specifically includes: inputting the image to be trained into the third model to be trained and obtaining the probability of the candidate resolution corresponding to the second model to be trained; selecting the target resolution corresponding to the second model to be trained from the candidate resolutions corresponding to the second model to be trained according to the magnitude of the probability.
[0032] In a possible implementation manner, obtaining a target loss according to the deviation between the expected value corresponding to the target resolution and the preset expected value, and the deviation between the features of the image to be trained with the target resolution and the true features of the image to be trained with the target resolution specifically includes: obtaining a target loss according to the deviation between the expected value of the probability of the target resolution and the preset expected value, the deviation between the expected value of the probability of the other candidate resolutions except the target resolution and the preset expected value, and the deviation between the features of the image to be trained with the target resolution and the true features of the image to be trained with the target resolution.
[0033] The fifth aspect of the embodiments of the present application provides a neural network compression device, and the device includes an acquisition module and a processing module; the acquisition module is used to acquire a target image; the processing module is used to input the target image into a first neural network to obtain the quantization bit number of a second neural network, where the second neural network is used to perform image processing on the target image, and the quantization bit number of the second neural network is positively correlated with the amount of computation required for image processing; the processing module is further used to perform quantization processing on the parameters of the second neural network according to the quantization bit number of the second neural network to obtain the second neural network after quantization processing.
[0034] As can be seen from the above device: after obtaining the target image, the target image can be input into the first neural network to obtain the quantization bits for compressing the second neural network. The quantization bits are positively correlated with the amount of computation required for image processing of the target image. If the amount of computation required for image processing of the target image is large, the quantization bits are large; if the amount of computation required for image processing of the target image is small, the quantization bits are small. It can be seen that due to the different processing difficulties of the target image, the quantization bits output by the first neural network are different, resulting in different degrees of quantization processing of the second neural network. In this way, for a target image with a small processing difficulty, the second neural network after quantization processing can perform image processing with a small accuracy on it; for a target image with a large processing difficulty, the second neural network after quantization processing can perform image processing with a large accuracy on it, thus ensuring the accuracy of image processing.
[0035] In a possible implementation manner, the quantization bits of the second neural network include the quantization bits of M layers of networks in the second neural network. The processing module is specifically configured to perform quantization processing on the parameters of the i-th layer of the network according to the quantization bits of the i-th layer of the network in the second neural network to obtain the second neural network after quantization processing, where i = 1, 2,..., M, and M is a positive integer.
[0036] In a possible implementation manner, the processing module is specifically configured to: input the target image into the first neural network to obtain the probability of the candidate bits of the i-th layer of the network in the second neural network; select the quantization bits of the i-th layer of the network from the candidate bits of the i-th layer of the network according to the magnitude of the probability.
[0037] In a possible implementation manner, the processing module is further configured to input the target image into the second neural network after quantization processing to obtain the features of the target image.
[0038] The sixth aspect of the embodiments of the present application provides a neural network compression device, which includes an acquisition module and a processing module; the acquisition module is configured to acquire a target image; the processing module is configured to input the target image into a third neural network to obtain the target resolution corresponding to the second neural network, where the second neural network is used for image processing of the target image, and the target resolution corresponding to the second neural network is positively correlated with the amount of computation required for image processing; the processing module is further configured to adjust the resolution of the target image to the target resolution to obtain a target image with the target resolution.
[0039] As can be seen from the above device: After obtaining the target image, the target image can be input into the third neural network to obtain the target resolution corresponding to the second neural network, which is used to compress the input of the second neural network, that is, the target image. This target resolution is positively correlated with the amount of computation required for image processing of the target image. If the amount of computation required for image processing of the target image is large, then this target resolution is large; if the amount of computation required for image processing of the target image is small, then this target resolution is small. It can be seen that due to the different processing difficulties of the target images, the target resolutions output by the third neural network are different, resulting in different target resolutions of the adjusted target images. In this way, for a target image with a small processing difficulty, its resolution can be made small, enabling the second neural network to perform image processing with a lower precision on it; for a target image with a large processing difficulty, its resolution can be made large, enabling the second neural network to perform image processing with a higher precision on it, thereby ensuring the accuracy of image processing.
[0040] In a possible implementation, the processing module is specifically configured to: input the target image into the third neural network to obtain the probability of the candidate resolution corresponding to the second neural network; select the target resolution corresponding to the second neural network from the candidate resolutions corresponding to the second neural network according to the magnitude of the probability.
[0041] In a possible implementation, the processing module is further configured to input the target image with the target resolution into the second neural network to obtain the features of the target image with the target resolution.
[0042] A seventh aspect of the embodiments of the present application provides a model training device, which includes an acquisition module and a training module; the acquisition module is used to acquire the image to be trained; the training module is used to input the image to be trained into the first model to be trained to obtain the quantization bit number of the second model to be trained; the training module is further used to perform quantization processing on the parameters of the second model to be trained according to the quantization bit number of the second model to be trained to obtain the second model to be trained after quantization processing; the training module is further used to input the image to be trained into the second model to be trained after quantization processing to obtain the features of the image to be trained; the training module is further used to update the parameters of the first model to be trained and the parameters of the second model to be trained according to the quantization bit number of the second model to be trained and the features of the image to be trained until the model training conditions are met, so as to obtain the first neural network and the second neural network.
[0043] As can be seen from the above device: By jointly training the first model to be trained and the second model to be trained, the first neural network and the second neural network can be obtained. The first neural network obtained by this device can accurately obtain the quantization bit number of the second neural network based on the input target image, and the second neural network obtained by this device can perform accurate image processing on the target image.
[0044] In a possible implementation, the training module is specifically configured to: obtain a target loss according to the deviation between the quantization bit number and the preset bit number of the second model to be trained, and the deviation between the features of the image to be trained and the true features of the image to be trained; update the parameters of the first model to be trained and the parameters of the second model to be trained according to the target loss until the model training condition is satisfied, so as to obtain the first neural network and the second neural network.
[0045] In a possible implementation, the quantization bit number of the second model to be trained includes the quantization bit numbers of M layers of networks in the second model to be trained. The training module is specifically configured to quantize the parameters of the i-th layer of network according to the quantization bit number of the i-th layer of network in the second model to be trained, so as to obtain the second model to be trained after quantization processing, where i = 1, 2,..., M, and M is a positive integer.
[0046] In a possible implementation, the training module is specifically configured to: input the image to be trained into the first model to be trained to obtain the probability of the candidate bit number of the i-th layer of network in the second model to be trained; select the quantization bit number of the i-th layer of network from the candidate bit numbers of the i-th layer of network according to the magnitude of the probability.
[0047] In a possible implementation, the deviation between the quantization bit number and the preset bit number of the second model to be trained includes the deviation between the quantization bit number of the i-th layer of network in the second model to be trained and the preset bit number.
[0048] The eighth aspect of the embodiments of the present application provides a model training device, which includes an acquisition module and a training module; the acquisition module is used to acquire the image to be trained; the training module is used to input the image to be trained into the third model to be trained to obtain the target resolution corresponding to the second model to be trained; the training module is further used to adjust the resolution of the image to be trained to the target resolution to obtain the image to be trained with the target resolution; the training module is further used to input the image to be trained with the target resolution into the second model to be trained to obtain the features of the image to be trained with the target resolution; the training module is further used to update the parameters of the third model to be trained and the parameters of the second model to be trained according to the target resolution and the features of the image to be trained with the target resolution until the model training condition is satisfied, so as to obtain the third neural network and the second neural network.
[0049] It can be seen from the above device that: by jointly training the third model to be trained and the second model to be trained, the third neural network and the second neural network can be obtained. The third neural network obtained by this device can accurately obtain the target resolution corresponding to the second neural network based on the input target image, and the second neural network obtained by this device can perform accurate image processing on the target image.
[0050] In a possible implementation, the training module is specifically configured to: obtain a target loss according to the deviation between the expected value corresponding to the target resolution and the preset expected value, and the deviation between the features of the training image of the target resolution and the true features of the training image of the target resolution; update the parameters of the third training model and the parameters of the second training model according to the target loss until the model training condition is satisfied, and obtain the third neural network and the second neural network.
[0051] In a possible implementation, the training module is specifically configured to: input the training image into the third training model to obtain the probability of the candidate resolution corresponding to the second training model; select the target resolution corresponding to the second training model from the candidate resolutions corresponding to the second training model according to the magnitude of the probability.
[0052] In a possible implementation, the training module is specifically configured to obtain a target loss according to the deviation between the expected value of the probability of the target resolution and the preset expected value, the deviation between the expected value of the probability of the other candidate resolutions except the target resolution and the preset expected value, and the deviation between the features of the training image of the target resolution and the true features of the training image of the target resolution.
[0053] A ninth aspect of the embodiments of the present application provides an image processing method, the method including: obtaining a target image; inputting the target image into a first neural network to obtain the quantization bits of a second neural network, where the second neural network is used to perform image processing on the target image, and the quantization bits of the second neural network are positively correlated with the amount of computation required for image processing; performing quantization processing on the parameters of the second neural network according to the quantization bits of the second neural network to obtain the second neural network after quantization processing; inputting the target image into the second neural network after quantization processing to obtain the features of the target image.
[0054] A tenth aspect of the embodiments of the present application provides an image processing method, the method including: obtaining a target image; inputting the target image into a third neural network to obtain the target resolution corresponding to the second neural network, where the second neural network is used to perform image processing on the target image, and the target resolution corresponding to the second neural network is positively correlated with the amount of computation required for image processing; adjusting the resolution of the target image to the target resolution to obtain the target image of the target resolution; inputting the target image of the target resolution into the second neural network to obtain the features of the target image of the target resolution.
[0055] The eleventh aspect of the embodiments of the present application provides a neural network compression device, which includes a memory and a processor; the memory stores code, and the processor is configured to execute the code. When the code is executed, the neural network compression device executes the method described in the first aspect, any possible implementation manner in the first aspect, the second aspect, or any possible implementation manner in the second aspect.
[0056] The twelfth aspect of the embodiments of the present application provides a model training device, which includes a memory and a processor; the memory stores code, and the processor is configured to execute the code. When the code is executed, the model training device executes the method described in the third aspect, any possible implementation manner in the third aspect, the fourth aspect, or any possible implementation manner in the fourth aspect.
[0057] The thirteenth aspect of the embodiments of the present application provides a circuit system, which includes a processing circuit configured to execute the method described in the first aspect, any possible implementation manner in the first aspect, the second aspect, any possible implementation manner in the second aspect, the third aspect, any possible implementation manner in the third aspect, the fourth aspect, or any possible implementation manner in the fourth aspect.
[0058] The fourteenth aspect of the embodiments of the present application provides a chip system, which includes a processor for calling a computer program or computer instruction stored in a memory, so that the processor executes the method described in the first aspect, any possible implementation manner in the first aspect, the second aspect, any possible implementation manner in the second aspect, the third aspect, any possible implementation manner in the third aspect, the fourth aspect, or any possible implementation manner in the fourth aspect.
[0059] In a possible implementation manner, the processor is coupled to the memory through an interface.
[0060] In a possible implementation manner, the chip system further includes a memory, and a computer program or computer instruction is stored in the memory.
[0061] The fifteenth aspect of the embodiments of the present application provides a computer storage medium, which stores a computer program. When the program is executed by a computer, the computer implements the method described in the first aspect, any possible implementation manner in the first aspect, the second aspect, any possible implementation manner in the second aspect, the third aspect, any possible implementation manner in the third aspect, the fourth aspect, or any possible implementation manner in the fourth aspect.
[0062] A sixteenth aspect of the embodiments of the present application provides a computer program product. The computer program product stores instructions that, when executed by a computer, cause the computer to implement the methods described in any one of the possible implementations of the first aspect, the second aspect, any one of the possible implementations of the second aspect, the third aspect, any one of the possible implementations of the third aspect, the fourth aspect, or any one of the possible implementations of the fourth aspect.
[0063] In the embodiments of the present application, after obtaining the target image, the target image can be input into the first neural network to obtain the quantization bits for compressing the second neural network. The quantization bits are positively correlated with the amount of computation required for image processing of the target image. If the amount of computation required for image processing of the target image is large, the quantization bits are large; if the amount of computation required for image processing of the target image is small, the quantization bits are small. It can be seen that due to the different processing difficulties of the target image, the quantization bits output by the first neural network are different, resulting in different degrees of quantization processing of the second neural network. In this way, for a target image with a small processing difficulty, the second neural network after quantization processing can perform image processing with a small accuracy on it; for a target image with a large processing difficulty, the second neural network after quantization processing can perform image processing with a large accuracy on it, thereby ensuring the accuracy of image processing. Description of the Drawings
[0064] Figure 1 It is a schematic structural diagram of a framework of an AI entity;
[0065] Figure 2a It is a schematic structural diagram of an image processing system provided by an embodiment of the present application;
[0066] Figure 2b It is another schematic structural diagram of an image processing system provided by an embodiment of the present application;
[0067] Figure 2c It is a schematic diagram of related equipment for image processing provided by an embodiment of the present application;
[0068] Figure 3 It is a schematic diagram of the architecture of system 100 provided by an embodiment of the present application;
[0069] Figure 4 It is a schematic flowchart of a neural network compression method provided by an embodiment of the present application;
[0070] Figure 5 It is a schematic diagram of an application example of a neural network compression method provided by an embodiment of the present application;
[0071] Figure 6Another schematic flowchart of the neural network compression method provided by the embodiments of the present application;
[0072] Figure 7 Another schematic diagram of an application example of the neural network compression method provided by the embodiments of the present application;
[0073] Figure 8 A schematic flowchart of the model training method provided by the embodiments of the present application;
[0074] Figure 9 Another schematic flowchart of the model training method provided by the embodiments of the present application;
[0075] Figure 10 A schematic structural diagram of the neural network compression device provided by the embodiments of the present application;
[0076] Figure 11 Another schematic structural diagram of the neural network compression device provided by the embodiments of the present application;
[0077] Figure 12 A schematic structural diagram of the model training device provided by the embodiments of the present application;
[0078] Figure 13 Another schematic structural diagram of the model training device provided by the embodiments of the present application;
[0079] Figure 14 A schematic structural diagram of the execution device provided by the embodiments of the present application;
[0080] Figure 15 A schematic structural diagram of the training device provided by the embodiments of the present application;
[0081] Figure 16 A schematic structural diagram of the chip provided by the embodiments of the present application. Detailed implementation manners
[0082] The embodiments of the present application provide a neural network compression method and related devices, which can enable the compressed neural network to perform image processing with different precisions on different images, thereby ensuring the accuracy of image processing.
[0083] In the description of the present application, the specification, the claims, and the above-mentioned drawings, the terms "first", "second", etc. are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that such terms can be interchanged under appropriate circumstances, which is only a way of distinguishing objects with the same attributes when describing embodiments of the present application. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion, so that a process, method, system, product, or device comprising a series of units does not have to be limited to those units, but may include other units not clearly listed or inherent to these processes, methods, products, or devices.
[0084] Artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a branch of computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. Artificial intelligence also studies the design principles and implementation methods of various intelligent machines, enabling the machines to have the functions of perception, reasoning, and decision-making.
[0085] First, the overall workflow of the artificial intelligence system will be described. Please refer to Figure 1 , Figure 1 FIG. is a schematic structural diagram of a framework of an artificial intelligence entity. The above-mentioned artificial intelligence subject framework will be elaborated from two dimensions: the "intelligent information chain" (horizontal axis) and the "IT value chain" (vertical axis). Among them, the "intelligent information chain" reflects a series of processes from data acquisition to processing. For example, it can be the general processes of intelligent information perception, intelligent information representation and formation, intelligent reasoning, intelligent decision-making, intelligent execution, and output. In this process, the data undergoes the process of "data - information - knowledge - wisdom" refinement. The "IT value chain" reflects the value brought by artificial intelligence to the information technology industry from the underlying infrastructure of artificial intelligence, information (providing and processing technology implementation) to the industrial ecological process of the system.
[0086] (1) Infrastructure
[0087] The infrastructure provides computing power support for the artificial intelligence system, enables communication with the external world, and is supported by the basic platform. It communicates with the external world through sensors; the computing power is provided by intelligent chips (such as hardware acceleration chips like CPU, NPU, GPU, ASIC, FPGA, etc.); the basic platform includes relevant platform guarantees and supports such as distributed computing frameworks and networks, and may include cloud storage and computing, interconnected networks, etc. For example, sensors communicate with the external world to obtain data, and these data are provided to the intelligent chips in the distributed computing system provided by the basic platform for computing.
[0088] (2) Data
[0089] The data at the upper layer of the infrastructure is used to represent the data sources in the field of artificial intelligence. The data involves graphics, images, voices, texts, and also involves the Internet of Things data of traditional devices, including the business data of existing systems and the sensed data such as force, displacement, liquid level, temperature, humidity, etc.
[0090] (3) Data Processing
[0091] Data processing usually includes data training, machine learning, deep learning, search, inference, decision-making, etc.
[0092] Among them, machine learning and deep learning can perform symbolic and formal intelligent information modeling, extraction, preprocessing, training, etc. on the data.
[0093] Inference refers to the process of simulating the intelligent reasoning method of humans in a computer or intelligent system, and using formalized information for machine thinking and problem-solving according to the inference control strategy. The typical function is search and matching.
[0094] Decision-making refers to the process of making decisions after the intelligent information is inferred, and usually provides functions such as classification, sorting, prediction, etc.
[0095] (4) General Capabilities
[0096] After the data is processed through the above-mentioned data processing, some general capabilities can be formed based on the results of the data processing. For example, it can be an algorithm or a general system. For example, translation, text analysis, computer vision processing, speech recognition, image recognition, etc.
[0097] (5) Intelligent Products and Industry Applications
[0098] Intelligent products and industry applications refer to the products and applications of the artificial intelligence system in various fields. It is the encapsulation of the overall artificial intelligence solution, productizes the intelligent information decision-making, and realizes the landing application. Its application fields mainly include: intelligent terminals, intelligent transportation, intelligent healthcare, autonomous driving, smart cities, etc.
[0099] Next, several application scenarios of this application will be introduced.
[0100] Figure 2a FIG. 4 is a schematic structural diagram of an image processing system provided by an embodiment of this application. The image processing system includes a user device and a data processing device. Among them, the user device includes intelligent terminals such as mobile phones, personal computers, or information processing centers. The user device is the initiating end of image processing and, as the initiator of an image processing request, usually a user initiates a request through the user device.
[0101] The above data processing device can be a device or server with data processing functions such as a cloud server, a network server, an application server, and a management server. The data processing device receives an image enhancement request from the intelligent terminal through an interaction interface, and then performs image processing in ways such as machine learning, deep learning, searching, reasoning, and decision-making through a memory for storing data and a processor for data processing. The memory in the data processing device can be a general term, including local storage and a database for storing historical data. The database can be on the data processing device or on other network servers.
[0102] In Figure 2a In the image processing system shown in FIG. 4, the user device can receive a user's instruction. For example, the user device can obtain an image input / selected by the user, and then send a request to the data processing device, so that the data processing device executes an image semantic segmentation application for the image obtained by the user device, thereby obtaining a corresponding processing result for the image. Exemplarily, the user device can obtain an image to be processed input by the user, and then send an image processing request to the data processing device, so that the data processing device performs an image processing application (such as image classification, object detection, and image segmentation, etc.) on the image, thereby obtaining a processed image.
[0103] In Figure 2a FIG. 5, the data processing device can execute the neural network compression method and the model training method of the embodiment of this application.
[0104] Figure 2b FIG. 6 is another schematic structural diagram of the image processing system provided by the embodiment of this application. In Figure 2b FIG. 6, the user device directly serves as the data processing device. The user device can directly obtain an input from the user and directly process it by the hardware of the user device itself. The specific process is similar to Figure 2a FIG. 4 and can refer to the above description and will not be elaborated here.
[0105] In Figure 2bIn the image processing system shown, the user device can receive a user's instruction. For example, the user device can obtain a to-be-processed image selected by the user in the user device, and then the user device itself executes an image processing application (such as image classification, object detection, and image segmentation, etc.) on the image, so as to obtain a corresponding processing result for the image.
[0106] In Figure 2b , the user device itself can execute the neural network compression method and the model training method of the embodiments of the present application.
[0107] Figure 2c It is a schematic diagram of related devices for image processing provided by the embodiments of the present application.
[0108] The above Figure 2a and Figure 2b The user device in Figure 2c can specifically be the local device 301 or the local device 302 in Figure 2a The data processing device in Figure 2c can specifically be the execution device 210 in
[0109] Figure 2a and Figure 2b The processor in
[0110] can perform data training / machine learning / deep learning through a neural network model or other models (such as a model based on a support vector machine), and use the model finally trained or learned from the data to execute an image processing application on the image, so as to obtain a corresponding processing result.
[0111] Figure 3 It is a schematic diagram of the architecture of system 100 provided by the embodiments of the present application. In Figure 3In this case, the execution device 110 configures an input / output (I / O) interface 112 for data interaction with external devices. A user can input data to the I / O interface 112 through the client device 140. The input data in the embodiments of this application may include: various tasks to be scheduled, invocable resources, and other parameters.
[0112] During the preprocessing of the input data by the execution device 110, or during the related processing such as calculation by the calculation module 111 of the execution device 110 (such as implementing the functions of the neural network in this application), the execution device 110 can call data, code, etc. in the data storage system 150 for corresponding processing, and can also store the data, instructions, etc. obtained from the corresponding processing into the data storage system 150.
[0113] Finally, the I / O interface 112 returns the processing result to the client device 140 and provides it to the user.
[0114] It should be noted that the training device 120 can generate corresponding target models / rules based on different training data for different targets or tasks. The corresponding target models / rules can be used to achieve the above targets or complete the above tasks, so as to provide the required results for the user. Among them, the training data can be stored in the database 130 and comes from the training samples collected by the data acquisition device 160.
[0115] In Figure 3 the case shown, the user can manually give input data, and this manual giving can be operated through the interface provided by the I / O interface 112. In another case, the client device 140 can automatically send input data to the I / O interface 112. If the client device 140 is required to automatically send input data and user authorization is required, the user can set the corresponding permissions in the client device 140. The user can view the results output by the execution device 110 on the client device 140, and the specific presentation forms can be display, sound, action, etc. The client device 140 can also be used as a data acquisition end to collect the input data input to the I / O interface 112 and the output result output from the I / O interface 112 as new sample data and store them in the database 130. Of course, it can also be collected without going through the client device 140, but the input data input to the I / O interface 112 and the output result output from the I / O interface 112 as shown are directly stored in the database 130 as new sample data.
[0116] It should be noted that Figure 3 is only a schematic diagram of a system architecture provided by the embodiments of this application. The positional relationships between the devices, components, modules, etc. shown in the figure do not constitute any limitation. For example, inFigure 3 In this case, the data storage system 150 is an external memory relative to the execution device 110. In other cases, the data storage system 150 can also be placed in the execution device 110. As Figure 3 shown, a neural network can be trained according to the training device 120.
[0117] An embodiment of the present application further provides a chip, which includes a neural network processor NPU. The chip can be set in the execution device 110 as Figure 3 shown, to complete the calculation work of the calculation module 111. The chip can also be set in the training device 120 as Figure 3 shown, to complete the training work of the training device 120 and output a target model / rule.
[0118] The neural network processor NPU is mounted on the main central processing unit (CPU) (host CPU) as a coprocessor, and tasks are allocated by the main CPU. The core part of the NPU is the arithmetic circuit, and the controller controls the arithmetic circuit to extract data from the memory (weight memory or input memory) and perform operations.
[0119] In some implementations, the arithmetic circuit includes multiple processing units (PEs) inside. In some implementations, the arithmetic circuit is a two-dimensional systolic array. The arithmetic circuit can also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit is a general matrix processor.
[0120] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit fetches the corresponding data of matrix B from the weight memory and caches it on each PE in the arithmetic circuit. The arithmetic circuit fetches the data of matrix A from the input memory and performs matrix operations with matrix B, and the partial results or final results of the obtained matrix are stored in the accumulator.
[0121] The vector calculation unit can further process the output of the arithmetic circuit, such as vector multiplication, vector addition, exponential operation, logarithmic operation, magnitude comparison, etc. For example, the vector calculation unit can be used for network calculations in non-convolutional / non-FC layers of a neural network, such as pooling, batch normalization, local response normalization, etc.
[0122] In some implementations, the vector computing unit can store the processed output vectors in the unified buffer. For example, the vector computing unit can apply a non-linear function to the output of the arithmetic circuit, such as a vector of accumulated values, to generate activation values. In some implementations, the vector computing unit generates normalized values, combined values, or both. In some implementations, the processed output vectors can be used as activation inputs to the arithmetic circuit, such as for use in subsequent layers in a neural network.
[0123] The unified memory is used to store input data and output data.
[0124] The weight data directly transfers the input data in the external memory to the input memory and / or the unified memory, stores the weight data in the external memory into the weight memory, and stores the data in the unified memory into the external memory through the direct memory access controller (DMAC).
[0125] The bus interface unit (BIU) is used to implement the interaction between the main CPU, DMAC, and the instruction fetch memory through the bus.
[0126] The instruction fetch buffer connected to the controller is used to store the instructions used by the controller;
[0127] The controller is used to call the instructions cached in the instruction fetch memory to control the working process of the arithmetic accelerator.
[0128] Generally, the unified memory, the input memory, the weight memory, and the instruction fetch memory are all on-chip memories, and the external memory is the memory outside the NPU. The external memory can be a double data rate synchronous dynamic random access memory (DDRSDRAM), a high bandwidth memory (HBM), or other readable and writable memories.
[0129] Since the embodiments of the present application involve a large number of neural network applications, for ease of understanding, the relevant terms and related concepts such as neural networks involved in the embodiments of the present application will be introduced below.
[0130] (1) Neural Network
[0131] A neural network can be composed of neural units. A neural unit can refer to an arithmetic unit with xs and intercept 1 as inputs, and the output of the arithmetic unit can be:
[0132]
[0133] Among them, s = 1, 2, …… n, where n is a natural number greater than 1, Ws is the weight of xs, and b is the bias of the neuron. f is the activation function of the neuron, which is used to introduce non-linear characteristics into the neural network to convert the input signal in the neuron into an output signal. The output signal of this activation function can be used as the input of the next convolutional layer. The activation function can be the sigmoid function. A neural network is a network formed by connecting many such single neurons together, that is, the output of one neuron can be the input of another neuron. The input of each neuron can be connected to the local receptive field of the previous layer to extract the predicted label of the local receptive field, and the local receptive field can be a region composed of several neurons.
[0134] The operation of each layer in the neural network can be described by the mathematical expression y = a(Wx + b): From a physical perspective, the operation of each layer in the neural network can be understood as completing the transformation from the input space (the set of input vectors) to the output space (i.e., from the row space to the column space of the matrix) through five operations on the input space. These five operations include: 1. Dimension increase / dimension reduction; 2. Magnification / minification; 3. Rotation; 4. Translation; 5. "Bending". Among them, operations 1, 2, and 3 are completed by Wx, operation 4 is completed by +b, and operation 5 is implemented by a(). The reason for using the word "space" here is that the object to be classified is not a single thing, but a class of things, and space refers to the set of all individuals of this class of things. Among them, W is the weight vector, and each value in this vector represents the weight value of a neuron in this layer of the neural network. This vector W determines the space transformation from the input space to the output space described above, that is, the weight W of each layer controls how to transform the space. The purpose of training the neural network, that is, ultimately obtaining the weight matrix of all layers of the trained neural network (the weight matrix formed by vectors W of many layers). Therefore, the training process of the neural network is essentially to learn the way to control the space transformation, and more specifically, to learn the weight matrix.
[0135] Since we hope that the output of the neural network is as close as possible to the value we really want to predict, we can compare the predicted value of the current network with the real target value, and then update the weight vector of each layer of the neural network according to the difference between the two. (Of course, there is usually an initialization process before the first update, that is, pre-configuring parameters for each layer in the neural network). For example, if the predicted value of the network is too high, adjust the weight vector to make it predict lower, and keep adjusting until the neural network can predict the real target value. Therefore, it is necessary to pre-define "how to compare the difference between the predicted value and the target value", which is the loss function or objective function. They are important equations for measuring the difference between the predicted value and the target value. Among them, taking the loss function as an example, the higher the output value (loss) of the loss function, the greater the difference. Then the training of the neural network becomes a process of minimizing this loss as much as possible.
[0136] (2) Backpropagation algorithm
[0137] The neural network can use the backpropagation (BP) algorithm to correct the size of the parameters in the initial neural network model during the training process, so that the reconstruction error loss of the neural network model becomes smaller and smaller. Specifically, the forward propagation of the input signal until the output will generate an error loss, and the initial neural network model parameters are updated by backpropagating the error loss information, so that the error loss converges. The backpropagation algorithm is a backpropagation movement dominated by the error loss, aiming to obtain the optimal parameters of the neural network model, such as the weight matrix.
[0138] The following describes the method provided in this application from the training side of the neural network and the application side of the neural network.
[0139] The model training method provided in the embodiments of this application is related to image processing, and can be specifically applied to data processing methods such as data training, machine learning, and deep learning. It performs symbolic and formal intelligent information modeling, extraction, preprocessing, training, etc. on the training data (such as the to-be-trained images in this application), and finally obtains a trained neural network; moreover, the neural network compression method provided in the embodiments of this application can use the above-mentioned trained neural network, input the input data (such as the target image in this application) into the trained neural network, and obtain the output data (such as the quantization bit number, features of the target image in this application). It should be noted that the model training method and the neural network compression method provided in the embodiments of this application are inventions generated based on the same concept, and can also be understood as two parts of a system, or two stages of an overall process: such as the model training stage and the model application stage.
[0140] Figure 4 This is a schematic flowchart of the neural network compression method provided by the embodiments of this application. As Figure 4 shown, this method includes:
[0141] 401. Obtain a target image.
[0142] When image processing needs to be performed on the target image, the target image to be processed can be obtained first. Further, a first neural network and a second neural network can also be obtained. Among them, the first neural network is used to compress the second neural network, that is, the first neural network is used to perform quantization processing on the parameters of the second neural network, so as to reduce the device resources occupied by the parameters of the second neural network. The second neural network is used to perform image processing on the target image. For example, the second neural network can perform image classification on the target image to determine the category to which the object in the target image belongs. Another example is that the second neural network can perform image segmentation on the target image to divide the objects of different categories in the target image, and so on.
[0143] It should be understood that the first neural network can be any one of models such as a multi-layer perceptron (MLP), a convolutional neural network (CNN), a recursive neural network, and a recurrent neural network (RNN). The second neural network can also be any one of models such as MLP, CNN, recursive neural network, and RNN, and no limitation is made here.
[0144] It should also be understood that the first neural network and the second neural network in the embodiments of this application are both trained neural network models, and the training processes of the first neural network and the second neural network will not be introduced in detail here.
[0145] 402. Input the target image into the first neural network to obtain the quantization bit number of the second neural network. Among them, the second neural network is used to perform image processing on the target image, and the quantization bit number of the second neural network is positively correlated with the amount of computation required for image processing.
[0146] For target images containing different contents, the difficulty of image processing is also different. For example, taking the image processing of the target image to recognize the object in the target image as an example for illustration. Suppose there are image A and image B. The object in image A is a dog, and the object in image B is a dragonfly parked on a flower. Since the dog in image A is easy to recognize and the dragonfly in image B is not easy to recognize, the recognition difficulty of image A is lower than that of image A.
[0147] It can be seen that if the image processing difficulty of the target image is small, even if the amount of computation required for the second neural network to process the target image is small (i.e., the accuracy of image processing is low), the accuracy of image processing can be ensured. If the image processing difficulty of the target image is large, the amount of computation required for the second neural network to process the target image needs to be large (i.e., the accuracy of image processing is high), so as to ensure the accuracy of image processing. Based on this, when quantifying the parameters of the second neural network, it needs to change with the change of the image processing difficulty of the target image. If the parameters of the second neural network are represented and stored as high-bit values, the image processing accuracy that the second neural network can achieve is high. If the parameters of the second neural network are represented and stored as low-bit values, the image processing accuracy that the second neural network can achieve is low.
[0148] Specifically, after obtaining the target image, the target image can be input into the first neural network, so that the first neural network performs image processing on the target image to obtain the quantization bit number of the second neural network. Among them, the quantization bit number of the second neural network is associated with the amount of computation required for the second neural network to perform image processing on the target image (i.e., the difficulty of image processing of the second neural network on the target image). Generally, if the image processing difficulty of the target image is small, the quantization bit number of the second neural network is small. If the image processing difficulty of the target image is large, the quantization bit number of the second neural network is large. Still as the above example, if image A is input into the first neural network, the quantization bit number of the second neural network output by the first neural network is 2 bits. If image B is input into the second neural network, the quantization bit number of the second neural network output by the first neural network is 4 bits.
[0149] Furthermore, the quantization bit number of the second neural network can include the quantization bit numbers of M layers of networks in the second neural network (for example, M convolutional layers of the second neural network), and M is a positive integer. For the second neural network, the quantization bit numbers of these M layers of networks can be exactly the same or partially the same. Then, the first neural network can obtain the quantization bit numbers of the M layers of networks in the second neural network in the following way:
[0150] Input the target image into the first neural network to obtain the probability of the candidate bit number of the i-th layer of network in the second neural network, where i = 1, 2,..., M. Then, according to the magnitude of the probability of the candidate bit number of the i-th layer of network, select the quantization bit number of the i-th layer of network from the candidate bit numbers of the i-th layer of network. Therefore, the quantization bit numbers of each layer of network in the M layers of networks of the second neural network can be obtained.
[0151] To further understand the foregoing process, the following is combined with Figure 5 for further explanation. Figure 5A schematic diagram of an application example of the neural network compression method provided by an embodiment of this application is shown as follows Figure 5 Assume that there are image A and image B (the image processing difficulty of image A is lower than that of image B), and the candidate bit numbers of each convolutional layer of the second neural network are 2 bits, 3 bits, 4 bits, and 5 bits.
[0152] After inputting image A into the first neural network, the probabilities of the candidate bit numbers of each convolutional layer of the second neural network can be obtained. Among them, in the probability of the candidate bit number of the first convolutional layer, the probability of 2 bits is the largest; in the probability of the candidate bit number of the second convolutional layer, the probability of 3 bits is the largest; in the probability of the candidate bit number of the third convolutional layer, the probability of 2 bits is the largest, and so on. Then, it can be determined that the quantization bit number of the first convolutional layer of the second neural network is 2 bits, the quantization bit number of the second convolutional layer is 3 bits, the quantization bit number of the third convolutional layer is 2 bits,..., and the quantization bit number of the Mth convolutional layer is 2 bits.
[0153] After inputting image B into the first neural network, the probabilities of the candidate bit numbers of each convolutional layer of the second neural network can be obtained. Among them, in the probability of the candidate bit number of the first convolutional layer, the probability of 5 bits is the largest; in the probability of the candidate bit number of the second convolutional layer, the probability of 4 bits is the largest; in the probability of the candidate bit number of the third convolutional layer, the probability of 4 bits is the largest, and so on. Then, it can be determined that the quantization bit number of the first convolutional layer of the second neural network is 5 bits, the quantization bit number of the second convolutional layer is 4 bits, the quantization bit number of the third convolutional layer is 4 bits,..., and the quantization bit number of the Mth convolutional layer is 5 bits.
[0154] 403. According to the quantization bit number of the second neural network, perform quantization processing on the parameters of the second neural network to obtain the quantized second neural network.
[0155] After obtaining the quantization bit number of the second neural network, according to the quantization bit number of the second neural network, perform quantization processing on the parameters of the second neural network to obtain the quantized second neural network. Specifically, for the M-layer network of the second neural network, according to the quantization bit number of the ith layer network in the second neural network, perform quantization processing on the parameters of the ith layer network to obtain the quantized second neural network, where i = 1, 2,..., M, and M is a positive integer.
[0156] Still as Figure 5An example is described below. For image A, after determining that the quantization bit number of the first convolutional layer of the second neural network is 2 bits, the quantization bit number of the second convolutional layer is 3 bits, the quantization bit number of the third convolutional layer is 2 bits, …, and the quantization bit number of the Mth convolutional layer is 2 bits, the parameters of the first convolutional layer of the second neural network are represented by 2-bit values, the parameters of the second convolutional layer are represented by 3-bit values, the parameters of the third convolutional layer are represented by 2-bit values, …, and the parameters of the Mth convolutional layer are represented by 2-bit values, thereby completing the quantization process of the second neural network and obtaining the quantized second neural network.
[0157] For image B, after determining that the quantization bit number of the first convolutional layer of the second neural network is 5 bits, the quantization bit number of the second convolutional layer is 4 bits, the quantization bit number of the third convolutional layer is 4 bits, …, and the quantization bit number of the Mth convolutional layer is 5 bits, the parameters of the first convolutional layer of the second neural network are represented by 5-bit values, the parameters of the second convolutional layer are represented by 4-bit values, the parameters of the third convolutional layer are represented by 4-bit values, …, and the parameters of the Mth convolutional layer are represented by 5-bit values, thereby completing the quantization process of the second neural network and obtaining the quantized second neural network.
[0158] In addition, in this embodiment, the parameters of each layer of the second neural network usually refer to the weights of that layer. While quantizing the weights of each layer of the network, corresponding quantization processing can also be performed on the input of each layer. For example, while representing the weights of the first convolutional layer by 2-bit values, the input of the first convolutional layer can also be represented by 2-bit values.
[0159] 404. Input the target image into the quantized second neural network to obtain the features of the target image.
[0160] After obtaining the quantized second neural network, the target image can be input into the quantized second neural network so that the quantized second neural network performs image processing (such as image classification, object detection, and image segmentation, etc.) on the target image, thereby obtaining the features of the target image.
[0161] In this embodiment, after obtaining the target image, the target image can be input into the first neural network to obtain the quantization bit number for compressing the second neural network. This quantization bit number is positively correlated with the amount of computation required for image processing of the target image. If the amount of computation required for image processing of the target image is large, then this quantization bit number is large; if the amount of computation required for image processing of the target image is small, then this quantization bit number is small. It can be seen that due to the different processing difficulties of the target images, the quantization bit numbers output by the first neural network are different, resulting in different degrees of quantization processing for the second neural network. In this way, for target images with relatively low processing difficulty, the second neural network after quantization processing can perform image processing with relatively low precision on them; for target images with relatively high processing difficulty, the second neural network after quantization processing can perform image processing with relatively high precision on them, thereby ensuring the accuracy of image processing.
[0162] In addition, the neural network in the embodiment of the present application can be compared with the neural networks obtained by other compression methods (such as the neural network compression method in the background art) in terms of performance. The comparison results are shown in Table 1 (corresponding to the first image dataset):
[0163] Table 1
[0164]
[0165] Other compression methods (including Method 1 and Method 2) adopt the method of static fixed bit numbers (that is, fixing the bit numbers of the parameters of the neural network), while the present application adopts the method of dynamically adjusting the bit numbers (that is, automatically adjusting the bit numbers of the parameters of the neural network according to different images). Based on Table 1, it can be known that the compression method provided by the embodiment of the present application can ensure that when the neural network processes the images in the first type of image dataset, higher accuracy can be achieved with relatively low computational cost.
[0166] Furthermore, the neural network in the embodiment of the present application can be further compared with the neural networks obtained by other compression methods (such as the neural network compression method in the background art) in terms of performance. The comparison results are shown in Table 2 (corresponding to the second image dataset):
[0167] Table 2
[0168]
[0169]
[0170] Based on Table 2, it can be known that the compression method provided by the embodiment of the present application can ensure that when the neural network processes the images in the second type of image dataset, higher accuracy can also be achieved with relatively low computational cost.
[0171] Figure 6 Another process schematic diagram of the neural network compression method provided by the embodiment of the present application. As Figure 6 shown, the method includes:
[0172] 601. Obtain a target image.
[0173] When image processing needs to be performed on the target image, the target image to be processed can be obtained first. Further, a third neural network and a second neural network can also be obtained. Among them, the third neural network is used to compress the input of the second neural network, that is, the third neural network is used to adjust the resolution of the target image, so as to reduce the device resources occupied when the second neural network processes the image target. The second neural network is used to perform image processing on the target image. For example, the second neural network can perform image classification on the target image to determine the category to which the object in the target image belongs. Another example is that the second neural network can perform image segmentation on the target image to divide different categories of objects in the target image, and so on.
[0174] It should be understood that the third neural network can be any one of models such as MLP, CNN, recurrent neural network, RNN, etc., and the second neural network can also be any one of models such as MLP, CNN, recurrent neural network, RNN, etc., and no limitation is made here.
[0175] It should also be understood that the third neural network and the second neural network in the embodiment of the present application are both trained neural network models, and the training processes of the third neural network and the second neural network will not be introduced in detail here.
[0176] 602. Input the target image into the third neural network to obtain the target resolution corresponding to the second neural network, where the second neural network is used to perform image processing on the target image, and the target resolution corresponding to the second neural network is positively correlated with the amount of computation required for image processing.
[0177] For target images containing different contents, the difficulty of image processing is also different. For example, taking the image processing of the target image to recognize the object in the target image as an example for illustration. Suppose there are image A and image B. The object in image A is a dog, and the object in image B is a dragonfly parked on a flower. Since the dog in image A is easy to recognize and the dragonfly in image B is not easy to recognize, the recognition difficulty of image A is lower than that of image A.
[0178] It can be seen that if the image processing difficulty of the target image is small, even if the amount of computation required for the second neural network to process the target image is small (i.e., the accuracy of image processing is low), the accuracy of image processing can be ensured. If the image processing difficulty of the target image is large, the amount of computation required for the second neural network to process the target image is large (i.e., the accuracy of image processing is high), so as to ensure the accuracy of image processing. Based on this, when the target image is input into the second neural network, the resolution of the target image can be changed according to the change of the image processing difficulty of the target image. If the resolution of the input target image is small, the amount of computation required for the second neural network to implement image processing is small. If the resolution of the input target image is large, the amount of computation required for the second neural network to implement image processing is large.
[0179] Specifically, after obtaining the target image, the target image is input into the third neural network to obtain the target resolution corresponding to the second neural network. Among them, the target resolution corresponding to the second neural network is associated with the amount of computation required for the second neural network to perform image processing on the target image (i.e., the difficulty of the second neural network to perform image processing on the target image). Generally, if the image processing difficulty of the target image is small, the target resolution corresponding to the second neural network (the target resolution of the target image input into the second neural network) is small. If the image processing difficulty of the target image is large, the target resolution corresponding to the second neural network is large. Still taking the above example, if image A is input into the first neural network, the target resolution corresponding to the second neural network output by the first neural network is 168×168. If image B is input into the second neural network, the target resolution corresponding to the second neural network output by the first neural network is 224×224.
[0180] Furthermore, the first neural network can obtain the target resolution corresponding to the second neural network in the following way:
[0181] The target image is input into the third neural network to obtain the probability of the candidate resolution corresponding to the second neural network. According to the magnitude of the probability of the candidate resolution corresponding to the second neural network, the target resolution corresponding to the second neural network is selected from the candidate resolutions corresponding to the second neural network.
[0182] To further understand the foregoing process, the following is combined with Figure 7 for further illustration. Figure 7 FIG. is a schematic diagram of another application example of the neural network compression method provided by the embodiment of the present application. As Figure 7 shown, it is assumed that there are image A and image B (the image processing difficulty of image A is lower than that of image B), and the candidate resolutions corresponding to the second neural network are 168×168, 200×200, and 224×224.
[0183] After inputting Image A into the first neural network, the probabilities of the candidate resolutions corresponding to the second neural network can be obtained, that is, the probabilities of 168×168, 200×200, and 224×224. Among them, the probability of 168×168 is the largest, so 168×168 can be determined as the target resolution corresponding to the second neural network.
[0184] After inputting Image B into the first neural network, the probabilities of the candidate resolutions corresponding to the second neural network can be obtained, that is, the probabilities of 168×168, 200×200, and 224×224. Among them, the probability of 224×224 is the largest, so 224×224 can be determined as the target resolution corresponding to the second neural network.
[0185] 603. Adjust the resolution of the target image to the target resolution to obtain the target image with the target resolution.
[0186] After obtaining the target resolution corresponding to the second neural network, the target image can be adjusted from the original resolution to the target resolution, so as to obtain the target image with the target resolution. For example, assume the target resolution is 168×168 and the original resolution of the target image is 400×400. Then, the resolution of the target image is adjusted from 400×400 to 168×168 to obtain the target image with the resolution of 168×168.
[0187] 604. Input the target image with the target resolution into the second neural network to obtain the features of the target image with the target resolution.
[0188] After obtaining the target image with the target resolution, the target image with the target resolution can be input into the second neural network so that the second neural network performs image processing on the target image with the target resolution (such as image classification, object detection, and image segmentation, etc.), thereby obtaining the features of the target image with the target resolution.
[0189] In this embodiment, after obtaining the target image, the target image can be input into the third neural network to obtain the target resolution corresponding to the second neural network, which is used to compress the input of the second neural network, that is, the target image. The target resolution is positively correlated with the amount of computation required for image processing of the target image. If the amount of computation required for image processing of the target image is large, the target resolution is large; if the amount of computation required for image processing of the target image is small, the target resolution is small. It can be seen that due to the different processing difficulties of the target images, the target resolutions output by the third neural network are different, so that the target resolutions of the adjusted target images are also different. In this way, for target images with relatively low processing difficulty, their resolutions can be made smaller, enabling the second neural network to perform image processing with relatively low precision on them. For target images with relatively high processing difficulty, their resolutions can be made larger, enabling the second neural network to perform image processing with relatively high precision on them, thereby ensuring the accuracy of image processing.
[0190] It should be noted that, in combination with Figure 4 the embodiment shown in Figure 6 and the embodiment of
[0191] After obtaining the target image with the target resolution, the target image with the target resolution can also be input into the second neural network after quantization processing, so that the second neural network after quantization processing performs image processing on the target image with the target resolution, thereby obtaining the features of the target image with the target resolution.
[0192] Table 3
[0193]
[0194] Based on Table 3, it can be seen that when the neural network provided in the embodiment of the present application performs image processing, it can achieve higher accuracy while occupying a lower amount of computation.
[0195] The above is a detailed description of the neural network compression method provided in the embodiment of the present application. Next, the model training method provided in the embodiment of the present application will be introduced. Figure 8 is a schematic flowchart of a model training method provided in an embodiment of the present application. As Figure 8 shown, the method includes:
[0196] 801. Obtain the image to be trained.
[0197] When model training is to be performed, the training images to be used for model training can be obtained first. It should be noted that the true features of the training images to be used are known. Further, the first model to be trained and the second model to be trained can also be obtained to perform joint training on these two models.
[0198] 802. Input the training images to be used into the first model to be trained to obtain the quantization bit number of the second model to be trained.
[0199] After obtaining the training images to be used and the first model to be trained, the training images to be used can be input into the first model to be trained to obtain the quantization bit number of the second model to be trained.
[0200] Specifically, the quantization bit number of the second model to be trained includes the quantization bit numbers of the M-layer networks in the second model to be trained, where M is a positive integer. Then, the first model to be trained can obtain the quantization bit numbers of the M-layer networks in the second model to be trained in the following manner:
[0201] Input the training images to be used into the first model to be trained to obtain the probability of the candidate bit numbers of the i-th layer network in the second model to be trained; according to the magnitude of the probability of the candidate bit numbers of the i-th layer network in the second model to be trained, select the quantization bit number of the i-th layer network from the candidate bit numbers of the i-th layer network, where i = 1, 2,..., M.
[0202] 803. According to the quantization bit number of the second model to be trained, perform quantization processing on the parameters of the second model to be trained to obtain the second model to be trained after quantization processing.
[0203] After obtaining the quantization bit number of the second model to be trained, quantization processing can be performed on the parameters of the second model to be trained according to the quantization bit number of the second model to be trained to obtain the second model to be trained after quantization processing. Specifically, quantization processing can be performed on the parameters of the i-th layer network according to the quantization bit number of the i-th layer network in the second model to be trained to obtain the second model to be trained after quantization processing.
[0204] 804. Input the training images to be used into the second model to be trained after quantization processing to obtain the features of the training images to be used.
[0205] After obtaining the second model to be trained after quantization processing, the training images to be used can be input into the second model to be trained after quantization processing so that the second model to be trained after quantization processing performs image processing on the training images to be used, thereby obtaining the features (predicted features) of the training images to be used.
[0206] Regarding the descriptions of steps 802 to 804, reference can be made to Figure 4 the relevant description parts of steps 402 to 404 in the illustrated embodiments, which will not be elaborated here.
[0207] 805. Obtain the target loss according to the deviation between the quantization bit number and the preset bit number of the second model to be trained, and the deviation between the features of the image to be trained and the true features of the image to be trained.
[0208] After obtaining the quantization bit number of the second model to be trained and the features of the image to be trained, the target loss can be obtained according to the deviation between the quantization bit number and the preset bit number of the second model to be trained, and the deviation between the features of the image to be trained and the true features of the image to be trained. Specifically, by constructing a standard cross-entropy loss and introducing a regularization term, the target loss is obtained, where the standard cross-entropy loss is used to indicate the deviation between the features of the image to be trained and the true features of the image to be trained, and the regularization term is used to indicate the deviation between the quantization bit number of the i-th layer network in the second model to be trained and the preset bit number (that is, the variation between the quantization bit numbers of each layer network and the threshold bit number in the M-layer network of the second model to be trained).
[0209] The target loss in this embodiment can be obtained through formula (2):
[0210]
[0211] In the above formula, L is Figure 8 the target loss in the embodiment shown, L cls is the standard cross-entropy loss, which is generated based on the features of the image to be trained and the true features of the image to be trained, α is a preset hyperparameter, B i is the amount of computation corresponding to the quantization bit number of the i-th layer network (that is, the amount of computation occupied when the parameters of the i-th layer network are represented by the values of the corresponding quantization bit numbers), B tar is the amount of computation corresponding to the preset bit number.
[0212] 806. Update the parameters of the first model to be trained and the parameters of the second model to be trained according to the target loss until the model training conditions are met, and obtain the first neural network and the second neural network.
[0213] After obtaining the target loss, it can be judged whether the target loss converges. If the target loss has not converged, update the parameters of the first model to be trained and the parameters of the second model to be trained (the non-quantized second model to be trained), and re-perform joint training on the first model to be trained and the second model to be trained with new images to be trained until the target loss converges, that is, the model training conditions are met, and obtain Figure 4 the first neural network and the second neural network in the embodiment shown.
[0214] In this embodiment, by jointly training the first model to be trained and the second model to be trained, a first neural network and a second neural network can be obtained. The first neural network obtained through this embodiment can accurately obtain the quantization bits of the second neural network based on the input target image, and the second neural network obtained through this embodiment can accurately perform image processing on the target image.
[0215] Figure 9 Another process schematic diagram of the model training method provided by the embodiments of the present application is shown in Figure 9 As shown, the method includes:
[0216] 901. Obtain the image to be trained;
[0217] When model training is to be performed, the image to be trained for model training can be obtained first. It should be noted that the true features of the image to be trained with the target resolution are known. Further, a third model to be trained and a second model to be trained can also be obtained to jointly train these two models.
[0218] 902. Input the image to be trained into the third model to be trained to obtain the target resolution corresponding to the second model to be trained.
[0219] After obtaining the image to be trained and the third model to be trained, input the image to be trained into the third model to be trained to obtain the target resolution corresponding to the second model to be trained.
[0220] Specifically, the third model to be trained can obtain the target resolution corresponding to the second model to be trained in the following manner:
[0221] Input the image to be trained into the third model to be trained to obtain the probability of the candidate resolution corresponding to the second model to be trained; according to the magnitude of the probability of the candidate resolution corresponding to the second model to be trained, select the target resolution corresponding to the second model to be trained from the candidate resolutions corresponding to the second model to be trained.
[0222] 903. Adjust the resolution of the image to be trained to the target resolution to obtain the image to be trained with the target resolution.
[0223] After obtaining the target resolution corresponding to the second model to be trained, the image to be trained can be adjusted from the original resolution to the target resolution to obtain the image to be trained with the target resolution.
[0224] 904. Input the image to be trained with the target resolution into the second model to be trained to obtain the features of the image to be trained with the target resolution.
[0225] After obtaining the to-be-trained image with the target resolution, the to-be-trained image with the target resolution can be input into the second to-be-trained model, so that the second to-be-trained model performs image processing on the to-be-trained image with the target resolution, thereby obtaining the features (predicted features) of the to-be-trained image with the target resolution.
[0226] For the descriptions of steps 902 to 904, reference can be made to Figure 6 the relevant description parts of steps 602 to 604 in the embodiment shown, which will not be elaborated here.
[0227] 905. Obtain the target loss according to the deviation between the expected value corresponding to the target resolution and the preset expected value, and the deviation between the features of the to-be-trained image with the target resolution and the true features of the to-be-trained image with the target resolution.
[0228] After obtaining the target resolution corresponding to the second to-be-trained model and the features of the target image with the target resolution, the target loss can be obtained according to the deviation between the expected value corresponding to the target resolution and the preset expected value, and the deviation between the features of the to-be-trained image with the target resolution and the true features of the to-be-trained image with the target resolution. Specifically, by constructing a standard cross-entropy loss and introducing a regularization term, the target loss is obtained, where the standard cross-entropy loss is used to indicate the deviation between the features of the to-be-trained image with the target resolution and the true features of the to-be-trained image with the target resolution, and the regularization term is used to indicate the deviation according to the expected value of the probability of the target resolution and the preset expected value, and the deviation between the expected value of the probability of the remaining candidate resolutions except the target resolution and the preset expected value.
[0229] The target loss in this embodiment can be obtained through formula (3):
[0230]
[0231] In the above formula, L0 is Figure 9 the target loss in the embodiment shown, L ce is the standard cross-entropy loss, which is generated based on the features of the to-be-trained image with the target resolution and the true features of the to-be-trained image with the target resolution, η is a preset hyperparameter, β is the preset expected value (preset penalty coefficient), E(h j ) is the expected value (mathematical expectation) of the probability of the jth candidate resolution, and N is the number of candidate resolutions. It can be seen that if the expected value of the probability of a certain candidate resolution is less than β, then L reg will increase.
[0232] 906. Update the parameters of the third to-be-trained model and the parameters of the second to-be-trained model according to the target loss until the model training conditions are met, and obtain the third neural network and the second neural network.
[0233] After obtaining the target loss, it can be determined whether the target loss converges. If the target loss has not converged, the parameters of the third model to be trained and the parameters of the second model to be trained are updated, and the third model to be trained and the second model to be trained are jointly trained again with new images to be trained until the target loss converges, that is, the model training conditions are met, and the Figure 6 third neural network and the second neural network in the illustrated embodiment are obtained.
[0234] In this embodiment, by jointly training the third model to be trained and the second model to be trained, the third neural network and the second neural network can be obtained. The third neural network obtained through this embodiment can accurately obtain the target resolution corresponding to the second neural network based on the input target image, and the second neural network obtained through this embodiment can perform accurate image processing on the target image.
[0235] The above is a detailed description of the model training method provided by the embodiments of the present application. Next, the neural network compression device provided by the embodiments of the present application will be introduced. Figure 10 is a schematic structural diagram of the neural network compression device provided by the embodiments of the present application. As Figure 10 shown, the device includes: an acquisition module 1001 and a processing module 1002;
[0236] The acquisition module 1001 is configured to acquire a target image;
[0237] The processing module 1002 is configured to input the target image into the first neural network to obtain the quantization bit number of the second neural network, where the second neural network is used to perform image processing on the target image, and the quantization bit number of the second neural network is positively correlated with the amount of computation required for image processing;
[0238] The processing module 1002 is further configured to perform quantization processing on the parameters of the second neural network according to the quantization bit number of the second neural network to obtain the second neural network after quantization processing.
[0239] As can be seen from the above device: after obtaining the target image, the target image can be input into the first neural network to obtain the quantization bits for compressing the second neural network. The quantization bits are positively correlated with the amount of computation required for image processing of the target image. If the amount of computation required for image processing of the target image is large, the quantization bits are large; if the amount of computation required for image processing of the target image is small, the quantization bits are small. It can be seen that due to the different processing difficulties of the target images, the quantization bits output by the first neural network are different, resulting in different degrees of quantization processing for the second neural network. In this way, for target images with relatively low processing difficulty, the second neural network after quantization processing can perform image processing with relatively low precision; for target images with relatively high processing difficulty, the second neural network after quantization processing can perform image processing with relatively high precision, thus ensuring the accuracy of image processing.
[0240] In a possible implementation, the quantization bits of the second neural network include the quantization bits of M layers of the second neural network. The processing module 1002 is specifically configured to perform quantization processing on the parameters of the i-th layer of the second neural network according to the quantization bits of the i-th layer of the second neural network to obtain the second neural network after quantization processing, where i = 1, 2,..., M and M is a positive integer.
[0241] In a possible implementation, the processing module 1002 is specifically configured to: input the target image into the first neural network to obtain the probability of the candidate bits of the i-th layer of the second neural network; select the quantization bits of the i-th layer of the second neural network from the candidate bits of the i-th layer according to the magnitude of the probability.
[0242] In a possible implementation, the processing module 1002 is further configured to input the target image into the second neural network after quantization processing to obtain the features of the target image.
[0243] Figure 11 Another structural schematic diagram of the neural network compression device provided by the embodiments of the present application is as Figure 11 shown. The device includes an acquisition module 1101 and a processing module 1102;
[0244] The acquisition module 1101 is configured to acquire a target image;
[0245] The processing module 1102 is configured to input the target image into the third neural network to obtain the target resolution corresponding to the second neural network, where the second neural network is used for image processing of the target image, and the target resolution corresponding to the second neural network is positively correlated with the amount of computation required for image processing;
[0246] The processing module 1102 is further configured to adjust the resolution of the target image to the target resolution to obtain the target image with the target resolution.
[0247] As can be seen from the above device: after obtaining the target image, the target image can be input into the third neural network to obtain the target resolution corresponding to the second neural network, which is used to compress the input of the second neural network, that is, the target image. This target resolution is positively correlated with the amount of computation required for image processing of the target image. If the amount of computation required for image processing of the target image is large, then this target resolution is large; if the amount of computation required for image processing of the target image is small, then this target resolution is small. It can be seen that due to the different processing difficulties of the target images, the target resolutions output by the third neural network are different, so that the target resolutions of the adjusted target images are also different. In this way, for target images with relatively low processing difficulty, their resolutions can be made relatively small, so that the second neural network performs image processing with relatively low precision on them; for target images with relatively high processing difficulty, their resolutions can be made relatively large, so that the second neural network performs image processing with relatively high precision on them, thus ensuring the accuracy of image processing.
[0248] In a possible implementation, the processing module 1102 is specifically configured to: input the target image into the third neural network to obtain the probability of the candidate resolutions corresponding to the second neural network; select the target resolution corresponding to the second neural network from the candidate resolutions corresponding to the second neural network according to the magnitude of the probability.
[0249] In a possible implementation, the processing module 1102 is further configured to input the target image with the target resolution into the second neural network to obtain the features of the target image with the target resolution.
[0250] The above is a detailed description of the neural network compression device provided by the embodiments of the present application. The model training device provided by the embodiments of the present application will be introduced below. Figure 12 FIG. is a schematic structural diagram of the model training device provided by the embodiments of the present application, as Figure 12 shown, the device includes an acquisition module 1201 and a training module 1202;
[0251] The acquisition module 1201 is configured to acquire the image to be trained;
[0252] The training module 1202 is configured to input the image to be trained into the first model to be trained to obtain the quantization bits of the second model to be trained;
[0253] The training module 1202 is further configured to perform quantization processing on the parameters of the second model to be trained according to the quantization bits of the second model to be trained to obtain the second model to be trained after quantization processing;
[0254] The training module 1202 is further configured to input the image to be trained into the second model to be trained after quantization processing to obtain the features of the image to be trained;
[0255] The training module 1202 is further configured to update the parameters of the first model to be trained and the parameters of the second model to be trained according to the quantization bits of the second model to be trained and the features of the image to be trained until the model training conditions are met, so as to obtain the first neural network and the second neural network.
[0256] It can be seen from the above device that: by jointly training the first model to be trained and the second model to be trained, the first neural network and the second neural network can be obtained. The first neural network obtained by this device can accurately obtain the quantization bits of the second neural network based on the input target image, and the second neural network obtained by this device can perform accurate image processing on the target image.
[0257] In a possible implementation, the training module 1202 is specifically configured to: obtain a target loss according to the deviation between the quantization bits of the second model to be trained and the preset bits, and the deviation between the features of the image to be trained and the true features of the image to be trained; update the parameters of the first model to be trained and the parameters of the second model to be trained according to the target loss until the model training conditions are met, so as to obtain the first neural network and the second neural network.
[0258] In a possible implementation, the quantization bits of the second model to be trained include the quantization bits of the M layers of the network in the second model to be trained. The training module 1202 is specifically configured to perform quantization processing on the parameters of the i-th layer of the network according to the quantization bits of the i-th layer of the network in the second model to be trained, so as to obtain the second model to be trained after quantization processing, where i = 1, 2,..., M, and M is a positive integer.
[0259] In a possible implementation, the training module 1202 is specifically configured to: input the image to be trained into the first model to be trained to obtain the probability of the candidate bits of the i-th layer of the network in the second model to be trained; select the quantization bits of the i-th layer of the network from the candidate bits of the i-th layer of the network according to the magnitude of the probability.
[0260] In a possible implementation, the deviation between the quantization bits of the second model to be trained and the preset bits includes the deviation between the quantization bits of the i-th layer of the network in the second model to be trained and the preset bits.
[0261] Figure 13 Another structural schematic diagram of the model training device provided by the embodiment of the present application is as Figure 13 shown. The device includes an acquisition module 1301 and a training module 1302;
[0262] The acquisition module 1301 is configured to acquire the image to be trained;
[0263] The training module 1302 is configured to input the image to be trained into the third model to be trained, and obtain the target resolution corresponding to the second model to be trained.
[0264] The training module 1302 is further configured to adjust the resolution of the image to be trained to the target resolution, and obtain the image to be trained with the target resolution.
[0265] The training module 1302 is further configured to input the image to be trained with the target resolution into the second model to be trained, and obtain the features of the image to be trained with the target resolution.
[0266] The training module 1302 is further configured to update the parameters of the third model to be trained and the parameters of the second model to be trained according to the target resolution and the features of the image to be trained with the target resolution, until the model training condition is satisfied, and obtain the third neural network and the second neural network.
[0267] It can be seen from the above device that: by jointly training the third model to be trained and the second model to be trained, the third neural network and the second neural network can be obtained. The third neural network obtained by this device can accurately obtain the target resolution corresponding to the second neural network based on the input target image, and the second neural network obtained by this device can accurately perform image processing on the target image.
[0268] In a possible implementation manner, the training module 1302 is specifically configured to: obtain the target loss according to the deviation between the expected value corresponding to the target resolution and the preset expected value, and the deviation between the features of the image to be trained with the target resolution and the true features of the image to be trained with the target resolution; update the parameters of the third model to be trained and the parameters of the second model to be trained according to the target loss, until the model training condition is satisfied, and obtain the third neural network and the second neural network.
[0269] In a possible implementation manner, the training module 1302 is specifically configured to: input the image to be trained into the third model to be trained, and obtain the probability of the candidate resolution corresponding to the second model to be trained; select the target resolution corresponding to the second model to be trained from the candidate resolutions corresponding to the second model to be trained according to the magnitude of the probability.
[0270] In a possible implementation manner, the training module 1302 is specifically configured to obtain the target loss according to the deviation between the expected value of the probability of the target resolution and the preset expected value, the deviation between the expected value of the probability of the other candidate resolutions except the target resolution and the preset expected value, and the deviation between the features of the image to be trained with the target resolution and the true features of the image to be trained with the target resolution.
[0271] It should be noted that the information interaction, execution process, etc. between the modules / units of the above-mentioned device are based on the same concept as the method embodiment of the present application, and the technical effects they bring are the same as those of the method embodiment of the present application. For specific contents, please refer to the description in the method embodiment shown above in the embodiment of the present application, and will not be repeated here.
[0272] The embodiment of the present application also relates to an execution device, Figure 14 A schematic diagram of the structure of the execution device provided in the embodiment of the present application. Figure 14 As shown, the execution device 1400 can be specifically a mobile phone, a tablet, a laptop, a smart wearable device, a server, etc., which is not limited here. Figure 10 or Figure 11 The neural network compression device described in the corresponding embodiment is used to implement Figure 4 or Figure 6 The neural network compression function and the image processing function in the corresponding embodiment. Specifically, the execution device 1400 includes: a receiver 1401, a transmitter 1402, a processor 1403 and a memory 1404 (wherein the number of the processor 1403 in the execution device 1400 can be one or more, Figure 14 In the example of FIG. 1403 , the processor 1403 may include an application processor 14031 and a communication processor 14032. In some embodiments of the present application, the receiver 1401, the transmitter 1402, the processor 1403 and the memory 1404 may be connected via a bus or other means.
[0273] The memory 1404 may include a read-only memory and a random access memory, and provides instructions and data to the processor 1403. A portion of the memory 1404 may also include a non-volatile random access memory (NVRAM). The memory 1404 stores processor and operation instructions, executable modules or data structures, or subsets thereof, or extended sets thereof, wherein the operation instructions may include various operation instructions for implementing various operations.
[0274] The processor 1403 controls the operation of the execution device. In a specific application, the various components of the execution device are coupled together through a bus system, wherein the bus system includes not only a data bus but also a power bus, a control bus, and a status signal bus, etc. However, for the sake of clarity, various buses are referred to as bus systems in the figure.
[0275] The method disclosed in the embodiments of the present application can be applied to or implemented by the processor 1403. The processor 1403 can be an integrated circuit chip with signal processing capabilities. In the implementation process, each step of the above method can be completed by the integrated logic circuit in the hardware of the processor 1403 or the instructions in the form of software. The above-mentioned processor 1403 can be a general-purpose processor, a digital signal processor (DSP), a microprocessor or a microcontroller, and can further include an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. The processor 1403 can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed and completed by the hardware decoding processor, or executed and completed by the combination of the hardware and software modules in the decoding processor. The software module can be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory 1404, and the processor 1403 reads the information in the memory 1404 and combines its hardware to complete the steps of the above method.
[0276] The receiver 1401 can be used to receive input digital or character information, and generate signal inputs related to the relevant settings and function controls of the execution device. The transmitter 1402 can be used to output digital or character information through the first interface; the transmitter 1402 can also be used to send instructions to the disk group through the first interface to modify the data in the disk group; the transmitter 1402 can also include a display device such as a display screen.
[0277] In an embodiment of the present application, in one case, the processor 1403 is used to execute Figure 4 or Figure 6 the neural network compression method in the corresponding embodiment.
[0278] The embodiment of the present application also relates to a training device, Figure 15 which is a schematic structural diagram of the training device provided by the embodiment of the present application. As Figure 15As shown, the training device 1500 is implemented by one or more servers. The training device 1500 can vary significantly due to configuration or performance differences, and may include one or more central processing units (CPUs) 1515 (e.g., one or more processors) and a memory 1532, and one or more storage media 1530 (e.g., one or more mass storage devices) storing application programs 1542 or data 1544. Among them, the memory 1532 and the storage media 1530 can be transient storage or persistent storage. The program stored in the storage media 1530 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the training device. Further, the central processing unit 1515 can be configured to communicate with the storage media 1530 and execute a series of instruction operations in the storage media 1530 on the training device 1500.
[0279] The training device 1500 may further include one or more power supplies 1526, one or more wired or wireless network interfaces 1550, one or more input / output interfaces 1558; or, one or more operating systems 1541, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, and so on.
[0280] Specifically, the training device can execute Figure 8 or Figure 9 the steps in the corresponding embodiments.
[0281] The embodiment of the present application also relates to a computer storage medium. The computer-readable storage medium stores a program for signal processing. When it runs on a computer, it causes the computer to execute the steps performed by the aforementioned execution device, or causes the computer to execute the steps performed by the aforementioned training device.
[0282] The embodiment of the present application also relates to a computer program product. The computer program product stores instructions. When the instructions are executed by a computer, they cause the computer to execute the steps performed by the aforementioned execution device, or cause the computer to execute the steps performed by the aforementioned training device.
[0283] The execution device, training device, or terminal device provided by the embodiments of the present application may specifically be a chip, and the chip includes: a processing unit and a communication unit. The processing unit may be a processor, for example, and the communication unit may be an input / output interface, a pin, a circuit, or the like. The processing unit may execute the computer execution instructions stored in the storage unit to cause the chip in the execution device to execute the data processing method described in the above embodiments, or to cause the chip in the training device to execute the data processing method described in the above embodiments. Optionally, the storage unit is a storage unit in the chip, such as a register, a cache, etc., and the storage unit may also be a storage unit outside the chip in the radio access device, such as a read-only memory (ROM) or other types of static storage devices that can store static information and instructions, a random access memory (RAM), etc.
[0284] Specifically, please refer to Figure 16 , Figure 16 which is a schematic structural diagram of the chip provided by the embodiments of the present application. The chip may be embodied as a neural network processor NPU 1600. The NPU 1600 is mounted on the main CPU (Host CPU) as a coprocessor, and tasks are assigned by the Host CPU. The core part of the NPU is the arithmetic circuit 1603, and the arithmetic circuit 1603 is controlled by the controller 1604 to extract matrix data from the memory and perform multiplication operations.
[0285] In some implementations, the arithmetic circuit 1603 includes multiple processing units (Process Engine, PE) inside. In some implementations, the arithmetic circuit 1603 is a two-dimensional systolic array. The arithmetic circuit 1603 may also be a one-dimensional systolic array or other electronic circuits capable of performing mathematical operations such as multiplication and addition. In some implementations, the arithmetic circuit 1603 is a general matrix processor.
[0286] For example, assume there is an input matrix A, a weight matrix B, and an output matrix C. The arithmetic circuit fetches the corresponding data of matrix B from the weight memory 1602 and caches it on each PE in the arithmetic circuit. The arithmetic circuit fetches the data of matrix A from the input memory 1601 and performs matrix operations with matrix B, and the partial results or final results of the obtained matrix are saved in the accumulator 1608.
[0287] The unified memory 1606 is used to store input data and output data. The weight data is directly transferred through the Direct Memory Access Controller (DMAC) 1605 to the weight memory 1602. The input data is also transferred through the DMAC to the unified memory 1606.
[0288] The BIU is the Bus Interface Unit, i.e., the bus interface unit 1610, which is used for the interaction between the AXI bus, the DMAC, and the Instruction Fetch Buffer (IFB) 1609.
[0289] The bus interface unit 1610 (Bus Interface Unit, abbreviated as BIU) is used for the instruction fetch buffer 1609 to obtain instructions from the external memory, and is also used for the storage unit access controller 1605 to obtain the original data of the input matrix A or the weight matrix B from the external memory.
[0290] The DMAC is mainly used to transfer the input data in the external memory DDR to the unified memory 1606, or transfer the weight data to the weight memory 1602, or transfer the input data to the input memory 1601.
[0291] The vector calculation unit 1607 includes multiple arithmetic processing units, which, if necessary, further process the output of the arithmetic circuit 1603, such as vector multiplication, vector addition, exponential operation, logarithmic operation, magnitude comparison, etc. It is mainly used for non-convolution / full connection layer network calculations in neural networks, such as Batch Normalization, pixel-level summation, upsampling of the feature plane, etc.
[0292] In some implementations, the vector calculation unit 1607 can store the processed output vector in the unified memory 1606. For example, the vector calculation unit 1607 can apply a linear function; or, a non-linear function to the output of the arithmetic circuit 1603, such as performing linear interpolation on the feature plane extracted by the convolutional layer, or, for example, a vector of accumulated values, to generate activation values. In some implementations, the vector calculation unit 1607 generates normalized values, pixel-level summation values, or both. In some implementations, the processed output vector can be used as the activation input to the arithmetic circuit 1603, for example, for use in subsequent layers in the neural network.
[0293] The instruction fetch buffer 1609 connected to the controller 1604 is used to store the instructions used by the controller 1604;
[0294] The unified memory 1606, the input memory 1601, the weight memory 1602, and the instruction fetch memory 1609 are all On-Chip memories. The external memory is private to the NPU hardware architecture.
[0295] Wherein, the processor mentioned anywhere above can be a general-purpose central processing unit, a microprocessor, an ASIC, or one or more integrated circuits for controlling the execution of the above programs.
[0296] In addition, it should be noted that the device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. In addition, in the attached drawings of the device embodiments provided in this application, the connection relationships between the modules indicate that they have communication connections, which can be specifically implemented as one or more communication buses or signal lines.
[0297] Through the description of the above embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general hardware, and of course, it can also be implemented by dedicated hardware including dedicated integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be various, such as analog circuits, digital circuits, or dedicated circuits. However, for this application, in more cases, software program implementation is a better implementation method. Based on such an understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a readable storage medium, such as a floppy disk, a USB flash drive, a mobile hard disk, a ROM, a RAM, a magnetic disk, or an optical disc of a computer, and includes several instructions for causing a computer device (which can be a personal computer, a training device, or a network device, etc.) to execute the methods described in the various embodiments of this application.
[0298] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0299] The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, training device, or data center to another website, computer, training device, or data center via wired (such as coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that a computer can store or a data storage device such as a training device or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)), etc.
Claims
1. A neural network compression method, characterized in that, The method includes: Input the target image into a first neural network to obtain the quantization bit number of a second neural network, where the second neural network is used to perform image processing on the target image, and the quantization bit number of the second neural network is positively correlated with the amount of computation required for the image processing; Quantize the parameters of the second neural network according to the quantization bit number of the second neural network to obtain the second neural network after quantization processing.
2. The method according to claim 1, wherein The quantization bit number of the second neural network includes the quantization bit numbers of M layers of networks in the second neural network. The step of quantizing the parameters of the second neural network according to the quantization bit number of the second neural network to obtain the second neural network after quantization processing specifically includes: Quantize the parameters of the i-th layer of network in the second neural network according to the quantization bit number of the i-th layer of network in the second neural network to obtain the second neural network after quantization processing, where i = 1, 2,..., M, and M is a positive integer.
3. The method according to claim 2, wherein The step of inputting the target image into the first neural network to obtain the quantization bit number of the second neural network specifically includes: Input the target image into the first neural network to obtain the probability of the candidate bit number of the i-th layer of network in the second neural network; Select the quantization bit number of the i-th layer of network from the candidate bit numbers of the i-th layer of network according to the magnitude of the probability.
4. The method according to any one of claims 1 to 3, characterized in that The method further includes: Input the target image into the second neural network after quantization processing to obtain the features of the target image.
5. A neural network compression method, characterized in that, The method includes: Input the target image into a third neural network to obtain the target resolution corresponding to the second neural network, where the second neural network is used to perform image processing on the target image, and the target resolution corresponding to the second neural network is positively correlated with the amount of computation required for the image processing; Adjust the resolution of the target image to the target resolution to obtain the target image with the target resolution.
6. The method according to claim 5, wherein The step of inputting the target image into the third neural network to obtain the target resolution corresponding to the second neural network specifically includes: Input the target image into the third neural network to obtain the probability of the candidate resolution corresponding to the second neural network; Select the target resolution corresponding to the second neural network from the candidate resolutions corresponding to the second neural network according to the magnitude of the probability.
7. The method according to claim 5 or 6, characterized in that, After the step of adjusting the resolution of the target image to the target resolution to obtain the target image with the target resolution, the method further includes: Input the target image with the target resolution into the second neural network to obtain the features of the target image with the target resolution.
8. A model training method, characterized in that, The method includes: Input the image to be trained into a first model to be trained to obtain the quantization bit number of a second model to be trained; Quantize the parameters of the second model to be trained according to the quantization bit number of the second model to be trained to obtain the second model to be trained after quantization processing; Input the image to be trained into the second model to be trained after quantization processing to obtain the features of the image to be trained; Update the parameters of the first model to be trained and the parameters of the second model to be trained according to the quantization bits of the second model to be trained and the features of the image to be trained until the model training conditions are met, and obtain the first neural network and the second neural network.
9. The method according to claim 8, wherein The step of updating the parameters of the first model to be trained and the parameters of the second model to be trained according to the quantization bits of the second model to be trained and the features of the image to be trained until the model training conditions are met, and obtaining the first neural network and the second neural network specifically includes: Obtain the target loss according to the deviation between the quantization bits of the second model to be trained and the preset bits, and the deviation between the features of the image to be trained and the true features of the image to be trained. Update the parameters of the first model to be trained and the parameters of the second model to be trained according to the target loss until the model training conditions are met, and obtain the first neural network and the second neural network.
10. The method according to claim 8 or 9, characterized in that, The quantization bits of the second model to be trained include the quantization bits of the M layers of networks in the second model to be trained. The step of quantizing the parameters of the second model to be trained according to the quantization bits of the second model to be trained to obtain the quantized second model to be trained specifically includes: Quantize the parameters of the i-th layer of network in the second model to be trained according to the quantization bits of the i-th layer of network in the second model to be trained, and obtain the quantized second model to be trained, where i = 1, 2,..., M, and M is a positive integer.
11. The method according to claim 10, characterized in that, The step of inputting the image to be trained into the first model to be trained to obtain the quantization bits of the second model to be trained specifically includes: Input the image to be trained into the first model to be trained to obtain the probability of the candidate bits of the i-th layer of network in the second model to be trained. Select the quantization bits of the i-th layer of network from the candidate bits of the i-th layer of network according to the magnitude of the probability.
12. The method according to claim 9, wherein The deviation between the quantization bits of the second model to be trained and the preset bits includes the deviation between the quantization bits of the i-th layer of network in the second model to be trained and the preset bits.
13. A model training method, characterized in that, The method includes: Input the image to be trained into the third model to be trained to obtain the target resolution corresponding to the second model to be trained. Adjust the resolution of the image to be trained to the target resolution to obtain the image to be trained with the target resolution. Input the image to be trained with the target resolution into the second model to be trained to obtain the features of the image to be trained with the target resolution. Update the parameters of the third model to be trained and the parameters of the second model to be trained according to the target resolution and the features of the image to be trained with the target resolution until the model training conditions are met, and obtain the third neural network and the second neural network.
14. The method according to claim 13, wherein The step of updating the parameters of the third model to be trained and the parameters of the second model to be trained according to the target resolution and the features of the image to be trained with the target resolution until the model training conditions are met, and obtaining the third neural network and the second neural network specifically includes: Obtain a target loss according to the deviation between the expected value corresponding to the target resolution and the preset expected value, and the deviation between the features of the training image of the target resolution and the true features of the training image of the target resolution. Update the parameters of the third model to be trained and the parameters of the second model to be trained according to the target loss until the model training conditions are met, and obtain a third neural network and a second neural network.
15. The method according to claim 13 or 14, characterized in that, The step of inputting the training image into the third model to be trained to obtain the target resolution corresponding to the second model to be trained specifically includes: Input the training image into the third model to be trained to obtain the probability of the candidate resolution corresponding to the second model to be trained. Select the target resolution corresponding to the second model to be trained from the candidate resolutions corresponding to the second model to be trained according to the magnitude of the probability.
16. The method according to claim 14, wherein The step of obtaining the target loss according to the deviation between the expected value corresponding to the target resolution and the preset expected value, and the deviation between the features of the training image of the target resolution and the true features of the training image of the target resolution specifically includes: Obtain the target loss according to the deviation between the expected value of the probability of the target resolution and the preset expected value, the deviation between the expected value of the probability of the other candidate resolutions except the target resolution and the preset expected value, and the deviation between the features of the training image of the target resolution and the true features of the training image of the target resolution.
17. A neural network compression device, characterized in that The device includes a processing module; The processing module is configured to input a target image into a first neural network to obtain the quantization bit number of a second neural network, where the second neural network is used to perform image processing on the target image, and the quantization bit number of the second neural network is positively correlated with the amount of computation required for the image processing. The processing module is further configured to perform quantization processing on the parameters of the second neural network according to the quantization bit number of the second neural network to obtain a second neural network after quantization processing.
18. The device according to claim 17, wherein The quantization bit number of the second neural network includes the quantization bit numbers of M layers of networks in the second neural network. Specifically, the processing module is configured to perform quantization processing on the parameters of the i-th layer of network in the second neural network according to the quantization bit number of the i-th layer of network in the second neural network to obtain a second neural network after quantization processing, where i = 1, 2,..., M, and M is a positive integer.
19. The device according to claim 18, characterized in that, The processing module is specifically configured to: Input the target image into the first neural network to obtain the probability of the candidate bit numbers of the i-th layer of network in the second neural network. Select the quantization bit number of the i-th layer of network from the candidate bit numbers of the i-th layer of network according to the magnitude of the probability.
20. The device according to any one of claims 17 to 19, characterized in that The processing module is further configured to input the target image into the second neural network after quantization processing to obtain the features of the target image.
21. A neural network compression device, characterized in that, The device includes a processing module; The processing module is configured to input the target image into a third neural network to obtain the target resolution corresponding to a second neural network, where the second neural network is used to perform image processing on the target image, and the target resolution corresponding to the second neural network is positively correlated with the amount of computation required for the image processing; The processing module is further configured to adjust the resolution of the target image to the target resolution to obtain a target image with the target resolution.
22. The device according to claim 21, wherein, Specifically, the processing module is configured to: Input the target image into a third neural network to obtain the probability of candidate resolutions corresponding to the second neural network; Select the target resolution corresponding to the second neural network from the candidate resolutions corresponding to the second neural network according to the magnitude of the probability.
23. The device according to claim 21 or 22, characterized in that The processing module is further configured to input the target image with the target resolution into the second neural network to obtain the features of the target image with the target resolution.
24. A model training device, characterized in that, The apparatus includes a training module; The training module is configured to input a training image into a first model to be trained to obtain the quantization bit number of a second model to be trained; The training module is further configured to perform quantization processing on the parameters of the second model to be trained according to the quantization bit number of the second model to be trained to obtain a second model to be trained after quantization processing; The training module is further configured to input the training image into the second model to be trained after quantization processing to obtain the features of the training image; The training module is further configured to update the parameters of the first model to be trained and the parameters of the second model to be trained according to the quantization bit number of the second model to be trained and the features of the training image until the model training condition is satisfied, to obtain a first neural network and a second neural network.
25. The device according to claim 24, wherein Specifically, the training module is configured to: Obtain a target loss according to the deviation between the quantization bit number of the second model to be trained and a preset bit number, and the deviation between the features of the training image and the true features of the training image; Update the parameters of the first model to be trained and the parameters of the second model to be trained according to the target loss until the model training condition is satisfied, to obtain a first neural network and a second neural network.
26. The device according to claim 24 or 25, characterized in that, The quantization bit number of the second model to be trained includes the quantization bit numbers of M layers of networks in the second model to be trained. Specifically, the training module is configured to perform quantization processing on the parameters of the i-th layer of network in the second model to be trained according to the quantization bit number of the i-th layer of network in the second model to be trained to obtain a second model to be trained after quantization processing, where i = 1, 2,..., M, and M is a positive integer.
27. The device according to claim 26, characterized in that, Specifically, the training module is configured to: Input the training image into the first model to be trained to obtain the probability of candidate bit numbers of the i-th layer of network in the second model to be trained; Select the quantization bit number of the i-th layer of network from the candidate bit numbers of the i-th layer of network according to the magnitude of the probability.
28. The device according to claim 25, wherein The deviation between the quantization bit number of the second model to be trained and the preset bit number includes the deviation between the quantization bit number of the i-th layer of network in the second model to be trained and the preset bit number.
29. A model training device, characterized in that, The apparatus includes a training module; The training module is configured to input the image to be trained into the third model to be trained, and obtain the target resolution corresponding to the second model to be trained; The training module is further configured to adjust the resolution of the image to be trained to the target resolution, and obtain the image to be trained with the target resolution; The training module is further configured to input the image to be trained with the target resolution into the second model to be trained, and obtain the features of the image to be trained with the target resolution; The training module is further configured to update the parameters of the third model to be trained and the parameters of the second model to be trained according to the target resolution and the features of the image to be trained with the target resolution, until the model training condition is satisfied, and obtain the third neural network and the second neural network.
30. The device according to claim 29, wherein Specifically, the training module is configured to: Obtain the target loss according to the deviation between the expected value corresponding to the target resolution and the preset expected value, and the deviation between the features of the image to be trained with the target resolution and the true features of the image to be trained with the target resolution; Update the parameters of the third model to be trained and the parameters of the second model to be trained according to the target loss, until the model training condition is satisfied, and obtain the third neural network and the second neural network.
31. The device according to claim 29 or 30, characterized in that, Specifically, the training module is configured to: Input the image to be trained into the third model to be trained, and obtain the probability of the candidate resolution corresponding to the second model to be trained; Select the target resolution corresponding to the second model to be trained from the candidate resolutions corresponding to the second model to be trained according to the magnitude of the probability.
32. The device according to claim 30, characterized in that Specifically, the training module is configured to obtain the target loss according to the deviation between the expected value of the probability of the target resolution and the preset expected value, the deviation between the expected value of the probability of the other candidate resolutions except the target resolution and the preset expected value, and the deviation between the features of the image to be trained with the target resolution and the true features of the image to be trained with the target resolution.
33. A neural network compression device, characterized in that, It includes a memory and a processor; the memory stores code, and the processor is configured to execute the code. When the code is executed, the neural network compression device executes the method according to any one of claims 1 to 7.
34. A model training device, characterized in that, It includes a memory and a processor; the memory stores code, and the processor is configured to execute the code. When the code is executed, the model training device executes the method according to any one of claims 8 to 16.
35. A computer storage medium, characterized in that, The computer storage medium stores a computer program, and when the program is executed by a computer, the computer implements the method according to any one of claims 1 to 16.
36. A computer program product, characterized in that, The computer program product stores instructions, and when the instructions are executed by a computer, the computer implements the method according to any one of claims 1 to 16.