Training Method, Apparatus and Device for Neural Network

By comprehensively considering the loss value and sparseness value in neural network training and optimizing neural network parameters, the problem of information loss during sparse processing is solved, the balance between accuracy and sparseness is achieved, and the bandwidth and calculation amount is reduced.

CN114202052BActive Publication Date: 2025-06-24GUANGDONG GAOHANG INTELLECTUAL PROPERTY OPERATION CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202010911838.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2020-09-02
Publication Date
2025-06-24
Estimated Expiration
2040-09-02

AI Technical Summary

Technical Problem

When existing neural networks sparsely process feature maps, they can easily lead to loss of key information, thereby reducing the accuracy of the result.

Method used

During the neural network training process, the parameters of the neural network are optimized to balance the accuracy and sparseness by comprehensively considering the difference between the predicted label and the sample label (loss value) and the sparsity of the feature map (sparseness value).

Benefits of technology

It is achieved to improve the sparsity of the feature map while ensuring the accuracy of the result, thereby reducing bandwidth and calculation amount, without manually inputting sparse parameters or adding additional sparse units.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114202052B_ABST
    Figure CN114202052B_ABST
Patent Text Reader

Abstract

The present invention provides a training method, device, equipment, and storage medium for a neural network, which can better ensure the accuracy of the result and the sparsity of the feature map. The method includes: inputting image sample data into a first neural network to obtain a predicted label output by the first neural network and a feature map output by at least one specified processing layer of the first neural network. The feature map is output during the process of the first neural network determining the predicted label, and the feature map contains multiple feature values; determining a loss value according to the predicted label and the sample label corresponding to the image sample data by using a preset loss function, where the loss value is used to characterize the difference between the predicted label and the sample label; determining a sparsity value used to characterize the sparsity of the feature map according to the feature map; determining an optimization parameter value for optimizing the first neural network according to the loss value and the sparsity value; and optimizing the first neural network according to the optimization parameter value to obtain a second neural network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and particularly relates to a method, device and equipment for training a neural network. Background Art

[0002] An artificial neural network (ANN) is a research hotspot in the field of artificial intelligence. It abstracts the human brain neuron network from the perspective of information processing, establishes a simple model, and forms different networks according to different connection methods. In the engineering and academic circles, it is often simply referred to as a neural network or a neural-like network. Neural networks are usually deployed on chips and can be used to implement image processing, including object detection, image classification, image recognition, etc. When the chip calls the neural network to implement image processing, it requires a large bandwidth and a relatively large amount of computation, which is the main bottleneck for the chip.

[0003] Generally speaking, sparse processing of the feature map through a neural network can effectively reduce the bandwidth of the chip and also save the amount of computation. For example, after obtaining a feature map (the size of the feature map is, for example, c*w*h, where c is the number of channels, w is the width, h is the height, and a feature value is represented by 32 bits, so the data volume is, for example, c*w*h*32) in a certain processing layer of the neural network, the feature map needs to be stored in a specified cache, and the next processing layer reads the feature map from the specified cache. The amount of data required to be transmitted by the chip bus for one storage and one retrieval is 2*c*w*h*32, and the required bandwidth is relatively large. Through sparse processing, some feature values in the feature map to be stored are set to 0. By adopting a compressed storage method, the feature values with a size of 0 in the feature map do not need to be stored and thus do not occupy bandwidth, which can reduce the bandwidth.

[0004] In related methods, the neural network usually sparsifies the feature map according to artificially input sparse parameters. The sparse methods are, for example: 1) The sparse parameter is a sparse threshold, and the feature values in the feature map that are less than the sparse threshold are set to 0; 2) The sparse parameter is a sparsity rate, and the feature values in the feature map are sorted in descending order, and the feature values at the tail are set to 0 according to the sparsity rate, so as to reduce the amount of data to be stored in the feature map, and thus reduce the bandwidth required for data access.

[0005] In this method, sparse processing of the feature map using artificially input sparse parameters often causes the loss of key information in the feature map. For example, in the case of object detection scenarios, target information will be lost, resulting in a decrease in the accuracy of the results. Summary of the Invention

[0006] In view of this, the present invention provides a training method, apparatus, device and storage medium for a neural network, which can better ensure the accuracy of the result and the sparsity of the feature map.

[0007] The first aspect of the present invention provides a training method for a neural network, including:

[0008] Inputting image sample data into a first neural network to obtain a predicted label output by the first neural network and a feature map output by at least one specified processing layer of the first neural network, where the feature map is output during the process of the first neural network determining the predicted label, and the feature map includes multiple feature values;

[0009] Determining a loss value according to the predicted label and the sample label corresponding to the image sample data and using a preset loss function, where the loss value is used to characterize the difference between the predicted label and the sample label;

[0010] Determining a sparsity value for characterizing the sparsity of the feature map according to the feature map;

[0011] Determining an optimization parameter value for optimizing the first neural network according to the loss value and the sparsity value;

[0012] Optimizing the first neural network according to the optimization parameter value to obtain a second neural network.

[0013] According to an embodiment of the present invention, the determining a sparsity value for characterizing the sparsity of the feature map according to the feature map includes:

[0014] Modifying each feature value in the feature map into a corresponding transformed value according to a set transformation method; wherein, in the set transformation method, the larger the absolute value of the feature value, the larger the corresponding transformed value, and when the feature value is 0, the corresponding transformed value is a set minimum value;

[0015] Determining the sparsity value according to the transformed values in the feature map.

[0016] According to an embodiment of the present invention, modifying each feature value in the feature map into a corresponding transformed value according to a set transformation method includes:

[0017] For each feature value in the feature map, determining the absolute value of the feature value as the corresponding transformed value; or,

[0018] For each feature value in the feature map, determining the Nth power of the feature value as the corresponding transformed value, where N is a positive even number; or,

[0019] For each feature value in the feature map, determining the value with a specified constant as the base and the absolute value of the feature value or the Nth power of the feature value as the power as the corresponding transformed value; or,

[0020] For each eigenvalue in the feature map, the M-th power of the absolute value of the eigenvalue is determined as the corresponding transformation value, where M is a positive integer.

[0021] According to an embodiment of the present invention,

[0022] When the number of the specified processing layers is 1, the number of the feature maps is also 1. Determining the optimization parameter value for optimizing the first neural network based on the loss value and the sparsity value includes:

[0023] Calculating the product of the sparsity value and a first set hyperparameter to obtain a first result, and determining the sum of the loss value and the first result as the optimization parameter value;

[0024] Or,

[0025] When the number of the specified processing layers is greater than 1, the number of the feature maps is also greater than 1. Determining the optimization parameter value for optimizing the first neural network based on the loss value and the sparsity value includes:

[0026] For each feature map, calculating the product of the sparsity value corresponding to the feature map and a second set hyperparameter corresponding to the specified processing layer to obtain a second result corresponding to the feature map, where the smaller the size of the feature map output by the specified processing layer, the larger the corresponding second set hyperparameter;

[0027] Determining the sum of the loss value and the second results corresponding to the feature maps as the optimization parameter value.

[0028] According to an embodiment of the present invention, optimizing the first neural network based on the optimization parameter value to obtain a second neural network includes:

[0029] Determining the network parameter update amount according to the set backpropagation algorithm and the optimization parameter value, and optimizing the network parameters in the first neural network based on the network parameter update amount;

[0030] When the current set training end condition is satisfied, determining the optimized first neural network as the second neural network.

[0031] A second aspect of the present invention provides a training device for a neural network, including:

[0032] An image sample data input module, configured to input image sample data into the first neural network to obtain a predicted label output by the first neural network and feature maps output by at least one specified processing layer of the first neural network, where the feature maps are output during the process of the first neural network determining the predicted label, and the feature maps include multiple eigenvalues;

[0033] A loss value determination module, configured to determine a loss value according to the predicted label and the sample label corresponding to the image sample data by using a preset loss function, where the loss value is used to characterize the difference between the predicted label and the sample label;

[0034] A sparsity value determination module, configured to determine a sparsity value for characterizing the sparsity of the feature map according to the feature map;

[0035] An optimization parameter value determination module, configured to determine an optimization parameter value for optimizing the first neural network according to the loss value and the sparsity value;

[0036] A neural network optimization module, configured to optimize the first neural network according to the optimization parameter value to obtain a second neural network.

[0037] According to an embodiment of the present invention, when the sparsity value determination module determines a sparsity value for characterizing the sparsity of the feature map according to the feature map, it is specifically configured to:

[0038] Modify each feature value in the feature map to a corresponding transformation value according to a set transformation method; wherein, in the set transformation method, the larger the absolute value of the feature value, the larger the corresponding transformation value, and when the feature value is 0, the corresponding transformation value is a set minimum value;

[0039] Determine the sparsity value according to the transformation values in the feature map.

[0040] According to an embodiment of the present invention, when the sparsity value determination module modifies each feature value in the feature map to a corresponding transformation value according to a set transformation method, it is specifically configured to:

[0041] For each feature value in the feature map, determine the absolute value of the feature value as the corresponding transformation value; or,

[0042] For each feature value in the feature map, determine the Nth power of the feature value as the corresponding transformation value, where N is a positive even number; or,

[0043] For each feature value in the feature map, determine the value with a specified constant as the base and the absolute value of the feature value or the Nth power of the feature value as the power as the corresponding transformation value; or,

[0044] For each feature value in the feature map, determine the Mth power of the absolute value of the feature value as the corresponding transformation value, where M is a positive integer.

[0045] According to an embodiment of the present invention,

[0046] When the number of the specified processing layers is 1, the number of the feature maps is also 1. When the optimization parameter value determination module determines the optimization parameter value for optimizing the first neural network based on the loss value and the sparsity value, it is specifically configured to:

[0047] Calculate the product of the sparsity value and a first set hyperparameter to obtain a first result, and determine the sum of the loss value and the first result as the optimization parameter value;

[0048] Or,

[0049] When the number of the specified processing layers is greater than 1, the number of the feature maps is also greater than 1. When the optimization parameter value determination module determines the optimization parameter value for optimizing the first neural network based on the loss value and the sparsity value, it is specifically configured to:

[0050] For each feature map, calculate the product of the sparsity value corresponding to the feature map and a second set hyperparameter corresponding to the specified processing layer to obtain a second result corresponding to the feature map, where the smaller the size of the feature map output by the specified processing layer, the larger the corresponding second set hyperparameter;

[0051] Determine the sum of the loss value and the second results corresponding to the respective feature maps as the optimization parameter value.

[0052] According to an embodiment of the present invention, when the neural network optimization module optimizes the first neural network based on the optimization parameter value to obtain a second neural network, it is specifically configured to:

[0053] Determine a network parameter update amount according to a set backpropagation algorithm and the optimization parameter value, and optimize the network parameters in the first neural network based on the network parameter update amount;

[0054] When the current satisfies the set training end condition, determine the optimized first neural network as the second neural network.

[0055] A third aspect of the present invention provides an electronic device, including a processor and a memory; the memory stores a program that can be called by the processor; wherein, when the processor executes the program, the neural network training method described in the foregoing embodiment is implemented.

[0056] A fourth aspect of the present invention provides a machine-readable storage medium, on which a program is stored, and when the program is executed by a processor, the neural network training method described in the foregoing embodiment is implemented.

[0057] The embodiments of the present invention have the following beneficial effects:

[0058] In the embodiments of the present invention, when calculating the optimization parameter values for optimizing the first neural network, not only the loss value determined based on the difference between the predicted labels output by the first neural network and the sample labels corresponding to the image sample data is considered, but also the sparsity value determined based on the feature maps output by at least one specified processing layer in the first neural network during the process of determining the predicted labels is considered. The optimization parameter values comprehensively consider the loss value and the sparsity value, and can comprehensively reflect the difference between the predicted labels and the sample labels, as well as the sparsity of the feature maps. Then, when optimizing the first neural network based on the optimization parameter values, in addition to the constraint on accuracy, the constraint on the sparsity of the feature maps is also added. That is to say, not only high accuracy is desired, but also it is desired to be as sparse as possible. In this way, a second neural network with balanced accuracy and sparsity can be obtained, that is, the sparsity and accuracy can be guaranteed by the training process. When using the second neural network to perform corresponding tasks, the accuracy of the results and the sparsity of the feature maps can be better guaranteed. Moreover, there is no need to manually input sparse parameters, nor to add additional sparse units, and it can be applied to any neural network that needs to achieve sparse feature maps. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 is a schematic flowchart of a method for training a neural network according to an embodiment of the present invention;

[0060] Figure 2 is a structural block diagram of a first neural network according to an embodiment of the present invention;

[0061] Figure 3 is a schematic diagram of a feature map according to an embodiment of the present invention;

[0062] Figure 4 is a structural block diagram of a training device for a neural network according to an embodiment of the present invention;

[0063] Figure 5 is a structural block diagram of an electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0064] Here, exemplary embodiments will be described in detail, and examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present invention as detailed in the appended claims.

[0065] The terms used in this invention are for the purpose of describing specific embodiments only and are not intended to limit the invention. The singular forms "a", "the", and "said" used in this invention and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.

[0066] It should be understood that although the terms first, second, third, etc. may be used in this invention to describe various devices, such information should not be limited to these terms. These terms are only used to distinguish devices of the same type from each other. For example, without departing from the scope of this invention, the first device may also be referred to as the second device, and similarly, the second device may also be referred to as the first device. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to a determination".

[0067] To make the description of this invention clearer and more concise, some technical terms in this invention are explained below:

[0068] Bandwidth: In a chip bus system, the total amount of data transmission allowed per unit time.

[0069] The training method of the neural network according to the embodiments of this invention can be applied to various neural network application scenarios, such as object detection, object recognition, object classification and other scenarios. Specifically, for example, there are traffic scenarios that require vehicle detection or license plate recognition, and access control scenarios that require face recognition, fingerprint recognition, etc. It can be understood that the scenarios here are not restrictive, and any scenario that can apply to a neural network is applicable.

[0070] The training method of the neural network according to the embodiments of this invention is described in more detail below, but should not be limited thereto. In one embodiment, see Figure 1 , a training method of a neural network, the method may include the following steps:

[0071] S100: Input the image sample data into the first neural network to obtain the predicted label output by the first neural network and the feature map output by at least one specified processing layer of the first neural network. The feature map is output during the process of the first neural network determining the predicted label, and the feature map contains multiple feature values;

[0072] S200: Determine the loss value according to the predicted label and the sample label corresponding to the image sample data and using a preset loss function. The loss value is used to characterize the difference between the predicted label and the sample label;

[0073] S300: Determine a sparsity value for characterizing the sparsity of the feature map based on the feature map;

[0074] S400: Determine an optimization parameter value for optimizing the first neural network based on the loss value and the sparsity value;

[0075] S500: Optimize the first neural network based on the optimization parameter value to obtain a second neural network.

[0076] In an embodiment of the present invention, the execution subject of the training method of the neural network is an electronic device. The electronic device may be a computer device, a mobile terminal, an imaging device, etc. The specific type is not limited, as long as it is a device with a certain processing capacity.

[0077] In step S100, input the image sample data into the first neural network to obtain a predicted label output by the first neural network and a feature map output by at least one specified processing layer of the first neural network. The feature map is output during the process of the first neural network determining the predicted label, and the feature map includes multiple feature values.

[0078] An image sample data set may be established in advance. The image sample data set includes multiple image sample data, and each image sample data has been calibrated with a corresponding sample label. The sample label may be manually calibrated. Of course, the calibration method is not limited to this, and other calibration methods may also be used, such as part being manually calibrated and part being calibrated by a neural network.

[0079] The image sample data may be an image. For example, it may be an image collected from an application scenario. Of course, it may also be an image obtained through other means, and the specific is not limited. The sample label may be determined according to the task that the neural network needs to perform. For example, when the task is object detection, the sample label may include the category and location information of the object in the image sample data; when the task is object classification, the sample label may include the category of the object in the image sample data; when the task is semantic segmentation, the sample label may include the contour information of the object in the image sample data, and so on.

[0080] The object here may include a person, a fingerprint, a face, a motor vehicle, a non-motor vehicle, a license plate, text, a ship, and / or a flying object, etc. The specific is not limited and may be determined according to the application scenario.

[0081] During training, one or more image sample data may be selected from the image sample data set, and the selected image sample data is input into the first neural network to obtain a predicted label output by the first neural network and a feature map output by at least one specified processing layer of the first neural network.

[0082] When the first neural network is used for object detection, the predicted label can be the predicted category and predicted location information of the object; or, when the first neural network is used for object classification, the predicted label can be the predicted category of the object; or, when the first neural network is used for semantic segmentation, the predicted label is the predicted contour information of the object. Of course, the tasks that the first neural network is used to perform are not limited to this, and can be determined according to actual needs.

[0083] Here, the first neural network can be an established initial model, or it can be a neural network in which the initial model has been trained to a certain extent but does not yet meet the requirements and needs to be further trained, and the specific details are not limited.

[0084] The neural network architectures used to perform different tasks can be different. Taking object detection as an example, the neural network can adopt architectures such as Faster-RCNN (a deep learning-based object detection technology), YOLO (You Only Look Once, which uses a single CNN model to achieve end-to-end object detection), SSD (single shot multibox detector, an object detection algorithm that directly predicts the coordinates and categories of object boxes), etc., and the specific details are not limited to this.

[0085] The first neural network can include multiple processing layers. The types of processing layers are not limited, as long as all processing layers cooperate to perform the corresponding tasks. The specified processing layer can be a processing layer other than the processing layer (i.e., the last layer) used to output the predicted label in the neural network. Specifically, it can be specified according to the actual situation. For example, it can be specified that during the process of determining the predicted label, the processing layer that needs to store the output feature map into the specified cache is used as the specified processing layer, and there are certain sparsity requirements for the output feature map.

[0086] There can be one or more specified processing layers in the first neural network. For example, if the first neural network includes 5 processing layers, the 1st, 2nd, 3rd, and / or 4th processing layers in the neural network can be used as the specified processing layer. The specified processing layer can be, for example, a convolutional layer in the first neural network. Of course, it can also be other layers such as a classification layer, a pooling layer, etc., and the specific type is not limited.

[0087] For example, referring to Figure 2 , the first neural network 200 can include:

[0088] The convolutional layer 201 is used to extract features from the input image sample data to obtain a feature map;

[0089] The convolutional layer 202 is used to determine multiple candidate regions in the feature map output by the convolutional layer 201;

[0090] The binary classification layer 203 is used to classify the multiple candidate regions determined by the convolutional layer 202 to determine whether the candidate region belongs to the foreground category or the background category, determine the candidate region belonging to the foreground category as the region of interest where the target is located, and output the position information of the region of interest after correction;

[0091] The pooling layer 204 is used to intercept the corresponding region of interest from the feature map output by the convolutional layer 201 according to the position information output by the binary classification layer 203, and downsample the region of interest (downsampling can make the size of the region of interest unified into a fixed size) to obtain and output the downsampled feature map;

[0092] At least one convolutional layer 205 is used to determine and output a feature vector based on the feature map output by the pooling layer 204;

[0093] The target classification layer 206 is used to determine and output the predicted category of the target based on the characteristic vector output by the convolutional layer 205.

[0094] Among them, the convolutional layer 201 and the pooling layer 204 are specified processing layers, and the predicted category output by the target classification layer 206 is the predicted label. The embodiment of the present invention can optimize the first neural network 200 based on the feature maps output by the convolutional layer 201 and the pooling layer 204 and the predicted category output by the target classification layer 206.

[0095] It can be understood that the architecture of the above first neural network 200 is only an example and is not a limitation.

[0096] In this embodiment, during training, not only the predicted label finally output by the first neural network is concerned, but also the feature map output by the specified processing layer of the first neural network is concerned, and the first neural network is optimized according to these two output information, rather than the usual practice of only optimizing the neural network according to the predicted label.

[0097] The feature map is output during the process of the first neural network determining the predicted label, and the feature map contains multiple feature values. Generally speaking, during processing, the feature map can be represented by a matrix, and each element in the matrix is a feature value. The feature values in the feature map can be positive, negative or 0.

[0098] In step S200, a loss value is determined according to the predicted label and the sample label corresponding to the image sample data by using a preset loss function, and the loss value is used to characterize the difference between the predicted label and the sample label.

[0099] The loss function can be preset according to the task performed by the first neural network. The loss functions corresponding to different tasks can be the same or different. For example, when the task is object classification, a loss function for determining the classification loss can be preset; when the task is semantic segmentation, a loss function for determining the segmentation loss can be preset; when the task is object detection, a loss function for determining the classification loss and the regression loss can be preset. Specifically, the loss function can be, for example, the cross-entropy loss function. Of course, this is just an example here, and in fact, it can be selected according to the task requirements.

[0100] When determining the loss value based on the predicted label and the sample label and using the preset loss function, the predicted label and the sample label can be substituted into the loss function to calculate the loss value. The loss value can represent the difference between the predicted label and the sample label. Generally speaking, the greater the difference, the greater the loss value. For example, in a classification task, the predicted label includes the predicted category of the object, and the sample label includes the true category of the object. Substituting the two into the loss function can determine the loss value, which represents the difference between the predicted category and the true category and also reflects the accuracy or precision of the neural network prediction.

[0101] In step S300, a sparsity value for characterizing the sparsity of the feature map is determined based on the feature map.

[0102] The sparsity can be represented by the number or proportion of feature values with a size of 0 in the feature map. The more feature values with a size of 0 (or the higher the proportion), the higher the sparsity, and vice versa.

[0103] For example, the feature map is as Figure 3 shown. In this feature map, the number of feature values with a size of 0 is 10, and the total number of feature values in the feature map is 16. Then the sparsity is the ratio of 10 to 16: 0.625. Of course, this is just an example for easy understanding of the sparsity, and the actual feature map is not limited to this.

[0104] The sparsity value can characterize the sparsity of the feature map. When determining the sparsity value based on the feature map, a set mathematical operation can be performed on the feature values in the feature map, and the resulting mathematical operation value is used as the sparsity value. Of course, the specific method is not limited.

[0105] In step S400, an optimization parameter value for optimizing the first neural network is determined based on the loss value and the sparsity value.

[0106] In this embodiment, different from the usual training method, the training is not directly based on the loss value, but also the corresponding sparsity value is determined based on the feature map. The sparsity value is used to characterize the sparsity of the feature map, and the optimization parameter value for training is determined based on the loss value and the sparsity value corresponding to the feature map.

[0107] When determining the optimization parameter value based on the loss value and the sparsity value, for example, the sum of the loss value and the sparsity value can be determined as the optimization parameter value, or the loss value and the sparsity value can be weighted and summed, or the value obtained by adjusting (such as adding or subtracting a set value) the above two summation results can be used as the optimization parameter value, etc. There is no specific limitation, as long as the optimization parameter value can comprehensively reflect the difference between the predicted label and the sample label, and the sparsity of the feature map.

[0108] In step S500, the first neural network is optimized based on the optimization parameter value to obtain a second neural network.

[0109] Since the loss value represents the difference between the predicted label and the sample label, and the sparsity value represents the sparsity of the feature map, the optimization parameter value can comprehensively reflect the difference between the predicted label and the sample label, and the sparsity of the feature map.

[0110] The difference between the predicted label and the sample label can reflect the accuracy of the first neural network. In terms of accuracy alone, it is expected that the features extracted by the neural network are as rich as possible, that is, the fewer the feature values of size 0 in the feature map, the higher the accuracy. And in terms of the sparsity of the feature map alone, the more the feature values of size 0 in the feature map, the greater the sparsity, and the less bandwidth required to access the feature map.

[0111] Therefore, the demands for the feature map from accuracy and sparsity are actually contradictory. In the embodiments of the present invention, the optimization parameter value is the comprehensive result of the loss value and the sparsity value, and thus comprehensively reflects the situations of accuracy and sparsity. By optimizing the first neural network with the optimization parameter value, the accuracy and sparsity of the first neural network can be balanced, and a second neural network with relatively better accuracy and sparsity can be obtained.

[0112] In the embodiments of the present invention, when calculating the optimization parameter values for optimizing the first neural network, not only the loss value determined based on the difference between the predicted labels output by the first neural network and the sample labels corresponding to the image sample data is considered, but also the sparsity value determined based on the feature maps output by at least one specified processing layer in the first neural network during the process of determining the predicted labels is considered. The optimization parameter values comprehensively consider the loss value and the sparsity value, and can comprehensively reflect the difference between the predicted labels and the sample labels, as well as the sparsity of the feature maps. Then, when optimizing the first neural network based on the optimization parameter values, in addition to the constraint on accuracy, a constraint on the sparsity of the feature maps is added. That is to say, not only high accuracy is desired, but also high sparsity is desired. In this way, a second neural network with balanced accuracy and sparsity can be obtained, that is, the sparsity and accuracy can be guaranteed by the training process. When using the second neural network to perform corresponding tasks, the accuracy of the result and the sparsity of the feature maps can be better guaranteed. Moreover, there is no need to manually input sparsity parameters, nor to add additional sparsity units, and it can be applied to any neural network that needs to achieve sparse feature maps.

[0113] The architecture of the second neural network is the same as that of the first neural network, except for the network parameters. In comparison, the network parameters of the second neural network are better than those of the first neural network, and this advantage is reflected in the accuracy and the sparsity of the feature maps.

[0114] In one example, when the second neural network performs a task, the specified processing layer can determine a feature map based on the data input to it, and under the constraint of the network parameters of the specified processing layer, the sparsity of the feature map output by the specified processing layer meets the specified requirements, such as the sparsity is less than the specified sparsity, etc.; or, the specified processing layer can check whether the sparsity of the feature map meets the specified requirements when obtaining the feature map. If not, the feature map is then sparsely processed according to the network parameters (such as the optimized sparsity parameters), and the sparsely processed feature map is output.

[0115] Optionally, the feature information of the feature values with non-zero magnitudes in the feature map can be stored in a specified storage format. Specifically, a 0-1 matrix can be generated. The size of the 0-1 matrix is the same as that of the feature map. 0 in the 0-1 matrix indicates that the magnitude of the feature value at the corresponding position in the feature map is 0, and 1 in the 0-1 matrix indicates that the magnitude of the feature value at the corresponding position in the feature map is non-zero. The 0-1 matrix is stored in a specified cache, and the feature values with non-zero magnitudes in the feature map are arranged in a specified order to obtain feature information and stored at the position corresponding to the 0-1 matrix in the specified cache.

[0116] Generally speaking, each eigenvalue in the feature map needs to be represented by multiple bits, such as 32 bits. In this example, only the non-zero eigenvalues in the feature map need to be stored. Assuming the size of the feature map is c*h*w, the data volume is c*h*w*32. The data transfer volume generated by one storage and one retrieval is 2*c*h*w*32. If the non-zero eigenvalues account for half, only c*h*w*16 of data volume needs to be stored in this example. At the same time, 0 or 1 in the 0-1 matrix only needs 1 bit to represent. Therefore, a matrix with a data volume of c*h*w*1 is stored again. The overall data volume is c*h*w*17, and the data volume is greatly reduced. The data transfer volume generated by one storage and one retrieval is 2*c*h*w*17, which is nearly half less than 2*c*h*w*32, thus greatly reducing the bandwidth required for transmission.

[0117] Of course, the specified storage format is not limited to this. For example, it can include: compressed column storage (CCS) or compressed row storage (CRS), and the specific is not limited.

[0118] When the specified recovery time arrives, the feature map can be restored according to the stored feature information to restore the feature map output by the specified processing layer. Specifically, the stored 0-1 matrix can be read out, and the stored eigenvalues are read in sequence according to the specified order, and the 1 in the 0-1 matrix is updated to the read eigenvalue in sequence.

[0119] In one embodiment, in step S300, the sparsity value used to characterize the sparsity of the feature map determined according to the feature map may include the following steps:

[0120] S301: Modify each eigenvalue in the feature map to the corresponding transformed value according to the set transformation method; wherein, in the set transformation method, the larger the absolute value of the eigenvalue, the larger the corresponding transformed value, and if the eigenvalue is 0, the corresponding transformed value is the set minimum value;

[0121] S302: Determine the sparsity value according to the transformed values in the feature map.

[0122] Optionally, the set minimum value may be greater than or equal to 0, and the specific is not limited.

[0123] Optionally, when determining the sparsity value according to the transformed values in the feature map, the sum of the transformed values in the feature map can be calculated, and this sum is determined as the sparsity value. Of course, this is only a preferred example here. The actual method of determining the sparsity value based on the transformed value is not limited to this, and it can also be other mathematical operation methods. For example, certain adjustments can also be made on the basis of this sum to obtain the sparsity value.

[0124] The larger the absolute value of the eigenvalue, the larger the corresponding transformation value. If the eigenvalue is 0, the corresponding transformation value is the set minimum value. Based on the above constraints, the more eigenvalues with a size of 0 in the feature map, the more transformation values that are the set minimum value, and the smaller the determined sparsity value. Therefore, the sparsity value can, to a certain extent, reflect the quantity of eigenvalues with a size of 0 in the feature map.

[0125] Ideally, the more eigenvalues with a size of 0, the better. So, according to the above operation method, the smaller the sparsity value, the better. For example, when the sparsity value approaches the minimum value such as 0, the sparsity is optimal. However, if all the eigenvalues in the feature map have a size of 0, the accuracy of the result output by the first neural network is very poor. Therefore, the first neural network does not allow this situation to occur, and the first neural network will make a balance between accuracy and sparsity during training.

[0126] In the case where the smaller the loss value, the smaller the difference between the predicted label and the sample label, and the smaller the sparsity value, the more eigenvalues with a size of 0, through training, the optimization parameter value can be gradually reduced to approach a certain ideal value, that is, during the training process, the predicted label output by the neural network and the feature map can make the optimization parameter value gradually approach the ideal value.

[0127] In one embodiment, in step S301, modifying each eigenvalue in the feature map to the corresponding transformation value according to the set transformation method includes:

[0128] For each eigenvalue in the feature map, determining the absolute value of the eigenvalue as the corresponding transformation value; or,

[0129] For each eigenvalue in the feature map, determining the Nth power of the eigenvalue as the corresponding transformation value, where N is a positive even number; or,

[0130] For each eigenvalue in the feature map, determining the value with a specified constant as the base and the absolute value of the eigenvalue or the Nth power of the eigenvalue as the power as the corresponding transformation value; or,

[0131] For each eigenvalue in the feature map, determining the Mth power of the absolute value of the eigenvalue as the corresponding transformation value, where M is a positive integer.

[0132] The specified constant here is, for example, the natural constant e (an infinite non - repeating decimal and a transcendental number, whose value is approximately 2.71828), and it is not specifically limited.

[0133] It can be understood that the above - mentioned ways of determining the transformation value are only examples. In fact, there can be other ways as long as the above - mentioned constraints are met.

[0134] In one embodiment, when the number of the specified processing layers is 1, the number of the feature maps is also 1. In step S400, determining the optimization parameter value for optimizing the first neural network according to the loss value and the sparsity value includes:

[0135] Calculating the product of the sparsity value and a first set hyperparameter to obtain a first result, and determining the sum of the loss value and the first result as the optimization parameter value.

[0136] In other words, the calculation formula of the optimization parameter value Loss can be as follows:

[0137] Loss = L old + λ||activation||

[0138] Wherein, L old is the loss value, λ is the first set hyperparameter, and ||activation|| is the sparsity value.

[0139] In another embodiment, when the number of the specified processing layers is greater than 1, the number of the feature maps is also greater than 1. In step S400, determining the optimization parameter value for optimizing the first neural network according to the loss value and the sparsity value includes:

[0140] For each feature map, calculating the product of the sparsity value corresponding to the feature map and a second set hyperparameter corresponding to the specified processing layer, wherein the smaller the size of the feature map output by the specified processing layer, the larger the corresponding second set hyperparameter;

[0141] Determining the sum of the loss value and the second results corresponding to the feature maps as the optimization parameter value.

[0142] In other words, the calculation formula of the optimization parameter value Loss is as follows:

[0143]

[0144] Wherein, L old is the loss value, λ i is the second set hyperparameter corresponding to the feature map output by the i-th specified processing layer, and ||activation i || is the sparsity value corresponding to the feature map output by the i-th specified processing layer.

[0145] When the number of specified processing layers is greater than 1, each specified processing layer outputs a feature map. Moreover, since the depths of the specified processing layers in the neural network are different (the deeper it is, the closer it is to the output end), the sizes of the feature maps output by different specified processing layers are generally different. Usually, the deeper the depth, the smaller the size of the feature map, and the closer the expression is to the result. The first neural network itself does not want the feature map to be sparse, so it is also relatively difficult to be sparse.

[0146] Therefore, in this embodiment, the second set of hyperparameters corresponding to each feature map can be determined according to the size to adjust the influence of different sparsity values on the optimization parameter values. Among them, the smaller the size of the feature map output by the specified processing layer, the larger the corresponding second set of hyperparameters. Such a setting can be more conducive to the sparsity of the feature map with a smaller size.

[0147] In one embodiment, in step S500, optimizing the first neural network according to the optimization parameter values to obtain a second neural network includes:

[0148] Determining the network parameter update amount according to the set backpropagation algorithm and the optimization parameter values, and optimizing the network parameters in the first neural network according to the network parameter update amount;

[0149] When the set training end condition is currently satisfied, the optimized first neural network is determined as the second neural network.

[0150] The backpropagation algorithm (i.e., the BP algorithm) is a learning algorithm suitable for multi-layer neuron networks. It is based on the gradient descent method and mainly consists of two links (activation propagation, weight update) that are repeatedly iterated until the response of the network to the input reaches a predetermined target range.

[0151] In this embodiment, the response of the network to the input is reflected by the optimization parameter values. The network parameter update amounts of each layer in the first neural network are determined layer by layer as the input data of the backpropagation algorithm, and the network parameters in the first neural network are optimized according to the network parameter update amounts. For example, the sum of the current network parameters of each layer and the network parameter update amount is used as the new network parameter.

[0152] There can be various set training end conditions. For example, in one example, the performance of the optimized first neural network in performing tasks is tested. If the performance meets the set requirements, it means that the current training end condition is satisfied and the iteration can be ended; in another example, if there is no unselected image sample data in the image sample dataset, it means that the current training end condition is satisfied and the iteration can be ended; in another example, a training times threshold can be set, the current training times are calculated, and when the current training times reach the training times threshold, it means that the current training end condition is satisfied and the training can be ended.

[0153] The above several examples are not restrictive. As long as the currently set training end condition is met, the training can be ended, and the optimized first neural network is determined as the second neural network.

[0154] Optionally, when the currently set training end condition is not met, another image sample data in the image sample dataset can be selected, and the step of inputting the image sample data into the first neural network is returned.

[0155] The present invention also provides a training device for a neural network. Refer to Figure 4 , the training device 100 of the neural network includes:

[0156] An image sample data input module 101, configured to input image sample data into the first neural network to obtain a predicted label output by the first neural network and a feature map output by at least one specified processing layer of the first neural network. The feature map is output during the process of the first neural network determining the predicted label, and the feature map includes multiple feature values;

[0157] A loss value determination module 102, configured to determine a loss value according to the predicted label and the sample label corresponding to the image sample data by using a preset loss function. The loss value is used to characterize the difference between the predicted label and the sample label;

[0158] A sparsity value determination module 103, configured to determine a sparsity value for characterizing the sparsity of the feature map according to the feature map;

[0159] An optimization parameter value determination module 104, configured to determine an optimization parameter value for optimizing the first neural network according to the loss value and the sparsity value;

[0160] A neural network optimization module 105, configured to optimize the first neural network according to the optimization parameter value to obtain a second neural network.

[0161] In one embodiment, when the sparsity value determination module determines the sparsity value for characterizing the sparsity of the feature map according to the feature map, it is specifically configured to:

[0162] Modify each feature value in the feature map into a corresponding transformation value according to a set transformation method; wherein, in the set transformation method, the larger the absolute value of the feature value, the larger the corresponding transformation value, and when the feature value is 0, the corresponding transformation value is a set minimum value;

[0163] Determine the sparsity value according to the transformation values in the feature map.

[0164] In one embodiment, when the sparsity value determination module modifies each eigenvalue in the feature map to a corresponding transformed value according to a set transformation method, it is specifically used for:

[0165] For each eigenvalue in the feature map, determining the absolute value of the eigenvalue as the corresponding transformed value; or,

[0166] For each eigenvalue in the feature map, determining the Nth power of the eigenvalue as the corresponding transformed value, where N is a positive even number; or,

[0167] For each eigenvalue in the feature map, determining the value with a specified constant as the base and the absolute value of the eigenvalue or the Nth power of the eigenvalue as the exponent as the corresponding transformed value; or,

[0168] For each eigenvalue in the feature map, determining the Mth power of the absolute value of the eigenvalue as the corresponding transformed value, where M is a positive integer.

[0169] In one embodiment,

[0170] When the number of the specified processing layers is 1, the number of the feature maps is also 1. When the optimization parameter value determination module determines the optimization parameter value for optimizing the first neural network based on the loss value and the sparsity value, it is specifically used for:

[0171] Calculating the product of the sparsity value and a first set hyperparameter to obtain a first result, and determining the sum of the loss value and the first result as the optimization parameter value;

[0172] Or,

[0173] When the number of the specified processing layers is greater than 1, the number of the feature maps is also greater than 1. When the optimization parameter value determination module determines the optimization parameter value for optimizing the first neural network based on the loss value and the sparsity value, it is specifically used for:

[0174] For each feature map, calculating the product of the sparsity value corresponding to the feature map and a second set hyperparameter corresponding to the specified processing layer to obtain a second result corresponding to the feature map, where the smaller the size of the feature map output by the specified processing layer, the larger the corresponding second set hyperparameter;

[0175] Determining the sum of the loss value and the second results corresponding to each feature map as the optimization parameter value.

[0176] In one embodiment, when the neural network optimization module optimizes the first neural network based on the optimization parameter value to obtain a second neural network, it is specifically used for:

[0177] Determine the network parameter update amount according to the set backpropagation algorithm and the optimized parameter values, and optimize the network parameters in the first neural network according to the network parameter update amount;

[0178] When the set training end condition is currently satisfied, determine the optimized first neural network as the second neural network.

[0179] The implementation processes of the functions and roles of each unit in the above device are specifically described in the implementation processes of the corresponding steps in the above method, and will not be elaborated here.

[0180] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units.

[0181] The present invention also provides an electronic device, including a processor and a memory; the memory stores a program that can be called by the processor; wherein, when the processor executes the program, the neural network training method described in the foregoing embodiments is implemented.

[0182] The embodiment of the neural network training device of the present invention can be applied to an electronic device. Taking software implementation as an example, as a logically meaningful device, it is formed by the processor of the electronic device where it is located reading the corresponding computer program instructions in the non-volatile memory into the specified cache for running. From the hardware level, as Figure 5 shown, Figure 5 FIG. 19 is a hardware structure diagram of an electronic device where the neural network training device 100 of the present invention is located according to an exemplary embodiment. In addition to Figure 5 the processor 510, memory 530, network interface 520, and non-volatile memory 540 shown in FIG. 19, the electronic device where the neural network training device 100 is located in the embodiment usually includes other hardware according to the actual functions of the electronic device, which will not be elaborated here.

[0183] The present invention also provides a machine-readable storage medium, on which a program is stored, and when the program is executed by a processor, the neural network training method described in the foregoing embodiments is implemented.

[0184] The present invention may be embodied in the form of a computer program product implemented on one or more storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain program code. Machine-readable storage media include permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information may be computer-readable instructions, data structures, program modules, or other data. Examples of machine-readable storage media include but are not limited to: phase change designated cache (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other designated cache technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information that can be accessed by a computing device.

[0185] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A training method for a neural network, characterized in that, Including: Inputting image sample data into a first neural network to obtain a predicted label output by the first neural network and a feature map output by at least one specified processing layer of the first neural network. The feature map is output during the process of the first neural network determining the predicted label, and the feature map contains multiple feature values. If the first neural network is used for object detection, the sample label includes the category and location information of the object in the image sample data. If the first neural network is used for object classification, the sample label includes the category of the object in the image sample data. If the first neural network is used for object recognition, the sample label includes the contour information of the object in the image sample data. Determining a loss value according to the predicted label and the sample label corresponding to the image sample data by using a preset loss function. The loss value is used to characterize the difference between the predicted label and the sample label. Modifying each feature value in the feature map to a corresponding transformed value according to a set transformation method. In the set transformation method, the larger the absolute value of the feature value, the larger the corresponding transformed value, and when the feature value is 0, the corresponding transformed value is the set minimum value. Determining a sparsity value for characterizing the sparsity of the feature map according to the transformed values in the feature map. Determining an optimization parameter value for optimizing the first neural network according to the loss value and the sparsity value. The loss value is used to constrain the accuracy during the training process, the sparsity is used to constrain the sparsity during the training process, and the optimization parameter value is used to ensure the balance of accuracy and sparsity during the training process. Optimizing the first neural network according to the optimization parameter value to obtain a second neural network. The second neural network is used to perform at least one of the following image processing operations: object detection, object classification, and object recognition.

2. The training method of the neural network according to claim 1, wherein Modifying each feature value in the feature map to a corresponding transformed value according to a set transformation method includes: For each feature value in the feature map, determining the absolute value of the feature value as the corresponding transformed value; or For each feature value in the feature map, determining the Nth power of the feature value as the corresponding transformed value, where N is a positive even number; or For each feature value in the feature map, determining the value with a specified constant as the base and the absolute value of the feature value or the Nth power of the feature value as the exponent as the corresponding transformed value; or For each feature value in the feature map, determining the Mth power of the absolute value of the feature value as the corresponding transformed value, where M is a positive integer.

3. The training method of the neural network according to any one of claims 1-2, characterized in that When the number of the specified processing layers is 1, the number of the feature maps is also 1. Determining the optimization parameter value for optimizing the first neural network according to the loss value and the sparsity value includes: Calculating the product of the sparsity value and a first set hyperparameter to obtain a first result, and determining the sum of the loss value and the first result as the optimization parameter value; Or When the number of the specified processing layers is greater than 1, the number of the feature maps is also greater than 1. Determining the optimization parameter value for optimizing the first neural network according to the loss value and the sparsity value includes: For each feature map, calculating the product of the sparsity value corresponding to the feature map and the second set hyperparameter corresponding to the specified processing layer to obtain a second result corresponding to the feature map, wherein the smaller the size of the feature map output by the specified processing layer, the larger the corresponding second set hyperparameter; Determining the sum of the loss value and the second results corresponding to the feature maps as the optimization parameter value.

4. The training method of the neural network according to claim 1, wherein Optimizing the first neural network according to the optimization parameter value to obtain a second neural network, including: Determining the network parameter update amount according to the set backpropagation algorithm and the optimization parameter value, and optimizing the network parameters in the first neural network according to the network parameter update amount; When the set training end condition is currently satisfied, determining the optimized first neural network as the second neural network.

5. A training device for a neural network, characterized in that, Including: An image sample data input module, configured to input image sample data into the first neural network to obtain a predicted label output by the first neural network and feature maps output by at least one specified processing layer of the first neural network, where the feature maps are output during the process of the first neural network determining the predicted label, and the feature maps include multiple feature values; if the first neural network is used for object detection, the sample label includes the category and location information of the object in the image sample data; if the first neural network is used for object classification, the sample label includes the category of the object in the image sample data; if the first neural network is used for object recognition, the sample label includes the contour information of the object in the image sample data; A loss value determination module, configured to determine a loss value according to the predicted label and the sample label corresponding to the image sample data by using a preset loss function, where the loss value is used to represent the difference between the predicted label and the sample label; A sparsity value determination module, configured to modify each feature value in the feature map into a corresponding transformed value according to a set transformation method; wherein, in the set transformation method, the larger the absolute value of the feature value, the larger the corresponding transformed value, and if the feature value is 0, the corresponding transformed value is the set minimum value; Determining a sparsity value for representing the sparsity of the feature map according to the transformed values in the feature map; An optimization parameter value determination module, configured to determine an optimization parameter value for optimizing the first neural network according to the loss value and the sparsity value; the loss value is used to constrain the accuracy during training, the sparsity is used to constrain the sparsity during training, and the optimization parameter value is used to ensure the balance between accuracy and sparsity during training; A neural network optimization module, configured to optimize the first neural network according to the optimization parameter value to obtain a second neural network; the second neural network is used to perform at least one of the following image processing tasks: object detection, object classification, and object recognition; wherein, the second neural network can balance accuracy and sparsity when performing tasks.

6. The training device of the neural network according to claim 5, wherein, When the sparsity value determination module modifies each eigenvalue in the feature map to a corresponding transformed value according to a set transformation method, it specifically is used for: For each eigenvalue in the feature map, determining the absolute value of the eigenvalue as the corresponding transformed value; or, For each eigenvalue in the feature map, determining the Nth power of the eigenvalue as the corresponding transformed value, where N is a positive even number; or, For each eigenvalue in the feature map, determining the value with a specified constant as the base and the absolute value of the eigenvalue or the Nth power of the eigenvalue as the exponent as the corresponding transformed value; or, For each eigenvalue in the feature map, determining the Mth power of the absolute value of the eigenvalue as the corresponding transformed value, where M is a positive integer.

7. The training device of the neural network according to any one of claims 5-6, characterized in that When the number of the specified processing layers is 1, the number of the feature maps is also 1. When the optimization parameter value determination module determines the optimization parameter value for optimizing the first neural network based on the loss value and the sparsity value, it specifically is used for: Calculating the product of the sparsity value and a first set hyperparameter to obtain a first result, and determining the sum of the loss value and the first result as the optimization parameter value; Or, When the number of the specified processing layers is greater than 1, the number of the feature maps is also greater than 1. When the optimization parameter value determination module determines the optimization parameter value for optimizing the first neural network based on the loss value and the sparsity value, it specifically is used for: For each feature map, calculating the product of the sparsity value corresponding to the feature map and a second set hyperparameter corresponding to the specified processing layer to obtain a second result corresponding to the feature map, where the smaller the size of the feature map output by the specified processing layer, the larger the corresponding second set hyperparameter; Determining the sum of the loss value and the second results corresponding to each feature map as the optimization parameter value.

8. An electronic device, characterized in that, Comprising a processor and a memory; the memory stores a program that can be called by the processor; wherein, when the processor executes the program, it implements the neural network training method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Convolutional neural network pruning method based on feature map sparsification

    CN110874631A