Training Method of Image Classification Model, Image Classification Method and Device
By using multiple rounds of training and sparse weight matrix cutting technology during the training process of image classification model, the problem of high computing and storage resources in the prior art is solved, and a faster training process and lower computing time is achieved.
Patent Information
- Application Number
- CN202210536970.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-17
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-05-17
AI Technical Summary
The prior art has problems of high computing and storage resources consumption in the training process of image classification models, and the structured sparse method fails to effectively utilize sparse characteristics in the pre-training stage, and the acceleration of the inference stage requires specific hardware support.
A training method for image classification model is proposed. Through the multi-round training process, the sparse weight matrix is used to perform forward calculation, the loss function is calculated, the dense weight matrix is calculated in the reverse gradient, and it is cropped to obtain the sparse weight matrix for the next round, until the value of the loss function is less than the threshold.
By automatically sparse the weight matrix during the training process of the image classification model, the consumption of calculation and storage resources is reduced, the calculation speed during the training process is accelerated, and the calculation time is reduced.
Smart Images

Figure CN114998649B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, specifically to technical fields such as deep learning and computer vision, and in particular to a training method for an image classification model, an image classification method and a device. Background Art
[0002] With the development of computer technology, image classification technology has emerged. Image classification is a basic task in the field of computer vision. Image classification technology can quickly identify the category to which each object in the image belongs.
[0003] In order to realize automatic classification of objects in an image, when an image classification model is constructed on an existing sample image, the image classification model needs to be trained, so as to classify the image to be classified based on the trained image classification model.
[0004] In order to improve the prediction effect of the model, it is very important to train the image classification model. Summary of the invention
[0005] The present disclosure provides a training method for an image classification model, an image classification method and a device.
[0006] According to one aspect of the present disclosure, a training method for an image classification model is provided, comprising: acquiring a sample image; using the sample image to perform multiple rounds of training processes on the image classification model, wherein any round of training process comprises: performing forward calculation of the image classification model according to a set sparse weight matrix to obtain a predicted category of the sample image in this round; determining a loss function of this round according to a difference between the predicted category and a category annotation label on the sample image; performing reverse gradient calculation on the loss function of this round to obtain a dense weight matrix of this round; clipping the dense weight matrix of this round to obtain a sparse weight matrix for the next round of training process; and stopping the training process when the value of the loss function of any round is less than a threshold.
[0007] According to another aspect of the present disclosure, there is provided an image classification method, comprising: acquiring an image to be classified; and using an image classification model obtained by the training method of the image classification model described in the first aspect embodiment to perform category prediction on an object in the image to be classified, so as to obtain a predicted category of the image to be classified.
[0008] According to another aspect of the present disclosure, a training device for an image classification model is provided, comprising: an acquisition module for acquiring sample images; a training module for using the sample images to perform multiple rounds of training processes on the image classification model, wherein any round of the training process comprises: performing forward calculation of the image classification model according to a set sparse weight matrix to obtain a predicted category of the sample image in this round; determining a loss function of this round according to a difference between the predicted category and a category annotation label on the sample image; performing reverse gradient calculation on the loss function of this round to obtain a dense weight matrix of this round; clipping the dense weight matrix of this round to obtain a sparse weight matrix for the next round of training process; and a stop module for stopping the training process when the value of the loss function of any round is less than a threshold.
[0009] According to another aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method described in the embodiment of the first aspect of the present disclosure, or execute the method described in the embodiment of the second aspect of the present disclosure.
[0010] According to another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the method described in the embodiment of the first aspect of the present disclosure, or to execute the method described in the embodiment of the second aspect of the present disclosure.
[0011] According to another aspect of the present disclosure, a computer program product is provided. When the computer program is executed by a processor, the computer program implements the method described in the embodiment of the first aspect of the present disclosure, or executes the method described in the embodiment of the second aspect of the present disclosure.
[0012] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it intended to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The accompanying drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure.
[0014] Figure 1 A flowchart of a training method for an image classification model provided in the first embodiment of the present disclosure;
[0015] Figure 2 A flowchart of a training method for an image classification model provided in the second embodiment of the present disclosure;
[0016] Figure 3 A schematic diagram of converting a dense weight matrix provided in an embodiment of the present disclosure into a sparse weight matrix;
[0017] Figure 4 A flowchart of a training method for an image classification model provided in Embodiment 3 of the present disclosure;
[0018] Figure 5 A flowchart of a training method for an image classification model provided in Embodiment 4 of the present disclosure;
[0019] Figure 6 A schematic diagram of matrix storage provided by an embodiment of the present disclosure;
[0020] Figure 7 A flowchart of a training method for an image classification model provided in Embodiment 5 of the present disclosure;
[0021] Figure 8 A schematic diagram of a sparse weight matrix for the next round provided by an embodiment of the present disclosure;
[0022] Fig. 9 A flowchart of a training method for an image classification model provided in Embodiment 6 of the present disclosure;
[0023] Fig.10 A flowchart of a method for training an image classification model provided in Embodiment 7 of the present disclosure;
[0024] Fig.11 A schematic diagram of an image classification model used in this round and the next round of the embodiment of the present disclosure;
[0025] Fig.12 A flowchart of a method for training an image classification model provided in Embodiment 8 of the present disclosure;
[0026] Fig.13 A schematic diagram of the process of training an image classification model according to an embodiment of the present disclosure;
[0027] Fig.14 A flowchart of an image classification method provided in Embodiment 9 of the present disclosure;
[0028] Fig.15 A schematic diagram of the structure of a training device for an image classification model provided in Embodiment 10 of the present disclosure;
[0029] Fig.16 A schematic diagram of the structure of an image classification device provided in the eleventh embodiment of the present disclosure;
[0030] Fig.17 A schematic block diagram of an example electronic device that can be used to implement embodiments of the present disclosure is shown. DETAILED DESCRIPTION
[0031] The following is a description of exemplary embodiments of the present disclosure in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be recognized by those of ordinary skill in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0032] At present, with the development of deep learning technology, deep neural networks can be used as image classification models for image classification. As the number of layers of deep neural networks becomes deeper and deeper, the number of parameters continues to increase, resulting in the need for a large amount of computing and storage resources to be consumed in the training process of deep neural networks. Since there are many small values of the parameters of deep neural networks, and many zero values are generated during the training process, for example, the activation function will generate zero values. In related technologies, the structured sparse method can be used to remove useless or insignificant parameters in the model, which can effectively reduce the computing and storage resources required during the network training process.
[0033] However, the structured sparse method has the following three main problems: 1. The sparsification method is not used to accelerate training in the pre-training stage; 2. In the pre-training stage, only the weight parameters are sparse, and the actual storage and calculation do not utilize the sparse characteristics; 3. It must be a fixed-mode structured sparse method with limited sparsity, and acceleration in the inference stage requires specific hardware support.
[0034] In order to solve the above problems, the present invention provides an image classification model training method, an image classification method and an image classification device.
[0035] The following describes the image classification model training method, image classification method and device of the embodiments of the present disclosure with reference to the accompanying drawings.
[0036] Figure 1 A flowchart of the training method of the image classification model provided in the first embodiment of the present disclosure.
[0037] The embodiment of the present disclosure takes the training method of the image classification model as an example in which the training method of the image classification model is configured in the training device of the image classification model. The training device of the image classification model can be applied to any electronic device so that the electronic device can perform the training function of the image classification model.
[0038] Among them, the electronic device can be any device with computing capabilities, such as a personal computer (PC), a mobile terminal, etc. The mobile terminal can be, for example, a mobile phone, a tablet computer, a personal digital assistant, a wearable device, and other hardware devices with various operating systems, touch screens and / or display screens.
[0039] like Figure 1 As shown, the training method of the image classification model may include the following steps:
[0040] Step 101: Obtain a sample image.
[0041] In the embodiments of the present disclosure, multiple sample images can be obtained, wherein the sample images can be obtained from an existing training set, or the sample images can be collected online, such as by using a web crawler technology, or the sample images can be collected offline, such as by taking images of each object, or the sample images can be artificially synthesized images to obtain the sample images, etc. The present disclosure does not impose any restrictions on this.
[0042] In addition, the objects in each sample image can be labeled by category to obtain the category label corresponding to each category.
[0043] It should be noted that, in order to improve the training effect of the model, the labels can be manually annotated, or, in order to reduce labor costs and improve the training efficiency of the model, the labels can also be automatically annotated. For example, the object in the sample image can be automatically annotated by category through the annotation model, and the present disclosure does not limit this. Furthermore, after the objects in the sample image are automatically annotated, the labels annotated in the sample image can also be reviewed by manual review to improve the accuracy of the sample annotation results, thereby improving the training effect of the model.
[0044] Step 102, using sample images, performing multiple rounds of training processes on the image classification model, wherein any round of training process includes: performing forward calculations on the image classification model according to a set sparse weight matrix to obtain a predicted category of the sample image in this round; determining a loss function for this round according to a difference between the predicted category and the category annotation label on the sample image; performing reverse gradient calculations on the loss function for this round to obtain a dense weight matrix for this round; and clipping the dense weight matrix for this round to obtain a sparse weight matrix for the next round of training process.
[0045] In the disclosed embodiment, sample images may be used to perform multiple rounds of training processes on the image classification model, wherein any round of training process may include: performing forward calculation of the image classification model according to a set sparse weight matrix, and obtaining the predicted category of the sample image in this round, and then determining the loss function of this round according to the difference between the predicted category and the category annotation label on the sample image, and performing reverse gradient calculation on the loss function of this round. Since a dense weight matrix may be generated after the reverse gradient calculation of the loss function, the dense weight matrix of this round obtained by the reverse gradient calculation may be cropped, and the cropped sparse weight matrix is used as the set sparse weight matrix for the forward calculation of the image classification model in the next round of training process.
[0046] Step 103: When the value of the loss function in any round is less than a threshold, the training process is stopped.
[0047] As an example, a loss function can be generated based on the difference between the predicted category corresponding to the sample image and the category annotation label on the sample image, wherein the value of the loss function is positively correlated with the above difference, that is, the smaller the difference, the smaller the value of the loss function, and conversely, the larger the difference, the larger the value of the loss function. Therefore, in the present disclosure, the image classification model can be trained according to the value of the loss function to minimize the value of the loss function.
[0048] It should be noted that the above only takes the termination condition of the image classification model training as the minimization of the value of the loss function as an example. In actual application, other termination conditions can also be set. For example, the termination condition can also be that the number of training times reaches a set threshold, the training time is greater than a set threshold, etc. The present disclosure does not impose any restrictions on this.
[0049] In summary, by acquiring sample images; using the sample images, performing multiple rounds of training processes on the image classification model, wherein any round of training process includes: performing forward calculation of the image classification model according to the set sparse weight matrix to obtain the predicted category of the sample image in this round; determining the loss function of this round according to the difference between the predicted category and the category annotation label on the sample image; performing reverse gradient calculation on the loss function of this round to obtain the dense weight matrix of this round; cropping the dense weight matrix of this round to obtain the sparse weight matrix for the next round of training process; stopping the training process when the value of the loss function of any round is less than a threshold value, thereby automatically sparseening the weight matrix of the image classification model during the entire training process of the image classification model, so that the weight matrix is kept sparse during the training process, speeding up the calculation speed during the training process, and effectively reducing the calculation time during the training process.
[0050] In order to clearly illustrate how the dense weight matrix of the current round is pruned in the above embodiment of the present disclosure to obtain the sparse weight matrix of the next round, the present disclosure also proposes a training method for an image classification model.
[0051] Figure 2 A flowchart of the training method of the image classification model provided in the second embodiment of the present disclosure.
[0052] like Figure 2 As shown, the training method of the image classification model may include the following steps:
[0053] Step 201: Obtain a sample image.
[0054] Step 202: Using sample images, perform multiple rounds of training on the image classification model.
[0055] Step 203, the training process of any round includes: comparing each weight value in the dense weight matrix of this round with the set weight threshold to determine a target weight value greater than the set weight threshold from each weight value; setting the remaining weight values in each weight value except the target weight value to zero to obtain a sparse weight matrix for the next round.
[0056] Among them, the dense weight matrix of this round is calculated by forward calculation of the image classification model according to the set sparse weight matrix to obtain the predicted category of the sample image in this round; the loss function of this round is determined according to the difference between the predicted category and the category annotation label on the sample image; and the loss function of this round is obtained by reverse gradient calculation.
[0057] As an example, after performing reverse gradient calculation on the loss function of this round to obtain the dense weight matrix of this round, each weight value in the dense weight matrix of this round can be compared with the set weight threshold, and the weight values less than the set weight threshold can be set to zero to obtain the sparse weight matrix of the next round.
[0058] It should be noted that the above only uses the example of setting the weight values in the dense weight matrix that are less than the set weight threshold to zero to sparse the dense weight matrix. In actual application, a sparse matrix algorithm can also be used to sparse the dense weight matrix. Among them, the sparse matrix algorithm may include: graph search algorithm and sparse linear equation system solution.
[0059] For example, Figure 3 As shown, the weight values less than 0.002 in the weight matrix can be set to zero, and the weight matrix after being set to zero is used as the sparse weight matrix of the next round.
[0060] Step 204: When the value of the loss function in any round is less than a threshold, the training process is stopped.
[0061] It should be noted that the execution process of steps 201 to 202 and step 204 can be implemented in any way in the embodiments of the present disclosure, and the embodiments of the present disclosure do not limit this and will not be described in detail.
[0062] In summary, by comparing each weight value in the dense weight matrix of this round with the set weight threshold, a target weight value greater than the set weight threshold is determined from each weight value; the remaining weight values except the target weight value are set to zero to obtain the sparse weight matrix of the next round. Thus, by setting the weight values in the dense weight matrix that are less than the set weight threshold to zero, the sparse weight matrix of the next round can be obtained.
[0063] In order to obtain the corresponding sparse weight matrix in time during the next round of training, such as Figure 4 As shown, Figure 4 The flowchart of the training method of the image classification model provided in the third embodiment of the present disclosure is shown in FIG. In the embodiment of the present disclosure, the sparse weight matrix of the next round can be stored so as to obtain the corresponding sparse weight matrix when executing the next round of training process. Figure 4 The illustrated embodiment may include the following steps:
[0064] Step 401, obtaining a sample image.
[0065] Step 402, using sample images, performing multiple rounds of training processes on the image classification model, wherein any round of training process also includes: storing the sparse weight matrix of the next round, so as to obtain the corresponding sparse weight matrix when performing the next round of training process.
[0066] Among them, the sparse weight matrix of the next round is calculated by forward calculation of the image classification model according to the set sparse weight matrix to obtain the predicted category of the sample image in this round; the loss function of this round is determined according to the difference between the predicted category and the category annotation label on the sample image; the loss function of this round is reversely calculated by gradient calculation to obtain the dense weight matrix of this round; and the dense weight matrix of this round is obtained by clipping.
[0067] In the embodiment of the present disclosure, in order to obtain the corresponding sparse weight matrix in time when executing the next round of training process, the dense weight matrix of this round can be clipped to obtain the sparse weight matrix for the next training process, and the sparse weight matrix of the next round can be stored. It should be noted that in order to save storage space and facilitate accelerated calculation, the sparse weight matrix of the next round can be stored in a sparse format, for example, it can be stored in coordinate form or in a compressed row information manner.
[0068] Step 403: When the value of the loss function in any round is less than a threshold, the training process is stopped.
[0069] It should be noted that the execution process of step 401 and step 403 can be implemented in any way in the embodiments of the present disclosure, and the embodiments of the present disclosure do not limit this and will not be described in detail.
[0070] In summary, by storing the sparse weight matrix of the next round, the corresponding sparse weight matrix can be obtained in time during the next round of training.
[0071] In order to clearly illustrate how the sparse weight matrix of the next round is stored in the above embodiment of the present disclosure, as an example, the embodiment of the present disclosure proposes a training method for an image classification model.
[0072] Figure 5 A flowchart of the training method of the image classification model provided in the fourth embodiment of the present disclosure.
[0073] like Figure 5 As shown, the training method of the image classification model may include the following steps:
[0074] Step 501: Acquire a sample image.
[0075] Step 502: Using sample images, perform multiple rounds of training on the image classification model.
[0076] Step 503: During any round of training, each non-zero weight value in the sparse weight matrix of the next round and the position of each non-zero weight value in the sparse weight matrix of the next round are obtained.
[0077] Among them, the sparse weight matrix of the next round is calculated by forward calculation of the image classification model according to the set sparse weight matrix to obtain the predicted category of the sample image in this round; the loss function of this round is determined according to the difference between the predicted category and the category annotation label on the sample image; the loss function of this round is reversely calculated by gradient calculation to obtain the dense weight matrix of this round; and the dense weight matrix of this round is obtained by clipping.
[0078] In an embodiment of the present disclosure, each non-zero weight value in the sparse weight matrix of the next round and the position of each non-zero weight value in the sparse weight matrix of the next round can be obtained, for example, the row index and column index of each non-zero weight value in the sparse weight matrix of the next round.
[0079] Step 504: storing each non-zero weight value according to its position in the sparse weight matrix of the next round.
[0080] For example, Figure 6As shown, the row index, column index and corresponding non-zero weight value of each non-zero weight value are stored.
[0081] Step 505: When the value of the loss function in any round is less than a threshold, the training process is stopped.
[0082] It should be noted that the execution process of step 501 to step 502 and step 505 can be implemented in any way in the embodiments of the present disclosure, and the embodiments of the present disclosure do not limit this and will not be described in detail.
[0083] In summary, during any round of training, each non-zero weight value in the sparse weight matrix of the next round and the position of each non-zero weight value in the sparse weight matrix of the next round are obtained; each non-zero weight value is stored according to the position of each non-zero weight value in the sparse weight matrix of the next round. Thus, only the non-zero weight values in the sparse weight matrix of the next round are stored, which can save storage space and facilitate accelerated calculation during training.
[0084] In order to clearly illustrate how to store the sparse weight matrix of the next round in the above embodiment of the present disclosure, as another example, the embodiment of the present disclosure proposes a training method for an image classification model.
[0085] Figure 7 FIG. 5 is a flow chart of the training method of the image classification model provided in the fifth embodiment of the present disclosure. Figure 7 As shown, the training method of the image classification model may include the following steps:
[0086] Step 701, obtaining a sample image.
[0087] Step 702: Using sample images, perform multiple rounds of training on the image classification model.
[0088] Step 703, during any round of training, obtain each non-zero weight value in the sparse weight matrix of the next round, and the row index and column index of each non-zero weight value in the sparse weight matrix of the next round.
[0089] Among them, the sparse weight matrix of the next round is calculated by forward calculation of the image classification model according to the set sparse weight matrix to obtain the predicted category of the sample image in this round; the loss function of this round is determined according to the difference between the predicted category and the category annotation label on the sample image; the loss function of this round is reversely calculated by gradient calculation to obtain the dense weight matrix of this round; and the dense weight matrix of this round is obtained by clipping.
[0090] For example, Figure 8 As shown, Figure 8Taking the matrix shown in as the sparse weight matrix of the next round as an example, the non-zero weight values in the sparse weight matrix of the next round that can be obtained are 1, 2, 3, 4, 5 and 6.
[0091] Step 704: Use elements in the first storage array to store each non-zero weight value.
[0092] In the embodiment of the present disclosure, each element in the first storage array may be used to store each non-zero element.
[0093] Step 705 , determining the first target element corresponding to each row from the first storage array according to the row index of each non-zero weight value.
[0094] Among them, the first target element of each row is the first non-zero weight value in the corresponding row of the sparse weight matrix of the next round.
[0095] Furthermore, according to the row index of each non-zero weight value, the first non-zero weight value in each row of the next round of sparse weight matrix can be screened out from the first storage array. For example, Figure 8 The first target element corresponding to the row with row index 0 in the matrix is 1, the first target element corresponding to the row with row index 1 is 2, the first target element corresponding to the row with row index 2 is 4, and the first target element corresponding to the row with row index 3 is 5.
[0096] Step 706: Use the elements in the second storage array to store the position of the first target element in the first storage data group, and use the last element in the second storage array to store the number of elements contained in the first storage array.
[0097] Continue with Figure 8 Take the matrix shown as an example, Figure 8 The non-zero weight values in the matrix shown are 1, 2, 3, 4, 5 and 6, and the non-zero weight values 1, 2, 3, 4, 5 and 6 are stored in the first storage array, and the first storage array is {1, 2, 3, 4, 5, 6}, wherein it can be determined from the first array that the first target element corresponding to the row with row index 0 is 1, the first target element corresponding to the row with row index 1 is 2, the first target element of the row with row index 2 is 4, and the first target element of the row with row index 3 is 5. The position of the first target element 1 in the first storage array is 0, the position of the first target element 2 in the first storage array is 1, the position of the first target element 4 in the first storage array is 3, and the position of the first target element 5 in the first storage array is 4. Furthermore, the elements in the second storage array are used to store the position of the first target element in the first storage array, and the last element in the second storage array is used to store the number of elements contained in the first storage array, such as, the last element in the second storage array is 6, such as, the second storage array is {0, 1, 3, 4, 6}.
[0098] Step 707: Use the elements in the third storage array to store the column index corresponding to each element in the first storage array in the sparse weight matrix of the next round.
[0099] Continue with Figure 8 Taking the matrix shown in FIG. 1 as an example, the column indexes corresponding to the elements in the first storage array {1, 2, 3, 4, 5, 6} in the next round of sparse weight matrix are 1, 0, 3, 2, 1 and 3 respectively. The third storage array is {1, 0, 3, 2, 1, 3}.
[0100] Step 708: When the value of the loss function in any round is less than a threshold, the training process is stopped.
[0101] It should be noted that the execution process of step 701 to step 702 and step 708 can be implemented in any way in the embodiments of the present disclosure, and the embodiments of the present disclosure do not limit this and will not be repeated.
[0102] In summary, through any round of training process, each non-zero weight value in the sparse weight matrix of the next round is obtained, as well as the row index and column index of each non-zero weight value in the sparse weight matrix of the next round; the elements in the first storage array are used to store each non-zero weight value; according to the row index of each non-zero weight value, the first target element corresponding to each row is determined from the first storage array; the elements in the second storage array are used to store the position of the first target element in the first storage data group, and the last element in the second storage array is used to store the number of elements contained in the first storage array; the elements in the third storage array are used to store the column index corresponding to each element in the first storage array in the sparse weight matrix of the next round, thereby, the sparse weight matrix of the next round is stored in a compressed row information manner, which can save storage space and facilitate accelerated calculation during training.
[0103] In order to clearly illustrate how to perform the forward calculation of the image classification model according to the sparse weight matrix of this round in any round of training, and perform reverse gradient calculation on the loss function of this round to obtain the dense weight matrix of this round, the present disclosure proposes a training method for an image classification model.
[0104] Fig. 9 This is a flow chart of the training method of the image classification model provided in the sixth embodiment of the present disclosure. Fig. 9 As shown, the training method of the image classification model may include the following steps:
[0105] Step 901, obtaining a sample image.
[0106] Step 902: Using sample images, perform multiple rounds of training on the image classification model.
[0107] Step 903: During any round of training, forward calculation is performed based on each operator matrix in the image classification model of this round and the sparse weight matrix of this round to obtain the predicted category of this round.
[0108] It should be noted that the image classification model may include multiple network layers, each network layer may include multiple operators, and each operator may include at least one operator matrix.
[0109] In the disclosed embodiment, during any round of training of the image classification model, the sample image is input into the image classification model of this round, and each operator matrix in the image classification model of this round and the sparse weight matrix of this round are forward calculated to obtain the predicted category of the sample image in this round.
[0110] Step 904, performing reverse gradient calculation on the loss function of this round according to each operator matrix in the image classification model of this round, so as to obtain a dense weight matrix of this round.
[0111] The loss function is obtained based on the difference between the predicted category and the category annotation label on the sample image.
[0112] Step 905: Each operator matrix in the current round image classification model performs reverse gradient calculation on the loss function of this round to obtain a dense weight matrix of this round.
[0113] Step 906, pruning the dense weight matrix of this round to obtain a sparse weight matrix for the next round of training process.
[0114] Step 907: When the value of the loss function in any round is less than a threshold, the training process is stopped.
[0115] It should be noted that the execution process of step 901 to step 902 and step 905 to step 907 can be implemented in any way in the embodiments of the present disclosure, and the embodiments of the present disclosure do not limit this and will not be repeated.
[0116] In summary, during any round of training, forward calculation is performed according to each operator matrix in this round of image classification model and the sparse weight matrix of this round to obtain the prediction category of this round; reverse gradient calculation is performed on the loss function of this round according to each operator matrix in this round of image classification model to obtain the dense weight matrix of this round. Thus, the prediction category and dense weight matrix in any round of training can be obtained.
[0117] In order to further speed up the calculation, Fig.10 As shown, Fig.10This is a flow chart of the training method of the image classification model provided in the seventh embodiment of the present disclosure. A sparse operator can be used to calculate the weight matrix to speed up the calculation. Fig.10 The illustrated embodiment may include the following steps:
[0118] Step 1001, obtaining a sample image.
[0119] Step 1002, using sample images, execute multiple rounds of training process on the image classification model.
[0120] Step 1003: During any round of training, at least one target operator matrix is determined from the operator matrices in the image classification model of this round.
[0121] Step 1004, among the operator matrices in the current round of image classification model, at least one target operator matrix is sparsely processed to obtain the operator matrices used in the next round of training.
[0122] The operator matrices corresponding to the first round are obtained by sparsely processing the initial operator matrices of the image classification model, and the initial operator matrices can be set operator matrices.
[0123] For example, Fig.11 As shown, for example, the image classification model includes a first convolution layer, a first activation function layer, a first maximum pooling layer, a second convolution layer, a second activation function layer and a linear layer, and the operator matrices of the first convolution layer, the second convolution layer and the linear layer can be used as target operator matrices.
[0124] Then, from each operator matrix in the current image classification model, the target operator matrix is sparsed, and the sparse target operator and the operators of each operator matrix in the current image classification model except the target operator matrix are used as the operator matrices used in the next round of training, where: Fig.11 On the left are the network layers corresponding to the operator matrices used in this round of training. Fig.11 The right side shows the network layers corresponding to the operator matrices used in the next round of training, and, Fig.11 The shaded part on the right is the network layer corresponding to the target operator matrix.
[0125] Step 1005, forward calculation is performed according to each operator matrix in the current round image classification model and the current round sparse weight matrix to obtain the predicted category of the sample image in this round.
[0126] Step 1006, performing reverse gradient calculation on the loss function of this round according to each operator matrix in the image classification model of this round, so as to obtain the dense weight matrix of this round.
[0127] The loss function is obtained based on the difference between the predicted category and the category annotation label on the sample image.
[0128] Step 1007, pruning the dense weight matrix of this round to obtain a sparse weight matrix for the next round of training process.
[0129] Step 1008: When the value of the loss function in any round is less than a threshold, the training process is stopped.
[0130] It should be noted that the execution process of step 1001 to step 1003 and step 1005 to step 1008 can be implemented in any way in the embodiments of the present disclosure, and the embodiments of the present disclosure do not limit this and will not be repeated.
[0131] In summary, during any round of training, at least one target operator matrix is determined from each operator matrix in the current round of image classification model; among each operator matrix in the current round of image classification model, at least one target operator matrix is sparsely constructed to obtain each operator matrix used in the next round of training. Thus, the sparsely constructed operator matrix is used for calculation in each round of training, which can speed up the calculation and does not require specific hardware support.
[0132] In order to clearly explain how the set sparse weight matrix is obtained during the first round of training of the image classification model, the present disclosure proposes a training method for the image classification model.
[0133] Fig.12 This is a flow chart of the training method of the image classification model provided in the eighth embodiment of the present disclosure. Fig.12 As shown, the training method of the image classification model may include the following steps:
[0134] Step 1201, obtaining a sample image.
[0135] Step 1202, using sample images, performing multiple rounds of training processes on the image classification model, wherein an initial weight matrix of the image classification model is determined based on the sample images.
[0136] In an embodiment of the present disclosure, a sample image may be input into an image classification model to initialize a weight matrix of the image classification model, thereby obtaining an initial weight matrix of the image classification model.
[0137] Step 1203, sparse the initial weight matrix to obtain a set sparse weight matrix used in the first round of training process.
[0138] Furthermore, by sparsifying the initial weight matrix, a set sparse weight matrix used in the first round of training can be obtained. For example, the matrix values in the initial weight matrix that are less than the set weight threshold can be set to zero, and the weight matrix after being set to zero can be used as the set sparse weight matrix used in the first round of training.
[0139] Step 1204, in any round of training process, forward calculation of the image classification model is performed according to the set sparse weight matrix to obtain the predicted category of the sample image in this round; the loss function of this round is determined according to the difference between the predicted category and the category annotation label on the sample image; the loss function of this round is reversely calculated to obtain the dense weight matrix of this round; the dense weight matrix of this round is cropped to obtain the sparse weight matrix for the next round of training process.
[0140] Step 1205: When the value of the loss function in any round is less than a threshold, the training process is stopped.
[0141] It should be noted that the execution process of step 1201 and step 1204 to step 1205 can be implemented in any way in the embodiments of the present disclosure, and the embodiments of the present disclosure do not limit this and will not be repeated.
[0142] In summary, by sparsely computing the initial weight matrix of the image classification model obtained according to the sample image, a set sparse weight matrix used in the first round of training can be obtained.
[0143] In order to explain the above embodiment clearly, an example is given for illustration.
[0144] For example, Fig.13 As shown in the figure, first, the initial weight matrix of the image classification model is converted into a sparse format through the sparsification module, and the operators that support sparse calculations are converted into corresponding sparse operators, thereby obtaining a new sparse network, and then the normal training process is performed, but a clipping operation is added after each step of training. This is because dense weight data may be generated after the reverse calculation of the gradient, so the clipping operation is inserted here to clip the weight data again so that the weight data remains sparse, and then the next round of training is performed. This method always keeps the weight data sparse during the training process, saving storage space, and uses sparse operators for calculation. When the sparsity reaches a certain level, it can effectively reduce the calculation time.
[0145] The training method of the image classification model of the embodiment of the present disclosure obtains sample images, uses the sample images, and performs multiple rounds of training processes on the image classification model, wherein any round of training process includes: performing forward calculation of the image classification model according to the set sparse weight matrix to obtain the predicted category of the sample image in this round; determining the loss function of this round according to the difference between the predicted category and the category annotation label on the sample image; performing reverse gradient calculation on the loss function of this round to obtain the dense weight matrix of this round; and cropping the dense weight matrix of this round to obtain the sparse weight matrix for the next round of training process. Thus, the weight matrix of the image classification model is automatically sparse during the entire training process of the image classification model, so that the weight matrix is kept sparse during the training process, the calculation speed during the training process is accelerated, and the calculation time during the training process can be effectively reduced.
[0146] The above are various embodiments corresponding to the training method of the image classification model. The present disclosure also proposes an application method of the image classification model, namely, an image classification method.
[0147] Fig.14 This is a flowchart of the image classification method provided in Embodiment 9 of the present disclosure.
[0148] like Fig.14 As shown, the image classification method may include the following steps:
[0149] Step 1401, obtaining an image to be classified.
[0150] In the embodiments of the present disclosure, the images to be classified can be obtained from an existing test set, or the images to be classified can be collected online, for example, by using web crawler technology, or the images to be classified can be collected offline, or the images to be classified can be images input by the user, etc. The embodiments of the present disclosure do not limit this.
[0151] Step 1402: Use the trained image classification model to predict the category of the object in the image to be classified to obtain the predicted category of the image to be classified.
[0152] The image classification model may be trained by using any of the above method embodiments.
[0153] In the disclosed embodiment, the image to be classified may be input into a trained image classification model, and the image classification model may perform category prediction on the object in the image to be classified to obtain the predicted category of the image to be classified.
[0154] The image classification method of the embodiment of the present disclosure obtains an image to be classified; uses a trained image classification model to predict the category of the object in the image to be classified to obtain a predicted category of the image to be classified, thereby performing category prediction on the image to be classified based on deep learning technology, which can improve the accuracy and reliability of the category prediction result.
[0155] With the above Figures 1 to 13 Corresponding to the training method of the image classification model provided in the embodiment, the present disclosure also provides a training device for the image classification model. Figures 1 to 13 The training method of the image classification model provided in the embodiment corresponds to the embodiment, so the implementation method of the image classification model training method is also applicable to the training device of the image classification model provided in the embodiment of the present disclosure, and will not be described in detail in the embodiment of the present disclosure.
[0156] Fig.15 A schematic diagram of the structure of a training device for an image classification model provided in Embodiment 10 of the present disclosure.
[0157] like Fig.15 As shown, the training device 1500 for the image classification model includes: an acquisition module 1510, a training module 1520 and a stop module 1530.
[0158] Among them, the acquisition module 1510 is used to acquire sample images; the training module 1520 is used to use the sample images to perform multiple rounds of training processes on the image classification model, wherein any round of training process includes: performing forward calculation of the image classification model according to the set sparse weight matrix to obtain the predicted category of the sample image in this round; determining the loss function of this round according to the difference between the predicted category and the category annotation label on the sample image; performing reverse gradient calculation on the loss function of this round to obtain the dense weight matrix of this round; cropping the dense weight matrix of this round to obtain the sparse weight matrix for the next round of training process; and the stop module 1530 is used to stop the training process when the value of the loss function of any round is less than a threshold.
[0159] As a possible implementation method of the embodiment of the present disclosure, the training module 1520 is also used to: compare each weight value in the dense weight matrix of this round with the set weight threshold to determine a target weight value greater than the set weight threshold from each weight value; set the remaining weight values in each weight value except the target weight value to zero to obtain a sparse weight matrix for the next round.
[0160] As a possible implementation of the embodiment of the present disclosure, the training device 1500 for the image classification model further includes: a storage module.
[0161] The storage module is used to store the sparse weight matrix of the next round so as to obtain the corresponding sparse weight matrix when executing the next round of training process.
[0162] As a possible implementation method of an embodiment of the present disclosure, the storage module is also used to: obtain each non-zero weight value in the sparse weight matrix of the next round, and the position of each non-zero weight value in the sparse weight matrix of the next round; store each non-zero weight value according to the position of each non-zero weight value in the sparse weight matrix of the next round.
[0163] As a possible implementation method of the embodiment of the present disclosure, the storage module is also used to: obtain each non-zero weight value in the sparse weight matrix of the next round, and the row index and column index of each non-zero weight value in the sparse weight matrix of the next round; use elements in the first storage array to store each non-zero weight value; determine the first target element corresponding to each row from the first storage array according to the row index of each non-zero weight value; wherein the first target element of each row is the first non-zero weight value in the corresponding row of the sparse weight matrix of the next round; use elements in the second storage array to store the position of the first target element in the first storage data group, and use the last element in the second storage array to store the number of elements contained in the first storage array; use elements in the third storage array to store the column index corresponding to each element in the first storage array in the sparse weight matrix of the next round.
[0164] As a possible implementation method of the embodiment of the present disclosure, the training module is also used to: perform forward calculation according to each operator matrix in the current round image classification model and the current round sparse weight matrix to obtain the predicted category of the sample image in this round; perform reverse gradient calculation on the loss function of this round according to each operator matrix in the current round image classification model to obtain the current round dense weight matrix.
[0165] As a possible implementation of the embodiment of the present disclosure, the training device 1500 for the image classification model further includes: a first determination module and a first sparse module.
[0166] Among them, the first determination module is used to determine at least one target operator matrix from the operator matrices in the current round of image classification model; the first sparse module is used to sparse at least one target operator matrix among the operator matrices in the current round of image classification model to obtain the operator matrices used in the next round of training, wherein the operator matrices corresponding to the first round are obtained by sparsely storing the initial operator matrices of the image classification model.
[0167] As a possible implementation manner of the embodiment of the present disclosure, the training device 1500 for the image classification model further includes: a second determination module and a second sparse module.
[0168] Among them, the second determination module is used to determine the initial weight matrix of the image classification model according to the sample image; the second sparse module is used to sparse the initial weight matrix to obtain the set sparse weight matrix used in the first round of training process.
[0169] The training device of the image classification model of the embodiment of the present disclosure obtains sample images and uses the sample images to perform multiple rounds of training processes on the image classification model, wherein any round of training process includes: performing forward calculation of the image classification model according to the set sparse weight matrix to obtain the predicted category of the sample image in this round; determining the loss function of this round according to the difference between the predicted category and the category annotation label on the sample image; performing reverse gradient calculation on the loss function of this round to obtain the dense weight matrix of this round; and cropping the dense weight matrix of this round to obtain the sparse weight matrix for the next round of training process. Therefore, by automatically sparseening the weight matrix of the image classification model during the entire training process of the image classification model, the weight matrix is kept sparse during the training process, which speeds up the calculation speed during the training process and can effectively reduce the calculation time during the training process.
[0170] With the above Fig.14 Corresponding to the image classification method provided in the embodiment, the present disclosure also provides an image classification device. Fig.14 The image classification method provided in the embodiment corresponds to the above, so the implementation of the image classification method is also applicable to the image classification device provided in the embodiment of the present disclosure, and will not be described in detail in the embodiment of the present disclosure.
[0171] Fig.16 This is a schematic diagram of the structure of the image classification device provided in the eleventh embodiment of the present disclosure.
[0172] like Fig.16 As shown, the image classification device 1600 includes: an acquisition module 1610 and a prediction module 1620 .
[0173] The acquisition module 1610 is used to acquire the image to be classified; the prediction module 1620 is used to adopt Fig.15 The image classification model obtained by the image classification model training device performs category prediction on the object in the image to be classified to obtain the predicted category of the image to be classified.
[0174] The image classification device of the embodiment of the present disclosure obtains an image to be classified; and uses a trained image classification model to predict the category of the object in the image to be classified to obtain a predicted category of the image to be classified. Thus, based on deep learning technology, the category prediction of the image to be classified is performed, which can improve the accuracy and reliability of the category prediction result.
[0175] In order to implement the above embodiments, the present disclosure also provides an electronic device, which may include at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the image classification model training method or image classification method proposed in any of the above embodiments of the present disclosure.
[0176] In order to implement the above embodiments, the present disclosure also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to enable a computer to execute the image classification model training method or image classification method proposed in any of the above embodiments of the present disclosure.
[0177] In order to implement the above embodiments, the present disclosure also provides a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the image classification model training method or image classification method proposed in any of the above embodiments of the present disclosure.
[0178] It should be noted that in the technical solution of the present disclosure, the collection, storage, use, processing, transmission, provision and disclosure of user personal information are all carried out with the user's consent, and are in compliance with relevant laws and regulations and do not violate public order and good morals.
[0179] According to an embodiment of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium and a computer program product.
[0180] Fig.17 A schematic block diagram of an example electronic device 1700 that can be used to implement an embodiment of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.
[0181] like Fig.17As shown, the device 1700 includes a computing unit 1701, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 1702 or a computer program loaded from a storage unit 1708 into a random access memory (RAM) 1703. In the RAM 1703, various programs and data required for the operation of the device 1700 can also be stored. The computing unit 1701, the ROM 1702, and the RAM 1703 are connected to each other via a bus 1704. An input / output (I / O) interface 1705 is also connected to the bus 1704.
[0182] A number of components in the device 1700 are connected to the I / O interface 1705, including: an input unit 1706, such as a keyboard, a mouse, etc.; an output unit 1707, such as various types of displays, speakers, etc.; a storage unit 1708, such as a disk, an optical disk, etc.; and a communication unit 1709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1709 allows the device 1700 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.
[0183] The computing unit 1701 may be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 1701 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 1701 performs the various methods and processes described above, such as the training method of the image classification model, or the image classification method. For example, in some embodiments, the training method of the image classification model, or the image classification method may be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as a storage unit 1708. In some embodiments, part or all of the computer program may be loaded and / or installed on the device 1700 via the ROM 1702 and / or the communication unit 1709. When the computer program is loaded into the RAM 1703 and executed by the computing unit 1701, the training method of the image classification model described above, or one or more steps of the image classification method, may be executed. Alternatively, in other embodiments, the computing unit 1701 may be configured to execute a training method for an image classification model, or an image classification method, in any other appropriate manner (eg, by means of firmware).
[0184] Various implementations of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), load programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various implementations can include: being implemented in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0185] The program code for implementing the method of the present disclosure may be written in any combination of one or more programming languages. These program codes may be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device, so that the program code, when executed by the processor or controller, enables the functions / operations specified in the flow chart and / or block diagram to be implemented. The program code may be executed entirely on the machine, partially on the machine, partially on the machine and partially on a remote machine as a stand-alone software package, or entirely on a remote machine or server.
[0186] In the context of the present disclosure, a machine-readable medium may be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, device, or equipment. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium may include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0187] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0188] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0189] A computer system may include a client and a server. The client and the server are generally remote from each other and usually interact through a communication network. The relationship of client and server is generated by computer programs running on respective computers and having a client-server relationship to each other. The server may also be a server of a distributed system, or a server combined with a blockchain.
[0190] It should be noted that artificial intelligence is a discipline that studies how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.), and includes both hardware-level and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing; artificial intelligence software technologies mainly include computer vision technology, speech recognition technology, natural language processing technology, as well as machine learning / deep learning, big data processing technology, knowledge graph technology, and other major directions.
[0191] It should be understood that the various forms of processes shown above can be used to reorder, add or delete steps. For example, the steps recorded in this disclosure can be executed in parallel, sequentially or in different orders, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and this document does not limit this.
[0192] The above specific implementations do not constitute a limitation on the protection scope of the present disclosure. It should be understood by those skilled in the art that various modifications, combinations, sub-combinations and substitutions can be made according to design requirements and other factors. Any modification, equivalent substitution and improvement made within the spirit and principle of the present disclosure shall be included in the protection scope of the present disclosure.
Claims
1. A training method for an image classification model, include: Get a sample image; Using the sample image, perform multiple rounds of training processes on the image classification model, wherein any round of training process includes: performing forward calculation of the image classification model according to a set sparse weight matrix to obtain the predicted category of the sample image in this round; determining the loss function of this round according to the difference between the predicted category and the category annotation label on the sample image; performing reverse gradient calculation on the loss function of this round to obtain a dense weight matrix of this round; and pruning the dense weight matrix of this round to obtain a sparse weight matrix for the next round of training process, wherein the image classification model includes multiple network layers, each network layer includes multiple operators, and each operator may include at least one operator matrix; When the value of the loss function of any round is less than a threshold, stopping the training process; The forward calculation of the image classification model is performed according to the sparse weight matrix of this round, including: Perform forward calculation according to each operator matrix in the current round image classification model and the current round sparse weight matrix to obtain the predicted category of the sample image in this round; Accordingly, the loss function of the current round is reversely calculated to obtain the dense weight matrix of the current round, including: The loss function of this round is reversely calculated based on each operator matrix in the image classification model of this round to obtain a dense weight matrix of this round.
2. The method according to claim 1, in, The dense weight matrix of the current round is trimmed to obtain a sparse weight matrix of the next round, including: Comparing each weight value in the dense weight matrix of the current round with a set weight threshold to determine a target weight value greater than the set weight threshold from each of the weight values; The remaining weight values in the weight values except the target weight value are set to zero to obtain a sparse weight matrix for the next round.
3. The method according to claim 2, in, The method further comprises: The sparse weight matrix of the next round is stored so as to obtain the corresponding sparse weight matrix when executing the next round of training process.
4. The method according to claim 3, in, The storing of the sparse weight matrix of the next round includes: Obtaining each non-zero weight value in the sparse weight matrix of the next round, and the position of each non-zero weight value in the sparse weight matrix of the next round; The non-zero weight values are stored according to their positions in the sparse weight matrix of the next round.
5. The method according to claim 3, in, The storing of the sparse weight matrix of the next round includes: Obtaining each non-zero weight value in the sparse weight matrix of the next round, and the row index and column index of each non-zero weight value in the sparse weight matrix of the next round; Using elements in the first storage array to store each of the non-zero weight values; According to the row index of each non-zero weight value, determine the first target element corresponding to each row from the first storage array; wherein the first target element of each row is the first non-zero weight value in the corresponding row of the sparse weight matrix of the next round; Using elements in the second storage array to store the position of the first target element in the first storage array, and using the last element in the second storage array to store the number of elements contained in the first storage array; The elements in the third storage array are used to store the column index corresponding to each element in the first storage array in the sparse weight matrix of the next round.
6. The method according to any one of claims 1 to 5, in, Any round of training process also includes: Determine at least one target operator matrix from each operator matrix in the current round of image classification model; Among the operator matrices in the current round of image classification model, at least one target operator matrix is sparsely constructed to obtain the operator matrices used in the next round of training, wherein the operator matrices corresponding to the first round are obtained by sparsely constructing the initial operator matrices of the image classification model.
7. The method according to any one of claims 1 to 5, in, The method further comprises: Determining an initial weight matrix of the image classification model according to the sample image; The initial weight matrix is sparsely constructed to obtain the set sparse weight matrix used in the first round of the training process.
8. A method for image classification, include: Get the image to be classified; An image classification model obtained by the training method of the image classification model described in any one of claims 1 to 7 is used to perform category prediction on the object in the image to be classified to obtain the predicted category of the image to be classified.
9. A training device for an image classification model, include: An acquisition module, used for acquiring a sample image; A training module, used to use the sample image to perform multiple rounds of training processes on the image classification model, wherein any round of training process includes: performing forward calculation of the image classification model according to a set sparse weight matrix to obtain the predicted category of the sample image in this round; determining the loss function of this round according to the difference between the predicted category and the category annotation label on the sample image; performing reverse gradient calculation on the loss function of this round to obtain a dense weight matrix of this round; and pruning the dense weight matrix of this round to obtain a sparse weight matrix for the next round of training process, wherein the image classification model includes multiple network layers, each network layer includes multiple operators, and each operator may include at least one operator matrix; A stopping module, used for stopping the training process when the value of the loss function of any round is less than a threshold; Wherein, the device further comprises: A first determination module, used to determine at least one target operator matrix from each operator matrix in the current round of image classification model; The first sparse module is used to sparse the at least one target operator matrix in the operator matrices in the current round of image classification model to obtain the operator matrices used in the next round of training, wherein the operator matrices corresponding to the first round are obtained by sparsely processing the initial operator matrices of the image classification model.
10. The device according to claim 9, in, The training module is also used for: Comparing each weight value in the dense weight matrix of the current round with a set weight threshold to determine a target weight value greater than the set weight threshold from each of the weight values; The remaining weight values in the weight values except the target weight value are set to zero to obtain a sparse weight matrix for the next round.
11. The device according to claim 10, in, The device also includes: The storage module is used to store the sparse weight matrix of the next round so as to obtain the corresponding sparse weight matrix when executing the next round of training process.
12. The device according to claim 11, in, The storage module is further used for: Obtaining each non-zero weight value in the sparse weight matrix of the next round, and the position of each non-zero weight value in the sparse weight matrix of the next round; The non-zero weight values are stored according to their positions in the sparse weight matrix of the next round.
13. The device according to claim 11, in, The storage module is further used for: Obtaining each non-zero weight value in the sparse weight matrix of the next round, and the row index and column index of each non-zero weight value in the sparse weight matrix of the next round; Using elements in the first storage array to store each of the non-zero weight values; According to the row index of each non-zero weight value, determine the first target element corresponding to each row from the first storage array; wherein the first target element of each row is the first non-zero weight value in the corresponding row of the sparse weight matrix of the next round; Using elements in the second storage array to store the position of the first target element in the first storage array, and using the last element in the second storage array to store the number of elements contained in the first storage array; The elements in the third storage array are used to store the column index corresponding to each element in the first storage array in the sparse weight matrix of the next round.
14. The device according to claim 9, in, The training module is also used for: Perform forward calculation according to each operator matrix in the current round image classification model and the current round sparse weight matrix to obtain the predicted category of the sample image in this round; The loss function of this round is reversely calculated based on each operator matrix in the image classification model of this round to obtain a dense weight matrix of this round.
15. The device according to any one of claims 9 to 14, in, The device also includes: A second determination module, used to determine an initial weight matrix of the image classification model according to the sample image; The second sparse module is used to sparse the initial weight matrix to obtain the set sparse weight matrix used in the first round of the training process.
16. An image classification device, include: An acquisition module, used for acquiring images to be classified; A prediction module is used to use the image classification model obtained by the training device of the image classification model described in any one of claims 9 to 15 to perform category prediction on the object in the image to be classified to obtain the predicted category of the image to be classified.
17. An electronic device, include: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method of any one of claims 1 to 7, or the method of claim 8.
18. A non-transitory computer-readable storage medium storing computer instructions, in, The computer instructions are used to enable the computer to execute the method according to any one of claims 1 to 7, or to execute the method according to claim 9.
19. A computer program product comprising a computer program, in, When the computer program is executed by a processor, the computer program implements the method according to any one of claims 1 to 7, or executes the method according to claim 8.
Citation Information
Patent Citations
FPGA-based sparsity neural network accelerating system
CN108932548A
Structured sparse parameter processing method, device and equipment and storage medium
CN112508190A