Neural network model training method and device, storage medium and electronic equipment

By performing sparse processing and iterative training in the neural network model, the problem of decreasing computing speed of large-scale neural network models is solved, and the model is lightweight and high-precision computing speed is improved.

CN120124698APending Publication Date: 2025-06-10NANJING HORIZON INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510192295.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-20
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

In deep learning technology, as the scale of neural network models increases, the number of model parameters and calculations increases, resulting in a decrease in the calculation speed. How to improve the calculation speed of model has become an important issue.

Method used

By determining multiple operators that support sparseness in the neural network model, evaluating the sparse sensitivity corresponding to each operator, determining the sparse parameters, sparse the operators, obtaining the initial sparse model, and optimizing the model parameters through iterative training to obtain a high-precision target sparse model.

Benefits of technology

It realizes the lightweight, reduced complexity of neural network models, while maintaining high accuracy, thereby improving the computing speed of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120124698A_ABST
    Figure CN120124698A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a neural network model training method and device, a storage medium and electronic equipment. The method comprises the following steps: determining a plurality of operators supporting rarefaction in a neural network model; determining at least one sparse sensitivity of each operator in the plurality of operators corresponding to at least one sparse condition; based on the at least one sparse sensitivity corresponding to each operator in the plurality of operators, determining sparse parameters of the plurality of operators; based on the sparse parameters, sparse processing is carried out on at least part of operators in the multiple operators, and an initial sparse model is obtained; and performing iterative training on the initial sparse model to obtain a target sparse model. According to the embodiment of the invention, the calculation speed of the model can be improved through the sparsification of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to deep learning technology, and in particular to a training method, device, storage medium and electronic device for a neural network model. Background Art

[0002] At present, the development of deep learning technology has made the scale of neural network models larger and larger, which can enhance the learning and expression capabilities of neural network models, but it will increase the number of model parameters and the amount of computation, resulting in a decrease in the model's computing speed. How to improve the model's computing speed is a problem worthy of attention for those skilled in the art. Summary of the Invention

[0003] To solve the above technical problems, the present disclosure provides a training method, device, storage medium and electronic device for a neural network model.

[0004] According to one aspect of the embodiments of the present disclosure, there is provided a training method for a neural network model, including:

[0005] Determine a plurality of operators in the neural network model that support sparsification;

[0006] Determine at least one sparsity sensitivity corresponding to at least one sparsity condition for each operator in the plurality of operators;

[0007] Based on at least one sparsity sensitivity corresponding to each operator in the plurality of operators, determine the sparsity parameters of the plurality of operators;

[0008] Based on the sparsity parameters, perform sparsification processing on at least some of the plurality of operators to obtain an initial sparse model;

[0009] Perform iterative training on the initial sparse model to obtain a target sparse model.

[0010] According to another aspect of the embodiments of the present disclosure, there is provided a training device for a neural network model, including:

[0011] A first determination module, configured to determine a plurality of operators in the neural network model that support sparsification;

[0012] A second determination module, configured to determine at least one sparsity sensitivity corresponding to at least one sparsity condition for each operator in the plurality of operators determined by the first determination module;

[0013] A third determination module, configured to determine the sparsity parameters of the plurality of operators based on at least one sparsity sensitivity corresponding to each operator in the plurality of operators determined by the second determination module;

[0014] A processing module, configured to perform sparse processing on at least some of the multiple operators based on the sparse parameters determined by the third determination module to obtain an initial sparse model;

[0015] A training module, configured to iteratively train the initial sparse model obtained by the processing module to obtain a target sparse model.

[0016] According to another aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium storing a computer program for executing the above-mentioned training method of the neural network model.

[0017] According to another aspect of the embodiments of the present disclosure, there is provided an electronic device, including:

[0018] A processor;

[0019] A memory for storing executable instructions of the processor;

[0020] The processor is configured to read the executable instructions from the memory and execute the instructions to implement the above-mentioned training method of the neural network model.

[0021] According to another aspect of the embodiments of the present disclosure, there is provided a computer program product, when the instructions in the computer program product are executed by a processor, the above-mentioned training method of the neural network model is executed.

[0022] Based on the training method, device, storage medium, electronic device and program product of the neural network model provided in the above embodiments of the present disclosure, for multiple operators in the neural network model that support sparsification, at least one sparsity sensitivity corresponding to each operator for at least one sparsity condition can be evaluated to clarify the influence degree of different operators on the output of the neural network model under each sparsity condition. Based on these influence degrees, the sparse parameters of multiple operators can be reasonably determined, so as to be used for the sparse processing of the neural network model. It should be noted that there are usually some redundant parameters in the neural network model. By performing sparse processing on the neural network model based on the sparse parameters, it is beneficial to remove this part of redundant parameters and obtain an initial sparse model that is more lightweight and has lower complexity than the neural network model. By iteratively training the initial sparse model, the model parameters can be continuously optimized, and the adverse effects that sparse processing may bring to the model accuracy can be eliminated, so as to obtain a target sparse model with high accuracy for deployment on hardware (such as a neural network accelerator). Therefore, by adopting the embodiments of the present disclosure, the finally deployed model on the hardware has the characteristics of lightweight, low complexity, high accuracy, etc., thereby improving the model calculation speed. Description of the Drawings

[0023] Figure 1It is one of the schematic flowcharts of the training method of the neural network model provided by some exemplary embodiments of the present disclosure.

[0024] Figure 2 It is the second of the schematic flowcharts of the training method of the neural network model provided by some exemplary embodiments of the present disclosure.

[0025] Figure 3 It is the third of the schematic flowcharts of the training method of the neural network model provided by some exemplary embodiments of the present disclosure.

[0026] Figure 4 It is the fourth of the schematic flowcharts of the training method of the neural network model provided by some exemplary embodiments of the present disclosure.

[0027] Figure 5 It is the fifth of the schematic flowcharts of the training method of the neural network model provided by some exemplary embodiments of the present disclosure.

[0028] Figure 6 It is the sixth of the schematic flowcharts of the training method of the neural network model provided by some exemplary embodiments of the present disclosure.

[0029] Figure 7 It is the seventh of the schematic flowcharts of the training method of the neural network model provided by some exemplary embodiments of the present disclosure.

[0030] Figure 8 It is the eighth of the schematic flowcharts of the training method of the neural network model provided by some exemplary embodiments of the present disclosure.

[0031] Figure 9 It is the ninth of the schematic flowcharts of the training method of the neural network model provided by some exemplary embodiments of the present disclosure.

[0032] Figure 10-1 It is the tenth of the schematic flowcharts of the training method of the neural network model provided by some exemplary embodiments of the present disclosure.

[0033] Figure 10-2 It is the eleventh of the schematic flowcharts of the training method of the neural network model provided by some exemplary embodiments of the present disclosure.

[0034] Figure 10-3 It is the twelfth of the schematic flowcharts of the training method of the neural network model provided by some exemplary embodiments of the present disclosure.

[0035] Figure 11 It is the first of the structural schematic diagrams of the training device of the neural network model provided by some exemplary embodiments of the present disclosure.

[0036] Figure 12 It is the second of the structural schematic diagrams of the training device of the neural network model provided by some exemplary embodiments of the present disclosure.

[0037] Figure 13 It is the third schematic structural diagram of the training device for a neural network model provided by some exemplary embodiments of the present disclosure.

[0038] Figure 14 It is the fourth schematic structural diagram of the training device for a neural network model provided by some exemplary embodiments of the present disclosure.

[0039] Figure 15 It is the fifth schematic structural diagram of the training device for a neural network model provided by some exemplary embodiments of the present disclosure.

[0040] Figure 16 It is the sixth schematic structural diagram of the training device for a neural network model provided by some exemplary embodiments of the present disclosure.

[0041] Figure 17 It is the seventh schematic structural diagram of the training device for a neural network model provided by some exemplary embodiments of the present disclosure.

[0042] Figure 18 It is the schematic structural diagram of an electronic device provided by some exemplary embodiments of the present disclosure. Detailed implementation manners

[0043] To explain the present disclosure, exemplary embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present disclosure, rather than all embodiments. It should be understood that the present disclosure is not limited by the exemplary embodiments.

[0044] It should be noted that: unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions and values set forth in these embodiments do not limit the scope of the present disclosure.

[0045] Application Overview

[0046] Neural network models are widely used in many fields such as computer vision, natural language processing, speech recognition, and autonomous driving. For example, using a neural network model, environmental images collected by an in-vehicle camera can be subjected to object detection, object tracking, instance segmentation, etc.

[0047] It can be understood that if a neural network model is to implement a computing function, it generally includes multiple computing units, and these computing units can be called operators, such as convolution operators, pooling operators, deconvolution operators, rectified linear unit (ReLU) operators, elementwise operators, etc.

[0048] In the context of the development of deep learning technology, which has led to an increasing scale of neural network models, how to improve the model's computational speed is an issue worthy of attention for those skilled in the art.

[0049] Exemplary System

[0050] The following text involves the sparsification of the model. It can be understood that sparsification is a compression method for the model, which avoids unnecessary storage and computation by reducing the number of non-zero elements in the model parameters, thereby improving the model's computational speed.

[0051] In the embodiments of the present disclosure, for multiple operators in the neural network model that support sparsification, at least one sparsity sensitivity corresponding to at least one sparsity condition can be evaluated for each operator, so as to clarify the influence degree of different operators on the output of the neural network model under various sparsity conditions. Based on these influence degrees, the sparsity parameters of multiple operators can be reasonably determined, and thus used for the sparsification process of the neural network model. By iteratively training the sparsified neural network model, a target sparse model for deployment on hardware can be obtained. In this way, through the sparsification of the model, the model's computational speed can be effectively improved, that is, the model inference performance can be improved.

[0052] Exemplary Method

[0053] Figure 1 It is a schematic flowchart of a training method for a neural network model provided by some exemplary embodiments of the present disclosure. Figure 1 The method shown may include step 110, step 120, step 130, step 140, and step 150.

[0054] Step 110, determine multiple operators in the neural network model that support sparsification.

[0055] Optionally, the neural network model can be any model to be sparsified. The neural network model may include N operators arranged in sequence, where N is an integer greater than or equal to 2. From the N operators included in the neural network model, M operators that support sparsification can be selected, where M is an integer less than or equal to N.

[0056] The following gives an example of the method for selecting M operators that support sparsification from N operators.

[0057] First, determine the target operators carrying learnable parameters. Assume that among N operators, some operators carry learnable parameters while others do not. Then, these operators carrying learnable parameters can be determined from the N operators. For example, operators that do not carry learnable parameters (i.e., only carry non-learnable parameters) can be pooling operators, ReLU operators, Elementwise operators, etc.

[0058] Secondly, according to the hardware units for executing the operators carrying learnable parameters, determine the operators that support sparsification. A neural network model usually needs to run on hardware such as a Neural network Processing Unit (NPU). For any one of the operators carrying learnable parameters among the N operators (for ease of description, it will be referred to as the target operator hereinafter), the target hardware unit for running the target operator can be determined from the neural network accelerator first. For example, if the target operator is a convolution operator, the target hardware unit can be a tensor calculation unit. Another example is that if the target operator is a point-to-point element-wise operator, the target hardware unit can be a vector calculation unit. The target hardware unit usually has requirements for the operator regarding sparsification. For example, the target hardware unit may require that the size of the convolution kernel is different from 1*1, or the target hardware unit may require that the number of input channels is greater than or equal to 4. If the target operator meets the sparsification requirements of the target hardware unit for the operator, it can be determined that the target operator supports sparsification. If the target operator does not meet the sparsification requirements of the target hardware unit for the operator, it can be determined that the target operator does not support sparsification. In the above way, each operator that supports sparsification can be determined from the operators carrying learnable parameters. In this way, M operators that support sparsification can be efficiently and reliably screened out from the N operators.

[0059] Step 120, determine at least one sparsity sensitivity of each operator among the multiple operators corresponding to at least one sparsity condition.

[0060] Optionally, the at least one sparsity condition can be expressed as K sparsity conditions, where K is an integer greater than or equal to 1. The sparsity condition can be understood as information indicating the sparsification method. The sparsification method can be roughly divided into a structured sparsity method (which usually needs to follow a specific pattern or structure) and an unstructured sparsity method (which usually does not need to follow a specific pattern or structure). The structured sparsity method can be further divided into channel sparsity, layer sparsity, block sparsity, etc. The unstructured sparsity method can be further divided into sparsification at different sparsity rates. The sparsity rate can be, for example, 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9, etc.

[0061] Optionally, for each of the M operators that support sparsification, K sparsity sensitivities corresponding to K sparsity conditions for the operator can be determined. The sparsity sensitivity of any operator corresponding to any sparsity condition is used to characterize the impact of sparsifying the operator according to the sparsity condition on the output of the neural network model. For example, the greater the sparsity sensitivity of the operator corresponding to the sparsity condition, the greater the impact of sparsifying the operator according to the sparsity condition on the accuracy of the output of the neural network model; the smaller the sparsity sensitivity of the operator corresponding to the sparsity condition, the smaller the impact of sparsifying the operator according to the sparsity condition on the accuracy of the output of the neural network model.

[0062] The following gives an example of the method for determining the sparsity sensitivity of any operator corresponding to any sparsity condition.

[0063] First, input data can be obtained. The input data can be, for example, an image captured by a camera, or a sequence of images captured by a camera. Next, the input data can be computed through the neural network model after sparsifying the operator according to the sparsity condition to obtain first output data, and the input data can be computed through the neural network model without sparsification to obtain second output data. After that, the similarity between the first output data and the second output data can be computed. Based on the similarity, the sparsity sensitivity of the operator corresponding to the sparsity condition can be determined. The sparsity sensitivity and the similarity can be negatively correlated. For example, a first function with a negative correlation between the independent variable and the dependent variable (such as a linear function with a slope less than 0, an exponential function with a base between 0 and 1, etc.) can be preset, and the similarity can be used as the value of the independent variable to be substituted into the first function for operation, and the corresponding value of the dependent variable can be obtained, which can be used as the sparsity sensitivity of the operator corresponding to the sparsity condition. Alternatively, the cross entropy between the first output data and the second output data can be computed. Based on the cross entropy, the sparsity sensitivity of the operator corresponding to the sparsity condition can be determined. The sparsity sensitivity and the cross entropy can be positively correlated. For example, a second function with a positive correlation between the independent variable and the dependent variable (such as a linear function with a slope greater than 0, an exponential function with a base greater than 1, etc.) can be preset, and the cross entropy can be used as the value of the independent variable to be substituted into the second function for operation, and the corresponding value of the dependent variable can be obtained, which can be used as the sparsity sensitivity of the operator corresponding to the sparsity condition.

[0064] Step 130: Determine the sparsity parameters of multiple operators based on at least one sparsity sensitivity corresponding to each of the multiple operators.

[0065] Optionally, in the manner introduced above, the K sparse sensitivities corresponding to each of the M operators that support sparsification can be determined. Based on these sparse sensitivities, the sparse parameters of the M operators can be determined. The sparse parameters can be used to indicate the operators among the M operators that need to be sparsified. Additionally, for the operators among the M operators that need to be sparsified, the sparse parameters can further indicate the corresponding sparsification methods adapted to them.

[0066] Step 140: Based on the sparse parameters, perform sparsification on at least some of the multiple operators to obtain an initial sparse model.

[0067] As introduced above, the sparse parameters are used to indicate the operators among the M operators that need to be sparsified and the sparsification methods adapted to the operators that need to be sparsified. Then, for the neural network model, the operators that need to be sparsified indicated by the sparse parameters can be actually sparsified according to the sparsification methods indicated by the sparse parameters to obtain a sparsified neural network model, and the sparsified neural network model can be used as the initial sparse model.

[0068] Step 150: Perform iterative training on the initial sparse model to obtain a target sparse model.

[0069] Optionally, a large amount of training data can be used to perform iterative training on the initial sparse model. The training data can be, for example, images or image sequences with annotation information. Through the initial sparse model, the training data can be calculated to obtain third output data. The third output data can be regarded as prediction data, and the annotation information carried by the training data can be regarded as ground truth data. By comparing the prediction data and the ground truth data, the model loss value can be determined. Using the model loss value, the parameters of the initial policy model can be optimized through gradient backpropagation. After several rounds of parameter optimization, the initial policy model can converge, and at this time, the initial policy model after parameter optimization can be used as the target sparse model. The target sparse model can be used for deployment on hardware such as neural network accelerators.

[0070] In embodiments of the present disclosure, for multiple operators in a neural network model that support sparsification, at least one sparsity sensitivity corresponding to each operator for at least one sparsity condition can be evaluated to clarify the degree of influence of different operators on the output of the neural network model under each sparsity condition. Based on these degrees of influence, the sparsity parameters of the multiple operators can be reasonably determined and thus used for the sparsification process of the neural network model. It should be noted that there are usually some redundant parameters in the neural network model. By sparsifying the neural network model based on the sparsity parameters, it is beneficial to remove this part of redundant parameters and obtain an initial sparse model that is lighter and has lower complexity than the neural network model. By iteratively training the initial sparse model, the model parameters can be continuously optimized to eliminate the possible adverse effects of the sparsification process on the model accuracy, thereby obtaining a target sparse model with high accuracy for deployment on hardware (such as a neural network accelerator). Therefore, by adopting the embodiments of the present disclosure, the model finally deployed on the hardware has characteristics such as light weight, low complexity, and high accuracy, thereby improving the model calculation speed.

[0071] In some alternative examples, Figure 1 At least one sparsity condition involved in the illustrated embodiments can be multiple sparsity conditions, and the multiple sparsity conditions can include multiple reference sparsity rates for unstructured sparsity. Determining at least one sparsity sensitivity corresponding to each operator among the multiple operators for at least one sparsity condition includes: determining multiple sparsity sensitivities corresponding to each operator among the multiple operators based on the multiple reference sparsity rates.

[0072] As introduced above, at least one sparsity condition can be expressed as K sparsity conditions. Here, K can be an integer greater than or equal to 2, and each sparsity condition among the K sparsity conditions can be a reference sparsity rate, and the K sparsity conditions can be K reference sparsity rates. As an example, K can be 9, and the 9 reference sparsity rates can be 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 respectively. Or, K can be 5, and the 5 reference sparsity rates can be 0.1, 0.3, 0.5, 0.7, 0.9 respectively. Or, K can be 4, and the 4 reference sparsity rates can be 0.2, 0.4, 0.6, 0.8 respectively.

[0073] For each of the M operators that support sparsification, K sparsity sensitivities corresponding one-to-one to K reference sparsity rates can be determined. For example, the neural network model obtained by sparsifying the operator according to the K reference sparsity rates respectively can be used to calculate the input data to obtain K first output data corresponding one-to-one to the K reference sparsity rates. The input data can be calculated by the neural network model without sparsification to obtain the second output data. Based on the K first output data and the second output data, by calculating the similarity or cross entropy, K sparsity sensitivities corresponding one-to-one to the K reference sparsity rates can be determined.

[0074] In the embodiments of the present disclosure, by determining the multiple sparsity sensitivities corresponding to each of the multiple operators based on multiple reference sparsity rates, the influence degree of each operator on the output of the neural network model using each reference sparsity rate can be clarified. In this way, it is beneficial to determine an appropriate sparsity rate for each operator to obtain the initial sparse model, so as to improve the rationality and reliability of the sparse parameters.

[0075] In some optional examples, as Figure 2 shown, determining the sparse parameters of multiple operators based on at least one sparsity sensitivity corresponding to each of the multiple operators includes steps 210, 220, 230, 240, and 250.

[0076] Step 210, determine the baseline sparsity rate among the multiple reference sparsity rates.

[0077] Among them, the baseline sparsity rate can be any one of the K reference sparsity rates. Optionally, the reference sparsity rate with the middle numerical value among the K reference sparsity rates can be used as the baseline sparsity rate. For example, if the K reference sparsity rates are 9 reference sparsity rates, which are 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 respectively, then 0.5 can be used as the baseline sparsity rate. Of course, any one of 0.4, 0.6, 0.7 can also be used as the baseline sparsity rate, and the present disclosure does not limit this.

[0078] Step 220, based on the multiple sparsity sensitivities corresponding to each of the multiple operators, using the baseline sparsity rate as the starting sparsity rate, successively decreasing the sparsity rate according to the screening rule of the first number of operators with the highest sparsity sensitivity and increasing the sparsity rate of the second number of operators with the lowest sparsity sensitivity, select the sparsity rate to be used corresponding to each of the multiple operators from the multiple reference sparsity rates.

[0079] Optionally, the first number can be represented as m, the second number can be represented as n, and both m and n are integers greater than or equal to 1. The values of m and n are the same or different. In addition, the values of m and n can be preset or determined according to actual needs.

[0080] After determining the K sparse sensitivities corresponding to each of the M operators that support sparsification and determining the baseline sparsity rate from the K sparse sensitivities, the baseline sparsity rate can be used as the starting sparsity rate, and the sparsity rate can be successively decreased according to the m operators with higher sparse sensitivities and increased according to the n operators with lower sparse sensitivities. From the K reference sparsity rates, the sparsity rate to be used corresponding to each of the M operators is screened. Among them, the sparsity rate to be used corresponding to any operator is the reference sparsity rate to be verified for its suitability with the operator.

[0081] In an optional example, K is 9, and the 9 reference sparsity rates are 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, and 0.9 respectively. M is 5, and the 5 operators are Operator 1, Operator 2, Operator 3, Operator 4, and Operator 5. Each of Operator 1, Operator 2, Operator 3, Operator 4, and Operator 5 corresponds to 9 sparse sensitivities, so there can be a total of 45 sparse sensitivities, and the 45 sparse sensitivities can form a first set. Assuming the baseline sparsity rate is 0.5, then 0.5 can be used as the starting sparsity rate, and 5 sparse sensitivities corresponding to 0.5 as the starting sparsity rate are selected from the first set to form a second set. Assuming that among the second set, the sparse sensitivity corresponding to Operator 1 is denoted as s1, the sparse sensitivity corresponding to Operator 2 is denoted as s2, the sparse sensitivity corresponding to Operator 3 is denoted as s3, the sparse sensitivity corresponding to Operator 4 is denoted as s4, and the sparse sensitivity corresponding to Operator 5 is denoted as s5, then s1 to s5 can be arranged in descending order, and the sorting result can be, for example, s2, s3, s4, s5, s1. Assuming that both m and n are 2, obviously, for the second set, the first number of operators with higher sparse sensitivities are Operator 2 corresponding to s2 and Operator 3 corresponding to s3, and the first number of operators with lower sparse sensitivities are Operator 5 corresponding to s5 and Operator 1 corresponding to s1. Then, the sparsity rate can be decreased for Operator 2 and Operator 3 to obtain the sparsity rates to be used corresponding to Operator 2 and Operator 3 respectively. For example, 0.4 can be determined as the sparsity rates to be used corresponding to Operator 2 and Operator 3 respectively. In addition, the sparsity rate can be increased for Operator 5 and Operator 1 to obtain the sparsity rates to be used corresponding to Operator 5 and Operator 1 respectively. For example, 0.6 can be determined as the sparsity rates to be used corresponding to Operator 5 and Operator 1 respectively. In addition, 0.5 can also be determined as the sparsity rate to be used corresponding to Operator 4. In this way, the first operation of determining the sparsity rates to be used corresponding to Operator 1 to Operator 5 is completed.

[0082] For the introduction in the previous paragraph, when increasing or decreasing the sparsity for any operator, the step size can be regarded as 1. For example, the sparsity rate of the operator increases from 0.5 to 0.6, or decreases from 0.5 to 0.4. In specific implementation, the step size can also be 2, 3, etc. For example, the sparsity rate of the operator can increase from 0.5 to 0.7, or decrease from 0.5 to 0.3.

[0083] If it is necessary to perform the operation of determining the corresponding sparsity rate to be used for operators 1 to 5 for the second time, then the sparsity sensitivities corresponding to operators 1 to 5 for the corresponding sparsity rates to be used determined in the previous paragraph can be selected from the first set to form the third set. Assume that in the third set, the sparsity sensitivity of operator 1 corresponding to the reference sparsity rate of 0.6 is denoted as r1, the sparsity sensitivity of operator 2 corresponding to the reference sparsity rate of 0.4 is denoted as r2, the sparsity sensitivity of operator 3 corresponding to the reference sparsity rate of 0.4 is denoted as r3, the sparsity sensitivity of operator 4 corresponding to the reference sparsity rate of 0.5 is denoted as r4, and the sparsity sensitivity of operator 5 corresponding to the reference sparsity rate of 0.6 is denoted as r5. Then, r1 to r5 can be arranged in descending order, and the sorting result can be, for example, r4, r5, r2, r1, r3. Assume that both m and n are 2. Obviously, for the third set, the first number of operators with the highest sparsity sensitivity are operator 4 corresponding to r4 and operator 5 corresponding to r5, and the first number of operators with the lowest sparsity sensitivity are operator 1 corresponding to r1 and operator 3 corresponding to r3. Then, the sparsity rate can be decreased for operator 4 and operator 5 to obtain the sparsity rates to be used corresponding to operator 4 and operator 5 respectively. For example, 0.4 can be determined as the sparsity rate to be used corresponding to operator 4, and 0.5 can be determined as the sparsity rate to be used corresponding to operator 5. In addition, the sparsity rate can be increased for operator 1 and operator 3 to obtain the sparsity rates to be used corresponding to operator 1 and operator 3 respectively. For example, 0.7 can be determined as the sparsity rate to be used corresponding to operator 1, and 0.5 can be determined as the sparsity rate to be used corresponding to operator 3. In addition, 0.4 can also be determined as the sparsity rate to be used corresponding to operator 2. In this way, the operation of determining the corresponding sparsity rates to be used for operators 1 to 5 for the second time is completed.

[0084] Subsequently, the operation of determining the corresponding sparsity rates to be used for operators 1 to 5 can be performed again. The specific determination method can refer to the introduction in the above two paragraphs and will not be elaborated here.

[0085] Step 230, for each sparsity rate to be used selected, determine the first accuracy evaluation value and the first performance evaluation value of the neural network model.

[0086] Among them, for any sparsity rate to be used selected each time, the corresponding first accuracy evaluation value is a value used to characterize the accuracy of the neural network model in the target state, and the corresponding first performance evaluation value is a value used to characterize the computing speed of the neural network model in the target state; among them, the target state refers to: the state after unstructured sparsification of M operators supporting sparsification in the neural network model according to the sparsity rate to be used selected this time.

[0087] Optionally, for any sparsity rate to be used selected each time, according to these sparsity rates to be used, unstructured sparsification processing can be performed on M operators supporting sparsification in the neural network model to obtain the neural network model after this sparsification processing. For example, after completing the first operation of determining the corresponding sparsity rates to be used for operators 1 to 5, since the sparsity rates to be used corresponding to operators 1 to 5 are 0.6, 0.4, 0.4, 0.5, and 0.6 in sequence, unstructured sparsification can be performed on operators 1 to 5 in the neural network model, and the sparsity rate of operator 1 is 0.6, the sparsity rate of operator 2 is 0.4, the sparsity rate of operator 3 is 0.4, the sparsity rate of operator 4 is 0.5, and the sparsity rate of operator 5 is 0.6. Thus, the neural network model after this sparsification processing can be obtained. Through the neural network model after this sparsification processing, the input data can be calculated to obtain the fourth output data. Next, the similarity between the fourth output data and the second output data in the above text can be calculated, and the calculated similarity can be used as the first accuracy evaluation value. In addition, the first quantity of input data that the neural network model after this sparsification processing can process within a unit time (such as 1 second) can be determined, and the second quantity of input data that the neural network model without sparsification processing can process within a unit time can be determined. The ratio of the first value to the second quantity can be used as the first performance evaluation value.

[0088] Step 240: Based on the first accuracy evaluation value, the first performance evaluation value, and the sparsity rate to be used selected this time, determine the target sparsity rate adapted to each operator among the multiple operators.

[0089] Optionally, for any sparsity rate to be used selected each time, based on the first accuracy evaluation value and the first performance evaluation value, it can be inferred whether each sparsity rate to be used among these sparsity rates to be used is suitable for the corresponding operator. According to the inference result, the corresponding determination method can be adopted to determine the target sparsity rate adapted to each operator among the multiple operators.

[0090] In some optional embodiments of the present disclosure, as Figure 3 shown, step 240 may include step 310 and step 320.

[0091] Step 310, in response to the first precision evaluation value meeting the first preset precision condition and the first performance evaluation value meeting the first preset performance condition, for each operator among the multiple operators, determine the to-be-used sparsity rate corresponding to the current operator screening as the target sparsity rate adapted to the operator.

[0092] Step 320, in response to the first precision evaluation value meeting the second preset precision condition, and / or the first performance evaluation value meeting the second preset performance condition, trigger the next operation of screening the to-be-used sparsity rate corresponding to each operator among the multiple operators from the multiple reference sparsity rates.

[0093] As introduced above, for the to-be-used sparsity rate screened out in any screening, the fourth output data can be calculated from the input data through the to-be-used sparsity rate, and the similarity between the fourth output data and the second output data can be used as the first precision evaluation value. The first precision evaluation value can be compared with the preset similarity. The preset similarity can be, for example, 85%, 90%, 95%, etc., which will not be listed one by one here. If the first precision evaluation value is greater than the preset similarity, it indicates that the precision of the neural network model after the current sparse processing based on these to-be-used sparsity rates meets the requirements. Then, it can be determined that the first precision evaluation value meets the first preset precision condition. If the first precision evaluation value is less than or equal to the preset similarity, it indicates that the precision of the neural network model after the current sparse processing based on these to-be-used sparsity rates does not meet the requirements. Then, it can be determined that the first precision evaluation value does not meet the first preset precision condition, but meets the second preset precision condition.

[0094] As introduced above, for the to-be-used sparsity rate screened out in any screening, the first quantity and the second quantity can be determined, and the ratio of the first quantity to the second quantity can be used as the first performance evaluation value. The first performance evaluation value can be compared with the preset ratio. The preset ratio can be, for example, 1, 1.2, 1.5, etc., which will not be listed one by one here. If the first performance evaluation value is greater than the preset ratio, it indicates that the performance (specifically, the calculation speed) of the neural network model after the current sparse processing based on these to-be-used sparsity rates meets the requirements. Then, it can be determined that the first performance evaluation value meets the first preset performance condition. If the first performance evaluation value is less than or equal to the preset ratio, it indicates that the calculation speed of the neural network model after the current sparse processing based on these to-be-used sparsity rates does not meet the requirements. Then, it can be determined that the first performance evaluation value does not meet the first preset performance condition, but meets the second preset performance condition.

[0095] If, for the sparsity rate to be used selected in a certain screening, it is determined that the first precision evaluation value meets the first preset precision condition and the first performance evaluation value meets the first preset performance condition, that is, the precision and computational speed of the neural network model after the current sparsity processing based on these sparsity rates to be used can both meet the requirements, it can be inferred that each sparsity rate to be used among these sparsity rates to be used is adapted to the corresponding operator. Then, for each of the M operators, the sparsity rate to be used selected for this operator in the current screening can be determined as the target sparsity rate corresponding to this operator.

[0096] If, for the sparsity rate to be used selected in a certain screening, it is determined that the first precision evaluation value meets the second preset precision condition, and / or the first performance evaluation value meets the second preset performance condition, that is, the precision and / or computational speed of the neural network model after the current sparsity processing based on these sparsity rates to be used cannot meet the requirements, it can be inferred that at least some of the sparsity rates to be used among these sparsity rates to be used are not adapted to the corresponding operators. Then, the operation of screening the sparsity rate to be used for the next time can be performed according to the screening rule of gradually decreasing the sparsity rate for the first number of operators with the highest sparsity sensitivity and gradually increasing the sparsity rate for the second number of operators with the lowest sparsity sensitivity. The subsequent process can be inferred by analogy and will not be elaborated in detail here.

[0097] In this way, by combining the evaluation values in the two dimensions of precision and computational speed, it can be determined whether the precision and computational speed of the neural network model can meet the requirements under the influence of the sparsity rate to be used selected in the current screening. Based on this, it can be determined whether to use the sparsity rate to be used selected in the current screening as the target sparsity rate, or to perform the operation of screening the sparsity rate to be used for the next time to search for other sparsity rates to be used as the target sparsity rate. In this way, it is beneficial to improve the rationality and reliability of the target sparsity rate.

[0098] Step 250: Determine the sparse parameters based on the target sparsity rate adapted to each of the multiple operators.

[0099] Optionally, the sparse parameters may include the target sparsity rate adapted to each of the M operators that support sparsification.

[0100] In some embodiments, the sparse parameters may further include the sparsity rates of the remaining operators in the neural network model other than these M operators, and these sparsity rates may all be 0.

[0101] In the embodiments of the present disclosure, a baseline sparsity rate among multiple reference sparsity rates can be determined, and the baseline sparsity rate can be used as the starting sparsity rate. Then, according to the screening rule of successively reducing the sparsity rate by the first number of operators with the highest sparsity sensitivity and increasing the sparsity rate by the second number of operators with the lowest sparsity sensitivity, the sparsity rate to be used corresponding to each operator among multiple operators that support sparsification is screened from the multiple reference sparsity rates. Additionally, for each sparsity rate to be used screened each time, model evaluation can be performed from two dimensions of accuracy and computing speed to obtain corresponding evaluation values, thereby determining whether each sparsity rate to be used is suitable for the corresponding operator. On this basis, for each operator among the multiple operators, a target sparsity rate that is suitable can be determined. For example, for an operator that has a greater impact on accuracy, its corresponding target sparsity rate can be relatively small to avoid a drop in accuracy; for an operator that has a smaller impact on accuracy, its corresponding target sparsity rate can be relatively large to improve the computing speed. In this way, using the sparsity parameters determined based on the target sparsity rate for the sparsification of the neural network model is beneficial to improving the accuracy and computing speed of the finally obtained target sparse model. It should be noted that since specific screening rules are adopted in the embodiments of the present disclosure to screen the sparsity rate to be used and thereby determine the target sparsity rate, there is no need to use a brute-force search method to determine the target sparsity rate, which is beneficial to finding better sparsity parameters in a relatively short time, thus helping to save time and computing power costs.

[0102] In some alternative examples, Figure 1 At least one sparsity condition involved in the illustrated embodiments can be a single sparsity condition, and the sparsity condition can include target parameters for structured sparsity. Determining at least one sparsity sensitivity corresponding to each operator among the multiple operators for at least one sparsity condition includes: determining the sparsity sensitivity corresponding to each operator among the multiple operators based on the target parameters.

[0103] Optionally, the sparsity condition can include only one parameter, for example, only the target parameters. The target parameter can be, for example, Structured Sparsity. In this way, the sparsity processing method indicated by the sparsity condition can be a structured sparsity method.

[0104] Optionally, for each of the M operators that support sparsification, one sparsity sensitivity corresponding to the operator can be determined based on the target parameters. For example, the neural network model after sparsifying the operator according to the target parameters can be used to calculate the input data to obtain first output data. The neural network model without sparsification can be used to calculate the input data to obtain second output data. Based on the first output data and the second output data, the sparsity sensitivity corresponding to the operator can be determined by calculating the similarity or cross-entropy.

[0105] In the embodiments of the present disclosure, by determining the sparse sensitivity corresponding to each operator among multiple operators based on target parameters, the influence degree of each operator performing structured sparsity on the output of the neural network model can be clarified. In this way, it is beneficial to evaluate whether each operator should perform structured sparsity in order to obtain the initial sparse model.

[0106] In some alternative examples, as Figure 4 shown, determining the sparse parameters of multiple operators based on at least one sparse sensitivity corresponding to each operator among multiple operators includes step 410, step 420, and step 430.

[0107] Step 410, based on the sparse sensitivity corresponding to each operator among multiple operators, successively screen non-sparse operators from multiple operators in the order of decreasing sparse sensitivity, and successively increase the number of non-sparse operators.

[0108] Among them, a non-sparse operator refers to an operator that does not perform structured sparsity.

[0109] Optionally, based on the sparse sensitivity corresponding to each operator among M operators that support sparsification, non-sparse operators can be successively screened from the M operators in the order of decreasing sparse sensitivity, and the number of non-sparse operators can be successively increased. For example, when screening non-sparse operators from the M operators for the first time, the operator with the largest sparse sensitivity among the M operators can be used as the non-sparse operator. When screening non-sparse operators from the M operators for the second time, the operator with the largest sparse sensitivity and the operator with the second largest sparse sensitivity among the M operators can be used as a non-sparse operator respectively. When screening non-sparse operators from the M operators for the third time, the operator with the largest sparse sensitivity, the operator with the second largest sparse sensitivity, and the operator with the third largest sparse sensitivity among the M operators can be used as a non-sparse operator respectively. The subsequent process can be inferred by analogy, and will not be elaborated in detail here.

[0110] Step 420, for each time the non-sparse operators are screened, determine the second accuracy evaluation value and the second performance evaluation value of the neural network model.

[0111] Optionally, for any non-sparse operators screened out, structured sparsity can be performed on the remaining operators among the M operators that support sparsification in the neural network model except for these non-sparse operators to obtain the neural network model after the current sparse processing. Based on the neural network model after the current sparse processing, the neural network model without sparse processing, and the input data, the second accuracy evaluation value and the second performance evaluation value can be determined. The specific determination methods and meanings of the second accuracy evaluation value and the second performance evaluation value can refer to Figure 2 the introduction of the first accuracy evaluation value and the first performance evaluation value in the embodiments shown, and will not be elaborated here.

[0112] Step 430: Determine the sparse parameters of multiple operators based on the second precision evaluation value and the second performance evaluation value.

[0113] Optionally, for any non-sparse operator selected in a screening, it can be determined whether the accuracy and computational speed of the neural network model can meet the requirements by only performing structured sparsity on the remaining operators among the M operators except these non-sparse operators based on the second precision evaluation value and the second performance evaluation value. According to the determination result, a corresponding determination method can be adopted to determine the sparse parameters.

[0114] In some alternative embodiments of the present disclosure, as Figure 5 shown, step 430 may include step 510 and step 520.

[0115] Step 510: In response to the second precision evaluation value satisfying the first preset precision condition and the second performance evaluation value satisfying the first preset performance condition, determine the remaining operators among the multiple operators except the non-sparse operators selected in this screening, and determine the sparse parameters used to represent that the remaining operators are to be subjected to structured sparsity.

[0116] Step 520: In response to the second precision evaluation value satisfying the second preset precision condition, and / or the second performance evaluation value satisfying the second preset performance condition, trigger the next operation of screening non-sparse operators from the multiple operators.

[0117] Optionally, the method of determining whether the second precision evaluation value satisfies the first preset precision condition or the second preset precision condition, and the method of determining whether the second performance evaluation value satisfies the first preset performance condition or the second preset performance condition can refer to Figure 3 the introduction of the relevant determination methods in the shown embodiments, and will not be elaborated here.

[0118] For any non-sparse operator selected in a screening, if the second precision evaluation value satisfies the first preset precision condition and the second performance evaluation value satisfies the first preset performance condition, that is, the accuracy and performance of the neural network model after this sparse processing obtained by performing structured sparsity on the remaining operators among the M operators except these non-sparse operators can both meet the requirements, the sparse parameters can be used to represent that these remaining operators are to be subjected to structured sparsity.

[0119] For any non-sparse operator selected in a screening, if the second precision evaluation value meets the second preset precision condition, and / or the second performance evaluation value meets the second preset performance condition, that is, for the M operators, the precision and / or performance of the neural network model after the current sparse processing obtained by performing structured sparsity on the remaining operators except these non-sparse operators cannot meet the requirements. Then, the next operation of screening non-sparse operators can be performed in the order from the largest to the smallest according to the sparse sensitivity, and the number of non-sparse operators is increased. The subsequent process is carried out in the same way and will not be elaborated here.

[0120] In this way, by combining the evaluation values in the two dimensions of precision and calculation speed, it can be determined whether the precision and calculation speed of the neural network model can meet the requirements when only performing structured sparsity on the remaining operators except the non-sparse operators selected in the current screening. Based on this, it can be determined whether to determine the sparse parameters used to represent the corresponding remaining operators to be structurally sparse based on the non-sparse operators selected in the current screening, or to perform the next operation of screening non-sparse operators to increase the number of non-sparse operators. In this way, the optimal number of non-sparse operators can be found, so as to determine the sparse parameters accordingly, which is beneficial to improving the rationality and reliability of the sparse parameters.

[0121] In the above introduction, for two adjacent operations of screening non-sparse operators, the number of non-sparse operators obtained in the latter screening is 1 greater than the number of non-sparse operators obtained in the previous screening. Optionally, if the value of M is relatively large, such as 50, 100, etc., the number of non-sparse operators obtained in the latter screening can also be 2, 3, 4, etc. greater than the number of non-sparse operators obtained in the previous screening, and will not be listed one by one here.

[0122] In the embodiments of the present disclosure, based on the sparse sensitivity corresponding to each of the multiple operators, non-sparse operators can be sequentially screened from the multiple operators in descending order of sparse sensitivity, and the number of non-sparse operators can be sequentially increased. In addition, for each screened non-sparse operator, the model can be evaluated from two dimensions of accuracy and computing speed to obtain corresponding evaluation values, so as to determine whether the accuracy and computing speed of the neural network model can meet the requirements when only the remaining operators except these non-sparse operators among the multiple operators are structured sparsified. On this basis, each operator that should be structured sparsified can be searched to obtain corresponding sparse parameters. For example, for an operator with a greater impact on accuracy, it can be determined that it should not be structured sparsified; for an operator with a smaller impact on accuracy, it can be determined that it should be structured sparsified. In this way, using the sparse parameters for sparsifying the neural network model is beneficial to improving the accuracy and computing speed of the finally obtained target sparse model. It should be noted that since specific screening rules are adopted in the embodiments of the present disclosure for screening non-sparse operators, there is no need to use a brute-force search method to determine non-sparse operators, which is beneficial to finding better sparse parameters in a shorter time, thereby saving time and computing power costs.

[0123] In some alternative examples, as Figure 6 shown, step 140 includes step 610, step 620, step 630, and step 640.

[0124] Step 610, based on the neural network model, determine the first parameter matrix of each of the multiple operators.

[0125] Optionally, if the neural network model carries the first parameter matrix of each of the M operators that support sparsification, then the first parameter matrix of each of the M operators can be extracted from the neural network model; wherein, all relevant parameters of an operator can be recorded in the first parameter matrix of the operator, for example, all weight parameters of the operator can be recorded. Each parameter recorded in each first parameter matrix can also be called an element. Each parameter recorded in each first parameter matrix can be a learnable parameter.

[0126] Step 620, based on the sparse parameters, determine the elements to be sparsified in the first parameter matrix of each of the multiple operators.

[0127] As introduced above, the sparse parameters can indicate the operators that need to be sparsified among the M operators, and indicate the sparsification method adapted to the operators that need to be sparsified. For each of the M operators, there can be the following possible situations:

[0128] (1) If the sparse parameters indicate that the operator does not need to be sparsified, then there are no elements to be sparsified in the first parameter matrix of the operator.

[0129] (2) The sparse parameter includes a target sparsity rate adapted to the operator. Assuming the target sparsity rate is 0.4, 40% of the elements can be selected from the first parameter matrix of the operator, and each of the selected 40% of the elements can be used as a to-be-sparse element.

[0130] (3) The sparse parameter is used to characterize that the operator is to perform structured sparsity. Assuming the number of channels of the first parameter matrix of the operator is 4, and the sparse parameter is specifically used to characterize that the operator is to perform semi-structured sparsity of selecting 2 out of 4, then for every 4 elements in the first parameter matrix of the operator that are in the same row, the same column, and different channels, 2 elements can be selected from them, and each of the selected 2 elements can be used as a to-be-sparse element.

[0131] Step 630: Determine the mask matrices corresponding to the multiple operators based on the to-be-sparse elements in the respective first parameter matrices of the multiple operators; wherein, in the mask matrix corresponding to any one of the multiple operators, the positions corresponding to the to-be-sparse elements in the corresponding first parameter matrix are the first numerical values used to characterize the zeroing type, and the remaining positions are the second numerical values used to characterize the retention type.

[0132] Optionally, for each of the M operators that support sparsification, the mask matrix corresponding to the operator can be determined according to the following rule: the mask matrix corresponding to the operator has the same size as the first parameter matrix of the operator, and if the element at the i-th row, j-th column, and c-th channel in the first parameter matrix of the operator is a to-be-sparse element, then the element at the i-th row, j-th column, and c-th channel in the mask matrix corresponding to the operator is the first numerical value used to characterize the zeroing type. After all the positions where the first numerical value needs to be placed in the mask matrix corresponding to the operator are determined, the remaining positions are all filled with the second numerical values used to characterize the retention type. Here, each mask matrix can be a Boolean-type matrix, so the first numerical value used to characterize the zeroing type can be 0, and the second numerical value used to characterize the retention type can be 1.

[0133] Step 640: Update the respective first parameter matrices of the multiple operators by using the respective mask matrices of the multiple operators to obtain an initial sparse model.

[0134] Optionally, for each of the M operators, the mask matrix corresponding to the operator can be multiplied by the first parameter matrix of the operator to obtain the updated parameter matrix of the operator. After the parameter matrices of all the M operators in the neural network model are updated, it can be considered that a neural network model after sparse processing is obtained, and the neural network model after sparse processing can be used as the initial sparse model.

[0135] In the embodiments of the present disclosure, based on sparse parameters, the elements to be sparsified in the first parameter matrices of multiple operators can be determined efficiently and reliably. Accordingly, the mask matrices corresponding to the multiple operators can be reasonably determined. By using the mask matrices corresponding to the multiple operators, the update of the first parameter matrices of the multiple operators can be efficiently and reliably implemented through operations such as multiplication to set the elements to be sparsified to zero, thus realizing the sparsification of the neural network model, and thereby an initial sparse model can be efficiently and reliably obtained.

[0136] As Figure 7 shown, after step 630, the method provided by the embodiments of the present disclosure may further include step 710, step 720, step 730, and step 740.

[0137] Step 710: Perform an inversion process on the mask matrices corresponding to the multiple operators to obtain the inversion results corresponding to the multiple operators.

[0138] Optionally, for each of the M operators that support sparsification, the inversion result corresponding to the operator may be determined according to the following rule: update the elements that are 0 in the mask matrix corresponding to the operator to 1, and update the elements that are 1 in the mask matrix corresponding to the operator to 0 to obtain the updated result of the mask matrix, and the updated result of the mask matrix may be used as the inversion result corresponding to the operator.

[0139] Step 720: Store the inversion results corresponding to the multiple operators in a memory.

[0140] Optionally, the memory may include but is not limited to a static random access memory (SRAM), a register bank, etc. The inversion results corresponding to the M operators can all be stored in the memory, and these inversion results can be read from the memory subsequently for further use.

[0141] Step 730: Based on the initial sparse model, determine the second parameter matrices of the multiple operators.

[0142] Optionally, the specific implementation manner of step 730 may refer to the relevant introduction to step 610 and will not be elaborated here.

[0143] Step 740: Use the inversion results stored in the memory to perform a restoration process on the second parameter matrices of the multiple operators to obtain the elements to be sparsified corresponding to the multiple operators.

[0144] Optionally, for each of the M operators, the corresponding negated result of the operator can be read from the memory. By multiplying the negated result corresponding to the operator by the second parameter matrix of the operator, the first parameter matrix of the operator can be obtained, and the elements to be sparsified corresponding to the operator exist in the first parameter matrix of the operator. In this way, it is equivalent to restoring the elements to be sparsified corresponding to the operator through the restoration process.

[0145] In the embodiments of the present disclosure, by negating the respective mask matrices corresponding to multiple operators, the corresponding negated results of the multiple operators can be obtained efficiently and reliably for storage in the memory. Subsequently, using the negated results stored in the memory, through the restoration process, the elements to be sparsified corresponding to the multiple operators can be restored efficiently and reliably. On this basis, a neural network model without sparse processing can be obtained for use of the neural network model without sparse processing.

[0146] In some alternative examples, as Figure 8 shown, the iterative training of the initial sparse model in step 150 includes step 810, step 820, and step 830.

[0147] Step 810: Based on the first training data and the initial sparse model, determine the first gradient matrices corresponding to the multiple operators for backpropagation.

[0148] Optionally, the first training data may include a large amount of training data. The training data may be, for example, an image or an image sequence with annotation information. Through the initial sparse model, the training data can be calculated to obtain the third output data. The third output data can be regarded as prediction data, and the annotation information carried by the training data can be regarded as ground-truth data. By comparing the prediction data and the ground-truth data, the model loss value can be determined. Using the model loss value, the first gradient matrices corresponding to the M operators that support sparsification for backpropagation can be determined.

[0149] Step 820: Use the respective mask matrices corresponding to the multiple operators to update the first gradient matrices corresponding to the multiple operators to obtain the second gradient matrices corresponding to the multiple operators.

[0150] Optionally, for each of the M operators, the mask matrix corresponding to the operator can be multiplied by the first gradient matrix corresponding to the operator to set the gradients corresponding to the elements to be sparsified in the first gradient matrix to zero, thereby obtaining the second gradient matrix corresponding to the operator.

[0151] Step 830: Optimize the parameters of the initial sparse model according to the second gradient matrices corresponding to the multiple operators.

[0152] Optionally, according to the second gradient matrices respectively corresponding to the M operators, the gradients can be backpropagated through a conventional optimizer for model optimization to achieve parameter optimization of the initial sparse model.

[0153] In the embodiments of the present disclosure, by using the mask matrices respectively corresponding to the multiple operators to update the first gradient matrices respectively corresponding to the multiple operators, the gradients corresponding to the elements to be sparsified in the first gradient matrices can be set to zero. In this way, during the process of backpropagating the gradients according to the updated second gradient matrices, the positions corresponding to the elements to be sparsified can remain zero and will not change to non-zero due to the backpropagation of the gradients, which is beneficial to avoiding the loss of sparsity of the model.

[0154] In some alternative examples, as Figure 9 shown, the iterative training of the initial sparse model further includes step 910, step 920, step 930, and step 940.

[0155] Step 910, after optimizing the parameters of the initial sparse model according to the second gradient matrices respectively corresponding to the multiple operators, based on the initial sparse model after parameter optimization, determine the third parameter matrices respectively corresponding to the multiple operators.

[0156] Optionally, the specific implementation manner of step 910 can refer to the relevant introduction to step 610, which will not be elaborated here.

[0157] Step 920, use the mask matrices respectively corresponding to the multiple operators to update the third parameter matrices respectively corresponding to the multiple operators to obtain the initial sparse model with the updated parameter matrices.

[0158] Optionally, the specific implementation manner of step 920 can refer to the relevant introduction to step 640, which will not be elaborated here.

[0159] Step 930, based on the second training data and the initial sparse model with the updated parameter matrices, determine the third gradient matrices respectively corresponding to the multiple operators for backpropagation.

[0160] Optionally, the specific implementation manner of step 930 can refer to the relevant introduction to step 810 above, which will not be elaborated here.

[0161] Step 940, based on the third gradient matrices respectively corresponding to the multiple operators and the mask matrices respectively corresponding to the multiple operators, optimize the parameters of the initial sparse model with the updated parameter matrices.

[0162] Optionally, for each of the M operators, the mask matrix corresponding to the operator can be multiplied by the third gradient matrix corresponding to the operator to set the gradients corresponding to the elements to be sparsified in the third gradient matrix to zero, so as to obtain the fourth gradient matrix corresponding to the operator. According to the fourth gradient matrices respectively corresponding to the M operators, the gradients can be backpropagated through an optimizer to implement parameter optimization of the initial sparse model with the parameter matrix updated.

[0163] It should be noted that during the forward process of the model, for some operators, the elements that were originally set to zero in the parameter matrix may become non-zero. Such an operator can be, for example, a batch normalization (BN) operator. This situation may cause the model to lose sparsity. In view of this, in the embodiments of the present disclosure, after parameter optimization of the initial sparse model according to the second gradient matrices respectively corresponding to the multiple operators, the third parameter matrices respectively corresponding to the multiple operators can be determined based on the initial sparse model with parameters optimized, and the third parameter matrices respectively corresponding to the multiple operators can be updated by using the mask matrices respectively corresponding to the multiple operators to set the positions corresponding to the elements to be sparsified to zero again, so as to obtain the initial sparse model with the parameter matrix updated. The initial sparse model with the parameter matrix updated can maintain sparsity. Combining the third gradient matrices respectively corresponding to the multiple operators and the mask matrices respectively corresponding to the multiple operators to perform parameter optimization on the initial sparse model with the parameter matrix updated can, on the basis of the model maintaining sparsity, obtain the target sparse model through iterative training, so that the target sparse model has sparsity.

[0164] In some alternative examples, a converged floating-point model can be trained according to a conventional model training method, for example, obtaining the Figure 10-1 floating-point model in Figure 10-1 The floating-point model in can be used as the neural network model to be sparsified. Next, M operators in the neural network model that support sparsification can be determined, and the K sparse sensitivities corresponding to each operator for K sparse conditions can be evaluated to clarify the influence degree of each operator on the output of the neural network model under each sparse condition. Based on these sparse sensitivities, a sparsification configuration can be determined, that is, the sparse parameters mentioned above can be determined (for example, the sparse parameters can be determined in the manner shown in the embodiment of Figure 2 or determined in the manner introduced in the embodiment shown in Figure 4 ). According to the sparse parameters, at least some of the M operators in the neural network model can be sparsified to perform model sparsification (for example, the method shown in Figure 6Implement model sparsification in the manner of the illustrated embodiment to obtain an initial sparse model. Subsequently, the initial sparse model can be iteratively trained to converge the model to the best accuracy, thereby obtaining a target sparse model for deployment on hardware. Compared with the neural network model, the accuracy of the target sparse model has almost no loss, but the amount of computation and the number of parameters are significantly reduced, and the inference performance can be significantly improved.

[0165] In some other alternative examples, Figure 10-2 the floating-point model in can be used as the neural network model to be sparsified. Next, all operators in the neural network model can be traversed in sequence, and it can be determined whether the traversed operator supports sparsification.

[0166] If the traversed operator supports sparsification, the K sparsity sensitivities corresponding to K sparsity conditions of the traversed operator can be determined. For example, the sparsity sensitivity of the traversed operator at each of the sparsity rates of 0.1, 0.2, 0.3, 0.4, 0.5, 0.6, 0.7, 0.8, 0.9 can be determined to obtain 9 sparsity sensitivities. Next, it can be determined whether the traversed operator is the last operator. If the traversed operator is not the last operator, then continue to traverse the next operator; if the traversed operator is the last operator, then based on all the obtained sparsity sensitivities, the sparsity parameter can be determined (for example, the sparsity parameter can be determined in the manner of the illustrated embodiment). Figure 2 the manner of the illustrated embodiment to determine the sparsity parameter).

[0167] If the traversed operator does not support sparsification, it can be determined whether the traversed operator is the last operator. If the traversed operator is not the last operator, then continue to traverse the next operator; if the traversed operator is the last operator, then based on all the obtained sparsity sensitivities, the sparsity parameter can be determined (for example, the sparsity parameter can be determined in the manner of the illustrated embodiment). Figure 2 the manner of the illustrated embodiment to determine the sparsity parameter).

[0168] According to the sparsity parameter, at least some of the M operators in the neural network model can be sparsified to implement model sparsification (for example, the model sparsification can be implemented in the manner of the illustrated embodiment) to obtain an initial sparse model. Subsequently, the initial sparse model can be iteratively trained to converge the model to the best accuracy, thereby obtaining a target sparse model for deployment on hardware. Compared with the neural network model, the accuracy of the target sparse model has almost no loss, but the amount of computation and the number of parameters are significantly reduced, and the inference performance can be significantly improved. Figure 6 the manner of the illustrated embodiment to implement model sparsification) to obtain an initial sparse model. Subsequently, the initial sparse model can be iteratively trained to converge the model to the best accuracy, thereby obtaining a target sparse model for deployment on hardware. Compared with the neural network model, the accuracy of the target sparse model has almost no loss, but the amount of computation and the number of parameters are significantly reduced, and the inference performance can be significantly improved.

[0169] In still some other alternative examples, Figure 10-3The floating-point model therein can be a neural network model to be sparsified. Next, all operators in the neural network model can be traversed in sequence, and it can be determined whether the traversed operator supports sparsification.

[0170] If the traversed operator supports sparsification, K sparsity sensitivities corresponding to K sparsity conditions for the traversed operator can be determined. To determine the sparsity sensitivity of the traversed operator corresponding to any sparsity condition, the neural network model after sparsification of the operator according to the sparsity condition and the neural network model without sparsification can be used to calculate the input data respectively to obtain the first output data and the second output data, and the corresponding sparsity sensitivity can be determined based on the first output data and the second output data. Optionally, the way to obtain the neural network model after sparsification of the operator according to the sparsity condition can be: determining the first parameter matrix of the operator; based on the sparsity condition, determining the elements to be sparsified in the first parameter matrix of the operator to obtain the corresponding mask matrix; by multiplying the mask matrix with the first parameter matrix of the operator in the neural network model, the update of the parameter matrix of the operator in the neural network model is realized, so as to obtain the neural network model after sparsification of the operator according to the sparsity condition. By performing an inversion process on the mask matrix, an inversion result can be obtained, and the inversion result can be stored in the memory. In addition, it can be determined whether the traversed operator is the last operator. If the traversed operator is not the last operator, continue to traverse the next operator; if the traversed operator is the last operator, the sparsity parameter can be determined based on all the obtained sparsity sensitivities (for example, the sparsity parameter can be determined in the manner shown in the embodiment Figure 2 or the way introduced in the embodiment shown in Figure 4 ).

[0171] If the traversed operator does not support sparsification, it can be determined whether the traversed operator is the last operator. If the traversed operator is not the last operator, continue to traverse the next operator; if the traversed operator is the last operator, the sparsity parameter can be determined based on all the obtained sparsity sensitivities (for example, the sparsity parameter can be determined in the manner shown in the embodiment Figure 2 or the way introduced in the embodiment shown in Figure 4 ).

[0172] According to the sparsity parameter, at least some of the M operators in the neural network model can be sparsified to obtain an initial sparse model through model sparsification. After that, the initial sparse model can be iteratively trained to obtain a target sparse model. During the iterative training process, the gradient can be backpropagated through the optimizer, and the gradient used for backpropagation can be corrected to avoid the loss of model sparsity (for example, it can be adoptedFigure 8 The illustrated embodiments avoid sparsity loss). Additionally, after the optimizer backpropagates the gradients, for the operators in the M operators that have undergone sparsification processing, the mask matrix can be multiplied by the current parameter matrix to avoid the loss of sparsity in the forward process operators.

[0173] In summary, the embodiments of the present disclosure introduce the concept of sparse sensitivity. Based on the influence degree of different operators on the output of the neural network model under different sparse conditions, better sparse parameters can be found in a relatively short time. Sparse processing is performed on the neural network model according to the found sparse parameters, and further iterative training is carried out, which can improve the accuracy and calculation speed of the finally obtained target sparse model. Additionally, the sparse parameters in the embodiments of the present disclosure do not depend on manual experience but are obtained based on some specific screening rules, and are interpretable. The determination method of the sparse parameters can be transferred to the sparse tasks of different models.

[0174] Exemplary Device

[0175] Figure 11 It is a schematic structural diagram of a training device for a neural network model provided by some exemplary embodiments of the present disclosure. Figure 11 The illustrated device may include:

[0176] A first determination module 1110, configured to determine multiple operators in the neural network model that support sparsification;

[0177] A second determination module 1120, configured to determine at least one sparse sensitivity corresponding to at least one sparse condition for each operator among the multiple operators determined by the first determination module 1110;

[0178] A third determination module 1130, configured to determine sparse parameters for the multiple operators based on at least one sparse sensitivity corresponding to each operator among the multiple operators determined by the second determination module 1120;

[0179] A processing module 1140, configured to perform sparse processing on at least some of the multiple operators based on the sparse parameters determined by the third determination module 1130 to obtain an initial sparse model;

[0180] A training module 1150, configured to perform iterative training on the initial sparse model obtained by the processing module 1140 to obtain a target sparse model.

[0181] In some alternative examples, the at least one sparse condition is multiple sparse conditions, and the multiple sparse conditions include multiple reference sparse rates for unstructured sparsity;

[0182] The second determination module 1120 is configured to determine multiple sparse sensitivities corresponding to each operator among the multiple operators based on the multiple reference sparse rates.

[0183] In some alternative examples, such as Figure 12 shown, the second determination module 1120 includes:

[0184] The first determination sub-module 1210 is configured to determine a benchmark sparsity rate among multiple reference sparsity rates;

[0185] The second determination sub-module 1220 is configured to, based on the multiple sparsity sensitivities corresponding to each operator among the multiple operators determined by the second determination module 1120, use the benchmark sparsity rate determined by the first determination sub-module 1210 as the starting sparsity rate, and successively decrease the sparsity rate according to the screening rule of the first number of operators with the highest sparsity sensitivity and increase the sparsity rate of the second number of operators with the lowest sparsity sensitivity, so as to screen the sparsity rate to be used corresponding to each operator among the multiple operators from the multiple reference sparsity rates;

[0186] The third determination sub-module 1230 is configured to determine a first accuracy evaluation value and a first performance evaluation value of the neural network model for the sparsity rate to be used screened by the second determination sub-module 1220 each time;

[0187] The fourth determination sub-module 1240 is configured to determine the target sparsity rate adapted to each operator among the multiple operators based on the first accuracy evaluation value and the first performance evaluation value determined by the third determination sub-module 1230, and the sparsity rate to be used screened by the second determination sub-module 1220 this time;

[0188] The fifth determination sub-module 1250 is configured to determine the sparsity parameter based on the target sparsity rate adapted to each operator among the multiple operators determined by the fourth determination sub-module 1240.

[0189] In some alternative examples, the fifth determination sub-module 1250 includes:

[0190] The first determination unit is configured to, in response to the first accuracy evaluation value determined by the third determination sub-module 1230 satisfying the first preset accuracy condition and the first performance evaluation value determined by the third determination sub-module 1230 satisfying the first preset performance condition, determine the sparsity rate to be used corresponding to the operator screening this time as the target sparsity rate adapted to each operator;

[0191] The second determination unit is configured to, in response to the first accuracy evaluation value determined by the third determination sub-module 1230 satisfying the second preset accuracy condition, and / or the first performance evaluation value determined by the third determination sub-module 1230 satisfying the second preset performance condition, trigger the next operation of screening the sparsity rate to be used corresponding to each operator among the multiple operators from the multiple reference sparsity rates.

[0192] In some alternative examples, at least one sparse condition is a sparse condition, and the sparse condition includes target parameters for structured sparsity;

[0193] A second determination module 1120, configured to determine the sparse sensitivity corresponding to each operator among multiple operators based on the target parameters.

[0194] In some alternative examples, as Figure 13 shown, the second determination module 1120 includes:

[0195] A screening sub-module 1310, configured to successively screen non-sparse operators from multiple operators in descending order of sparse sensitivity based on the sparse sensitivity corresponding to each operator among the multiple operators determined by the second determination module 1120, and successively increase the number of non-sparse operators;

[0196] A sixth determination sub-module 1320, configured to determine a second accuracy evaluation value and a second performance evaluation value of the neural network model for each time the non-sparse operators screened by the screening sub-module 1310;

[0197] A seventh determination sub-module 1330, configured to determine the sparse parameters of multiple operators based on the second accuracy evaluation value and the second performance evaluation value determined by the sixth determination sub-module 1320.

[0198] In some alternative examples, the seventh determination sub-module 1330 includes:

[0199] A third determination unit, configured to, in response to the second accuracy evaluation value determined by the sixth determination sub-module 1320 satisfying a first preset accuracy condition and the second performance evaluation value determined by the sixth determination sub-module 1320 satisfying a first preset performance condition, determine the remaining operators among the multiple operators except the non-sparse operators screened this time, and determine the sparse parameters used to characterize the remaining operators to be subject to structured sparsity;

[0200] A fourth determination unit, configured to, in response to the second accuracy evaluation value determined by the sixth determination sub-module 1320 satisfying a second preset accuracy condition, and / or the second performance evaluation value determined by the sixth determination sub-module 1320 satisfying a second preset performance condition, trigger the next operation of screening non-sparse operators from multiple operators.

[0201] In some alternative examples, as Figure 14 shown, the processing module 1140 includes:

[0202] An eighth determination sub-module 1410, configured to determine the first parameter matrix of each of the multiple operators based on the neural network model;

[0203] The ninth determination sub-module 1420 is configured to determine the elements to be sparsified in the first parameter matrix of each of the multiple operators based on the sparsity parameters determined by the third determination module 1130;

[0204] The tenth determination sub-module 1430 is configured to determine the mask matrix corresponding to each of the multiple operators based on the elements to be sparsified in the first parameter matrix of each of the multiple operators determined by the ninth determination sub-module 1420; wherein, in the mask matrix corresponding to any one of the multiple operators, the positions corresponding to the elements to be sparsified in the corresponding first parameter matrix are the first numerical values used to represent the zeroing type, and the remaining positions are the second numerical values used to represent the retention type;

[0205] The first update sub-module 1440 is configured to update the first parameter matrix of each of the multiple operators by using the mask matrix corresponding to each of the multiple operators determined by the tenth determination sub-module 1430 to obtain an initial sparse model.

[0206] In some alternative examples, as Figure 15 shown, the apparatus provided by the embodiments of the present disclosure further includes:

[0207] The negation module 1510 is configured to perform a negation process on the mask matrix corresponding to each of the multiple operators after the tenth determination sub-module 1430 determines the mask matrix corresponding to each of the multiple operators to obtain the negation result corresponding to each of the multiple operators;

[0208] The storage module 1520 is configured to store the negation result corresponding to each of the multiple operators obtained by the negation module 1510 in a memory;

[0209] The fourth determination module 1530 is configured to determine the second parameter matrix of each of the multiple operators based on the initial sparse model obtained by the processing module 1140;

[0210] The restoration module 1540 is configured to perform a restoration process on the second parameter matrix of each of the multiple operators determined by the fourth determination module 1530 by using the negation result stored by the storage module 1520 in the memory to obtain the elements to be sparsified corresponding to each of the multiple operators.

[0211] In some alternative examples, as Figure 16 shown, the training module 1150 includes:

[0212] The eleventh determination sub-module 1610 is configured to determine the first gradient matrix for backpropagation corresponding to each of the multiple operators based on the first training data and the initial sparse model obtained by the processing module 1140;

[0213] A second update sub-module 1620, configured to update a first gradient matrix corresponding to each of the multiple operators determined by an eleventh determination sub-module 1610 by using a mask matrix corresponding to each of the multiple operators determined by a tenth determination sub-module 1430, so as to obtain a second gradient matrix corresponding to each of the multiple operators;

[0214] A first optimization sub-module 1630, configured to optimize parameters of an initial sparse model obtained by a processing module 1140 according to the second gradient matrix corresponding to each of the multiple operators obtained by the second update sub-module 1620.

[0215] In some alternative examples, as Figure 17 shown, the training module 1150 further includes:

[0216] A twelfth determination sub-module 1710, configured to determine a third parameter matrix corresponding to each of the multiple operators based on the initial sparse model after parameter optimization by a parameter optimization sub-module;

[0217] A third update sub-module 1720, configured to update the third parameter matrix corresponding to each of the multiple operators determined by the twelfth determination sub-module 1710 by using the mask matrix corresponding to each of the multiple operators determined by the tenth determination sub-module 1430, so as to obtain an initial sparse model with updated parameter matrix;

[0218] A thirteenth determination sub-module 1730, configured to determine a third gradient matrix corresponding to each of the multiple operators for backpropagation based on the second training data and the initial sparse model with updated parameter matrix obtained by the third update sub-module 1720;

[0219] A second optimization sub-module 1740, configured to optimize parameters of the initial sparse model with updated parameter matrix obtained by the third update sub-module 1720 based on the third gradient matrix corresponding to each of the multiple operators determined by the thirteenth determination sub-module 1730 and the mask matrix corresponding to each of the multiple operators determined by the tenth determination sub-module 1430.

[0220] In the apparatus of the present disclosure, the various alternative embodiments, alternative implementation manners, and alternative examples disclosed above can be flexibly selected and combined as needed to achieve corresponding functions and effects, and the present disclosure does not list them one by one.

[0221] Exemplary Electronic Device

[0222] Figure 18 The block diagram of an electronic device according to an embodiment of the present disclosure is illustrated. The electronic device 1800 includes one or more processors 1810 and a memory 1820.

[0223] The processor 1810 can be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and can control other components in the electronic device 1800 to perform desired functions.

[0224] The memory 1820 can include one or more computer program products, and the computer program products can include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory can include, for example, random access memory (RAM) and / or cache memory, etc. Non-volatile memory can include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions can be stored on the computer-readable storage media, and the processor 1810 can run one or more computer program instructions to implement the methods of the various embodiments of the present disclosure described above and / or other desired functions.

[0225] In one example, the electronic device 1800 can further include: an input device 1830 and an output device 1840, and these components are interconnected through a bus system and / or other form of connection mechanism (not shown).

[0226] The input device 1830 can further include, for example, a keyboard, a mouse, etc.

[0227] The output device 1840 can output various information to the outside, which can include, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.

[0228] Of course, for simplicity, Figure 18 only some of the components in the electronic device 1800 related to the present disclosure are shown, and components such as a bus, an input / output interface, etc. are omitted. In addition, according to specific application scenarios, the electronic device 1800 can further include any other appropriate components.

[0229] Exemplary Computer Program Product and Computer Readable Storage Medium

[0230] In addition to the above methods and devices, an embodiment of the present disclosure can also be a computer program product, which includes computer program instructions that, when run by a processor, cause the processor to execute the steps in the methods according to various embodiments of the present disclosure described in the "Exemplary Methods" section above of this specification.

[0231] A computer program product may write program code for performing the operations of the embodiments of the present disclosure in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and also including conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computing device, partially on the user's device, executed as a stand-alone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.

[0232] In addition, an embodiment of the present disclosure may also be a computer-readable storage medium having computer program instructions stored thereon, and when the computer program instructions are run by a processor, the processor is caused to execute the steps in the methods according to various embodiments of the present disclosure described in the "Exemplary Method" section above of this specification.

[0233] The computer-readable storage medium may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. The readable storage medium may, for example, include but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0234] The basic principles of the present disclosure have been described above in connection with specific embodiments. However, the advantages, benefits, effects, etc. mentioned in the present disclosure are only examples and not limitations, and it cannot be considered that these advantages, benefits, effects, etc. are essential for each embodiment of the present disclosure. The specific details disclosed above are only for the purposes of illustration and easy understanding, and not for limitation. The above details do not limit the present disclosure to necessarily adopt the above specific details for implementation.

[0235] Those skilled in the art can make various changes and modifications to the present disclosure without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present disclosure and their equivalent technologies, the present disclosure also intends to include these changes and modifications.

Claims

1. A training method for a neural network model, comprising: Identify multiple operators in the neural network model that support sparsification; determining at least one sparse sensitivity of each operator of the plurality of operators corresponding to at least one sparse condition; Determining sparse parameters of the multiple operators based on at least one sparse sensitivity corresponding to each operator of the multiple operators; Based on the sparse parameters, sparse processing is performed on at least part of the multiple operators to obtain an initial sparse model; The initial sparse model is iteratively trained to obtain a target sparse model.

2. The method according to claim 1, wherein: The determining of at least one sparse sensitivity of each operator of the plurality of operators corresponding to at least one sparse condition comprises: The at least one sparse condition is a plurality of sparse conditions, and the plurality of sparse conditions include a plurality of reference sparse rates for unstructured sparseness; Based on the multiple reference sparsity rates, multiple sparsity sensitivities corresponding to each of the multiple operators are determined.

3. The method according to claim 2, wherein: The determining, based on at least one sparse sensitivity corresponding to each of the multiple operators, sparse parameters of the multiple operators comprises: determining a base thinning rate among the plurality of reference thinning rates; Based on the multiple sparse sensitivities corresponding to each of the multiple operators, taking the reference sparse rate as the starting sparse rate, successively reducing the sparse rate of a first number of operators with the highest sparse sensitivity and increasing the sparse rate of a second number of operators with the lowest sparse sensitivity according to the screening rule, and screening the sparse rate to be used corresponding to each of the multiple operators from the multiple reference sparse rates; Determining a first accuracy evaluation value and a first performance evaluation value of the neural network model for each screened sparsity rate to be used; Determine a target sparsity rate adapted for each of the multiple operators based on the first accuracy evaluation value, the first performance evaluation value, and the sparsity rate to be used in this screening; The sparse parameter is determined based on a target sparsity rate adapted for each of the multiple operators.

4. The method according to claim 3, wherein: The determining, based on the first precision evaluation value, the first performance evaluation value, and the sparsity rate to be used in this screening, a target sparsity rate adapted for each of the multiple operators includes: In response to the first accuracy evaluation value satisfying a first preset accuracy condition, the first performance evaluation value satisfying a first preset performance condition, for each of the multiple operators, determining the to-be-used sparsity rate corresponding to the operator screening this time as a target sparsity rate adapted for the operator; In response to the first precision evaluation value satisfying the second preset precision condition, and / or the first performance evaluation value satisfying the second preset performance condition, the next operation of screening the sparse rate to be used corresponding to each operator in the multiple operators from the multiple reference sparsity rates is triggered.

5. The method according to claim 1, wherein: The determining of at least one sparse sensitivity of each operator of the plurality of operators corresponding to at least one sparse condition comprises: The at least one sparse condition is a sparse condition, wherein the sparse condition includes a target parameter for structured sparsity; Based on the target parameter, a sparse sensitivity corresponding to each operator of the multiple operators is determined.

6. The method according to claim 5, wherein: The determining, based on at least one sparse sensitivity corresponding to each of the multiple operators, sparse parameters of the multiple operators comprises: Based on the sparse sensitivity corresponding to each operator in the multiple operators, non-sparse operators are selected from the multiple operators in descending order of the sparse sensitivity, and the number of the non-sparse operators is increased successively; For each non-sparse operator screened, determining a second accuracy evaluation value and a second performance evaluation value of the neural network model; Based on the second accuracy evaluation value and the second performance evaluation value, sparse parameters of the plurality of operators are determined.

7. The method according to claim 6, wherein: The determining, based on the second precision evaluation value and the second performance evaluation value, the sparse parameters of the plurality of operators comprises: In response to the second precision evaluation value satisfying the first preset precision condition, the second performance evaluation value satisfying the first preset performance condition, determining the remaining operators in the multiple operators except the non-sparse operators screened this time, and determining the sparse parameters for characterizing the remaining operators to be structured sparsified; In response to the second precision evaluation value satisfying a second preset precision condition, and / or the second performance evaluation value satisfying a second preset performance condition, the next operation of screening non-sparse operators from the multiple operators is triggered.

8. The method according to any one of claims 1 to 7, wherein: The step of performing sparse processing on at least part of the multiple operators based on the sparse parameters to obtain an initial sparse model includes: Based on the neural network model, determining a first parameter matrix of each of the plurality of operators; Based on the sparse parameters, determining the elements to be sparse in the first parameter matrix of each of the plurality of operators; Based on the elements to be sparse in the first parameter matrices of the multiple operators, respectively, mask matrices corresponding to the multiple operators are determined; wherein, in the mask matrix corresponding to any operator of the multiple operators, the positions corresponding to the elements to be sparse in the corresponding first parameter matrix are first values ​​for characterizing the zeroing type, and the remaining positions are second values ​​for characterizing the retention type; Using the mask matrices corresponding to the multiple operators, the first parameter matrices of the multiple operators are updated to obtain an initial sparse model.

9. The method according to claim 8, wherein: After determining the mask matrices corresponding to the multiple operators based on the elements to be sparse in the first parameter matrices of the multiple operators, the method further includes: Performing inversion processing on the mask matrices corresponding to the multiple operators to obtain inversion results corresponding to the multiple operators; Storing the inverted results corresponding to each of the plurality of operators in a memory; Based on the initial sparse model, determining a second parameter matrix of each of the plurality of operators; The inversion results stored in the memory are used to restore the second parameter matrices of the multiple operators to obtain the to-be-sparse elements corresponding to the multiple operators.

10. The method according to claim 8, wherein: The iterative training of the initial sparse model comprises: Based on the first training data and the initial sparse model, determining a first gradient matrix for back propagation corresponding to each of the plurality of operators; Using the mask matrices corresponding to the multiple operators, respectively, the first gradient matrices corresponding to the multiple operators are updated to obtain the second gradient matrices corresponding to the multiple operators, respectively. Parameters of the initial sparse model are optimized according to the second gradient matrices corresponding to the multiple operators.

11. The method according to claim 10, wherein: The iterative training of the initial sparse model further includes: After performing parameter optimization on the initial sparse model according to the second gradient matrices corresponding to the multiple operators respectively, determining the third parameter matrices of the multiple operators respectively based on the initial sparse model after parameter optimization; Using the mask matrices corresponding to the multiple operators, respectively, the third parameter matrices of the multiple operators are updated to obtain the initial sparse model after the parameter matrix is ​​updated; Based on the second training data and the initial sparse model after the parameter matrix is ​​updated, determining a third gradient matrix for back propagation corresponding to each of the plurality of operators; Based on the third gradient matrices corresponding to the multiple operators and the mask matrices corresponding to the multiple operators, parameter optimization is performed on the initial sparse model after the parameter matrix is ​​updated.

12. A training device for a neural network model, comprising: A first determination module is used to determine multiple operators supporting sparsification in a neural network model; a second determination module, configured to determine at least one sparse sensitivity of each operator of the plurality of operators determined by the first determination module corresponding to at least one sparse condition; a third determination module, configured to determine sparse parameters of the multiple operators based on at least one sparse sensitivity corresponding to each of the multiple operators determined by the second determination module; A processing module, configured to perform sparse processing on at least part of the multiple operators based on the sparse parameters determined by the third determining module to obtain an initial sparse model; The training module is used to iteratively train the initial sparse model obtained by the processing module to obtain a target sparse model.

13. A computer-readable storage medium storing a computer program for executing the training method of the neural network model described in any one of claims 1 to 11.

14. An electronic device, comprising: processor; a memory for storing instructions executable by the processor; The processor is used to read the executable instructions from the memory and execute the instructions to implement the training method of the neural network model described in any one of claims 1-11 above.