A method, device and storage medium for accelerating reasoning of deep neural network

Through data-driven pruning and low-bit width quantization technology, the problem of inference efficiency of deep learning models in resource-constrained environments is solved, significantly reducing computational complexity and storage requirements, while maintaining model performance, and suitable for embedded systems and mobile devices.

CN119476356BActive Publication Date: 2025-05-16SHENZHEN EXTREME VISION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510073434.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-17
Publication Date
2025-05-16
Estimated Expiration
2045-01-17

AI Technical Summary

Technical Problem

With the expansion of the scale of deep learning models, the consumption of computing resources during training and inference has increased significantly, making it difficult to apply in resource-constrained environments. At the same time, the huge volume of the model increases storage and transmission costs, and the inference speed cannot meet the real-time application needs.

Method used

Through data-driven pruning and low-bit width quantization, the computational complexity and storage requirements of the model are significantly reduced. The specific steps include: determining the standard data set that matches the target task, building a deep neural network model, performing weight distribution and L1 norm recording, pruning neurons or connections with importance below the preset threshold, performing quantization processing, and optimizing model parameters using a joint loss function.

Benefits of technology

It significantly reduces the computational complexity and storage requirements of the model, while maintaining the model performance to the greatest extent, adapting to diverse task requirements, improving the model inference speed and resource utilization efficiency, and is especially suitable for resource-constrained environments such as embedded systems and mobile devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119476356B_ABST
    Figure CN119476356B_ABST
Patent Text Reader

Abstract

The present application discloses a method, device and storage medium for accelerating inference of a deep neural network. The method of the present application includes: preprocessing a standard data set, training a deep neural network model with a training set; recording the weight distribution and L1 norm of each layer of neurons or connections; determining the importance value of neurons or connections based on the weight distribution and L1 norm recorded during the training process; pruning neurons or connections whose importance values ​​are lower than a preset importance threshold according to a preset ratio, and dynamically adjusting the structure of the deep neural network model after each pruning; determining the quantization bit width and the upper and lower limits of quantization; for the pruned deep neural network model, based on the quantization bit width and the upper and lower limits of quantization, performing pseudo-quantization processing on weights and activation values; calculating the task loss and quantization error loss based on the result after pseudo-quantization, and updating the full-precision weights; optimizing the model parameters using a joint loss function.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer data processing technology, and in particular to a method, device and storage medium for accelerating inference of a deep neural network. Background Art

[0002] With the rapid development of deep learning technology, neural networks have achieved remarkable results in many fields, such as image recognition, natural language processing, speech recognition, etc. These technological breakthroughs rely heavily on large-scale neural network models that can automatically extract features and perform complex data analysis. However, as the scale of the model continues to expand, the complex model structure and the large number of parameters also bring many challenges.

[0003] First, the computing resource consumption during training and inference has increased significantly. These large models usually require a lot of computing power and high-performance hardware support, such as GPUs and TPUs, which makes them difficult to apply in resource-constrained environments. For example, edge devices and mobile devices usually have limited computing power and cannot process such large models, thus limiting the widespread application of deep learning technology in these fields.

[0004] Secondly, the large size of the model leads to increased storage and transmission costs. In some real-time application scenarios, such as voice assistants for smartphones and real-time image processing for driverless vehicles, fast response and low latency are required, and the inference speed of large models often cannot meet these requirements. In addition, as data continues to grow, the training time of the model is also increasing, increasing the difficulty of development and deployment. Summary of the invention

[0005] In order to solve the above technical problems, the present application provides a deep neural network accelerated reasoning method, device and storage medium.

[0006] The technical solution provided in this application is described below:

[0007] The first aspect of the present application provides a method for accelerating reasoning of a deep neural network, the method comprising:

[0008] Determine a standard data set that matches the target task, preprocess the standard data set, and divide it into a training set, a validation set, and a test set;

[0009] Constructing a deep neural network model, and using the training set to train the deep neural network model;

[0010] Using the validation set to validate the deep neural network model, and recording the weight distribution and L1 norm of each layer of neurons or connections;

[0011] Determining the importance value of the neuron or connection based on the weight distribution and L1 norm recorded during the training process;

[0012] Prune neurons or connections whose importance values ​​are lower than a preset importance threshold according to a preset ratio, and dynamically adjust the structure of the deep neural network model after each pruning;

[0013] Determine the quantization bit width and the upper and lower limits of quantization;

[0014] For the pruned deep neural network model, during the forward propagation process, the weights and activation values ​​are pseudo-quantized based on the quantization bit width and the upper and lower limits of quantization;

[0015] During the back propagation process, based on the pseudo-quantization results, the task loss and quantization error loss are calculated, and the full-precision weights are updated;

[0016] The model parameters are optimized using a joint loss function, where the joint loss function includes task loss and quantization error loss.

[0017] Optionally, determining the importance value of the neuron or connection based on the weight distribution and L1 norm recorded during the training process includes:

[0018] Record the statistical information of each layer weight, including mean, variance, maximum and minimum values;

[0019] The L1 norm of the weight is calculated by the following formula;

[0020] ;

[0021] Among them, n represents the number of weights, and w represents the weight;

[0022] The sparsity index S of the weight is calculated by the following formula:

[0023] ;

[0024] ;

[0025] Among them, S is the sparsity index of the weight, which is used to indicate the importance of the weight.

[0026] Optionally, pruning neurons or connections whose importance values ​​are lower than a preset importance threshold according to a preset ratio, and dynamically adjusting the structure of the deep neural network model after each pruning includes:

[0027] Get the importance value of each layer of neurons or connections and sort them from low to high;

[0028] The number of pruning n is determined according to the preset ratio p through the following formula:

[0029] n = (p × N);

[0030] Where p represents the preset ratio, and N represents the total number of neurons or connections in the current layer;

[0031] Remove neurons or connections whose importance is below a preset importance threshold;

[0032] Update the structure of the deep neural network model and delete the corresponding weights and biases.

[0033] Optionally, determining the quantization bit width and the quantization upper and lower limits includes:

[0034] For a given bit width b, use the training set or pruned weights and activation values ​​to calculate the maximum value x_max and the minimum value x_min;

[0035] The quantization upper and lower limits are the dynamic range, which is calculated by the following formula:

[0036] l=x_min,u=x_max;

[0037] Where l represents the lower limit of quantization, and u represents the upper limit of quantization.

[0038] Optionally, for the pruned deep neural network model, during the forward propagation process, pseudo-quantization processing is performed on weights and activation values ​​based on the quantization bit width and the quantization upper and lower limits, including:

[0039] The quantization step size is calculated by the following formula:

[0040] ;

[0041] in, Indicates the quantization step size;

[0042] Clip weights and activations to the range [l,u];

[0043] According to the quantization step size The clipped weights and activation values ​​are discretized. The discretization formula is:

[0044] ;

[0045] in, Represents discrete values;

[0046] The discrete values ​​obtained by the discretization are mapped back to floating point numbers by the following formula:

[0047] ;

[0048] in, Represents a floating point number.

[0049] Optionally, in the back propagation process, based on the result after pseudo quantization, calculating the task loss and the quantization error loss, and updating the full-precision weights include:

[0050] The loss of the classification task is calculated through the cross entropy function;

[0051] The loss of the regression task is calculated by the mean squared error;

[0052] Calculate the quantization error loss using the full-precision weights and the fake-quantized weights;

[0053] The loss of the classification task, the loss of the regression task, and the quantization error loss are fused through a predefined joint loss function;

[0054] Based on the joint loss function, the gradient is calculated by back-propagation and the full-precision weights are updated.

[0055] Optionally, the quantization error loss is calculated in the following manner:

[0056] ;

[0057] in, represents the quantization error loss, represents the full precision weight, represents the pseudo-quantized weight, Calculate using the following formula:

[0058] ;

[0059] Wherein, b represents the quantization bit width, and (u, l) represents the upper and lower limits of the quantization.

[0060] Optionally, the use of a joint loss function to optimize model parameters, wherein the joint loss function includes a task loss and a quantization error loss, comprises:

[0061] According to a preconfigured weight coefficient, the task loss and the quantization error loss are fused into a joint loss through weighted calculation;

[0062] In back-propagation, the joint loss is used to calculate the gradient and update the parameters of the deep neural network model.

[0063] A second aspect of the present application provides a deep neural network accelerated reasoning device, the device comprising:

[0064] A preprocessing unit, used to determine a standard data set matching the target task, preprocess the standard data set, and divide it into a training set, a validation set, and a test set;

[0065] A model building unit, used to build a deep neural network model and train the deep neural network model using the training set;

[0066] A verification unit, used to verify the deep neural network model using the verification set, and record the weight distribution and L1 norm of each layer of neurons or connections;

[0067] An importance evaluation unit, used to determine the importance value of the neuron or connection based on the weight distribution and L1 norm recorded during the training process;

[0068] A pruning unit, used to prune neurons or connections whose importance values ​​are lower than a preset importance threshold according to a preset ratio, and dynamically adjust the structure of the deep neural network model after each pruning;

[0069] A quantization determination unit, used to determine the quantization bit width and the quantization upper and lower limits;

[0070] A pseudo-quantization processing unit is used to perform pseudo-quantization processing on weights and activation values ​​of the pruned deep neural network model based on the quantization bit width and the upper and lower limits of quantization during the forward propagation process;

[0071] The loss calculation unit is used to calculate the task loss and quantization error loss based on the pseudo-quantization results during the back-propagation process, and update the full-precision weights;

[0072] An optimization unit is used to optimize model parameters using a joint loss function, wherein the joint loss function includes a task loss and a quantization error loss.

[0073] The third aspect of the present application provides a deep neural network accelerated reasoning device, the device comprising:

[0074] Processor, memory, input-output unit, and bus;

[0075] The processor is connected to the memory, the input and output unit, and the bus;

[0076] The memory stores a program, and the processor calls the program to execute the first aspect and any optional method in the first aspect.

[0077] A fourth aspect of the present application provides a computer-readable storage medium, on which a program is stored. When the program is executed on a computer, the program executes the first aspect and any optional method in the first aspect.

[0078] It can be seen from the above technical solutions that this application has the following advantages:

[0079] This deep neural network accelerated inference method significantly reduces the computational complexity and storage requirements of the model through data-driven pruning and low-bit-width quantization, while maximizing the model performance under the optimization of the joint loss function; its ability to dynamically adjust the structure adapts to diverse task requirements, improves the model inference speed and resource utilization efficiency, and is particularly suitable for low-power and efficient operation in resource-constrained environments such as embedded systems and mobile devices, and has good scalability and ease of deployment. BRIEF DESCRIPTION OF THE DRAWINGS

[0080] In order to more clearly illustrate the technical solution in the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0081] Figure 1 A schematic diagram of an embodiment of a deep neural network accelerated reasoning method provided in this application;

[0082] Figure 2 A schematic diagram of an embodiment of the flow chart of updating weights based on a joint loss function in a method for accelerating inference of a deep neural network provided in this application;

[0083] Figure 3 A flowchart of another embodiment of the method for accelerating inference of a deep neural network provided in this application;

[0084] Figure 4 This is a schematic diagram of the structure of an embodiment of a deep neural network acceleration inference device provided in this application;

[0085] Figure 5 This is a schematic diagram of the structure of an embodiment of another deep neural network accelerated inference device provided in this application. DETAILED DESCRIPTION

[0086] It should be noted that the deep neural network accelerated reasoning method provided in this application can be applied to a terminal, a system, or a server. For example, the terminal can be a smart phone or a computer, a tablet computer, a smart TV, a smart watch, a portable computer terminal, or a fixed terminal such as a desktop computer. For the convenience of explanation, this application uses the terminal as the execution subject for example.

[0087] See also Figure 1 , the present application first provides an embodiment of a method for accelerating reasoning of a deep neural network, the embodiment comprising:

[0088] S101, determining a standard data set matching the target task, preprocessing the standard data set, and dividing the standard data set into a training set, a validation set, and a test set;

[0089] In this embodiment, a standard dataset that can represent the characteristics of the target task is first selected. Data that can cover the characteristics of the target task can be selected from public databases (such as ImageNet, CIFAR-10, COCO) or custom collection datasets. Ensure that the data covers diverse characteristics related to the task, and then preprocess the data, where the data preprocessing may include:

[0090] Data cleaning, specifically:

[0091] Check the dataset for duplicate, corrupted, incomplete, or anomalous samples. Remove duplicate samples by hashing or direct comparison. Use file detection tools such as imghdr or Pandas to check for unreadable files and remove them. Check that the labels and distribution of the samples are consistent with expectations and remove samples with incorrect labels.

[0092] Data Augmentation:

[0093] The purpose of data augmentation is to expand the data set and improve the robustness of the model to the data.

[0094] For image-based tasks, these can include:

[0095] Geometric transformations: rotation, scaling, translation, cropping, flipping.

[0096] Color dithering: adjust brightness, contrast, saturation.

[0097] Noise addition: Add Gaussian noise or salt and pepper noise.

[0098] For time series tasks, these can include:

[0099] Add noise, time warp, interpolate data.

[0100] For text-based tasks, these can include:

[0101] Synonym replacement, random deletion, random insertion.

[0102] This can be implemented using Python libraries such as Albumentations, Augmentor, TorchVision, etc.

[0103] You can also normalize the data, for example, using the StandardScaler or MinMaxScaler functions of NumPy, Pandas, or Scikit-learn.

[0104] In this step, the preprocessed data also needs to be divided into a training set, a validation set, and a test set. In one possible implementation, the data can be divided according to a specific ratio, for example, the training set accounts for 60%-80%, the validation set accounts for 10%-20%, and the test set accounts for 10%-20%.

[0105] Specifically, you can divide the data by random division, for example, using train_test_split (from Scikit-learn) or a custom function to ensure that the samples are evenly distributed. Here is a code example:

[0106] from sklearn.model_selection import train_test_split

[0107] train_data, temp_data = train_test_split(data, test_size=0.3, random_state=42)

[0108] val_data, test_data = train_test_split(temp_data, test_size=0.5, random_state=42).

[0109] For classification tasks, you can also use a hierarchical division method. Here is a code example:

[0110] train_data, temp_data = train_test_split(data, test_size=0.3, stratify=data.labels, random_state=42).

[0111] After the division, the corresponding data sets can be organized into specific files. For example, the data can be assigned to different folders (such as train / , val / , test / ). For an implementation method suitable for image or document data, a script can be used to automatically move files. The following is an example of a script:

[0112] import os

[0113] import shutil

[0114] def move_files(file_list, source_dir, target_dir):

[0115] os.makedirs(target_dir, exist_ok=True)

[0116] for file in file_list:

[0117] shutil.move(os.path.join(source_dir, file), os.path.join(target_dir, file)).

[0118] S102, constructing a deep neural network model, and using the training set to train the deep neural network model;

[0119] In this embodiment, the structure of the deep neural network model is designed according to the task requirements (classification, regression, target detection, etc.), and the appropriate layer type and topology can be selected.

[0120] First select the network type:

[0121] Convolutional Neural Networks (CNN): used for image processing tasks.

[0122] Recurrent Neural Network (RNN) or Transformer: for sequential data (such as text, time series).

[0123] Multilayer Perceptron (MLP): used for regression or classification tasks with numerical features.

[0124] Key modules also need to be designed:

[0125] Convolutional layer: extract features.

[0126] Pooling layer: downsampling to reduce the amount of computation (such as maximum pooling, average pooling).

[0127] Fully connected layer: used for feature classification or regression.

[0128] Regularization layer: such as Dropout and BatchNormalization, to prevent overfitting.

[0129] In a specific implementation, a framework (such as PyTorch, TensorFlow) can be used to implement the network structure. The following is a code example:

[0130] import torch.nn as nn

[0131] class CustomCNN(nn.Module):

[0132] def __init__(self):

[0133] super(CustomCNN, self).__init__()

[0134] self.conv_layers = nn.Sequential(

[0135] nn.Conv2d(3, 32, kernel_size = 3, stride = 1, padding = 1),

[0136] nn.ReLU(),

[0137] nn.MaxPool2d(kernel_size = 2, stride = 2),

[0138] nn.Conv2d(32, 64, kernel_size = 3, stride = 1, padding = 1),

[0139] nn.ReLU(),

[0140] nn.MaxPool2d(kernel_size = 2, stride = 2) )

[0142] self.fc_layers = nn.Sequential(

[0143] nn.Linear(64 * 8 * 8, 128),

[0144] nn.ReLU(),

[0145] nn.Dropout(0.5),

[0146] nn.Linear(128, 10) # Assume 10 classes )

[0148] def forward(self, x):

[0149] x = self.conv_layers(x)

[0150] x = x.view(x.size(0), -1) # Flatten

[0151] x = self.fc_layers(x)

[0152] return x.

[0153] S103, using the verification set to verify the deep neural network model, and recording the weight distribution and L1 norm of each layer of neurons or connections;

[0154] Use the validation set data to evaluate the trained model and record its accuracy, loss value and other indicators. To avoid overfitting, the generalization ability of the model can be enhanced through early stopping mechanism and regularization. Count the distribution of weights in each layer and analyze the changes in weight values. Calculate the L1 norm of the weight (that is, the sum of the absolute values ​​of the weights) to measure the importance of the weights.

[0155] In this year's application, the role of the validation set is to evaluate the performance of the model and monitor the generalization ability of the model to avoid overfitting. First, you need to prepare the validation set data to ensure that the validation set is not involved in the training. Then calculate the performance indicators of the model. For different processing tasks, the calculation method can be different. For classification tasks, you can calculate accuracy, F1 score, precision, and recall.

[0156] For regression tasks, you can calculate the mean square error (MSE), mean absolute error (MAE), and R2.

[0157] Here is a code example:

[0158] model.eval()#Switch to evaluation mode, turn off Dropout and BatchNorm

[0159] correct = 0

[0160] total = 0

[0161] with torch.no_grad(): #Disable gradient calculation to improve reasoning efficiency

[0162] for inputs, labels in validation_loader:

[0163] outputs = model(inputs)

[0164] _, predicted = torch.max(outputs, 1)# Get the predicted category

[0165] total += labels.size(0)

[0166] correct += (predicted == labels).sum().item()

[0167] accuracy = correct / total

[0168] print(f"Validation Accuracy: {accuracy:4f}").

[0169] Use the validation set to calculate the validation loss, determine the optimization trend of the model, and count the weight distribution of each layer in the neural network for subsequent pruning and optimization analysis.

[0170] First, you need to extract the weight data, traverse each layer of the model, and extract its weight. You can use a histogram or density plot to observe the distribution of weights.

[0171] For L1 norm, the L1 norm is the sum of the absolute values ​​of the weights and is used to measure the importance of the weights.

[0172] For a single layer weight L1 norm, here is a code example:

[0173] for name, param in model.named_parameters():

[0174] if 'weight' in name:

[0175] l1_norm = param.abs().sum().item()

[0176] print(f"{name} L1 Norm: {l1_norm:4f}").

[0177] For all weight L1 norms, it summarizes the L1 norms of all layers and analyzes the importance of the overall weight. Here is a code example:

[0178] total_l1_norm=sum(param.abs().sum().item() for param inmodel.parameters() if 'weight' in name)

[0179] print(f"Total L1 Norm: {total_l1_norm: 4f}").

[0180] By analyzing the relationship between weight distribution and L1 norm, we can provide a basis for pruning. For example, weights with large L1 norms indicate high importance and are more likely to be retained.

[0181] In an optional embodiment, in order to enhance the robustness of the model, a weight_decay parameter may be introduced into the optimizer, and the implementation is as follows: optimizer = torch.optim.Adam (model.parameters(), lr = 0.001, weight_decay = 1e-4);

[0182] In the above example, Dropout is used to randomly discard some neurons to prevent overfitting.

[0183] In this embodiment, the L1 norm can be calculated by the following formula:

[0184] ;

[0185] Among them, n represents the number of weights and w represents the weight.

[0186] S104, determining the importance value of the neuron or connection based on the weight distribution and L1 norm recorded during the training process;

[0187] In this step, the L1 norm can be used as an importance indicator. The larger the absolute value of the weight, the more important it is to the network output. The weights of the same layer or the same type can be normalized in combination with distribution statistics for comparison. The importance values ​​of the relevant weights are summarized to determine the comprehensive importance of the corresponding neurons or connections. The neurons and connections of each layer are sorted to clarify the priority parts to be retained.

[0188] In this embodiment, the L1 norm represents the sum of the absolute values ​​of the weights, and the larger the value is, the more important the weight or neuron is to the network output.

[0189] For each layer's weight parameter, calculate the sum of their absolute values ​​one by one. The weight ranges of different layers may vary greatly, so they can be normalized for comparison.

[0190] Combined with the distribution statistics of each layer's weights, the significance of the weights is measured. First, the statistical information of the weights is extracted, and statistical indicators such as the mean and variance are analyzed. According to the weight distribution, the importance value below a certain percentile (such as 10%) is set as the pruning target.

[0191] Summarize and rank the importance of neurons or connections to determine which ones to retain first.

[0192] For each layer, summarize the importance values ​​of all weights associated with the neurons. Here is a code example:

[0193] for name, param in model.named_parameters():

[0194] if 'weight' in name:

[0195] neuron_importance = param.abs().sum(dim=0) # Sum the weights of all connections

[0196] print(f"Neuron importance for {name}: {neuron_importance}").

[0197] Sort the importance values ​​of neurons or connections in ascending order and identify the low-importance ones. A code example is as follows:

[0198] sorted_indices = torch.argsort(neuron_importance) # Sort by importance

[0199] prune_indices = sorted_indices[:int(0.2 * len(sorted_indices))]# Pruning ratio 20%

[0200] print(f"Indices for pruning: {prune_indices}").

[0201] In order to intuitively analyze the importance distribution of weights and neurons, visualization operations can be performed.

[0202] S105, pruning neurons or connections whose importance values ​​are lower than a preset importance threshold according to a preset ratio, and dynamically adjusting the structure of the deep neural network model after each pruning;

[0203] Set the importance threshold and filter out neurons or connections with lower importance based on the sorting results. The pruning ratio can be adjusted dynamically, such as a small ratio in the initial stage and gradually increased in the later stage. Delete the pruned connections or neurons and rebuild the network topology. Retrain or fine-tune the pruned model to recover some performance losses.

[0204] This step aims to prune the deep neural network model, remove neurons or connections with lower importance, and dynamically adjust the model structure. First, set the threshold or ratio of pruning based on the calculated importance value. For example, prune the importance values ​​below the 10% quantile. The pruning ratio can increase with the number of iterations. For example, it is set to 10% in the initial stage and then gradually increased to 50%. The importance value is used to filter the neuron or connection index below the threshold, and the first certain proportion of the less important parts are selected according to the sorting.

[0205] During pruning, rows or columns related to low-importance neurons or connections are directly removed from the weight matrix to dynamically adjust the network structure. For example, some channels of the convolutional layer or neurons of the fully connected layer are removed, and the model configuration file or network architecture is updated to ensure that the new network structure matches the pruned weights.

[0206] In an optional embodiment, the importance of the weight may be evaluated by calculating the sparsity index of the weight. A method for calculating the sparsity index is provided below:

[0207] ;

[0208] in, , S is the sparsity index of the weight, which is used to indicate the importance of the weight.

[0209] 106. Determine the quantization bit width and the upper and lower limits of quantization;

[0210] In this embodiment, quantization is the process of mapping floating point weights and activation values ​​to low-precision representations (such as integers), thereby reducing the storage and computational costs of the model.

[0211] S107. For the pruned deep neural network model, during the forward propagation process, the weights and activation values ​​are pseudo-quantized based on the quantization bit width and the quantization upper and lower limits. First, a commonly used quantization bit width is selected:

[0212] 8-bit: used for mobile devices and embedded hardware, a compromise between accuracy and performance.

[0213] 4-bit: Lower storage requirements, but may affect model accuracy.

[0214] 16-bit: Suitable for scenarios that require higher precision. The following is a code example using 8-bit bit width:

[0215] # Example bit width setting

[0216] bit_width=8#Use 8-bit quantization

[0217] quantization_levels=2**bit_width#quantization level

[0218] print(f"QuantizationLevels:{quantization_levels}").

[0219] Next, determine the upper and lower limits, collect the activation value distribution of the training or validation data, and count the minimum and maximum values. Symmetric or asymmetric quantization can be used:

[0220] Symmetric quantization: The upper and lower limits are the same, usually including zero (zero value symmetry).

[0221] Asymmetric quantization: The upper and lower limits are different, which is suitable for biased data distribution.

[0222] Calculate the minimum and maximum values ​​of the weights and adjust the range appropriately to improve the dynamic range coverage. The floating point value can be mapped to the quantized integer range using the quantization step size, where the quantization step size is as follows:

[0223] ;

[0224] in, represents the quantization step size, l represents the lower limit of quantization, and u represents the upper limit of quantization.

[0225] Here is a specific implementation method:

[0226] The quantization step size is calculated by the following formula:

[0227] ;

[0228] in, Indicates the quantization step size;

[0229] Clip weights and activations to the range [l,u];

[0230] According to the quantization step size The clipped weights and activation values ​​are discretized. The discretization formula is:

[0231] ;

[0232] in, Represents discrete values;

[0233] The discrete values ​​obtained by the discretization are mapped back to floating point numbers by the following formula:

[0234] ;

[0235] in, Represents a floating point number.

[0236] Assume parameter: b=8.

[0237] l = −1.0 (lower limit of quantization).

[0238] u=1.0 (quantization upper limit).

[0239] Calculation step length ≈0.007843;

[0240] Assume the following weights or activation values:

[0241] x=[0.5, −1.2, 0.9, −0.8, 1.2, −0.5]

[0242] The quantization range is [−1.0, 1.0], after clipping:

[0243] = [0.5, −1.0, 0.9, −0.8, 1.0, −0.5].

[0244] Then discretize:

[0245] After cutting, Discretize to get discrete values:

[0246] q=[191, 0, 242, 26, 255, 64];

[0247] Then dequantize and map the discrete value q back to a floating point number to get:

[0248] = [0.497213, −1.0, 0.897206, −0.796098, 1.0, −0.498848].

[0249] The following table shows the numerical comparison of the results through "Table 1. Example table of quantization and dequantization of weights and activation values":

[0250] Table 1. Example table of quantization and dequantization of weights and activation values

[0251]

[0252] Through this process, weights and activation values ​​are quantized into integer form (discrete value q), and can be restored to floating-point numbers close to the original values ​​through the inverse quantization formula. This quantization method can significantly reduce storage and computational costs, while minimizing the impact on model performance through appropriate bit width and clipping range selection.

[0253] S108. In the back propagation process, based on the result after pseudo quantization, the task loss and the quantization error loss are calculated, and the full-precision weight is updated;

[0254] The purpose of this step is to incorporate the quantization error into the optimization objective while maintaining the full-precision weights, and gradually adjust the model parameters to make them insensitive to the performance degradation after quantization.

[0255] See also Figure 2 In a specific implementation, this may include:

[0256] S1081. Calculate the loss of the classification task through the cross entropy function;

[0257] For classification tasks, the cross entropy loss function is used to measure the gap between the model prediction value and the true label.

[0258] S1082. Calculate the loss of the regression task by using the mean square error.

[0259] For regression tasks, the mean squared error (MSE) is used as the loss function, which measures the sum of squared errors between the predicted value and the true value.

[0260] S1083, using the full-precision weight and the pseudo-quantized weight to calculate the quantization error loss;

[0261] The quantization error loss is calculated using the full-precision weights and the fake quantization weights. The specific calculation method is as follows:

[0262] ;

[0263] in, represents the quantization error loss, represents the full precision weight, represents the pseudo-quantized weight, Calculate using the following formula:

[0264] ;

[0265] Wherein, b represents the quantization bit width, and (u, l) represents the upper and lower limits of the quantization.

[0266] S1084, fusing the loss of the classification task, the loss of the regression task, and the quantization error loss through a predefined joint loss function;

[0267] The classification task loss, regression task loss, and quantization error loss are fused through a predefined joint loss function.

[0268] In a specific implementation, the fusion method can be implemented through weighted calculation. The following provides an implementation example of a specific joint loss function:

[0269] L_total=αL_CE+βL_MSE+γL_quant;

[0270] Among them, α, β, γ are weight coefficients, indicating the importance of each loss in the joint loss, L_total represents the joint loss function, L_CE represents the cross entropy loss function, and L_MSE represents the mean square error, quantization error loss.

[0271] S1085. Based on the joint loss function, calculate the gradient through back propagation and update the full-precision weight.

[0272] Based on the joint loss function L_total, the gradient is calculated through back propagation to update the full-precision weights of the model. Through this implementation, the full-precision weight and quantized weight training can be effectively combined to improve the quantization adaptability of the model while maintaining task performance.

[0273] S109. Optimizing model parameters using a joint loss function, wherein the joint loss function includes task loss and quantization error loss.

[0274] In this embodiment, the reverse gradient is calculated by joint loss, and the full-precision weight is updated to make it more adaptable to low-bitwidth quantization in the next quantization. After each training cycle, the quantization bit width can be dynamically adjusted (such as gradually reducing from a high bit width to a target bit width) or the upper and lower limits of quantization can be adjusted to further enhance quantization adaptability.

[0275] See also Figure 3 , the present application provides another method embodiment, which includes:

[0276] S301, determining a standard data set matching the target task, preprocessing the standard data set, and dividing the standard data set into a training set, a validation set, and a test set;

[0277] S302, constructing a deep neural network model, and using the training set to train the deep neural network model;

[0278] S303, using the verification set to verify the deep neural network model, and recording the weight distribution and L1 norm of each layer of neurons or connections;

[0279] S304, determining the importance value of the neuron or connection based on the weight distribution and L1 norm recorded during the training process;

[0280] S305, pruning neurons or connections whose importance values ​​are lower than a preset importance threshold according to a preset ratio, and dynamically adjusting the structure of the deep neural network model after each pruning;

[0281] S306, performing sparse regularization during the pruning process;

[0282] In this embodiment, sparse regularization is to promote the realization of pruning effect by constraining the model weights to be sparse. During the training process, a sparse regularization term is added to make the unimportant weights approach zero, which is convenient for pruning operations. The Group Lasso method can be used to impose sparse constraints on the weight group. The specific formula is as follows:

[0283] ;

[0284] in, Represents a weight group, such as a convolution channel.

[0285] The sparse regularization term is combined with the task loss to form a new objective optimization function:

[0286] L_total=L_task+λ⋅Regularization Loss;

[0287] Among them, λ represents the regularization coefficient, which is used to balance the weight of task loss and sparse regularization term.

[0288] In this embodiment, a smaller regularization coefficient λ is set in the initial training to avoid excessive interference with the model weights. As the training progresses, λ is gradually increased to increase the constraint on sparsity.

[0289] Through the above steps and implementation, the sparsification effect can be significantly improved, providing a better basis for pruning, while reducing redundant computing and storage costs.

[0290] S307, determining the quantization bit width and the upper and lower limits of quantization;

[0291] S308, for the pruned deep neural network model, in the forward propagation process, based on the quantization bit width and the quantization upper and lower limits, perform pseudo-quantization processing on the weights and activation values;

[0292] S309. In the back propagation process, based on the pseudo-quantization result, the task loss and the quantization error loss are calculated, and the full-precision weight is updated;

[0293] S310, optimizing model parameters using a joint loss function, wherein the joint loss function includes task loss and quantization error loss.

[0294] With the advent of the big data era, data processing and analysis have become increasingly important. In the field of computer vision, data labeling, as an important part of data processing, is crucial for the training of deep learning algorithms such as object detection. However, most existing data labeling methods rely on manual operations, which are not only inefficient but also costly. Especially for large-scale data sets, the time-consuming and labor-intensive manual labeling makes it almost an impossible task.

[0295] In recent years, although some automatic labeling methods based on small models have been proposed, these methods are only for certain limited ranges of data. When processing complex data and large-scale visual image data sets, the labeling accuracy and efficiency are often unsatisfactory. In addition, these methods are usually inflexible and difficult to adapt to the data labeling needs of different fields and types. Therefore, the present application also provides an embodiment of automatic batch labeling of data, that is, a specific implementation method of preprocessing data in step S101, which aims to solve the problems of low labeling efficiency, high cost and insufficient labeling accuracy in the prior art, so as to accelerate the thrust efficiency of the model. The implementation method includes:

[0296] Step S1: data preprocessing;

[0297] In an embodiment of the present invention, data preprocessing includes data pre-labeling and interactive prompt word extraction. Data preprocessing refers to the selection of two to three thousand representative images for manual pre-labeling. The annotation content includes but is not limited to the category, location (such as bounding box), attributes, etc. of the object. The pre-labeled data should cover the main changes and scenarios of the target category to improve the generalization ability of the model. Interactive prompt word extraction refers to extracting concise and clear prompt words that can accurately describe the target object or scene according to the specific needs of the annotation task. For example, in the scene of helmet detection, the prompt words can be "Category 1: Head wearing a helmet; Category 2: Head not wearing a helmet; Category 3: Human body". These prompt words will be used as input instructions when interacting with the large model to help the model understand the requirements of the annotation task.

[0298] Step S2 large model training:

[0299] The large model in this embodiment has been fully trained on large-scale text and image data, can fully understand the object information in the real world, and can accurately detect most common categories of objects in the real world. The data pre-labeled in step S1 is input into the model for training so that the large model can better understand the scene information that needs to be labeled data. The large model used in the embodiment of the present invention includes three main parts: 1) Image encoder: The image encoder divides the input image into multiple small blocks and encodes these visual small blocks into high-dimensional features. 2) Prompt encoder: Encode various forms of prompts (points, boxes, text, etc.) provided by the user into vectors. Points and boxes use position encoding plus embedding of each prompt type, and free-form text is encoded using a text encoder. 3) Detection decoder: Combine image features and prompt information to generate detection box information of the target object. The decoder is designed to be lightweight so that it can respond quickly to each prompt.

[0300] Step S3: large model interaction;

[0301] The refined interactive prompt words are input into the trained large model as instructions for the model to understand the requirements of the annotation task. The image data to be annotated is input into the model in batches, and the model will automatically annotate according to the prompt words and image content. During the annotation process, the output of the model is monitored in real time to ensure the accuracy and stability of the annotation results.

[0302] Step S4: The large model outputs the annotation results;

[0303] The large model will organize the output annotation results, including associating the annotation information (such as category, location, attributes, etc.) with the original image to form a complete annotation data set. Users can choose to convert the annotation results into the required format (such as JSON, XML, etc.) according to actual needs for subsequent processing and analysis.

[0304] Step S5: quality inspection of labeling results;

[0305] In the embodiment of the present invention, data quality inspection includes automatic quality inspection and manual review. Automatic quality inspection uses preset quality inspection rules (such as the overlap of annotation boxes, image clarity, occlusion, size of annotation boxes, etc.) to automatically check the annotation results and filter out possible errors or non-compliance. Manual review manually reviews suspected errors found in automatic quality inspection to ensure the accuracy of the annotation results. At the same time, some data can also be randomly selected for manual inspection to evaluate the overall annotation quality.

[0306] Step S6: Acceptance of marking results;

[0307] In the embodiment of the present invention, detailed acceptance criteria and indicators are formulated according to project requirements and industry standards, including accuracy, completeness, consistency, etc. of the annotation. The project leader or relevant experts shall inspect and accept the annotation results to ensure that all annotations meet the acceptance criteria. The embodiment of the present invention will also automatically generate an acceptance report to summarize the quality of the annotation results, existing problems, and improvement measures, providing a reference for subsequent work.

[0308] The above describes in detail the embodiments of the method provided in the present application. The following describes the embodiments of the device provided in the present application:

[0309] See also Figure 4 , the present application provides an embodiment of a deep neural network acceleration reasoning device, the embodiment comprising:

[0310] A preprocessing unit 401 is used to determine a standard data set that matches the target task, preprocess the standard data set, and divide it into a training set, a validation set, and a test set;

[0311] A model building unit 402 is used to build a deep neural network model and train the deep neural network model using the training set;

[0312] A verification unit 403 is used to verify the deep neural network model using the verification set, and record the weight distribution and L1 norm of each layer of neurons or connections;

[0313] An importance evaluation unit 404, used to determine the importance value of the neuron or connection based on the weight distribution and L1 norm recorded during the training process;

[0314] A pruning unit 405 is used to prune neurons or connections whose importance values ​​are lower than a preset importance threshold according to a preset ratio, and dynamically adjust the structure of the deep neural network model after each pruning;

[0315] A quantization determination unit 406, used to determine the quantization bit width and the quantization upper and lower limits;

[0316] A pseudo-quantization processing unit 407 is used to perform pseudo-quantization processing on weights and activation values ​​of the pruned deep neural network model based on the quantization bit width and the quantization upper and lower limits during the forward propagation process;

[0317] A loss calculation unit 408, used to calculate the task loss and the quantization error loss based on the pseudo-quantized result during the back-propagation process, and update the full-precision weight;

[0318] The optimization unit 409 is used to optimize the model parameters using a joint loss function, wherein the joint loss function includes a task loss and a quantization error loss.

[0319] Optionally, the importance evaluation unit 404 is specifically used for:

[0320] Record the statistical information of each layer weight, including mean, variance, maximum and minimum values;

[0321] The L1 norm of the weight is calculated by the following formula;

[0322] ;

[0323] Among them, n represents the number of weights, and w represents the weight;

[0324] The sparsity index S of the weight is calculated by the following formula:

[0325] ;

[0326] ;

[0327] Among them, S is the sparsity index of the weight, which is used to indicate the importance of the weight.

[0328] Optionally, the quantization determination unit 406 is specifically configured to:

[0329] For a given bit width b, use the training set or pruned weights and activation values ​​to calculate the maximum value x_max and the minimum value x_min;

[0330] The quantization upper and lower limits are the dynamic range, which is calculated by the following formula:

[0331] l=x_min,u=x_max;

[0332] Where l represents the lower limit of quantization, and u represents the upper limit of quantization.

[0333] Optionally, the pseudo quantization processing unit 407 is specifically used for:

[0334] The quantization step size is calculated by the following formula:

[0335] ;

[0336] in, Indicates the quantization step size;

[0337] Clip weights and activations to the range [l,u];

[0338] According to the quantization step size The clipped weights and activation values ​​are discretized. The discretization formula is:

[0339] ;

[0340] in, Represents discrete values;

[0341] The discrete values ​​obtained by the discretization are mapped back to floating point numbers by the following formula:

[0342] ;

[0343] in, Represents a floating point number.

[0344] Optionally, the loss calculation unit 408 is specifically used for:

[0345] The loss of the classification task is calculated through the cross entropy function;

[0346] The loss of the regression task is calculated by the mean squared error;

[0347] Calculate the quantization error loss using the full-precision weights and the fake-quantized weights;

[0348] The classification task loss, regression task loss, and quantization error loss are fused through a predefined joint loss function;

[0349] Based on the joint loss function, the gradient is calculated by back-propagation and the full-precision weights are updated.

[0350] Optionally, the quantization error loss is calculated in the following manner:

[0351] ;

[0352] in, represents the quantization error loss, represents the full precision weight, represents the pseudo-quantized weight, Calculate using the following formula:

[0353] ;

[0354] Wherein, b represents the quantization bit width, and (u, l) represents the upper and lower limits of the quantization.

[0355] Optionally, the loss calculation unit 408 is specifically used for:

[0356] According to a preconfigured weight coefficient, the task loss and the quantization error loss are fused into a joint loss through weighted calculation;

[0357] In back-propagation, the joint loss is used to calculate the gradient and update the parameters of the deep neural network model.

[0358] See also Figure 5 , the present application also provides a deep neural network accelerated reasoning device, comprising:

[0359] Processor 501, memory 502, input and output unit 503, bus 504;

[0360] The processor 501 is connected to the memory 502, the input and output unit 503 and the bus 504;

[0361] The memory 502 stores a program, and the processor 501 calls the program to execute any of the above methods.

[0362] The present application also relates to a computer-readable storage medium on which a program is stored, wherein when the program is run on a computer, the computer is caused to execute any of the above methods.

[0363] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0364] In the several embodiments provided in the present application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic. For example, the division of the units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0365] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0366] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of software functional units.

[0367] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, read-only memory), random access memory (RAM, random access memory), disk or optical disk and other media that can store program code.

Claims

1. A method for accelerating inference of a deep neural network, characterized in that: The method comprises: Determine a standard data set that matches the target task, preprocess the standard data set, and divide it into a training set, a validation set, and a test set; for image processing tasks, the preprocessing includes: rotation, scaling, translation, cropping, or flipping; for time series tasks, the preprocessing includes: adding noise, time warping, or data interpolation; for text tasks, the preprocessing includes: synonym replacement, random deletion, or random insertion; Construct a deep neural network model, and use the training set to train the deep neural network model. For image processing tasks, the deep neural network model is a convolutional neural network; for time series tasks, the deep neural network model is a recurrent neural network; for regression or classification tasks of numerical features, the deep neural network model is a multilayer perceptron; Using the validation set to validate the deep neural network model, and recording the weight distribution and L1 norm of each layer of neurons or connections; Determining the importance value of the neuron or connection based on the weight distribution and L1 norm recorded during the training process; Prune neurons or connections whose importance values ​​are lower than a preset importance threshold according to a preset ratio, and dynamically adjust the structure of the deep neural network model after each pruning; Determine the quantization bit width and the upper and lower limits of quantization; For the pruned deep neural network model, during the forward propagation process, the weights and activation values ​​are pseudo-quantized based on the quantization bit width and the upper and lower limits of quantization; During the back propagation process, based on the pseudo-quantization results, the task loss and quantization error loss are calculated, and the full-precision weights are updated; The model parameters are optimized using a joint loss function, where the joint loss function includes task loss and quantization error loss.

2. The method for accelerating inference of a deep neural network according to claim 1, characterized in that: Based on the weight distribution and L1 norm recorded during the training process, determining the importance value of the neuron or connection includes: Record the statistical information of each layer weight, including mean, variance, maximum and minimum values; The L1 norm of the weight is calculated by the following formula; ; Among them, n represents the number of weights, and w represents the weight; The sparsity index S of the weight is calculated by the following formula: ; ; Among them, S is the sparsity index of the weight, which is used to indicate the importance of the weight.

3. The method for accelerating inference of a deep neural network according to claim 1, characterized in that: Pruning neurons or connections whose importance values ​​are lower than a preset importance threshold according to a preset ratio, and dynamically adjusting the structure of the deep neural network model after each pruning includes: Get the importance value of each layer of neurons or connections and sort them from low to high; The number of pruning n is determined according to the preset ratio p through the following formula: n = (p × N); Where p represents the preset ratio, and N represents the total number of neurons or connections in the current layer; Remove neurons or connections whose importance is below a preset importance threshold; Update the structure of the deep neural network model and delete the corresponding weights and biases.

4. The method for accelerating inference of a deep neural network according to claim 1, characterized in that: Determining the quantization bit width and the upper and lower limits of quantization includes: For a given bit width b, use the training set or pruned weights and activation values ​​to calculate the maximum value x_max and the minimum value x_min; The quantization upper and lower limits are the dynamic range, which is calculated by the following formula: l=x_min,u=x_max; Where l represents the lower limit of quantization, and u represents the upper limit of quantization.

5. The method for accelerating inference of a deep neural network according to claim 4, characterized in that: For the pruned deep neural network model, during the forward propagation process, the weights and activation values ​​are pseudo-quantized based on the quantization bit width and the upper and lower limits of quantization, including: The quantization step size is calculated by the following formula: ; in, Indicates the quantization step size; Clip weights and activations to the range [l,u]; According to the quantization step size The pruned weights and activation values ​​are discretized; the discretization formula is: ; in, Represents discrete values; The discrete values ​​obtained by the discretization are mapped back to floating point numbers by the following formula: ; in, Represents a floating point number.

6. The method for accelerating inference of a deep neural network according to claim 1, characterized in that: In the back propagation process, based on the pseudo-quantization result, the task loss and quantization error loss are calculated, and the full-precision weight is updated, including: The loss of the classification task is calculated through the cross entropy function; The loss of the regression task is calculated by the mean squared error; Calculate the quantization error loss using the full-precision weights and the fake-quantized weights; The loss of the classification task, the loss of the regression task, and the quantization error loss are fused through a predefined joint loss function; Based on the joint loss function, the gradient is calculated by back-propagation and the full-precision weights are updated.

7. The method for accelerating inference of a deep neural network according to claim 6, characterized in that: The quantization error loss is calculated as follows: ; in, represents the quantization error loss, represents the full precision weight, represents the pseudo-quantized weight, Calculate using the following formula: ; Wherein, b represents the quantization bit width, and (u, l) represents the upper and lower limits of the quantization.

8. The method for accelerating inference of a deep neural network according to claim 7, characterized in that: The use of a joint loss function to optimize model parameters, wherein the joint loss function includes task loss and quantization error loss, comprises: According to a preconfigured weight coefficient, the task loss and the quantization error loss are fused into a joint loss through weighted calculation; In back-propagation, the joint loss is used to calculate the gradient and update the parameters of the deep neural network model.

9. A deep neural network accelerated reasoning device, characterized in that: The device comprises: A preprocessing unit is used to determine a standard data set that matches the target task, preprocess the standard data set, and divide it into a training set, a validation set, and a test set. For image processing tasks, the preprocessing includes: rotation, scaling, translation, cropping, or flipping; for time series tasks, the preprocessing includes: adding noise, time warping, or data interpolation; for text tasks, the preprocessing includes: synonym replacement, random deletion, or random insertion; A model building unit, used to build a deep neural network model, and use the training set to train the deep neural network model. For image processing tasks, the deep neural network model is a convolutional neural network; for time series tasks, the deep neural network model is a recurrent neural network; for regression or classification tasks of numerical features, the deep neural network model is a multilayer perceptron; A verification unit, used to verify the deep neural network model using the verification set, and record the weight distribution and L1 norm of each layer of neurons or connections; An importance evaluation unit, used to determine the importance value of the neuron or connection based on the weight distribution and L1 norm recorded during the training process; A pruning unit, used to prune neurons or connections whose importance values ​​are lower than a preset importance threshold according to a preset ratio, and dynamically adjust the structure of the deep neural network model after each pruning; A quantization determination unit, used to determine the quantization bit width and the quantization upper and lower limits; A pseudo-quantization processing unit is used to perform pseudo-quantization processing on weights and activation values ​​of the pruned deep neural network model based on the quantization bit width and the upper and lower limits of quantization during the forward propagation process; The loss calculation unit is used to calculate the task loss and quantization error loss based on the pseudo-quantization results during the back-propagation process, and update the full-precision weights; An optimization unit is used to optimize model parameters using a joint loss function, wherein the joint loss function includes a task loss and a quantization error loss.

10. A deep neural network accelerated reasoning device, characterized in that: The device comprises: Processor, memory, input-output unit, and bus; The processor is connected to the memory, the input and output unit, and the bus; The memory stores a program, and the processor calls the program to execute the method according to any one of claims 1 to 8.

11. A computer-readable storage medium having a program stored thereon, wherein the program, when executed on a computer, performs the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Channel pruning and fast connection layer pruning method and system

    CN113222142A

  • Deep neural network compression method based on joint dynamic pruning

    CN118468968A