Image classification method and system based on quantitative perception and neural architecture search

By adopting quantitative perception and neural architecture search methods in image object detection tasks, the search space for lightweight candidate operations is designed and quantized perception training is carried out, which solves the problems of complex structures and high parameter quantities in image object detection tasks, and realizes efficient and highly generalized model training.

CN120047742APending Publication Date: 2025-05-27SHANDONG NORMAL UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510137353.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

Existing deep neural networks have problems such as complex structure, huge parameters and high design difficulty in image object detection tasks, which leads to slow training speed and difficulty in improving.

Method used

The image classification method based on quantitative perception and neural architecture search is adopted to design a search space for lightweight candidate operations to reduce the complexity of the model, and use quantitative perception training during the training process to solve the problem of non-differentiation in the quantization process, while improving the generalization ability of the model.

Benefits of technology

It effectively reduces the number of parameters and computational complexity (FLOPs) of the model, improves the generalization ability and training speed of the model, and reduces the negative impact of quantization on model accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120047742A_ABST
    Figure CN120047742A_ABST
Patent Text Reader

Abstract

The invention discloses an image classification method and system based on quantitative perception and neural architecture search, and the method comprises the steps: inputting a training set into a neural network, training the neural network, and obtaining a trained neural network; the construction of the neural network architecture comprises the following steps: constructing a dynamic channel adjustment convolution candidate operation model, and taking the dynamic channel adjustment convolution candidate operation model as a generated candidate operation; combining the generated candidate operation with the original candidate operation to obtain a candidate operation set; based on the candidate operation set, setting an initial Normal unit and an initial Reduction unit; based on the initial Normal unit and the initial Reduction unit, constructing a first neural network architecture; the first neural network architecture is trained based on the training set, and in the training process, if the loss function value is lower than a set threshold value, it is indicated that the current first neural network architecture is the final neural network architecture.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of artificial intelligence and machine learning, and particularly to an image classification method and system based on quantization awareness and neural architecture search. Background Art

[0002] In recent years, with the continuous improvement of the computing power of hardware devices, convolutional neural network models based on deep learning have emerged in the field of image target detection tasks.

[0003] Although convolutional neural network models based on deep learning have achieved good results in image target detection tasks, researchers still encounter new difficulties. The main reasons for these difficulties are the complex neural network structure, huge number of parameters, and high design difficulty.

[0004] Existing deep neural networks need to manually design the network architecture based on expert experience and build a deep neural network that meets the image detection task by continuously trial and error to adjust the parameters. This process is time-consuming and difficult to improve the training speed. Summary of the Invention

[0005] To solve the deficiencies of the prior art, the present invention provides an image classification method and system based on quantization awareness and neural architecture search; by designing the search space of lightweight candidate operations, the complexity of the model is reduced to a certain extent. During the training process, quantization-aware training is used, which not only solves the non-differentiable problem in the quantization process but also improves the generalization ability of the model. In addition, by adopting multi-objective optimization technology, the number of parameters and FLOPs of the model can be reduced while maintaining the model performance.

[0006] On the one hand, an image classification method based on quantization awareness and neural architecture search is provided, including:

[0007] Obtain a training set, where the training set is an image with a known image classification result; based on the training set, construct a neural network architecture; input the training set into the neural network and train the neural network to obtain a trained neural network; input the image to be classified into the trained neural network to obtain an image classification result; wherein, constructing a neural network architecture based on the training set includes:

[0008] (1) Construct a dynamic channel adjustment convolution candidate operation model, and use the dynamic channel adjustment convolution candidate operation model as the generated candidate operation; merge the generated candidate operation with the original candidate operation to obtain a candidate operation set;

[0009] (2) Based on the candidate operation set, set an initial Normal unit and an initial Reduction unit;

[0010] (3) Based on the initial Normal unit and the initial Reduction unit, construct the first neural network architecture;

[0011] (4) Train the first neural network architecture based on the training set. During the training process, if the loss function value is higher than the set threshold, return to (2); if the loss function value is lower than the set threshold, it means that the current first neural network architecture is the final neural network architecture, and end.

[0012] On the other hand, an image classification system based on quantization awareness and neural architecture search is provided, including:

[0013] An acquisition module, which is configured to: acquire a training set, where the training set is an image with a known image classification result; based on the training set, construct a neural network architecture; input the training set into the neural network to train the neural network, and obtain a trained neural network; input the image to be classified into the trained neural network to obtain an image classification result; among them, based on the training set, constructing a neural network architecture includes:

[0014] (1) Construct a dynamic channel adjustment convolution candidate operation model, and use the dynamic channel adjustment convolution candidate operation model as the generated candidate operation; merge the generated candidate operation with the original candidate operation to obtain a candidate operation set;

[0015] (2) Based on the candidate operation set, set the initial Normal unit and the initial Reduction unit;

[0016] (3) Based on the initial Normal unit and the initial Reduction unit, construct the first neural network architecture;

[0017] (4) Train the first neural network architecture based on the training set. During the training process, if the loss function value is higher than the set threshold, return to (2); if the loss function value is lower than the set threshold, it means that the current first neural network architecture is the final neural network architecture, and end.

[0018] On yet another aspect, an electronic device is further provided, including:

[0019] A memory for non-temporarily storing computer-readable instructions; and

[0020] A processor for running the computer-readable instructions,

[0021] wherein, when the computer-readable instructions are run by the processor, the method described in the first aspect above is executed.

[0022] In another aspect, a storage medium is also provided, which non-temporarily stores computer-readable instructions. When the non-temporary computer-readable instructions are executed by a computer, the method described in the first aspect is executed.

[0023] In another aspect, a computer program product is also provided, including a computer program, which is used to implement the method described in the above first aspect when running on one or more processors.

[0024] The above technical solutions have the following advantages or beneficial effects:

[0025] By designing the search space of lightweight candidate operations, the complexity of the model is reduced to a certain extent. During the training process, quantization-aware training is used, which not only solves the non-differentiable problem of the quantization process but also improves the generalization ability of the model. In addition, by adopting the multi-objective optimization technology, the number of model parameters and FLOPs can be reduced while maintaining the model performance.

[0026] The DCAC structure in the present invention reduces the amount of computation by adjusting the number of channels and the depth of the feature map, and at the same time uses the SE block to enhance the feature representation ability. By dynamically selecting a part of important channels for convolution calculation, for the unselected channels, they are multiplied by an approximate value of 0 to suppress, while keeping the spatial dimension of the feature map unchanged. This structure effectively reduces the number of model parameters while maintaining the information flow and enhancing the learning ability of the model.

[0027] The QPS proposed by the present invention shows significant beneficial effects in the field of neural architecture search (NAS). By increasing the number of layers and complexity of the network architecture in stages, this strategy allows for effective search in resource-constrained environments, especially suitable for scenarios with a large problem space or limited computing resources. The improved PS introduces a performance feedback mechanism, which dynamically adjusts the search strategy according to performance metrics such as accuracy and FLOPs at each stage, ensuring the efficiency of the search process and avoiding premature convergence to local optima.

[0028] In addition, the present invention effectively reduces the dependencies between layers and neurons in the network, reduces the risk of overfitting, and enhances the generalization ability of the model by restricting the number of skip connections and adopting the regularization strategy of operation-level Dropout. In terms of multi-objective optimization, the present invention not only considers model accuracy but also incorporates the number of parameters and FLOPs into the optimization strategy, and achieves a network architecture with high performance and low computational complexity by balancing the weights between these factors.

[0029] The QAT in the present invention considers the impact of quantization during the training process, which can significantly reduce the negative impact of quantization on model accuracy, so that the model adapts to quantization errors during the training stage, which helps to reduce accuracy loss when actually applying quantization. Brief Description of the Drawings

[0030] The accompanying drawings forming a part of this invention are used to provide a further understanding of the invention. The schematic embodiments of the invention and their descriptions are used to explain the invention and do not unduly limit the invention.

[0031] Figure 1 It is the DCAC structure diagram of the present invention;

[0032] Figures 2(a) and 2(b) are the evaluable progressive search strategies of the present invention;

[0033] Figures 3(a) and 3(b) are the unit structure diagrams searched by the present invention;

[0034] Figure 4 It is the directed acyclic graph of the present invention. Detailed Description of the Invention

[0035] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present invention belongs.

[0036] Neural Architecture Search (NAS), as a method for automatically designing neural network architectures, improves the performance and efficiency of models by searching for optimal or near-optimal neural network architectures. NAS mainly consists of three parts: the search space, the search strategy, and the evaluation strategy. The search strategies of NAS can be classified into the following categories: reinforcement learning (RL), gradient optimization, evolutionary algorithms, Bayesian optimization, and some hybrid strategies, etc.

[0037] Although NAS has made significant progress in automatically designing neural networks, its search process often requires a large amount of computing resources, which becomes a major challenge in resource-constrained environments. To address this issue, Lightweight Neural Architecture Search (L-NAS) emerged. L-NAS effectively reduces the demand for computing resources by optimizing the search process while maintaining the accuracy of the model. It focuses on reducing the computational complexity and the number of parameters of the model, thereby reducing the demand for storage and computing resources and making NAS more efficient and practical.

[0038] In terms of pursuing model search efficiency, the DARTS algorithm conducts architecture search in a differentiable manner, significantly shortening the search time and effectively achieving the search acceleration goal of L-NAS. However, the DARTS algorithm has some drawbacks, such as a large number of operation parameters, which leads to an increase in model complexity, and a large number of multiplication operations in the convolutional layer consume a large amount of computing resources. To overcome these problems, researchers have proposed a quantization-aware lightweight differentiable architecture search method. This method reduces the model complexity by designing a search space for lightweight candidate operations. At the same time, through quantization-aware training (QAT), the non-differentiable problem of the quantization process is solved, and the generalization ability of the model is improved. In addition, multi-objective optimization techniques are adopted to reduce the number of model parameters and FLOPs while maintaining the model performance, further enhancing the model efficiency.

[0039] QAT is a technique that introduces quantization effects during model training, aiming to make the model consider the errors brought by quantization during training, so as to maintain higher performance when performing inference with low precision during actual deployment. Different from traditional post-training quantization (PTQ), QAT simulates quantization operations during training. By inserting pseudo-quantization nodes (FakeQuant) in the model to simulate the errors introduced by quantization, and minimizing these errors during training, a model adapted to the quantization environment is finally obtained.

[0040] Example 1

[0041] This example provides an image classification method based on quantization awareness and neural architecture search;

[0042] An image classification method based on quantization awareness and neural architecture search includes:

[0043] Obtain a training set, where the training set is images with known image classification results; based on the training set, construct a neural network architecture; input the training set into the neural network and train the neural network to obtain a trained neural network; input the image to be classified into the trained neural network to obtain an image classification result; among them, based on the training set, constructing a neural network architecture includes:

[0044] (1) Construct a dynamic channel adjustment convolution candidate operation model, and use the dynamic channel adjustment convolution candidate operation model as the generated candidate operation; merge the generated candidate operation with the original candidate operations to obtain a candidate operation set;

[0045] (2) Based on the candidate operation set, set an initial Normal unit and an initial Reduction unit;

[0046] (3) Based on the initial Normal unit and the initial Reduction unit, construct a first neural network architecture;

[0047] (4) Train the first neural network architecture based on the training set. During the training process, if the value of the loss function is higher than the set threshold, return to (2); if the value of the loss function is lower than the set threshold, it means that the current first neural network architecture is the final neural network architecture, and end.

[0048] Furthermore, the step of constructing a neural network architecture based on the training set further includes:

[0049] (5) Determine whether the performance of the final neural network architecture exceeds the set threshold. If so, end; if not, proceed to (6);

[0050] (6) Based on the candidate operation set, set the initial Normal unit and the initial Reduction unit;

[0051] (7) Based on the initial Normal unit and the initial Reduction unit, construct the second neural network architecture;

[0052] (8) Train the second neural network architecture based on the training set. During the training process, if the objective function is higher than the set threshold, return to (6); if the value of the objective function is lower than the set threshold, it means that the current second neural network architecture is the final neural network architecture, and end.

[0053] Furthermore, (1) the original candidate operations include: 3×3 depthwise separable convolution SepConv, 3×3 dilated convolution DilConv, 5×5 depthwise separable convolution SepConv, 5×5 dilated convolution DilConv, 3×3 max pooling MaxPooling, 3×3 average pooling Avg Pooling, zeroing operation Zero, and identity mapping Identity.

[0054] Furthermore, (1) the dynamic channel adjustment convolution candidate operation model includes: an input layer, a channel importance evaluation module, an adaptive channel selection module, a depth convolution module, and an output layer connected in sequence; the input layer is used to input the images of the training set, and the output layer is used to output the feature maps of the images;

[0055] Among them, the channel importance evaluation module includes: an average pooling layer, a first fully connected layer, a first activation function layer, a second fully connected layer, and a second activation function layer connected in sequence; the channel importance evaluation module is used to output a scoring vector;

[0056] Among them, the adaptive channel selection module selects the top k channels with the highest scoring vectors;

[0057] Among them, the depth convolution module includes: a depth convolution layer, a batch normalization layer, and a third activation function layer connected in sequence; the third activation function layer of the depth convolution module outputs the feature map of the image.

[0058] Among them, the input layer is also connected to the input end of the third activation function layer through a residual connection.

[0059] It should be understood that the depth convolution layer independently applies a convolution kernel to each input channel, which means that if the input feature map has multiple channels, there will be the same number of convolution kernels, and each convolution kernel only acts on one channel. In this way, the features within each channel can be extracted without mixing the information between channels.

[0060] Relative to the depth convolution layer, the ordinary convolution layer uses a convolution kernel of a fixed size to perform convolution calculations at each position of the input feature map. Each convolution kernel needs to learn multiple weight values, and there is no shared weight between different convolution kernels.

[0061] Furthermore, the working process of the (1) dynamic channel adjustment convolution candidate operation model includes:

[0062] The input layer inputs the images of the training set, and the channel importance evaluation module outputs the evaluation value a;

[0063] a = σ(W 2 ·ReLU(W 1 ·GAP(X)));

[0064] Among them, the channel importance evaluation vector GAP is the global average pooling operation, which converts the input feature map into a vector of size C in , and are the weight matrices of the attention network, and σ is the sigmoid activation function;

[0065] According to the evaluation value a output by the channel importance evaluation module, select the channels with evaluation values higher than the set threshold for convolution calculation. For the unselected channels, multiply them by 0 to suppress, and keep the spatial dimension of the feature map unchanged; the calculation formula is as follows:

[0066] X selected = X[:, :, top-k(a, C selected )];

[0067] According to the scoring vector a, select the channels with the top-k(a, C selected ) with high scores for convolution calculation, X selectedThe selected feature map only contains the top k channels with the highest scores selected by the a scoring vector, where ":" indicates selecting all rows and columns;

[0068] Finally, the adaptive channel selection module performs channel selection, only retaining the top k channels with the highest scores;

[0069] The depth convolution module performs depth convolution operations on the top k channels with the highest scores.

[0070] It should be understood that the Dynamic Channel Adjustment Convolution Candidate Operation Model (DCAC) improves the inference speed while reducing the computational amount of the model. The Dynamic Channel Adjustment Convolution Candidate Operation Model consists of adaptive channel selection, depth convolution, and residual connection. This part reduces the computational amount by adjusting the number of channels and the depth of the feature map, while using the SE block to enhance the feature representation ability and using the SE block to evaluate the importance of each channel before convolution.

[0071] To maintain the information flow and enhance the learning ability of the model, a residual connection is introduced in the Dynamic Channel Adjustment Convolution Candidate Operation Model. If the dimensions of the input and output feature maps are the same, the residual connection can be directly performed; if they are different, 1×1 convolution is used for matching. The structure diagram of the Dynamic Channel Adjustment Convolution Candidate Operation Model is shown by Figure 1 as follows.

[0072] The full name of the SE block is: Squeeze-and-Excitation Block. Chinese explanation: The SE block is a module used in deep learning models, mainly used to enhance the representation ability of feature channels in convolutional neural networks. Its core idea is to dynamically adjust the importance weights of each feature channel through learning, thereby enhancing the model's attention to key features.

[0073] Squeeze (compression): Compress the spatial dimension of the feature map into a single feature vector through global average pooling, retaining the global information of each channel.

[0074] Excitation (activation): Process the compressed feature vector through two fully connected layers to generate the weights of each channel. After these weights are normalized by the Sigmoid activation function, they are re-weighted to the original feature map.

[0075] Furthermore, (2) based on the candidate operation set, set the initial Normal unit and the initial Reduction unit, including:

[0076] The construction process of the initial Normal unit includes: sorting the candidate operations in the candidate operation set from large to small based on the probability value, selecting the top M candidate operations, and connecting the selected candidate operations in sequence to obtain a directed acyclic graph, such as Figure 4 As shown in the figure, each node in the directed acyclic graph represents a feature graph, and each edge represents a candidate operation; the probability value is calculated based on the Softmax function; the input value of the Softmax function is a vector containing all candidate operations; the Softmax function converts each element of the input vector into a probability;

[0077]

[0078] in represents a vector of candidate operations, K is the number of candidate operations, yes The exponential function of .

[0079] The construction process of the initial Reduction unit includes: sorting the candidate operations in the candidate operation set from large to small based on the probability value, selecting the top M candidate operations, and connecting the selected candidate operations in sequence to obtain a directed acyclic graph, in which each node represents a feature graph and each edge represents a candidate operation. When constructing the directed acyclic graph of the initial Reduction unit, at least one candidate operation is a convolution operation with stride = 2 to reduce the spatial resolution of the feature graph.

[0080] The initial Normal unit is used to maintain the resolution of the feature map, that is, the size of the input and output feature maps is the same. The initial Normal unit is used in most layers of the network, and its purpose is to learn the combination and transformation of features without changing the size of the feature map.

[0081] The initial Reduction unit is used to reduce the resolution of the feature map. This is achieved by using a step size greater than 1 in the convolution operation, which halves the width and height of the feature map, thereby reducing the number of parameters and the amount of computation. The position of this unit in the network is fixed, and this method is used to reduce the size of the feature map at one-third and two-thirds of the network.

[0082] The design principle of unit-based neural architecture search is to insert a Reduction unit after every N Normal units to maintain network performance and effectively control the complexity and computational cost of the network.

[0083] Furthermore, the (3) constructs a first neural network architecture based on the initial Normal unit and the initial Reduction unit; wherein the first neural network architecture includes:

[0084] The first initial Normal unit, the first initial Reduction unit, the second initial Normal unit, the second initial Reduction unit, and the third initial Normal unit connected in sequence.

[0085] Further, (4) train the first neural network architecture based on the training set. During the training process, if the value of the loss function is higher than the set threshold, return to (2); if the value of the loss function is lower than the set threshold, it means that the current first neural network architecture is the final neural network architecture, and end; the loss function is implemented using the cross-entropy loss function.

[0086] Further, (5) determine whether the performance of the final neural network architecture exceeds the set threshold. If so, end; if not, enter (6); where the performance is a performance score calculated based on the accuracy rate and the number of floating-point operations.

[0087] The performance score QPS, the formula is as follows:

[0088] QPS = γ·Acc - δ(Lat + FLOPs)

[0089] Where γ is the weight of the accuracy rate, indicating the impact of the accuracy rate on the performance, δ is the weight of the complexity, used to represent the penalty degree of the latency and the computational complexity on the performance; Acc is the accuracy rate, and FLOPs is the number of floating-point operations. Lat represents the latency, and the latency is the time required for the model to make a prediction on the input data.

[0090] Further, (6) set the initial Normal unit and the initial Reduction unit based on the candidate operation set; the process is the same as that of (2), and will not be elaborated here.

[0091] Further, (7) construct the second neural network architecture based on the initial Normal unit and the initial Reduction unit; where the second neural network architecture includes:

[0092] Four initial Normal units, one initial Reduction unit, four initial Normal units, one initial Reduction unit, and four initial Normal units connected in sequence.

[0093] Further, (8) train the second neural network architecture based on the training set. During the training process, if the objective function is higher than the set threshold, return to (6); if the value of the objective function is lower than the set threshold, it means that the current second neural network architecture is the final neural network architecture, and end; where the specific expression of the objective function is:

[0094]

[0095] Among them, CE is the cross-entropy loss between the evaluation layer structures, used to limit the size of the parameters, used to limit the number of floating-point operations, λ 1 and λ 2 are the weights that balance accuracy, parameter size, and FLOPs.

[0096] It should be understood that the traditional progressive search is improved so that it can dynamically adjust the next operation according to the performance of the model, which not only shortens the search time but also avoids waste of resources.

[0097] Progressive search measurement is a strategy that gradually expands the search scope and complexity during the search process, especially suitable for scenarios with resource constraints or overly large problem spaces. The progressive search strategy mentioned in P-DARTS divides the search process into multiple stages. In each stage, the number of layers of the generated network architecture is gradually increased, while the number of choices in the candidate operations is reduced. In the initial stage, operations with lower complexity and smaller search scope are selected to quickly screen out potential high-quality candidate operations. As the number of layers of the generated network architecture increases, the candidate operations are searched at a finer level.

[0098] The present invention improves this progressive search strategy (PS) so that it can rely on a performance feedback mechanism to dynamically adjust the search strategy (as shown in FIGS. 2(a) and 2(b)). Specifically, after each stage ends, the system evaluates the current search results and performance metrics (such as accuracy Acc, number of floating-point operations FLOPs, etc.). Based on this feedback information, QPS decides whether to skip or merge subsequent search stages. FIGS. 3(a) and 3(b) are the unit structure diagrams searched by the present invention.

[0099] Using progressive search may lead to getting stuck in a local optimal solution because during the search process, the system usually only considers the optimal solution under the current complexity and ignores the potentially better solution that may exist at a higher complexity. Experimental observations show that candidate operations tend to choose skip connections rather than convolutions or pooling because skip connections can lead to fast gradient descent, especially on small datasets, which often leads to overfitting. Therefore, it is necessary to limit the number of skip connections and use a regularization strategy to solve this problem. Inserting operation-level Dropout after each skip connection can reduce the dependencies between layers and neurons in the network, thereby indirectly controlling the number of skip connections. This approach not only enhances the generalization ability of the model but also helps to effectively control overfitting.

[0100] To construct a lightweight network architecture, the present invention also adopts a multi-objective optimization strategy, which not only considers the accuracy of the model, but also incorporates the number of parameters P(o (i,j) ) and the number of floating-point operations F(o (i,j) ) of the selected operations in the architecture into the optimization strategy. To maintain the differentiability of the objective function, the weighted parameter size and FLOPs between a pair of nodes can be calculated.

[0101] Further, (4) the first neural network architecture is trained based on the training set, and a quantization-aware training method is adopted.

[0102] Further, (8) the second neural network architecture is trained based on the training set, and also a quantization-aware training method is adopted.

[0103] During the quantization-aware training process, pseudo-quantization nodes are added during the model training process to simulate the errors caused during the quantization process. The pseudo-quantization nodes are used to simulate the quantization behavior of the hardware during the model training process. It inserts pseudo-quantization operations during the floating-point training, enabling the model to perceive in advance the performance impact after quantization, thereby reducing the accuracy loss caused by quantization;

[0104] Time node: The pseudo-quantization nodes are usually inserted into the computational graph at the beginning stage of the model training. Based on the pre-trained model, after inserting the pseudo-quantization operator into the computational graph, training or fine-tuning begins;

[0105] During the forward propagation process, quantization operations are performed on the weights W and activation values A of the model, converting them from the floating-point format to the integer format. This process can be expressed as:

[0106] y = WA(x);

[0107] y q = o(G(x,b),Q(W,b));

[0108] where y represents the linear transformation of the output of the activation function g(x) by the weight matrix W, b represents the quantization bit width, and y q is the quantized output, which is obtained by combining the quantized activation G(x,b) and the quantized weight Q(W,b) through the convolutional layer o;

[0109] The quantization function G(x,b) is a half-wave Gaussian quantization function, and its expression is:

[0110]

[0111] The quantization function Q(x,b) is uniform quantization, and its expression is:

[0112]

[0113] Δ(b) is the quantization step calculated based on the bit width b, rounded to the nearest integer by round, and μ and σ are the mean and standard deviation of the signal respectively.

[0114] During the backpropagation stage, the impact of quantization is processed. During this process, the straight-through estimator (STE) is used to approximate the gradient of the quantization operation, ignoring the non-differentiability of the quantization operation, and directly taking the loss function The gradient of the quantization output As the gradient of the loss function with respect to the original input So that the model containing the quantization step can be trained.

[0115] Through quantization-aware training, the simulated error can be used to adjust the original weights, enabling the model to learn to adapt to the possible impact of quantization. This method allows the model to take into account the quantization effect during training while maintaining the normal backpropagation of the gradient flow.

[0116]

[0117] Among them, is the gradient of the loss function with respect to the quantized weights, is the gradient of the quantization function with respect to the original weights.

[0118] It should be understood that by using quantization-aware training, the weights and activations of the model are simulated for quantization during the forward propagation, enabling the model to not only maintain relative accuracy but also reduce the computational amount of the model.

[0119] In order to maintain a relatively high accuracy rate while reducing the computational amount and FLOPs of the model, quantization-aware training (QAT) is proposed.

[0120] This embodiment provides a lightweight differentiable architecture search method based on quantization awareness, which is divided into three parts as a whole: (1) Search space: During the process of constructing a lightweight neural network architecture, dynamic channel adjustment convolution is introduced as a candidate operation, aiming to optimize the search space; (2) Search strategy: An innovative improvement is made to the traditional progressive search strategy. The present invention implements an intelligent adjustment mechanism that allows the search process to flexibly adjust subsequent operations according to the real-time performance feedback of the model; (3) Quantization-aware training: By converting the weights and activation values from floating-point numbers to low-bit-width integer representations during the training stage, the model gradually adapts to the quantization operation during the training process.

[0121] The specific steps are as follows:

[0122] S1: Define the search space. The search space is the set of all potential neural network architectures, which encompasses various types of network layers, the arrangement order between layers, connection patterns, and hyperparameter configurations. Designing a reasonable search space is crucial for improving the efficiency of NAS because it directly affects the computational cost of the search process and the performance of the final model;

[0123] S2: Design the search strategy. The search strategy is the core part of exploring the search space to discover the optimal network architecture. The key lies in how to effectively explore and traverse the potential architecture space. This involves finding the right balance between exploration (discovering new possibilities) and exploitation (optimizing known solutions). The goal is to identify a network architecture with excellent performance in the shortest possible time while preventing the search process from prematurely fixing on a non-optimal solution;

[0124] S3: Performance evaluation. Evaluate the performance of each candidate network, which usually involves training the candidate model on a specific dataset and using metrics such as accuracy, loss value, and inference time to measure the effectiveness of the model. Performance evaluation is a key step in the NAS process because it determines which architectures are more likely to be further optimized or deployed in practical applications;

[0125] S4: Continuously optimize through iteration until the optimal network architecture is found.

[0126] In step S1, lightweight candidate operations are used to incorporate DCAC into the search space. Based on the search space of DARTS (including: 3×3 SepConv, 3×3 DilConv, 5×5 SepConv, 5×5 DilConv, 3×3 Max Pooling, 3×3 Avg Pooling, Zero, Identity), the computationally relatively large 5×5 dilated convolution operation is removed. The candidate operations contained in the search space of this example are: DCAC, 3×3 SepConv, 3×3 DilConv, 5×5 SepConv, 3×3 Max Pooling, 3×3 Avg Pooling, Zero, Identity.

[0127] In step S2, an evaluable progressive search measurement is adopted to balance the relationship of the architecture depth between search and evaluation. The total process is divided into two stages. In the initial stage, it is composed of 5 cells stacked together, and there are 4 operations in the candidate operation set, and each operation has a corresponding operation weight (corresponding to the number next to the connection line in the figure). This process will evaluate the search results and performance of the current stage. If the effect is good, it will enter the next stage, an architecture composed of 16 cells stacked together. At this time, there are only 2 operations left in the candidate set.

[0128] The performance feedback mechanism can dynamically adjust the search direction according to real-time performance data. At the end of each stage, the system evaluates the current architecture based on key performance indicators such as accuracy (Acc) and floating-point operations (FLOPs), and decides whether to continue the in-depth search or adjust the search path accordingly. This decision-making process is achieved by comparing the performance score with a preset threshold.

[0129] The performance evaluation formula is: PS = γ·Acc - δ(Lat + FLOPs);

[0130] The core idea of this formula is to balance the accuracy and complexity of the model. The higher the accuracy, the higher the performance score; while the higher the latency and computational cost, the lower the performance score. By adjusting the weights γ and δ, the relative importance of accuracy and complexity in performance evaluation can be controlled. This can help the NAS system decide whether to continue exploring a certain candidate architecture, or whether to adjust the search strategy to find a better architecture. For example, if a candidate architecture has a high accuracy but also high latency and computational cost, it may not get a high performance score, and the NAS system may choose to explore other more efficient architectures. On the contrary, if an architecture has a reasonable accuracy while having low latency and computational cost, it may get a high performance score and thus be selected as a better candidate architecture.

[0131] In this process, to avoid falling into local optima, a regularization strategy is adopted to limit the use of skip connections, because although they help with fast gradient descent, they may lead to overfitting. By introducing operation-level Dropout after skip connections, the dependencies between layers and neurons are reduced, effectively controlling the number of skip connections, thereby enhancing the generalization ability of the model.

[0132] In this example, during the model training process, pseudo-quantization nodes are added to simulate the errors caused during the quantization process. During the forward propagation process, quantization operations are performed on the weights and activations of the model. In the backpropagation stage, to handle the impact of quantization, the straight-through estimator (STE) is used to approximate the gradient of the quantization operation. Specifically, the straight-through estimator allows the network to perform quantization during forward propagation, while ignoring the quantization operation during backpropagation and directly passing the gradient from the output layer to the input layer. In this way, the model can learn how to adjust the weights under quantization conditions to minimize the loss function. Through this method, the model can gradually adapt to the possible impact of quantization and thus maintain high performance after quantization deployment.

[0133] In step S3, this example adopts a multi-objective optimization strategy, aiming to find the best balance among multiple objectives. This strategy considers multiple performance indicators such as accuracy, number of parameters, and floating-point operations. During this process, performance evaluation is a key link, which involves how to quantify and compare the trade-offs between different objectives.

[0134] In the operation of multi-objective optimization, multiple objective functions are first defined, and these functions represent different aspects that need to be optimized in the NAS process. For example, one objective may be to maximize the accuracy CE of the model, while another objective may be to minimize the number of model parameters and FLOPs: There may be conflicts between these objective functions. To solve this conflict, an attempt is made to find a balance point among the model accuracy, the number of parameters, and the computational complexity. By adjusting the weight coefficients λ 1 and λ 2 to explore different trade-off schemes, the objective function is:

[0135] The performance metrics measured in this step not only provide a quantitative measurement standard for the search algorithm to compare and judge the advantages and disadvantages of different candidate network architectures, but also ensure that the optimal network structure can be efficiently identified during the search process. The core of performance evaluation lies in balancing multiple performance metrics, so as to quickly screen out potential architectures and avoid resource waste under limited computing resources.

[0136] The present invention relates to the field of automated machine learning, aiming to automatically design and optimize lightweight neural network architectures to adapt to resource-constrained environments. The method is implemented through three main parts: the definition of the search space, the innovative improvement of the search strategy, and quantization-aware training. In the search space part, dynamic channel adjustment convolution is introduced as a candidate operation to optimize the structure of the search space. In the search strategy part, an intelligent adjustment mechanism is implemented, allowing the search process to flexibly adjust subsequent operations according to the real-time performance feedback of the model. In the quantization-aware training part, by adding pseudo-quantization nodes during the model training process to simulate the errors caused during the quantization process, the model gradually adapts to the quantization operation. The method is iteratively optimized until the optimal network architecture is found. In addition, the present invention also includes performance evaluation and multi-objective optimization strategies to find the best balance point among multiple performance metrics. The technical solution of the present invention improves the efficiency and effectiveness of neural network architecture search, and at the same time ensures the high performance of the model after quantization deployment, and is applicable to application scenarios that require lightweight and efficient neural network models.

[0137] Embodiment 2

[0138] This embodiment provides an image classification system based on quantization awareness and neural architecture search, including:

[0139] An acquisition module, which is configured to: acquire a training set, where the training set is an image with a known image classification result; construct a neural network architecture based on the training set; input the training set into the neural network to train the neural network, obtaining a trained neural network; input an image to be classified into the trained neural network to obtain an image classification result; wherein, constructing the neural network architecture based on the training set includes:

[0140] (1) Construct a dynamic channel adjustment convolution candidate operation model, and use the dynamic channel adjustment convolution candidate operation model as the generated candidate operation; merge the generated candidate operation with the original candidate operations to obtain a candidate operation set;

[0141] (2) Set an initial Normal unit and an initial Reduction unit based on the candidate operation set;

[0142] (3) Construct a first neural network architecture based on the initial Normal unit and the initial Reduction unit;

[0143] (4) Train the first neural network architecture based on the training set. During the training process, if the loss function value is higher than the set threshold, return to (2); if the loss function value is lower than the set threshold, it means that the current first neural network architecture is the final neural network architecture, and end.

[0144] In the above embodiments, the descriptions of each embodiment have their own focuses. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0145] The proposed system can be implemented in other ways. For example, the system embodiments described above are merely illustrative. For example, the above module division is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed.

[0146] Embodiment 3

[0147] This embodiment also provides an electronic device, including: one or more processors, one or more memories, and one or more computer programs; wherein, the processor is connected to the memory, and the above one or more computer programs are stored in the memory. When the electronic device runs, the processor executes the one or more computer programs stored in the memory, so that the electronic device executes the method described in Embodiment 1 above.

[0148] It should be understood that in this embodiment, the processor may be a central processing unit (CPU), or the processor may also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.

[0149] The memory may include a read-only memory and a random access memory, and provide instructions and data to the processor. A part of the memory may also include a non-volatile random access memory. For example, the memory may also store information about the device type.

[0150] In the implementation process, each step of the above method may be completed by the integrated logic circuit in the hardware of the processor or the instructions in the form of software.

[0151] The method in Embodiment 1 may be directly embodied as being executed and completed by the hardware processor, or executed and completed by the combination of the hardware and software modules in the processor. The software module may be located in a mature storage medium in the art such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. This storage medium is located in the memory, and the processor reads the information in the memory and combines its hardware to complete the steps of the above method. To avoid repetition, it will not be described in detail here.

[0152] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with this embodiment can be implemented by electronic hardware or the combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0153] Embodiment 4

[0154] This embodiment also provides a computer-readable storage medium for storing computer instructions. When the computer instructions are executed by the processor, the method described in Embodiment 1 is completed.

[0155] The above is only the preferred embodiment of the present invention and is not used to limit the present invention. For those skilled in the art, the present invention may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.

Claims

1. Image classification method based on quantization perception and neural architecture search, characterized by: include: Obtaining a training set, wherein the training set is an image of known image classification results; Based on the training set, build a neural network architecture; Input the training set into the neural network, train the neural network, and obtain the trained neural network; The image to be classified is input into the trained neural network to obtain the image classification result; wherein, based on the training set, the neural network architecture is constructed, including: (1) constructing a dynamic channel adjustment convolution candidate operation model, and using the dynamic channel adjustment convolution candidate operation model as the generated candidate operation; merging the generated candidate operation with the original candidate operation to obtain a candidate operation set; (2) Based on the candidate operation set, set the initial Normal unit and the initial Reduction unit; (3) constructing the first neural network architecture based on the initial Normal unit and the initial Reduction unit; (4) Train the first neural network architecture based on the training set. During the training process, if the loss function value is higher than the set threshold, return to (2); if the loss function value is lower than the set threshold, it means that the current first neural network architecture is the final neural network architecture, and the process ends.

2. The image classification method based on quantitative perception and neural architecture search as claimed in claim 1, characterized in that: The step of constructing a neural network architecture based on the training set further includes: (5) Determine whether the performance of the final neural network architecture exceeds the set threshold. If yes, end the process; if no, proceed to (6); (6) Based on the candidate operation set, set the initial Normal unit and the initial Reduction unit; (7) constructing a second neural network architecture based on the initial Normal unit and the initial Reduction unit; (8) Train the second neural network architecture based on the training set. During the training process, if the objective function is higher than the set threshold, return to (6); if the objective function value is lower than the set threshold, it means that the current second neural network architecture is the final neural network architecture, and the training ends.

3. The image classification method based on quantitative perception and neural architecture search as claimed in claim 1, characterized in that: The dynamic channel adjustment convolution candidate operation model comprises: an input layer, a channel importance evaluation module, an adaptive channel selection module, a deep convolution module and an output layer connected in sequence; the input layer is used to input an image of a training set, and the output layer is used to output a feature map of the image; Among them, the channel importance evaluation module includes: an average pooling layer, a first fully connected layer, a first activation function layer, a second fully connected layer, and a second activation function layer connected in sequence; the channel importance evaluation module is used to output a scoring vector; Among them, the adaptive channel selection module selects the top k channels with the highest score vectors; The deep convolution module includes: a deep convolution layer, a batch normalization layer and a third activation function layer connected in sequence; the third activation function layer of the deep convolution module outputs a feature map of the image; Among them, the input layer is also connected to the input end of the third activation function layer through a residual connection.

4. The image classification method based on quantitative perception and neural architecture search as claimed in claim 3, characterized in that: Dynamic channel adjustment convolution candidate operation model, the working process includes: The input layer inputs the image of the training set, and the channel importance evaluation module outputs the evaluation value a; a=σ(W2·ReLU(W1·GAP(X))); Among them, the channel importance evaluation vector GAP is a global average pooling operation that takes the input feature map Convert to size C in The vector of and is the weight matrix of the attention network, σ is the sigmoid activation function; According to the evaluation value a output by the channel importance evaluation module, channels with evaluation values ​​higher than the set threshold are selected for convolution calculation. For unselected channels, they are multiplied by 0 to suppress them, keeping the spatial dimension of the feature map unchanged. The calculation formula is as follows: X selected =X[:,:,top-k(a,C selected )]; According to the score vector a, select the top-k (a, C selected ) The channel with high score is convolutionally calculated, X selected is the selected feature map, which only contains the top k channels with the highest scores selected by the score vector a, where ":" means selecting all rows and columns; Finally, the adaptive channel selection module performs channel selection and only retains the top k channels with the highest scores; The deep convolution module performs deep convolution operations on the top k channels with the highest scores.

5. The image classification method based on quantitative perception and neural architecture search as claimed in claim 1, characterized in that: Based on the candidate operation set, set the initial Normal unit and initial Reduction unit. include: Among them, the initial Normal unit, the construction process includes: sorting the candidate operations in the candidate operation set from large to small based on the probability value, selecting the top M candidate operations, and connecting the selected candidate operations in sequence to obtain a directed acyclic graph, in which each node represents a feature graph, and each edge represents a candidate operation; the probability value is calculated according to the Softmax function; the input value of the Softmax function is a vector containing all candidate operations; the Softmax function converts each element of the input vector into a probability; Among them, the initial Reduction unit, the construction process includes: sorting the candidate operations in the candidate operation set from large to small based on the probability value, selecting the top M candidate operations, and connecting the selected candidate operations in sequence to obtain a directed acyclic graph, where each node in the directed acyclic graph represents a feature graph, and each edge represents a candidate operation; when constructing the directed acyclic graph of the initial Reduction unit, at least one candidate operation is a convolution operation with stride = 2.

6. The image classification method based on quantitative perception and neural architecture search as claimed in claim 1, characterized in that: Based on the initial Normal unit and the initial Reduction unit, a first neural network architecture is constructed; wherein the first neural network architecture includes: A first initial Normal unit, a first initial Reduction unit, a second initial Normal unit, a second initial Reduction unit, and a third initial Normal unit are connected in sequence.

7. The image classification method based on quantitative perception and neural architecture search as claimed in claim 2, characterized in that: (5) Determine whether the performance of the final neural network architecture exceeds the set threshold. If yes, end the process; if no, proceed to (6); wherein the performance is a performance score calculated based on the accuracy and the number of floating-point operations; Performance score QPS, the formula is as follows: QPS = γ·Acc-δ(Lat+FLOPs) Among them, γ is the weight of accuracy, which indicates the impact of accuracy on performance; δ is the weight of complexity, which is used to indicate the degree of penalty of latency and computational complexity on performance; Acc is accuracy, FLOPs is the number of floating-point operations; Lat represents latency, which is the time required for the model to make predictions on input data; (7) Based on the initial Normal unit and the initial Reduction unit, construct a second neural network architecture; wherein the second neural network architecture includes: Four initial Normal units, one initial Reduction unit, four initial Normal units, one initial Reduction unit, and four initial Normal units connected in sequence; (8) Train the second neural network architecture based on the training set. During the training process, if the objective function is higher than the set threshold, return to (6); if the objective function value is lower than the set threshold, it means that the current second neural network architecture is the final neural network architecture, and the process ends. The specific expression of the objective function is: Among them, CE is the cross entropy loss between evaluation layer structures, Used to limit the size of parameters, Used to limit the number of floating-point operations, λ1 and λ2 are the weights to balance accuracy, parameter size, and FLOPs.

8. Image classification system based on quantization perception and neural architecture search, characterized by, include: An acquisition module is configured to: acquire a training set, wherein the training set is an image of known image classification results; Based on the training set, build a neural network architecture; Input the training set into the neural network, train the neural network, and obtain the trained neural network; The image to be classified is input into the trained neural network to obtain the image classification result; wherein, based on the training set, the neural network architecture is constructed, including: (1) constructing a dynamic channel adjustment convolution candidate operation model, and using the dynamic channel adjustment convolution candidate operation model as the generated candidate operation; merging the generated candidate operation with the original candidate operation to obtain a candidate operation set; (2) Based on the candidate operation set, set the initial Normal unit and the initial Reduction unit; (3) constructing the first neural network architecture based on the initial Normal unit and the initial Reduction unit; (4) Train the first neural network architecture based on the training set. During the training process, if the loss function value is higher than the set threshold, return to (2); if the loss function value is lower than the set threshold, it means that the current first neural network architecture is the final neural network architecture, and the process ends.

9. An electronic device, comprising: a memory for non-transitory storage of computer readable instructions; as well as a processor for executing the computer readable instructions, When the computer-readable instructions are executed by the processor, the method described in any one of claims 1 to 7 is executed.

10. A storage medium, characterized in that: The computer-readable instructions are non-transitory stored, wherein when the non-transitory computer-readable instructions are executed by a computer, the method according to any one of claims 1 to 7 is performed.

Citation Information

Cited By

  • Heterogeneous graph neural network-based search model establishment method and system and application

    CN120764639A