A method for constructing an image classification model without training in the search phase

By calculating gradient signal-to-noise ratio proxy index on the image classification data set and evaluating and selecting neural network model structures, the problem of insufficient testing accuracy caused by the neglect of generalization in the existing technology is solved, and efficient neural network architecture search and better classification accuracy are achieved.

CN116310578BActive Publication Date: 2025-08-29INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310314946.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-28
Publication Date
2025-08-29
Estimated Expiration
2043-03-28

AI Technical Summary

Technical Problem

In the existing neural network architecture search methods that do not require training, proxy indicators mainly focus on the training of the network and ignore the generalization of the network, making it difficult to guarantee the test accuracy of the selected model after training.

Method used

In the search stage of the model structure, the performance of each model to be selected is evaluated by calculating the signal-to-noise ratio proxy index of the gradient mean to variance ratio of each image sample in the image classification dataset, thereby selecting the target network model and performing image classification training in the training stage.

Benefits of technology

It improves the performance and classification accuracy of the target network architecture, and greatly reduces the time and resource overhead of neural network architecture search, saving computing resources and energy consumption.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116310578B_ABST
    Figure CN116310578B_ABST
Patent Text Reader

Abstract

The present invention provides a method for constructing an image classification model that does not require training in a search phase, comprising: in a model structure search phase, executing steps A1-A4: A1, sampling multiple candidate model structures from multiple neural network model structures contained in a preset search space; A2, for each candidate model structure, using each image sample in an evaluation set to perform a forward propagation and a backward propagation in the candidate model structure, to obtain the gradient of each parameter corresponding to each image sample under the candidate model structure; A3, determining a signal-to-noise ratio proxy index for each candidate model structure based on the gradient of each parameter; A4, selecting a target network model from multiple candidate model structures based on the signal-to-noise ratio proxy indexes of all candidate model structures; in a training phase, performing image classification training on the target network model based on a training set extracted from an image classification data set to obtain a trained image classification model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of automated machine learning, specifically to the field of image classification technology, and more specifically to a method for constructing an image classification model that does not require training during the search phase. Background Art

[0002] The problem that automated machine learning aims to solve is to automatically build machine learning processes for one or more specific machine learning tasks, without human expert intervention and under limited computing resources. Research directions include automated feature extraction, automated model selection, automated model parameter tuning, automated model structure search, automated model evaluation, meta-learning, and transfer learning. Neural network architecture search, a key component, aims to automatically search within a predefined search space to obtain the optimal network structure. The performance of structures obtained through neural network architecture search has been verified to exceed that of manually designed network structures in multiple tasks. Therefore, automated network structure design has attracted widespread attention from researchers.

[0003] A key step in neural network architecture search is evaluating the performance of each architecture in the search space. Early methods required training each architecture individually to convergence and then validating their performance. This process was time-consuming and resource-intensive, requiring hundreds or even thousands of GPU days to complete the search. Later, a weight-sharing technique was proposed to reduce the time required for architecture performance evaluation. Weight-sharing involves sharing the weights of the same operations between different subnetworks. This allows only a single supernetwork to be trained, and the subnetworks can directly inherit the supernetwork's weights for performance verification. Specifically, gradient-trained supernetworks introduce architectural parameters, training network weights and architectural parameters alternately on training and validation sets, then evaluating the performance of candidate architectures based on the size of the architectural parameters. On the other hand, single-path sampling-based supernetworks, after training convergence, use an evolutionary algorithm to select a large number of subnetworks that inherit the supernetwork's weights for performance verification. In either case, weight-sharing-based supernetwork training schemes require training before performance verification to conduct a search based on the training results. After the search is complete, the resulting target network is retrained. This training, search, and retraining approach incurs significant computational overhead.

[0004] In the field of image classification, neural network models are becoming increasingly large. To achieve high-performance models, it may be necessary to identify the desired model structure from a large number of existing search spaces. Given the large search volume and computational complexity of the models, the continued use of a training, search, and retraining approach consumes significant computing power and time, resulting in inefficiency and a significant waste of resources.

[0005] To minimize the overhead of neural network performance evaluation, several training-free search methods have been proposed. These methods employ different proxy metrics, which are then calculated at neural network initialization and used as indicators of network performance to estimate network performance. This eliminates the need for any training, significantly improving the speed of neural network architecture search. However, the accuracy of performance evaluation is directly affected by the design of the proxy metric. Previous proxy metrics, such as SNIP, GraSP, and SynFlow, were either inspired by empirical experience with neural network pruning at initialization or lacked accurate performance prediction. These training-free proxy metrics perform differently across different search spaces and often fail to outperform the simplest proxy, namely the parameter count of the neural network. On the other hand, other proxy indicators are derived from theoretical analysis of neural network training. For example, the NASWOT method predicts the performance of the network by analyzing the linear region of the neural network; the neural network tangent kernel describes the training dynamics of the neural network. The NASI and TE-NAS methods predict the performance of the network by calculating the neural network tangent kernel indicator at initialization. However, the neural network tangent kernel can only be guaranteed under extremely wide networks, which has certain limitations. In addition, this proxy indicator designed based on the network training dynamics lacks consideration of generalization. There are also proxy indicators that analyze the gradient loss landscape of the neural network training process and then design new proxy indicators to predict network performance.

[0006] The proxy metrics used in existing non-training search methods primarily focus on network trainability while neglecting generalization, making it difficult to guarantee the test accuracy of the selected model after training. Therefore, improvements to existing technologies are needed. Summary of the Invention

[0007] Therefore, the purpose of the present invention is to overcome the above-mentioned defects of the prior art and provide a method for constructing an image classification model that does not require training during the search stage.

[0008] The purpose of the present invention is achieved through the following technical solutions:

[0009] According to a first aspect of the present invention, a method for constructing an image classification model that does not require training in the search phase is provided, comprising: in the search phase of the model structure, executing steps A1-A4: A1, sampling multiple model structures to be selected from multiple neural network model structures contained in a preset search space; A2, for each model structure to be selected, using each image sample in an evaluation set extracted from an image classification data set to perform a forward propagation in the model structure to obtain an image classification result, and calculating the gradient based on the classification loss of the image classification result and back-propagating to obtain the image classification result of each image sample under the model structure to be selected; A3. Determine the signal-to-noise ratio proxy index of each candidate model structure according to the gradient of each trainable parameter corresponding to each image sample under each candidate model structure, wherein the signal-to-noise ratio proxy index is positively correlated with the ratio of the square of the mean of the gradient of the parameter corresponding to each image sample to the variance of the gradient of the parameter; A4. Select a target network model from multiple candidate model structures according to the signal-to-noise ratio proxy index of all candidate model structures; in the training stage, perform image classification training on the target network model according to the training set extracted from the image classification dataset to obtain a trained image classification model.

[0010] In some embodiments of the present invention, the signal-to-noise ratio proxy indicator is the sum of the ratios of the square of the gradient of the parameter corresponding to each image sample in the evaluation set to the variance of the gradient of the parameter.

[0011] In some embodiments of the present invention, the signal-to-noise ratio proxy indicator is the sum of the ratios of the squares of the gradients of the parameters corresponding to each image sample in the evaluation set to the variance of the gradients of the corrected parameters, wherein the gradients of the corrected parameters are the sum of the variance of the gradients of the parameters corresponding to the image sample and a preset regularization value.

[0012] In some embodiments of the present invention, the signal-to-noise ratio proxy indicator is determined in the following manner:

[0013]

[0014] Where N represents the total number of image samples in the evaluation set, X i represents the i-th image sample in the evaluation set, Y i represents the label of the i-th image sample, θ j represents the jth parameter, represents the parameter θ calculated for the i-th image sample j The gradient, Represents the square of the mean of the gradient of the parameter corresponding to the i-th image sample, represents the variance of the gradient of the parameter corresponding to the i-th image sample, and ξ represents the preset regularization value.

[0015] In some embodiments of the present invention, during the model structure search phase, for each candidate model structure, each image sample in the evaluation set corresponds to only one forward propagation and one backward propagation to determine the gradient of each trainable parameter corresponding to the image sample under the candidate model structure, and to shield the process of updating the parameters according to the gradient of the parameters.

[0016] In some embodiments of the present invention, in step A4, the model structure with the highest value of the signal-to-noise ratio proxy indicator is selected from all candidate model structures as the target network model.

[0017] According to a second aspect of the present invention, there is provided a training device for an image classification model for implementing the method described in the first aspect, comprising: a neural network sampling module for sampling a plurality of candidate model structures from a plurality of neural network model structures contained in a preset search space; a neural network performance estimation module for, for each candidate model structure, performing a forward propagation on the candidate model structure using each image sample in an evaluation set extracted from an image classification data set to obtain an image classification result, and calculating the gradient based on the classification loss of the image classification result and performing back propagation to obtain each trainable performance of each image sample corresponding to the candidate model structure; The invention also provides a method for determining a signal-to-noise ratio proxy index for each candidate model structure based on the gradient of each trainable parameter corresponding to each image sample under each candidate model structure, wherein the signal-to-noise ratio proxy index is positively correlated with the ratio of the square of the mean of the gradient of the parameter corresponding to each image sample to the variance of the gradient of the parameter; a target network selection module is used to select a target network model from multiple candidate model structures based on the signal-to-noise ratio proxy index of all candidate model structures; a target network training module is used to perform image classification training on the target network model based on a training set extracted from an image classification data set to obtain a trained image classification model.

[0018] According to a third aspect of the present invention, there is provided an image classification method, comprising: acquiring an image to be predicted; inputting the image to be predicted into a trained image classification model trained by the method described in the first aspect or the device described in the second aspect, and outputting an image classification result.

[0019] According to the fourth aspect of the present invention, an electronic device is provided, comprising: one or more processors; and a memory, wherein the memory is used to store executable instructions; the one or more processors are configured to implement the steps of the method described in the first aspect and / or the third aspect by executing the executable instructions.

[0020] Compared with the prior art, the advantages of the present invention are:

[0021] In the search stage of the model structure, the present invention sets an evaluation set extracted from the image classification data set, uses each image sample in the evaluation set to perform a forward propagation and a corresponding back propagation on the model structure to be selected, and determines the signal-to-noise ratio proxy index of each model structure to be selected based on the gradient of each trainable parameter corresponding to each image sample under each model structure to be selected, wherein the signal-to-noise ratio proxy index is positively correlated with the ratio of the square of the mean of the gradient of the parameter corresponding to different image samples to the variance of the gradient of the parameter. The signal-to-noise ratio proxy index thus determined takes into account the generalization of the model structure, improves the performance of the selected target network architecture, and the selected target network has better classification accuracy when used for image classification. At the same time, since the search process does not require training the neural network to directly evaluate the performance, the time and resource overhead of the neural network architecture search is greatly reduced, saving computing resources, energy consumption and model development efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The embodiments of the present invention are further described below with reference to the accompanying drawings, in which:

[0023] Figure 1 2. A flowchart of a method for constructing an image classification model that does not require training during the search phase according to an embodiment of the present invention;

[0024] Figure 2 A simplified schematic diagram of a search space according to an embodiment of the present invention;

[0025] Figure 3 is a simplified schematic diagram of a candidate model structure according to an embodiment of the present invention;

[0026] Figure 4 A flowchart of a method for constructing an image classification model that does not require training during the search phase according to another embodiment of the present invention;

[0027] Figure 5 2 is a module diagram of a training device for an image classification model according to an embodiment of the present invention. DETAILED DESCRIPTION

[0028] In order to make the purpose, technical solutions and advantages of the present invention more clearly understood, the present invention is further described in detail below through specific embodiments in conjunction with the accompanying drawings. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0029] As mentioned in the background technology section, the training, search and retraining methods adopted by the existing technology require a large computational overhead. The proxy indicators of the existing search methods that do not require training mainly focus on the trainability of the network and ignore the generalization of the network, making it difficult to guarantee the test accuracy of the selected model after training. The inventors have determined through research that only considering trainability may cause network overfitting, while generalization can ensure that high test accuracy can be achieved even after training. Therefore, when designing a search proxy indicator for a neural network architecture that does not require training, considering generalization is a critical part. By analyzing the relationship between the generalization of the neural network and the mean and variance of one-step gradient descent, the inventors deduced that the gradient signal-to-noise ratio plays a key role in reducing the generalization gap, that is, the larger the gradient signal-to-noise ratio, the smaller the generalization gap, which means that the generalization of the neural network is better. In this regard, during the search stage of the model structure, the present application sets an evaluation set extracted from the image classification data set, and uses each image sample in the evaluation set to perform a forward propagation and a corresponding back propagation on the model structure to be selected. According to the gradient of each trainable parameter corresponding to each image sample under each model structure to be selected, the signal-to-noise ratio proxy index of each model structure to be selected is determined, wherein the signal-to-noise ratio proxy index is positively correlated with the ratio of the square of the mean of the gradient of the parameter corresponding to different image samples to the variance of the gradient of the parameter. The signal-to-noise ratio proxy index thus determined takes into account the generalization of the model structure, improves the performance of the selected target network architecture, and the selected target network has better classification accuracy when used for image classification. At the same time, since the search process does not require training the neural network to directly evaluate the performance, the time and resource overhead of the neural network architecture search is greatly reduced, saving computing resources, energy consumption and model development efficiency.

[0030] According to one embodiment of the present invention, a method for constructing an image classification model that does not require training during the search phase (or referred to as a method for constructing an image classification model or a method for training an image classification model) is provided, comprising:

[0031] 1. The model structure search phase, which is used to select the target network model;

[0032] 2. The training phase is used to train the selected target network model to obtain a trained image classification model.

[0033] In order to better understand the present invention, the contents of these two parts are introduced separately in conjunction with specific embodiments below.

[0034] 1. Model structure search phase

[0035] In the search phase of the model structure, see Figure 1 , execute steps K1-K4. Where:

[0036] Step K1: Sampling multiple candidate model structures from multiple neural network model structures contained in a preset search space.

[0037] According to one embodiment of the present invention, the preset search space includes a variety of different model structures, and the structures of the convolution kernels in the processing layers of the various model structures and / or the connection methods between the processing layers are different. For example, different convolution kernel sizes, different connection methods between layers, etc. Taking the search space based on the unit structure as an example, Figure 2 As shown in a, each unit structure in the search space is defined as a directed acyclic graph, with multiple different arrows between any two nodes (for example, there are multiple arrows between feature graphs 0→1, 0→2, 0→3, 1→2, and 2→3), and each arrow corresponds to a specific processing layer structure; in addition, there may be some scenarios that make the connection between processing layers different, such as comparing Figure 2 b and Figure 2 a, where Figure 2 b is missing some connections, Figure 2 In b, there is no connection between feature maps 0→3. To quickly verify the effectiveness of the present invention, we can use model structures proposed by researchers as search spaces, such as the search space NAS-Bench-201, which provides 15,625 neural network model structures, or the search space NAS-Bench-101. It should be understood that any other neural network model structure customized by the implementer can also be used to form the search space, and the present invention does not impose any restrictions on this.

[0038] Step K2: For each candidate model structure, use each image sample in the evaluation set extracted from the image classification dataset to perform a forward propagation on the candidate model structure to obtain the image classification result, and calculate the gradient based on the classification loss of the image classification result and backpropagate to obtain the gradient of each trainable parameter corresponding to each image sample under the candidate model structure.

[0039] According to one embodiment of the present invention, an image classification dataset includes multiple image samples and a label for each image sample, and the label indicates the true value of the image category of the corresponding image sample. The image classification dataset can adopt an existing image classification dataset or an image classification dataset customized by the implementer. Existing image classification datasets include, for example, CIFAR-10, CIFAR-100, and ImageNet datasets. The set of categories (label space) of the image classification dataset can be customized. For example, the label space of CIFAR-10 is airplane, car, bird, cat, deer, dog, frog, horse, ship, and truck, and the category (label) to which each image sample belongs is one of airplane, car, bird, cat, deer, dog, frog, horse, ship, and truck. In the image classification dataset customized by the implementer, the label indicates the true value of the category of the corresponding sample image (such as human, animal category, object category, road sign category, etc.).

[0040] According to one embodiment of the present invention, for the field of image classification, the model structure (image classification model) generally includes a feature extraction module and a prediction module. The feature extraction module includes one or more processing layers, which may be convolutional layers, pooling layers, fully connected layers, etc. The feature extraction module is used to extract image features from the input image, and the prediction module is used to perform image classification prediction based on the image features to obtain image classification results. During the forward propagation in step K2, each image sample is input into the model structure to be selected, and the feature extraction module of the model structure to be selected extracts image features from it. The model structure to be selected outputs the image classification result based on the image features; then a gradient backpropagation process is performed, wherein the gradient is calculated based on the classification loss of the image classification result and backpropagated to obtain the gradient of each trainable parameter corresponding to each image sample under the model structure to be selected.

[0041] Step K3: Determine the signal-to-noise ratio proxy indicator for each candidate model structure based on the gradient of each trainable parameter corresponding to each image sample under each candidate model structure, wherein the signal-to-noise ratio proxy indicator is positively correlated with the ratio of the square of the mean of the gradient of the parameter corresponding to different image samples to the variance of the gradient of the parameter.

[0042] According to one embodiment of the present invention, the signal-to-noise ratio proxy indicator is the sum of the ratios of the squares of the gradients of the parameters corresponding to each image sample in the evaluation set to the variance of the gradients of the parameters. Preferably, the signal-to-noise ratio proxy indicator is determined in the following manner:

[0043]

[0044] Where N represents the total number of image samples in the evaluation set, X i represents the i-th image sample in the evaluation set, Y i represents the label of the i-th image sample, θj represents the jth parameter, represents the parameter θ calculated for the i-th image sample j The gradient, Represents the square of the mean of the gradient of the parameter corresponding to the i-th image sample, Represents the variance of the gradient of the parameter corresponding to the i-th image sample.

[0045] The inventors discovered during experiments that if the variance of the gradient of a certain parameter is particularly small, it will dominate the variance of the gradients of other parameters, slightly affecting the evaluation of the signal-to-noise ratio proxy indicator, so it can be further improved. According to one embodiment of the present invention, the signal-to-noise ratio proxy indicator is the sum of the ratios of the squares of the mean of the gradients of the parameters corresponding to each image sample in the evaluation set and the variance of the gradients of the corrected parameters, where the gradient of the corrected parameters is the sum of the variance of the gradients of the parameters corresponding to the image sample and a preset regularization value. Preferably, the signal-to-noise ratio proxy indicator is determined as follows:

[0046]

[0047] Where N represents the total number of image samples in the evaluation set, X i represents the i-th image sample in the evaluation set, Y i represents the label of the i-th image sample, θ j represents the jth parameter, represents the parameter θ calculated for the i-th image sample j The gradient, Represents the square of the mean of the gradient of the parameter corresponding to the i-th image sample, Represents the variance of the gradient of the parameter corresponding to the i-th image sample, and ξ represents a preset regularization value. The preset regularization value is set to 1e-8, 5e-8 or 1e-9, for example. In this embodiment, a small ξ value is set in the denominator to stabilize the calculation, because if the gradient variance of a certain parameter is particularly small, it will dominate the gradient variance of other parameters. In this embodiment, by adding a small perturbation value ξ, the calculation process can be regularized to obtain a more stable proxy indicator. Experiments show that the signal-to-noise ratio proxy indicator improved by this embodiment can improve the positive correlation with the actual performance of the network, thereby illustrating the effectiveness of the signal-to-noise ratio proxy indicator.

[0048] The following is a description of the derivation process for the positive correlation between the signal-to-noise ratio proxy and the generalization of the neural network, including the following steps:

[0049] Step 1: Determine the generalization gap for one step of gradient descent;

[0050] Assume that the given dataset Extract the training set and the test set, where both the training set and the test set follow the same distribution as the sampling in this dataset. The parameters θ of the neural network are optimized by the gradient descent of the loss function L. In each step of the gradient descent, the one-step generalization ratio (OSGR) is defined as the ratio of the test loss of one-step gradient descent to the training loss of one-step gradient descent:

[0051]

[0052] This one-step generalization ratio OSGR can be used to represent the generalization gap. Generally, the neural network is optimized on the training set, so the training loss decreases faster than the test loss, resulting in 0 < OSGR < 1. The closer the one-step generalization ratio OSGR is to 1, the better the generalization performance. Among them, represents the test dataset, represents the training dataset, represents the loss decrease on the test set after one-step gradient descent, represents the loss decrease on the training set after one-step gradient descent, represents the expectation of the loss decrease on the test set, represents the expectation of the loss decrease on the training set.

[0053] Step 2: Verify the generalization guarantee during the gradient descent process.

[0054] Assume that as the training process progresses, the average gradients of the training set and the test set follow the same distribution. Then the means and variances of the gradients on the training set and the test set are:

[0055]

[0056]

[0057] Among them, represents the gradient on the training set, represents the gradient on the test set, E[·] represents the expectation, and Var[·] represents the variance.

[0058] Derived by Taylor expansion, the expectations of the loss decreases after one-step gradient descent on the training set and the test set are:

[0059]

[0060]

[0061] Among them, η represents the learning rate, represents the square of the gradient of the parameters on the training set.

[0062] Therefore, after one-step gradient descent, the one-step generalization ratio OSGR can be reformulated as:

[0063]

[0064] Where n represents the number of samples in the dataset, ρ 2 (θ j ) represents the parameter θ on the training set j The gradient variance of Represents the parameter θ on the training set j The expected square of the gradient, GSNR(θ j ) represents the parameter θ j The ratio of the gradient variance to the expected square;

[0065] This formula shows that generalization performance is related to the variance and expectation of the gradient (GSNR). Specifically, the closer the formula is to 1, the better the generalization, and a larger GSNR ensures that the formula is closer to 1. Therefore, it can be deduced that the generalization of the network is related to GNSR: a larger GSNR indicates better network generalization.

[0066] Step 3: Design the gradient signal-to-noise ratio;

[0067] The gradient signal-to-noise ratio (GSNR) of a neural network parameter is defined as the ratio of the mean square of the gradient of the parameter to its variance:

[0068]

[0069] Generally speaking, GSNR measures the consistency of the gradient update direction of the parameters across different data samples under a specific optimization state. As can be seen from the generalization guarantee of the gradient descent process described above, the larger the gradient signal-to-noise ratio (GSNR), the closer the one-step generalization rate is to 1, indicating better generalization of the neural network.

[0070] Step 4: Design a proxy metric for neural network performance that does not require training;

[0071] Based on the gradient signal-to-noise ratio (GSNR) of neural network parameters, the inventors further enhanced its relationship with the generalization of neural networks and designed a proxy indicator of ξ-GSNR:

[0072]

[0073] Compared with existing technologies of the same period, other technologies either use the amount of neural network parameters or use proxy indicators designed by the tangent kernel of the neural network, etc. The correlation between them and the actual performance of the neural network is very low, that is, they cannot accurately predict the performance of the network. In order to improve the accuracy of the proxy indicator, the present invention calculates the expectation of the ratio between the mean square of the parameter gradient and the gradient variance under different samples when the neural network is initialized. In addition, the inventor sets a small ξ value in the denominator to stabilize the calculation, because if the gradient variance of a certain parameter is particularly small, it will dominate the gradient variance of other parameters, resulting in the indicator value not being stable enough. By adding a small perturbation ξ value, the calculation process can be regularized to obtain a relatively stable proxy score. Experiments show that the proxy indicator improved by the present invention can improve the positive correlation with the actual performance of the network, thereby illustrating the effectiveness of the proxy indicator.

[0074] It should be understood that the implementer may also adjust the above embodiment to obtain other implementation methods, and the present invention does not impose any limitation on this. For example: Among them, γ is expressed as The weight of the setting.

[0075] Step K4: Select a target network model from multiple candidate model structures based on the signal-to-noise ratio proxy indicators of all candidate model structures.

[0076] According to one embodiment of the present invention, from all candidate model structures, the model structure with the highest value of the signal-to-noise ratio proxy index is selected as the target network model. For example, assuming that there are 15,625 neural network model structures in the search space, among which the candidate model structure X has the highest signal-to-noise ratio proxy index, it is selected as the target network model. Figure 3 As shown, a certain model structure is selected as the target network model.

[0077] According to one embodiment of the present invention, in addition to selecting a single target network model at a time, multiple preferred target network models may also be selected. Preferably, a preset number of model structures are selected from a plurality of candidate model structures in descending order of the signal-to-noise ratio proxy indicator values ​​to obtain a preset number of target network models. The preset number may be, for example, 3, 5, or 10.

[0078] According to the above two embodiments, it is determined that one or more target network models have better generalization and can achieve good image classification prediction accuracy after training.

[0079] 2. Training Phase

[0080] In the training phase, the target network model is trained for image classification based on the training set extracted from the image classification dataset to obtain a trained image classification model.

[0081] According to one embodiment of the present invention, after a target network model is selected, image classification training is performed on the target network model based on a training set extracted from an image classification dataset to obtain a trained image classification model. During training, image samples from the training set are input into the target network model, and image classification results are output. A classification loss is calculated based on the image classification results and the corresponding labels indicating the true values ​​of the image categories according to a preset loss function. The gradient of the classification loss is calculated and backpropagated to update the parameters of the feature extraction module and prediction module in the target network model. The preset loss function can be a cross-entropy loss function.

[0082] According to one embodiment of the present invention, when multiple target network models are selected, each target network model is trained for image classification based on a training set extracted from an image classification dataset to obtain multiple trained target network models. Preferably, the performance of the multiple trained target network models can be tested using a test set extracted from the image classification dataset, and the trained target network model with the best performance is selected as the trained image classification model.

[0083] See also Figure 4 , the technical solution of the present invention is described below through another embodiment.

[0084] According to one embodiment of the present invention, a method for constructing an image classification model that does not require training in the search phase is provided, comprising steps S11-S17, wherein:

[0085] Step S11: Define the target data set.

[0086] According to one embodiment of the present invention, the target dataset may adopt a mainstream dataset, such as food-101, CIFAR-10, CIFAR-100, ImageNet dataset, etc. for classification tasks.

[0087] Step S12: Define the search space.

[0088] According to one embodiment of the present invention, the search space defines all possible neural network structures (multiple different model structures), such as different convolution kernel sizes, different connection methods, etc. Taking the search space based on the unit structure as an example, see again Figure 2 , each unit structure in the search space is defined as a directed acyclic graph, and there are multiple different candidate operations (layer structures) between any two nodes. The purpose of neural network architecture search is to determine the optimal operation on each edge.

[0089] Step S13: Sample a single neural network from the search space.

[0090] According to one embodiment of the present invention, it is necessary to sample a single neural network structure (corresponding to the model structure to be selected) from a predefined search space to calculate the signal-to-noise ratio proxy indicator. Taking the search space based on the unit structure as an example, Figure 3 As shown, it is a schematic diagram of a sampled single neural network, in which only one operation is retained between any two intermediate nodes (only one arrow is left between the two nodes, corresponding to a specific layer structure), forming a single neural network structure.

[0091] Step S14: Randomly initialize the parameters of a single neural network structure and determine its signal-to-noise ratio proxy indicator.

[0092] According to one embodiment of the present invention, the neural network parameters sampled in step S13 are then used to sample a small batch size of data from an image classification dataset as an evaluation set. Forward propagation and backpropagation are then performed to obtain the gradient values ​​of all parameters. Based on the gradient values ​​of all parameters, a proxy signal-to-noise ratio (ξ-GSNR) proxy indicator for a single neural network structure is calculated. This proxy indicator reflects the performance of a single neural network structure; a larger proxy indicator indicates better network performance.

[0093] Step S15: Determine whether the search space has been sampled or the algorithm has reached the specified number of iterations.

[0094] According to one embodiment of the present invention, step S15 is used to determine whether all candidate neural network structures in the search space have been sampled. If not, the process proceeds to step S13 to sample the remaining neural network structures from the search space. Alternatively, the process proceeds to the next step after the algorithm reaches a specified number of iterations, such as when the number of sampled neural network structures reaches a certain threshold.

[0095] Step S16: Select the optimal target network based on the signal-to-noise ratio proxy indicator.

[0096] According to one embodiment of the present invention, the signal-to-noise ratio proxy indicators corresponding to all sampled neural networks are counted, and then the neural network structure with the largest signal-to-noise ratio proxy indicator is selected to obtain the optimal target network (target network model) in the search space.

[0097] Step S17: Use the training set to train the target network to obtain a trained image classification model.

[0098] According to one embodiment of the present invention, the target network obtained in step S16 is trained on the training set until convergence or a predetermined number of training times is reached, thereby obtaining a trained image classification model, and its performance indicators are verified on the test set.

[0099] According to one embodiment of the present invention, see Figure 5, providing a training device for an image classification model for implementing the method of the aforementioned embodiment, comprising:

[0100] A neural network sampling module 51 is used to sample multiple candidate model structures from multiple neural network model structures included in a preset search space;

[0101] A neural network performance estimation module 52 is configured to perform a forward propagation on each candidate model structure using each image sample in an evaluation set extracted from an image classification dataset to obtain an image classification result, calculate a gradient based on the classification loss of the image classification result, and perform backpropagation to obtain the gradient of each trainable parameter corresponding to each image sample under the candidate model structure; and determine a signal-to-noise ratio proxy indicator for each candidate model structure based on the gradient of each trainable parameter corresponding to each image sample under each candidate model structure, wherein the signal-to-noise ratio proxy indicator is positively correlated with the ratio of the square of the mean of the gradient of the parameter corresponding to each image sample to the variance of the gradient of the parameter;

[0102] A target network selection module 53 is configured to select a target network model from a plurality of candidate model structures based on a signal-to-noise ratio proxy indicator of all candidate model structures;

[0103] a target network training module 54, configured to perform image classification training on the target network model based on a training set extracted from the image classification data set to obtain a trained image classification model;

[0104] The performance testing module 55 is used to test the performance of the trained image classification model based on the test set.

[0105] According to one embodiment of the present invention, there is provided an image classification method, which is characterized by comprising: acquiring an image to be predicted; inputting the image to be predicted into a trained image classification model trained by the method or apparatus of the aforementioned embodiment, and outputting an image classification result.

[0106] To verify the accuracy of the signal-to-noise ratio proxy metric designed in this paper in predicting network performance, the inventors also conducted a comparative experiment. This comparative experiment was conducted on the NAS-Bench-201 search space. This search space provides 15,625 neural network architectures and the test accuracy of each network architecture on the CIFAR-10 dataset. At the same time, using different proxy metrics, corresponding proxy scores were calculated for each neural network architecture. The proxy scores were then correlated with the provided test accuracy. A higher correlation indicates a more accurate proxy metric in predicting network performance.

[0107] As shown in the table below, compared with other existing technologies of the same period, the improved signal-to-noise ratio proxy indicator (ξ-GSNR proxy indicator) of the present invention achieved the highest correlation, which shows its effectiveness in estimating network performance. The selected target network has higher generalization. Therefore, the trained image classification model obtained after training the target network has higher classification accuracy.

[0108]

[0109] In summary, this paper proposes a method and training device for constructing an image classification model that does not require training during the search phase. Through this improvement, the signal-to-noise ratio proxy metric is strongly correlated with the actual network performance. This not only improves the ranking of different candidate neural networks, but also significantly reduces the time and resource overhead of searching for neural network architectures, as performance can be directly evaluated without requiring neural network training, thereby improving the performance of the target network architecture.

[0110] It should be noted that although the above describes the various steps in a specific order, it does not mean that the steps must be performed in the above specific order. In fact, some of these steps can be executed concurrently or even in a different order as long as the required functions can be achieved.

[0111] The present invention may be a system, a method and / or a computer program product. The computer program product may include a computer-readable storage medium carrying computer-readable program instructions for causing a processor to implement various aspects of the present invention.

[0112] Computer-readable storage media can be a tangible device that holds and stores the instructions used by an instruction execution device. Computer-readable storage media can, for example, include, but are not limited to, an electrical storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination thereof. More specific examples (non-exhaustive list) of computer-readable storage media include: a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanical encoding device, a punch card or a raised structure in a groove on which instructions are stored, for example, and any suitable combination thereof.

[0113] While various embodiments of the present invention have been described above, the above descriptions are intended to be illustrative, non-exhaustive, and not limited to the disclosed embodiments. Many modifications and variations will be apparent to those skilled in the art without departing from the scope and spirit of the described embodiments. The terminology used herein is selected to best explain the principles of the embodiments, their practical applications, or technological improvements in the marketplace, or to enable others skilled in the art to understand the embodiments disclosed herein.

Claims

1. A method for constructing an image classification model that does not require training in the search phase, characterized in that: include: During the model structure search phase, steps A1-A4 are performed: A1. Sampling multiple candidate model structures from multiple neural network model structures included in a preset search space; A2. For each candidate model structure, use each image sample in the evaluation set extracted from the image classification dataset to perform a forward propagation on the candidate model structure to obtain the image classification result. Based on the classification loss of the image classification result, the gradient is calculated and back-propagated to obtain the gradient of each trainable parameter corresponding to each image sample under the candidate model structure; A3. Determine a signal-to-noise ratio proxy indicator for each candidate model structure based on the gradient of each trainable parameter corresponding to each image sample under each candidate model structure, wherein the signal-to-noise ratio proxy indicator is positively correlated with the ratio of the square of the mean of the gradient of the parameter corresponding to each image sample to the variance of the gradient of the parameter. The signal-to-noise ratio proxy indicator is calculated in any of the following ways: Calculation method 1: The signal-to-noise ratio proxy indicator is the sum of the ratios of the square of the gradient of the parameter corresponding to each image sample in the evaluation set to the variance of the gradient of the parameter; Calculation method 2: The signal-to-noise ratio proxy indicator is the sum of the ratios of the square of the mean of the gradient of the parameter corresponding to each image sample in the evaluation set to the variance of the gradient of the modified parameter, where the gradient of the modified parameter is the sum of the variance of the gradient of the parameter corresponding to the image sample and a preset regularization value. The signal-to-noise ratio proxy indicator is determined as follows: in, represents the total number of image samples in the evaluation set, Indicates the first image samples, Indicates the The labels of the image samples, Indicates the parameters, , Indicates the The square of the mean value of the gradient of the parameter corresponding to the image samples, Indicates the The variance of the gradient of the parameter corresponding to the image samples, Represents the preset regularization value; A4. Selecting a target network model from multiple candidate model structures based on the signal-to-noise ratio proxy indicator of all candidate model structures; In the training phase, the target network model is trained for image classification based on the training set extracted from the image classification dataset to obtain a trained image classification model.

2. The method according to claim 1, characterized in that During the model structure search phase, for each candidate model structure, each image sample in the evaluation set corresponds to only one forward propagation and one backward propagation to determine the gradient of each trainable parameter corresponding to the image sample under the candidate model structure, and to shield the process of updating the parameters according to the gradient of the parameters.

3. The method according to claim 1, characterized in that In step A4, the model structure with the highest value of the signal-to-noise ratio proxy indicator is selected from all candidate model structures as the target network model.

4. A training device for an image classification model for implementing the method according to any one of claims 1 to 3, characterized in that: include: A neural network sampling module is used to sample multiple candidate model structures from multiple neural network model structures contained in a preset search space; The neural network performance estimation module is used to perform a forward propagation on each candidate model structure using each image sample in the evaluation set extracted from the image classification dataset to obtain the image classification result. The gradient of the classification loss of the image classification result is calculated and back-propagated to obtain the gradient of each trainable parameter corresponding to each image sample under the candidate model structure. and determining a signal-to-noise ratio proxy indicator for each candidate model structure based on the gradient of each trainable parameter corresponding to each image sample under each candidate model structure, wherein the signal-to-noise ratio proxy indicator is positively correlated with the ratio of the square of the mean of the gradient of the parameter corresponding to each image sample to the variance of the gradient of the parameter; A target network selection module is used to select a target network model from multiple candidate model structures based on the signal-to-noise ratio proxy indicator of all candidate model structures; The target network training module is used to perform image classification training on the target network model according to the training set extracted from the image classification dataset to obtain a trained image classification model.

5. An image classification method, characterized in that: include: Obtain the image to be predicted; The image to be predicted is input into a trained image classification model obtained by training using the method according to any one of claims 1 to 3 or the apparatus according to claim 4, and an image classification result is output.

6. A computer-readable storage medium, characterized in that A computer program is stored thereon, and the computer program can be executed by a processor to implement the steps of the method according to any one of claims 1 to 3 and 5.

7. An electronic device, characterized in that: include: one or more processors; as well as a memory, wherein the memory is used to store executable instructions; The one or more processors are configured to implement the steps of the method of any one of claims 1 to 3 and 5 by executing the executable instructions.

Citation Information

Patent Citations

  • Image classification method for neural network architecture search based on evolutionary strategy

    CN114373101A

  • Weighted cascading convolutional neural networks

    US20190138888A1