Image classification model training method and device and image classification method and device

By predefined uniformly distributed class center points of the training samples of the image classification model, subsets with different classification difficulties are obtained and iterative training is carried out, the problem of improper fitting caused by the unoptimized training samples of the image classification model in the existing technology is solved, and more uniform sample distribution and more comprehensive feature learning are achieved, which improves the generalization ability of the model.

CN120198730APending Publication Date: 2025-06-24CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510280570.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-10
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The existing image classification model has not been optimized for training sample processing, which leads to inappropriate fitting problems during model training.

Method used

By obtaining the training sample set and multiple predefined uniformly distributed class center points, the training sample set is divided to obtain a subset of training samples with different classification difficulties, and these subsets are used to iteratively train the initial image classification model.

Benefits of technology

This method makes the sample distribution in the feature space more uniform, and the model can learn features more comprehensively, avoiding the problems of low learning efficiency and poor classification performance caused by uneven data distribution, and enhancing the model's ability to generalize data without seeing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120198730A_ABST
    Figure CN120198730A_ABST
Patent Text Reader

Abstract

The invention discloses an image classification model training method and device and an image classification method and device. The method comprises the steps that a training sample set is acquired, a plurality of predefined uniformly-distributed class center points are acquired, and the training sample set comprises a plurality of groups of training samples composed of a plurality of image samples and classification labels corresponding to the image samples; dividing the training sample set based on the plurality of predefined uniformly distributed class center points to obtain at least two training sample subsets with different classification difficulty; and performing iterative training on the initial image classification model by using the at least two training sample subsets to obtain a trained target image classification model. According to the method and the device, the technical problem that improper fitting is easy to occur during model training due to the fact that a related image classification model does not optimize a training sample is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of image recognition. Specifically, it relates to a method and device for training an image classification model, and a method and device for image classification. Background Art

[0002] Currently, most image classification models use model structures such as support vector machines, multi-layer perceptrons, decision trees, etc. for image classification, and these methods all obtain the best recognition performance by extracting the features of pattern samples and minimizing the intra-class distance and maximizing the inter-class distance in the feature space. Therefore, these methods do not pay attention to the regularity of the features themselves, so the potential of sample features remains to be explored. In addition, existing training samples are often unevenly distributed in the feature space, which may lead to some sample features being fully learned while some other sample features are ignored, thus affecting the overall performance of the classifier. For the above problems, no effective solution has been proposed yet. Summary of the Invention

[0003] Embodiments of the present application provide a method and device for training an image classification model, and a method and device for image classification, so as to at least solve the technical problem that the relevant image classification model does not optimize the training samples, resulting in improper fitting during model training.

[0004] According to one aspect of the embodiments of the present application, a method for training an image classification model is provided, including: obtaining a training sample set and obtaining a plurality of predefined uniformly distributed class center points, where the training sample set includes: multiple groups of training samples composed of multiple image samples and the classification label corresponding to each image sample; dividing the training sample set based on the plurality of predefined uniformly distributed class center points to obtain at least two training sample subsets with different classification difficulties; and iteratively training an initial image classification model using the at least two training sample subsets to obtain a target image classification model that has completed training.

[0005] Optionally, obtaining a plurality of predefined uniformly distributed class center points includes: determining the feature vector corresponding to each group of training samples in the training sample set, and forming a feature space from the feature vectors corresponding to multiple groups of training samples in the training sample set; determining the spatial dimension of the feature space, and determining a plurality of points uniformly distributed on a preset unit hypersphere based on the spatial dimension; and mapping the plurality of points back into the feature space to obtain a plurality of predefined uniformly distributed class center points.

[0006] Optionally, the training sample set is divided based on multiple predefined uniformly distributed class center points to obtain at least two training sample subsets with different classification difficulties, including: determining the distances between each feature vector in the feature space and each predefined uniformly distributed class center point; dividing the feature space according to the distances to obtain multiple clustering clusters, and determining the target class center point corresponding to each clustering cluster; dividing the multiple groups of training samples corresponding to the feature space based on the multiple target class center points to obtain at least two training sample subsets with different classification difficulties.

[0007] Optionally, dividing the multiple groups of training samples corresponding to the feature space based on the multiple target class center points, at least two training sample subsets with different classification difficulties, including: according to the multiple target class center points, using the following threshold function to determine the modulus value corresponding to each feature vector in the feature space, including:

[0008]

[0009] where f(h i ) represents the modulus value corresponding to the feature vector h i , μ ij represents the target class center point corresponding to the j-th clustering cluster to which the feature vector h i belongs, j ∈ [1, 2,..., M] represents the number of clustering clusters, i ∈ [1, 2,..., N] represents the number of feature vectors contained in the clustering cluster corresponding to the j-th target class center point, cosθ represents the cosine value of the angle between the feature vector h i and the target class center point corresponding to the j-th clustering cluster on the closed spherical surface, and the closed spherical surface is the closed space surrounded by the tangent planes where each target class center point is located; the first training sample subset is composed of the training samples corresponding to the feature vectors with modulus values not less than 1, and the second training sample subset is composed of the training samples corresponding to the feature vectors with modulus values less than 1, where the classification difficulty of the first training sample subset is lower than that of the second training sample subset.

[0010] Optionally, the initial image classification model is iteratively trained using at least two training sample subsets to obtain the trained target image classification model, including: preprocessing the first training sample subset and the second training sample subset respectively, where the preprocessing includes at least one of the following: dimensionality reduction processing, normalization processing; constructing an initial image classification model with a multi-level classifier tree architecture; batch inputting the groups of training samples in the preprocessed first training sample subset and second training sample subset into the initial image classification model to obtain the predicted classification labels corresponding to the image samples in each group of training samples output by the initial image classification model; constructing a target loss function based on the classification labels and the corresponding predicted classification labels in each group of training samples, and optimizing the target loss function until the preset convergence condition is satisfied, to obtain the trained target image classification model.

[0011] Optionally, an initial image classification model with a multi-level classifier tree architecture is constructed, including: determining the dimension of the feature space corresponding to the training sample set; when the number of dimensions is not less than a preset threshold value, determining that each level in the multi-level corresponding to the initial image classification model includes: at least one neural network-based classifier; when the number of dimensions is lower than the threshold value, determining that each level in the multi-level corresponding to the initial image classification model includes: at least one machine learning classifier.

[0012] Optionally, the neural network-based classifier includes at least one of the following: multi-layer perceptron, convolutional neural network, recurrent neural network; the machine learning classifier includes at least one of the following: support vector machine, decision tree, analytical learning classifier.

[0013] According to another aspect of the embodiments of the present application, an image classification method is further provided, including: obtaining a target image to be classified; analyzing the target image by using a pre-trained target image classification model to obtain a target classification label corresponding to the target image, where the target image classification model is trained by the above-mentioned image classification model training method.

[0014] According to another aspect of the embodiments of the present application, an image classification model training device is further provided, including: a first acquisition module, configured to acquire a training sample set and acquire a plurality of predefined uniformly distributed class center points, where the training sample set includes: multiple groups of training samples composed of multiple image samples and classification labels corresponding to each image sample; a division module, configured to divide the training sample set based on the multiple predefined uniformly distributed class center points to obtain at least two training sample subsets with different classification difficulties; a training module, configured to iteratively train an initial image classification model by using the at least two training sample subsets to obtain a trained target image classification model.

[0015] According to another aspect of the embodiments of the present application, an image classification device is further provided, including: a second acquisition module, configured to acquire a target image to be classified; a classification module, configured to analyze the target image by using a pre-trained target image classification model to obtain a target classification label corresponding to the target image, where the target image classification model is trained by the above-mentioned image classification model training method.

[0016] According to another aspect of the embodiments of the present application, a computer program product is further provided, including: a computer program, where when the computer program is executed by a processor, the above-mentioned image classification model training method is implemented.

[0017] In the embodiments of the present application, first, a training sample set is obtained, and a plurality of predefined uniformly distributed class centroids are obtained. Among them, the training sample set includes multiple groups of training samples composed of a plurality of image samples and the classification labels corresponding to each image sample. Then, based on the plurality of predefined uniformly distributed class centroids, the training sample set is divided to obtain at least two training sample subsets with different classification difficulties. This division method makes the sample distribution in the feature space more uniform, helps the model learn more comprehensive features, and avoids the problems of low learning efficiency and poor classification performance caused by uneven data distribution in the prior art. Finally, the initial image classification model is iteratively trained using at least two training sample subsets to obtain the trained target image classification model, enabling the model to obtain more balanced and comprehensive learning. Especially for samples with high classification difficulty, the model will invest more resources in learning, which helps to enhance the generalization ability of the model for unseen data. Furthermore, it solves the technical problem that the relevant image classification model does not optimize the training samples, resulting in improper fitting during model training. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:

[0019] Figure 1 is a schematic flowchart of an optional method for training an image classification model according to an embodiment of the present application;

[0020] Figure 2 is a schematic diagram of the division of an optional feature space according to an embodiment of the present application;

[0021] Figure 3 is a schematic diagram of the architecture of an optional multi-level classifier tree according to an embodiment of the present application;

[0022] Figure 4 is a schematic diagram of the structure of an optional analytical learning classifier according to an embodiment of the present application;

[0023] Figure 5 is a schematic diagram of the influence of noise on classification results according to an embodiment of the present application;

[0024] Figure 6 is a schematic diagram of the structure of an optional multi-layer perceptron according to an embodiment of the present application;

[0025] Figure 7 is a schematic diagram of the structure of an optional convolutional neural network according to an embodiment of the present application;

[0026] Figure 8It is a schematic flowchart of an optional image classification method according to an embodiment of the present application;

[0027] Figure 9 It is a schematic structural diagram of an optional image classification model training device according to an embodiment of the present application;

[0028] Figure 10 It is a schematic structural diagram of an optional image classification device according to an embodiment of the present application;

[0029] Figure 11 It is a schematic structural diagram of an optional electronic device according to an embodiment of the present application. Detailed implementation manners

[0030] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0031] It should be noted that the terms "first", "second", etc. in the specification, claims and drawings of the present application are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0032] To better understand the embodiments of the present application, some nouns or terms that appear in the description process of the embodiments of the present application are first translated and explained as follows:

[0033] Predefined Evenly-Distributed Class Centroids (PEDCC): The clustering centers of latent variables are artificially set by predefined optimal random clustering centers, and the distances between these clustering centers are far enough.

[0034] Predefined Optimal-Distribution Loss (POD Loss): It is a Softmax-free loss function based on a predefined optimal distribution, mainly used for classification tasks in deep learning. Its main idea is to optimize the classification performance by constraining the latent feature distribution of samples based on predefined evenly distributed class centers (PEDCC). Therefore, it does not rely on the Softmax layer but directly constrains the features in the intermediate layer.

[0035] Embodiment 1

[0036] According to an embodiment of the present application, a method for training an image classification model is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0037] Figure 1 is a flowchart diagram of a method for training an image classification model provided according to an embodiment of the present application. As Figure 1 shown, the method includes the following steps S102 - S106, where:

[0038] Step S102, obtain a training sample set and obtain multiple predefined evenly distributed class centers.

[0039] Among them, the above training sample set includes: multiple groups of training samples composed of multiple image samples and the classification label corresponding to each image sample.

[0040] In addition, the above predefined evenly distributed class centers (PEDCC, Predefined Evenly Distributed Class Centers) are a concept in pattern recognition and machine learning, mainly used for classification tasks, especially when designing and optimizing classifiers. These points are predefined class centers in the feature space, and they are designed to be evenly distributed in the high-dimensional space to ensure the best distribution of the centers of each category in the feature space, so as to promote the performance optimization of the classifier.

[0041] As an optional implementation manner, in the technical solution provided in the above step S102, multiple predefined evenly distributed class centers can be obtained by the following method, including:

[0042] The first step: Determine the feature vector corresponding to each group of training samples in the training sample set, and form a feature space from the feature vectors corresponding to multiple groups of training samples in the training sample set.

[0043] Specifically, the above steps can be carried out by preprocessing and feature encoding each group of training samples in the training sample set in sequence to obtain the feature vectors of each group of training samples, and the feature space is composed of the feature vectors corresponding to multiple groups of training samples in the training sample set. Among them, preprocessing includes but is not limited to missing value processing, data cleaning, standardization or normalization, etc., to ensure data quality and facilitate subsequent feature extraction and space construction; feature encoding is to convert non-numerical features into numerical forms for representation in the feature space, and common methods include but are not limited to one-hot encoding, embedding, label encoding, etc.

[0044] It should be noted that before feature encoding, feature selection and / or feature construction can also be carried out. Among them, feature selection is to select meaningful features from the original data, and these features can reflect the inherent attributes of the samples and contribute to the classification task. Common methods include methods such as correlation analysis and principal component analysis; while feature construction is to construct new features when the features in the original data are insufficient to describe the complexity of the samples. The common method is to create cross features or high-order features by combining or transforming existing features.

[0045] Step 2: Determine the spatial dimension of the feature space, and based on the spatial dimension, determine multiple predefined uniformly distributed class center points that are uniformly distributed on the preset unit hypersphere.

[0046] Specifically, in the process of determining the predefined uniformly distributed class center points above, a charge repulsion model can usually be used to obtain points with equal charges determined on the unit hypersphere corresponding to the spatial dimension. These points will repel each other until they reach an equilibrium state, in which the distance between the points is maximized and the distribution is uniform. In addition, the process of determining the points above can also be regarded as a Thomson problem, that is, the problem of uniformly distributing multiple predefined uniformly distributed class center points on the unit sphere with the least energy (defined as the sum of the reciprocals of the distances between all pairs of points). This problem can be approximately solved by a simulated annealing algorithm, a charge repulsion algorithm or other heuristic algorithms.

[0047] In addition to the two implementation schemes listed above, based on the basic concept of the present invention, those skilled in the art can also determine multiple predefined uniformly distributed class center points that are uniformly distributed on the unit hypersphere through other technical solutions. For example, those skilled in the art transform the above implementation schemes, and they should also be within the protection scope of the present invention.

[0048] Step 3: Map the multiple points back into the feature space to obtain multiple predefined uniformly distributed class center points.

[0049] That is to say, after determining the points evenly distributed on the unit hypersphere, by statistically analyzing the training sample set, the average value of the feature vectors of each image category is determined, and then multiple points evenly distributed on the unit hypersphere are mapped to these average values through appropriate linear transformations to obtain multiple pre-defined evenly distributed class center points.

[0050] Step S104: Based on multiple pre-defined evenly distributed class center points, divide the training sample set to obtain at least two training sample subsets with different classification difficulties.

[0051] In the technical solution provided in the above step S104, it is considered that training samples are often unevenly distributed in the feature space, which may lead to some sample features being fully learned while some other sample features are ignored, thus affecting the overall performance of the classifier. At the same time, existing classifiers, especially deep learning models, are prone to overfitting on the training set, that is, they perform well on the training data but have poor generalization ability on unseen test data. Therefore, the embodiments of this application propose to divide the training sample set through evenly distributed class centers. On the one hand, it solves the problem of non-uniform data distribution and ensures that all sample features can be fully considered and optimized. On the other hand, it can train the model on subsets with different difficulties, avoiding overlearning of simple samples and ensuring sufficient training of difficult samples, thus balancing the complexity and generalization ability of the model.

[0052] As an optional implementation manner, in the technical solution provided in the above step S104, the method may include:

[0053] Step S1041: Determine the distances between each feature vector in the feature space and each pre-defined evenly distributed class center point. Among them, the above distances can be measured by cosine similarity or Euclidean distance.

[0054] Step S1042: Divide the feature space according to the distances to obtain multiple clustering clusters, and determine the target class center point corresponding to each clustering cluster.

[0055] Specifically, use a clustering algorithm to group the feature vectors to form multiple clustering clusters. Among them, the samples within each clustering cluster have similar features and the shortest distance to one or more specific clustering center points. The target class center point is determined by calculating the average value of all sample feature vectors within each clustering cluster.

[0056] Step S1043: Divide the multiple groups of training samples corresponding to the feature space according to multiple target class center points to obtain at least two training sample subsets with different classification difficulties.

[0057] Specifically, in the technical solution provided in step S1043, multiple target class center points determined in the above step S1042 are used to divide multiple groups of training samples in the feature space again, so as to identify at least two training sample subsets with different classification difficulties. Optionally, the specific implementation manner of this step includes the following steps:

[0058] The first step: According to multiple target class center points, use the following threshold function to determine the modulus value corresponding to each feature vector in the feature space, including:

[0059]

[0060] In the formula, f(h i ) represents the modulus value corresponding to the feature vector h i , μ ij represents the target class center point corresponding to the j-th clustering cluster to which the feature vector h i belongs, j ∈ [1, 2,..., M] represents the number of clustering clusters, i ∈ [1, 2,..., N] represents the number of feature vectors included in the clustering cluster corresponding to the j-th target class center point, cosθ represents the cosine value of the angle between the feature vector h i and the target class center point corresponding to the j-th clustering cluster on the closed spherical surface, and the closed spherical surface is the closed space surrounded by the tangent planes where each target class center point is located.

[0061] The second step: The first training sample subset is composed of training samples corresponding to feature vectors with a modulus value not less than 1, and the second training sample subset is composed of training samples corresponding to feature vectors with a modulus value less than 1.

[0062] Among them, a modulus value not less than 1 indicates that the projection length of the training sample in the direction of the target class center is greater than the length from the target class center to the origin, that is, the position of the training sample in the feature space is relatively far from the origin. For such training samples, the classifier can easily classify them correctly based on their features, so they can be regarded as easily distinguishable samples; on the contrary, a modulus value less than 1 indicates that the position of the training sample in the feature space is relatively far from the origin. For such training samples, the classifier is difficult to classify them correctly based on their features. Therefore, the classification difficulty of the first training sample subset is lower than that of the second training sample subset. Therefore, through the division method based on the modulus value, the training sample set can be divided into two parts with different recognition difficulties, so that the model can better learn the laws of the features of each subset through training sample subsets with different recognition difficulties, and further improve the performance of the classifier.

[0063] Taking the IRIS data set containing 150 training samples and a total of three types of image categories as an example, the division result of its training samples is as Figure 2 shown. Figure 2Each point in is the sample supervision signal (i.e., classification label) of each training sample in the IRIS dataset, and the sample supervision signal of each training sample is regarded as a point in the feature space corresponding to the IRIS dataset. The class center points of each class are regarded as predefined uniformly distributed class center points distributed on the unit hypersphere, and tangent planes are made along these three class center points, so that a closed space can be obtained. Points in this space are regarded as points with a modulus value less than 1; and points outside this space are regarded as points with a modulus value not less than 1.

[0064] For example, there are six commonly used training sample sets, namely: Mnist dataset (consisting of 60,000 training samples and 10,000 test samples, with a total of 10 categories of pictures, and each picture is a 28*28 grayscale handwritten font picture), FashionMnist dataset (consisting of grayscale pictures of clothes, shoes and other clothing), Didits dataset (consisting of 1,797 samples, a total of 10 categories, and each picture is 8*8 grayscale handwritten font pictures), Pima Indian Diabetes dataset (consisting of 768 samples, a total of two categories, each containing 8 features), Waveform dataset (a total of 5,000 samples, a total of three categories, each containing 21 features), Iris dataset (consisting of 150 samples, a total of three categories, each containing 4 features). The above training sample sets are divided by using the threshold function, and the divided training sample subsets are connected to the analytical learning classifier. The recognition rate of the analytical learning classifier is tested on the original training sample set, the training sample subset with high recognition difficulty and the training sample subset with low training difficulty, and the comparison results are shown in Table 1 below.

[0065] Table 1

[0066]

[0067] It can be seen from Table 1 above that the recognition rate of the training sample subset with low recognition difficulty is significantly higher than the recognition rate of the training sample subset with high recognition difficulty.

[0068] Step S106, iteratively training the initial image classification model using at least two training sample subsets to obtain a trained target image classification model.

[0069] In the technical solution provided in the above step S106, since each training sample subset contains samples with similar classification difficulty, at least two training sample subsets divided based on classification difficulty can be used in different stages of the training model to optimize the model parameters in a targeted manner to make it more suitable for the feature distribution of the subset. This targeted training enables the model to gradually learn features from simple to complex, improving the overall training efficiency and classification performance.

[0070] As an alternative implementation, in the technical solution provided in step S106 above, the method may include:

[0071] Step S1061, preprocess the first training sample subset and the second training sample subset respectively.

[0072] Among them, the preprocessing includes but is not limited to: dimensionality reduction processing, normalization processing, etc. For example, the dimensionality reduction processing of the training sample subset is performed by the Principal Component Analysis (PCA) algorithm to maximize the retention of information in the data while reducing the dimensions and eliminating the correlation between features. After the dimensionality reduction processing, the principal components are further processed by whitening processing, so that they are statistically independent and have equal variances (that is, the data is converted into a Gaussian distribution with a mean of 0 and a variance of 1 for each dimension, similar to the characteristics of white noise). That is, the whitening processing normalizes the data after PCA conversion to ensure that the variance of each principal component is 1, thereby eliminating the influence of the dimension and making all features have similar importance, which helps to simplify the subsequent model training and improve its stability and efficiency.

[0073] Step S1062, construct an initial image classification model with a multi-level classifier tree architecture.

[0074] Among them, the multi-level classifier tree structure is similar to the decision tree in terms of structural form, as Figure 3 shown. However, the principles and structures are completely different, and the specific differences are as follows:

[0075] (1) The division of the multi-level classifier tree structure depends on the distribution of training samples in the feature space at each layer. The distribution of training samples in the feature space is optimized at each layer, thus bringing an improvement in the recognition rate; while the division methods such as decision trees depend on the decision discrimination of the original features from top to bottom when constructing the tree structure and cannot optimize the features.

[0076] (2) In the multi-level classifier tree structure, each node is an independent classifier. The input is the data set determined and divided by the upper-level classifier based on the features of this layer. While outputting the classification result of this layer, the data set is divided and passed to the next-level classifier; while each node in the decision tree is a judgment criterion, and the entire decision tree formed is a classifier.

[0077] (3) The last layer in the multi-level classifier tree structure is the output node, and the training purpose of the remaining nodes is to divide the data set and optimize the feature distribution.

[0078] (4) The multi-level classifier tree can control the number of levels through hyperparameters and can stop training when the data of a certain level is too small to be divided or overfitting occurs.

[0079] Step S1063: Batch input each group of training samples in the preprocessed first training sample subset and second training sample subset into the initial image classification model, and obtain the predicted classification labels corresponding to the image samples in each group of training samples output by the initial image classification model.

[0080] Step S1064: Construct an objective loss function based on the classification labels and the corresponding predicted classification labels in each group of training samples, and optimize the objective loss function until the preset convergence condition is met, thereby obtaining the trained target image classification model.

[0081] In the above embodiments, when using two training sample subsets with different classification difficulties for model training, the training of the model on the two subsets can be carried out alternately, or start from the first training sample subset with lower difficulty first, and then transition to the second training sample subset with higher difficulty. Through multiple iterations, the model can be gradually optimized until satisfactory classification performance is achieved on both subsets.

[0082] For example, taking the case of starting from the first training sample subset with lower difficulty first and then transitioning to the second training sample subset with higher difficulty as an example, the specific iteration process is as follows:

[0083] First, batch input each group of training samples in the preprocessed first training sample subset into the initial image classification model, and obtain the predicted classification labels corresponding to the image samples in each group of training samples output by the initial image classification model;

[0084] Next, construct an objective loss function based on the classification labels and the corresponding predicted classification labels in each group of training samples in the first training sample subset, and calculate the objective loss function through the backpropagation algorithm;

[0085] Then, determine whether the objective loss function meets the preset convergence condition, where the convergence condition includes: the number of iterations reaches the maximum number of iterations, and the gradient of the loss function approaches zero;

[0086] When the objective loss function meets this convergence condition, determine the target image classification model;

[0087] When the objective loss function does not meet this convergence condition, batch input each group of training samples in the preprocessed second training sample subset into the initial image classification model, and obtain the predicted classification labels corresponding to the image samples in each group of training samples output by the initial image classification model; update the objective loss function based on the classification labels and the corresponding predicted classification labels in each group of training samples in the second training sample subset until the objective loss function meets this convergence condition, thereby obtaining the target image classification model.

[0088] Optionally, in the technical solution provided in step S1062, the type of the classifier in each layer in the initial image classification model can be determined by the following method:

[0089] Determine the number of dimensions of the feature space corresponding to the training sample set;

[0090] In the case where the number of dimensions is not less than a preset threshold value, determining that each level in the multiple levels corresponding to the initial image classification model includes: at least one classifier based on a neural network;

[0091] When the number of dimensions is lower than a threshold value, it is determined that each level within the multiple levels corresponding to the initial image classification model includes: at least one machine learning classifier.

[0092] That is to say, if the number of dimensions is small, the classifiers at each level in the initial image classification model can choose simpler classifiers, such as simple neural network models such as decision trees, support vector machines (SVM), and analytical learning classifiers (ALC). These models usually have fewer parameters and do not require a large amount of data, so they usually have higher efficiency and accuracy when processing small category sets. On the contrary, when the number of dimensions is large, especially when processing complex image classification tasks, the classifiers at each level in the initial image classification model can choose more complex classifiers, such as multilayer perceptrons (MLP), convolutional neural networks (CNN), recursive neural networks, etc., to capture complex features in the image and subtle differences between categories.

[0093] It should be noted that the type of classifier in each layer of the initial image classification model can also be adjusted according to the classification difficulty of the training sample set. For example, a deeper neural network or a more complex classifier tree is designed for a subset of training samples with high classification difficulty; while for a subset with low classification difficulty, a shallower model or fewer classifier layers is used to save computing resources and speed up the training process.

[0094] The following sets of experimental data will specifically illustrate the impact of training sample sets with different dimensionalities on the type of classifiers in the multi-level classifier tree.

[0095] Experiment 1: The parsing learning classifier is suitable for analyzing training sample sets with a small number of dimensions.

[0096] There are six commonly used training sample sets, namely: Mnist dataset (consisting of 60,000 training samples and 10,000 test samples, with a total of 10 categories of pictures, and each picture is a 28*28 grayscale handwritten font picture), FashionMnist dataset (composed of grayscale pictures of clothing such as clothes and shoes), Didits dataset (consisting of 1,797 samples, with a total of 10 categories, and each picture is an 8*8 grayscale handwritten font picture), Pima Indian Diabetes dataset (consisting of 768 samples, with a total of two categories, and each category contains 8 features), Waveform dataset (with a total of 5,000 samples, with a total of three categories, and each category contains 21 features), Iris dataset (consisting of 150 samples, with a total of three categories, and each category contains 4 features). Using the threshold function to divide each of the above datasets, and inputting the divided subsets into the image classification model with a single-level classifier tree architecture constructed by the parsing classification learner as shown in Figure 4 shown, the image classification model with a multi-level classifier tree architecture constructed by the parsing classification learner as shown in Figure 4 shown, support vector machine, decision tree, and obtain the recognition rate comparison results in Table 2 below.

[0097] Table 2

[0098]

[0099]

[0100] It can be seen from Table 2 above that the recognition rate of the multi-level parsing learning classifier has been significantly improved compared with that of the single-level parsing learning classifier. Except for the slightly lower recognition rate on the Mnist dataset, the recognition rate on the large sample dataset is comparable to that of the support vector machine, while it has an obvious advantage on the small sample dataset. Therefore, the parsing learning classifier is suitable for model training with fewer classification labels.

[0101] It should be noted that during the model classification process, noise often interferes with the classification results. In the multi-level classifier tree structure, due to the continuous deepening of the structure, it is necessary to consider whether the influence of noise will also become larger and larger. Based on this, the embodiments of this application conduct a noise experiment. In the experiment, Gaussian noise, that is, the features whose probability density function conforms to the Gaussian distribution, are added to the image classification model with a multi-level parsing classifier tree structure, and the influence of noise on the recognition rate at different signal-to-noise ratios is compared, and the experimental results as shown in Figure 5 shown are obtained. It can be seen from Figure 5 that the influence of low signal-to-noise ratio on the recognition rate is relatively large. When the signal-to-noise ratio is greater than 10 dB, the influence on the recognition rate is relatively small. Therefore, in some cases, the recognition rate can even be improved.

[0102] Experiment 2: The multi-layer perceptron is suitable for analyzing training sample sets with a relatively large number of dimensions.

[0103] The above-mentioned Mnist dataset and FashionMnist dataset are divided using a threshold function, and each of the divided subsets is respectively input into an image classification model with a single-level classifier tree architecture constructed by a multi-layer perceptron as shown in Figure 6 , an image classification model with a multi-level classifier tree architecture constructed by a multi-layer perceptron as shown in Figure 6 , and a support vector machine, and the recognition rate comparison results in Table 3 below are obtained. Among them, Figure 6 The multi-layer perceptron shown in includes two hidden layers, and the ReLU activation function is used between each layer. BatchNorm is used to maintain the gradient stability and accelerate convergence. The output features of the second hidden layer are first used for the backpropagation of POD-Loss, and then the weights are fixed by PEDCC and classified and output using the cosine distance. The optimizer of the network uses stochastic gradient, and the learning rate optimization strategy is the adaptive learning rate.

[0104] Table 3

[0105]

[0106]

[0107] It is not difficult to see from Table 3 above that for grayscale large datasets such as Mnist and FashionMnist, the recognition performance of the multi-level multi-layer perceptron is further optimized compared to the single-level multi-layer perceptron, and it has obvious advantages over the support vector machine.

[0108] In addition, the activation function is an important part that affects the recognition performance of the multi-layer perceptron. It can increase the non-linear expression ability of features and enable the network to learn sample features more fully. Therefore, a suitable activation function should be selected for different networks. Table 4 below shows the recognition rate comparison results of multi-layer multi-level perceptrons using different activation functions on the FashionMnist dataset and the Cifar10 dataset.

[0109] Table 4

[0110] Activation function FashionMnist dataset Cifar10 dataset Sigmoid activation function 89.53% 97.71% Tanh activation function 90.36% 87.93% ReLU activation function 91.39% 89.97%

[0111] It can be seen from Table 4 above that the recognition rates of the Sigmod activation function and the Tanh activation function are relatively low. Therefore, to ensure the unity of the overall structure of the classifier in the multi-level neural network, the same network structure and parameter settings are used for each layer. Therefore, ReLU is used as the hidden layer activation function for each layer.

[0112] Experiment 3: The convolutional neural network is also suitable for analyzing training sample sets with a relatively large number of dimensions.

[0113] The above-mentioned Mnist dataset and FashionMnist dataset are partitioned using a threshold function, and each of the partitioned subsets is respectively input into an image classification model with a single-level classifier tree architecture constructed by a convolutional neural network as shown in Figure 7 and an image classification model with a multi-level classifier tree architecture constructed by a convolutional neural network as shown in Figure 7 to obtain the recognition rate comparison results in Table 5 below. Among them, Figure 7 the convolutional neural network shown in uses four layers of convolution, and the size of each convolution kernel is 3×3; the pooling method selects max pooling, that is, the maximum value in the local area is selected as the output value, so as to retain the most significant features. At the same time, max pooling can learn the edge and texture features of the image; the network optimizer selects stochastic gradient descent, and an adaptive learning rate is used to ensure the optimal recognition performance; the fully connected layer uses two layers of fully connected layers, and the ReLU activation function is used as the activation function.

[0114] Table 5

[0115]

[0116]

[0117] It can be easily seen from Table 5 above that the convolutional neural network can extract features more effectively for recognition, and for grayscale large datasets such as Mnist and FashionMnist, the recognition performance of the multi-level convolutional neural network is further optimized compared to the single-level convolutional neural network.

[0118] In addition, during the construction and recognition process of the convolutional neural network, backpropagation uses the POD Loss function based on predefined uniform distribution class centers, and backpropagation involves the pooling layer and the convolutional layer, with high computational complexity. Therefore, the hidden feature dimension has a greater impact on the recognition rate. Therefore, the multi-layer convolutional neural network under different hidden feature dimensions has different effects on the recognition rate on different datasets, and the specific influence results are shown in Table 6 below.

[0119] Table 6

[0120] Hidden feature dimension FashionMnist dataset Cifar10 dataset 9 89.69% 92.35% 32 89.93% 92.78% 64 91.69% 93.19% 128 91.46% 93.42% 256 91.22% 93.34% 512 91.19% 93.27%

[0121] It can be easily seen from Table 6 above that within a certain range, the recognition rate will increase with the increase of the hidden feature dimension, but there is an optimal hidden feature dimension, that is, when the dimension continues to increase, the recognition performance will decline. Therefore, the optimal hidden feature dimension for each dataset can be determined through experiments. At the same time, for each node classifier in the multi-level classifier tree structure, the selected optimal hidden feature dimension should be the same as that of the first level to ensure the unity and stability of the overall structure.

[0122] Therefore, based on the solution defined in the above steps S102 to S106, it can be known that compared with the existing model training solution, the solution of the present application has the following technical advantages:

[0123] (1) Optimize the division of the feature space based on the threshold function of the predefined uniform distribution class center points, so as to divide the training sample set into subsets with different classification difficulties. This division method makes the sample distribution in the feature space more uniform, allowing the model to quickly learn basic features on the subsets that are easy to classify, and conduct more in-depth and detailed learning on the subsets with higher classification difficulties, greatly improving the pertinence and efficiency of training, and effectively avoiding overfitting and underfitting problems. Among them, overfitting usually occurs when the model over-learns simple training data, while underfitting is because the model lacks sufficient learning on complex data.

[0124] (2) After the training of each subset with different classification difficulties is completed, the model can obtain a more balanced and comprehensive learning. Especially for samples with high classification difficulties, the model will invest more resources in learning, which helps to enhance the generalization ability of the model for unseen data.

[0125] (3) The image classification model with a multi-level classifier tree structure can be flexibly adjusted according to the characteristics of the data set, and the classifier of each node can be independently optimized, so that the most suitable model structure can be adopted for data sets of different sizes and numbers of categories, improving the adaptability and flexibility of the model.

[0126] Embodiment 2

[0127] According to the embodiment of the present application, an image classification method is also provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0128] Figure 8 is a schematic flowchart of an image classification model training method provided according to the embodiment of the present application. As Figure 8 shown, the method includes the following steps S802 - S804, where:

[0129] Step S802, obtain the target image to be classified.

[0130] Step S804, use the pre-trained target image classification model to analyze the target image to obtain the target classification label corresponding to the target image.

[0131] Among them, the above-mentioned target image classification model is trained by the image classification model training method in Embodiment 1, so the training process of the target image classification model will not be elaborated in the embodiments of this application.

[0132] Embodiment 3

[0133] According to the embodiments of this application, there is also provided an image classification model training device for implementing the image classification model training method in Embodiment 1. As Figure 9 shown, the image classification model training device at least includes: a first acquisition module 92, a division module 94, and a training module 96, where:

[0134] The first acquisition module 92 is configured to acquire a training sample set and acquire a plurality of predefined uniformly distributed class center points. Among them, the training sample set includes: multiple groups of training samples composed of multiple image samples and the classification label corresponding to each image sample;

[0135] The division module 94 is configured to divide the training sample set based on the multiple predefined uniformly distributed class center points to obtain at least two training sample subsets with different classification difficulties;

[0136] The training module 96 is configured to perform iterative training on the initial image classification model by using at least two training sample subsets to obtain a target image classification model that has completed training.

[0137] It should be noted that each module in the image classification model training device in the embodiments of this application corresponds one by one to each implementation step of the image classification model training method in Embodiment 1. Since Embodiment 1 has been described in detail, some details not shown in this embodiment can be referred to Embodiment 1 and will not be elaborated here.

[0138] Embodiment 4

[0139] According to the embodiments of this application, there is also provided an image classification device for implementing the image classification method in Embodiment 2. As Figure 10 shown, the image classification device at least includes: a second acquisition module 12 and a classification module 14, where:

[0140] The second acquisition module 12 is configured to acquire a target image to be classified;

[0141] The classification module 14 is configured to analyze the target image by using the pre-trained target image classification model to obtain the target classification label corresponding to the target image.

[0142] Among them, the above-mentioned target image classification model is trained by the image classification model training method in Embodiment 1, so the training process of the target image classification model will not be elaborated in the embodiments of this application.

[0143] Example 5

[0144] According to an embodiment of the present application, there is also provided a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the image classification model training method in Embodiment 1 or the image classification method in Embodiment 2.

[0145] According to an embodiment of the present application, there is also provided a non-volatile storage medium, which includes a stored computer program. When the device where the non-volatile storage medium is located runs the computer program, it executes the image classification model training method in Embodiment 1 or the image classification method in Embodiment 2.

[0146] According to an embodiment of the present application, there is also provided a processor, which is used to run a computer program. When the computer program runs, it executes the image classification model training method in Embodiment 1 or the image classification method in Embodiment 2.

[0147] According to an embodiment of the present application, there is also provided an electronic device, which includes: a memory and a processor. The memory stores a computer program, and the processor is configured to execute the image classification model training method in Embodiment 1 or the image classification method in Embodiment 2 through the computer program.

[0148] Optionally, when the computer program runs, it executes the following steps: obtaining a training sample set and obtaining a plurality of predefined uniformly distributed class center points, where the training sample set includes: multiple groups of training samples composed of multiple image samples and the classification label corresponding to each image sample; dividing the training sample set based on the multiple predefined uniformly distributed class center points to obtain at least two training sample subsets with different classification difficulties; using the at least two training sample subsets to iteratively train an initial image classification model to obtain a trained target image classification model.

[0149] Optionally, when the computer program runs, it executes the following steps: obtaining a target image to be classified; using a pre-trained target image classification model to analyze the target image to obtain a target classification label corresponding to the target image, where the target image classification model is trained by the image classification model training method provided in Embodiment 1.

[0150] As an optional implementation manner, the above electronic device may exist in the form of a mobile terminal, a computer terminal, or a similar computing device. Figure 11 Shows a hardware structure block diagram of an electronic device for implementing the image classification model training method. As Figure 11As shown, the electronic device 110 may include one or more processors 1102 (illustrated as 1102a, 1102b, ……, 1102n in the figure) (the processor 1102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 1104 for storing data, and a transmission device 1106 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 11 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the electronic device 110 may further include more or fewer components than Figure 11 shown in, or have a different configuration from Figure 11 that shown.

[0151] It should be noted that the above one or more processors 1102 and / or other data processing circuits may generally be referred to as "data processing circuits" herein. The data processing circuit may be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the electronic device 110. As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistance terminal path connected to an interface).

[0152] The memory 1104 may be used to store software programs and modules of application software, such as the program instructions / data storage device corresponding to the image classification model training method in the embodiments of the present application. The processor 1102 executes various functional applications and data processing by running the software programs and modules stored in the memory 1104, that is, implements the vulnerability detection method of the above-mentioned application program. The memory 1104 may include a high-speed random access memory, and may further include a non-volatile memory, such as one or more magnetic storage devices, a flash memory, or other non-volatile solid-state memories. In some instances, the memory 1104 may further include a memory remotely set relative to the processor 1102, and these remote memories may be connected to the electronic device 110 through a network. Examples of the above network include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0153] The transmission device 1106 is used to receive or send data via a network. Specific examples of the above-mentioned network may include a wireless network provided by a communication provider of the electronic device 110. In one example, the transmission device 1106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station so as to communicate with the Internet. In one example, the transmission device 1106 can be a Radio Frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0154] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables a user to interact with the user interface of the electronic device 110.

[0155] The above-mentioned serial numbers of the embodiments are only for description and do not represent the advantages or disadvantages of the embodiments.

[0156] In the above embodiments of the present application, the descriptions of the various embodiments have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0157] In several embodiments provided in the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0158] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0159] In addition, in each embodiment of the present application, the various functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.

[0160] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.

[0161] The above are only the preferred embodiments of this application. It should be noted that for those of ordinary skill in the art of this technology, without departing from the principle of this application, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of this application.

Claims

1. A method for training an image classification model, characterized in that: include: Obtaining a training sample set and obtaining a plurality of predefined uniformly distributed class center points, wherein the training sample set includes: a plurality of groups of training samples consisting of a plurality of image samples and a classification label corresponding to each of the image samples; Dividing the training sample set based on the multiple predefined uniformly distributed class center points to obtain at least two training sample subsets with different classification difficulties; The initial image classification model is iteratively trained using the at least two training sample subsets to obtain a trained target image classification model.

2. The method according to claim 1, characterized in that Get multiple predefined uniformly distributed center points, including: Determine a feature vector corresponding to each group of training samples in the training sample set, and form a feature space by the feature vectors corresponding to multiple groups of training samples in the training sample set; Determining a spatial dimension of the feature space, and determining a plurality of points uniformly distributed on a preset unit hypersphere based on the spatial dimension; The multiple points are mapped back into the feature space to obtain the multiple predefined uniformly distributed class center points.

3. The method according to claim 2, characterized in that The training sample set is divided based on the multiple predefined uniform distribution class center points to obtain at least two training sample subsets with different classification difficulties, including: Determine the distance between each feature vector in the feature space and each of the predefined uniformly distributed class center points; Dividing the feature space according to the distance to obtain a plurality of clusters, and determining the center point of the target class corresponding to each cluster; The plurality of groups of training samples corresponding to the feature space are divided according to the plurality of target class center points to obtain at least two training sample subsets with different classification difficulties.

4. The method according to claim 3, characterized in that Dividing the multiple groups of training samples corresponding to the feature space according to the multiple target class center points to obtain at least two training sample subsets with different classification difficulties, including: According to the plurality of target class center points, the modulus corresponding to each feature vector in the feature space is determined using the following threshold function, including: Where f(h i ) represents the feature vector h i The corresponding modulus value, μ ij Denotes the feature vector h i The target class center point corresponding to the jth cluster belongs to, j∈[1,2,…,M] represents the number of clusters, i∈[1,2,…,N] represents the number of feature vectors contained in the cluster corresponding to the jth target class center point, and cosθ represents the feature vector h i The cosine value of the angle between the target class center point corresponding to the j-th cluster and the closed sphere, and the closed sphere is a closed space surrounded by the tangent planes where the target class center points are located; A first training sample subset is composed of training samples corresponding to feature vectors whose modulus values ​​are not less than 1, and a second training sample subset is composed of training samples corresponding to feature vectors whose modulus values ​​are less than 1, wherein the classification difficulty of the first training sample subset is lower than that of the second training sample subset.

5. The method according to claim 4, characterized in that Iteratively training the initial image classification model using the at least two training sample subsets to obtain a trained target image classification model, including: Preprocessing the first training sample subset and the second training sample subset respectively, wherein the preprocessing includes at least one of the following: dimensionality reduction processing and standardization processing; Constructing the initial image classification model of a multi-level classifier tree architecture; Inputting the preprocessed training samples in the first training sample subset and the second training sample subset into the initial image classification model in batches, and obtaining the predicted classification labels corresponding to the image samples in each group of the training samples output by the initial image classification model; A target loss function is constructed based on the classification labels in each group of training samples and the corresponding predicted classification labels, and the target loss function is optimized until a preset convergence condition is met to obtain a trained target image classification model.

6. The method according to claim 5, characterized in that Constructing the initial image classification model of a multi-level classifier tree architecture, comprising: Determining the dimension of the feature space corresponding to the training sample set; In the case where the number of dimensions is not less than a preset threshold value, determining that each level in the multiple levels corresponding to the initial image classification model includes: at least one classifier based on a neural network; When the number of dimensions is lower than the threshold value, it is determined that each level within the multiple levels corresponding to the initial image classification model includes: at least one machine learning classifier.

7. The method according to claim 6, characterized in that The neural network-based classifier includes at least one of the following: a multi-layer perceptron, a convolutional neural network, and a recursive neural network; The machine learning classifier includes at least one of the following: a support vector machine, a decision tree, and an analytical learning classifier.

8. An image classification method, characterized in that: include: Obtain a target image to be classified; The target image is analyzed using a pre-trained target image classification model to obtain a target classification label corresponding to the target image, wherein the target image classification model is trained using the image classification model training method described in any one of claims 1 to 7.

9. An image classification model training device, characterized in that: include: A first acquisition module is used to acquire a training sample set and acquire a plurality of predefined uniformly distributed class center points, wherein the training sample set includes: a plurality of groups of training samples consisting of a plurality of image samples and a classification label corresponding to each of the image samples; A division module, used for dividing the training sample set based on a plurality of predefined uniformly distributed class center points to obtain at least two training sample subsets with different classification difficulties; The training module is used to iteratively train the initial image classification model using the at least two training sample subsets to obtain a trained target image classification model.

10. An image classification method, characterized in that: include: A second acquisition module is used to acquire a target image to be classified; A classification module is used to analyze the target image using a pre-trained target image classification model to obtain a target classification label corresponding to the target image, wherein the target image classification model is trained by the image classification model training method described in any one of claims 1 to 7.

11. A computer program product, characterized in that include: A computer program, wherein when the computer program is executed by a processor, it implements the image classification model training method described in any one of claims 1 to 7 or the image classification method described in claim 8.

12. An electronic device, characterized in that: include: A memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the image classification model training method described in any one of claims 1 to 7 or the image classification method described in claim 8 through the computer program.