Image classification method, training method of image classification model and related equipment
By using dynamic correction technology in the image classification method, the target base class sample is selected from the initial base class sample based on the similarity matrix, and the initial new class sample is corrected, which solves the problem of feature distribution deviation in learning in few samples and improves the accuracy of the image classification model.
Patent Information
- Application Number
- CN202510020560.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-06
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-06
AI Technical Summary
The prior art ignores the real feature distribution of the target category group represented by the few samples in the learning image classification of few samples, resulting in a deviation from the generated samples and the few samples, which in turn causes poor image classification accuracy of the trained models.
An image classification method is proposed. By obtaining the initial sample, including the initial new class sample and multiple initial base class samples, it is input to the initial image classification model, and the similarity matrix is determined based on the preset composite metric function, the target base class sample is selected from multiple initial base class samples, dynamically corrects the initial new class sample, obtain the target new class sample, and train the initial image classification model using the target new class sample.
By accurately capturing more similar target base class samples, the problem of scarcity of effective training features in the initial new class samples is improved, and the authenticity and accuracy of feature distribution is improved, thereby improving the image classification accuracy of the trained model.
Smart Images

Figure CN120070944A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and particularly relates to an image classification method, a training method for an image classification model, and related devices. Background Art
[0002] Image classification refers to assigning an input image to one of a predefined set of classes. Among them, few-shot learning image classification is a common image classification task, and its purpose is to train a model to achieve efficient task execution with only a small number of labeled samples.
[0003] Related technologies have focused on sample generation, generating more samples to help the model access more sample data during training. However, such an image classification method ignores the true feature distribution of the target class group represented by the few samples, resulting in a deviation between the generated samples and the few samples, and further causing poor image classification accuracy of the trained model. Summary of the Invention
[0004] The main purpose of the embodiments of the present application is to propose an image classification method, a training method for an image classification model, and related devices, aiming to improve the accuracy of image classification.
[0005] To achieve the above object, a first aspect of the embodiments of the present application proposes an image classification method, including:
[0006] Obtain initial samples, where the initial samples include initial new-class samples and a plurality of initial base-class samples;
[0007] Input the initial samples into an initial image classification model, and determine a corresponding similarity matrix of the initial samples based on a preset composite metric function, where the similarity matrix is used to characterize the feature matching degree between the initial new-class samples and each initial base-class sample in multiple dimensions;
[0008] Select a target base-class sample from the plurality of initial base-class samples based on the similarity matrix;
[0009] Dynamically correct the initial new-class samples according to the target base-class sample and the corresponding similarity matrix of the target base-class sample to obtain target new-class samples;
[0010] Use the target new-class samples to train the initial image classification model until a training stop condition is reached to obtain a trained image classification model;
[0011] Obtain a target image to be classified, input the target image into the trained image classification model, and obtain a corresponding target classification result.
[0012] In some embodiments, the initial new class samples include a plurality of initial new class features, and each initial base class sample includes a plurality of initial base class features;
[0013] Determining a similarity matrix based on a composite metric function includes:
[0014] Determining a new class feature matrix of the initial new class samples and a base class feature matrix of the initial base class samples respectively;
[0015] Based on the plurality of initial new class features and the plurality of initial base class features, determining a first similarity value of the initial new class sample and each initial base class sample in the first dimension, and determining a second similarity value of the initial new class sample and each initial base class sample in the second dimension;
[0016] Based on the new class feature matrix and the base class feature matrix, determining a third similarity value of the initial new class sample and each initial base class sample in the third dimension;
[0017] According to the composite metric function set with preset hyperparameters, fusing the first similarity value, the second similarity value and the third similarity value to obtain a similarity matrix.
[0018] In some embodiments, determining a new class feature matrix of the initial new class samples and a base class feature matrix of the initial base class samples respectively includes:
[0019] Performing a transpose process on each initial new class feature to obtain a corresponding new class transposed feature;
[0020] Based on the cross product of each initial new class feature and the corresponding new class transposed feature, obtaining a new class feature matrix;
[0021] Based on the plurality of initial base class features, obtaining the corresponding initial base class means of each initial base class sample;
[0022] Performing a transpose process on each initial base class mean to obtain a corresponding base class transposed mean;
[0023] Based on the cross product of each initial base class mean and the corresponding base class transposed mean, obtaining a base class feature matrix.
[0024] In some embodiments, based on the plurality of initial new class features and the plurality of initial base class features, determining a first similarity value of the initial new class sample and each initial base class sample in the first dimension, and determining a second similarity value of the initial new class sample and each initial base class sample in the second dimension includes:
[0025] Determining a first probability distribution set of the initial new class sample and determining a second probability distribution set of the initial base class sample;
[0026] According to the first probability distribution set and the second probability distribution set, forming a joint distribution set;
[0027] Determine a first similarity value based on the expected value of the distance between any initial new class feature and any initial base class feature under the joint distribution set;
[0028] Determine a second similarity value based on the divergence value between the first probability distribution set and the second probability distribution set.
[0029] In some embodiments, the new class feature matrix includes a plurality of first matrix elements, and the base class feature matrix includes a plurality of second matrix elements;
[0030] Based on the new class feature matrix and the base class feature matrix, determine the third similarity value of the initial new class sample and each initial base class sample in the third dimension, including:
[0031] Calculate the sum of squares between any first matrix element and any second matrix element to obtain an element square value;
[0032] Perform a square root operation on the element square value to obtain an element square root value;
[0033] Determine the third similarity value based on the element square root value.
[0034] In some embodiments, based on the similarity matrix, select a target base class sample from a plurality of initial base class samples, including:
[0035] Based on a preset normalization operation operator, normalize the similarity matrix of all initial base class samples to obtain an adjusted similarity matrix, where the adjusted similarity matrix includes a plurality of third matrix elements;
[0036] For any adjusted similarity matrix, if all third matrix elements are greater than a preset similarity threshold, determine the corresponding initial base class sample as the target base class sample.
[0037] In some embodiments, according to the target base class sample and the corresponding similarity matrix of the target base class sample, dynamically correct the initial new class sample to obtain a target new class sample, including:
[0038] Based on a plurality of initial base class features, obtain the initial base class covariance values corresponding to each initial base class sample;
[0039] Based on the similarity matrix, update the initial base class mean and the initial base class covariance value respectively to obtain a target base class mean and a target base class covariance value;
[0040] Based on the target base class mean and the target base class covariance value, transfer the target base class sample to the initial new class sample to obtain a dynamically corrected target new class sample.
[0041] In some embodiments, before determining the similarity matrix based on the composite metric function, further include:
[0042] If the preset distribution hyperparameter value of the initial image classification model is not equal to the preset parameter threshold, perform logarithmic processing on all the initial new-class samples to obtain the updated initial new-class samples;
[0043] If the preset distribution hyperparameter value of the initial image classification model is equal to the preset parameter threshold, perform exponential processing on all the initial new-class samples to obtain the updated initial new-class samples.
[0044] To achieve the above object, a second aspect of the embodiments of the present application proposes a method for training an image classification model, including:
[0045] Obtain initial samples, where the initial samples include initial new-class samples and a plurality of initial base-class samples;
[0046] Input the initial samples into an initial image classification model, and determine a corresponding similarity matrix of the initial samples based on a preset composite metric function, where the similarity matrix is used to characterize the feature matching degrees of the initial new-class samples and each of the initial base-class samples in multiple dimensions;
[0047] Based on the similarity matrix, select target base-class samples from the plurality of initial base-class samples;
[0048] According to the target base-class samples and the corresponding similarity matrix of the target base-class samples, perform dynamic correction on the initial new-class samples to obtain target new-class samples;
[0049] Use the target new-class samples to train the initial image classification model until a training stop condition is reached to obtain a trained image classification model.
[0050] To achieve the above object, a third aspect of the embodiments of the present application proposes an image classification device, including:
[0051] An acquisition module, configured to obtain initial samples, where the initial samples include initial new-class samples and a plurality of initial base-class samples;
[0052] A similarity matrix determination module, configured to input the initial samples into an initial image classification model, and determine a corresponding similarity matrix of the initial samples based on a preset composite metric function, where the similarity matrix is used to characterize the feature matching degrees of the initial new-class samples and each of the initial base-class samples in multiple dimensions;
[0053] A target base-class sample selection module, configured to select target base-class samples from the plurality of initial base-class samples based on the similarity matrix;
[0054] A calibration module, configured to dynamically calibrate the initial new class samples according to the target base class samples and the corresponding similarity matrix of the target base class samples, so as to obtain target new class samples;
[0055] A training module, configured to train the initial image classification model by using the target new class samples until a training stop condition is reached, so as to obtain a trained image classification model;
[0056] An application module, configured to obtain a target image to be classified, and input the target image into the trained image classification model to obtain a corresponding target classification result.
[0057] To achieve the above object, a fourth aspect of the embodiments of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the image classification method of the first aspect or the training method of the image classification model of the second aspect is implemented.
[0058] To achieve the above object, a fifth aspect of the embodiments of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the image classification method of the first aspect or the training method of the image classification model of the second aspect is implemented.
[0059] The image classification method, the training method of the image classification model, and related devices provided by the present application. The image classification method includes obtaining initial samples, where the initial samples include initial new class samples and multiple initial base class samples; inputting the initial samples into an initial image classification model, and determining a corresponding similarity matrix of the initial samples based on a preset composite metric function. The similarity matrix is used to characterize the feature matching degree between the initial new class samples and each initial base class sample in multiple dimensions; the similarity matrix determined according to the composite metric function provides more position and distribution information, improving the accuracy of the feature distribution description between the two types of samples; then, based on the similarity matrix, target base class samples are selected from multiple initial base class samples; according to the target base class samples and the corresponding similarity matrix of the target base class samples, the initial new class samples are dynamically calibrated to make up for the problem of scarce effective training features in the initial new class samples by accurately capturing more similar target base class samples, so as to obtain target new class samples; further, it is convenient to use the target new class samples with more real and accurate feature distribution to help improve the image classification accuracy of the trained model; then, the initial image classification model is trained by using the target new class samples until a training stop condition is reached, so as to obtain a trained image classification model; after that, a target image to be classified is obtained, and the target image is input into the trained image classification model to obtain a corresponding target classification result. Description of the Drawings
[0060] Figure 1 It is a schematic diagram of the application scenario of the image classification device provided by an embodiment of the present application;
[0061] Figure 2 It is an optional flowchart of the image classification method provided by an embodiment of the present application;
[0062] Figure 3 It is an optional schematic diagram of sample processing of the image classification method provided by an embodiment of the present application;
[0063] Figure 4 It is an optional schematic diagram of calibration of the image classification method provided by an embodiment of the present application;
[0064] Figure 5 It is an optional flowchart of the training method of the image classification model provided by an embodiment of the present application;
[0065] Figure 6 It is an optional flowchart of the image classification device provided by an embodiment of the present application;
[0066] Figure 7 It is a schematic diagram of the hardware structure of the electronic device provided by an embodiment of the present application. Detailed implementation manners
[0067] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0068] It should be noted that although functional module division is performed in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first" and "second" in the description, claims and the above drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.
[0069] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application and are not intended to limit the present application.
[0070] First, several nouns involved in the present application are analyzed:
[0071] Artificial Intelligence (AI): It is a new technical science that studies and develops theories, methods, technologies, and application systems for simulating, extending, and expanding human intelligence; artificial intelligence is a branch of computer science. Artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. The research in this field includes robots, speech recognition, image recognition, natural language processing, and expert systems, etc. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence also refers to the theory, method, technology, and application system that uses a digital computer or a machine controlled by a digital computer to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.
[0072] Among them, machine learning is a subfield of artificial intelligence. Machine learning aims to enable relevant models to automatically learn and improve their performance through data without explicit programming. In other words, the goal of machine learning is to let the computer "learn" the rules from the data and use these rules for intelligent prediction or decision-making.
[0073] Furthermore, few-shot learning (small-sample learning): It is a task in the field of machine learning, aiming to solve classification or regression problems when the training data is very limited. Different from traditional supervised learning methods, few-shot learning assumes that the model only encounters a small number of samples (usually 1 to 5) of each category during the training phase, and needs to reason and classify new categories during the test phase.
[0074] Image classification refers to assigning the input image to one of the predefined categories. Among them, few-shot learning image classification is a common image classification task, and its purpose is to train the model to achieve efficient task execution with only a small number of labeled samples.
[0075] The related technologies mainly focus on the generation of samples, and help the model to encounter more sample data during the training process by generating more samples. However, such an image classification method ignores the true feature distribution of the target category group represented by the few samples, resulting in a deviation between the generated samples and the few samples, and further causing poor classification accuracy of the model images obtained by training.
[0076] Based on this, the embodiments of the present application provide an image classification method, a training method for an image classification model, and related devices, aiming to improve the accuracy of image classification.
[0077] Exemplarily, as Figure 1 shown, Figure 1It is a schematic diagram of the application scenario of the image classification device provided by the embodiments of the present application. In an optional application scenario, the client 11 is communicatively connected to the server 12. The image classification device proposed by the embodiments of the present application is deployed in the server 12. Moreover, an image classification model is set in the image classification device. The image classification model is pre-trained based on new class samples. The new class samples refer to samples that the image classification model has never seen before. The new class samples only include a small number of labeled samples available for training. The user inputs a target image to be classified through the client 11. The trained image classification model performs image recognition on the target image and determines which category in the preset categories the target image belongs to based on the recognized image features, thereby completing the image classification task. Since the image classification model proposed by the embodiments of the present application performs calibration processing on the new class samples based on the image classification model training method proposed by the embodiments of the present application, improving the accuracy of the feature distribution of the target category group indicated by the new class samples in the feature space. Therefore, the image classification model set in the server 12 and trained based on the calibrated new class samples can perform image classification processing on the input target image to obtain a corresponding highly accurate target classification result.
[0078] It should be noted that in the embodiments of the present application, when it comes to information related to user characteristics such as user basic information or user identity, user permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of the present application need to obtain sensitive personal information of users, separate permission or separate consent of the users will be obtained first. After clearly obtaining the separate permission or separate consent of the users, the necessary data for the normal operation of the embodiments of the present application will be obtained. For example, when the embodiments of the present application obtain initial samples, the authorization or consent of relevant personnel will be obtained first. Otherwise, initial samples that cannot be used in the embodiments of the present application will be obtained. Additionally, other relevant data obtained by the image classification device of the present application are all authorized data, which will not be elaborated here one by one.
[0079] In the embodiments of the present application, the description will be made from the dimension of the image classification device. The image classification device can be integrated in a computer device, such as a server. As Figure 2 shown, Figure 2 is an optional flowchart of the image classification method provided by the embodiments of the present application. Figure 2 The method in Figure 2 may include but is not limited to the following steps 101 to step 106. When the image classification device executes the image classification method, the specific process is as follows. It should be noted first that the order of steps 101 to 106 in
[0080] Step 101: Obtain an initial sample, which includes an initial new-class sample and multiple initial base-class samples.
[0081] The following provides a detailed description of Step 101.
[0082] Among them, the initial sample refers to a dataset used to train or evaluate a machine learning model. Generally, the initial sample includes a large number of initial base-class samples with known classification labels under different base-class categories, and a small number of initial new-class samples with known classification labels under different new-class categories, which represent what the initial image classification model has never seen or is not familiar with. Generally, the number of base-class categories is much larger than the number of new-class categories. Each base-class category includes multiple initial base-class samples, and each new-class category includes one or two initial new-class samples. For ease of description, in the embodiments of the present application, it is exemplified that the base-class category to which each initial base-class sample belongs and the new-class category to which each initial new-class sample belongs are all different.
[0083] Furthermore, in the image classification task of the embodiments of the present application, both the initial base-class samples and the initial new-class samples are image samples. Among them, the classification granularity of the classification labels of the initial samples can be set according to the actual situation, and the embodiments of the present application do not limit this. For example, an initial base-class sample a has a species classification label of "cat", and another initial base-class sample b has a human body part classification label of "human eye".
[0084] Among them, the initial image classification model can already perform good image classification processing on the initial base-class samples. However, without being trained by the graphic classification method proposed in the embodiments of the present application, the accuracy of the image classification results for the initial new-class samples is relatively poor. The embodiments of the present application take the initial image classification model as a starting point and perform further optimization training to improve the recognition ability of the trained image classification model for the initial new-class samples.
[0085] As Figure 3 shown, Figure 3 is an optional sample processing schematic diagram of the image classification method provided by the embodiments of the present application. The samples for training the initial image classification model include initial base-class samples and initial new-class samples. The initial new-class samples further include support-class samples and query-class samples. The support-class samples include one sample each under five new-class categories (Class), and the support-class samples are used to train the initial image classification model; the target category object represented by the query-class samples is one of the five new-class category samples, and the query-class samples are used to test and verify the trained image classification model.
[0086] It should be noted that the initial base class samples are usually obtained from an open-source database. The initial new class samples can be obtained from an open-source database or captured in real time. The acquisition method of the initial samples is not unique. The embodiments of the present application only give examples and do not limit it.
[0087] Step 102: Input the initial samples into the initial image classification model, and determine the corresponding similarity matrix based on a preset composite metric function. The similarity matrix is used to characterize the feature matching degree between the initial new class samples and each initial base class sample in multiple dimensions.
[0088] The following gives a detailed description of Step 102.
[0089] Among them, the initial image classification model can be a wide residual network, a convolutional neural network model, a recurrent neural network model, a convolutional recurrent neural network model, etc., which can be specifically adjusted according to the actual situation. The embodiments of the present application do not limit this. As Figure 3 shown, in the present application, a wide residual network is taken as an example. The wide residual network (WideResNet) is a deep convolutional neural network improved on the basis of the residual network. Its main feature is to use wide convolutional layers, increasing the number of channels of the network, thereby improving the expression ability and performance of the network.
[0090] Among them, the composite metric function is a function used to comprehensively evaluate the similarity degree between the initial base class samples and the initial new class samples (two types of samples) from multiple dimensions. The matrix elements in the similarity matrix generated based on the result of the composite metric function are used to represent the similarity scores between different feature points of the two types of samples. Compared with the feature distribution description obtained by the traditional method under a single-dimensional description, the similarity matrix determined according to the composite metric function in the embodiments of the present application provides more position and distribution information, thereby improving the accuracy of the feature distribution description between the two types of samples.
[0091] Furthermore, as Figure 3 shown, after inputting the initial samples into the initial image classification model, the initial image classification model respectively performs feature extraction processing on the initial base class samples and the support class samples to obtain corresponding multiple initial base class features and multiple initial new class samples. The feature extraction processing includes data augmentation processing, convolutional processing, feature extraction processing, activation and other processing steps. The specific steps of the feature extraction processing can be adaptively adjusted according to the actual situation. The embodiments of the present application do not limit this.
[0092] In some embodiments, before determining the similarity matrix based on the composite metric function, the following steps are further included:
[0093] (102.a.1) If the preset distribution hyperparameter value of the initial image classification model is not equal to the preset parameter threshold, perform logarithmic processing on all initial new class samples to obtain the updated initial new class samples.
[0094] (102.a.2) If the preset distribution hyperparameter value of the initial image classification model is equal to the preset parameter threshold, perform exponential processing on all initial new class samples to obtain the updated initial new class samples.
[0095] The following provides a detailed description of steps (102.a.1) to (102.a.2).
[0096] In some embodiments, since the Gaussian distribution can simplify the computational complexity of the initial image classification model, and the feature distribution of the initial new class samples does not necessarily follow the Gaussian distribution, as shown in Equation <1> below, the embodiments of the present application will first perform Tukey's Ladder of Power Transformation (which can also be referred to as "Tukey transformation") on the support class samples S and the query class samples Q:
[0097]
[0098] where α is the preset distribution hyperparameter value, which is used to adjust the distribution deviation; in the embodiments of the present application, 0 is the preset parameter threshold. Of course, the preset parameter threshold can be adaptively adjusted according to the actual situation; x i represents the i-th feature from the support class samples or the query class samples.
[0099] where logx i represents performing logarithmic processing on all initial new class samples; represents performing logarithmic processing on all initial new class samples.
[0100] Further, after the Tukey transformation, the support class samples or the query class samples are respectively updated to and where y i is the classification label of the support class samples or the query class samples, N is the number of the support class samples or the query class samples, K is the feature dimension of the support class samples, and q is the feature dimension of the query class samples.
[0101] In some embodiments, determining the similarity matrix based on the composite metric function includes the following steps:
[0102] (102.b.1) Determine the new class feature matrix of the initial new class samples and the base class feature matrix of the initial base class samples respectively.
[0103] The following provides a detailed description of step (102.b.1).
[0104] Among them, both the new class feature matrix and the base class feature matrix are in matrix form structures. The new class feature matrix is used to represent the positions and feature attributes of multiple initial new class features in the feature space, and the base class feature matrix is used to represent the positions and feature attributes of multiple initial base class features in the feature space.
[0105] Furthermore, both the new class feature matrix and the base class feature matrix are the products of the complex relationships between the sample features captured under the conditions of a high-dimensional feature space. The unified feature representation facilitates subsequent comparison of the similarity degrees of the initial base class samples and the initial new class samples in the same feature space.
[0106] Next, as Figure 3 shown, after performing the Tukey transformation processing on the initial new class samples, it is necessary to determine the feature matrices of the initial new class samples and the initial base class samples respectively, so as to perform similarity comparison through the feature matrices of the two types of samples subsequently, and then select the target base class samples for correcting the initial new class samples from the database containing multiple initial base class samples.
[0107] In some embodiments, determining the new class feature matrix of the initial new class samples and the base class feature matrix of the initial base class samples respectively includes the following steps:
[0108] (A.1) Perform transposition processing on each initial new class feature to obtain the corresponding new class transposed feature.
[0109] (A.2) Based on the cross product of each initial new class feature and the corresponding new class transposed feature, obtain the new class feature matrix.
[0110] (A.3) Based on multiple initial base class features, obtain the corresponding initial base class means of each initial base class sample.
[0111] (A.4) Perform transposition processing on each initial base class mean to obtain the corresponding base class transposed mean.
[0112] (A.5) Based on the cross product of each initial base class mean and the corresponding base class transposed mean, obtain the base class feature matrix.
[0113] The following will describe steps (A.1) to (A.5) in detail.
[0114] In some embodiments, the new class feature matrix M is determined by the following formula <2> i :
[0115]
[0116] where, v i represents the i-th initial new class feature; T represents transposition processing; for v iPerform a transpose operation to obtain v i The corresponding new class transpose feature · is the cross product (also known as the vector product or outer product).
[0117] Furthermore, compared with the traditional method that only describes the initial new class features from a single dimensional direction, the new class feature matrix in the embodiments of the present application includes the feature description distributions of multiple initial new class features from the original dimensional direction, and also includes the feature description distributions in other dimensional directions after transposition. Therefore, more comprehensive and richer feature distribution information of the initial new class features in multiple dimensions can be obtained.
[0118] Furthermore, since the number of initial base class samples is much larger than the number of initial new class samples, for each initial base class sample, if the corresponding base class feature matrix is calculated according to Equation <2>, the initial image classification model will consume a large amount of computing power resources due to high-complexity calculations. Based on this, the embodiments of the present application first obtain the corresponding initial base class mean μ of each initial base class sample through the following Equation <3> j :
[0119]
[0120] where n j represents the number of initial base class features of the j-th initial base class sample, and x i represents the i-th initial base class feature in the j-th initial base class sample; Σ represents the summation process.
[0121] In some other embodiments, if there are multiple initial base class samples under one base class category, then n j represents the number of initial base class samples of the j-th base class category, and x i represents the i-th initial base class sample in the j-th base class category.
[0122] Then, determine the base class feature matrix M′ through the following Equation <4> j :
[0123]
[0124] where μ j represents the initial base class mean corresponding to the j-th initial base class sample; T represents the transpose operation; perform a transpose operation on μ j to obtain the transposed mean μ j of the corresponding base class · is the cross product.
[0125] In this way, the computational complexity of the base class feature matrix is reduced from O(n 2Reduce to O(1). When the initial image classification model processes the initial base class samples, especially a large number of initial base class samples, it can improve the sample processing speed and reduce the waste of computing power resources.
[0126] Next, the following continues to describe steps (102.b.2) to (102.b.4) in detail.
[0127] (102.b.2) Based on multiple initial new class features and multiple initial base class features, determine the first similarity value of the initial new class sample and each initial base class sample in the first dimension, and determine the second similarity value of the initial new class sample and each initial base class sample in the second dimension.
[0128] (102.b.3) Based on the new class feature matrix and the base class feature matrix, determine the third similarity value of the initial new class sample and each initial base class sample in the third dimension.
[0129] (102.b.4) According to the composite metric function set with preset hyperparameters, fuse the first similarity value, the second similarity value, and the third similarity value to obtain a similarity matrix.
[0130] In some embodiments, the similarity matrix D is determined by the following formula <5> of the composite metric function set with preset hyperparameters cm :
[0131] D cm =β·W + ò·F + η·JS <5>
[0132] where β, ò, and η are preset hyperparameters, and their specific values can be adaptively adjusted according to the actual situation, and the embodiments of the present application do not limit this. W, and JS are the similarity values of the initial new class sample and each initial base class sample calculated based on the earth mover's distance (Wasserstein distance), Frobenius norm (Frobenius norm), and Jensen-Shannon divergence (JS divergence) methods, respectively.
[0133] It can be understood that the traditional method relying on the Euclidean distance in a single dimension as the similarity metric standard has limitations. Especially in the high-dimensional feature space, such a metric method ignores the shape, variance, and other high-order statistical attributes of the data distribution, and it is difficult to comprehensively reflect the actual similarity between the initial base class samples and the initial new class samples. In contrast, the embodiments of the present application integrate different metric standards to more comprehensively capture the similarity between the two types of samples.
[0134] It should be noted that a composite metric function for determining the similarity matrix in the embodiments of the present application can also be constructed based on metric methods such as chi-square divergence, Hellinger distance, Bhattacharyya distance, etc., and can be specifically selected according to actual situations, and the embodiments of the present application do not limit this.
[0135] In some embodiments, based on a plurality of initial new class features and a plurality of initial base class features, determining a first similarity value of an initial new class sample and each initial base class sample in a first dimension, and determining a second similarity value of the initial new class sample and each initial base class sample in a second dimension, includes the following steps:
[0136] (B.1) Determine a first probability distribution set of the initial new class sample, and determine a second probability distribution set of the initial base class sample.
[0137] (B.2) According to the first probability distribution set and the second probability distribution set, form a joint distribution set.
[0138] (B.3) Based on the expected value of the distance between any initial new class feature and any initial base class feature under the joint distribution set, determine the first similarity value.
[0139] (B.4) Based on the divergence value between the first probability distribution set and the second probability distribution set, determine the second similarity value.
[0140] The following makes a detailed description of steps (B.1) to (B.4).
[0141] Among them, the first similarity value W between the initial base class sample and the initial new class sample is measured by the following formula <6>:
[0142]
[0143] Among them, ∏(p 1 , p 2 ) is a joint distribution set composed of the first probability distribution set P 1 and the second probability distribution set P 2 ; for each possible joint distribution γ, a sample pair (x, y) ∼ γ can be obtained, that is, a sample pair composed of any initial new class feature and any initial base class feature; then, evaluate the distance ||x - y|| between these sample pairs, and calculate the expected value of the distance of this sample pair under the joint distribution γ inf represents "infimum", and further, the Wasserstein distance solves the minimum value of this expected value among all potential joint distributions; intuitively, it can be understood as moving a pile of soil from position P 1 to position P 2The cost, the Wasserstein distance represents the minimum cost under the optimal path planning. One of the key advantages of the Wasserstein distance measurement method is that it can measure the distance between two probability distributions, even if their support sets do not overlap.
[0144] Among them, the first probability distribution set and the second probability distribution set can be obtained by methods such as parameter estimation method, maximum likelihood estimation method, Bayesian estimation method, nonparametric estimation method, histogram method, etc., and can be specifically selected according to the actual situation. The embodiments of the present application do not limit this.
[0145] Further, the second similarity value JS between the initial base class samples and the initial new class samples is measured by the following formula <7>:
[0146]
[0147] Among them, KL is the Kullback-Leibler Divergence formula. The KL divergence is also called relative entropy and is an asymmetric measure of the difference between two probability distributions. Formula <7> processes the KL divergence by variant and uses the JS divergence as one of the measurement methods. Therefore, when there is an intersection between two probability distributions, the similarity result obtained by the symmetric JS divergence measurement method will be more accurate.
[0148] In some embodiments, based on the new class feature matrix and the base class feature matrix, determining the third similarity value between the initial new class samples and each initial base class sample in the third dimension includes the following steps:
[0149] (C.1) Calculate the sum of squares between any first matrix element and any second matrix element to obtain the element square value.
[0150] (C.2) Take the square root of the element square value to obtain the element square root value.
[0151] (C.3) Determine the third similarity value based on the element square root value.
[0152] The following details steps (C.1) to (C.3).
[0153] Among them, the third similarity value F between the initial base class samples and the initial new class samples is measured by the following formula <8>:
[0154]
[0155] Among them, M i is the new class feature matrix, M′ jis the base class feature matrix; the Frobenius norm is a generalization of the Euclidean norm (i.e., the length or magnitude of a vector) in the context of matrices. Therefore, the third similarity value in Equation <8> is obtained by taking the square root after squaring the sum of any first matrix element and any second matrix element. Moreover, the Frobenius norm measurement method is regarded as the Euclidean length of the matrix in the high-dimensional space, making the output result easy to understand and interpret.
[0156] Step 103: Based on the similarity matrix, select a target base class sample from multiple initial base class samples.
[0157] The following details Step 103.
[0158] In some embodiments, the initial base class samples that are more similar to the initial new class sample have more similar feature representation information. That is, different initial base class samples have different correction contributions to the initial new class sample. In traditional correction methods, no weights or fixed weights are used to measure the contribution degree of the initial base class samples, resulting in the initial base class samples with low similarity occupying a larger weight instead, thereby causing unreasonable correction of the initial new class sample. In contrast, in the embodiments of the present application, based on the similarity matrix obtained by composite measurement, and based on the dynamically adjusted K value, a target base class sample that is more similar and better matched to the initial new class sample is dynamically selected from multiple initial base class samples, improving the authenticity of the feature distribution of the subsequent obtained target new class sample.
[0159] In some embodiments, selecting a target base class sample from multiple initial base class samples based on the similarity matrix includes the following steps:
[0160] (103.a.1) Based on a preset normalization operation operator, normalize the similarity matrix of all initial base class samples to obtain an adjusted similarity matrix, where the adjusted similarity matrix includes multiple third matrix elements.
[0161] (103.a.2) For any adjusted similarity matrix, if all third matrix elements are greater than a preset similarity threshold, determine the corresponding initial base class sample as the target base class sample.
[0162] The following details Steps (103.a.1) to (103.a.2).
[0163] Among them, the adjusted similarity matrix is obtained by the following Equation <9>:
[0164] W m =f(D cm ) <9>
[0165] Among them, D cmis the similarity matrix; the similarity matrix of each initial base class sample is normalized to obtain D cm correspondingly adjust the similarity matrix W m ; f(·) is a preset normalization operation operator.
[0166] Further, W m includes multiple third matrix elements. If all third matrix elements are greater than a preset similarity threshold T, the corresponding initial base class sample is determined as the target base class sample. In the embodiments of the present application, the number of target base class samples is set to K. If the number of adjusted similarity matrices in which all third matrix elements are greater than the similarity threshold T is less than K, the corresponding adjusted similarity matrices can be selected in descending order according to the number of third matrix elements greater than the preset similarity threshold T until the number of selected adjusted similarity matrices is equal to K. Among them, the specific value of K can be adaptively adjusted according to the actual situation, and the similarity threshold T is usually set between 0 and 1.
[0167] Alternatively, first select the most matching k from multiple adjusted similarity matrices. Then, if there are adjusted similarity matrices lower than the similarity threshold T among the selected k adjusted similarity matrices, the ones with lower similarity are removed in turn until the number of target base class samples is K, where K < k.
[0168] It can be understood that if there are samples with very low similarity in the target base class samples, it will lead to a large deviation between the actual target new class samples obtained by subsequent calibration and the expected target new class samples. In the embodiments of the present application, the target base class samples are determined based on the dynamic K value, which can select more matching target base class samples for calibration for the initial new class samples on the basis of ensuring similarity.
[0169] As Figure 4 shown, Figure 4 is an optional calibration schematic diagram of the image classification method provided by the embodiments of the present application. The initial new class samples are within the dotted circle. To ensure the authenticity of the feature distribution of the target new class samples after calibration, it is necessary to select initial base class samples with higher similarity (characterized as closer distance in the figure) from multiple initial base class samples as the target base class samples. Among them, the initial base class samples within the dotted box are the target base class samples, and the initial base class samples outside the dotted box are characterized as unselected samples, which are not shown in the figure. For the target base class samples with higher similarity to the initial new class samples, the weight values represented by the elements in their similarity matrices are correspondingly higher. Thus, samples with higher similarity can provide more contributions to the calibration of the initial new class samples, making the feature distribution of the calibrated target new class samples more real and accurate.
[0170] Step 104: Dynamically correct the initial new class samples according to the target base class samples and the corresponding similarity matrix thereof to obtain the target new class samples.
[0171] The following provides a detailed description of Step 104.
[0172] In some embodiments, in the few-shot learning image classification task, dynamically correcting the initial new class samples based on the target base class samples and their corresponding similarity matrix can make up for the problem of scarce effective training features in the initial new class samples by accurately capturing more similar target base class samples, and further facilitate using the target new class samples with more real and accurate feature distributions to help improve the image classification accuracy of the trained model.
[0173] In some embodiments, dynamically correcting the initial new class samples according to the target base class samples and the corresponding similarity matrix thereof to obtain the target new class samples includes the following steps:
[0174] (104.a.1) Based on multiple initial base class features, obtain the initial base class covariance values corresponding to each initial base class sample.
[0175] (104.a.2) Based on the similarity matrix, update the initial base class mean and the initial base class covariance value respectively to obtain the target base class mean and the target base class covariance value.
[0176] (104.a.3) Based on the target base class mean and the target base class covariance value, transfer the target base class samples to the initial new class samples to obtain the dynamically corrected target new class samples.
[0177] The following provides a detailed description of Step (104.a.1) to Step (104.a.3).
[0178] Among them, the initial base class covariance value matrix ∑ corresponding to each initial base class sample is determined by the following formula <10> j :
[0179]
[0180] where n j represents the number of initial base class features of the j-th initial base class sample, and x i represents the i-th initial base class feature in the j-th initial base class sample; Σ represents the summation process; T represents the transpose process.
[0181] Furthermore, the target base class mean and the target base class covariance value
[0182]
[0183] where λ j is the j-th third matrix element in the adjusted similarity matrix W m and the eigenvector comes from the updated support class samples S, and δ is the compensation for the intra-class variation of the initial new class samples, and δ can be adaptively adjusted according to the actual situation.
[0184] Furthermore, based on the target base class mean and the target base class covariance value, the feature statistical information of the K target base class samples is migrated to the initial new class samples, and the target new class samples y i obtained after correction and the target base class samples satisfy the following equations <11> and <12> in terms of feature distribution:
[0185]
[0186] where is the learned feature distribution set; N is the feature distribution between the target new class samples and any target base class sample.
[0187] Among them, the migration of feature statistical information can be to migrate the classification labels of the target base class samples to the initial new class samples to supplement more descriptions of the initial new class samples and help the training of the initial image classification model; or, it can be to generate target new class samples that conform to the true distribution characteristics of the initial new class samples according to the feature statistical information. A large number of target new class samples with correct feature distributions can help improve the image recognition and discrimination classification capabilities of the initial image classification model for new class samples in the subsequent process.
[0188] Step 105, use the target new class samples to train the initial image classification model until the training stop condition is reached, and obtain the trained image classification model.
[0189] The following gives a detailed description of Step 105.
[0190] Among them, the training stop condition can be that the number of training rounds reaches a preset maximum value; or, when the change of the loss function used to evaluate the training degree is less than a certain threshold, it is considered that the initial image classification model has converged and the training stops; or, during the training process, an independent validation set is used to evaluate the performance of the model. If the performance (such as accuracy, F1 score, etc.) on the validation set no longer improves, the training stops. Of course, the training stop condition can be set according to the actual situation, and the embodiments of the present application do not limit this.
[0191] Furthermore, as Figure 3As shown, after obtaining the trained image classification model, in order to evaluate the generalization ability of the initial image classification model, it is also necessary to use query class samples to verify the trained image classification model. If the image classification model can accurately classify the target object represented by the query class samples, it is generally considered that the training result of the image classification model is good.
[0192] As Figure 3 shown, the target object represented by the query class samples adopted in the embodiments of the present application is one of the new class categories to which the support class samples belong. Thus, the image classification processing ability of the image classification model for new class samples can be tested. The number of query class samples can be set according to the actual situation. For example, a corresponding query class sample can be set for each new class category, and the training is completed when each query class sample is correctly classified; or, the training is completed when the classification accuracy rate of all query class samples reaches a preset accuracy rate threshold.
[0193] Step 106, obtain the target image to be classified, input the target image into the trained image classification model, and obtain the corresponding target classification result.
[0194] The following will describe step 106 in detail.
[0195] Among them, the target image refers to the image for which the category needs to be predicted by the trained image classification model. The target image can be obtained by the user's real-time shooting or from an open-source database. The embodiments of the present application do not limit the acquisition method of the target image.
[0196] Among them, the target classification result refers to the predicted category obtained by the image classification model after classifying the input target image. The target classification result is usually a category label or a probability distribution of a group of categories. Depending on the task form, the target classification result has different output methods: for a single-label classification task, the target classification result is a category label; for a multi-label classification task, the target classification result is multiple category labels and their corresponding confidence levels; the specific output method can be adjusted adaptively according to the actual situation.
[0197] As Figure 5 shown, Figure 5 is an optional flowchart of the training method of the image classification model provided by the embodiments of the present application. Figure 5 The method in Figure 5 may include but is not limited to the following steps 101 to 106. When the image classification device executes the training method of the image classification model, the specific process is as follows. It should be noted first that the embodiments of the present application do not make specific limitations on
[0198] Step 201: Obtain an initial sample, where the initial sample includes an initial new-class sample and multiple initial base-class samples;
[0199] Step 202: Input the initial sample into an initial image classification model, and determine a corresponding similarity matrix for the initial sample based on a preset composite metric function. The similarity matrix is used to characterize the feature matching degree between the initial new-class sample and each of the initial base-class samples in multiple dimensions;
[0200] Step 203: Based on the similarity matrix, select a target base-class sample from the multiple initial base-class samples;
[0201] Step 204: Dynamically correct the initial new-class sample according to the target base-class sample and the similarity matrix corresponding to the target base-class sample to obtain a target new-class sample;
[0202] Step 205: Use the target new-class sample to train the initial image classification model until a training stop condition is reached, and obtain a trained image classification model.
[0203] The following provides a detailed description of Steps 201 to 205.
[0204] Among them, the initial sample refers to a data set used to train or evaluate a machine learning model. Generally, the initial sample includes a large number of initial base-class samples with known classification labels under different base-class categories, and a small number of initial new-class samples with known classification labels under different new-class categories, which represent that the initial image classification model has never seen or is not familiar with. Generally, the number of base-class categories is much larger than the number of new-class categories. Each base-class category includes multiple initial base-class samples, and each new-class category includes one or two initial new-class samples. For ease of description, in the embodiments of the present application, it is exemplified that the base-class category to which each initial base-class sample belongs and the new-class category to which each initial new-class sample belongs are all different.
[0205] Further, in the image classification task of the embodiments of the present application, both the initial base-class samples and the initial new-class samples are image samples. Among them, the classification granularity of the classification labels of the initial samples can be set according to the actual situation, and the embodiments of the present application do not limit this. For example, an initial base-class sample a has a species classification label "cat", and another initial base-class sample b has a human body part classification label "human eye".
[0206] Among them, the initial image classification model can already perform good image classification processing on the initial base class samples. However, without being trained by the graph classification method proposed in the embodiments of the present application, the accuracy of the image classification result for the initial new class samples is relatively poor. The embodiments of the present application use the initial image classification model as a starting point and perform further optimization training to improve the recognition ability of the trained image classification model for the initial new class samples.
[0207] It should be noted that the initial base class samples are usually obtained from an open-source database. The initial new class samples can be obtained from an open-source database or can be obtained by real-time shooting. The acquisition method of the initial samples is not unique. The embodiments of the present application only give examples and do not limit them.
[0208] Among them, the composite metric function is a function used to comprehensively evaluate the similarity between two types of samples from multiple dimensions. The matrix elements in the similarity matrix generated based on the result of the composite metric function are used to represent the similarity scores between different feature points of the two types of samples. Compared with the feature distribution description obtained under the single-dimensional description of the traditional method, the similarity matrix determined according to the composite metric function in the embodiments of the present application provides more position and distribution information, thereby improving the accuracy of the feature distribution description between the two types of samples.
[0209] In some embodiments, the initial base class samples that are more similar to the initial new class samples have more similar feature representation information. That is, different initial base class samples have different correction contributions to the initial new class samples. In traditional correction methods, no weights or fixed weights are used to measure the contribution degree of the initial base class samples, resulting in the fact that the initial base class samples with low similarity instead occupy a larger weight, thereby causing unreasonable correction of the initial new class samples. In contrast, in the embodiments of the present application, based on the similarity matrix obtained by the composite metric, and based on the dynamically adjusted K value, target base class samples that are more similar and better matched to the initial new class samples are dynamically selected from multiple initial base class samples, improving the authenticity of the feature distribution of the subsequent obtained target new class samples.
[0210] In some embodiments, in the image classification task of few-shot learning, dynamically correcting the initial new class samples based on the target base class samples and their corresponding similarity matrix can effectively address problems such as data scarcity, class imbalance, and insufficient feature representation in few-shot learning, thereby improving the image classification accuracy of the finally trained model.
[0211] Among them, the training stop condition can be that the number of training rounds reaches a preset maximum value; or, when the change of the loss function used to evaluate the training degree is less than a certain threshold, it is considered that the initial image classification model has converged and the training stops; or, during the training process, an independent validation set is used to evaluate the performance of the model. If the performance on the validation set (such as accuracy, F1 score, etc.) no longer improves, the training stops. Of course, the training stop condition can be set according to the actual situation, and the embodiments of the present application do not limit this.
[0212] Further, as Figure 3 shown, after obtaining the trained image classification model, in order to evaluate the generalization ability of the initial image classification model, it is also necessary to use query class samples to verify the trained image classification model. If the image classification model can accurately classify the target object represented by the query class samples, it is generally considered that the training result of the image classification model is good.
[0213] In addition, steps 201 to 205 are similar to steps 101 to 105. Therefore, other details of steps 201 to 205 can be seen from steps 101 to 105 and will not be elaborated here.
[0214] As Figure 6 shown, Figure 6 is an optional flowchart of an image classification device provided by an embodiment of the present application. The image classification device includes the following modules 301 to 306:
[0215] An acquisition module 301, configured to acquire initial samples, where the initial samples include initial new class samples and a plurality of initial base class samples;
[0216] A similarity matrix determination module 302, configured to input the initial samples into an initial image classification model, and determine a corresponding similarity matrix of the initial samples based on a preset composite metric function, where the similarity matrix is used to characterize the feature matching degrees of the initial new class samples and each of the initial base class samples in multiple dimensions;
[0217] A target base class sample selection module 303, configured to select target base class samples from the plurality of initial base class samples based on the similarity matrix;
[0218] A correction module 304, configured to dynamically correct the initial new class samples according to the target base class samples and the corresponding similarity matrix of the target base class samples to obtain target new class samples;
[0219] A training module 305, configured to use the target new class samples to train the initial image classification model until a training stop condition is reached to obtain a trained image classification model;
[0220] An application module 306 is configured to obtain a target image to be classified, and input the target image into the trained image classification model to obtain a corresponding target classification result.
[0221] The image classification method, the training method of the image classification model, and related devices proposed in this application. The image classification method includes obtaining initial samples, where the initial samples include initial new-class samples and multiple initial base-class samples; inputting the initial samples into an initial image classification model, and determining a corresponding similarity matrix of the initial samples based on a preset composite metric function, where the similarity matrix is used to characterize the feature matching degree between the initial new-class samples and each initial base-class sample in multiple dimensions; the similarity matrix determined according to the composite metric function provides more position and distribution information, improving the accuracy of the feature distribution description between the two types of samples; then, based on the similarity matrix, target base-class samples are selected from multiple initial base-class samples; according to the target base-class samples and the corresponding similarity matrix of the target base-class samples, the initial new-class samples are dynamically corrected to make up for the problem of scarce effective training features in the initial new-class samples by accurately capturing more similar target base-class samples, obtaining target new-class samples; thus facilitating subsequent use of the target new-class samples with more real and accurate feature distributions to help improve the image classification accuracy of the trained model; then, using the target new-class samples to train the initial image classification model until the training stop condition is reached, obtaining a trained image classification model; after that, obtaining a target image to be classified, and inputting the target image into the trained image classification model to obtain a corresponding target classification result.
[0222] In addition, a Graphics Processing Unit (GPU) is also deployed in the image classification device. The GPU is a hardware specifically designed for efficiently processing image and graphics-related calculations. Its powerful parallel computing ability can significantly improve the image processing performance, thereby helping to improve the training efficiency of the initial image classification model and the image classification efficiency of the trained image classification model in actual applications.
[0223] Among them, the GPU can be:
[0224] (1) NVIDIA: such as Quadro / RTX A series, Tesla / V100 / H100 series, Jetson series, etc.;
[0225] (2) Intel: such as Arc series, Ponte Vecchio / Gaudi2 series, etc.
[0226] It should be noted that the specific type of the GPU can be set according to the actual situation. The above is only an example and does not represent a limitation in the embodiments of the present application.
[0227] The specific implementation manner of the image classification device is basically the same as the specific embodiments of the above image classification method, and will not be elaborated here.
[0228] The embodiments of the present application also provide an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above image classification method is implemented. The electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.
[0229] As Figure 7 shown, Figure 7 is a schematic diagram of the hardware structure of the electronic device provided by the embodiments of the present application. The electronic device includes:
[0230] A processor 401, which can be implemented by using a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided by the embodiments of the present application;
[0231] A memory 402, which can be implemented in the form of a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), etc. The memory 402 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 402, and the processor 401 is called to execute the image classification method of the embodiments of the present application;
[0232] An input / output interface 403, which is used to implement information input and output;
[0233] A communication interface 404, which is used to implement communication interaction between this device and other devices, and can implement communication through a wired manner (such as USB, network cable, etc.) or through a wireless manner (such as mobile network, WIFI, Bluetooth, etc.);
[0234] A bus 405, which transmits information between various components of the device (such as the processor 401, the memory 402, the input / output interface 403, and the communication interface 404);
[0235] Among them, the processor 401, the memory 402, the input / output interface 403, and the communication interface 404 are communicatively connected to each other inside the device through the bus 405.
[0236] An embodiment of the present application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above image classification method.
[0237] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include memories remotely located relative to the processor, and these remote memories can be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0238] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.
[0239] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or combine certain steps, or different steps.
[0240] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0241] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in the systems and devices can be implemented as software, firmware, hardware, and appropriate combinations thereof.
[0242] In the description of the present application and the above-mentioned drawings, terms such as "first", "second", "third", "fourth", etc. (if any) are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.
[0243] It should be understood that in the present application, "at least one (item)" means one or more, and "a plurality" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that there can be three relationships. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one) of the following" or its similar expression refers to any combination of these items, including any combination of single items (ones) or plural items (ones). For example, at least one (one) of a, b or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or plural.
[0244] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above-mentioned division of units is only a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.
[0245] The units described above as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0246] In addition, in each embodiment of the present application, the functional units may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated unit may be implemented in the form of hardware or in the form of a software functional unit.
[0247] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present application. The foregoing storage medium includes: various media that can store programs, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0248] The preferred embodiments of the embodiments of the present application have been described above with reference to the accompanying drawings, and thus do not limit the scope of rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of rights of the embodiments of the present application.
Claims
1. An image classification method, characterized in that: include: Acquire initial samples, where the initial samples include an initial new class sample and a plurality of initial base class samples; Inputting the initial sample into the initial image classification model, and determining the similarity matrix corresponding to the initial sample based on a preset composite metric function, wherein the similarity matrix is used to characterize the degree of feature matching between the initial new class sample and each of the initial base class samples in multiple dimensions; Based on the similarity matrix, selecting a target base class sample from a plurality of the initial base class samples; According to the target base class samples and the similarity matrix corresponding to the target base class samples, dynamically correct the initial new class samples to obtain target new class samples; The initial image classification model is trained using the target new class samples until a training stop condition is reached to obtain a trained image classification model; Obtain a target image to be classified, input the target image into the trained image classification model, and obtain a corresponding target classification result.
2. The image classification method according to claim 1, characterized in that: The initial new class sample includes a plurality of initial new class features, and each of the initial base class samples includes a plurality of initial base class features; The determining of the similarity matrix corresponding to the initial sample based on a preset composite metric function includes: Respectively determining a new class feature matrix of the initial new class samples and a base class feature matrix of the initial base class samples; Based on the multiple initial new class features and the multiple initial base class features, determine a first similarity value between the initial new class sample and each of the initial base class samples in a first dimension, and determine a second similarity value between the initial new class sample and each of the initial base class samples in a second dimension; Based on the new class feature matrix and the base class feature matrix, determining a third similarity value between the initial new class sample and each of the initial base class samples in a third dimension; According to the composite metric function provided with preset hyperparameters, the first similarity value, the second similarity value and the third similarity value are fused to obtain the similarity matrix.
3. The image classification method according to claim 2, characterized in that: The step of respectively determining the new class feature matrix of the initial new class sample and the base class feature matrix of the initial base class sample comprises: Transpose each of the initial new class features to obtain a corresponding new class transposed feature; Obtaining the new class feature matrix based on the cross product of each of the initial new class features and the corresponding new class transposed features; Based on the multiple initial base class features, obtaining the initial base class mean corresponding to each of the initial base class samples; Transposing each of the initial base class means to obtain a corresponding base class transposed mean; The base class feature matrix is obtained based on the cross product of each of the initial base class means and the corresponding base class transposed mean.
4. The image classification method according to claim 2, characterized in that: The step of determining a first similarity value between the initial new class sample and each of the initial base class samples in a first dimension based on the plurality of initial new class features and the plurality of initial base class features, and determining a second similarity value between the initial new class sample and each of the initial base class samples in a second dimension, comprises: Determine a first probability distribution set of the initial new class samples, and determine a second probability distribution set of the initial base class samples; According to the first probability distribution set and the second probability distribution set, a joint distribution set is formed; Determine the first similarity value based on the expected value of the distance between any of the initial new class features and any of the initial base class features under the joint distribution set; The second similarity value is determined based on a divergence value between the first probability distribution set and the second probability distribution set.
5. The image classification method according to claim 2, characterized in that: The new class characteristic matrix includes a plurality of first matrix elements, and the base class characteristic matrix includes a plurality of second matrix elements; The determining, based on the new class feature matrix and the base class feature matrix, a third similarity value between the initial new class sample and each of the initial base class samples in a third dimension comprises: Calculate the sum of squares between any element of the first matrix and any element of the second matrix to obtain a square value of the element; Performing a square root processing on the square value of the element to obtain a square root value of the element; The third similarity value is determined based on the square root of the element.
6. The image classification method according to claim 1, characterized in that: The selecting a target base class sample from a plurality of the initial base class samples based on the similarity matrix comprises: Based on a preset normalization operation operator, the similarity matrix of all the initial base class samples is normalized to obtain an adjusted similarity matrix, wherein the adjusted similarity matrix includes a plurality of third matrix elements; For any of the adjusted similarity matrices, if all of the third matrix elements are greater than a preset similarity threshold, the corresponding initial base class sample is determined to be the target base class sample.
7. The image classification method according to claim 3, characterized in that: The dynamically correcting the initial new class sample according to the target base class sample and the similarity matrix corresponding to the target base class sample to obtain the target new class sample includes: Based on the multiple initial base class features, obtaining the initial base class covariance value corresponding to each of the initial base class samples; Based on the similarity matrix, respectively updating the initial base class mean and the initial base class covariance value to obtain a target base class mean and a target base class covariance value; Based on the target base class mean and the target base class covariance value, the target base class samples are migrated to the initial new class samples to obtain the dynamically corrected target new class samples.
8. The image classification method according to claim 1, characterized in that: Before determining the similarity matrix corresponding to the initial sample based on the preset composite metric function, the method further includes: If the distribution hyperparameter value preset in the initial image classification model is not equal to the preset parameter threshold, logarithmically processing all the initial new class samples to obtain updated initial new class samples; If the preset distribution hyperparameter value of the initial image classification model is equal to the preset parameter threshold, all the initial new class samples are subjected to exponential processing to obtain updated initial new class samples.
9. A method for training an image classification model, characterized in that: include: Acquire initial samples, where the initial samples include an initial new class sample and a plurality of initial base class samples; Inputting the initial sample into the initial image classification model, and determining the similarity matrix corresponding to the initial sample based on a preset composite metric function, wherein the similarity matrix is used to characterize the degree of feature matching between the initial new class sample and each of the initial base class samples in multiple dimensions; Based on the similarity matrix, selecting a target base class sample from a plurality of the initial base class samples; According to the target base class samples and the similarity matrix corresponding to the target base class samples, dynamically correct the initial new class samples to obtain target new class samples; The initial image classification model is trained using the target new class samples until a training stop condition is reached to obtain a trained image classification model.
10. An image classification device, characterized in that: include: An acquisition module, used for acquiring initial samples, wherein the initial samples include an initial new class sample and a plurality of initial base class samples; A similarity matrix determination module, used for inputting the initial sample into the initial image classification model, and determining the similarity matrix corresponding to the initial sample based on a preset composite metric function, wherein the similarity matrix is used for characterizing the degree of feature matching between the initial new class sample and each of the initial base class samples in multiple dimensions; A target base class sample selection module, used for selecting a target base class sample from a plurality of the initial base class samples based on the similarity matrix; A correction module, used for dynamically correcting the initial new class sample according to the target base class sample and the similarity matrix corresponding to the target base class sample to obtain a target new class sample; A training module, used to train the initial image classification model using the target new class samples until a training stop condition is reached to obtain a trained image classification model; The application module is used to obtain the target image to be classified, input the target image into the trained image classification model, and obtain the corresponding target classification result.
11. An electronic device, characterized in that: The electronic device includes a memory and a processor, the memory stores a computer program, and the processor implements the image classification method according to any one of claims 1 to 8 or the image classification model training method according to claim 9 when executing the computer program.
12. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, it implements the image classification method described in any one of claims 1 to 8 or the image classification model training method described in claim 9.
Citation Information
Patent Citations
Three-dimensional human body behavior recognition method under small sample condition
CN115188068A
Small sample image classification method based on intra-class deviation migration
CN115631382A
Image processing apparatus, image processing method, and device-readable storage medium
JP2023120158A