Method and apparatus for training image classification model, electronic device, and storage medium

By performing multiple processing and feature extraction on the initial image samples, combining the loss function of covariance matrix and cosine similarity, the problem of false negative samples in unsupervised small sample learning is solved, and the accuracy and optimization efficiency of the image classification model are improved.

WO2025152066A1PCT designated stage expired Publication Date: 2025-07-24BOE TECHNOLOGY GROUP CO LTD +1
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/072773
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-01-17
Publication Date
2025-07-24

AI Technical Summary

Technical Problem

Existing unsupervised small sample learning methods fail to effectively solve the pseudo-negative sample problem in image classification, resulting in limited common semantic knowledge of the model learning between different instances, and using only instance-level consistency constraints lead to optimization difficulties and slowing convergence speed.

Method used

By performing multiple image processing on the initial image sample, the first image and the second image are obtained, and feature extraction is performed respectively, the first loss function and the second loss function are calculated, and the covariance matrix and cosine similarity are combined to enhance the attribute correlation consistency constraints, reduce background influence, and improve the image classification accuracy of the model.

Benefits of technology

It improves the accuracy of the image classification model, especially in small sample tasks, reduces the negative impact of false negative samples, and enhances the optimization ability and convergence speed of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024072773_24072025_PF_FP_ABST
    Figure CN2024072773_24072025_PF_FP_ABST
Patent Text Reader

Abstract

Provided are a method and apparatus for training an image classification model, an electronic device, and a storage medium. The training method comprises: performing multi-time image processing on an initial image sample to obtain a first image and a second image; performing feature extraction on the initial image sample to obtain an initial image feature; performing feature extraction on the first image to obtain a first image feature; performing feature extraction on the second image to obtain a second image feature; obtaining a first loss function and a second loss function on the basis of the initial image feature, the first image feature, and the second image feature; and training an image classification model on the basis of the first loss function and the second loss function. The image classification model is trained on the basis of the first loss function and the second loss function, and the association between the initial image sample, the first image, and the second image can be obtained from more dimensions, thus the accuracy of image classification by the trained image classification model can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Image classification model training method, device, electronic device and storage medium Technical Field

[0001] At least one embodiment of the present disclosure relates to a method, apparatus, electronic device, and storage medium for training an image classification model. Background Art

[0002] Few-Shot Learning (FSL) image classification algorithms are a method for classifying images by training a neural network model using a small number of samples. Since collecting large auxiliary sets of labeled data can be expensive and difficult, unsupervised few-shot learning (UFSL) using unlabeled auxiliary sets has attracted widespread attention and research.

[0003] Summary of the Invention

[0004] At least one embodiment of the present disclosure provides a method, apparatus, electronic device, and storage medium for training an image classification model.

[0005] At least one embodiment of the present disclosure provides a method for training an image classification model, comprising: performing multiple image processing on an initial image sample to obtain a first image and a second image; performing feature extraction on the initial image sample to obtain initial image features; performing feature extraction on the first image to obtain first image features; performing feature extraction on the second image to obtain second image features; obtaining a first loss function and a second loss function based on the initial image features, the first image features, and the second image features; and training the image classification model based on the first loss function and the second loss function.

[0006] For example, according to the training method described in an embodiment of the present disclosure, performing image processing on the initial image sample to obtain the first image includes: performing data augmentation on the initial image sample to obtain the first image.

[0007] For example, according to the training method described in an embodiment of the present disclosure, the training method further includes: taking the initial image sample and the first image as a first positive sample pair; and obtaining the first loss function based on the first positive sample pair.

[0008] For example, according to the training method described in an embodiment of the present disclosure, the initial image feature includes a first attribute feature, and the first image feature includes a second attribute feature; the first loss function is obtained based on the first attribute feature and the second attribute feature.

[0009] For example, according to the training method described in an embodiment of the present disclosure, the first attribute feature is the semantic feature channel dimension of the initial image sample; and the second attribute feature is the semantic feature channel dimension of the first image feature.

[0010] For example, according to the training method described in an embodiment of the present disclosure, the first loss function is obtained based on the first attribute feature and the second attribute feature, including: obtaining a covariance matrix based on the semantic feature channel dimension of the initial image sample and the semantic feature channel dimension of the first image feature; obtaining the first loss function based on the covariance matrix.

[0011] For example, according to the training method described in an embodiment of the present disclosure, performing image processing on the initial image sample to obtain the second image includes: performing target area extraction on the initial image sample to obtain the second image.

[0012] For example, according to the training method described in an embodiment of the present disclosure, the target area extraction of the initial image sample to obtain the second image includes: performing salient area detection on the initial image sample to obtain a grayscale image; and binarizing the grayscale image to obtain the second image.

[0013] For example, according to the training method described in an embodiment of the present disclosure, the binarization processing of the grayscale image to obtain the second image includes: performing the binarization processing on the grayscale image to obtain a binarization result; multiplying the binarization result with the initial image sample to obtain the second image; the second image includes the target area.

[0014] For example, according to the training method described in an embodiment of the present disclosure, the image processing of the initial image sample to obtain the first image includes: performing data augmentation processing on the initial image sample to obtain the first image; the training method also includes: taking the first image and the second image corresponding to the first image as a second positive sample pair; and obtaining the second loss function based on the second positive sample pair.

[0015] For example, according to the training method described in the embodiment of the present disclosure, the covariance matrix is ​​expressed as: C i,q =Z i,q ×Z i T

[0016] Among them, Z i,q represents the semantic feature channel dimension of the i-th image in the first image, Z i represents the semantic feature channel dimension of the i-th image in the initial image sample.

[0017] For example, according to the training method described in an embodiment of the present disclosure, the first loss function is expressed as:

[0018] Among them, l AUC (i,q) represents the normalized classification score, d(*) represents the cosine similarity, represents the semantic feature channel dimension of the i-th image in the second image.

[0019] For example, according to the training method described in an embodiment of the present disclosure, the second loss function is expressed as:

[0020] Among them, l OAS (i,q) represents the normalized classification score, d(*) represents the cosine similarity, f(*) represents the feature embedding function of the image, and x i represents the i-th image in the initial image sample, Among them, g(x i ) represents the binarization result, x i,q represents the second positive sample pair.

[0021] For example, according to the training method described in an embodiment of the present disclosure, the training of the image classification model based on the first loss function and the second loss function includes: obtaining a total loss function based on the first loss function and the second loss function; and training the image classification model based on the total loss function.

[0022] For example, according to the training method described in the embodiment of the present disclosure, the total loss function is expressed as:

[0023] Wherein, N represents the number of the initial image samples, M represents the number of the first images, and l AUC (i,q) represents the first loss function, l OAS (i,q) represents the second loss function, and λ represents a preset parameter.

[0024] At least one embodiment of the present disclosure provides an image classification method, comprising: obtaining an image to be classified; classifying the image to be classified using an image classification model to determine a category label of the image to be classified; wherein at least a portion of the image classification model is trained according to the above-mentioned training method.

[0025] At least one embodiment of the present disclosure provides a training device for an image classification model, comprising: an image processing module, configured to perform multiple image processing on an initial image sample to obtain a first image and a second image; a feature extraction module, configured to perform feature extraction on the initial image sample to obtain initial image features; configured to perform feature extraction on the first image to obtain first image features; and configured to perform feature extraction on the second image to obtain second image features; an analysis module, configured to obtain a first loss function and a second loss function based on the initial image features, the first image features, and the second image features; and a training module, configured to train the image classification model based on the first loss function and the second loss function to obtain a trained image classification model.

[0026] At least one embodiment of the present disclosure provides a training device, comprising: one or more memories, which non-transitorily store computer-executable instructions; and one or more processors, which are configured to execute the computer-executable instructions, wherein the computer-executable instructions are implemented according to the above-mentioned training method when executed by the one or more processors.

[0027] At least one embodiment of the present disclosure provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions implement the above-mentioned training method when executed by a processor.

[0028] At least one embodiment of the present disclosure provides a training method, device, electronic device, and storage medium for an image classification model, which obtains a first image and a second image corresponding to the initial image sample by performing multiple image processing on the initial image sample. By performing feature extraction on the initial image sample, the first image, and the second image, the initial image features of the initial image sample, as well as the first image features of the first image and the second image features obtained after image processing, can be obtained. A first loss function and a second loss function are determined based on the initial image features, the first image features, and the second image features. The image classification model is trained based on the first loss function and the second loss function, and the association between the initial image sample, the first image, and the second image can be obtained from more dimensions, thereby improving the image classification accuracy of the trained image classification model. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the drawings of the embodiments will be briefly introduced below. Obviously, the drawings in the following description only relate to some embodiments of the present disclosure, rather than limiting the present disclosure.

[0030] FIG1A is a schematic diagram of a support set.

[0031] FIG1B is a schematic diagram showing a false negative sample.

[0032] FIG2 is a flow chart of a method for training an image classification model provided by at least one embodiment of the present disclosure.

[0033] FIG3 is a schematic diagram of an image classification model provided by at least one embodiment of the present disclosure.

[0034] FIG4 is a schematic diagram of the attribute relationship between the first positive sample pairs provided by at least one embodiment of the present disclosure.

[0035] FIG5 is a flowchart of an image classification method provided by at least one embodiment of the present disclosure.

[0036] FIG6 is a schematic block diagram of a training apparatus for an image classification model provided by at least one embodiment of the present disclosure.

[0037] FIG7 is a schematic block diagram of an image classification apparatus provided by at least one embodiment of the present disclosure.

[0038] FIG8 is a schematic block diagram of an electronic device provided by at least one embodiment of the present disclosure.

[0039] FIG9 is a schematic block diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure. DETAILED DESCRIPTION

[0040] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions of the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the described embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.

[0041] Unless otherwise defined, technical or scientific terms used in this disclosure should have the ordinary meanings understood by people with ordinary skills in the field to which this disclosure belongs. The words "first", "second" and similar terms used in this disclosure do not indicate any order, quantity or importance, but are simply used to distinguish different components. The words "include" or "comprising" and similar terms mean that the elements or objects preceding the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects.

[0042] The features such as "perpendicular", "parallel" and "same" used in this disclosure include the features such as "perpendicular", "parallel" and "same" in the strict sense, as well as the cases where "approximately perpendicular", "approximately parallel" and "approximately the same" include certain errors, taking into account the errors associated with the measurement and the measurement of specific quantities (that is, the limitations of the measurement system), and are expressed as being within the acceptable deviation range for a specific value determined by a person of ordinary skill in the art. The "center" in the embodiments of the present disclosure can include a position strictly at the geometric center and a position approximately at the center of a small area around the geometric center. For example, "approximately" can mean within one or more standard deviations, or within 10% or 5% of the value.

[0043] Small sample image classification refers to accurately classifying images in the query set based on the existing support set. The support set is similar to the training set, including N classification labels, each label has K images. The query set is similar to the test set, including q unclassified images. The small sample image classification task based on meta-learning is used as an example. The model first uses a data set D (basic class) with sufficient data. base ={(x i ,y i )}. Among them, x i Represents the image, y i Represents the category label corresponding to an image. After pre-training, the model is transferred to the downstream small sample classification task This task only contains the new class D novel At the same time, there is no overlap between the basic class and the new class, that is, satisfying The downstream meta-learning process for new classes requires constructing a training set, which consists of a support set and a query set. The support set is represented as The query set is represented as The support set S includes N categories, each category has K images, and the query set Q includes N categories, each category has M images. s and y q Represents the image x s and image x q The category label of y. q It is only used to calculate the evaluation criteria. Therefore, the above task can be called an "N-way K-shot" (pseudo N-classification K-sample) image classification task. When the value of K is very small (for example, K < 10), the task is a small-shot image classification task.

[0044] In supervised few-shot learning image classification algorithms, model training is often performed using an auxiliary dataset with a large amount of labeled data that is disjoint from the target dataset. However, large labels are expensive, require significant human effort, and are sometimes difficult to obtain, such as in certain rare medical fields. Furthermore, labeling a large dataset (even if it is disjoint from the target dataset) is practically intractable in addressing the problem of few labeled examples in few-shot image classification. Therefore, unsupervised few-shot learning has been proposed.

[0045] Unlike supervised few-shot learning, unsupervised few-shot learning does not require any labeled data for training. It aims to use a large-scale unlabeled auxiliary dataset to train a network that can be transferred to a model that recognizes new classes with only a few labeled samples. Specifically, in supervised few-shot learning, the dataset D base ={(x i ,y i )} is fully labeled. In the pre-training (or meta-training) stage of unsupervised small sample learning, there is only an unlabeled dataset D base ={(x i )} is available. Therefore, the primary challenge in unsupervised few-shot learning is how to obtain pseudo-labels. The key lies in learning useful representations from unlabeled data and generalizing them to new classes with limited labeled data. Once pseudo-labels are obtained, the model can be trained using the standard learning paradigm of supervised few-shot learning, using the contextual training mechanism.

[0046] For example, some unsupervised small-shot learning is based on example discrimination tasks, taking each original image as a pseudo class. In addition, data augmentation techniques are used to augment the anchor image to obtain multiple transformed samples corresponding to each pseudo class. That is, the changed image obtained after data augmentation is also considered to be a sample of the pseudo class corresponding to the original image. In this way, with the help of pseudo labels, the contextual training mechanism is utilized in supervised small-shot learning to train the model for unsupervised small-shot learning. In other words, thousands of N-way K-shot small-shot tasks can be constructed from unlabeled auxiliary sets and further used for training. For example, given a small batch of N unlabeled images, an N-way 1-shot small-shot training classification task can be obtained by M times of data augmentation. Among them, the support set Query set Q = {(x q ,y q )} N*M , N represents N pseudo-classes. Then, some existing small sample learning methods, such as Prototypical Networks, can be used to train the entire model.

[0047] FIG1A is a schematic diagram of a support set, and FIG1B is a schematic diagram of a false negative sample.

[0048] During the study, the inventors of the present application found that the problem of false negative samples is rarely considered in the current research on unsupervised small-shot learning. In the instance discrimination task, as shown in FIG1A , in the support set of 5-way 1-shot, each original image is regarded as a separate pseudo-class, and different original images will be regarded as different pseudo-classes. As a result, as shown in FIG1B , two original images sharing the same semantic category will be roughly regarded as two different pseudo-classes, resulting in the problem of false negative samples. On the one hand, this problem will weaken the unsupervised small-shot learning model to learn the common semantic knowledge of the true category between different instances. On the other hand, when this problem occurs in mini-batch data during training (which is usually unavoidable), the confusion of false negative samples will make model optimization difficult and slow down the convergence of the model.

[0049] In addition, the inventors of this application also found that unsupervised small-sample learning models usually only use instance-level consistency to constrain the similarity between the feature vectors of two positive samples, which also imposes certain limitations on the optimization of the model.

[0050] At least one embodiment of the present disclosure provides a method for training an image classification model, comprising: performing multiple image processing on an initial image sample to obtain a first image and a second image; performing feature extraction on the initial image sample to obtain initial image features; performing feature extraction on the first image to obtain first image features; performing feature extraction on the second image to obtain second image features; obtaining a first loss function and a second loss function based on the initial image features, the first image features, and the second image features; and training the image classification model based on the first loss function and the second loss function.

[0051] At least one embodiment of the present disclosure provides an image classification method, comprising: acquiring an image to be classified; performing classification processing on the image to be classified using an image classification model to determine a category label of the image to be classified; wherein at least a portion of the image classification model is trained according to the above-mentioned training method.

[0052] At least one embodiment of the present disclosure provides a training device for an image classification model, comprising: an image processing module, configured to perform multiple image processing on an initial image sample to obtain a first image and a second image; a feature extraction module, configured to perform feature extraction on the initial image sample to obtain initial image features; configured to perform feature extraction on the first image to obtain first image features; and configured to perform feature extraction on the second image to obtain second image features; an analysis module, configured to obtain a first loss function and a second loss function based on the initial image features, the first image features, and the second image features; and a training module, configured to train the image classification model based on the first loss function and the second loss function to obtain a trained image classification model.

[0053] At least one embodiment of the present disclosure provides a training device, comprising: one or more memories, which non-transitorily store computer-executable instructions; and one or more processors, which are configured to execute the computer-executable instructions, wherein the computer-executable instructions are implemented according to the above-mentioned training method when executed by the one or more processors.

[0054] At least one embodiment of the present disclosure provides a non-transitory computer-readable storage medium, wherein the non-transitory computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, the above-mentioned training method is implemented.

[0055] At least one embodiment of the present disclosure provides a training method, device, electronic device, and storage medium for an image classification model, which obtains a first image and a second image corresponding to the initial image sample by performing multiple image processing on the initial image sample. By performing feature extraction on the initial image sample, the first image, and the second image, the initial image features of the initial image sample, as well as the first image features of the first image and the second image features obtained after image processing, can be obtained. A first loss function and a second loss function are determined based on the initial image features, the first image features, and the second image features. The image classification model is trained based on the first loss function and the second loss function, and the association between the initial image sample, the first image, and the second image can be obtained from more dimensions, thereby improving the image classification accuracy of the trained image classification model.

[0056] The following describes the image classification model training method, device, electronic device, and storage medium with reference to the accompanying drawings and through some embodiments.

[0057] Figure 2 is a flow chart of a method for training an image classification model according to at least one embodiment of the present disclosure. Figure 3 is a schematic diagram of an image classification model according to at least one embodiment of the present disclosure.

[0058] As shown in FIG. 2 and FIG. 3 , the image classification model training method provided by at least one embodiment of the present disclosure includes the following steps S110 to S160 .

[0059] Step S110: performing multiple image processing on the initial image sample to obtain a first image and a second image;

[0060] Step S120: extracting features from the initial image sample to obtain initial image features;

[0061] Step S130: extracting features from the first image to obtain first image features;

[0062] Step S140: extracting features from the second image to obtain second image features;

[0063] Step S150: obtaining a first loss function and a second loss function based on the initial image feature, the first image feature, and the second image feature;

[0064] Step S160: training an image classification model based on the first loss function and the second loss function.

[0065] As shown in Figures 2 and 3, the training method of the image classification model provided by the embodiment of the present disclosure obtains a first image and a second image corresponding to the initial image sample respectively by performing multiple image processing on the initial image sample. By performing feature extraction on the initial image sample, the first image, and the second image respectively, the initial image features of the initial image sample, as well as the first image features of the first image and the second image features of the second image obtained after image processing can be obtained. The first loss function and the second loss function are determined according to the initial image features, the first image features, and the second image features, and the image classification model is trained based on the first loss function and the second loss function, so that the association between the initial image sample, the first image, and the second image can be obtained from more dimensions, thereby improving the image classification accuracy of the trained image classification model.

[0066] As shown in Figures 2 and 3, in step S110, for example, the initial image sample can be an image obtained by photographing or scanning an object. For example, the initial image sample can be various types of images. The object can be, for example, a lesion site or a product defect site, and the initial image sample can be an image of the lesion site or a product defect site.

[0067] As shown in Figures 2 and 3, for example, the initial image sample can be an image captured by an image acquisition device (e.g., a digital camera or a camera on a mobile phone). The initial image sample can be a grayscale image, a black and white image, or a color image. It should be noted that the initial image sample refers to a form of visually presenting an object, such as a picture of the object. For another example, the initial image sample can also be obtained by scanning or other methods. For example, the initial image sample can be different types of medical images, such as magnetic resonance imaging images and X-ray computed tomography (CT) images.

[0068] As shown in FIG. 2 and FIG. 3 , for example, in steps S120 to S140 , feature extraction processing may be performed on the input initial image sample, the first sample, and the second sample by an encoder or a feature extractor.

[0069] As shown in Figures 2 and 3, for example, in step S150, the first and second loss functions can represent the difference between the first and / or second images and the initial image. During the model training phase, after each batch of training data is fed into the model, a predicted value is output through forward propagation. The loss function calculates the difference between the predicted value and the true value, that is, the loss value. After obtaining the loss value, the model updates various parameters through backpropagation to reduce the loss between the true value and the predicted value, so that the predicted value generated by the model is closer to the true value, thereby achieving the purpose of learning.

[0070] As shown in Figures 2 and 3, for example, in step S160, the parameters in the image classification model can be corrected according to the loss value of the first loss function and the loss value of the second loss function. When the first loss function and the second loss function meet the predetermined conditions, a trained image classification model is obtained. When the first loss function and the second loss function do not meet the predetermined conditions, the above training process is repeated. For example, the predetermined condition may be that the loss of the first loss function converges and the loss of the second loss function converges. That is, the loss value of the first loss function and the loss value of the second loss function no longer decrease significantly. For example, the predetermined condition may be that the number of training times or training cycles reaches a predetermined number. For example, the predetermined number may be millions.

[0071] As shown in Figures 2 and 3, in some examples, performing multiple image processing on the initial image sample to obtain the first image and the second image, that is, step S110, may include: performing data augmentation on the initial image sample to obtain the first image. For example, data augmentation, also known as data enhancement, refers to making limited data produce the equivalent value of more data without substantially increasing the data. For example, the initial image sample can be cropped, the contrast and saturation of the color of the initial image sample can be transformed, the initial image sample can be grayscale transformed, the initial image sample can be pixel-filled, the initial image sample can be randomly affine transformed, the initial image sample can be randomly horizontally flipped, the initial image sample can be randomly rotated, the initial image sample can be randomly vertically flipped, the initial image sample can be scaled or distorted, etc. For example, N unlabeled initial image samples (unlabeled images) in a small sample can be given, and after performing M data augmentation on each initial image sample, M transformed query images (query images) are generated, that is, M first images are obtained.

[0072] As shown in FIG2 , in some examples, the training method further includes taking the initial image sample and the first image as a first positive sample pair; and obtaining a first loss function based on the first positive sample pair. For example, based on the first positive sample pair formed by the initial image sample and the first image, the first loss function is obtained to represent the difference between the initial image sample and the first image.

[0073] In some unsupervised few-shot learning and semi-supervised few-shot learning methods, classification is performed based on instance-level consistency, that is, based on feature consistency of one-dimensional labels (such as one-dimensional vectors or one-dimensional scalar values). This approach ignores the consistency between features and sample attributes.

[0074] In some examples, the initial image features include first attribute features, and the first image features include second attribute features; a first loss function is obtained based on the first attribute features and the second attribute features. For example, a feature (eigenvalue) refers to an N-dimensional array output by an N-dimensional convolutional layer, such as a representation of an input image at a certain level in spatial dimensions (width and height). For example, an attribute is used to represent a feature of a data object. The disclosed embodiment also calculates more intermediate layer information and extracts sample attribute relationships between more subjects. Thus, the attribute-related consistency of the initial image sample and the query image is taken into account, thereby enhancing the instance-level consistency constraint and simultaneously expanding the inconsistency between negative sample pairs. Thus, the attribute-related consistency can be supplemented as instance-level consistency, thereby promoting the representation learning of the image classification model.

[0075] FIG4 is a schematic diagram of the attribute relationship between the first positive sample pairs provided by at least one embodiment of the present disclosure.

[0076] As shown in Figure 4, each unlabeled initial image sample and the first image (i.e., the image obtained after data augmentation of the initial image sample) are considered as a class, and its foreground is used as a positive sample. Within the same training batch, the data augmented and foreground samples of other samples are used as negative samples for self-supervised contrastive learning training. At the same time, for a three-dimensional feature matrix (c, h, w) (channel, height, weight) obtained by the encoder (feature extractor), different channels represent different attributes of the sample. Based on this, it is converted into a two-dimensional feature matrix (channel, h*w), and the covariance matrix is ​​calculated with its transposed matrix to obtain a two-dimensional attribute relationship matrix (c, c), thereby constraining the attribute consistency of the first image after sample augmentation and the initial image sample.

[0077] As shown in Figures 3 and 4, in some examples, the first attribute feature is the channel dimension of the semantic features of the initial image sample, and the second attribute feature is the channel dimension of the semantic features of the first image features. By extracting the semantic features of the initial image sample and the first image, semantic information of the initial image sample and the first image, such as the category, structure, and spatial layout of objects, can be learned, which facilitates image recognition, image segmentation, and scene understanding. The channel dimension of the semantic feature refers to the position information of the corresponding feature value of the attribute feature of an object in an N-dimensional array. For example, the attributes of the initial image sample and the first image can be represented by the channel of the embedding, which is the output of the neural network. For example, the channel dimension is the mapping of attribute types in the image matrix. The channel dimension can include multiple information, such as image location information, texture features, edge features, color features, semantic features, abstract features, etc., that is, the attributes are expressed through the channels of the image matrix.

[0078] As shown in Figures 3 and 4, in some examples, a first loss function is obtained based on the first attribute feature and the second attribute feature, including: obtaining a covariance matrix based on the semantic feature channel dimension of the initial image sample and the semantic feature channel dimension of the first image feature; obtaining the first loss function based on the covariance matrix, thereby performing classification by calculating the relationship between the attributes of the first positive sample pair. For example, h (height) is the height, which represents the number of pixels in the vertical dimension of the image, w (width) is the width, which represents the number of pixels in the horizontal dimension of the image, and c (channel) is the number of channels, which represents the number of channels in an image.

[0079] As shown in Figures 3 and 4, in some examples, the covariance matrix is ​​expressed as: C i,q =Z i,q ×Z iT

[0080] Among them, Z i,q represents the semantic feature channel dimension of the i-th image in the first image, Z i Represents the semantic feature channel dimension of the i-th image in the initial image sample. The first attribute feature Z i The shape of the transpose is (h*w, c), the second attribute feature Z i,q The shape of is (c, h*w). From this we can get that the first attribute feature Z i and the second attribute feature Z i,q The covariance matrix C i,q The shape of is (c, c), which represents the attribute relationship of the first positive sample pair.

[0081] For example, in combination with the following example, the covariance matrix between the first image (query image) and the second image (foreground image) can also be calculated, and the formula can be expressed as:

[0082] in, represents the semantic feature channel dimension of the i-th image in the second image. It can be understood that C i,q It means xi ,q The corresponding covariance matrix is, It means The corresponding covariance matrix.

[0083] In some examples, the covariance matrix C i,q After flattening, the cosine similarity is calculated in the classifier, so that the first loss function that can represent the consistency loss of the attribute relationship can be obtained. For example, the first loss function is the uniformity classification loss function. The first loss function is expressed as:

[0084] Among them, l AUC (i,q) represents the normalized classification score, d(*) represents the cosine similarity, Represents the semantic feature channel dimension of the i-th image in the second image.

[0085] For example, in some unsupervised small sample learning algorithms, such as UMTRA, AAL, and ProtoTransfer, the loss is calculated by calculating the similarity between features, so that the positive sample pair (for example, the original image and the query image after data augmentation) is more similar in feature representation and farther away from the negative sample in the feature space. For example, the cross entropy loss is calculated for each query image. The cross entropy loss function describes the similarity between the actual output probability and the expected output probability, that is, the smaller the cross entropy value, the closer the two probability distributions are. For example, the formula can be expressed as:

[0086] Among them, l(i,q) represents the normalized classification score, d(*) represents the cosine similarity, f(*) represents the feature embedding function of the image, and x i represents the i-th image in the original image, x i,q represents the enhanced query image.

[0087] Often, samples from different pseudo-classes may come from the same true class. Therefore, the differences between classes may be small, potentially misleading the learning of the encoder (the feature embedding function of the image). For a single image (one class), the gap between the foreground and background is large, and the foreground and background typically represent different semantic information and should be farther apart in feature space. This results in the fact that in unsupervised few-shot learning, the intra-class differences are often greater than the inter-class differences.

[0088] As shown in Figures 2 and 3, in some examples, the initial image sample is processed to obtain a second image, that is, step S110, including: extracting a target area from the initial image sample to obtain the second image. For example, the target area refers to an area in the image that is of interest or focused on, such as the area where the main target in the image is located, or the area where the object to be extracted or identified is located during the image processing process. For example, the second image is a foreground image. By removing the background and extracting the foreground, the initial image sample (anchor image) can be enhanced. As a result, the diversity of the anchor image and the positive sample (the image obtained after enhancement from the same anchor image) is increased, which is conducive to the learning of unsupervised representations. Moreover, through the second image, the influence of the background is eliminated, and the image classification model will be constrained to learn more semantic features of the true category, so that the image classification model focuses on the target object itself. In addition, compared with the false negative sample obtained by only performing data augmentation, the second image obtained after removing the background is very different from it, so that the negative impact of the false negative sample can be significantly reduced.

[0089] As shown in Figures 2 and 3, in some examples, target region extraction is performed on the initial image sample to obtain a second image, including: performing salient region detection on the initial image sample to obtain a grayscale image. For example, converting the initial image sample from a color image to a grayscale image can simplify the image information, thereby facilitating subsequent image processing and analysis. For example, a salient object detection (SOD) model can be used to perform salient region detection on the initial image sample to distinguish the most obvious area in the image and obtain more effective information. For example, the salient object detection model can be a pre-trained and fixed model. For example, the network structure of the BASNet algorithm can be used. For example, the initial image sample can be used as a support image, and the support image can be input into a fixed salient target detection model.

[0090] In some examples, a grayscale image is binarized to produce a second image. Binarization involves converting the grayscale values ​​of pixels in a grayscale image to either 0 or 255, essentially converting the image into only two colors: black and white. This process typically involves setting a threshold, with pixels below the threshold set to 0 (black) and pixels above the threshold set to 255 (white). This allows the foreground and background of the initial image sample to be distinguished, resulting in a second image.

[0091] In some examples, binarizing a grayscale image to obtain a second image includes: binarizing the grayscale image to obtain a binarized result; and multiplying the binarized result with samples of the initial image to obtain a second image; wherein the second image includes the target area. For example, after binarizing the grayscale image, the foreground area of ​​the image is assigned a value of 1 and the background area is assigned a value of 0, thereby obtaining a binarized result. Thus, multiplying the binarized result with the samples of the initial image yields the second image, i.e., a foreground image including the target area.

[0092] In some examples, a first image and a second image corresponding to the first image are used as a second positive sample pair; a second loss function is derived based on the second positive sample pair. In the disclosed embodiments, the second image is obtained by extracting the target region, so that the first image and the second image corresponding to the first image are used as a positive sample pair. This enhances the distance between the second image (i.e., the foreground image) and the corresponding enhanced first image (i.e., the query image), while maintaining the distance between other foreground images, thereby reducing the impact of the background on the classification results. For example, the second loss function is a variance classification loss function.

[0093] In some examples, the second loss function is expressed as:

[0094] Among them, l OAS (i,q) represents the normalized classification score, d(*) represents the cosine similarity, f(*) represents the feature embedding function of the image, and x i represents the i-th image in the initial image sample, Among them, g(x i ) represents the binarization result, x i,q Represents the second positive sample pair.

[0095] In some examples, training the image classification model based on the first loss function and the second loss function, i.e., step S160, includes: obtaining a total loss function based on the first loss function and the second loss function; and training the image classification model based on the total loss function. For example, the total loss function of the image classification model can be obtained based on the first loss function and the second loss function in the aforementioned example to train the image classification model.

[0096] In some examples, the total loss function is expressed as:

[0097] Where N represents the number of initial image samples, M represents the number of first images, and l AUC (i,q) represents the first loss function, l OAS (i,q) represents the second loss function, and λ represents a preset parameter. For example, λ represents a hyperparameter. For example, λ can be set to 0.5.

[0098] For example, in the Self-Supervised Learning (SSL) algorithm, since there are no labels, each sample in the initial image sample is regarded as a class (pseudo-class). The query image obtained after data augmentation is regarded as a positive sample pair of the same type as the initial image sample, while other samples in the same batch or samples in the memory library are treated as negative samples. In the SimCLR model architecture and the UniSiam model architecture, consistency regularization can be used to calculate the difference between positive and negative samples. Consistency regularization is an operation in semi-supervised learning in addition to pseudo-labeling. It assumes that if a certain perturbation is made to the input, the model predictions should be similar. This loss calculation method can make the positive and negative samples farther apart in the feature space, thereby avoiding the model from suffering significant dimensionality collapse. The loss function can be expressed as:

[0099] Among them, Neg(z i ) represents sample z i negative samples, and λ represents the weighted hyperparameter.

[0100] In other embodiments, the first loss function and the second loss function can also be applied to self-supervised learning algorithms. For example, in the SimCLR model architecture and the UniSiam model architecture, an asymmetric alignment term is usually used to calculate the consistency between positive samples, so that the positive sample pairs are closer in the feature space. After substituting the first loss function and the second loss function, the total loss function of the SimCLR model can be obtained, which is expressed as follows:

[0101] For example, Simple Siamese Self-Attention (SimSiam for short) is another deep learning model for self-supervised learning, which is used to train deep neural networks without using labeled data. Unlike unsupervised small-sample learning, SimSiam usually uses a stopping gradient, which calculates the total self-supervised loss by exchanging the results of the projection head (Projection Head) and the prediction head (Prediction Head) of two positive sample pairs. The projection head is a small neural network that maps features to a low-dimensional space to produce a projection vector; the prediction head is another small neural network that maps the vector back to the same dimension of the original feature space to generate a prediction vector. The total loss function formula of SimSiam can be expressed as:

[0102] Among them, z i represents the projection head, p i Represents the prediction head.

[0103] The following is based on specific experimental results to illustrate the improvement in classification and recognition accuracy of the image classification model trained according to the training method of the image classification model disclosed in the present invention.

[0104] To make a fair comparison, we only conducted training experiments on the basic dataset without using any additional (unlabeled) data. Specifically, we conducted extensive experiments including multiple ablation studies to demonstrate the effectiveness of each component of the proposed method.

[0105] The experiments were based on two datasets: miniImageNet and tieredImageNet. MiniImageNet is a small dataset for image recognition and a subset of ImageNet. It consists of 100 categories, each with 600 images of 84x84 pixels. The inventors split the 100 classes into 64 for training, 16 for validation, and 20 for testing. TieredImageNet is a large-scale, high-resolution, hierarchical image classification dataset and another subset of ImageNet. TieredImageNet has a higher resolution and richer categories, including not only common objects but also some rare or difficult-to-understand categories. TieredImageNet contains a total of 779,165 images from 608 classes. Each class has 600 training images, 200 validation images, and 500 test images of 224x224 pixels. The 608 classes in tieredImageNet were initially grouped into 34 high-level categories. 351 classes in 34 categories are used for training, 97 classes in 6 categories are used for validation, and 160 classes in 8 categories are used for testing.

[0106] During training, the inventors used a four-layer convolutional neural network and a residual neural network (ResNet-18) as the backbone network. For the four-layer convolutional neural network, each layer included a 64-filter (3×3 kernel) convolution layer, a batch normalization layer, a nonlinear activation (ReLU) layer, and a 2×2 maximum pooling layer. For ResNet-18, the inventors chose Unisiam as the baseline and used the same projection head and prediction head as Unisiam and self-supervised learning work, and maintained the same MLP settings. For the four-layer convolutional neural network, the inventors set the initial learning rate to 0.001, and for the residual network, the inventors set the initial learning rate to 0.01.

[0107] During the testing process, the inventors evaluated the performance of the training method disclosed herein based on the two aforementioned datasets (miniImageNet and tieredImageNet), and compared it with some other learning methods (e.g., unsupervised small-sample learning methods, self-supervised learning methods, and supervised small-sample learning methods).

[0108] The training method disclosed in this disclosure achieved an accuracy of (48.61±0.61)% on the 5-way-1shot task using ProtoTransfer as the method and Conc-4-64 as the backbone network, and an accuracy of (63.49±0.59)% on the 5-way-5shot task. The method also achieved an accuracy of (64.63±0.36)% on the 5-way-1shot task using UniSiam as the method and ResNet-18 as the backbone network, and an accuracy of (81.35±0.25)% on the 5-way-5shot task. The method also achieved an accuracy of (66.94±0.41)% on the 5-way-1shot task using UniSiam as the method and ResNet-18 as the backbone network, and an accuracy of (82.52±0.31)% on the 5-way-5shot task using UniSiam as the method and ResNet-18 as the backbone network.

[0109] For example, on the 5-way-1-shot task on miniImageNet, an accuracy gain of 1.5% was achieved. Moreover, compared with metric-based few-shot learning methods and metric-based small sample learning methods (e.g., protoTransfer), the present disclosure can improve performance by using cosine distance instead of Euclidean distance as a metric. In addition, experimental results show that since the 5-way 1-shot task is more challenging, the training method of the present disclosure uses a more complex feature representation, which is more helpful for more complex tasks. Therefore, the accuracy is improved more on the 5-way 1-shot task than on the 5-way 5-shot task.

[0110] The inventors of the present disclosure also compared the effects of the training method of the present disclosure, the unsupervised small sample learning method and the self-supervised small sample learning method on small sample tasks. The experimental results show that since the training method of the self-supervised learning method is not targeted at specific downstream tasks, the downstream tasks are more applicable, while the training method of the present disclosure is more suitable for tasks with very small sample sizes. In addition, the self-supervised learning method usually requires a larger image size, a larger batch size and a deeper network structure to achieve better performance. In contrast, the training method provided by the present disclosure usually only requires the use of fewer resources and is lower in cost.

[0111] Furthermore, adding the first and / or second loss functions to the UniSiam and ProtoTransfer baselines on miniImageNet improves neural network performance. For example, on a 5-way 1-shot task with a 224-image input size and UniSiam as the method, a ResNet-18 as the backbone, and an accuracy of 63.26% on miniImageNet, the accuracy increases to 64.29% after adding the first loss function, 63.82% after adding the second loss function, and 64.34% after adding both the first and second loss functions. On a 5-way 1-shot task with a 84-image input size and a Conv-4-64 as the backbone, the accuracy is 47.12%. After adding the first loss function, the accuracy rate reached 48.32%, after adding the second loss function, the accuracy rate reached 48.5%, and after adding the first and second loss functions at the same time, the accuracy rate reached 48.61%.

[0112] FIG5 is a flowchart of an image classification method provided by at least one embodiment of the present disclosure.

[0113] 5 , at least one embodiment of the present disclosure further provides an image classification method, including the following steps S210 to S220 .

[0114] Step S210: Obtain an image to be classified. For example, the image to be classified can be an image obtained by photographing or scanning an object. For example, the image to be classified can be various types of images. The object can be, for example, a lesion site or a product defect site. Thus, the image to be classified can be an image of a lesion site or a product defect site.

[0115] Step S220: Classify the image to be classified using the image classification model to determine a category label for the image to be classified; wherein at least a portion of the image classification model is trained according to any of the training methods described in the aforementioned examples. The training method for the image classification model is as described above and will not be further described here. For example, after training the image classification model to be trained according to the aforementioned method, the trained image classification model can be used to classify the image to be classified to obtain a category label for the image to be classified.

[0116] FIG6 is a schematic block diagram of a training apparatus for an image classification model provided by at least one embodiment of the present disclosure.

[0117] As shown in Figure 6, at least one embodiment of the present disclosure further provides a training device 300 for an image classification model. The training device 300 for an image classification model may include an image processing module 301, a feature extraction module 302, an analysis module 303, and a training module 304. These components are interconnected via a bus system and / or other forms of connection mechanisms (not shown). For example, these modules can be implemented by hardware (e.g., circuit) modules, software modules, or any combination of the two. The following embodiments are the same and will not be repeated here. For example, these units can be implemented by a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a field programmable gate array (FPGA), or other forms of processing units with data processing capabilities and / or instruction execution capabilities, as well as corresponding computer instructions. It should be noted that the components and structures of the training device 300 shown in Figure 6 are exemplary only and not restrictive. The training device 300 may also have other components and structures as needed.

[0118] As shown in FIG6 , for example, in some examples, the image processing module 301 is configured to perform multiple image processing operations on the initial image sample to obtain a first image and a second image.

[0119] As shown in Figure 6, for example, in some examples, the feature extraction module 302 is configured to perform feature extraction on the initial image sample to obtain initial image features; is configured to perform feature extraction on the first image to obtain first image features; and is configured to perform feature extraction on the second image to obtain second image features.

[0120] As shown in FIG. 6 , for example, in some examples, the analysis module 303 is configured to obtain a first loss function and a second loss function based on the initial image feature, the first image feature, and the second image feature.

[0121] As shown in FIG6 , for example, in some examples, the training module 304 is configured to train the image classification model based on the first loss function and the second loss function to obtain a trained image classification model.

[0122] As shown in FIG6 , for example, the image processing module 301, the feature extraction module 302, the analysis module 303, and the training module 304 may include codes and programs stored in a memory; the processor may execute the codes and programs to implement some or all of the functions of the above-mentioned image processing module 301, the feature extraction module 302, the analysis module 303, and the training module 304. For example, the image processing module 301, the feature extraction module 302, the analysis module 303, and the training module 304 may be dedicated hardware devices used to implement some or all of the functions of the above-mentioned image processing module 301, the feature extraction module 302, the analysis module 303, and the training module 304. For example, the image processing module 301, the feature extraction module 302, the analysis module 303, and the training module 304 may be a circuit board or a combination of multiple circuit boards used to implement the above-mentioned functions. In an embodiment of the present application, the circuit board or the combination of multiple circuit boards may include: (1) one or more processors; (2) one or more non-transitory memories connected to the processors; and (3) firmware stored in the memory that is executable by the processor.

[0123] It should be noted that the image processing module 301 can be used to implement step S110 shown in FIG2 , the feature extraction module 302 can be used to implement steps S120 to S140 shown in FIG2 , the analysis module 303 can be used to implement step S150 shown in FIG2 , and the training module 304 can be used to implement step S160 shown in FIG2 . Therefore, for a detailed description of the functions implemented by the image processing module 301, the feature extraction module 302, the analysis module 303, and the training module 304, reference can be made to the description of steps S110 to S160 in the aforementioned embodiment of the image processing method, and any repetitions will not be repeated here. Furthermore, the training device 300 can achieve similar technical effects as the aforementioned training method, and therefore will not be repeated here.

[0124] It should be noted that in the embodiments of the present disclosure, the training device may include more or fewer circuits or units, and the connection relationship between the various circuits or units is not limited and can be determined according to actual needs. The specific configuration of each circuit or unit is not limited and can be composed of analog devices based on circuit principles, or can be composed of digital chips, or constructed in other applicable ways.

[0125] As shown in FIG6 , for example, in some examples, the image processing module 301 may also be configured to perform data amplification on the initial image sample to obtain a first image.

[0126] As shown in FIG6 , for example, in some examples, the analysis module 303 may also be configured to take the initial image sample and the first image as a first positive sample pair; and obtain a first loss function based on the first positive sample pair.

[0127] As shown in FIG. 6 , for example, in some examples, the analysis module 303 may also be configured to obtain a first loss function based on the first attribute feature and the second attribute feature.

[0128] As shown in Figure 6, for example, in some examples, the analysis module 303 can also be configured to obtain a covariance matrix based on the semantic feature channel dimension of the initial image sample and the semantic feature channel dimension of the first image feature; and obtain a first loss function based on the covariance matrix.

[0129] As shown in FIG6 , for example, in some examples, the image processing module 301 may also be configured to extract a target region from the initial image sample to obtain a second image.

[0130] As shown in FIG6 , for example, in some examples, the image processing module 301 may also be configured to perform salient region detection on the initial image sample to obtain a grayscale image; and perform binarization processing on the grayscale image to obtain a second image.

[0131] As shown in FIG6 , for example, in some examples, the image processing module 301 can also be configured to perform binarization on the grayscale image to obtain a binarization result; multiply the binarization result with the initial image sample to obtain a second image; and the second image includes the target area.

[0132] As shown in FIG6 , for example, in some examples, the analysis module 303 may also be configured to take the first image and the second image corresponding to the first image as a second positive sample pair; and obtain a second loss function based on the second positive sample pair.

[0133] As shown in FIG6 , for example, in some examples, the training module 304 may also be configured to obtain a total loss function based on the first loss function and the second loss function; and train the image classification model based on the total loss function.

[0134] As shown in FIG6 , for example, the image processing module 301 , the feature extraction module 302 , the analysis module 303 and the training module 304 may also implement more or further functions.

[0135] FIG7 is a schematic block diagram of an image classification apparatus provided by at least one embodiment of the present disclosure.

[0136] As shown in Figure 7, at least one embodiment of the present disclosure further provides an image classification device 400, and the image classification device 400 may include an acquisition module 401 and a classification module 402. These components are interconnected through a bus system and / or other forms of connection mechanisms (not shown). For example, these modules can be implemented by hardware (such as circuit) modules, software modules, or any combination of the two. The following embodiments are the same and will not be repeated here. For example, these units can be implemented by a central processing unit (CPU), a graphics processing unit (GPU), a tensor processing unit (TPU), a field programmable gate array (FPGA), or other forms of processing units with data processing capabilities and / or instruction execution capabilities, as well as corresponding computer instructions. It should be noted that the components and structures of the image classification device 400 shown in Figure 7 are only exemplary and not restrictive. The image classification device 400 may also have other components and structures as needed.

[0137] As shown in FIG. 7 , for example, in some examples, the acquisition module 401 may be configured to acquire an image to be classified.

[0138] As shown in Figure 7, for example, in some examples, the classification module 402 can be configured to perform classification processing on the image to be classified to determine the category label of the image to be classified; wherein, at least part of the image classification model is trained according to the training method in any of the aforementioned examples.

[0139] As shown in Figure 7, for example, the acquisition module 401 and the classification module 402 may include codes and programs stored in a memory; the processor may execute the codes and programs to implement some or all of the functions of the acquisition module 401 and the classification module 402 as described above. For example, the acquisition module 401 and the classification module 402 may be dedicated hardware devices used to implement some or all of the functions of the acquisition module 401 and the classification module 402 as described above. For example, the acquisition module 401 and the classification module 402 may be a circuit board or a combination of multiple circuit boards used to implement the functions as described above. In an embodiment of the present application, the circuit board or the combination of multiple circuit boards may include: (1) one or more processors; (2) one or more non-transitory memories connected to the processor; and (3) firmware stored in the memory that is executable by the processor.

[0140] FIG8 is a schematic block diagram of an electronic device provided by at least one embodiment of the present disclosure.

[0141] Some embodiments of the present disclosure further provide an electronic device. For example, as shown in FIG8 , an electronic device 600 includes a memory 601 and a processor 602. It should be noted that the components of the electronic device 600 shown in FIG8 are merely exemplary and non-limiting. The electronic device 600 may also include other components as required by actual applications.

[0142] As shown in FIG8 , for example, memory 601 non-transitorily stores computer-executable instructions, and processor 602 is configured to execute the computer-executable instructions. When executed by processor 602, the computer-executable instructions implement the training method for an image classification apparatus according to any of the aforementioned embodiments. The specific implementation and related explanations of each step of the training method can be found in the aforementioned embodiments of the training method, and any repetitive details are omitted here.

[0143] As shown in FIG8 , for example, the memory 601 may include any combination of one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), a hard disk, an erasable programmable read-only memory (EPROM), a portable compact disk read-only memory (CD-ROM), a USB memory, a flash memory, etc. One or more computer-readable instructions may be stored on the computer-readable storage medium, and the processor 602 may execute the computer-readable instructions to implement various functions of the electronic device 600. Various applications and various data may also be stored in the storage medium.

[0144] As shown in FIG8 , for example, processor 602 can control other components in electronic device 600 to perform desired functions. Processor 602 can be a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The central processing unit (CPU) can be of X86 or ARM architecture, etc.

[0145] FIG9 is a schematic block diagram of a non-transitory computer-readable storage medium provided by at least one embodiment of the present disclosure.

[0146] For example, as shown in FIG9 , a non-transitory computer-readable storage medium 700 stores computer-executable instructions 701 , which, when executed by a processor, can implement the training method for an image classification device according to any one of the above items.

[0147] As shown in Figure 9, for example, the non-transitory computer-readable storage medium 700 can be applied to the above-mentioned electronic device. For example, the non-transitory computer-readable storage medium 700 can include a memory in the electronic device.

[0148] For example, the description of the non-transitory computer-readable storage medium may refer to the description of the memory in the embodiment of the electronic device, and the repeated parts will be omitted.

[0149] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0150] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0151] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

[0152] Regarding this disclosure, the following points need to be explained:

[0153] (1) The drawings of the embodiments of the present disclosure only involve structures related to the embodiments of the present disclosure, and other structures can refer to general designs.

[0154] (2) In the absence of conflict, features in the same embodiment and different embodiments of the present disclosure may be combined with each other.

[0155] The foregoing description is merely an exemplary embodiment of the present disclosure and is not intended to limit the scope of protection of the present disclosure. The scope of protection of the present disclosure is determined by the appended claims.

Claims

1. A training method for an image classification model, comprising: Performing multiple image processing operations on an initial image sample to obtain a first image and a second image; Extracting features from the initial image sample to obtain initial image features; Extracting features from the first image to obtain first image features; Extracting features from the second image to obtain second image features; Obtaining a first loss function and a second loss function based on the initial image features, the first image features, and the second image features; Training the image classification model based on the first loss function and the second loss function.

2. The training method according to claim 1, wherein The performing image processing on the initial image sample to obtain the first image includes: Performing data augmentation on the initial image sample to obtain the first image.

3. The training method according to claim 2, further comprising: Taking the initial image sample and the first image as a first positive sample pair; Obtaining the first loss function based on the first positive sample pair.

4. The training method according to claim 3, wherein The initial image features include first attribute features, and the first image features include second attribute features; Obtaining the first loss function based on the first attribute features and the second attribute features.

5. The training method according to claim 4, wherein, The first attribute feature is the semantic feature channel dimension of the initial image sample; the second attribute feature is the semantic feature channel dimension of the first image features.

6. The training method according to claim 5, wherein, The obtaining the first loss function based on the first attribute features and the second attribute features includes: Obtaining a covariance matrix based on the semantic feature channel dimension of the initial image sample and the semantic feature channel dimension of the first image features; Obtaining the first loss function based on the covariance matrix.

7. The training method according to any one of claims 1 to 6, wherein, The performing image processing on the initial image sample to obtain the second image includes: Performing target region extraction on the initial image sample to obtain the second image.

8. The training method according to claim 7, wherein, The performing target region extraction on the initial image sample to obtain the second image includes: Performing saliency region detection on the initial image sample to obtain a grayscale image; Performing binarization processing on the grayscale image to obtain the second image.

9. The training method according to claim 8, wherein The performing binarization processing on the grayscale image to obtain the second image includes: Performing the binarization processing on the grayscale image to obtain a binarization result; Performing a multiplication operation on the binarization result and the initial image sample to obtain the second image; the second image includes the target region.

10. The training method according to claim 9, wherein, The performing image processing on the initial image sample to obtain the first image includes: Performing data augmentation processing on the initial image sample to obtain the first image; The training method further comprises: Taking the first image and the second image corresponding to the first image as a second positive sample pair; Obtaining the second loss function based on the second positive sample pair.

11. The training method according to claim 9 or 10, wherein, The covariance matrix is expressed as: C i,q = Z i,q × Z i T Among them, Z i,q represents the semantic feature channel dimension of the i-th image in the first image, and Z i represents the semantic feature channel dimension of the i-th image in the initial image sample.

12. The training method according to claim 11, wherein, The first loss function is expressed as: where l AUC (i, q) represents the normalized classification score, and d(*) represents the cosine similarity, Denoting the semantic feature channel dimension of the i-th image in the second image.

13. The training method according to any one of claims 10 to 12, wherein, The second loss function is expressed as: where l OAS (i, q) represents the normalized classification score, d(*) represents the cosine similarity, f(*) represents the feature embedding function of the image, and x i represents the i-th image in the initial image sample, Among them, g(x i ) represents the binarization result, and x i,q represents the second positive sample pair.

14. The training method according to any one of claims 1 to 13, wherein, The training the image classification model based on the first loss function and the second loss function includes: Obtaining a total loss function based on the first loss function and the second loss function; Training the image classification model based on the total loss function.

15. The training method according to claim 14, wherein, The total loss function is expressed as: Where N represents the number of the initial image samples, M represents the number of the first images, and l AUC (i, q) represents the first loss function, and l OAS (i, q) represents the second loss function, and λ represents a preset parameter.

16. An image classification method, comprising: Obtain an image to be classified; Use an image classification model to perform classification processing on the image to be classified to determine the class label of the image to be classified; wherein, at least part of the image classification model is trained according to the training method described in any one of claims 1 to 15.

17. A training device for an image classification model, comprising: An image processing module configured to perform multiple image processing operations on an initial image sample to obtain a first image and a second image; A feature extraction module configured to extract features from the initial image sample to obtain initial image features; Configured to extract features from the first image to obtain first image features; and configured to extract features from the second image to obtain second image features; An analysis module configured to obtain a first loss function and a second loss function based on the initial image features, the first image features, and the second image features; A training module configured to train the image classification model based on the first loss function and the second loss function to obtain a trained image classification model.

18. An electronic device, comprising: One or more memories storing computer-executable instructions non-transiently; One or more processors configured to run the computer-executable instructions, wherein the computer-executable instructions, when run by the one or more processors, implement the training method described in any one of claims 1 to 15.

19. A non-transitory computer-readable storage medium, wherein, The non-transitory computer-readable storage medium stores computer-executable instructions, and the computer-executable instructions, when executed by a processor, implement the training method described in any one of claims 1 to 15.

Citation Information

Patent Citations

  • Convolutional neural network training and video processing method and apparatus, and electronic device

    CN108122234A

  • Deep learning-based melanoma skin disease classification and recognition method

    CN108427963A

  • Model determination method and related device

    CN117218358A

  • Target detection method and apparatus, storage medium, and terminal

    WO2022111352A1