A classification method, device, terminal and storage medium based on cross-modal feature projection learning

Through the cross-modal feature projection learning method and combined with the CLIP model, the problem of too many parameters and long training time in the small sample classification method of the visual-language pre-trained model is solved, and efficient and robust image classification is achieved.

CN118968157BActive Publication Date: 2025-07-18SOUTHERN UNIVERSITY OF SCIENCE AND TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411019821.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-29
Publication Date
2025-07-18
Estimated Expiration
2044-07-29

AI Technical Summary

Technical Problem

The existing small sample classification method of visual-language pretrained models has problems such as too many learning parameters, too long training time, low timeliness and insufficient robustness, which affects the classification effect.

Method used

The method based on cross-modal feature projection learning is adopted, and the feature maps of multiple sets of sample subsets of the target sample set are obtained, combined with the CLIP model to predict, and the category of the target image is obtained by using the cross-modal feature projection learning method.

Benefits of technology

It improves the accuracy and robustness of small sample learning image classification, reduces data annotation costs, and improves the prediction speed of classification models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118968157B_ABST
    Figure CN118968157B_ABST
Patent Text Reader

Abstract

The present invention discloses a classification method, device, terminal and storage medium based on cross-modal feature projection learning. The method includes: obtaining a target sample set, where the target sample set includes multiple groups of sample subsets, and the target sample set is a small-sample image set; obtaining multiple category feature maps based on the target sample set, each of the category feature maps corresponding to one category; receiving a target image, and obtaining a first prediction probability based on the target feature map of the target image, where the first prediction probabilities are multiple and each first prediction probability corresponds to one category; obtaining a second prediction probability based on the target image and the CLIP model, where the second prediction probabilities are multiple and each second prediction probability corresponds to one category; and obtaining the category of the target image based on the first prediction probability and the second prediction probability. The classification method based on cross-modal feature projection learning provided by the present invention can effectively accelerate the model prediction speed and avoid the problem of high data annotation cost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image detection, and particularly relates to a classification method, device, terminal and storage medium based on cross-modal feature projection learning. Background Art

[0002] Currently, the few-shot classification methods for vision-language pre-trained models mainly fall into two learning strategies: prompt-based fine-tuning and adapter-based modules. The prompt-based method relies on the powerful function of large vision-language models in image-text matching, and extracts knowledge related to specific tasks through carefully designed prompts and added learnable contexts; while the adapter-based method attaches a small neural network after the encoder of the vision-language large model to finely adjust the feature representation, so as to more effectively encode and represent image or text information. These two methods have their own advantages and provide different optimization routes when dealing with few-shot learning problems.

[0003] However, these methods either have limited performance capabilities, or have too many learnable parameters, or require too long training time. All these factors affect their timeliness in actual classification. Moreover, the models constructed by existing methods have insufficient robustness and are easily interfered by noise, seriously affecting the final classification effect.

[0004] Therefore, the existing technology still needs to be improved and enhanced. Summary of the Invention

[0005] In view of the above-mentioned defects of the existing technology, a classification method, device, terminal and storage medium based on cross-modal feature projection learning are provided, aiming to solve the problems of low timeliness and poor classification effect in the few-shot classification method for vision-language pre-trained models due to excessive learning parameters and too long training time.

[0006] In the first aspect of the present invention, a classification method based on cross-modal feature projection learning is provided, including:

[0007] Obtain a target sample set, where the target sample set includes multiple groups of sample subsets, and the labels of the images in each group of sample subsets are of the same category. Among them, the target sample set is a few-shot image set;

[0008] Obtain multiple category feature maps based on the target sample set, and each category feature map corresponds to one category;

[0009] Receive a target image, and obtain a first prediction probability based on the target feature map of the target image. Among them, there are multiple first prediction probabilities, and each first prediction probability corresponds to one category;

[0010] Obtain a second prediction probability based on the target image and the CLIP model, where there are multiple second prediction probabilities and each second prediction probability corresponds to a category;

[0011] Obtain the category of the target image based on the first prediction probability and the second prediction probability.

[0012] The classification method based on cross-modal feature projection learning, wherein obtaining multiple category feature maps based on the target sample set includes:

[0013] Obtain the feature map corresponding to each image in the target sample set;

[0014] Taking the sample subset as a unit, combine and reshape the feature maps corresponding to the images in the sample subset to obtain the category feature map corresponding to each sample subset.

[0015] The classification method based on cross-modal feature projection learning, wherein obtaining the first prediction probability based on the target feature map of the target image includes:

[0016] Based on the target feature map, obtain a target matrix corresponding to the target category feature map, such that the similarity between the target reconstruction matrix and the target feature map is greater than the similarity between any other reconstruction matrix and the target feature map, where the target category feature map is the feature map corresponding to the target category, the target reconstruction matrix is the matrix obtained by multiplying the target matrix and the target category feature map, and any other reconstruction matrix is the matrix obtained by multiplying any other matrix and the target category feature map;

[0017] Calculate the target scalar probability between the target reconstruction matrix and the target feature map, and obtain the first prediction probability corresponding to the target category based on the target scalar probability.

[0018] The classification method based on cross-modal feature projection learning, wherein obtaining the target matrix corresponding to the target category feature map based on the target feature map includes:

[0019] Obtain the target matrix based on the first formula;

[0020] The first formula is:

[0021]

[0022] where θ * is the target matrix, ∥·∥ represents the Frobenius norm, M is the target feature map, F d is the target category feature map, and δ is a parameter controlling the strength of the ridge regression penalty term.

[0023] The classification method based on cross-modal feature projection learning, wherein calculating the target scalar probability of the target reconstruction matrix and the target feature map includes:

[0024] Calculating the square of the negative average Euclidean distance between the target reconstruction matrix and all positions in the target feature map based on the second formula to obtain the target scalar probability;

[0025] The second formula is:

[0026]

[0027] where M is the target feature map, is the target reconstruction matrix, and H and W respectively represent the height and width of the target feature map.

[0028] The classification method based on cross-modal feature projection learning, wherein obtaining the second prediction probability based on the target image and the CLIP model includes:

[0029] Extracting the target label embedding representation of the target category;

[0030] Extracting the target image embedding representation of the target image;

[0031] After performing regularization processing on the target label embedding representation and the target image embedding representation, obtaining the second prediction probability corresponding to the target category based on the CLIP model.

[0032] The classification method based on cross-modal feature projection learning, wherein obtaining the category of the target image based on the first prediction probability and the second prediction probability includes:

[0033] Obtaining the target prediction probability corresponding to each category based on the first prediction probability and the second prediction corresponding to each category;

[0034] Extracting the category with the maximum target prediction probability as the category of the target image.

[0035] In the second aspect of the present invention, there is provided a classification device based on cross-modal feature projection learning, including:

[0036] A sample acquisition module, configured to acquire a target sample set, where the target sample set includes multiple groups of sample subsets, and the labels of the images in each group of sample subsets are of the same category, and wherein the target sample set is a small-sample image set;

[0037] A feature editing module, configured to obtain multiple category feature maps based on the target sample set, and each category feature map corresponds to one category;

[0038] A first prediction module, configured to receive a target image and obtain a first prediction probability based on a target feature map of the target image, where there are multiple first prediction probabilities and each first prediction probability corresponds to a category;

[0039] A second prediction module, configured to obtain a second prediction probability based on the target image and a CLIP model, where there are multiple second prediction probabilities and each second prediction probability corresponds to a category;

[0040] An output module, configured to obtain the category of the target image based on the first prediction probability and the second prediction probability.

[0041] In a third aspect of the present invention, there is provided a terminal, including: a processor and a storage medium communicatively connected to the processor. The storage medium is adapted to store multiple instructions, and the processor is adapted to call the instructions in the storage medium to execute the steps of implementing the classification method based on cross-modal feature projection learning described in any one of the above.

[0042] In a fourth aspect of the present invention, there is provided a storage medium, where the storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the classification method based on cross-modal feature projection learning described in any one of the above.

[0043] Beneficial effects: Compared with the prior art, the present invention provides a classification method, device, terminal and storage medium based on cross-modal feature projection learning. In the classification method based on cross-modal feature projection learning provided by the present invention, by obtaining a target sample set, the target sample set includes multiple groups of sample subsets, and the labels of the images in each group of sample subsets are of the same category. Among them, the target sample set is a small-sample image set. Then, based on the target sample set, multiple category feature maps are obtained, and each category feature map corresponds to one category. After that, based on the received target image, the target image is output as a target feature map to obtain a first prediction probability, where the first prediction probability is multiple and each first prediction probability corresponds to one category; based on the target image and the CLIP model, a second prediction probability is obtained, where the second prediction probability is multiple and each second prediction probability corresponds to one category. Finally, based on the first prediction probability and the second prediction probability, the category of the target image is obtained. In this way, based on the classification method based on cross-modal feature projection learning provided by the present invention, it can solve the problems in the prior art that the small-sample classification methods for vision-language pre-training models either have limited performance capabilities, or have too many learnable parameters, or require too long training time, which affects their timeliness in actual classification. Moreover, the models constructed by the existing methods have insufficient robustness and are easily interfered by noise. Through the classification method based on cross-modal feature projection learning proposed by the method provided in this embodiment, it can effectively improve the accuracy and robustness of small-sample learning image classification, and can accelerate the prediction speed of the classification model, avoiding the problem of high data annotation costs. Description of the Drawings

[0044] Figure 1 It is a flowchart of an embodiment of the classification method based on cross-modal feature projection learning provided by the present invention;

[0045] Figure 2 It is an overall model structure diagram in an embodiment of the classification method based on cross-modal feature projection learning provided by the present invention;

[0046] Figure 3 It is a schematic structural diagram of an embodiment of the classification device based on cross-modal feature projection learning provided by the present invention;

[0047] Figure 4 It is a schematic structural diagram of an embodiment of the terminal provided by the present invention. Detailed Embodiments

[0048] To make the objectives, technical solutions and effects of the present invention clearer and more definite, the following further describes the present invention in detail with reference to the accompanying drawings and by way of examples. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0049] Those skilled in the art of the present technology can understand that, unless specifically stated, the singular forms "a", "an", "the", and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the description of the present invention means the presence of the described features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or their groups. It should be understood that when we say an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.

[0050] Those skilled in the art of the present technology can understand that, unless otherwise defined, all terms used herein (including technical terms and scientific terms) have the same meaning as the general understanding of those of ordinary skill in the art to which the present invention pertains. It should also be understood that terms such as those defined in a general dictionary should be understood to have a meaning consistent with the meaning in the context of the prior art, and will not be interpreted with an idealized or overly formal meaning unless specifically defined as here.

[0051] A classification method based on cross-modal feature projection learning provided by the present invention can be applied to a terminal with computing capabilities, and the terminal can execute the classification method based on cross-modal feature projection learning provided by the present invention to classify the image to be processed.

[0052] Embodiment 1

[0053] In this embodiment, a classification method based on cross-modal feature projection learning is provided. As Figure 1 shown, the classification method based on cross-modal feature projection learning provided by the present invention includes the steps:

[0054] S100. Obtain a target sample set, where the target sample set includes multiple groups of sample subsets, and the labels of the images in each group of sample subsets are of the same category. Among them, the target sample set is a small-sample image set.

[0055] The purpose of small-sample classification is to train a classifier that can recognize new categories with a very small number of labeled samples. This method effectively avoids the problem of high data annotation costs. However, traditional small-sample learning usually relies on training from the basic categories in the original domain, which often limits the generality of the model.

[0056] The development of the latest vision-language models provides another possibility. It doesn't have to rely on a specific source domain training dataset. By simply locking the weights of the pre-trained model and training an additional module specifically for a particular downstream task, excellent performance can be achieved with very few training samples.

[0057] Referring to Figure 2 , Figure 2 which is the overall structure diagram of the classification model provided in this embodiment. This embodiment is implemented based on a small number of labeled samples combined with zero-shot reasoning of CLIP. Specifically, it includes creating a new projection model that accurately maps the image features in the support set to the feature space of the query image, transforming the classification problem into a feature projection problem to improve the accuracy and robustness of the entire model.

[0058] Specifically, in this embodiment, the support set is the target sample set. The target sample set contains multiple groups of sample subsets, and the number of sample subsets depends on the number of categories of the classification model that the user wants to obtain. And this embodiment is a small-sample classification method. The number of samples in each group of the sample subsets is small. In this embodiment, the number of samples in each group of the sample subsets can be 1 or 2 or 4 or 8 or 16 samples. In this way, the final classification result can be obtained based on a relatively small support set, which can effectively speed up the model prediction speed and avoid the problem of high data annotation cost.

[0059] S200. Obtain multiple category feature maps based on the target sample set, and each category feature map corresponds to one category.

[0060] The obtaining of multiple category feature maps based on the target sample set includes:

[0061] Obtain the feature map corresponding to each image in the target sample set;

[0062] Taking the sample subset as a unit, combine and reshape the feature maps corresponding to the images in the sample subset to obtain the category feature map corresponding to each sample subset.

[0063] Specifically, referring to Figure 2 , the model adopted in this embodiment is constructed based on CLIP, using E t as the text encoder and E v as the image encoder. Taking the ResNet encoder as an example, CLIP makes a minor modification by attaching an attention pooling layer after the traditional ResNet. In this embodiment, the encoder without the attention pooling layer is represented as the visual editor E vm , and the output of the visual editor E vm retains sufficient spatial information, so it can be used to output the feature map.

[0064] In this embodiment, the target sample set includes D sample subsets, that is, it contains images of D categories in total, each category corresponding to a group of the sample subsets, and each group of the sample subsets includes N image samples, that is, there are N labeled images in each of the D categories.

[0065] Input the support set, that is, the sample subset, into the visual editor E. vm After that, the feature map corresponding to each image can be obtained. The height, width, and number of channels of the edited feature map are H, W, and C respectively. Based on each group of the sample subsets, the feature maps corresponding to the images in the sample subsets are combined and reshaped. Specifically, for each category d ∈ D, through the visual editor E vm Aggregate the features extracted from N available image sets, and then combine and reshape them into a unified feature matrix. That is, for the target category d, the corresponding target category feature map can be expressed as

[0066] S300. Receive the target image, and obtain the first prediction probability based on the target feature map of the target image, where there are multiple first prediction probabilities and each first prediction probability corresponds to a category.

[0067] When performing picture classification based on the model provided in this embodiment, obtain a query image. In this embodiment, the query image is the target image. After inputting the target image into the classification model, first input the target image into the visual editor E vm To extract image features and obtain the target feature map.

[0068] The height, width, and number of channels of the target feature map M are the same as those of the category feature Figure 1 That is, they are H, W, and C respectively. That is, the target feature map

[0069] Specifically, obtaining the first prediction probability based on the target feature map of the target image includes:

[0070] S310. Obtain the target matrix corresponding to the target category feature map based on the target feature map, such that the similarity between the target reconstruction matrix and the target feature map is greater than the similarity between any other reconstruction matrix and the target feature map, where the target category feature map is the feature map corresponding to the target category, the target reconstruction matrix is the matrix obtained by multiplying the target matrix by the target category feature map, and any other reconstruction matrix is the matrix obtained by multiplying any other matrix by the target category feature map.

[0071] Refer to again Figure 2The upper part is the specific structure of the projection model. Based on the projector in this structure, the target category feature map can be projected into the query image feature space, which is expressed as:

[0072]

[0073] To obtain the target matrix corresponding to the target category feature map based on the target feature map. That is, a matrix needs to be found such that the product θF d can be as close as possible to the target feature map M. In this embodiment, the target matrix is defined as θ * .

[0074] That is, based on the target matrix θ * , the similarity between the target reconstruction matrix and the target feature map is greater than the similarity between any other reconstruction matrix and the target feature map. Among them, the target category feature map F d is the feature map corresponding to the target category d, and the target reconstruction matrix is the matrix obtained by multiplying the target matrix θ * and the target category feature map F d , and the any other reconstruction matrix is the matrix obtained by multiplying any other matrix and the target category feature map.

[0075] In this embodiment, the optimal target matrix θ * is obtained based on the target linear least squares problem. The linear least squares problem is represented by the first formula, and the first formula is:

[0076]

[0077] where θ * is the target matrix, ∥·∥ represents the Frobenius norm, M is the target feature map, and F d is the target category feature map, and δ is a parameter that controls the strength of the ridge regression penalty term.

[0078] One of the main advantages of the ridge regression formula is that it can provide a definite closed-form solution for the matrix θ and the optimal reconstruction M d , as shown in the target ridge regression formula:

[0079]

[0080] After obtaining the target reconstruction matrix corresponding to the target category d, the first prediction probability corresponding to the target category d can be obtained based on the target reconstruction matrix.

[0081] Specifically, in step S320, calculate the target scalar probability of the target reconstruction matrix and the target feature map, and obtain the first prediction probability corresponding to the target category based on the target scalar probability.

[0082] The calculation of the target scalar probability of the target reconstruction matrix and the target feature map includes:

[0083] Calculate the square of the negative average Euclidean distance at all positions in the target reconstruction matrix and the target feature map based on the second formula to obtain the target scalar probability;

[0084] The second formula is:

[0085]

[0086] where M is the target feature map, is the target reconstruction matrix, and H and W respectively represent the height and width of the target feature map.

[0087] Specifically, in this embodiment, the target scalar probability logical value is calculated as the negative average Euclidean distance squared between M and at all feature map positions. In this embodiment, the target scalar probability is obtained based on the second formula;

[0088] The second formula is:

[0089]

[0090] where M is the target feature map, is the target reconstruction matrix, and H and W respectively represent the height and width of the target feature map.

[0091] After obtaining the target scalar probability, it is necessary to obtain the first prediction probability corresponding to the target category based on the target scalar probability.

[0092] In this embodiment, instead of relying on heuristic methods, a more adaptive method is selected, allowing the network to learn the parameter δ that controls the strength of the ridge regression penalty term. It is not difficult to see that δ is a regularizer. Based on the adaptive method, allowing the network to learn δ can endow the network with the ability to determine the optimal regularization level, promoting discriminative projection rather than rigidly adhering to the optimality of least squares.

[0093] In this embodiment, to ensure non-negativity, use e μ to parameterize δ, with the initial value set to zero. The target reconstruction matrix after parameterizing the regularizer δ is expressed as:

[0094]

[0095] In this embodiment, a trainable temperature factor is also introduced, denoted as ∈, for predicting the first prediction probability.

[0096] The expression of the first predicted probability is:

[0097]

[0098] x q is the target image, P R (y q =d|x q ) represents the probability that the target image belongs to the target type d, that is, the first predicted probability corresponding to type d, D represents the set of all categories, and d′ represents any type in D.

[0099] S400: Acquire a second prediction probability based on the target image and the CLIP model, wherein there are multiple second prediction probabilities and each second prediction probability corresponds to a category.

[0100] The acquiring a second prediction probability based on the target image and the CLIP model comprises:

[0101] Extract target label embedding representation of target category;

[0102] extracting a target image embedding representation of the target image;

[0103] After regularizing the target label embedding representation and the target image embedding representation, a second prediction probability corresponding to the target category is obtained based on the CLIP model.

[0104] It can be seen that this embodiment performs visual text reasoning based on the CLIP model to obtain the second prediction probability.

[0105] Specifically, for the label of the target category d∈D, it is placed in a manual hint template, such as “aphoto of{class}”, denoted as Π d Then, the text encoder E t Get the text features corresponding to the target category d Denoted as the target label embedding representation

[0106] For the target image, based on the image encoder E v Extract the target image x q The target image embedding representation f of the target image is obtained v Then, the target image is embedded into the representation fv and the target label embedding representation f t is processed by L2 regularization. Based on the original CLIP, the expression of the second prediction probability that the target image belongs to the target category d is:

[0107]

[0108] where T is the temperature parameter learned by CLIP.

[0109] S500. Obtain the category of the target image based on the first prediction probability and the second prediction probability.

[0110] The obtaining the category of the target image based on the first prediction probability and the second prediction probability includes:

[0111] Obtain the target prediction probability corresponding to each category based on the first prediction probability and the second prediction corresponding to each category;

[0112] Extract the category with the largest target prediction probability as the category of the target image.

[0113] Specifically, when the first prediction probability inferred from the visual representation reconstruction model and the second prediction probability inferred from the original CLIP are fused, better prediction results can be obtained.

[0114] Specifically, calculate the final probability that the target image is of the target type d based on the third formula;

[0115] The third formula is:

[0116] P total (y q =d|x q ) = P CLIP (y q =d|x q ) + ηP R (y q =d|x q );

[0117] P total (y q =d|x q ) is the final probability that the target image is of the target type d, and η is a hyperparameter used to control the scaling of the residual connection.

[0118] Based on this, calculate the final probability P total (y q =d'|x q) If d′ ∈ D, through comparison, the type with the largest final probability can be obtained and output as the category of the target image.

[0119] In summary, this embodiment provides a classification method based on cross-modal feature projection learning. In the classification method based on cross-modal feature projection learning provided in this embodiment, by obtaining a target sample set, where the target sample set includes multiple groups of sample subsets, and the labels of the images in each group of sample subsets are of the same category. Among them, the target sample set is a small-sample image set. Then, based on the target sample set, multiple category feature maps are obtained, and each category feature map corresponds to one category. After that, based on the received target image, the target image is output as a target feature map to obtain a first prediction probability, where the first prediction probability is multiple and each first prediction probability corresponds to one category; based on the target image and the CLIP model, a second prediction probability is obtained, where the second prediction probability is multiple and each second prediction probability corresponds to one category. Finally, based on the first prediction probability and the second prediction probability, the category of the target image is obtained. In this way, based on the classification method based on cross-modal feature projection learning provided in this embodiment, it can solve the problems in the prior art that the small-sample classification methods for vision-language pre-training models either have limited performance capabilities, or have too many learnable parameters, or require too long training time, which affects their timeliness in actual classification. Moreover, the models constructed by the existing methods have insufficient robustness and are easily interfered by noise. Through the classification method based on cross-modal feature projection learning proposed by the method provided in this embodiment, the accuracy and robustness of small-sample learning image classification can be greatly improved, and the prediction speed of the classification model can be effectively accelerated, avoiding the problem of high data annotation costs.

[0120] It should be understood that although each step in the flowchart given in the accompanying drawings of the present invention is shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of the steps in the present invention does not have a strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps of the present invention may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be executed alternately or alternately with at least a part of other steps or sub-steps or stages of other steps.

[0121] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided by the present invention can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0122] Embodiment 2

[0123] Based on the above embodiments, the present invention further provides a classification device based on cross-modal feature projection learning. The schematic diagram of its functional modules is as Figure 3 shown. The classification device based on cross-modal feature projection learning includes:

[0124] A sample acquisition module, configured to acquire a target sample set, where the target sample set includes multiple groups of sample subsets, and the labels of the images in each group of sample subsets are of the same category. Among them, the target sample set is a small-sample image set, specifically as described in Embodiment 1;

[0125] A feature editing module, configured to obtain multiple category feature maps based on the target sample set, and each category feature map corresponds to one category, specifically as described in Embodiment 1;

[0126] A first prediction module, configured to receive a target image and obtain a first prediction probability based on the target feature map of the target image. Among them, there are multiple first prediction probabilities, and each first prediction probability corresponds to one category, specifically as described in Embodiment 1;

[0127] A second prediction module, configured to obtain a second prediction probability based on the target image and the CLIP model, where there are multiple second prediction probabilities and each second prediction probability corresponds to one category, as specifically described in Embodiment 1;

[0128] An output module, configured to obtain the category of the target image based on the first prediction probability and the second prediction probability, as specifically described in Embodiment 1.

[0129] Embodiment 3

[0130] Based on the classification method based on cross-modal feature projection learning described in the above Embodiment 1, the present invention also provides a terminal, and its principle block diagram can be as Figure 4 shown. The terminal includes a memory 10 and a processor 20. A classification program based on cross-modal feature projection learning is stored in the memory 10. When the processor 10 executes the computer program, at least the following steps can be implemented:

[0131] Obtain a target sample set, where the target sample set includes multiple groups of sample subsets, and the labels of the images in each sample subset are of the same category, where the target sample set is a small sample image set;

[0132] Obtain multiple category feature maps based on the target sample set, and each category feature map corresponds to one category;

[0133] Receive a target image, and obtain a first prediction probability based on the target feature map of the target image, where there are multiple first prediction probabilities and each first prediction probability corresponds to one category;

[0134] Obtain a second prediction probability based on the target image and the CLIP model, where there are multiple second prediction probabilities and each second prediction probability corresponds to one category;

[0135] Obtain the category of the target image based on the first prediction probability and the second prediction probability.

[0136] Wherein, the obtaining multiple category feature maps based on the target sample set includes:

[0137] Obtain the feature map corresponding to each image in the target sample set;

[0138] Taking the sample subset as a unit, combine and reshape the feature maps corresponding to the images in the sample subset to obtain the category feature map corresponding to each sample subset.

[0139] Wherein, the obtaining the first prediction probability based on the target feature map of the target image includes:

[0140] Obtain a target matrix corresponding to the target class feature map based on the target feature map, such that the similarity between the target reconstruction matrix and the target feature map is greater than the similarity between any other reconstruction matrix and the target feature map, where the target class feature map is the feature map corresponding to the target class, the target reconstruction matrix is the matrix obtained by multiplying the target matrix by the target class feature map, and any other reconstruction matrix is the matrix obtained by multiplying any other matrix by the target class feature map;

[0141] Calculate the target scalar probability of the target reconstruction matrix and the target feature map, and obtain the first prediction probability corresponding to the target class based on the target scalar probability.

[0142] Among them, the obtaining the target matrix corresponding to the target class feature map based on the target feature map includes:

[0143] Obtain the target matrix based on the first formula;

[0144] The first formula is:

[0145]

[0146] Among them, θ* is the target matrix, ||·|| represents the Frobenius norm, M is the target feature map, F d is the target class feature map, and 6 is a parameter for controlling the strength of the ridge regression penalty term.

[0147] Among them, the calculating the target scalar probability of the target reconstruction matrix and the target feature map includes:

[0148] Calculate the square of the negative average Euclidean distance at all positions in the target reconstruction matrix and the target feature map based on the second formula to obtain the target scalar probability;

[0149] The second formula is:

[0150]

[0151] Among them, M is the target feature map, is the target reconstruction matrix, and H and W respectively represent the height and width of the target feature map.

[0152] Among them, the obtaining the second prediction probability based on the target image and the CLIP model includes:

[0153] Extract the target label embedding representation of the target class;

[0154] Extract the target image embedding representation of the target image;

[0155] After regularizing the target label embedding representation and the target image embedding representation, obtain the second prediction probability corresponding to the target category based on the CLIP model.

[0156] Among them, obtaining the category of the target image based on the first prediction probability and the second prediction probability includes:

[0157] Based on the first prediction probability and the second prediction corresponding to each category, obtain the target prediction probability corresponding to each category;

[0158] Extract the category with the largest target prediction probability as the category of the target image.

[0159] Embodiment 4

[0160] The present invention also provides a storage medium storing one or more programs, which can be executed by one or more processors to implement the steps of the classification method based on cross-modal feature projection learning described in the above embodiments.

[0161] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A classification method based on cross-modal feature projection learning, characterized in that, Including: Obtain a target sample set, where the target sample set includes multiple groups of sample subsets, and the labels of the images in each group of sample subsets are of the same category. Among them, the target sample set is a small-sample image set; Obtain multiple category feature maps based on the target sample set, and each category feature map corresponds to one category; Receive a target image, and obtain a first prediction probability based on the target feature map of the target image. Among them, there are multiple first prediction probabilities, and each first prediction probability corresponds to one category; Obtain a second prediction probability based on the target image and the CLIP model. Among them, there are multiple second prediction probabilities, and each second prediction probability corresponds to one category; Obtain the category of the target image based on the first prediction probability and the second prediction probability; The obtaining the first prediction probability based on the target feature map of the target image includes: obtaining a target matrix corresponding to the target category feature map based on the target feature map, such that the similarity between the target reconstruction matrix and the target feature map is greater than the similarity between any other reconstruction matrix and the target feature map. Among them, the target category feature map is the feature map corresponding to the target category, the target reconstruction matrix is the matrix obtained by multiplying the target matrix by the target category feature map, and any other reconstruction matrix is the matrix obtained by multiplying any other matrix by the target category feature map; calculating the target scalar probability between the target reconstruction matrix and the target feature map, and obtaining the first prediction probability corresponding to the target category based on the target scalar probability.

2. The classification method based on cross-modal feature projection learning according to claim 1, wherein The obtaining multiple category feature maps based on the target sample set includes: Obtain the feature map corresponding to each image in the target sample set; Taking the sample subset as a unit, combine and reshape the feature maps corresponding to the images in the sample subset to obtain the category feature map corresponding to each sample subset.

3. The classification method based on cross-modal feature projection learning according to claim 1, wherein The obtaining the target matrix corresponding to the target category feature map based on the target feature map includes: Obtain the target matrix based on the first formula; The first formula is: where θ * is the target matrix, ‖·‖ represents the Frobenius norm, M is the target feature map, F d is the target class feature map, and δ is a parameter that controls the strength of the ridge regression penalty term.

4. The classification method based on cross-modal feature projection learning according to claim 1, wherein The calculating the target scalar probability between the target reconstruction matrix and the target feature map includes: Calculating the square of the negative average Euclidean distance at all positions between the target reconstruction matrix and the target feature map based on the second formula to obtain the target scalar probability; The second formula is: Among them, M is the target feature map, is the target reconstruction matrix, and H and W respectively represent the height and width of the target feature map.

5. The classification method based on cross-modal feature projection learning according to claim 1, wherein The obtaining the second prediction probability based on the target image and the CLIP model includes: Extract the target label embedding representation of the target category; Extract the target image embedding representation of the target image; After performing regularization processing on the target label embedding representation and the target image embedding representation, obtain the second prediction probability corresponding to the target category based on the CLIP model.

6. The classification method based on cross-modal feature projection learning according to claim 1, characterized in that The obtaining the category of the target image based on the first prediction probability and the second prediction probability includes: Obtain the target prediction probability corresponding to each category based on the first prediction probability and the second prediction corresponding to each category; Extract the category with the largest target prediction probability as the category of the target image.

7. A classification device based on cross-modal feature projection learning, characterized in that, The device includes: A sample acquisition module for acquiring a target sample set, where the target sample set includes multiple groups of sample subsets, and the labels of the images in each group of sample subsets are of the same category. Among them, the target sample set is a small sample image set; A feature editing module for obtaining multiple category feature maps based on the target sample set, where each category feature map corresponds to one category; A first prediction module for receiving a target image and obtaining a first prediction probability based on the target feature map of the target image. Among them, there are multiple first prediction probabilities, and each first prediction probability corresponds to one category; A second prediction module for obtaining a second prediction probability based on the target image and the CLIP model. Among them, there are multiple second prediction probabilities, and each second prediction probability corresponds to one category; An output module for obtaining the category of the target image based on the first prediction probability and the second prediction probability; The classification device based on cross-modal feature projection learning is further configured to obtain a target matrix corresponding to the target category feature map based on the target feature map, such that the similarity between the target reconstruction matrix and the target feature map is greater than the similarity between any other reconstruction matrix and the target feature map. Among them, the target category feature map is the feature map corresponding to the target category, the target reconstruction matrix is the matrix obtained by multiplying the target matrix by the target category feature map, and any other reconstruction matrix is the matrix obtained by multiplying any other matrix by the target category feature map; calculate the target scalar probability between the target reconstruction matrix and the target feature map, and obtain the first prediction probability corresponding to the target category based on the target scalar probability.

8. A terminal, characterized in that, The terminal includes: a processor and a storage medium communicatively connected to the processor. The storage medium is adapted to store multiple instructions, and the processor is adapted to call the instructions in the storage medium to execute the steps of implementing the classification method based on cross-modal feature projection learning according to any one of claims 1-6 above.

9. A storage medium, characterized in that, The storage medium stores one or more programs, and the one or more programs can be executed by one or more processors to implement the steps of the classification method based on cross-modal feature projection learning according to any one of claims 1-6.