A training method and device for combined zero-shot image classification and model

By adopting global primitive relationships and overall modeling of primitive experts in the combined zero-sample image classification model, the problem that local primitive relationships are difficult to obtain comprehensive primitive representations is solved, and the performance of image classification is improved.

CN119785129BActive Publication Date: 2025-06-24RENMIN ZHONGKE (JINAN) INTELLIGENT TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510292848.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-06-24
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

The current combined zero-sample image classification model is difficult to obtain comprehensive primitive representations because it is based on local primitive relationships, resulting in poor classification results.

Method used

A combined zero-sample image classification model based on global primitive relationship is proposed. By learning the experts of attributes and objects on the entire training set, the global primitive relationship is mined, and the decoupled primitive features are generated to perform image classification.

Benefits of technology

Through the exploration of global primitive relationships and overall modeling of primitive experts, a more comprehensive primitive representation is obtained, which improves the performance of combined zero-sample image classification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119785129B_ABST
    Figure CN119785129B_ABST
Patent Text Reader

Abstract

The present disclosure belongs to the field of computer vision technology, and particularly relates to a combined zero-shot image classification and a training method and device for a model. The training method of the combined zero-shot image classification model includes: obtaining an image classification data set and dividing it into a training set and a test set; constructing a neural network model, training the neural network model based on the training set to generate the combined zero-shot image classification model, wherein the neural network model at least includes a combined recognition branch and a primitive recognition branch, the combined recognition branch is used to obtain a combined feature representation of each sample based on the global features of the training set samples, and the primitive recognition branch is used to obtain decoupled primitive features by mining the global primitive relationships of the training samples for the recognition of primitives, and the primitives include attributes and objects. The present disclosure improves the performance of combined zero-shot image classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure belongs to the technical field of computer vision, and particularly relates to a combined zero-shot image classification and model training method and apparatus. Background Art

[0002] Combined zero-shot image classification aims to classify images into previously unseen attribute-object pair labels by learning visual primitives (i.e., attributes or objects) in images with known pair labels. For example, based on the learned images of "ripe apples" and "sliced bread", the image of "sliced apples" that has not been learned can be successfully classified.

[0003] The key to combined zero-shot image classification lies in separately modeling attributes and objects for transfer to unseen combinations. To this end, visual decoupling technology is required to separate the representations of attributes and objects from the images of visible combinations. The current visual decoupling technology extracts decoupled primitive representations by calculating the similarity of a pair of images with the same primitives.

[0004] However, this method only explores local primitive relationships with prerequisites (i.e., when another primitive is the same). For complex combinations, it is difficult for the model to obtain a comprehensive primitive representation, which easily leads to classification errors. Summary of the Invention

[0005] The embodiments of the present disclosure propose a combined zero-shot image classification scheme based on global primitive relationships to solve the problem that the current combined zero-shot image classification scheme has poor classification effects due to local primitive relationships.

[0006] The first aspect of the embodiments of the present disclosure provides a training method for a combined zero-shot image classification model, including:

[0007] Obtain an image classification data set and divide it into a training set and a test set;

[0008] Construct a neural network model, and train the neural network model based on the training set to generate the combined zero-shot image classification model, where the neural network model at least includes a combination recognition branch and a primitive recognition branch. The combination recognition branch is used to obtain the combination feature representation of each sample based on the global features of the training set samples, and the primitive recognition branch is used to obtain the decoupled primitive features by mining the global primitive relationships of the training samples for primitive recognition, where the primitives include attributes and objects.

[0009] In some embodiments of the present disclosure, the obtaining the combination feature representation of each sample based on the global features of the training set samples includes:

[0010] Extract the global features of the training set samples;

[0011] Enhance the feature representation of the global features based on self-attention to obtain the combined feature representation of each sample.

[0012] In some embodiments of the present disclosure, the extraction of the global features of the training set samples includes:

[0013] Use a ViT model pre-trained in a self-supervised manner on the ImageNet dataset as an image encoder to extract global features of the training set samples that contain more semantic information and are more suitable for the image classification task.

[0014] In some embodiments of the present disclosure, the obtaining of the decoupled primitive features by mining the global primitive relationships of the training samples for primitive recognition includes:

[0015] Learn primitive experts on the entire training set;

[0016] Use the primitive experts as queries, and the global features of the training set samples as keys and key values, perform cross-attention processing, and obtain the decoupled primitive feature representations for the training set samples.

[0017] In some embodiments of the present disclosure, the performing of the cross-attention processing further includes:

[0018] For each sample data, randomly select two auxiliary features in the training set, and constrain the decoupled primitive features at the attention level based on the EMD distance between the two auxiliary features, where one of the two auxiliary features has the same attribute but different objects, and the other has the same object but different attributes.

[0019] In some embodiments of the present disclosure, the neural network model is further used to extract text features for all combinations and primitive categories in the training set, and by separately calculating the cosine similarity between the visual features and the text features of each category in the combination recognition branch and the primitive recognition branch, and using the cross-entropy loss for constraint, obtain the classification results of the training samples.

[0020] In some embodiments of the present disclosure, the method further includes:

[0021] Test the combined zero-shot image classification model based on the test set.

[0022] The second aspect of the embodiments of the present disclosure provides a training device for a combined zero-shot image classification model, including:

[0023] An acquisition module, configured to acquire an image classification data set and divide it into a training set and a test set;

[0024] A training module for constructing a neural network model, training the neural network model based on the training set to generate the combined zero-shot image classification model, where the neural network model includes at least a combined recognition branch and a primitive recognition branch. The combined recognition branch is used to obtain the combined feature representation of each sample based on the global features of the training set samples, and the primitive recognition branch is used to obtain the decoupled primitive features by mining the global primitive relationships of the training samples for primitive recognition. The primitive includes attributes and objects.

[0025] The third aspect of the embodiments of the present disclosure provides a combined zero-shot image classification method, including:

[0026] Obtain image data;

[0027] Input the image data into the combined zero-shot image classification model trained by the method according to the first aspect of the embodiments of the present disclosure, obtain the cosine similarities between the corresponding visual features and the text features of each category in the combined recognition branch and the primitive recognition branch respectively, perform weighted fusion on the two, and select the category with the highest similarity as the classification result of the image data and output it.

[0028] The fourth aspect of the embodiments of the present disclosure provides a combined zero-shot image classification device, including:

[0029] An acquisition module for obtaining image data;

[0030] A classification module for inputting the image data into the combined zero-shot image classification model trained by the method according to the first aspect of the embodiments of the present disclosure, obtaining the cosine similarities between the corresponding visual features and the text features of each category in the combined recognition branch and the primitive recognition branch respectively, performing weighted fusion on the two, and selecting the category with the highest similarity as the classification result of the image data and output it.

[0031] In summary, the training method and device of the combined zero-shot image classification model and the combined zero-shot image classification method and device provided by the embodiments of the present disclosure, by globally modeling two learnable primitive experts (i.e., attribute expert and object expert) on the entire training set to mine global primitive relationships, avoid the problem of only mining local primitive relationships for limited data with prerequisites in the prior art, so that a more comprehensive primitive representation can be obtained, and further improve the performance of combined zero-shot image classification. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] The features and advantages of the present disclosure will be more clearly understood by referring to the accompanying drawings. The drawings are schematic and should not be construed as limiting the present disclosure in any way. In the drawings:

[0033] Figure 1It is the algorithm framework of the combined zero-shot image classification method based on the global primitive relationship proposed by the present disclosure;

[0034] Figure 2 It is a flowchart of a training method for a combined zero-shot image classification model shown according to some embodiments of the present disclosure;

[0035] Figure 3 It is the network architecture of the neural network model used to train the combined zero-shot image classification model in the training and testing phases shown according to some embodiments of the present disclosure;

[0036] Figure 4 It is a schematic diagram of a training device for a combined zero-shot image classification model shown according to some embodiments of the present disclosure;

[0037] Figure 5 It is a flowchart of a combined zero-shot image classification method shown according to some embodiments of the present disclosure;

[0038] Figure 6 It is a schematic diagram of a combined zero-shot image classification device shown according to some embodiments of the present disclosure. Detailed implementation manners

[0039] In the following detailed description, many specific details of the present disclosure are set forth by way of example in order to provide a thorough understanding of the relevant disclosure. However, it will be apparent to those of ordinary skill in the art that the present disclosure may be practiced without these details. It should be understood that the terms "system", "device", "unit" and / or "module" used in the present disclosure are a way to distinguish different components, elements, parts or components at different levels in a sequential arrangement. However, if other expressions can achieve the same purpose, these terms may be replaced by other expressions.

[0040] It should be understood that when a device, unit or module is referred to as "on", "connected to" or "coupled to" another device, unit or module, it may be directly on the other device, unit or module, connected or coupled to or communicate with other devices, units or modules, or there may be intermediate devices, units or modules, unless the context clearly indicates an exception. For example, the term "and / or" used in the present disclosure includes any and all combinations of one or more of the related listed items.

[0041] The terms used in this disclosure are only for describing specific embodiments and do not limit the scope of this disclosure. As shown in the specification and claims of this disclosure, unless the context clearly indicates otherwise, words such as "a", "an", "one", and / or "the" are not specifically singular and may also include the plural. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of the clearly identified features, wholes, steps, operations, elements, and / or components, and such expressions do not constitute an exclusive list, and other features, wholes, steps, operations, elements, and / or components may also be included.

[0042] Referring to the following description and the accompanying drawings, these or other features and characteristics of the present disclosure, the operating methods, the functions of the related elements of the structure, the combination of parts, and the economy of manufacture can be better understood, where the description and the drawings form a part of the specification. However, it can be clearly understood that the drawings are only for the purpose of illustration and description and are not intended to limit the protection scope of the present disclosure. It can be understood that the drawings are not drawn to scale.

[0043] A variety of structure diagrams are used in this disclosure to illustrate various deformations according to the embodiments of the present disclosure. It should be understood that the foregoing or following structures are not used to limit the present disclosure. The protection scope of the present disclosure is subject to the claims.

[0044] Humans are born with the ability to generalize different learned concepts (such as objects and attributes) into new combinations. For example, there are images of "ripe apples" and "sliced bread" in the mind. When an unknown image of "sliced apples" appears, humans can effortlessly understand such a new composite concept. Similar to language-based computer vision tasks (such as image captioning and navigation), in the case of insufficient training samples, recognizing new composite concepts is crucial for visual understanding. Considering this, compositional zero-shot learning, which aims to classify images into previously unseen attribute-object pair labels by learning visual primitives (i.e., attributes or objects) in images with known pair labels, has recently received increasing attention.

[0045] Since visual primitives appear in combinations of visible images, the key to compositional zero-shot image classification lies in how to model attributes and objects separately so as to transfer to unseen combinations. To solve this problem, increasing attention has been paid to visual decoupling, which aims to separate the representations of attributes and objects from the images of visible combinations. Among them, some previous studies extracted decoupled primitive representations by calculating the similarity of a pair of images with the same primitives. Their success largely depends on mining different attribute / different object relationships under the same object. However, these methods only explore local primitive relationships with preconditions (i.e., when another primitive is the same). For complex combinations, it is difficult for the model to obtain a comprehensive primitive representation, which easily leads to classification errors.

[0046] To solve the above problems, the present disclosure proposes a combined zero-shot image classification model based on global primitive relationships and performs combined zero-shot image classification based on this model. The combined zero-shot image classification method proposed by the present disclosure aims to explore two global primitive relationships to better solve the combined zero-shot problem. To globally model the primitive relationships, the present disclosure separately learns two experts for all attributes and all objects. In addition, to avoid incomplete learning of relationships for limited data (single or paired samples), the present disclosure learns the experts at the level of the entire dataset. Figure 1 is the algorithm framework of the combined zero-shot image classification method based on global primitive relationships proposed by the present disclosure.

[0047] In some embodiments, the training method of the combined zero-shot image classification model is as Figure 2 shown, and specifically includes the following steps:

[0048] S210, obtain an image classification dataset and divide it into a training set and a test set.

[0049] First, obtain an image classification dataset.

[0050] In some embodiments of the present disclosure, the obtained image classification dataset includes three benchmark datasets, namely UT-Zappos50K, Clothing16K, and C-GQA, with 16 / 12, 9 / 8, and 413 / 674 attributes and objects respectively. Specifically, UT-Zappos50K is a fine-grained dataset containing different types of shoes (such as boots, sandals) with texture attributes (such as canvas, cotton). Clothing16K consists of different kinds of clothes (such as shirts, trousers) with color attributes (such as blue, yellow). C-GQA is created for the VQA task based on the StanfordGQA dataset and consists of common attributes (such as red, dirty) and objects (such as pens, windows) in daily life.

[0051] Secondly, divide the semantic categories into known combination categories and unknown combination categories, and accordingly divide the training set and the test set (including two ways of dividing the test set, namely narrow sense and broad sense).

[0052] Some embodiments of the present disclosure divide UT-Zappos50K, Clothing16K, and C-GQA into 83, 18, and 5592 training set categories respectively.

[0053] S220. Construct a neural network model, and train the neural network model based on the training set to generate the combined zero-shot image classification model. Among them, the neural network model at least includes a combined recognition branch and a primitive recognition branch. The combined recognition branch is used to obtain the combined feature representation of each sample based on the global features of the training set samples. The primitive recognition branch is used to obtain the decoupled primitive features by mining the global primitive relationships of the training samples for the recognition of primitives. The primitives include attributes and objects.

[0054] The complete training process of the combined zero-shot image classification model includes two stages: training and testing.

[0055] In the training stage, the neural network model extracts text features for all combinations and primitive categories in the training set, and calculates the cosine similarity between the visual features and the text features of each category in the combined recognition branch and the primitive recognition branch respectively, and uses the cross-entropy loss for constraint to obtain the classification results of the training samples. In the testing stage, the test set is used to test the combined zero-shot image classification model based on global relationship mining, and the classification results are output and the model performance is evaluated. In some embodiments of the present disclosure, the network architecture of the neural network model used to train the combined zero-shot image classification model in the training and testing stages is as Figure 3 shown.

[0056] The specific steps of the training stage are as follows:

[0057] 1. Obtain the combined feature representation based on the global features and mine the global primitive relationships

[0058] Specifically, the combined recognition branch first uses the ViT model pre-trained in a self-supervised manner on the ImageNet dataset as an image encoder to extract the global features of the training set samples that contain more semantic information and are more suitable for the image classification task, which can be expressed as , where is the sample index, is the number of training set samples, is the th image sample, is the image encoder, is the th global feature of the image sample. Then, the global features of all samples are enhanced by the self-attention module to obtain the combined feature representation of each sample, specifically:

[0059] ;

[0060]

[0061] Among them, Q is the query, K is the key, V is the key value, and the scaling factor is the dimension of is the sample index, is the number of training set samples, is the global feature of the

[0062] The primitive recognition branch first separately learns two primitive experts (i.e., the attribute expert and the object expert ) to mine the global primitive relationship. and are randomly initialized and have the same dimension as . Since a single sample only contains one attribute and one object, to avoid incomplete learning of the relationships of limited data, and are learned across data batches, that is, on the entire dataset.

[0063] Then, the primitive experts are used as queries and the global features of the training samples are used as keys and key values and input into the cross-attention module at the same time to obtain the decoupled attribute and object features for the training sample features respectively represented as and , specifically:

[0064] ;

[0065]

[0066] Specifically, is the sample index, is the number of training set samples, is the global feature of the image sample, is the attribute expert, is the object expert. is to obtain the decoupled attribute feature representation for the training sample feature is to obtain the decoupled object feature representation for the training sample feature . To ensure that the primitive experts generate similar features for samples with the same primitive, for each , the present invention randomly selects two auxiliary features in the dataset: one has the same attribute but different objects, denoted as , and the other has the same object but different attributes, denoted as . Since the Earth Mover's Distance (EMD) is often used to evaluate the dissimilarity between two multi-dimensional distributions in the feature space, this disclosure introduces EMD to constrain the decoupled primitive features at the attention level. For convenience, use to represent . Therefore, the EMD regularization constraint can be expressed as:

[0067] where and are the attributes and object features decoupled from the training sample features respectively, and are the attributes and object features decoupled from the training sample features respectively.

[0068] 2. Use the word embedding Word2vec to extract text features for all combinations and primitive categories in the training set , where , as learnable text information. It should be noted that through word embedding, each and can be directly extracted. By inputting the connected embeddings of and through a linear layer, the word embedding of the corresponding primitive combination can be obtained.

[0069] 3. Calculate the cosine similarity between the visual features and various text features in the combination and primitive branches respectively. All and need to be adjusted to the same feature dimension as through two-layer MLP respectively, and the results are denoted as , . At this time, all visual features obtained through self-attention or cross-attention are collectively referred to as , and its classification probability can be calculated as:

[0070] where, is the natural exponential function, is the matrix multiplication operation, is the temperature coefficient, is the set of all label texts of combinations, attributes and objects.

[0071] 4. Calculate the cross-entropy loss function for the visual feature , and use the Adam optimizer to optimize the entire model. The total loss function can be expressed as:

[0072] ;

[0073]

[0074] Among them, 、 、 are the cross-entropy loss functions on the attribute branch, combination branch, and object branch respectively, is the set of all label texts of combinations, attributes, and objects.

[0075] The specific steps in the test phase are as follows:

[0076] 1. Use the pre-trained image encoder to extract the global features of the test samples, enhance the feature representation through the self-attention module, and obtain the combined feature representation of each sample; and respectively input the learned two primitive experts (i.e., the attribute expert and the object expert) and the global features of the test samples into the cross-attention module to obtain the decoupled primitive feature representation for the test samples.

[0077] 2. Use the learned primitive word embeddings to obtain the text features of the combination categories in all test sets.

[0078] 3. Calculate the cosine similarities between the visual features and the text features of each category in the combination and primitive branches respectively to obtain the classification results of the test samples.

[0079] In particular, all parameters of the model are frozen in the test phase, no loss function is calculated; nor is the update of the image encoder parameters performed;

[0080] Figure 4 is a schematic diagram of a training device for a combined zero-shot image classification model shown according to some embodiments of the present disclosure. As Figure 4 shown, the training device 400 for the combined zero-shot image classification model includes a partitioning module 410 and a training module 420, wherein:

[0081] The partitioning module 410 is configured to obtain an image classification data set and partition it into a training set and a test set;

[0082] The training module 420 is configured to construct a neural network model, train the neural network model based on the training set, and generate the combined zero-shot image classification model, wherein the neural network model at least includes a combination recognition branch and a primitive recognition branch, the combination recognition branch is configured to obtain the combined feature representation of each sample based on the global features of the training set samples, and the primitive recognition branch is configured to obtain the decoupled primitive features by mining the global primitive relationships of the training samples for primitive recognition, and the primitives include attributes and objects.

[0083] Figure 5 FIG. 1 is a flowchart of a combined zero-sample image classification method according to some embodiments of the present disclosure. Figure 5 As shown, the specific steps include:

[0084] S510, acquiring image data.

[0085] S520, input the image data into the Figure 2 The combined zero-sample image classification model trained by the method described in s210-s220 obtains the cosine similarity of the corresponding visual features and the text features of each category in the combined recognition branch and the primitive recognition branch, respectively, performs weighted fusion on the two, selects the category with the highest similarity, and outputs it as the classification result of the image data.

[0086] Figure 6 is a schematic diagram of a combined zero-sample image classification device according to some embodiments of the present disclosure. Figure 4 As shown, the combined zero-sample image classification model device 600 includes an acquisition module 610 and a classification module 620, wherein:

[0087] An acquisition module, used for acquiring image data;

[0088] A classification module is used to input the image data into a Figure 2 The combined zero-sample image classification model trained by the method described in s210-s220 obtains the cosine similarity of the corresponding visual features and the text features of each category in the combined recognition branch and the primitive recognition branch, respectively, performs weighted fusion on the two, selects the category with the highest similarity, and outputs it as the classification result of the image data.

[0089] In summary, the training method and device of the combined zero-shot image classification model and the combined zero-shot image classification method and device provided in the embodiments of the present disclosure, by holistically modeling two learnable primitive experts (i.e., attribute expert and object expert) on the entire training set to mine global primitive relationships, thereby avoiding the problem of the prior art of only mining local primitive relationships for limited data with prerequisites. Therefore, a more comprehensive primitive representation can be obtained, thereby improving the performance of combined zero-shot image classification.

[0090] Although the subject matter described herein is provided in the general context of execution in conjunction with the execution of an operating system and applications on a computer system, those skilled in the art will recognize that other implementations may also be performed in conjunction with other types of program modules. In general, program modules include routines, programs, components, data structures, and other types of structures that perform particular tasks or implement particular abstract data types. Those skilled in the art will understand that the subject matter described herein may be practiced using other computer system configurations, including handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, minicomputers, mainframe computers, etc., and may also be used in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in both local and remote memory storage devices.

[0091] Those of ordinary skill in the art will appreciate that the elements and method steps of the examples described in connection with the embodiments disclosed herein can be implemented in electronic hardware, or in a combination of computer software and electronic hardware. Whether such functions are implemented in hardware or software depends upon the particular application and design constraints of the technical solution. Skilled artisans may use different methods for each particular application to implement the described functions, but such implementations should not be considered to exceed the scope of the present disclosure.

[0092] It should be understood that the above specific embodiments of the present disclosure are merely for illustrative purposes or for explaining the principles of the present disclosure, and do not constitute a limitation to the present disclosure. Therefore, any modifications, equivalent replacements, improvements, etc. made without departing from the spirit and scope of the present disclosure shall be included within the protection scope of the present disclosure. In addition, the appended claims of the present disclosure are intended to cover all variations and modifications that fall within the scope and boundaries of the appended claims, or within the equivalent forms of such scope and boundaries.

Claims

1. A training method for a combined zero-shot image classification model, characterized in that: include: Obtain an image classification dataset and divide it into a training set and a test set; Constructing a neural network model, training the neural network model based on the training set, and generating the combined zero-sample image classification model, wherein the neural network model at least includes a combined recognition branch and a primitive recognition branch, the combined recognition branch is used to obtain a combined feature representation of each sample based on the global features of the training set samples, and the primitive recognition branch is used to obtain decoupled primitive features by mining the global primitive relationship of the training samples to identify the primitives, and the primitives include attributes and objects: The step of obtaining decoupled primitive features by mining the global primitive relationship of the training samples to identify the primitives includes: Learn primitive experts on the entire training set; Using the primitive expert as a query and the global features of the training set samples as keys and key values, cross-attention processing is performed to obtain decoupled primitive feature representations for the training set samples; The cross-attention processing further comprises: For each sample data, two auxiliary features are randomly selected in the training set, and the decoupled primitive features are constrained at the attention level based on the EMD distance of the two auxiliary features, wherein one of the two auxiliary features has the same attributes but different objects, and the other has the same object but different attributes.

2. The method according to claim 1, characterized in that: The step of obtaining a combined feature representation of each sample based on the global features of the training set samples includes: Extract global features of training set samples; The feature representation of the global feature is enhanced based on self-attention to obtain a combined feature representation of each sample.

3. The method according to claim 2, characterized in that: The global features of the training set samples are extracted including: The ViT model pre-trained in a self-supervised manner on the ImageNet dataset is used as the image encoder to extract global features of the training set samples that contain more semantic information and are more suitable for image classification tasks.

4. The method according to claim 1, characterized in that: The neural network model is also used to extract text features for all combinations and primitive categories in the training set, and obtain the classification results of the training samples by respectively calculating the cosine similarity between the visual features in the combination recognition branch and the primitive recognition branch and the text features of each category, and using cross entropy loss for constraints.

5. The method according to claim 1, characterized in that: The method further comprises: The combined zero-shot image classification model is tested based on the test set.

6. A training device for a combined zero-shot image classification model, characterized in that: include: An acquisition module is used to acquire an image classification data set and divide it into a training set and a test set; A training module, used for constructing a neural network model, training the neural network model based on the training set, and generating the combined zero-sample image classification model, wherein the neural network model at least includes a combined recognition branch and a primitive recognition branch, the combined recognition branch is used for obtaining a combined feature representation of each sample based on the global features of the training set samples, and the primitive recognition branch is used for obtaining decoupled primitive features by mining the global primitive relationship of the training samples to identify primitives, wherein the primitives include attributes and objects; The step of obtaining decoupled primitive features by mining the global primitive relationship of the training samples to identify the primitives includes: Learn primitive experts on the entire training set; Using the primitive expert as a query and the global features of the training set samples as keys and key values, cross-attention processing is performed to obtain decoupled primitive feature representations for the training set samples; The cross-attention processing further comprises: For each sample data, two auxiliary features are randomly selected in the training set, and the decoupled primitive features are constrained at the attention level based on the EMD distance of the two auxiliary features, wherein one of the two auxiliary features has the same attributes but different objects, and the other has the same object but different attributes.

7. A combined zero-shot image classification method, characterized in that: include: Get image data; The image data is input into the combined zero-sample image classification model trained according to the method according to any one of claims 1 to 5, and the cosine similarities of the corresponding visual features and the text features of each category are obtained in the combined recognition branch and the primitive recognition branch respectively, and the two are weightedly fused, and the category with the highest similarity is selected as the classification result of the image data and output.

8. A combined zero-sample image classification device, characterized in that: include: An acquisition module, used for acquiring image data; A classification module is used to input the image data into the combined zero-sample image classification model trained according to the method according to any one of claims 1 to 5, obtain the cosine similarity of the corresponding visual features and the text features of each category in the combined recognition branch and the primitive recognition branch, perform weighted fusion on the two, select the category with the highest similarity, and output it as the classification result of the image data.

Citation Information

Patent Citations

  • Combined zero sample image classification method based on progressive mutual guidance

    CN118379562A

  • Combined image recognition method based on out-of-distribution sample perception

    CN118710975A

  • Combined incremental learning method for image classification task

    CN119445327A