Interpretable open set identification method based on hierarchical visual semantic attribute space
By introducing a hierarchical visual semantic attribute space into the open set recognition method, a class hierarchical structure from fine-grained to coarse-grained is constructed, which solves the problem of insufficient discrimination of unknown categories in the prior art, and achieves higher interpretability and adaptability.
Patent Information
- Application Number
- CN202510022297.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-07
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-01-07
AI Technical Summary
When facing the goal of unknown categories, existing open set recognition methods can only determine them as ‘unknown’ and cannot provide more feature information, resulting in insufficient interpretability and difficulty in adapting to the real-world model deployment needs.
Using a method based on hierarchical visual semantic attribute space, the attribute matrix of fine-grained categories is constructed, the coarse-grained categories are clustered, the shared attribute frequency is calculated, and the mapping function is constructed to realize the hierarchical division of categories from fine-grained to coarse-grained, and attribute prediction and category prediction are carried out for the recognized image.
It can provide corresponding coarse-grained categories and fine-grained attributes for the goals of unknown categories, which significantly improves the interpretability of the model and adapts to the real-world model deployment needs.
Smart Images

Figure CN120071387A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of open set recognition, and more specifically, relates to an interpretable open set recognition method based on a hierarchical visual semantic attribute space. Background Art
[0002] Traditional deep learning models are mainly designed in a closed environment, assuming that the training and test data belong to the same categories and follow the independent and identically distributed (i.i.d.) assumption. However, in practical applications, the models in the training phase may encounter unseen categories. Therefore, how to effectively handle these unseen categories becomes a crucial issue. To address this challenge, open set recognition technology has emerged, aiming to identify and process samples of categories that the model has never seen.
[0003] Existing open set recognition methods mainly focus on single-level classification tasks, where the model receives an image as input and generates label predictions and their confidence scores. By comparing the confidence scores with a preset threshold, it is determined whether the input belongs to a known category or an unknown category. However, when faced with an unknown target, the existing methods can only classify the target as "unknown" and cannot provide more information about the characteristics of the unknown target, resulting in a lack of interpretability and difficulty in meeting the deployment requirements of models in the real world. Summary of the Invention
[0004] In view of the above deficiencies or improvement requirements of the prior art, the present invention provides an interpretable open set recognition method based on a hierarchical visual semantic attribute space, aiming to overcome the deficiency that existing open set recognition methods mainly focus on single-level classification tasks, which can only output single category information and cannot clarify the characteristics of unknown targets, resulting in insufficient interpretability of these models and difficulty in adapting to the model deployment requirements in the real world.
[0005] To achieve the above object, the present invention provides an interpretable open set recognition method based on a hierarchical visual semantic attribute space, including:
[0006] Collect the binary attributes of all fine-grained categories in the training set label space to construct an attribute space; in the attribute space, the attribute prototypes of all fine-grained categories form an attribute matrix; where all attributes belonging to the i-th fine-grained category are called the attribute prototypes of the i-th fine-grained category;
[0007] Cluster the attribute prototypes in the attribute matrix to obtain J clustering clusters and the attribute prototypes included in each clustering cluster; where each clustering cluster corresponds to a coarse-grained category, and each corresponding relationship from the fine-grained category to the coarse-grained category constitutes a mapping function;
[0008] According to the mapping function, calculate the frequency of the shared attributes of all fine-grained categories in each coarse-grained category. Take the attributes with higher shared frequencies as the coarse-grained attributes of the coarse-grained category, and the attributes with lower shared frequencies as the fine-grained attributes of the coarse-grained category; wherein, the shared attributes refer to the attributes existing in at least one attribute prototype.
[0009] Input the image to be recognized into the trained classification prediction model, and output the corresponding attribute prediction results and the fine-grained category prediction scores guided by the classifier. Based on the mapping function, map the fine-grained category prediction with the highest score to the coarse-grained category prediction of the image to be recognized, and mask the coarse-grained attributes included in the attribute prediction results based on the corresponding coarse-grained attributes, to obtain the fine-grained attributes predicted for the image to be recognized.
[0010] Further, after obtaining the fine-grained attributes predicted for the image to be recognized, it further includes:
[0011] According to the coarse-grained category prediction, allocate the image to be recognized to the corresponding fine-grained attribute space, calculate the similarity between the fine-grained attributes predicted for the image to be recognized and all the fine-grained attribute prototypes in the fine-grained attribute space, and splice the corresponding similarity values in sequence as the category prediction score S guided by the attributes of the image to be recognized. a ; wherein, all the fine-grained attributes of the coarse-grained category constitute the fine-grained attribute space.
[0012] For the category prediction score S guided by the attributes a and the fine-grained category prediction score S guided by the classifier c perform weighted processing to obtain the category prediction score S guided by the classifier-attributes. ca ; Determine whether the image to be recognized belongs to a known category or an unknown category according to the maximum value in the category prediction score S ca .
[0013] If it is a known category, then take the fine-grained category corresponding to the maximum value as the category prediction guided by the classifier-attributes of the image to be recognized.
[0014] If it is an unknown category, then output the coarse-grained category prediction of the image to be recognized and the corresponding fine-grained attributes.
[0015] Further, the similarity is the cosine distance between the fine-grained attributes predicted for the image to be recognized and each fine-grained attribute prototype in the corresponding fine-grained attribute space.
[0016] Based on the maximum value, use the maximum classifier-attributes guided open-set recognition score to determine whether the image to be recognized belongs to a known category or an unknown category.
[0017] Further, cluster the attribute prototypes in the attribute matrix to obtain J clusters and the attribute prototypes included in each cluster, including:
[0018] Initialize J cluster centers c j , where J represents the number of coarse-grained categories;
[0019] For each attribute prototype in the attribute matrix, use the method of minimizing the total within-cluster sum of squares for clustering to obtain J clusters and the attribute prototypes included in each cluster.
[0020] Further, according to the mapping function, calculate the frequency of the shared attributes of all fine-grained categories in each coarse-grained category, including:
[0021] For the j-th coarse-grained category, the frequency of its attribute appearance R j ={r j1 ,r j2 ,...,r jd ,...,r jD} is calculated as:
[0022]
[0023] where r jd represents the frequency of the appearance of attribute a d in the j-th coarse-grained category, is the number of fine-grained categories included in the j-th coarse-grained category, A i represents the attribute prototype of the i-th fine-grained category, and A i ={a i1 ,a i2 ,..,a id ,...,a iD}, a id represents the d-th attribute of the i-th fine-grained category, a id =1 indicates that the fine-grained category i has the attribute a d , a id =0 indicates that the fine-grained category i does not have the attribute a d ; D represents the dimension of the attribute space.
[0024] Further, take the attributes with higher shared frequencies as the coarse-grained attributes of this coarse-grained category, and take the attributes with lower shared frequencies as the fine-grained attributes of this coarse-grained category, including:
[0025] Based on the frequency of the shared attributes of all fine-grained categories in the j-th coarse-grained category, calculate the coarse-grained attribute prototype of the j-th coarse-grained category
[0026]
[0027] where θ is the threshold; the coarse-grained attribute prototypes included in the j-th coarse-grained category serve as all the coarse-grained attributes of the j-th coarse-grained category; The attributes included are used as all the coarse-grained attributes of the j-th coarse-grained category;
[0028] Calculate the fine-grained attribute prototypes of the i-th fine-grained category belonging to the j-th coarse-grained category
[0029]
[0030] where Mask j The elements equal to 1 in it are the fine-grained attributes of the j-th coarse-grained category, and the unit vectors formed by all the fine-grained attributes of the j-th coarse-grained category span the fine-grained attribute space; ⊙ represents the multiplication operator.
[0031] Furthermore, use the training images in the training set to train the classification prediction model to obtain the trained classification prediction model; wherein, the classification prediction model includes a data preprocessing module, a feature extractor, and a double-branch mapping network;
[0032] The data preprocessing module is used to adjust the training images to images suitable for the network size, perform normalization processing, and convert them into tensor form to obtain the preprocessed images;
[0033] The feature extractor is used to extract the visual features of the preprocessed images;
[0034] The double-branch mapping network includes a fine-grained category mapping module and an attribute mapping module; the fine-grained category mapping module is used to map the visual features of the images to the label space to obtain the classifier-guided fine-grained category prediction score p c ; the attribute mapping module is used to map the visual features of the images to the attribute space to obtain the attribute prediction result D represents the dimension of the attribute space.
[0035] The present invention also provides an interpretable open-set recognition system based on a hierarchical visual semantic attribute space, including a computer-readable storage medium and a processor;
[0036] The computer-readable storage medium is used to store executable instructions;
[0037] The processor is used to read the executable instructions stored in the computer-readable storage medium and execute the interpretable open-set recognition method described in any one of the above.
[0038] The present invention also provides a computer-readable storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, it implements the interpretable open-set recognition method described in any one of the above.
[0039] The present invention also provides a computer program product, including a computer program, which when running on a computer, causes the computer to execute the interpretable open-set recognition method described in any one of the above.
[0040] Generally speaking, through the above technical solutions conceived by the present invention, the following beneficial effects can be achieved:
[0041] (1) In the interpretable open-set recognition method based on the hierarchical visual semantic attribute space of the present invention, by hierarchically dividing fine-grained categories and their corresponding attribute prototypes into a coarse-grained hierarchical structure, this hierarchical structure realizes the hierarchical division of categories from fine-grained to coarse-grained. The constructed mapping function reflects the mapping relationship between coarse-grained categories, fine-grained categories, and each attribute prototype. For each coarse-grained level, according to the frequency of shared attributes of all fine-grained categories in each coarse-grained level, a hierarchical relationship is assigned to each attribute (whether the attribute is a coarse-grained attribute or a fine-grained attribute belonging to the corresponding coarse-grained category), thereby obtaining the coarse-grained attributes and fine-grained attribute spaces corresponding to each coarse-grained category. Based on this mapping function and the coarse-grained attributes and fine-grained attribute spaces corresponding to each coarse-grained category, for the image to be recognized, whether it is a known category or an unknown category, it can at least be divided into the corresponding coarse-grained category and the corresponding fine-grained attributes. Compared with the traditional method that can only identify an unknown category as "unknown", the method of the present invention can also provide the corresponding coarse-grained category and the corresponding fine-grained attributes of the unknown category, which has significant superiority compared with the traditional open-set recognition technology. It greatly improves the interpretability of the model and can meet the requirements of model deployment in the real world.
[0042] (2) Further, by calculating the similarity between the fine-grained attributes of the image to be recognized and all fine-grained attribute prototypes in the fine-grained attribute space, the category prediction score guided by the attributes of the sample is obtained, and the category prediction score guided by the classifier and the category prediction score guided by the attributes are weighted, and the fine-grained category corresponding to the maximum value in the prediction scores is selected as the classifier-attribute-guided category prediction of the image to be recognized. That is, for known categories, the method of the present invention can also accurately identify the corresponding fine-grained categories.
[0043] (3) Preferably, the double-branch mapping network in the present invention can be effectively connected after any feature extractor, and has the characteristics of simple structure and reduced parameters.
[0044] In summary, the method of the present invention can not only accurately identify whether the target belongs to a seen category or an unseen category, but also clarify its coarse-grained category and specific attributes. This method has significant superiority compared to traditional open-set recognition techniques. It greatly improves the interpretability of the model, enabling the model to not only give a simple discrimination of the unknown category when facing unknown targets, but also provide more valuable information, thus providing strong technical support for the effective deployment of the model in complex and ever-changing real-world scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0045] Figure 1 It is a flowchart of the interpretable open-set recognition method based on the hierarchical visual semantic attribute space in an embodiment of the present invention;
[0046] Figure 2 It is an overall framework diagram of the interpretable open-set recognition method based on the hierarchical visual semantic attribute space in an embodiment of the present invention;
[0047] Figure 3 It is a comparison of the open-set recognition experimental results of the interpretable open-set recognition method based on the hierarchical visual semantic attribute space in an embodiment of the present invention and traditional methods on different data sets. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0048] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0049] Embodiment 1
[0050] As Figure 1 , Figure 2 shown, an embodiment of the present invention discloses an interpretable open-set recognition method based on a hierarchical visual semantic attribute space. The method includes the following steps:
[0051] Step 1: Collect the binary attributes of all fine-grained categories in the training set label space to construct an attribute space. In this attribute space, the attribute prototypes of all fine-grained categories covered form an attribute matrix; among them, all attributes belonging to the i-th fine-grained category are called the attribute prototypes of the i-th fine-grained category.
[0052] In an embodiment of the present invention, the constructed attribute space is denoted as where a nRepresents a single binary attribute in the attribute space, D represents the dimension of the attribute space, i.e., the total number of attributes, and the attribute prototype A of each fine-grained category i Is located in the attribute space, i represents the fine-grained category index, i ∈ {1, 2, …, K}, and the attribute prototype A i (A i ={a i1 , a i2 ,.., a id ,..., a iD}) refers to all attributes belonging to the i-th fine-grained category and is a subset of, a id = 1 indicates that the fine-grained category i has the attribute a d , and a id = 0 indicates that the fine-grained category i does not have the attribute a d . For a classification task with K fine-grained categories, the attribute prototypes A of all fine-grained categories i Form an attribute matrix In this embodiment, the element A of the attribute matrix id Represents the d-th attribute of the i-th fine-grained category, i ∈ {1, 2, …, K}, d ∈ {1, 2, …, D}.
[0053] As a preferred implementation, the binary attributes possessed by all fine-grained categories in the training set can be collected by manual annotation or querying a large language model to ensure the accuracy and comprehensiveness of the attribute data.
[0054] Step 2: Cluster all the attribute prototypes in the attribute matrix to obtain J clusters and the attribute prototypes included in each cluster; where each cluster corresponds to a coarse-grained category, thus realizing the hierarchical division of categories from fine-grained to coarse-grained. Each corresponding relationship from the fine-grained category to the coarse-grained category constitutes a mapping function Represents the label space of the coarse-grained category.
[0055] In the embodiment of the present invention, the method of minimizing the total within-cluster sum of squares is used to cluster all the attribute prototypes in the attribute matrix to obtain J clusters and the attribute prototypes included in each cluster; specifically including: initializing J cluster centers c j , representing J levels, that is, J coarse-grained categories; for each attribute prototype (in this embodiment, each row in the attribute matrix, all attributes belonging to the i-th category), minimize the total within-cluster sum of squares (TWCSS), and its expression is:[[]]
[0056]
[0057] where J is the number of clusters, Cj represents the j-th clustering cluster, A i represents the i-th category attribute prototype in the clustering cluster C j Among them.
[0058] By clustering all the attribute prototypes in the attribute matrix, J clustering clusters are obtained, as well as the attribute prototypes included in each clustering cluster; each cluster in the clustering represents a coarse-grained hierarchical structure (representing a coarse-grained category; for example: for the fine-grained categories of bighead carp and grass carp, their corresponding coarse-grained category is fish; for the fine-grained categories of dog, cat, and horse, their corresponding coarse-grained category is mammal). For all the labels (fine-grained categories) in the label space According to the attribute clustering result, a coarse-grained label is assigned to each fine-grained category y, representing the label space of the coarse-grained category, and a mapping function composed of each corresponding relationship from the fine-grained category to the coarse-grained category is obtained Mapping function reflects the mapping relationship among the coarse-grained category, the fine-grained category, and the attribute prototype.
[0059] Step 3: According to the mapping function Calculate the frequency of the attributes shared by all the fine-grained categories in each coarse-grained category. The attributes with higher sharing frequencies are used as the coarse-grained attributes of the coarse-grained category, and the attributes with lower sharing frequencies are used as the fine-grained attributes of the coarse-grained category, so as to obtain the coarse-grained attributes and the fine-grained attribute space corresponding to each coarse-grained category; among them, the shared attributes refer to the attributes existing in at least one attribute prototype.
[0060] In the embodiment of the present invention, according to the hierarchical division of the coarse-grained category, the frequency of occurrence of all the attributes in each coarse-grained category is calculated. For the j-th coarse-grained category, the frequency of occurrence R j ={r j1 ,r j2 ,...,r jd ,...,r jD} is calculated as:
[0061]
[0062] Among them, r jd represents the frequency of occurrence of the attribute a d in the j-th coarse-grained category, is the number of fine-grained categories included in the j-th coarse-grained category, A i ={a i1 ,a i2 ,..,a id ,...,a iD}。Then, the coarse-grained attribute prototype of the j-th coarse-grained category can be obtained
[0063]
[0064] where θ is a threshold value, which is determined empirically; a higher value indicates that the attribute needs to be shared by more fine-grained categories to be classified as a coarse-grained attribute; r jd represents the frequency of occurrence of attribute a d in the j-th coarse-grained category. The attributes included in the coarse-grained attribute prototype of the j-th coarse-grained category are used as all the coarse-grained attributes of the j-th coarse-grained category.
[0065] For the fine-grained category i and its corresponding coarse-grained category j, mask its coarse-grained attributes to calculate its fine-grained attribute prototype (representing the attributes unique to the fine-grained category i):
[0066]
[0067] where m d = 1 - a jd , and the unit vectors formed by the elements equal to 1 in Mask j span the fine-grained attribute space (i.e., the attributes that are 0 in the coarse-grained attribute are all the fine-grained attributes of the j-th coarse-grained category, and all the fine-grained attributes constitute the fine-grained attribute space of the j-th coarse-grained category); ⊙ represents the multiplication operator.
[0068] By traversing all J coarse-grained categories, the corresponding J fine-grained attribute spaces can be obtained. Each fine-grained attribute space contains attribute prototypes ( categories); is the number of categories included in the j-th coarse-grained category.
[0069] Step 4: Input the image to be recognized into the trained classification prediction model, and output the corresponding attribute prediction results and the fine-grained category prediction scores guided by the classifier; based on the mapping function map the fine-grained category prediction with the highest score to the coarse-grained category prediction of the image to be recognized; based on this coarse-grained category prediction, assign the image to be recognized to the corresponding fine-grained attribute space, and mask the coarse-grained attributes included in the attribute prediction results according to the coarse-grained attributes of this coarse-grained category prediction to obtain the fine-grained attributes in the attribute prediction results (prediction of the image to be recognized)
[0070] Furthermore, calculate the fine-grained attributes predicted for the image to be recognized The similarity with all the fine-grained attribute prototypes in the fine-grained attribute space, and the corresponding similarity values are concatenated in sequence as the class prediction score S guided by the attributes of the image to be recognized. a ; The class prediction score S guided by this attribute a is weighted with the fine-grained class prediction score S guided by the corresponding classifier c to obtain the class prediction score S guided by the classifier-attribute. ca ; The fine-grained class corresponding to the maximum value in the class prediction score is used as the class prediction guided by the classifier-attribute of the image to be recognized, and it is determined whether the image to be recognized belongs to a known class or an unknown class according to this maximum value, thus completing the open-set recognition.
[0071] As a preferred implementation, in step four, the training images in the training set are used to train the classification prediction model to obtain the trained classification prediction model. In the embodiment of the present invention, the classification prediction model includes a data preprocessing module, a feature extractor Φ F and a dual-branch mapping network.
[0072] The data preprocessing module is used to adjust the input training image to an image suitable for the network size, and perform normalization processing, and convert it into a tensor form to obtain the preprocessed image. In the embodiment of the present invention, the data preprocessing module adjusts the input image to a size of 224x224, and performs normalization processing and converts it into a tensor form to ensure the standardization and consistency of the data. The feature extractor Φ F is used to extract the visual features of the preprocessed image.
[0073] The dual-branch mapping network includes a fine-grained class mapping module Φ C and an attribute mapping module Φ A ; These two modules share the feature extractor of the model, and each is composed of a mapping layer and a fully connected layer connected in series to obtain the output information of the attribute prediction result and the fine-grained class prediction score p c guided by the classifier:
[0074]
[0075] Specifically, the fine-grained class mapping module is used to map the visual features of the image to the label space to obtain the fine-grained class prediction score p c guided by the classifier; the attribute mapping module is used to map the visual features of the image to the attribute space to obtain the attribute prediction result The dual-branch mapping network in the embodiment of the present invention can be effectively connected after any feature extractor, and has the characteristics of simple structure and concise parameters.
[0076] As a preferred implementation manner, in the embodiment of the present invention, a cross-entropy loss function is adopted to calculate the category prediction loss of the training image and a binary cross-entropy loss function is used to calculate the attribute prediction loss of the training image. The parameters of the feature extractor and the dual-branch mapping network are updated through the backpropagation algorithm. Wherein, the loss function is calculated as shown in the following formula:
[0077]
[0078] In the formula, λ is a hyperparameter for controlling the weight; y and A represent the true fine-grained category and attribute of the current training image.
[0079] As a preferred implementation manner, the image to be recognized is input into the above-mentioned trained classification prediction model, and the corresponding attribute prediction result and the classifier-guided fine-grained category prediction score are output. In this embodiment, represents the test sample set. The category prediction result with the highest score among the classifier-guided category prediction scores is selected, and the coarse-grained category prediction j of the image to be recognized is obtained through the coarse-grained mapping function :
[0080]
[0081] Wherein, argmax represents the index for taking the maximum value.
[0082] Based on this coarse-grained category prediction, the image to be recognized is assigned to the fine-grained attribute space corresponding to the j-th coarse-grained category, and the coarse-grained attribute information included in the attribute prediction result is masked to obtain the fine-grained attribute predicted for the image to be recognized
[0083] Then, the fine-grained attribute predicted for the image to be recognized is measured with all the fine-grained attribute prototypes in the fine-grained attribute space; in this embodiment, the similarity calculation method is the cosine distance between two attribute vectors. After splicing each similarity value, the attribute-guided category prediction score S of the sample is obtained a :
[0084]
[0085] Wherein, Concat[k 1 ,k 2 ,…,k n represents concatenating the elements k 1 ,k 2,…, k n are concatenated; represents the fine-grained attributes predicted for the image to be recognized and the fine-grained attribute prototypes in the fine-grained attribute space The cosine distance between them is used to characterize and The similarity between them.
[0086] Assume that β is the attribute affinity factor (assigned according to experience), and the classifier-attribute-guided class prediction score S ca can be obtained: S ca = S c + βS a .
[0087] The finally predicted class is the fine-grained class corresponding to the maximum value of S ca : The maximum classifier-attribute-guided open-set recognition score (MCAS) is used to determine whether the image to be recognized belongs to a known class or an unknown class (argmax{S ca}):
[0088] MCAS = max(S ca ).
[0089] Among them, MCAS represents the maximum value of S ca ; when MCAS exceeds the preset threshold, the predicted class is a known class, and the corresponding fine-grained class is Otherwise, the predicted class is an unknown class, and at the same time, the coarse-grained class of this unknown class and the corresponding fine-grained attributes can be obtained.
[0090] The interpretable open-set recognition method based on the hierarchical visual semantic attribute space of the present invention hierarchically divides the fine-grained classes and the corresponding attribute prototypes into a coarse-grained hierarchical structure. This hierarchical structure realizes the hierarchical division of classes from fine-grained to coarse-grained, and the constructed mapping function It reflects the mapping relationship among coarse-grained categories, fine-grained categories, and the prototypes of each attribute. For each coarse-grained level, according to the frequency of the shared attributes of all fine-grained categories in each coarse-grained level, a hierarchical relationship is assigned to each attribute (whether the attribute is a coarse-grained attribute or a fine-grained attribute belonging to the corresponding coarse-grained category), thereby obtaining the coarse-grained attributes and fine-grained attribute spaces corresponding to each coarse-grained category. Based on this mapping function and the coarse-grained attributes and fine-grained attribute spaces corresponding to each coarse-grained category, for the image to be recognized, whether it is a known category or an unknown category, it can at least be divided into the corresponding coarse-grained category and the corresponding fine-grained attributes. Compared with the traditional method that can only identify an unknown category as "unknown", the method of the present invention can also provide the corresponding coarse-grained category and the corresponding fine-grained attributes of the unknown category, and has significant superiority compared with the traditional open-set recognition technology. It greatly improves the interpretability of the model and can meet the model deployment requirements in the real world.
[0091] Furthermore, by calculating the fine-grained attributes of the image to be recognized and the similarity between all the fine-grained attribute prototypes in the fine-grained attribute space, the category prediction score guided by the attributes of the sample is obtained. The category prediction score guided by the classifier and the category prediction score guided by the attributes are weighted, and the fine-grained category corresponding to the maximum value in the prediction scores is selected as the classifier-attribute guided category prediction of the image to be recognized. That is, for known categories, the method of the present invention can also accurately identify the corresponding fine-grained categories.
[0092] As Figure 3 shown, it shows the comparison of the open-set recognition performance of the method proposed by the present invention and the traditional method on the CUB, FGVC, LAD, and AWA datasets. Using the AUROC and OSCR metrics commonly used in the field of open-set recognition, on the four datasets, the method proposed by the present invention has improved by 1.81% in the AUROC metric and 2.33% in the OSCR metric.
[0093] The interpretable open-set recognition method based on the hierarchical visual semantic attribute space designed by the present invention additionally utilizes the hierarchical information of attributes and categories for assistance during the open-set recognition process. Compared with the defect of the existing method that cannot provide information about unknown samples, the method proposed by the present invention can not only mark unknown samples as unknown categories, but also provide their coarse-grained category information and related fine-grained attributes. And experiments prove that compared with the existing methods, this method has a significant improvement in open-set recognition performance, and this method has the advantages of high precision and strong interpretability.
[0094] Example 2
[0095] An embodiment of the present invention provides an interpretable open-set recognition system based on a hierarchical visual semantic attribute space, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the steps of the interpretable open-set recognition method based on the hierarchical visual semantic attribute space in Embodiment 1 above are implemented.
[0096] The related technical solutions are the same as above and will not be elaborated here.
[0097] Embodiment 3
[0098] An embodiment of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the interpretable open-set recognition method based on the hierarchical visual semantic attribute space in Embodiment 1 above are implemented.
[0099] Specifically, the memory may include a high-speed random access memory, and may also include a non-volatile memory, such as a hard disk, a memory, a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.
[0100] The related technical solutions are the same as above and will not be elaborated here.
[0101] Embodiment 4
[0102] An embodiment of the present application provides a computer program product, including a computer program. When the computer program runs on a computer, the computer is caused to execute the steps of the interpretable open-set recognition method based on the hierarchical visual semantic attribute space in Embodiment 1 above.
[0103] The related technical solutions are the same as above and will not be elaborated here.
[0104] Those skilled in the art can easily understand that the above are only preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention should be included in the protection scope of the present invention.
Claims
1. An interpretable open set recognition method based on a hierarchical visual semantic attribute space, characterized in that: include: Collect training set label space The binary attributes of all fine-grained categories in the ith fine-grained category are used to construct an attribute space; in the attribute space, the attribute prototypes of all fine-grained categories form an attribute matrix; wherein all the attributes belonging to the ith fine-grained category are called the attribute prototypes of the ith fine-grained category; Clustering the attribute prototypes in the attribute matrix to obtain J clusters and the attribute prototypes contained in each cluster; wherein each cluster corresponds to a coarse-grained category, and each corresponding relationship from the fine-grained category to the coarse-grained category constitutes a mapping function; According to the mapping function, the frequency of shared attributes of all fine-grained categories in each coarse-grained category is calculated, and the attribute with higher sharing frequency is used as the coarse-grained attribute of the coarse-grained category, and the attribute with lower sharing frequency is used as the fine-grained attribute of the coarse-grained category; wherein the shared attribute refers to an attribute existing in at least one attribute prototype; The image to be identified is input into the trained classification prediction model, and the corresponding attribute prediction results and the fine-grained category prediction scores guided by the classifier are output; based on the mapping function, the fine-grained category prediction with the highest score is mapped to the coarse-grained category prediction of the image to be identified, and the coarse-grained attributes contained in the attribute prediction results are masked based on the coarse-grained attributes corresponding to the coarse-grained category prediction to obtain the fine-grained attributes predicted for the image to be identified.
2. The interpretable open set identification method according to claim 1, characterized in that: After obtaining the fine-grained attributes of the image to be recognized, it also includes: According to the coarse-grained category prediction, the image to be identified is assigned to the corresponding fine-grained attribute space, and the similarity between the predicted fine-grained attribute of the image to be identified and all fine-grained attribute prototypes in the fine-grained attribute space is calculated, and the corresponding similarity values are sequentially spliced as the category prediction score S guided by the attribute of the image to be identified a ; Wherein, all fine-grained attributes of the coarse-grained category constitute the fine-grained attribute space; The attribute-guided category prediction score S a The fine-grained category prediction score S guided by the classifier c Perform weighted processing to obtain the classifier-attribute guided category prediction score S ca ; According to the category prediction score S ca The maximum value in determines whether the image to be identified belongs to a known category or an unknown category; If it is a known category, the fine-grained category corresponding to the maximum value is As a classifier of the image to be identified - attribute-guided category prediction; If it is an unknown category, the coarse-grained category prediction of the image to be identified and the corresponding fine-grained attributes are output.
3. The interpretable open set identification method according to claim 2, characterized in that: The similarity is the cosine distance between the predicted fine-grained attribute of the image to be identified and each fine-grained attribute prototype in the corresponding fine-grained attribute space; Based on the maximum value, the maximum classifier-attribute guided open set recognition score is used to determine whether the image to be recognized belongs to a known category or an unknown category.
4. The interpretable open set identification method according to any one of claims 1 to 3, characterized in that: The attribute prototypes in the attribute matrix are clustered to obtain J clusters and the attribute prototypes contained in each cluster, including: Initialize J cluster centers c j , where J represents the number of coarse-grained categories; For each attribute prototype in the attribute matrix, clustering is performed by minimizing the total intra-cluster square sum to obtain J clusters and the attribute prototypes contained in each cluster.
5. The interpretable open set identification method according to any one of claims 1 to 3, characterized in that: According to the mapping function, the frequency of shared attributes of all fine-grained categories in each coarse-grained category is calculated, including: For the jth coarse-grained category, its attribute occurrence frequency R j = {r j1 ,r j2 ,...,r jd ,...,r jD } is calculated as: Among them, r jd represents the attribute a in the jth coarse-grained category d The frequency of occurrence, is the number of fine-grained categories contained in the jth coarse-grained category, A i represents the attribute prototype of the i-th fine-grained category, and A i ={a i1 ,a i2 ,..,a id ,...,a iD }, a id represents the dth attribute of the i-th fine-grained category, a id =1 means that fine-grained category i has attribute a d , a id = 0 means that attribute a does not exist in fine-grained category i d ; D represents the dimension of the attribute space.
6. The interpretable open set identification method according to claim 5, characterized in that: The attributes with higher sharing frequency are taken as the coarse-grained attributes of the coarse-grained category, and the attributes with lower sharing frequency are taken as the fine-grained attributes of the coarse-grained category, including: Based on the frequency of shared attributes of all fine-grained categories in the j-th coarse-grained category, calculate the coarse-grained attribute prototype of the j-th coarse-grained category Where θ is the threshold; the coarse-grained attribute prototype of the jth coarse-grained category The included attributes are all the coarse-grained attributes of the j-th coarse-grained category; Compute the fine-grained attribute prototype of the i-th fine-grained category belonging to the j-th coarse-grained category in, Mask j The elements equal to 1 in are the fine-grained attributes of the j-th coarse-grained category, and the unit vector formed by all the fine-grained attributes of the j-th coarse-grained category forms the fine-grained attribute space; ⊙ Represents the multiplication operator.
7. The interpretable open set identification method according to claim 1, characterized in that: The classification prediction model is trained using training images in the training set to obtain the trained classification prediction model; wherein the classification prediction model includes a data preprocessing module, a feature extractor, and a dual-branch mapping network; The data preprocessing module is used to adjust the training image to an image suitable for the network size, perform normalization processing, convert it into a tensor form, and obtain a preprocessed image; The feature extractor is used to extract visual features of the preprocessed image; The dual-branch mapping network includes a fine-grained category mapping module and an attribute mapping module; the fine-grained category mapping module is used to map the visual features of the image to the label space Get the classifier-guided fine-grained category prediction score p c The attribute mapping module is used to map the visual features of the image to the attribute space to obtain the attribute prediction result. D represents the dimension of the attribute space.
8. An interpretable open set recognition system based on a hierarchical visual semantic attribute space, characterized in that: comprising a computer readable storage medium and a processor; The computer-readable storage medium is used to store executable instructions; The processor is used to read the executable instructions stored in the computer-readable storage medium to execute the interpretable open set identification method described in any one of claims 1-7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the explainable open set identification method as described in any one of claims 1 to 7 is implemented.
10. A computer program product, characterized in that It includes a computer program, which, when running on a computer, enables the computer to execute the interpretable open set identification method described in any one of claims 1 to 7.
Citation Information
Patent Citations
Fine-grained image classification method based on image attribute active learning
CN112528058A
Semantic exploration-based open set action recognition method
CN116129333A
Image classification method and device
CN116310589A
Open set identification method and system based on untrained open set simulator
CN117689951A
Hierarchical image classification method and device based on semantic knowledge guidance
CN117975151A
Cited By
Generalized classification method and device for domain awareness based on large language model
CN122286452A