A multi-task attribute scene recognition method based on category labels and attribute annotations

Through the multi-task attribute scene recognition network MASR, combined with category tags and attribute annotations, the problem of insufficient utilization of attribute information in the prior art is solved, and more efficient and accurate scene recognition is achieved.

CN114241380BActive Publication Date: 2025-06-06ZHEJIANG LAB
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111547952.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-16
Publication Date
2025-06-06
Estimated Expiration
2041-12-16

AI Technical Summary

Technical Problem

The prior art is difficult to effectively utilize attribute information in scene recognition, resulting in insufficient discrimination ability of visually similar images, and manual annotation of object attributes is time-consuming and challenging.

Method used

A multi-task attribute scene recognition method based on category labels and attribute annotation is proposed. The network MASR is used to identify the network MASR through multi-task attribute scenes, and features are extracted using CNN, combined with attribute labeling strategies and object filtering logic, classification prediction and attribute probability prediction are performed, and attribute tasks are accelerated through attribute task loss function and attribute layer.

Benefits of technology

Reduced manual supervision and intervention, improved task efficiency, enhanced credibility of attribute prediction, and achieved more discernible representation and competitive recognition performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114241380B_ABST
    Figure CN114241380B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of scene recognition technology, and in particular to a multi-task attribute scene recognition method based on category labels and attribute annotations. Based on a multi-task attribute scene recognition network MASR, object attribute scores are used and calculated to screen and simplify object attributes, simplify the attribute annotation process, and reduce the training bias caused by data. In addition, an attribute loss function and an attribute layer are designed and used in the MASR network to make full use of the above-mentioned attribute features after screening and simplification, and re-weight the object attributes according to the importance level of the object detection score. The present invention effectively annotates the attribute labels of four large-scale data sets. Experimental results show that compared with the most advanced methods, the present invention learns more discriminative representations and achieves competitive recognition performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of scene recognition, and in particular to a multi-task attribute scene recognition method based on category labels and attribute annotations. Background Art

[0002] Scene recognition, also known as scene classification, is a high-level computer vision task that aims to determine the overall scene category by emphasizing the understanding of its global properties. Contextual information such as semantic segmentation, structural layout, and object attributes are key to improving scene recognition accuracy. In particular, semantic attributes are used to achieve richer scene descriptions, while semantic segmentation can express the spatial relationship between objects in the scene. Similarly, attribute information is very important for distinguishing similar images and improving scene recognition performance. Using visual features alone, it is difficult to distinguish between visually similar images. On the other hand, attributes are semantically descriptive across classes. However, extracting object attributes or constructing effective semantic representations has proven to be very challenging, especially when object attribute annotations have to be performed manually. Semantic segmentation is also challenging given that the task of labeling scenes with accurate per-pixel labels is time-consuming. Summary of the invention

[0003] In order to solve the above technical problems existing in the prior art, the present invention proposes a multi-task attribute scene recognition method based on category labels and attribute annotations, and its specific technical solution is as follows:

[0004] A multi-task attribute scene recognition method based on category labels and attribute annotations, based on a multi-task attribute scene recognition network MASR, specifically includes the following steps:

[0005] 1) Given a scene image x i , using CNN network to extract its features θ I is the CNN network parameter;

[0006] 2) Use attribute annotation strategy to calculate object attribute scores, and use the object attribute scores to identify v i The attribute objects in are streamlined according to the object filtering logic;

[0007] 3) The simplified feature v i Input to the fully connected layer L |K| Perform classification prediction, where K is the number of scene classification categories; at the same time, the simplified feature v i Input fully connected layer L |A| Predicted attribute probability p att , where A is the detected attribute set;

[0008] 4) The predicted attribute probability p attCompared with the attribute representation learned from external data alone, the input attribute layer is i Redistribute the weights and use the attribute task loss function to accelerate the attribute layer tasks;

[0009] 5) The corrected v i Feedback to the fully connected layer L |K| , to improve the performance of multi-task attribute scene recognition tasks.

[0010] Furthermore, the attribute labeling strategy is to combine the two probability distributions p s With p t Simply merge and use the object detection score P as the confidence score, i.e. the object attribute score, as follows:

[0011] Collect object attributes and context information from COCO Object and COCO Panoptic datasets, and process the stuff and thing types independently. Let S and T be the sets of stuff and thing respectively, and F s With F t is the pre-trained CNN model for each task, let {x 1 ,x 2 ,...,x n}∈X represents a scene-centric dataset with only category labels, using F on X s With F t Predict the distribution on S and T, p s =F s (X) and p t =F t (X), where p s ∈R |S| With p t ∈R |T| They are the probability distribution predictions of S and T respectively. Given a data set X, the final stuff+thing prediction P∈R |S|+|T| , defined as P = p on a given scene dataset s ∪p t , where P does not increase to 1 and does not represent a probability distribution. For two probability distributions p s With p t Take the average to combine them, Among them, S and T do not always have intersections, representing different data sources.

[0012] Furthermore, the object screening is to further screen the objects in S and T according to the object detection score and the object frequency, specifically including:

[0013] Based on object detection score: Object instances with object detection scores less than the threshold are discarded. Only objects with object detection scores higher than the threshold are selected as scene attributes. In this process, P is redefined as:

[0014]

[0015] Where ξ is the threshold, when the detection score is 0, the object is considered not to exist in the scene;

[0016] Based on object frequency: We further consider the attribute frequency of a given scene category and remove uncommon objects. For each category c, we define the relative attribute frequency as the number of non-zero fractions covering the category images if {a 1 ,a 2 ,...,a m}∈A c is the detection attribute set of c, the optimal Defined as:

[0017]

[0018] where f c (a j ) is the value of a given category c j The relative frequency of the attribute, β is the minimum frequency, is the final attribute list of c.

[0019] Furthermore, the attribute task loss function is specifically:

[0020] Define the multi-class cross entropy loss function:

[0021]

[0022] Among them, p att (x i ,j) is the training sample x i The predicted category probability on the j-th attribute of is a property annotation, which is defined as:

[0023]

[0024] Let’s introduce the regularization term β j , which reflects the relative frequency of the jth attribute in the training data, that is, the ratio of its positive and negative attribute labels. Formula (3) can be transformed into:

[0025]

[0026] where ||a j|| is the number of samples holding the k-th category label of the j-th attribute, that is, the size of the j-th attribute of the k-th scene category, where classifiers of different attribute features are not shared.

[0027] Furthermore, the attribute layer is specifically:

[0028] A layer is introduced to reweight the attributes according to the detection score, which consists of a series of linear transformations that aggregate all attribute information into a vector v i In, use Represents the attribute classifier f A The attribute scores of The confidence score c i for:

[0029]

[0030] Where σ is the sigmoid activation function, W * ∈R m×m With b i ∈R m×1 is a trainable parameter, v i By c i with a i Element-wise multiplication.

[0031] The beneficial effects are: the present invention first proposes a partially supervised annotation strategy, which reduces manual supervision and intervention and improves the efficiency of the task; adopts an object screening logic based on a confidence score mechanism to improve the low credibility of attribute predictions caused by training data bias, and compared with the most advanced methods, the multi-task attribute scene recognition network MASR of the present invention learns more discriminative representations and achieves competitive recognition performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] Figure 1 is an overview diagram of the MASR architecture of the present invention;

[0033] Figure 2 It is a graph of the concatenated predictions obtained from each prediction before the attribute reweighting layer is applied to the sigmoid. DETAILED DESCRIPTION

[0034] In order to make the purpose, technical scheme and technical effect of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the accompanying drawings.

[0035] The present invention provides a multi-task attribute scene recognition method based on category labels and attribute annotations, based on a multi-task attribute scene recognition MASR network, such as Figure 1As shown, attribute information is obtained from a pre-trained object-centric model, and the obtained attribute information is used to support the learning of CNN features through regularization loss and reweighting layers. The recognition method for the multi-task attribute scene specifically includes the following steps:

[0036] 1) Given a scene image x i , using a CNN-like network to extract its features θ i is the CNN network parameter;

[0037] 2) Use the attribute annotation strategy to calculate the object attribute score, and use the object attribute score to identify v i The attribute objects in are streamlined according to the object screening logic. This step simplifies the attribute annotation process and reduces the training bias caused by data.

[0038] 3) The simplified feature v i Input to the fully connected layer L |K| Perform classification prediction, where K is the number of scene classification categories; at the same time, the simplified feature v i Input fully connected layer L |A| Predicted attribute probability p att , where A is the detected attribute set;

[0039] 4) The predicted attribute probability p att Compared with the attribute representation learned from external data alone, the input attribute layer is i Redistribute the weights and use the attribute task loss function to accelerate the attribute layer tasks;

[0040] 5) The corrected v i Feedback to the fully connected layer L |K| , to improve the performance of multi-task attribute scene recognition tasks.

[0041] The attribute labeling strategy is specifically as follows:

[0042] First, the present invention collects object attributes and contextual information from two popular object-centric datasets: COCO Object and COCOPanoptic, and independently processes the stuff and thing types to improve scene recognition capabilities. The instance examples included are shown in Table 1.

[0043] Table 1:

[0044] Groups Attributes Things bottle,cup,apple,sheep,dog,suitca|se,tv,toilet... Stuff sea,river,road,sand,snow,wall,window,wall...

[0045] Let S and T be the sets of stuff and thing respectively, F s With F tis a pre-trained CNN model for each task. Let {x 1 ,x 2 ,...,x n}∈X represents a scene-centric dataset with only category labels. The goal of this paper is to use F s With F t Predict the distribution on S and T, p s =F s (X) and p t =F t (X). Among them, p s ∈R |S| With p t ∈R |T| are the probability distribution predictions of S and T respectively. Given a dataset X, the final stuff+thing prediction P∈R |S|+|T| , defined as P = p on a given scene dataset s ∪p t , where P does not increase to 1 and does not represent a probability distribution. For two probability distributions p s With p t Take the average to combine them, Among them, S and T do not always have intersections, and they are usually used to represent different data sources. In general, the present invention will s With p t Simply merge and use the object detection score P as the confidence score.

[0046] The object screening is specifically:

[0047] When there is too much information about the attributes and relationships of the objects, it is not conducive to the scene recognition task. To overcome this problem, the present invention further screens the objects in S and T according to the object detection score and the object frequency, specifically including:

[0048] Based on object detection score: Object instances with object detection scores less than the threshold are discarded. Only objects with object detection scores higher than the threshold are selected as scene attributes. In this process, P is redefined as:

[0049]

[0050] where ξ is the threshold, when the detection score is 0, the object is considered not present in the scene.

[0051] Based on object frequency: We further consider the attribute frequency of a given scene category and remove uncommon objects. For each category c, we define the relative attribute frequency as the number of non-zero fractions covering the category images. If {a 1 ,a 2 ,...,a m}∈A c is the detection attribute set of c, the optimal Defined as:

[0052]

[0053] where f c (a j ) is the value of a given category c j The relative frequency of the attribute, β is the minimum frequency, is the final attribute list of c.

[0054] The attribute task loss function is specifically:

[0055] Since the attributes are not completely mutually exclusive, the prediction of multiple attributes is a multi-label classification problem. The layer structure of the predicted attribute is different from the traditional single-label classification layer containing the loss function. In order to make the attribute layer adaptable to the multi-label classification problem, the present invention proposes a multi-class cross entropy loss function defined as follows:

[0056]

[0057] Among them, p att (x i ,j) is the training sample x i The predicted category probability on the j-th attribute of is a property annotation, which is defined as:

[0058]

[0059] The loss in formula (3) is usually affected by the data skew problem of training data and cannot be simply compensated by data sampling, because balancing the frequency of occurrence of one attribute will change other attributes. To solve this problem, the present invention introduces a regularization term β j , which reflects the relative frequency of the jth attribute in the training data, that is, the ratio of its positive and negative attribute labels. Formula (3) can be transformed into:

[0060]

[0061] where ||a j || is the number of samples holding the k-th category label of the j-th attribute, that is, the size of the j-th attribute of the k-th scene category, where classifiers of different attribute features are not shared.

[0062] The attribute layer is specifically:

[0063] Since the attribute representation is learned on separate data, it is expected that some attributes are more important than others. The present invention introduces a layer that reweights the attributes according to the detection score, which consists of a series of linear transformations that aggregate all attribute information into a vector v i In, use Represents the attribute classifier f A The attribute scores of The confidence score c i for:

[0064]

[0065] Where σ is the sigmoid activation function, W * ∈R m×m With b i ∈R m×1 is a trainable parameter. i By c i with a i The above operations constitute the attribute reweighting layer ARL, and its operation process is as follows: Figure 2 shown.

[0066] The above is only a preferred implementation case of the present invention and does not limit the present invention in any form. Although the implementation process of the present invention is described in detail above, for those familiar with the art, they can still modify the technical solutions recorded in the above examples, or replace some of the technical features therein with equivalents. All modifications, equivalent replacements, etc. made within the spirit and principles of the present invention should be included in the protection scope of the present invention.

Claims

1. A multi-task attribute scene recognition method based on category labels and attribute annotations, based on the multi-task attribute scene recognition network MASR, It is characterized in that The specific steps include: 1) Given a scene image x i , using CNN network to extract its features θ I is the CNN network parameter; 2) Use attribute annotation strategy to calculate object attribute scores, and use the object attribute scores to identify v i The attribute objects in are streamlined according to the object filtering logic; 3) The simplified feature v i Input to the fully connected layer L |K| Perform classification prediction, where K is the number of scene classification categories; at the same time, the simplified feature v i Input fully connected layer L |A| Predicted attribute probability p att , where A is the detected attribute set; 4) The predicted attribute probability p att Compared with the attribute representation learned from external data alone, the input attribute layer is i Redistribute the weights and use the attribute task loss function to accelerate the attribute layer tasks; 5) The corrected v i Feedback to the fully connected layer L |K| .

2. A multi-task attribute scene recognition method based on category labels and attribute annotations as claimed in claim 1, It is characterized in that The attribute labeling strategy is to transform two probability distributions p s With p t Simply merge and use the object detection score P as the confidence score, i.e. the object attribute score, as follows: Collect object attributes and context information from COCO Object and COCO Panoptic datasets, and process the stuff and thing types independently. Let S and T be the sets of stuff and thing respectively, and F s With F t is the pre-trained CNN model for each task, let {x 1 ,x 2 ,...,x n }∈X represents a scene-centric dataset with only category labels, using F on X s With F t Predict the distribution on S and T, p s =F s (X) and p t =F t (X), where p s ∈R |S| With p t ∈R |T| They are the probability distribution predictions of S and T respectively. Given a data set X, the final stuff+thing prediction P∈R |S|+|T| , defined as P = p on a given scene dataset s ∪p t , where P does not increase to 1 and does not represent a probability distribution. For two probability distributions p s With p t Take the average to combine them, Among them, S and T do not always have intersections, representing different data sources.

3. A multi-task attribute scene recognition method based on category labels and attribute annotations as claimed in claim 2, It is characterized in that The object screening is to further screen the objects in S and T according to the object detection score and the object frequency, specifically including: Based on object detection score: Object instances with object detection scores less than the threshold are discarded. Only objects with object detection scores higher than the threshold are selected as scene attributes. In this process, P is redefined as: Where ξ is the threshold, when the detection score is 0, the object is considered not to exist in the scene; Based on object frequency: We further consider the attribute frequency of a given scene category and remove uncommon objects. For each category c, we define the relative attribute frequency as the number of non-zero fractions covering the category images if {a 1 ,a 2 ,...,a m }∈A c is the detection attribute set of c, the optimal Defined as: where f c (a j ) is the value of a given category c j The relative frequency of the attribute, β is the minimum frequency, is the final attribute list of c.

4. A multi-task attribute scene recognition method based on category labels and attribute annotations as claimed in claim 3, It is characterized in that The attribute task loss function is specifically: Define the multi-class cross entropy loss function: Among them, p att (x i ,j) is the training sample x i The predicted category probability on the j-th attribute of is a property annotation, which is defined as: Let’s introduce the regularization term β j , which reflects the relative frequency of the jth attribute in the training data, that is, the ratio of its positive and negative attribute labels. Formula (3) can be transformed into: where ||a j || is the number of samples holding the k-th category label of the j-th attribute, that is, the size of the j-th attribute of the k-th scene category, where classifiers of different attribute features are not shared.

5. A multi-task attribute scene recognition method based on category labels and attribute annotations as claimed in claim 4, It is characterized in that The attribute layer is specifically: A layer is introduced to reweight the attributes according to the detection score, which consists of a series of linear transformations that aggregate all attribute information into a vector v i In, use Represents the attribute classifier f A The attribute scores of The confidence score c i for: Where σ is the sigmoid activation function, W * ∈R m×m With b i ∈R m×1 is a trainable parameter, v i By c i with a i Element-wise multiplication.

Citation Information

Patent Citations

  • Zero sample learning method based on global semantic consistency network

    CN108846413A

  • Fine-grained image classification method based on image attribute active learning

    CN112528058A