Visual feature generation method and system based on category attribute vector and semantic matrix

By combining class attribute vectors and semantic matrices, using conditional generative adversarial networks to generate visual features, the limitations of the zero-sample learning method in the prior art knowledge transfer between known categories and unknown categories and the insufficient quality of generated samples is achieved, and more efficient distinction between invisible class recognition and generated samples is achieved.

CN120125892APending Publication Date: 2025-06-10SHANDONG GUOSHU DEV CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510196289.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-11-12
Filing Date
2025-02-21
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing zero-sample learning method has limitations in the knowledge transfer between known and unknown categories, resulting in the generated virtual visual samples of invisible classes being confused with similar visible classes, and failing to make full use of the information in the feature semantic matrix, affecting the quality of the generated samples.

Method used

By obtaining the visual feature description statement of the category to be generated, a semantic matrix is ​​generated using the word vector model, and inputting it into the trained generative model, combining the semantic matrix and the category attribute vector to generate more refined and distinguishable visual features. The generative model adopts a conditional generative adversarial network, generates local image features through semantic matrix and Gaussian noise, and constructs a convex hull to weight the synthesis of complete visual features.

Benefits of technology

It effectively improves the recognition ability of invisible classes, reduces the confusion between generative models between invisible classes and similar visible classes, improves the distinction and robustness of generated samples, and makes the model perform excellently on multiple benchmark data sets.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125892A_ABST
    Figure CN120125892A_ABST
Patent Text Reader

Abstract

The invention proposes a visual feature generation method and system based on a category attribute vector and a semantic matrix, and relates to the field of deep learning, and the method comprises the steps: obtaining a visual feature description statement of a to-be-generated category, and generating a semantic matrix through a word vector model; inputting the semantic matrix into the trained generative model to obtain visual features of the category to be generated; wherein the generation model generates local image features through a semantic matrix and Gaussian noise, constructs a convex hull, and performs weighted synthesis on the local image features by taking the synthesized visual features in a vertex space of the convex hull as a target and taking category attribute vectors as weights to obtain complete visual features; the visual features which are finer and rich in distinction degree are generated, and the problems that the category semantic attribute vectors are difficult to migrate between known categories (visible categories) and unknown categories (invisible categories) and the distinction degree between different categories of the synthesized virtual visual samples is insufficient are solved, so that the recognition capability of the invisible categories is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of deep learning, and particularly relates to a method and system for generating visual features based on a category attribute vector and a semantic matrix. Background Art

[0002] At present, deep learning technology has made remarkable progress in many fields, but most technologies rely on large-scale labeled datasets for model training. However, zero-shot learning is a method that uses known category (visible class) data for training to achieve the recognition of unknown category (invisible class) data, and this method has broad application prospects in scenarios where data is scarce or it is difficult to obtain labeled data.

[0003] Existing zero-shot learning methods are mainly divided into embedding-based methods and generation-based methods. Embedding-based methods embed the category attribute vector and the image visual feature into the same space, and project the unknown data into the class semantic space for classification during the test phase. Since the model cannot access the image data of unknown categories during the training phase, this method may be biased towards known categories during the test phase, resulting in incorrect synthesis of the image visual features of unknown categories. The generation-based method generates unknown class samples through a generative model, but as Figure 1 shown, this method also has the problem that the generative model is biased towards known categories during the test phase.

[0004] Therefore, there are limitations in the knowledge transfer between existing generative zero-shot learning models for known categories (visible classes) and unknown categories (invisible classes), resulting in confusion between the generated virtual visual samples of invisible classes and similar visible classes during recognition. In addition, existing technologies only use category attribute vectors to generate visual features and do not fully utilize the information in the feature semantic matrix, which will affect the quality of the generated samples. Summary of the Invention

[0005] To overcome the above deficiencies of the prior art, the present invention provides a method and system for generating visual features based on a category attribute vector and a semantic matrix, which can generate more refined and discriminative visual features, solve the problems of difficult transfer of category semantic attribute vectors between known categories (visible classes) and unknown categories (invisible classes) and insufficient discrimination between different categories of synthesized virtual visual samples, thereby effectively improving the recognition ability of invisible classes.

[0006] To achieve the above object, one or more embodiments of the present invention provide the following technical solutions:

[0007] The first aspect of the present invention provides a method for generating visual features based on a category attribute vector and a semantic matrix.

[0008] A method for generating visual features based on a category attribute vector and a semantic matrix, comprising:

[0009] Obtain a visual feature description statement of the category to be generated, and generate a semantic matrix through a word vector model;

[0010] Input the semantic matrix into the trained generation model to obtain the visual features of the category to be generated;

[0011] Wherein, the generation model generates local image features through the semantic matrix and Gaussian noise, constructs a convex hull, and targets the synthesized visual features in the vertex space of the convex hull, and weights and synthesizes the local image features with the category attribute vector as the weight to obtain complete visual features.

[0012] Furthermore, the semantic matrix extracts semantic description information representing visual features from the visual feature description statement through a pre-trained word vector model, encodes it into a high-dimensional feature vector, and forms a semantic matrix representing various features.

[0013] Furthermore, the generation model adopts a conditional generative adversarial network to guide the feature generation process with the semantic matrix as a conditional constraint;

[0014] The conditional generative adversarial network includes a generator and a discriminator. The generator takes the semantic matrix and Gaussian noise as inputs and outputs complete visual features; the discriminator takes the visual features and the category as inputs and outputs the probability that the visual features belong to the category.

[0015] Furthermore, the generator includes two units: a local generation unit and a weighted synthesis unit;

[0016] The local generation unit takes the semantic matrix and Gaussian noise as inputs and generates local image features through an adversarial generation network;

[0017] The weighted synthesis unit weights and synthesizes the local image features with the category attribute vector as the weight to obtain complete visual features.

[0018] Furthermore, the training of the generation model is based on the constructed convex hull and is trained through a masked contrast learning method to achieve the effect of generating realistic fake samples so that the discriminator cannot accurately distinguish between real and fake samples.

[0019] Furthermore, the construction of the convex hull uses the local image features as points, and the region connected by all points is the convex hull;

[0020] Targeting the synthesized visual features in the vertex space of the convex hull is to normalize the category attribute vector so that its sum is 1 to avoid the synthesized visual features falling outside the convex hull.

[0021] Furthermore, in the masked contrastive learning method, a masking mechanism is introduced to mask the local image features related to negative classes, ensuring that the generator only learns the details related to positive classes.

[0022] The second aspect of the present invention provides a visual feature generation system based on a class attribute vector and a semantic matrix.

[0023] The visual feature generation system based on a class attribute vector and a semantic matrix includes a semantic generation module and a feature generation module:

[0024] The semantic generation module is configured to: obtain a visual feature description statement of the class to be generated, and generate a semantic matrix through a word vector model;

[0025] The feature generation module is configured to: input the semantic matrix into the trained generation model to obtain the visual features of the class to be generated;

[0026] Wherein, the generation model generates local image features through the semantic matrix and Gaussian noise, constructs a convex hull, and targets the synthesized visual features in the vertex space of the convex hull, and performs weighted synthesis of the local image features with the class attribute vector as the weight to obtain complete visual features.

[0027] The third aspect of the present invention provides a computer-readable storage medium, on which a program is stored, and when the program is executed by a processor, the steps in the visual feature generation method based on a class attribute vector and a semantic matrix as described in the first aspect of the present invention are implemented.

[0028] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored on the memory and executable on the processor, and when the processor executes the program, the steps in the visual feature generation method based on a class attribute vector and a semantic matrix as described in the first aspect of the present invention are implemented.

[0029] The above one or more technical solutions have the following beneficial effects:

[0030] (1) Fine feature generation: By combining the semantic matrix and the class attribute vector, the present invention can generate more fine-grained and discriminative visual features, thereby effectively improving the recognition ability for unseen classes.

[0031] (2) Cross-class transfer: By adopting the convex hull modeling and contrastive learning strategies, unbiased transfer from visible classes to unseen classes can be achieved, reducing the confusion of the generation model between unseen classes and similar visible classes.

[0032] (3) Strong robustness: By combining instance contrastive learning and class-level contrastive learning, the present invention can effectively reduce model bias, improve the discriminability of generated samples, and enable the model to perform excellently on multiple benchmark datasets.

[0033] Advantages of additional aspects of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] The accompanying drawings forming a part of this specification are used to provide a further understanding of the present invention. The schematic embodiments and descriptions thereof of the present invention are used to explain the present invention and do not unduly limit the present invention.

[0035] Figure 1 It is an example diagram of erroneously synthesizing visual features of an unknown class.

[0036] Figure 2 It is a flowchart of the method for the first embodiment.

[0037] Figure 3 It is a schematic diagram of the word vector model for the first embodiment.

[0038] Figure 4 It is a structural diagram of the generation model for the first embodiment.

[0039] Figure 5 It is a schematic diagram of the convex hull for the first embodiment. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0040] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in the present invention have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.

[0041] It should be noted that the terms used herein are for the purpose of describing specific embodiments only and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0042] Embodiment 1

[0043] In an embodiment of the present disclosure, a method for generating visual features based on a category attribute vector and a semantic matrix is provided. As Figure 2 shown, it includes the following steps:

[0044] Step S1: Obtain a visual feature description statement of the category to be generated, and generate a semantic matrix through a word vector model;

[0045] Furthermore, the semantic matrix extracts semantic description information representing visual features from visual feature description sentences through a pre-trained word vector model, encodes the information into a high-dimensional feature vector, and forms a semantic matrix representing various features.

[0046] Specifically, Figure 3 As shown, the visual feature description sentence can use feature semantic vocabulary to express some semantic features of the generated category in terms of wings, color, head, etc., and generate a semantic matrix M{m1,m2,…} through the word vector model.

[0047] Step S2: input the semantic matrix into the trained generation model to obtain the visual features of the category to be generated;

[0048] Among them, the generative model generates local image features through semantic matrix and Gaussian noise, and constructs a convex hull. The synthesized visual features are targeted in the vertex space of the convex hull, and the local image features are weighted synthesized with the category attribute vector as the weight to obtain a complete visual feature.

[0049] Specifically, Figure 4 As shown, the generative model adopts a conditional generative adversarial network (cGAN), which uses a semantic matrix M as a conditional constraint to guide the feature generation process; the conditional generative adversarial network includes a generator and a discriminator. The generator takes the semantic matrix and Gaussian noise as input and outputs complete visual features; the discriminator takes visual features and categories as input and outputs the probability that the visual features belong to the category.

[0050] Further, the generator includes two units: a local generation unit and a weighted synthesis unit;

[0051] The local generation unit takes the semantic matrix M and Gaussian noise ε as input and generates local image features h through a generative adversarial network (GAN) i , expressed as:

[0052] h i =G(m i ,ε i ),i={i,…,d}

[0053] Among them, d is the number of features of the semantic matrix M, which is also the dimension of the category attribute vector a.

[0054] The weighted synthesis unit uses the category attribute vector as the weight to perform weighted synthesis on the local image features to obtain a complete visual feature. The formula is:

[0055]

[0056] Among them, Norm normalizes the category attribute vector a so that its sum is 1, thus avoiding the synthesized features from falling outside the convex hull. ReLU is the activation function, and a K represents the semantic attribute vector of the K-th category.

[0057] Here, the category attribute vector a is scored by experts on the features of known categories. For example, for the category of tigers, feature one - tail: 0.9, feature two - limbs: 0.9, …, and the final category attribute vector a of the tiger category is [0.9, 0.9, …].

[0058] The training of the generative model is based on the constructed convex hull and is trained through masked contrastive learning to generate realistic fake samples, making the discriminator unable to accurately distinguish between real and fake samples.

[0059] The construction of the convex hull uses local image features as points, and the region connected by all points is the convex hull.

[0060] Specifically, in the process of generating complete visual features, a convex hull structure is used to model the generated local features. As Figure 5 shown, the generated local image feature h i is mapped to the vertex space of the convex hull. The generator generates local image features through the local image feature points of the convex hull and gradually expands to the complete visual feature process. This process aims at the synthesized visual features in the vertex space of the convex hull, normalizes the category attribute vector so that its sum is 1, avoids the synthesized visual features from falling outside the convex hull, helps the model to refine the modeling of samples on local image features, and ensures that all points within the convex hull can effectively distinguish different categories through contrastive learning.

[0061] During the training process, the real images of visible categories and their semantic matrices are used as inputs to construct a generative model. The generative model synthesizes virtual image features similar to real samples through a conditional generative adversarial network. In this process, masked contrastive learning is used to ensure that the generated samples of different categories have sufficient distinctiveness. In the testing phase, the model takes the semantic matrix of unknown categories as input, generates virtual samples through the generative model, and then uses these virtual samples for fully supervised classification; the generated virtual samples can better adapt to the distribution of invisible categories, thereby improving the model's recognition ability for unknown categories.

[0062] The masked contrastive learning is to introduce a masking mechanism to mask the local image features related to negative classes, ensuring that the generator only learns the details related to positive classes.

[0063] Specifically, during the process of generating samples, the model introduces a masking mechanism to mask the local image features related to the negative class in the negative samples Xcc of X generated by the masking mechanism, ensuring that the generator only learns the details related to the positive class.

[0064] The generation of the mask is based on the semantic similarity between the visible class and the invisible class. Through contrastive learning, the generated samples of the unknown class by the generative model are made to better conform to their true data distribution, and the synthesized local image feature h i As the vertices of the convex hull and the class semantic attribute vector a as the weight, by combining a and h i The complete visual feature X is obtained by weighted summation; the weights of some hi can be set to 0 according to probability or according to one's own intention. In this way, the generated complete visual feature will lack some key or non-key features of this class, so as to control what kind of negative samples are synthesized; by synthesizing different negative samples, the synthesized X can be made to better conform to its true data distribution.

[0065] The mask contrastive learning here includes instance contrastive learning and class-level contrastive learning, specifically as follows:

[0066] 1. Instance contrastive learning:

[0067] Negative examples:

[0068] By covering part of hi to synthesize its specific instance contrastive learning negative sample XIC, thus forcing X - to be distributed closer to the true data distribution in the convex hull. For the synthesized sample X - k of the k-th class, its corresponding class semantic vector {ak(i)}di = 1 can be found. The specific value of ak(i) represents the score of the k-th class on the i-th feature attribute, and Norm(ak(i)) is the weight for the weighted fusion of hi, that is, the importance of the i-th attribute to the k-th class. Therefore, by covering the corresponding hi when fusing X - the corresponding specific negative samples are generated; for example, if the value of hi with a larger weight is set to 0, then negative samples with a large gap from X - can be generated; while when the value of hi with a smaller weight is set to 0, negative samples that are relatively similar to X - but still have differences can be obtained; so now the problem becomes which hi need to be covered to generate negative samples.

[0069] Manually setting the dimensions and quantities to be covered requires a large number of hyperparameter experiments and may not necessarily achieve the best results; in order to generate rich negative samples and also to make them have specific differences from the positive samples, Bernoulli sampling is used to cover hi. Therefore, the process of covering (masking) hi when generating the instance contrastive learning negative sample XIC of X - can be expressed as:

[0070]

[0071] Among them, denotes Bernoulli sampling, that is, for each hi, it has a probability of ak(i) being masked. By this method, richer negative samples can be generated. The masked hi value will be set to 0 to generate the negative sample XIC. This process can be expressed as:

[0072]

[0073] By using these positive and negative samples, we define the objective function of instance contrast (IC) as follows:

[0074]

[0075] Among them, pd is the data distribution of the real sample, and τ represents the temperature parameter; for an X generated by the generator - , the positive sample X is set as the positive sample for all samples in a training batch with the same category as it, and the negative sample is X IC .

[0076] 2. Class-level contrastive learning

[0077] In class-level contrastive learning, for X - , the positive sample is still set as all samples in a training batch with the same category as it, and the negative sample X CC is composed of samples in the training batch with different categories from it. Therefore, class-level contrastive learning can be formulated as:

[0078]

[0079] This can make X - more conform to the real sample data distribution and at the same time increase its distinguishability from real samples of different categories, thus being more conducive to the training of the classifier.

[0080] In this embodiment, M is used to generate the discretized intermediate layer visual feature h i , and then h i is regarded as the point of the convex hull, and the final class feature X - is generated by injecting class information; finally, the three loss functions in this article are used to continuously optimize this process. Among them, L WGAN can better train the generator (G); L IC generates diverse negative samples through Bernoulli sampling mask operation, which can better learn the real data distribution of the class in the convex hull and the differences between similar classes; L CCThis can better make the synthetic sample features conform to their true data distribution. Therefore, the overall goal is to learn a generator (G) that minimizes the loss function as much as possible, that is:

[0081]

[0082] Among them, the formula of LG is as follows:

[0083]

[0084] Among them, X^ s = αX s +(1 - α)X - s And α ~ U(0, 1).

[0085] It should be noted that in Figure 4 , in order to make the flow chart simple and easy to understand, the discriminator in WGAN is omitted, and L WGAN is directly used to replace this part.

[0086] Embodiment 2

[0087] In an embodiment of the present disclosure, a visual feature generation system based on a category attribute vector and a semantic matrix is provided, including a semantic generation module and a feature generation module:

[0088] The semantic generation module is configured to: obtain a visual feature description statement of the category to be generated, and generate a semantic matrix through a word vector model;

[0089] The feature generation module is configured to: input the semantic matrix into the trained generation model to obtain the visual features of the category to be generated;

[0090] Among them, the generation model generates local image features through the semantic matrix and Gaussian noise,

[0091] and constructs a convex hull, aiming at the synthesized visual features in the vertex space of the convex hull, and performs weighted synthesis of the local image features with the category attribute vector as the weight to obtain complete visual features.

[0092] Embodiment 3

[0093] The purpose of this embodiment is to provide a computer-readable storage medium.

[0094] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the steps in the visual feature generation method based on a category attribute vector and a semantic matrix as described in Embodiment 1 of the present disclosure.

[0095] Embodiment 4

[0096] The purpose of this embodiment is to provide an electronic device.

[0097] An electronic device, comprising a memory, a processor, and a program stored on the memory and executable on the processor, wherein when the processor executes the program, the steps in the visual feature generation method based on the category attribute vector and the semantic matrix as described in Embodiment 1 of the present disclosure are implemented.

[0098] The foregoing is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and changes. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A visual feature generation method based on category attribute vector and semantic matrix, characterized in that: include: Obtain the visual feature description sentence of the category to be generated, and generate a semantic matrix through the word vector model; Input the semantic matrix into the trained generative model to obtain the visual features of the category to be generated; Among them, the generative model generates local image features through semantic matrix and Gaussian noise, and constructs a convex hull. The synthesized visual features are targeted in the vertex space of the convex hull, and the local image features are weighted synthesized with the category attribute vector as the weight to obtain a complete visual feature.

2. The visual feature generation method based on category attribute vector and semantic matrix according to claim 1, characterized in that: The semantic matrix extracts the semantic description information representing the visual features from the visual feature description sentence through a pre-trained word vector model, encodes it into a high-dimensional feature vector, and forms a semantic matrix representing various features.

3. The visual feature generation method based on category attribute vector and semantic matrix according to claim 1, characterized in that: The generative model adopts a conditional generative adversarial network and uses a semantic matrix as a conditional constraint to guide the feature generation process; The conditional generative adversarial network includes a generator and a discriminator, wherein the generator takes a semantic matrix and Gaussian noise as input and outputs a complete visual feature; The discriminator takes visual features and categories as input and outputs the probability that the visual feature belongs to the category.

4. The visual feature generation method based on category attribute vector and semantic matrix according to claim 3, characterized in that: The generator includes two units: a local generation unit and a weighted synthesis unit; The local generation unit takes the semantic matrix and Gaussian noise as input and generates local image features through a generative adversarial network; The weighted synthesis unit performs weighted synthesis on local image features using the category attribute vector as a weight to obtain a complete visual feature.

5. The visual feature generation method based on category attribute vector and semantic matrix according to claim 1, characterized in that: The training of the generative model is based on the constructed convex hull and is trained by mask contrast learning to achieve the effect of generating realistic fake samples so that the discriminator cannot accurately distinguish between real and fake samples.

6. The visual feature generation method based on category attribute vector and semantic matrix according to claim 5, characterized in that: The convex hull is constructed by taking local image features as points, and the area connected by all points is the convex hull; The method of taking the synthesized visual features in the vertex space of the convex hull as the target is to normalize the category attribute vector so that the sum thereof is 1, thereby preventing the synthesized visual features from falling outside the convex hull.

7. The method for generating visual features based on category attribute vectors and semantic matrices according to claim 5, characterized in that: The mask contrast learning method introduces a mask mechanism to mask local image features related to the negative class, ensuring that the generator only learns details related to the positive class.

8. A visual feature generation system based on category attribute vector and semantic matrix, characterized in that: Including semantic generation module and feature generation module: The semantic generation module is configured to: obtain a visual feature description sentence of a category to be generated, and generate a semantic matrix through a word vector model; The feature generation module is configured to: input the semantic matrix into the trained generation model to obtain the visual features of the category to be generated; Among them, the generative model generates local image features through semantic matrix and Gaussian noise, and constructs a convex hull. The synthesized visual features are targeted in the vertex space of the convex hull, and the local image features are weighted synthesized with the category attribute vector as the weight to obtain a complete visual feature.

9. An electronic device, comprising: a memory for non-transitory storage of computer readable instructions; a processor for executing the computer-readable instructions; Wherein, when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 7 is executed.

10. A storage medium, characterized in that: The computer-readable instructions are non-transitorily stored, wherein when the computer-readable instructions are executed by a computer, the method of any one of claims 1 to 7 is performed.