Fine-grained image generation method and device based on generalized zero sample learning, terminal and medium
Through the coordinated work of dynamic frequency domain and airspace feature extraction and fusion modules and generative adversarial networks, the optimization generator generates invisible category visual features, solving the problem that generation features in the existing technology is difficult to capture deep semantic information and insufficient generalization capabilities, and achieving high-quality visual feature generation and classification performance improvement.
Patent Information
- Application Number
- CN202510107029.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-23
- Publication Date
- 2025-06-10
AI Technical Summary
When generating visual features of invisible categories, it is difficult for the prior art to accurately capture the deep semantic information of input semantic descriptions, making it difficult for the generated features to truly reflect subtle differences between categories, and when the complex fine-grained data distribution is distributed, the generalization ability is insufficient.
Through the coordinated work of dynamic frequency domain and airspace feature extraction and fusion modules and generative adversarial networks, the generator generation quality of visual features for invisible categories is optimized. The specific methods include obtaining the real visual feature map of visible categories and its semantic features, conducting adversarial training on semantic features and visual features through the generation of adversarial networks, and optimizing feature representation using dynamic frequency domain and airspace feature extraction and fusion modules.
It significantly improves the generator's generation quality of visual features for invisible categories, has good semantic consistency and category distinction ability of generated features, and improves the classification performance of invisible categories in generalized zero-sample learning task.
Smart Images

Figure CN120125684A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a fine-grained image generation method, device, terminal and medium based on generalized zero-shot learning, belonging to the field of computer vision. Background Art
[0002] Fine-grained image classification is an important task in the field of computer vision, and its goal is to distinguish objects that belong to the same large category but have slight visual differences. For example, birds, dogs or plants of different species often show subtle differences in morphology, color, etc. Due to the high similarity between categories and the high variability within categories, fine-grained classification poses extremely high requirements on the feature extraction ability and discrimination ability of the model.
[0003] Currently, fine-grained image classification methods usually rely on large-scale labeled data to learn fine image features through deep neural networks. However, in practical applications, it is often difficult to obtain a large amount of labeled data. Especially when facing image data of new categories or minority categories, this high dependence on labeling becomes particularly prominent. In addition, when class labels are not available, traditional image classification methods usually cannot effectively model the feature distribution of new categories, which makes fine-grained image classification in zero-shot learning (ZSL) and generalized zero-shot learning (GZSL) tasks a challenging topic.
[0004] In the framework of zero-shot learning and generalized zero-shot learning, fine-grained classification methods attempt to use semantic descriptions as a bridge to associate visible categories and invisible categories. However, the existing technologies still have the following problems in generating features of invisible categories:
[0005] When existing generative models generate visual features of invisible categories, they often cannot accurately capture the deep semantic information of the input semantic description, resulting in the generated features being difficult to truly reflect the subtle differences between categories. In the face of complex fine-grained data distributions, the generalization ability of generative models is insufficient, and the generated features are prone to deviate from the true distribution, resulting in low classification performance of invisible categories in generalized zero-shot learning tasks. In summary, there are still certain deficiencies in generative models, and further optimization is urgently needed to solve the problems of complex semantics and scarce samples in fine-grained classification. Summary of the Invention
[0006] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a fine-grained image generation method, device, terminal and medium based on generalized zero-shot learning. Through the collaborative work of the dynamic frequency-domain and spatial-domain feature extraction and fusion module and the generative adversarial network, the generation quality of the generator for visual feature maps of invisible categories is effectively improved.
[0007] To achieve the above object, the present invention is implemented by the following technical solutions:
[0008] In a first aspect, the present invention provides a fine-grained image generation method based on generalized zero-shot learning, including:
[0009] Input the semantic features of the invisible class into the trained generator to generate the corresponding optimized visual feature map;
[0010] The training method of the generator includes:
[0011] Obtain an image data set, which contains the real visual feature maps of the visible classes and their corresponding semantic features;
[0012] Input the semantic features of the visible classes into the generator to generate synthetic visual feature maps;
[0013] Input the real visual feature map and the synthetic visual feature map into the dynamic frequency domain and spatial domain feature extraction and fusion module respectively to obtain the optimized real visual feature map and the optimized synthetic visual feature map;
[0014] Based on the generative adversarial network, perform adversarial training on the semantic features and visual features of the images, where:
[0015] The discriminator uses the optimized real visual feature map, the optimized synthetic visual feature map, and the unoptimized visual feature map, and based on a preset first loss function, learns the distribution difference and class discrimination ability between the features;
[0016] The generator optimizes the class discrimination ability and semantic consistency of the generated visual feature map according to the feedback result provided by the discriminator and in combination with a preset second loss function.
[0017] Further, inputting the real visual feature map and the synthetic visual feature map into the dynamic frequency domain and spatial domain feature extraction and fusion module respectively to obtain the optimized real visual feature map and the optimized synthetic visual feature map includes:
[0018] The dynamic frequency domain and spatial domain feature extraction and fusion module respectively performs spatial domain feature extraction and frequency domain feature extraction on the input visual feature map to obtain the spatial domain feature representation S and the frequency domain feature representation F; where, , ; represents the number of input visual feature maps, represents the feature dimension;
[0019] Concatenate the spatial domain feature representation S and the frequency domain feature representation F to obtain the concatenated feature representation ; Then, a fully connected network is used to map the concatenated feature representation C to generate a new feature representation :
[0020] ;
[0021] In the formula: represents an activation function; represents a weight matrix; represents a weight term;
[0022] The gating function is used to analyze and learn the feature representation to dynamically adjust the contribution ratio of the spatial domain feature representation S and the frequency domain feature representation F, and an optimized visual feature map is obtained based on the output of the gating function as shown in the following formula:
[0023] ;
[0024] In the formula: is the output of the gating function, representing the weight corresponding to the spatial domain feature representation S.
[0025] Furthermore, the first loss function is:
[0026] ;
[0027] In the formula: represents the Wasserstein distance loss; represents the gradient penalty; represents the RINCE loss; represents the classification loss; , , and respectively represent the weight coefficients corresponding to each loss and the gradient penalty;
[0028] Among them, the Wasserstein distance loss , the RINCE loss and the classification loss are obtained through the following steps:
[0029] The semantic features of the true visual feature map of the visible classes and the random noise are input into the generator to generate a synthetic visual feature map ;
[0030] The visible visual feature map is combined with the synthetic visual feature map Input the dynamic frequency domain and spatial domain feature extraction and fusion modules respectively to obtain the optimized visible visual feature map and the synthetic visual feature map ;
[0031] Input the optimized visible visual feature map into the discriminator to calculate the RINCE loss as shown in the following formula:
[0032] ;
[0033] In the formula: , s represents the similarity score of the optimized visible visual feature map; is the similarity score of the positive sample pair, indicating the similarity of two optimized visible visual feature maps belonging to the same category; is the similarity score of the negative sample pair, indicating the similarity of two optimized visible visual feature maps belonging to different categories; K represents the number of negative sample pairs; q represents the parameter controlling the behavior of the loss function; represents the weight balance parameter of the positive and negative sample pairs; represents the base of the exponential function, which is an irrational number;
[0034] Input the optimized visible visual feature map into the discriminator to calculate the classification loss as shown in the following formula:
[0035] ;
[0036] In the formula: M represents the total number of input visible visual feature maps; C represents the total number of categories; represents the true label of the input visible visual feature map on category c; represents the discriminator predicting the probability that the visible visual feature map belongs to category c;
[0037] Input the said visible visual feature map and the synthetic visual feature map into the discriminator respectively to obtain the outputs and of the discriminator, where
[0038] L 1 = E x + ~ P r [ D ( x + , a ) ] ;
[0039] L 2 = E x − ~ P g [ D ( x − , a ) ] ;
[0040] In the formula: represents the distribution of the visible visual feature map ; represents the synthetic visual feature map generated by the generator ; In the formula: represents the score given by the discriminator to the visible visual feature map ; represents the score given by the discriminator to the visible visual feature map ;
[0041] According to the output of the discriminator, the Wasserstein distance loss is expressed as follows:
[0042] .
[0043] Furthermore, the second loss function is:
[0044] ;
[0045] In the formula: , and respectively represent and the weight coefficients corresponding to each loss;
[0046] Among them, the RINCE loss and the classification loss are obtained through the following steps:
[0047] Combining the optimized visible visual feature map with the synthetic visual feature map , calculate the RINCE loss as follows:
[0048] ;
[0049] In the formula: , represents the similarity score between the optimized visible visual feature map and the synthetic visual feature map ; is the similarity score of the positive sample pair, representing the similarity that the visible visual feature map and the synthetic visual feature map belong to the same category; is the similarity score of the negative sample pair, representing the visible visual feature map and the synthetic visual feature map Similarities belonging to different categories;
[0050] The synthetic visual feature map will be Input into the discriminator to calculate the classification loss as shown in the following formula:
[0051] ;
[0052] In the formula: represents the total number of input synthetic visual feature maps; represents the true label of the input synthetic visual feature map on category c; represents the discriminator predicts the probability that the synthetic visual feature map belongs to category c.
[0053] Furthermore, input the semantic features of invisible categories into the trained generator to generate corresponding optimized visual feature maps. After that, it further includes:
[0054] Combine the optimized visual feature map and the true visual feature map to perform classification training on the generalized zero-shot learning classifier to obtain an optimized classifier.
[0055] In a second aspect, the present invention provides a fine-grained image generation device based on generalized zero-shot learning, including:
[0056] A dynamic frequency-domain and spatial-domain feature extraction and fusion module, which is used to receive the true visual feature map and the synthetic visual feature map, extract the spatial-domain features and frequency-domain features of the image, and generate an optimized true visual feature map and an optimized synthetic visual feature map;
[0057] A discriminator, which is used to calculate the distribution difference and category discrimination ability between features based on the optimized true visual feature map, the optimized synthetic visual feature map, and the unoptimized visual feature map, and update the discriminator parameters through a preset first loss function;
[0058] A generator, which is used to generate a synthetic visual feature map according to the input semantic features, and optimize the category discrimination ability and semantic consistency of the generated visual feature map based on the feedback result of the discriminator through a preset second loss function;
[0059] A generative adversarial network, including the discriminator and the generator, constructs an adversarial training mechanism, where:
[0060] The generator generates an optimized visual feature map according to the input semantic features of invisible categories;
[0061] The discriminator optimizes the discrimination ability between the true visual feature and the synthetic visual feature through the distribution comparison between features.
[0062] Further, the fine-grained image generation device based on generalized zero-shot learning further includes:
[0063] A model training unit, configured to perform classification training on a generalized zero-shot learning classifier by combining the optimized visual feature map and the real visual feature map to obtain an optimized classifier.
[0064] In a third aspect, the present invention provides an electronic terminal, including a processor and a memory connected to the processor, and a computer program is stored in the memory. When the computer program is executed by the processor, the steps of the above-mentioned fine-grained image generation method based on generalized zero-shot learning are executed.
[0065] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the above-mentioned fine-grained image generation method based on generalized zero-shot learning are implemented.
[0066] Compared with the prior art, the beneficial effects achieved by the present invention are:
[0067] 1. The present invention provides a fine-grained image generation method based on generalized zero-shot learning. Through the collaborative work of the dynamic frequency domain and spatial domain feature extraction and fusion module and the generative adversarial network, the generation quality of the generator for visual features of invisible categories is effectively improved.
[0068] 2. By the generator generating an optimized visual feature map according to the semantic features of invisible categories, the problem of lack of training samples for invisible categories is solved. The generated features have good semantic consistency and category discrimination ability, providing reliable data support for the training of the generalized zero-shot learning classifier. The dynamic frequency domain and spatial domain feature extraction and fusion module realizes the deep fusion of spatial domain and frequency domain features, further improving the expression ability and adaptability of the generated features. Based on adversarial training, the discriminator optimizes the learning of the distribution difference between real and generated features, and the generator continuously improves the authenticity and semantic consistency of the generated features by using the feedback, realizing the efficient cooperation and performance complementarity of the two. BRIEF DESCRIPTION OF THE DRAWINGS
[0069] Figure 1 is a schematic flowchart of a training method of a generator provided in Embodiment 1 of the present invention;
[0070] Figure 2 is a schematic flowchart of obtaining an optimized visual feature map by using the dynamic frequency domain and spatial domain feature extraction and fusion module provided in Embodiment 1 of the present invention. DETAILED DESCRIPTION
[0071] The technical solution of the present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the embodiments of the present application and the specific features in the embodiments are detailed descriptions of the technical solution of the present application, rather than limitations on the technical solution of the present application. Without conflict, the technical features in the embodiments of the present application and the embodiments can be combined with each other.
[0072] The terms "first", "second", etc. are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first", "second", etc. may explicitly or implicitly include one or more of such features. In the description of the present disclosure / the present application, unless otherwise stated, the meaning of "a plurality" is two or more.
[0073] Embodiment 1:
[0074] Figure 1 It is a schematic flowchart of the training method of the generator in Embodiment 1 of the present invention. This flowchart only shows the logical order of the method described in this embodiment. On the premise of no conflict, in other possible embodiments of the present invention, the steps shown or described may be completed in a different Figure 1 order than that shown.
[0075] The training method of the generator specifically includes the following steps:
[0076] Obtain an image data set, which contains real visual feature maps of visible categories and their corresponding semantic features;
[0077] Input the semantic features of the visible categories into the generator to generate synthetic visual feature maps;
[0078] Input the real visual feature map and the synthetic visual feature map into the dynamic frequency domain and spatial domain feature extraction modules respectively to obtain an optimized real visual feature map and an optimized synthetic visual feature map;
[0079] Based on the generative adversarial network, perform adversarial training on the semantic features and visual features of the images, where:
[0080] The discriminator uses the optimized real visual feature map, the optimized synthetic visual feature map, and the unoptimized visual feature map, and based on a preset first loss function, learns the distribution difference and category discrimination ability between the features;
[0081] The generator optimizes the category discrimination ability and semantic consistency of the generated visual feature map according to the feedback result provided by the discriminator and in combination with a preset second loss function.
[0082] The training datasets used in this embodiment include AwA2 (Animals with Attributes 2), CUB (Caltech-UCSD Birds 200), and SUN (Scene Understanding). These datasets are widely used in zero-shot learning and generalized zero-shot learning tasks. Each dataset contains image data of visible and invisible classes, as well as corresponding semantic attribute descriptors. The data of visible classes is used for training, and the data of invisible classes is used for testing.
[0083] Taking the CUB dataset as an example, the steps for data processing and training the generator are as follows:
[0084] First, for the visible-class image data in the CUB dataset, image processing and preprocessing are performed. Specifically, data augmentation operations such as cropping, scaling, and rotating the images are carried out to increase the diversity of training samples, and the size of the images is uniformly adjusted to 224 × 224 pixels, and color normalization is performed to ensure data consistency.
[0085] In terms of semantic description generation, since the CUB dataset only contains image class labels, this embodiment uses models such as CNN-RNN sentence embeddings or CLIP to generate corresponding semantic descriptions for each image. These semantic descriptions, as additional information of the images, can help the generator understand the content of the images. To ensure the accuracy of the semantic descriptions, we removed incorrect, repetitive, or irrelevant description information. The semantic description of each image class is represented by a 1024-dimensional class-level embedded feature vector, which is the semantic feature corresponding to the true visual feature map of the visible class. This feature vector is embedded using models such as CLIP to capture the deep semantic information of the image class.
[0086] Next, the ResNet-101 model is used to extract visual features from the images in the CUB dataset. Based on the ResNet-101 pre-trained on ImageNet-1K, we obtained a 2048-dimensional visual feature vector for each image, which is the true visual feature map of the visible class. The true visual feature map of the visible class and its corresponding semantic features are used as inputs to the dynamic frequency-domain and spatial-domain feature extraction module, the generator, and the discriminator for subsequent training and optimization.
[0087] In this embodiment, the dynamic frequency-domain and spatial-domain feature extraction and fusion module is used to further extract the visual features of the images. Its workflow includes two main branches: spatial-domain feature extraction and frequency-domain feature extraction, as Figure 2 shown.
[0088] In the spatial domain feature extraction part, first, the input feature vector is processed by a multi-layer perceptron, and through a series of linear transformations and activation functions, it is mapped to a new feature space to obtain spatial structure information. During this process, to alleviate the problem of vanishing gradients or exploding gradients that may occur in deep networks and to promote information flow, the model introduces a residual connection mechanism. The residual connection adds the input data to the features processed by the multi-layer perceptron, enabling information to bypass some network layers and be directly transmitted, thereby enhancing the learning ability of the network. The features combined with the residual connection are further normalized. Finally, the self-attention mechanism is used to optimize the extracted spatial features so that it can focus on the key regions of the input data. The spatial domain feature extraction part combines the advantages of the multi-layer perceptron, residual connection, feature normalization, and self-attention mechanism, and can efficiently extract spatial information while stably training, and enhance the learning ability for complex patterns.
[0089] In the frequency domain feature extraction part, by performing a frequency domain transformation on the input signal, it is decomposed into high-frequency and low-frequency components, and local detail and global structure features are extracted respectively. In this embodiment, these frequency features are further preliminarily processed through a linear transformation, and the residual connection mechanism is combined to ensure the effective flow of information of high-frequency and low-frequency features and avoid feature loss. In addition, a cross-band interaction mechanism is introduced in the frequency domain feature extraction. By allowing high-frequency and low-frequency features to exchange information, the expressive ability of the features is enhanced. The frequency domain feature extraction part can achieve a balance between detail extraction and global structure capture, providing richer frequency domain information for subsequent processing.
[0090] After the spatial domain and frequency domain feature extraction are completed, the spatial domain feature representation S and the frequency domain feature representation F are obtained; among them, , ; represents the number of input visual feature maps, represents the feature dimension; this module performs feature splicing and fusion.
[0091] The spatial domain feature representation S and the frequency domain feature representation F are spliced column by column to obtain the spliced feature representation ; This splicing operation enables the model to simultaneously capture spatial and frequency domain information in a unified feature space, laying a foundation for the high-order learning of features.
[0092] On the spliced feature representation C, a multi-layer perceptron is further used to learn high-order feature expressions. Through a fully connected network layer, the joint features are mapped to a new feature representation space to generate a new feature representation :
[0093] ;
[0094] In the formula: denotes the activation function; denotes the weight matrix; denotes the weight term;
[0095] To dynamically adjust the contribution ratios of spatial features and frequency-domain features in different samples, this module introduces a gating mechanism. The gating mechanism generates a weighting coefficient through a learned gating function , and obtains an optimized visual feature map based on the output of the gating function as shown in the following formula:
[0096] ;
[0097] wherein, the gating function consists of a multi-layer perceptron, includes a linear transformation and a non-linear activation function, and the output is a scalar, representing the weighting ratio of the spatial features. Through the weighted summation method, the gating mechanism adaptively adjusts the fusion ratio of the spatial and frequency-domain features according to the characteristics of the input samples, so as to generate a more flexible and efficient joint feature representation.
[0098] The design of the dynamic frequency-domain and spatial-feature extraction and fusion module combines information from different domains through concatenation operations, uses the gating mechanism to dynamically adjust the feature weights, and establishes a complementary relationship between multi-dimensional features. This mechanism exhibits significant performance advantages when dealing with complex data and lays a solid foundation for generating high-quality visual features.
[0099] After the dynamic frequency-domain and spatial-feature extraction and fusion module completes the extraction and fusion of the spatial features and frequency-domain features of the input feature map, this embodiment further trains the discriminator and the generator through a generative adversarial network to optimize the visual features and ensure the semantic consistency and class discrimination ability of the generated visual feature map.
[0100] During the training process of the generative adversarial network, the discriminator and the generator promote each other through an alternating optimization method. Among them, the task of the discriminator is to learn the distribution difference and class discrimination ability between features based on the real visual feature map, the synthetic visual feature map and its optimized visual feature map; the generator gradually improves the generated visual feature map through the feedback of the discriminator to make it closer to the real visual feature map in distribution. To achieve this goal, this embodiment introduces multiple loss functions to guide the training:
[0101] The discriminator 's loss function is as shown in the following formula:
[0102] ;
[0103] In the formula: represents the Wasserstein distance loss; Represents gradient penalty; Represents the RINCE loss; Represents the classification loss; 、 、 and respectively represent the weight coefficients corresponding to each loss and the gradient penalty;
[0104] Among them, the Wasserstein distance loss is used to measure the distribution difference between the real visual feature map and the generated visual feature map, ensuring that the feature distribution of the generated visual feature map gradually approaches the real visual feature map; the gradient penalty ensures that the gradient of the discriminator changes smoothly between the real visual feature map and the generated visual feature map, avoiding unstable training caused by too strong gradients; the RINCE loss improves the category judgment ability of the discriminator by minimizing the distance between visual feature maps of the same category and maximizing the distance between visual feature maps of different categories; in this embodiment, the RINCE loss of the discriminator is calculated by comparing whether two optimized visible visual feature maps belong to the same category; the cross-entropy classification loss is used to measure the classification accuracy of the discriminator for the input optimized visible visual feature map category, and enhances the classification ability of the discriminator by maximizing the classification probability of .
[0105] The specific calculation steps of each loss and the gradient penalty are as follows:
[0106] (1)Wasserstein distance loss
[0107] Input the semantic features of the visible visual feature map and the random noise into the generator to generate a synthetic visual feature map ;
[0108] Input the visible visual feature map and the synthetic visual feature map into the discriminator respectively to obtain the outputs and of the discriminator, where
[0109] L 1 = E x + ~ P r [ D ( x + , a ) ] ;
[0110] L 2 = E x − ~ P g [ D ( x − , a ) ] ;
[0111] In the formula: represents the distribution of the visible visual feature map ; represents the synthetic visual feature map generated by the generator ; represents the score given by the discriminator to the visible visual feature map ; represents the score given by the discriminator to the visible visual feature map ;
[0112] According to the output of the discriminator, the expression of the Wasserstein distance loss is as follows:
[0113] ;
[0114] (2) Gradient penalty
[0115] ;
[0116] Among them: ; ;
[0117] In the formula: represents the intermediate sample linearly interpolated proportionally from the visible visual feature map and the synthetic visual feature map ; represents the random weight sampled from the uniform distribution from 0 to 1; represents the discriminator to the 2-norm of the gradient; represents the penalty strength for the gradient deviating from 1; represents the weight coefficient used to control the influence intensity of the gradient penalty term;
[0118] (3) RINCE loss
[0119] Input the visible visual feature map and the synthetic visual feature map into the dynamic frequency domain and spatial domain feature extraction and fusion module respectively, and obtain the optimized visible visual feature map and the synthetic visual feature map ;
[0120] Input the optimized visible visual feature map into the discriminator to calculate the RINCE loss as shown in the following formula:
[0121] ;
[0122] In the formula: , s represents the similarity score of the optimized visible visual feature map; is the similarity score of the positive sample pair, indicating the similarity of two optimized visible visual feature maps belonging to the same category; is the similarity score of the negative sample pair, indicating the similarity of two optimized visible visual feature maps belonging to different categories; K represents the number of negative sample pairs; q represents the parameter controlling the behavior of the loss function, q ∈ ( 0 , 1 ] ; when q is close to 0, RINCE is similar to the InfoNCE loss; when q is close to 1, RINCE is more robust and can better handle noise; represents the weight balance parameter of the positive and negative sample pairs.
[0123] (4) Classification loss
[0124] Input the optimized visible visual feature map into the discriminator to calculate the classification loss as shown in the following formula:
[0125] ;
[0126] In the formula: M represents the total number of input visible visual feature maps; C represents the total number of categories; represents the true label of the input visible visual feature map on category c; represents the discriminator predicting the probability that the visible visual feature map belongs to category c. If the prediction of the discriminator and the true label are exactly the same, the loss value approaches 0; if the predicted probability and the true label deviate greatly, the loss value will increase.
[0127] The training objective of the generator is to generate visual features close to the true visual feature distribution, and its loss function is as shown in the following formula:
[0128] ;
[0129] In the formula: , and represent the weight coefficients corresponding to each loss respectively;
[0130] In this embodiment, the RINCE loss of the generator is calculated by comparing the optimized visible visual feature map with the optimized synthetic visual feature map to determine whether they belong to the same category; the cross - entropy classification loss is used to measure the classification accuracy of the discriminator for the input optimized synthetic visual feature map by maximizing the classification probability of to enhance the classification ability of the discriminator .
[0131] In the training update of the generator, the specific calculation steps of the RINCE loss and the classification loss are as follows:
[0132] (1) RINCE loss
[0133] Combining the optimized visible visual feature map with the synthetic visual feature map , calculate the RINCE loss as shown in the following formula:
[0134] ;
[0135] In the formula: , represents the similarity score between the optimized visible visual feature map and the synthetic visual feature map ; is the similarity score of the positive sample pair, indicating the similarity that the visible visual feature map and the synthetic visual feature map belong to the same category; is the similarity score of the negative sample pair, indicating the similarity that the visible visual feature map and the synthetic visual feature map belong to different categories.
[0136] (2) Classification loss
[0137] Input the synthetic visual feature map into the discriminator , and calculate the classification loss as shown in the following formula:
[0138] ;
[0139] Wherein: represents the total number of input synthetic visual feature maps; represents the true label of the input synthetic visual feature map on class c; represents the discriminator predicts the probability that the synthetic visual feature map belongs to class c.
[0140] During the training process, the discriminator and the generator are alternately optimized as follows:
[0141] The discriminator updates using the visible visual feature map and the synthetic visual feature map or the optimized visible visual feature map and the synthetic visual feature map as inputs, and updates the parameters of the discriminator based on the Wasserstein distance loss, the RINCE loss, and the cross-entropy classification loss.
[0142] The generator updates using the feedback results provided by the discriminator , combines the RINCE loss and the cross-entropy classification loss, and optimizes the parameters of the generator to make the generated visual feature map gradually approach the true visual feature map.
[0143] Through the adversarial training of the discriminator and the generator, the generator can further improve the quality of the generated visual feature map based on the dynamic frequency-domain and spatial-domain features, making it both conform to the description of the input semantic features and have good class discrimination ability.
[0144] By introducing multiple loss functions in the adversarial training, including the Wasserstein distance, the RINCE loss, and the cross-entropy loss, etc., the differences in feature distributions between visible classes and invisible classes can be balanced, thus showing stronger robustness in complex data scenarios.
[0145] After the generator completes training, this embodiment uses the trained generator to generate corresponding synthetic visual feature maps according to the semantic features of invisible classes. Specifically, the input semantic features of invisible classes are mapped to the visual feature space through the generator to generate synthetic visual feature maps corresponding to the semantic features. These generated synthetic features can effectively make up for the lack of real samples of invisible classes, thereby expanding the dataset and providing support for the subsequent classifier training.
[0146] Next, combine the generated synthetic visual feature maps with the true visual feature maps of visible classes to jointly construct the training set of the generalized zero-shot learning classifier. The training process includes the following steps:
[0147] Combine the invisible category synthetic visual feature map generated by the generator with the corresponding semantic category label to form the training samples of the invisible category. Combine the real visual feature map with its corresponding semantic category label to form the training samples of the visible category. Merge the two parts of features to form a complete training set including visible and invisible categories.
[0148] Design a multi-classification model as a generalized zero-shot learning classifier. The input is the visual feature map, and the output is the category probability distribution. The optimization objective of the classifier is to improve the classification ability for both visible and invisible categories simultaneously.
[0149] Since there are more training samples for visible categories, while the training samples for invisible categories completely rely on the generator, there may be imbalances in quantity or quality. Therefore, a category balance strategy is added during the classifier training. For example:
[0150] Set balanced loss weights for different categories to ensure that the features of invisible categories are not dominated by the features of visible categories during training. Use the contrastive loss to further narrow the distance between the generated features and the real distribution, and improve the generalization ability of the classifier.
[0151] In the testing stage, input the visual feature maps of real samples, including visible and invisible categories, and predict the category labels of the samples through the classifier.
[0152] And combine evaluation metrics to measure the comprehensive performance of the classifier for visible and invisible categories. The evaluation metrics include:
[0153] Accuracy of visible categories acc_seen: The performance of the classifier on visible category samples;
[0154] Accuracy of invisible categories acc_unseen: The performance of the classifier on invisible category samples;
[0155] Harmonic mean H: Measure the comprehensive performance of the classifier for visible and invisible categories, as shown in the following formula:
[0156] ;
[0157] Through the above steps, this embodiment successfully constructs a generalized zero-shot learning classifier. This classifier combines the invisible category features generated by the generator and the real data features, and shows good classification performance on both visible and invisible categories.
[0158] The present invention uses the invisible class visual feature map generated by the generator and the real visible class feature map to jointly train a generalized zero-shot learning classifier, which not only effectively expands the diversity of training samples, but also significantly improves the classification performance of the classifier on invisible classes, while taking into account the performance of visible classes and enhancing the overall classification effect of the generalized zero-shot learning task.
[0159] Embodiment 2:
[0160] The embodiment of the present invention also provides a fine-grained image generation device based on generalized zero-shot learning, including:
[0161] A dynamic frequency domain and spatial domain feature extraction and fusion module, configured to receive a real visual feature map and a synthetic visual feature map, extract the spatial domain feature and frequency domain feature of the image, and generate an optimized real visual feature map and an optimized synthetic visual feature map;
[0162] A discriminator, configured to calculate the distribution difference and class discrimination ability between features based on the optimized real visual feature map, the optimized synthetic visual feature map, and the unoptimized visual feature map, and update the discriminator parameters through a preset first loss function;
[0163] A generator, configured to generate a synthetic visual feature map according to the input semantic feature, and optimize the class discrimination ability and semantic consistency of the generated visual feature map based on the feedback result of the discriminator through a preset second loss function;
[0164] A generative adversarial network, including the discriminator and the generator, constructs an adversarial training mechanism, wherein:
[0165] The generator generates an optimized visual feature map according to the input invisible class semantic feature;
[0166] The discriminator optimizes the discrimination ability between real visual features and synthetic visual features through the distribution comparison between features.
[0167] The device further includes:
[0168] A model training unit, configured to perform classification training on a generalized zero-shot learning classifier by combining the optimized visual feature map and the real visual feature map to obtain an optimized classifier.
[0169] Embodiment 3:
[0170] The embodiment of the present invention also provides an electronic terminal, including a processor and a memory connected to the processor. A computer program is stored in the memory. When the computer program is executed by the processor, the steps of the fine-grained image generation method based on generalized zero-shot learning described in Embodiment 1 above are executed.
[0171] Embodiment 4:
[0172] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it first implements the steps of the fine-grained image generation method based on generalized zero-shot learning described in the above-mentioned first embodiment.
[0173] The computer-readable storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.
[0174] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) containing computer-usable program codes.
[0175] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0176] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including instruction means, and the instruction means implements the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0177] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide means for implementing the functions in the processFigure 1 one process or multiple processes and / or boxes Figure 1 steps of the functions specified in one box or multiple boxes
[0178] The above are only the preferred embodiments of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the technical principle of the present invention, several improvements and modifications can be made, and these improvements and modifications should also be regarded as the protection scope of the present invention.
Claims
1. A fine-grained image generation method based on generalized zero-shot learning, characterized in that: include: Input the semantic features of the unseen categories into the trained generator to generate the corresponding optimized visual feature map; The training method of the generator includes: Obtain an image dataset containing ground-truth visual feature maps of visible categories and their corresponding semantic features; Inputting the semantic features of the visible category into a generator to generate a synthetic visual feature map; Inputting the real visual feature map and the synthetic visual feature map into the dynamic frequency domain and spatial domain feature extraction and fusion modules respectively to obtain an optimized real visual feature map and an optimized synthetic visual feature map; Based on the generative adversarial network, the semantic features and visual features of the image are trained adversarially, where: The discriminator uses the optimized real visual feature map, the optimized synthetic visual feature map and the unoptimized visual feature map to learn the distribution difference and category distinction ability between features based on the preset first loss function; The generator optimizes the category distinction ability and semantic consistency of the generated visual feature map based on the feedback provided by the discriminator and the preset second loss function.
2. The fine-grained image generation method based on generalized zero-shot learning according to claim 1, characterized in that: The real visual feature map and the synthetic visual feature map are respectively input into the dynamic frequency domain and spatial domain feature extraction and fusion modules to obtain an optimized real visual feature map and an optimized synthetic visual feature map, including: The dynamic frequency domain and spatial domain feature extraction and fusion module performs spatial domain feature extraction and frequency domain feature extraction on the input visual feature map to obtain a spatial domain feature representation S and a frequency domain feature representation F; wherein, , ; Represents the number of input visual feature maps, Represents feature dimension; The spatial domain feature representation S and the frequency domain feature representation F are concatenated to obtain the concatenated feature representation ; Then use the fully connected network to map the concatenated feature representation C to generate a new feature representation : ; Where: represents the activation function; represents the weight matrix; represents the weight term; Using gating functions to represent features Perform analysis and learning, dynamically adjust the contribution ratio of spatial domain feature representation S and frequency domain feature representation F, and obtain the optimized visual feature map based on the output of the gating function As shown below: ; Where: is the output of the gating function, which represents the weight corresponding to the spatial feature representation S.
3. The fine-grained image generation method based on generalized zero-shot learning according to claim 1, characterized in that: The first loss function for: ; Where: Represents Wasserstein distance loss; represents the gradient penalty; Indicates RINCE loss; represents the classification loss; , , and Respectively represent the weight coefficients corresponding to each loss and gradient penalty; Among them, the Wasserstein distance loss 、RINCE loss and classification loss Obtained through the following steps: The visible visual feature map Semantic features of and random noise Input Generator , generating synthetic visual feature maps ; The visible visual feature map With the synthetic visual feature map Input the dynamic frequency domain and spatial domain feature extraction and fusion modules respectively to obtain the optimized visible visual feature map and synthetic visual feature maps ; The optimized visible visual feature map Input Discriminator , calculate the RINCE loss As shown below: ; Where: , s represents the similarity score of the optimized visible visual feature map; is the similarity score of the positive sample pair, representing two optimized visible visual feature maps Similarity of belonging to the same category; is the similarity score of the negative sample pair, representing two optimized visible visual feature maps The similarity between different categories; K represents the number of negative sample pairs; q represents the parameter that controls the behavior of the loss function; Represents the weight balance parameter of positive and negative sample pairs; It represents the base of the exponential function, which is an irrational number; The optimized visible visual feature map Input Discriminator , calculate the classification loss As shown below: ; Where: M represents the total number of visible visual feature maps of the input; C represents the total number of categories; Represents the true label of the input visible visual feature map on category c; Representation Discriminator Predict the probability that the visible visual feature map belongs to category c; The visible visual feature map and synthetic visual feature maps Input discriminator Get the output of the discriminator and ,in, ; ; Where: Represents visible visual feature map Distribution of Representation Generator Generated synthetic visual feature map Distribution of Represents the discriminator giving visible visual feature map score; Represents the discriminator giving visible visual feature map score; According to the output of the discriminator, the Wasserstein distance loss is obtained The expression is as follows: 。 4. The fine-grained image generation method based on generalized zero-shot learning according to claim 3, characterized in that: The second loss function for: ; Where: , and Respectively And the weight coefficient corresponding to each loss; Among them, the RINCE loss and classification loss Obtained through the following steps: Combined with the optimized visible visual feature map and synthetic visual feature maps , calculate the RINCE loss As shown below: ; Where: , Represents the optimized visible visual feature map and synthetic visual feature maps Similarity score of is the similarity score of the positive sample pair, representing the visible visual feature map and synthetic visual feature maps Similarity of belonging to the same category; is the similarity score of the negative sample pair, indicating the visible visual feature map and synthetic visual feature maps Similarity between categories; Synthesize visual feature map Input Discriminator , calculate the classification loss As shown below: ; Where: Represents the total number of synthetic visual feature maps of the input; Represents the true label of the input synthetic visual feature map on category c; Representation Discriminator Predict the probability that the synthesized visual feature map belongs to category c.
5. The fine-grained image generation method based on generalized zero-shot learning according to claim 1, characterized in that: The semantic features of the unseen categories are fed into the trained generator to generate the corresponding optimized visual feature maps, followed by: The generalized zero-shot learning classifier is trained by combining the optimized visual feature map and the real visual feature map to obtain an optimized classifier.
6. A fine-grained image generation device based on generalized zero-shot learning, characterized in that: include: A dynamic frequency domain and spatial domain feature extraction and fusion module is used to receive a real visual feature map and a synthetic visual feature map, extract the spatial domain features and frequency domain features of the image, and generate an optimized real visual feature map and an optimized synthetic visual feature map; A discriminator, used to calculate the distribution difference and category distinction ability between features based on the optimized real visual feature map, the optimized synthetic visual feature map and the unoptimized visual feature map, and update the discriminator parameters through a preset first loss function; A generator, used to generate a synthetic visual feature map according to the input semantic features, and based on the feedback result of the discriminator, optimize the category distinction ability and semantic consistency of the generated visual feature map through a preset second loss function; Generate an adversarial network, including the discriminator and the generator, to build an adversarial training mechanism, wherein: The generator generates an optimized visual feature map based on the input invisible category semantic features; The discriminator optimizes the ability to distinguish between real visual features and synthetic visual features by comparing the distribution between features.
7. The fine-grained image generation device based on generalized zero-shot learning according to claim 6, characterized in that: Also includes: The model training unit is used to perform classification training on the generalized zero-shot learning classifier in combination with the optimized visual feature map and the real visual feature map to obtain an optimized classifier.
8. An electronic terminal, characterized in that: The method comprises a processor and a memory connected to the processor, wherein a computer program is stored in the memory, and when the computer program is executed by the processor, the steps of the method for generating fine-grained images based on generalized zero-shot learning as described in any one of claims 1 to 5 are performed.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the fine-grained image generation method based on generalized zero-shot learning described in any one of claims 1 to 5 are implemented.