A zero-shot learning classification method and device based on multi-modal feature fusion
Through multimodal feature fusion and generative adversarial network semantic and visual modal alignment, the problems of domain offset and visual feature offset in the zero-sample learning model are solved, the recognition accuracy of unseen classes is improved, and it is suitable for image recognition in the real world open domain environment.
Patent Information
- Application Number
- CN202210207381.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-03
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2042-03-03
AI Technical Summary
There are domain offset problems and visual feature domain offset problems in the existing zero-sample learning model based on generative models, resulting in low recognition accuracy of unseen classes.
By extracting semantic features and visual principal features from the training samples for multimodal fusion, the multimodal fusion condition features are obtained, and synthetic visual features are generated through the generation of adversarial networks, semantic and visual modal alignment are performed, generator parameters are optimized to reduce losses, and accurate unseen pseudo-samples are generated for training softmax classifiers.
It improves the recognition accuracy of the unknown class, alleviates the domain bias problem, ensures that the generator generates visual features of the unknown class that conform to semantic descriptions without deviating from the visual principal features, and solves the image recognition problem in the real world open domain environment.
Smart Images

Figure CN114821148B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image recognition, and in particular, to a zero-shot learning classification method and device based on multi-modal feature fusion. Background Art
[0002] In recent years, supervised learning has achieved remarkable success in image classification tasks. Thanks to the use of deep learning frameworks and the increasing availability of labeled datasets for training, models can achieve high-precision recognition results through sufficient training. However, there are two challenges in the current classification tasks: one is the high cost of collecting large-scale datasets, and the other is the huge difficulty and time-consuming of sample collection in the face of continuously emerging new classes. To solve this problem, researchers have obtained inspiration from the process of human cognition of new things and proposed zero-shot learning (ZSL) to achieve the recognition of novel classes. Zero-shot learning aims to classify unseen classes by using the knowledge learned from seen classes.
[0003] Semantic information serves as an intermediate bridge connecting seen classes and unseen classes, and uses the knowledge of seen classes learned during the training process to recognize unseen classes. The zero-shot learning model based on the generative model is a kind of zero-shot learning method, which can improve the recognition accuracy of unseen classes in generalized zero-shot learning. However, there are two problems in the existing zero-shot learning models based on the generative model: the visual features of unseen classes generated by the trained generator in the model according to semantic descriptions have insufficient inter-class distinguishability, resulting in a domain shift problem when the classifier classifies unseen classes; the visual features of samples on different datasets will have cross-domain biases due to the influence of human factors during the data collection process, resulting in differences in sample distributions between different datasets. Therefore, there will be a visual feature domain bias problem in the residual network features used in the zero-shot learning benchmark dataset. Summary of the Invention
[0004] The embodiments of the present application provide a zero-shot learning classification method and device based on multi-modal feature fusion to solve the following technical problems: there are domain shift problems and visual feature domain shift problems in the existing zero-shot learning models based on the generative model, which have a greater impact on the recognition accuracy of unseen classes in generalized zero-shot learning.
[0005] The embodiments of the present application adopt the following technical solutions:
[0006] On the one hand, an embodiment of the present application provides a zero-shot learning classification method based on multi-modal feature fusion. The method includes: obtaining multi-modal fusion conditional features according to the semantic features and visual principal component features of training samples; obtaining synthetic visual features according to the true features of the training samples and the multi-modal fusion conditional features, and calculating the encoding loss function and discriminator loss function of the synthetic visual features; mapping the synthetic visual features through a first encoder to obtain semantic embedding features, and calculating the cyclic consistency loss between the semantic features and the semantic embedding features to obtain a semantic modality alignment loss function; reconstructing the semantic embedding features through the generator of a generative adversarial network to obtain reconstructed sample visual features, and calculating a visual modality alignment loss function; optimizing relevant parameters in the generator according to the total model loss function until the value of the total model loss function is less than a first preset threshold; wherein, the total model loss function is determined by the encoding loss function, the discriminator loss function, the semantic modality alignment loss function, and the visual modality alignment loss function; classifying unseen class image samples according to the optimized generator of the generative adversarial network to obtain corresponding unseen class pseudo-samples, so as to use the unseen class pseudo-samples for training a classifier.
[0007] In the embodiment of the present application, semantic features and visual principal component features extracted from training samples are fused to obtain multi-modal fusion conditional features, and then they are calculated with the true features in the training samples. After passing through the encoder, random noise is obtained. Then the random noise is combined with the semantic features, and after the operation of the generator of the generative adversarial network, synthetic visual features are obtained. Then, according to the synthetic visual features, true features and semantic features, a discriminator loss function is obtained, and the encoding loss function of the second encoder is calculated in this process. Then, according to the first encoder, the synthetic visual features are mapped to obtain semantic embedding features, and then the cyclic consistency loss between the semantic features and the semantic embedding features is calculated to obtain the semantic modality alignment loss function in the bimodal alignment. Then, according to the reconstruction of the semantic embedding features by the generator of the generative adversarial network, reconstructed visual features are obtained, and then the Euclidean distance between the true features and the reconstructed features is calculated to obtain the loss between visual modalities, and the visual modality alignment loss function in the dual-modal alignment is obtained. Finally, the sum of the encoding loss function, discriminator loss function, semantic modality alignment loss function and visual modality alignment loss function is used to generate the total loss function of the model, and then the relevant parameters of the generator are iteratively optimized. Finally, the generator of the optimized generative adversarial network is used to classify the image samples of unseen classes to obtain the pseudo-samples of unseen classes, and then these pseudo-samples of unseen classes are used to train the softmax classifier. By training with the optimized pseudo-samples of unseen classes generated by the optimized generator, the recognition accuracy of the softmax classifier for unseen class samples can be improved, the domain bias problem in the generation of unseen class samples in the zero-shot learning method can be alleviated, and the generator can be constrained simultaneously in the semantic modality and the visual modality, so that the generator can randomly generate different visual features of unseen classes according to the semantic description without deviating from the visual principal component features, and the image recognition problem in the real-world open-domain environment can be solved.
[0008] In a feasible embodiment, according to the semantic features and visual principal component features of the training samples, multi-modal fusion conditional features are obtained, which specifically include: extracting the true features in the training samples through the pre-trained model ResNet-101; wherein, the true features are 2048-dimensional visual feature vectors; generalizing the class features of the training samples to extract the semantic features; extracting the visual principal component features in the training samples through a deep principal component feature extraction network; and performing feature extraction and feature fusion on the training samples according to the semantic features and the visual principal component features to obtain the multi-modal fusion conditional features.
[0009] In a feasible embodiment, according to the semantic features and the visual principal component features, performing feature extraction and feature fusion on the training samples to obtain the multi-modal fusion conditional features specifically includes: performing feature extraction on the training samples through a feature extraction function; according to Le = E[logθ(x)] to obtain the loss of the feature extraction process; where x is the true feature, θ(·) is the feature extraction function, and E is the expected value; through the feature layer fusion module, according to perform feature fusion on the semantic feature and the visual principal feature to obtain the multi-modal fusion conditional feature c; where x p is the visual principal feature, a is the semantic feature, is the concatenation symbol.
[0010] In a feasible implementation manner, according to the true feature of the training sample and the multi-modal fusion conditional feature, a synthetic visual feature is obtained, and the encoding loss function and the discriminator loss function of the synthetic visual feature are calculated, specifically including: through the second encoder, encoding the true feature and the multi-modal fusion conditional feature to obtain random noise; according to to obtain the encoding loss function where z is the random noise, E(x, c) is the expectation of the second encoder, logG(z, a) is the reconstruction error of the generator of the generative adversarial network, KL(•) is used to calculate the KL divergence distance, β is the weight parameter of the KL divergence, p(z|a) represents the prior probability of the Gaussian distribution, a is the semantic feature, c is the multi-modal fusion conditional feature, and E is the expectation; through the decoder of the variational autoencoder VAE, decoding the random noise and the semantic feature to obtain the synthetic visual feature; where the generator of the generative adversarial network shares the decoder of the variational autoencoder VAE; through the discriminator of the adversarial generative network, calculating the similarity between the true feature and the synthetic visual feature; according to to obtain the discriminator loss function where is the similarity between the true feature x and the synthetic visual feature λE[(||D(x′, a)|| 2 - 1) 2 is the gradient penalty term with Lipschitz constraint, λ is the penalty parameter, x′ is the joint distribution of the semantic-visual feature, where α ∼ U(0, 1).
[0011] In a feasible implementation manner, through the first encoder, mapping the synthetic visual feature to obtain a semantic embedding feature, and calculating the cyclic consistency loss between the semantic feature and the semantic embedding feature to obtain a semantic modality alignment loss function, specifically including: according to perform modality alignment on the semantic embedding feature and the semantic feature to obtain the semantic modality alignment loss function Among them, Enc() is an encoding operation used to obtain the semantic embedding features a is the semantic feature, and E is the expectation; among them, the modality alignment is used to increase the intra-class compactness and inter-class separability between the training samples.
[0012] In a feasible implementation manner, the generator of the generative adversarial network is used to reconstruct the semantic embedding features to obtain reconstructed sample visual features, and the visual modality alignment loss function is calculated. Specifically, it includes: using the generator of the generative adversarial network to reconstruct the semantic embedding features to obtain reconstructed sample visual features; performing modality alignment on the synthetic visual features and the reconstructed sample visual features; according to to obtain the visual modality alignment loss function; among them, is the reconstructed sample visual feature, Dec() is the decoding operation, x is the real feature, and E is the expectation.
[0013] In a feasible implementation manner, according to the total model loss function, the relevant parameters in the generator are optimized until the value of the total model loss function is less than a first preset threshold. Specifically, it includes: according to to obtain the total model loss function L; among them, is the bimodal alignment loss function, is the semantic modality alignment loss function, is the visual modality alignment loss function; among them, β is the hyperparameter of the dual-modal alignment loss, is the encoding loss function, is the discriminator loss function; according to the gradient descent method, multiple rounds of gradient descent are performed on the relevant parameters of the generator of the generative adversarial network to reduce the loss amount; after the value of the total model loss function is less than the first preset threshold, the optimal relevant parameters of the generator are obtained to realize the iterative optimization of the generator.
[0014] In a feasible implementation manner, according to the optimized generator of the generative adversarial network, the unseen class image samples are classified to obtain corresponding unseen class pseudo-samples, and the unseen class pseudo-samples are used to train the classifier. Specifically, it includes: extracting the semantic features of the unseen class images through the backbone network of the pre-trained model ResNet-101; decoding the semantic features and random noise through the optimized generator of the generative adversarial network to obtain the corresponding unseen class pseudo-samples; training the softmax classifier according to the unseen class pseudo-samples and the categories of the unseen class images to obtain the trained softmax classifier.
[0015] In a feasible implementation, performing modality alignment on the synthetic visual features and the reconstructed sample visual features specifically includes: calculating an error based on the Euclidean distance between the synthetic visual features and the reconstructed sample visual features to obtain a visual modality alignment loss value, thereby completing the modality alignment of the synthetic visual features and the reconstructed sample visual features.
[0016] On the other hand, an embodiment of the present application also provides a zero-shot learning classification device based on multi-modal feature fusion, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, so that the at least one processor can execute a zero-shot learning classification method based on multi-modal feature fusion according to any one of the above embodiments.
[0017] The embodiment of the present application provides a zero-shot learning classification method and device based on multi-modal feature fusion. In order to improve the accuracy of unseen class recognition, a multi-modal feature fusion module is added, which utilizes the complementarity between low-dimensional visual principal features and semantic features. The visual principal features are used to improve the distinguishability of the corresponding features of the semantic features, and the semantic features constrain different data sets to alleviate the domain bias problem in cross-datasets. Moreover, semantic modality alignment and visual modality alignment are adopted to increase intra-class compactness and inter-class separability, and the generator is optimized to generate unseen class pseudo-samples with higher accuracy, thereby training a softmax classifier. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments recorded in the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts. In the drawings:
[0019] Figure 1 is a flowchart of a zero-shot learning classification method based on multi-modal feature fusion provided by an embodiment of the present application;
[0020] Figure 2 is a schematic structural diagram of a multi-modal feature fusion model provided by an embodiment of the present application;
[0021] Figure 3 is a schematic structural diagram of a multi-modal feature fusion module provided by an embodiment of the present application;
[0022] Figure 4 is a schematic structural diagram of a dual modality alignment module provided by an embodiment of the present application;
[0023] Figure 5 Schematic diagram of a zero-shot learning classification device based on multi-modal feature fusion provided by an embodiment of the present application. Detailed implementation manners
[0024] In order to enable those skilled in the art to better understand the technical solutions in the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of this specification, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0025] An embodiment of the present application provides a zero-shot learning classification method based on multi-modal feature fusion. The execution subject of this method is a zero-shot learning classification device based on multi-modal feature fusion. As Figure 1 shown, this method specifically includes steps S101-S106:
[0026] S101. Obtain multi-modal fusion conditional features according to the semantic features and visual principal component features of the training samples.
[0027] Specifically, first, the real features in the training samples are extracted through the pre-trained model ResNet-101. Among them, the real features are 2048-dimensional visual feature vectors. Then, the category features of the training samples are generalized to extract semantic features. The visual principal component features in the training samples are extracted through the PCANet deep principal component feature extraction network. According to the semantic features and visual principal component features, feature extraction and feature fusion are performed on the training samples to obtain multi-modal fusion conditional features.
[0028] As a feasible implementation manner, the feature extraction function in the pre-trained model ResNet-101 is used to extract features from the training samples. According to L e = E[logθ(x)], the loss of the feature extraction process is obtained. Where x is the real feature, θ(·) is the feature extraction function, and E is the expected value. Through the feature layer fusion module, according to the semantic features and visual principal component features are fused to obtain the multi-modal fusion conditional feature c. Where x p is the visual principal component feature, a is the semantic feature, is the concatenation symbol.
[0029] In one embodiment, Figure 3 Schematic diagram of a multi-modal feature fusion module structure provided by an embodiment of the present application. As Figure 3As shown, the training samples images pass through the pre-trained model ResNet-101 to extract the true features x of the training samples. At the same time, through the PCANet deep principal feature extraction network, the visual principal features x of the training samples are extracted. p , and then the visual principal features x p are fused with the semantic features a to finally obtain the multi-modal fusion conditional features c. A multi-modal feature fusion module is added. By using the low-dimensional visual principal features x p to complement the semantic features a, the visual principal features x p are used to improve the discriminability of the corresponding features generated by the semantics. The semantic features a can constrain the visual features in different data sets to alleviate the domain bias problem existing across data sets.
[0030] S102. According to the true features and multi-modal fusion conditional features of the training samples, synthetic visual features are obtained, and the encoding loss function and discriminator loss function of the synthetic visual features are calculated.
[0031] Specifically, through the encoder of the variational autoencoder VAE (i.e., the second encoder), the true features and multi-modal fusion conditional features are encoded to obtain random noise. According to the encoding loss function is obtained where z is the random noise, E(x, c) is the expectation of the second encoder, logG(z, a) is the reconstruction error of the generator of the generative adversarial network, KL(·) is used to calculate the KL divergence distance, β is the weight parameter of the KL divergence, p(z|a) represents the prior probability of the Gaussian distribution, a is the semantic feature, c is the multi-modal fusion conditional feature, and E is the expectation.
[0032] Further, through the decoder of the variational autoencoder VAE, the random noise and semantic features are decoded to obtain synthetic visual features. Among them, the generator of the generative adversarial network shares the decoder of the variational autoencoder VAE.
[0033] Further, through the discriminator of the generative adversarial network, the similarity between the true features and the synthetic visual features is calculated; according to the discriminator loss function is obtained where, is the similarity between the true feature x and the synthetic visual feature , and λE[(||D(x′, a)|| 2 - 1) 2 ] is the gradient penalty term with Lipschitz constraint, λ is the penalty parameter, x′ is the joint distribution of semantic-visual features, where α ∼ U(0, 1).
[0034] In one embodiment,Figure 2 A schematic diagram of a multi-modal feature fusion model structure provided by an embodiment of the present application is as follows Figure 2 As shown, through the encoder E of the variational autoencoder VAE, that is, the second encoder in the present application, the real feature x extracted by the visual feature extraction network and the multi-modal fusion conditional feature c formed through the multi-modal feature fusion module are jointly input into the encoder E to generate a random noise z, and then the generated random noise z and the semantic feature a are input into the generator G. Among them, the generator G is the decoder of the variational autoencoder VAE in the present application, and then the synthetic visual feature is obtained In this process, according to the encoding loss function of the encoder E and the decoder G can be obtained
[0035] In Figure 2 , through the discriminator D, the similarity between the real feature x and the synthetic visual feature is calculated. The generative adversarial network GAN module consists of the generator G and the discriminator D. Because the training of the GAN module is not easy to converge, the Wasserstein distance in WGAN is introduced to calculate the similarity between the generated real feature x and the synthetic visual feature . According to
[0036] the discriminator loss function related to the parameters of the discriminator D is obtained Among them, is the similarity between the real feature x and the synthetic visual feature , λE[(||D(x′, a)|| 2 -1) 2 is the gradient penalty term with Lipschitz constraint, λ is the penalty parameter, x′ is the joint distribution of semantic-visual features where α~U(0, 1).
[0037] S103. Through the first encoder, the synthetic visual feature is mapped to obtain the semantic embedding feature, and the cyclic consistency loss between the semantic feature and the semantic embedding feature is calculated to obtain the semantic modality alignment loss function
[0038] Specifically, according to the semantic embedding feature and the semantic feature are modality-aligned to obtain the semantic modality alignment loss function where Enc() is the encoding operation for obtaining the semantic embedding feature a is the semantic feature, and E is the expectation. Among them, modality alignment is used to increase the intra-class compactness and inter-class separability between the training samples
[0039] In one embodiment, Figure 4 is a schematic structural diagram of a dual-modal alignment module provided by an embodiment of the present application. As Figure 4 shown, through the encoder Enc in the autoencoder, that is, the first encoder in the present application, the synthetic visual features are mapped to obtain semantic embedding features In order to enable the generator G to effectively learn the mapping function from semantics to the visual space and generate the corresponding visual features correctly according to the semantic feature a, the mapped semantic embedding features are aligned with the semantic feature a of the real sample, and the semantic-modal alignment loss function is obtained by calculating the two Thereby realizing the semantic-modal alignment between the semantic embedding features and the semantic feature a.
[0040] S104. Through the generator of the generative adversarial network, the semantic embedding features are reconstructed to obtain reconstructed sample visual features, and the visual-modal alignment loss function is calculated.
[0041] Specifically, through the generator of the generative adversarial network, the semantic embedding features are reconstructed to obtain reconstructed sample visual features. The synthetic visual features and the reconstructed sample visual features are aligned in modality. According to the visual-modal alignment loss function is obtained. Wherein, is the reconstructed sample visual feature, Dec() is the decoding operation, x is the real feature, and E is the expectation.
[0042] Furthermore, according to the Euclidean distance between the synthetic visual features and the reconstructed sample visual features, the error is calculated to obtain the value of the visual-modal alignment loss, and the modality alignment between the synthetic visual features and the reconstructed sample visual features is realized.
[0043] In one embodiment, as Figure 4 shown, through the decoder Dec, where the decoder Dec is the generator G in Figure 3 That is to say, the generator G of this generative adversarial network has different functions in different steps. Then the semantic embedding features are reconstructed to obtain reconstructed sample visual features To ensure that the synthetic visual features generated by the generator G can correspond to the semantic description, the reconstructed sample visual features are aligned with the synthetic visual features in modality, and by calculating the real feature x and the reconstructed sample visual features The Euclidean distance between them is used to obtain the loss between visual modalities, and the value of the visual modality alignment loss is obtained. According to the value of the loss, iterative optimization is continuously performed to reduce the loss error and make it closer to the visual features of real samples, so as to alleviate the cross-domain bias problem caused by cross-datasets, and then the visual modality alignment loss function is obtained.
[0044] S105. According to the total model loss function, relevant parameters in the generator of the generative adversarial network are optimized until the value of the total model loss function is less than the first preset threshold.
[0045] Specifically, according to the total model loss function L is obtained. Among them, is the bimodal alignment loss function, is the semantic modality alignment loss function, is the visual modality alignment loss function; among them, β is the hyperparameter of the bimodal alignment loss, is the encoding loss function, is the discriminator loss function. According to the gradient descent method, multiple rounds of gradient descent are performed on the relevant parameters of the generator of the generative adversarial network to reduce the loss amount. After the value of the total model loss function is less than the first preset threshold, the optimal relevant parameters of the generator are obtained, and iterative optimization of the generator is realized.
[0046] In one embodiment, the total model loss function includes an encoding loss function, a discriminator loss function, a semantic modality alignment loss function, and a visual modality alignment loss function. Among them, the semantic modality alignment loss function and the visual modality alignment loss function are bimodal alignment loss functions, which can increase the intra-class compactness and inter-class separability. The mapping alignment between the two modalities ensures that the model can generate visual features corresponding to semantic features. Then, according to the gradient descent method, when the value of the total model loss function is less than the first preset threshold, iterative optimization training is performed on the relevant parameters of the generator of the generative adversarial network, so that the generator obtains the optimal relevant parameters, gradually reduces the loss amount, and generates more accurate unseen class samples, thereby realizing the improvement of the accuracy of unseen class recognition.
[0047] S106. According to the optimized generator of the generative adversarial network, the unseen class image samples are classified to obtain the corresponding unseen class pseudo-samples, so as to use the unseen class pseudo-samples to train the classifier.
[0048] Specifically, the semantic features of unseen-class images are extracted through the backbone network of the pre-trained model ResNet-101. The generator of the optimized generative adversarial network decodes the semantic features and random noise to obtain the corresponding unseen-class pseudo-samples. The softmax classifier is trained based on the unseen-class pseudo-samples and the categories of the unseen-class images to obtain a trained softmax classifier.
[0049] In one embodiment, we use the test set to test the trained softmax classifier to classify unseen-class samples. The performance of the model is evaluated on 4 zero-shot learning recognition benchmark datasets: AWA1 (Animals with Attributes1), AWA2 (Animals with Attributes2), SUN (SUN Attribute), CUB (Caltech-UCSD-Birds). The detailed information of the datasets is shown in the table:
[0050] Dataset Visual feature dimension Attribute dimension Total number of samples Number of visible classes Number of unseen classes SUN 2048 102 14340 645 72 CUB 2048 312 11788 150 50 AWA1 2048 85 30475 40 10 AWA2 2048 85 37322 40 10
[0051] Table 1
[0052] Table 1 shows the detailed information of the samples and visible-class / invisible-class divisions of the 4 benchmark datasets. In the 4 commonly used zero-shot learning benchmark datasets, we refer to the visible-class / invisible-class division method of PS (Proposed Splits). Therefore, the PS V2.0 division scheme is adopted in all 4 benchmark datasets. At the same time, the 2048-dimensional visual features used in the model testing process are all extracted by the ResNet-101 pre-trained model. The test is carried out under this setting, and the experimental results all reach the best performance.
[0053]
[0054] Table 2
[0055] Table 2 shows the Top-1 recognition accuracy (%) of unseen classes on the benchmark datasets AWA1, AWA2, SUN, and CUB under the zero-shot learning setting.
[0056] Table 2 shows the latest comparison on four benchmark datasets. TF-VAEGAN is one of the existing best-performing zero-shot learning methods, and the top-1 recognition accuracy on the 4 benchmark datasets all exceeds that of TF-VAEGAN. It also improves by 2.3%, 1.3%, 0.8%, and 0.2% on AWA1, CUB, SUN, and AWA2 respectively. In past work, there was a deviation between the generator and the semantic description when generating pseudo-samples for unseen classes, resulting in a bias problem of visible classes in the generation method, leading to poor classification effects for unseen classes. After complementary fusion through the visual-semantic modality, the bias problem in the existing generation model can be effectively improved.
[0057]
[0058]
[0059] Table 3
[0060] Table 3 shows the test results of MMF in the GZSL setting on the benchmark datasets AWA1, AWA2, SUN, and CUB.
[0061] Table 3 shows the recognition accuracies of different methods on unseen classes, visible classes, and the corresponding harmonic means on the 4 benchmark datasets. Our method achieved the best results on the AWA1, AWA2, SUN, and CUB benchmark datasets, and the harmonic means H are 67.8%, 66.7%, 42.9%, and 58.1% respectively.
[0062] Through the tests on the 4 test sets, it can be clearly seen that the generator of the trained and optimized generative adversarial network has the effect of significantly improving the accuracy of pseudo-samples for unseen classes. The double alignment of the pseudo-samples in the visual modality and the semantic modality ensures that the generator of the generative adversarial network can generate the visual features required by the category to alleviate the domain shift problem.
[0063] In addition, the embodiment of the present application also provides a zero-shot learning classification device based on multi-modal feature fusion, as Figure 5 shown. The zero-shot learning classification device 500 based on multi-modal feature fusion specifically includes:
[0064] At least one processor 501. And a memory 502 communicatively connected to at least one processor 501. Among them, the memory 502 stores instructions that can be executed by at least one processor 501, so that at least one processor 501 can execute:
[0065] Obtain multi-modal fusion conditional features according to the semantic features and visual principal features of the training samples;
[0066] Based on the true features of the training samples and the multi-modal fusion conditional features, synthetic visual features are obtained, and the encoding loss function and discriminator loss function of the synthetic visual features are calculated;
[0067] Through the first encoder, the synthetic visual features are mapped to obtain semantic embedding features, and the cyclic consistency loss between the semantic features and the semantic embedding features is calculated to obtain the semantic modality alignment loss function;
[0068] Through the generator of the generative adversarial network, the semantic embedding features are reconstructed to obtain reconstructed sample visual features, and the visual modality alignment loss function is calculated;
[0069] According to the total model loss function, the relevant parameters in the generator are optimized until the value of the total model loss function is less than the first preset threshold; wherein, the total model loss function is determined by the encoding loss function, the discriminator loss function, the semantic modality alignment loss function, and the visual modality alignment loss function;
[0070] According to the optimized generator of the generative adversarial network, the unseen class image samples are classified to obtain corresponding unseen class pseudo-samples, so as to use the unseen class pseudo-samples to train the softmax classifier.
[0071] The embodiments of the present application can be trained with the optimized unseen class pseudo-samples generated by the optimized generator, which can improve the recognition accuracy of unseen classes, alleviate the domain bias problem in the generation of unseen class samples in the zero-shot learning method, and can simultaneously constrain the generator in the semantic modality and the visual modality, enabling the generator to randomly generate different unseen class visual features according to the semantic description without deviating from the main visual features, and solving the image recognition problem in the real-world open-domain environment.
[0072] The various embodiments in the present application are all described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device and non-volatile computer storage medium embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiments.
[0073] The specific embodiments of the present application are described above. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the specific order or continuous order shown to achieve the desired results. In certain embodiments, multi-tasking and parallel processing are also possible or may be advantageous.
[0074] The above are only embodiments of the present application and are not intended to limit the present application. For those skilled in the art, various changes and modifications can be made to the embodiments of the present application. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the embodiments of the present application shall be included within the scope of the claims of the present application.
Claims
1. A zero - shot learning classification method based on multi - modal feature fusion, characterized in that, the method includes: Obtaining multi - modal fusion conditional features according to the semantic features and visual principal component features of training samples; Obtaining synthetic visual features according to the true features of the training samples and the multi - modal fusion conditional features, and calculating the encoding loss function and discriminator loss function of the synthetic visual features; Mapping the synthetic visual features through a first encoder to obtain semantic embedding features, and calculating the cycle - consistency loss between the semantic features and the semantic embedding features to obtain a semantic modality alignment loss function; Reconstructing the semantic embedding features through the generator of the generative adversarial network to obtain reconstructed sample visual features, and calculating a visual modality alignment loss function; Optimizing the relevant parameters in the generator according to the total model loss function until the value of the total model loss function is less than a first preset threshold; wherein, the total model loss function is determined by the encoding loss function, the discriminator loss function, the semantic modality alignment loss function, and the visual modality alignment loss function; Classifying unseen - class image samples according to the optimized generator of the generative adversarial network to obtain corresponding unseen - class pseudo - samples, and using the unseen - class pseudo - samples to train a classifier.
2. The zero - shot learning classification method based on multi - modal feature fusion according to claim 1, characterized in that, Obtaining multi - modal fusion conditional features according to the semantic features and visual principal component features of training samples specifically includes: Extracting the true features in the training samples through a pre - trained model ResNet - 101; wherein, the true features are 2048 - dimensional visual feature vectors; Generalizing the class features of the training samples to extract the semantic features; Extracting the visual principal component features in the training samples through a deep principal component feature extraction network; Performing feature extraction and feature fusion on the training samples according to the semantic features and the visual principal component features to obtain the multi - modal fusion conditional features.
3. The zero - shot learning classification method based on multi - modal feature fusion according to claim 2, characterized in that, Performing feature extraction and feature fusion on the training samples according to the semantic features and the visual principal component features to obtain the multi - modal fusion conditional features specifically includes: Performing feature extraction on the training samples through a feature extraction function; According to L e = E[logθ(x)], the loss of the feature extraction process is obtained; where x is the true feature, θ(·) is the feature extraction function, and E is the expected value; Through the feature layer fusion module, according to perform feature fusion on the semantic feature and the visual principal feature to obtain the multi-modal fusion conditional feature c; where x p is the visual principal feature, a is the semantic feature,[[]] is the connection symbol.
4. The zero - shot learning classification method based on multi - modal feature fusion according to claim 1, characterized in that, Obtaining synthetic visual features according to the true features of the training samples and the multi - modal fusion conditional features, and calculating the encoding loss function and discriminator loss function of the synthetic visual features specifically includes: Encoding the true features and the multi - modal fusion conditional features through a second encoder to obtain random noise; According to obtain the encoding loss function where z is random noise, E(x, c) is the expectation of the second encoder, logG(z, a) is the reconstruction error of the generator of the generative adversarial network, KL(·) is used to calculate the KL divergence distance, β is the weight parameter of the KL divergence, p(z|a) represents the prior probability of the Gaussian distribution, a is the semantic feature, c is the multi-modal fusion conditional feature, and E is the expectation; Decode the random noise and the semantic features through the decoder of the variational autoencoder (VAE) to obtain the synthetic visual features, where the generator of the generative adversarial network shares the decoder of the variational autoencoder (VAE). Calculate the similarity between the real features and the synthetic visual features through the discriminator of the generative adversarial network. According to obtain the discriminator loss function wherein is the similarity between the real feature x and the synthetic visual feature , λE[(||D(x′, a)|| 2 - 1) 2 is a gradient penalty term with Lipschitz constraint, λ is a penalty parameter, x′ is the joint distribution of semantic-visual features where α ∼ U(0, 1).
5. A zero-shot learning classification method based on multi-modal feature fusion according to claim 1,[[]]END]] wherein,[[]]END]] Map the synthetic visual features through a first encoder to obtain semantic embedding features, and calculate the cyclic consistency loss between the semantic features and the semantic embedding features to obtain a semantic modality alignment loss function, specifically including:[[]]END]] According to perform modal alignment on the semantic embedding feature and the semantic feature to obtain the semantic modal alignment loss function Among them, Enc() is an encoding operation used to obtain the semantic embedding features a is the semantic feature, and E is the expectation; among them, the modality alignment is used to increase the intra-class compactness and inter-class separability among the training samples.
6. A zero-shot learning classification method based on multi-modal feature fusion according to claim 1,[[]]END]] wherein,[[]]END]] Reconstruct the semantic embedding features through the generator of the generative adversarial network to obtain reconstructed sample visual features, and calculate a visual modality alignment loss function, specifically including:[[]]END]] Reconstruct the semantic embedding features through the generator of the generative adversarial network to obtain reconstructed sample visual features; Perform modality alignment on the synthetic visual features and the reconstructed sample visual features; According to obtain the visual modality alignment loss function wherein is the visual feature of the reconstructed sample, Dec() is the decoding operation, x is the true feature, and E is the expectation.
7. A zero-shot learning classification method based on multi-modal feature fusion according to claim 1,[[]]END]] wherein,[[]]END]] Optimize the relevant parameters in the generator according to the total model loss function until the value of the total model loss function is less than a first preset threshold, specifically including:[[]]END]] According to obtain the total loss function L of the model; where is the bimodal alignment loss function,[[]] is the semantic modality alignment loss function,[[]] is the visual modality alignment loss function; where β is the hyperparameter of the bimodal alignment loss,[[]] is the encoding loss function,[[]] is the discriminator loss function.[[]] According to the gradient descent method, perform multiple rounds of gradient descent on the relevant parameters of the generator of the generative adversarial network to reduce the loss amount; After the value of the total model loss function is less than the first preset threshold, obtain the optimal relevant parameters of the generator to achieve iterative optimization of the generator.
8. A zero-shot learning classification method based on multi-modal feature fusion according to claim 1,[[]]END]] wherein,[[]]END]] Classify the unseen class image samples according to the optimized generator of the generative adversarial network to obtain corresponding unseen class pseudo-samples, and use the unseen class pseudo-samples to train a classifier, specifically including:[[]]END]] Extract the semantic features of the unseen class images through the backbone network of the pre-trained model ResNet-101; Decode the semantic features and random noise through the optimized generator of the generative adversarial network to obtain the corresponding unseen class pseudo-samples; Train the softmax classifier according to the unseen class pseudo-samples and the categories of the unseen class images to obtain the trained softmax classifier.
9. A zero-shot learning classification method based on multi-modal feature fusion according to claim 6,[[]]END]] wherein,[[]]END]] Perform modality alignment on the synthetic visual features and the reconstructed sample visual features, specifically including:[[]]END]] Calculate the error according to the Euclidean distance between the synthetic visual features and the reconstructed sample visual features to obtain a visual modality alignment loss value, and complete the modality alignment of the synthetic visual features and the reconstructed sample visual features.
10. A zero-shot learning classification device based on multi-modal feature fusion, characterized in that, the device comprises: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor to enable the at least one processor to execute a zero-shot learning classification method according to any one of claims 1-9.
Citation Information
Patent Citations
Zero-sample image classification method based on variational self-coding adversarial network
CN110580501A
Zero-sample image recognition method and system based on generative adversarial network
CN111476294A