Multi-mode-based generative generalized zero sample learning method

Through the improved Gaussian hybrid variational autoencoder and cross-modal comparison learning method, the generative generalized zero-sample learning method solves the problem of insufficient cross-modal feature alignment and generation feature diversity, and realizes efficient multimodal information fusion and high-quality representation of unknown category features.

CN120105202AInactive Publication Date: 2025-06-06GUANGDONG UNIV OF TECH
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202510265734.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-07
Publication Date
2025-06-06
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The existing cross-modal comparison learning method ignores the differences in cross-modal feature distribution when constructing positive and negative sample pairs, resulting in the model being unable to fully utilize the complementary information of vision and text when feature alignment. The features generated by the zero-sample learning method based on the generative model lack diversity and consistency, making it difficult to adapt to complex unknown category distributions.

Method used

A generative generalized zero-sample learning method based on multimodality is proposed to model the known category feature distribution through an improved Gaussian hybrid variational autoencoder, and generate high-quality unknown category features in combination with regularization constraints, maximum mean difference optimization and adversarial training strategies. At the same time, cross-modal contrast learning is introduced to optimize the alignment of visual and text features.

Benefits of technology

This method can efficiently integrate multimodal information in zero-sample task, improve the generalization ability and classification accuracy of unknown categories, and produce high quality of features and good diversity and consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120105202A_ABST
    Figure CN120105202A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-modal-based generative generalized zero sample learning method. The method comprises the following steps: S1, collecting and preprocessing image and text data; s2, optimizing image features by adopting self-supervised learning, and generating visual feature vectors; s3, text features are enhanced through context attention, and text feature vectors are generated; s4, constructing a cross-modal embedding space, aligning visual and text features, and constructing a positive and negative sample pair; s5, calculating the similarity of the sample pair, and optimizing cross-modal feature distribution; s6, introducing regularization constraint, modeling known category features by using the improved Gaussian mixture variational auto-encoder, and generating unknown category features; and S7, training a classification model, and performing category prediction in the optimized cross-modal embedding space. According to the method, visual and text features are fused, cross-modal alignment is optimized, the quality of unknown category features is improved, and the defects of an existing zero sample learning method in cross-modal alignment and generalization ability are overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of cross-modal representation learning and generative zero-shot classification, and in particular to a multi-modal based generative generalized zero-shot learning method. Background Art

[0002] In the development of computer vision and natural language processing, traditional supervised learning methods usually rely on a large amount of labeled data. However, in real application scenarios, it is often difficult to obtain labeled data, especially in some rare categories, emerging categories or open world environments, the model may not be able to obtain enough training samples to learn the representation of these categories. Therefore, zero-shot learning (ZSL) has become an important research direction in the field of machine learning in recent years, aiming to enable the model to correctly classify or predict the target category even if it has not seen the target category during the training phase. The core idea of ​​the zero-shot learning method is to transfer the knowledge of known categories to unknown categories through shared feature space or semantic information, thereby improving the generalization ability of the model.

[0003] In the research of zero-shot learning, common methods include attribute-based zero-shot learning, embedding-based zero-shot learning, and generative model-based zero-shot learning. Attribute-based methods map visual features to semantic space through predefined semantic attribute vectors to infer unknown categories. However, such methods rely on manually designed semantic labels and are difficult to scale to larger datasets. Embedding-based methods achieve zero-shot classification by learning a shared feature space so that visual features and semantic features are comparable in this space. However, such methods usually perform poorly in classification tasks of unknown categories. The main reason is the deviation of category distribution, that is, the distribution of known categories during training and unknown categories during testing in the feature space does not match, which makes the model difficult to generalize. Generative model-based methods attempt to simulate the missing data distribution by generating features of unknown categories, thereby converting it into a standard supervised learning problem. Such methods usually use variational autoencoders (VAE) or generative adversarial networks (GAN) to synthesize features of unknown categories, but since the generated feature distribution still has a certain deviation from the true distribution, the classification performance is still subject to certain limitations.

[0004] In addition, multimodal learning has become an important direction of artificial intelligence research in recent years. The core idea of ​​multimodal learning is to combine information from different modalities (such as vision and text) to improve the performance of tasks and overcome the limitations of single modal information. In zero-shot learning tasks, visual data and text data are the two main modalities. Visual data usually includes images or videos, while text data can be category names, descriptive information, or other semantic features. Traditional zero-shot learning methods often rely only on the features of a single modality for reasoning, which makes it difficult for the model to fully utilize cross-modal information. For example, zero-shot learning methods based on visual features are easily restricted by data distribution, while zero-shot learning methods based only on text descriptions are difficult to capture the fine-grained features of images. How to achieve efficient zero-shot learning in a multimodal environment has become one of the key challenges of current research.

[0005] At present, some studies have begun to try to optimize the alignment of multimodal features through cross-modal contrastive learning, so that visual features and text features can be matched in a shared embedding space. Such methods usually use a contrastive loss function to make visual and text features from the same category closer in the feature space, while features from different categories are farther away. However, existing cross-modal contrastive learning methods still face the following problems: First, when constructing positive and negative sample pairs, existing methods often ignore the differences in cross-modal feature distribution, resulting in the model being unable to fully utilize the complementary information of vision and text when aligning features. Second, cross-modal contrastive learning methods usually rely on a large number of samples of known categories, while in zero-shot learning scenarios, it is difficult to accurately represent features of unknown categories by directly optimizing contrastive losses.

[0006] On the other hand, zero-shot learning methods based on generative models usually rely on variational autoencoders (VAE) or generative adversarial networks (GAN) to generate features of unknown categories. However, these methods have two main problems in the generation process: first, the generated features may lack diversity, making it difficult for the model to adapt to complex unknown category distributions; second, existing generation methods usually ignore the constraints of cross-modal information and only synthesize features through information from a single modality, resulting in a lack of consistency in the generated features between different modalities. Therefore, how to find an effective combination between cross-modal learning and generative methods to enhance the feature representation capabilities of unknown categories is an important challenge facing current research.

[0007] Therefore, how to provide a multimodal-based generative generalized zero-shot learning method so that the model can efficiently integrate multimodal information in zero-shot tasks while improving the generalization ability of unknown categories is a problem that technical personnel in this field urgently need to solve. Summary of the invention

[0008] One purpose of the present invention is to propose a generative generalized zero-shot learning method based on multimodality. The present invention models the known category feature distribution through an improved Gaussian mixture variational autoencoder, and combines regularization constraints, maximum mean difference optimization and adversarial training strategies to generate high-quality unknown category features. At the same time, cross-modal contrastive learning is introduced to optimize the alignment of visual and text features. The present invention has the advantages of strong cross-modal information fusion capability, high quality of generated features, and strong generalization ability, which improves the classification accuracy of zero-shot learning tasks.

[0009] A generative generalized zero-shot learning method based on multimodality according to an embodiment of the present invention includes the following steps:

[0010] S1, collecting image data and text data, and preprocessing the image data and text data respectively;

[0011] S2, using a self-supervised learning strategy to optimize the features of the preprocessed image data, learning the association between local features and global features of the image through the mask reconstruction task, and generating an optimized visual feature vector;

[0012] S3, using the contextual attention mechanism to semantically enhance the preprocessed text data, extracting the global semantic features of the text by modeling the long-distance dependencies in the text sequence, and generating an optimized text feature vector;

[0013] S4. Construct a cross-modal embedding space, map the optimized visual feature vector and text feature vector to the same feature space, align the visual features and text features using a feature transformation network, and construct positive sample pairs of the same category and negative sample pairs of different categories;

[0014] S5. Calculate the feature similarity between the positive sample pairs and the negative sample pairs, adjust the feature mapping of the positive sample pairs and the negative sample pairs, minimize the distance difference between the positive sample pairs and maximize the distance difference between the negative sample pairs, and optimize the feature distribution of the cross-modal embedding space;

[0015] S6. Introducing regularization constraints, using improved Gaussian mixture variational autoencoders to model the feature distribution of known categories, combining Gaussian mixture models to cluster discrete latent variables, and generating new latent features to enhance the feature representation capabilities of unknown categories;

[0016] S7. Based on the enhanced unknown category features, the classification model is trained, category prediction is performed in the optimized cross-modal embedding space, feature extraction is performed on the input data, and category determination is performed based on the multimodal feature distribution.

[0017] Optionally, the preprocessing includes discretizing the image data, extracting discrete latent image features and converting them into visual feature vector representations, continuously encoding the text data, extracting continuous text semantic features, and converting them into text feature vector representations of the same dimension as the latent image features.

[0018] Optionally, the S4 specifically includes:

[0019] S41. Construct a cross-modal embedding space to optimize the visual feature vector and text feature vector Normalize and use the feature projection function f v (·) and f t (·) Perform feature projection and convert it into a vector representation of the same dimension:

[0020] V′=f v (V),T′=f t (T);

[0021] in, and are the converted visual features and text features respectively, d v ,d t They represent the dimensions of the original visual and text features, d is the common feature dimension after projection, and f v (·) and f t (·) Feature projection is achieved through linear transformation:

[0022] f v (V) = W v V+b v ,f t (T) = W t T+b t ;

[0023] in, and is the projection matrix, b v ,b t is the bias vector;

[0024] S42, use the feature transformation network to align the visual features and text features, through the mapping matrix Perform a linear transformation to make the visual features and text features in the same feature space:

[0025] V″=W m V′,T″=W m T′;

[0026] in, and are the converted visual features and text features respectively, Wm To map the matrix, ensure that the visual and textual features remain aligned after projection;

[0027] S43. Use the cross-attention mechanism to calculate the mutual influence between cross-modal features and obtain the fused visual features and text features:

[0028]

[0029]

[0030] Among them, softmax(·) is an activation function, which normalizes the numerical vector into a probability distribution vector, and the sum of each probability is 1, W q ,W k A is the query matrix and the key matrix. v and A t are the attention weights of visual features and text features respectively, and are the adjusted visual features and text features;

[0031] S44. Calculate the similarity of the fused features and adjust the feature transformation network parameters to keep the visual features and text features aligned in the cross-modal embedding space:

[0032]

[0033] Among them, α is the balance coefficient, which controls the degree of fusion of the original features and the attention-enhanced features. and Visual features and text features in the final cross-modal embedding space;

[0034] S45, construct positive sample pairs and negative sample pairs, let the visual features V″′ of the same category i and text features T″′ i Composed of positive sample pairs, visual features of different categories V″′ i and text features T″′ j Form a negative sample pair, where i≠j.

[0035] Optionally, the S5 specifically includes:

[0036] S51, visual feature vector V″′ i and text feature vector T″′ j Calculate the normalized cosine similarity and construct the positive sample pair set P and the negative sample pair set N:

[0037]

[0038] Among them, sim(V″′ i ,T″′j ) is the feature similarity score between sample i and sample j, indicating the matching degree between visual features and text features, y i and j Represent the category labels of sample i and sample j respectively;

[0039] S52, calculate the feature similarity of the positive sample pairs, optimize the feature mapping of the positive sample pairs, and reduce the feature distance between the positive sample pairs through optimization, and define the loss function of the positive sample pairs:

[0040]

[0041] Introduce the aggregation loss of the positive sample distribution to minimize the internal variance of the positive sample pairs:

[0042]

[0043] Among them, L pos Optimize the loss for positive pairs, L var Optimize the loss for the positive sample distribution, μ v and μ t They are the means of visual features and text features on positive sample pairs:

[0044]

[0045] S53. Optimize the discrimination of negative sample pairs

[0046] Calculate the similarity between negative sample pairs, optimize the feature mapping of negative sample pairs, and increase the feature distance between negative sample pairs through optimization. Define the loss function of negative sample pairs:

[0047]

[0048] Introduce the minimum distribution interval constraint between negative sample pairs:

[0049]

[0050] Among them, L neg Optimize the loss for negative sample pairs, L margin is the negative sample interval optimization loss, and m is the preset minimum interval hyperparameter;

[0051] S54, comprehensive positive sample optimization loss L pos , negative samples optimize the loss L neg , positive sample distribution optimization loss L var and negative sample interval optimization loss L margin , jointly optimize the cross-modal feature distribution and define the final loss function:

[0052] L=λ1 L pos +λ 2 L neg +λ 3 L var +λ 4 L margin ;

[0053] Among them, λ 1 ,λ 2 ,λ 3 ,λ 4 is the weight coefficient, which is used to control the optimization impact of different loss terms;

[0054] S55. Use the gradient descent method to quadratically optimize the feature distribution of the cross-modal embedding space and update the parameters of the feature transformation network so that the optimized visual features and text features remain aligned in the cross-modal embedding space and have good distinguishing ability.

[0055] Optionally, the S6 specifically includes:

[0056] S61. Extract visual feature vectors based on the optimized cross-modal embedding space and text feature vector Calculate the feature mean vector μ and covariance matrix ∑ of known categories:

[0057]

[0058] Where N is the number of samples of known categories, d is the dimension of the feature, μ is the mean vector of the known category features, and ∑ is the covariance matrix of the known category features;

[0059] In order to improve the consistency of cross-modal features, an inter-modal alignment regularization constraint is introduced to keep the distribution of visual features and text features consistent:

[0060]

[0061] Among them, L align is the cross-modal alignment loss, ||V″′ i -T″′ i || 2 Measure the Euclidean distance between the i-th visual feature and the text feature to ensure that the two are consistent;

[0062] S62. Construct an improved Gaussian mixture variational autoencoder and define the prior distribution of the Gaussian mixture model:

[0063]

[0064] Among them, p(z) is the prior distribution, is a latent variable, representing the implicit feature vector, K is the number of components of the Gaussian mixture distribution, that is, the number of clusters, π k is the mixing coefficient of the kth Gaussian component, satisfying represents the weight of the component, is the kth Gaussian distribution, μ k and∑ k are the mean vector and covariance matrix respectively;

[0065] In the Gaussian mixture latent space, a variational distribution is constructed to approximate the true posterior distribution:

[0066]

[0067] Among them, q(z|x) is the posterior distribution, is the kth Gaussian distribution, r k represents the probability that sample x belongs to the kth cluster, μ′ k and ∑′ k are the mean vector and covariance matrix of the variational distribution respectively;

[0068] Based on the prior distribution and the posterior distribution, the optimization objective adopts a variational lower bound:

[0069]

[0070] Among them, L ELBO is the optimization objective of the variational autoencoder, is the expected value of the probability distribution of log p(x|z) under the distribution of q(z|x), D KL (q(z|x)||p(z)) is the Kullback-Leibler divergence, which is used to measure the deviation of the approximate posterior distribution from the prior distribution;

[0071] S63. Sampling is performed in the optimized Gaussian mixture latent space to generate enhanced latent features z:

[0072]

[0073] The decoder D(·) is used to transform the latent features into new visual features.

[0074]

[0075] in, is the generated visual feature of unknown category, D(·) is the decoding function of Gaussian mixture variational autoencoder;

[0076] S64, introduce the maximum mean difference regularization constraint to optimize the generated feature distribution:

[0077]

[0078] Among them, L MMD is the maximum mean difference regularization loss, k(x,y) is the kernel function, and M is the number of generated samples;

[0079] The regularization term is introduced to constrain the distribution of generated features to make them closer to known categories:

[0080]

[0081] Among them, L reg To generate the feature distribution regularization loss, tr(·) represents the trace operation of the matrix, constraining the covariance matrix of the generated features;

[0082] S65. Introduce adversarial training loss to improve the authenticity of generated features:

[0083]

[0084] Among them, L adv To combat the loss, D adv (·) is the adversarial discriminative network;

[0085] Comprehensive maximum mean difference regularization loss L MMD , Generate feature distribution regularization loss L reg , adversarial loss L adv and the cross-modal alignment loss L align , define the final optimization target L:

[0086] L=λ 1 L ELBO +λ 2 L MMD +λ 3 L reg +λ 4 L adv +λ 5 L align ;

[0087] Among them, λ 1 ,λ 2 ,λ 3 ,λ 4 ,λ 5 is a hyperparameter used to control the weight of each loss;

[0088] S66. Use gradient descent to optimize the Gaussian mixture variational autoencoder so that the generated unknown category features are more consistent with the known category features, thereby enhancing the feature representation capability of the unknown category.

[0089] The beneficial effects of the present invention are:

[0090] First, the present invention adopts an improved Gaussian mixture variational autoencoder to model the feature distribution of known categories, and combines the unknown category features generated by regularized constraint optimization, so that the distribution of generated features is closer to the real category features, thereby improving the representation ability of unseen categories.

[0091] Secondly, through a cross-modal contrastive learning strategy, the present invention constructs positive sample pairs and negative sample pairs, optimizes the alignment of visual and textual features in a shared embedding space, and ensures that data of different modalities are more closely matched in the feature space, thereby reducing the deviation of cross-modal information fusion.

[0092] Finally, the present invention introduces the maximum mean difference and adversarial training strategies, so that the generated unseen category features are not only more consistent with the known category features in statistical distribution, but also improves the diversity and discrimination ability of the generated features, thereby improving the accuracy of zero-shot classification tasks.

[0093] In summary, the present invention combines multimodal feature optimization, generative feature enhancement and cross-modal contrastive learning to enable the model to achieve better recognition effect on unseen categories, effectively overcoming the shortcomings of traditional zero-shot learning methods in cross-modal alignment, feature generation and generalization capabilities. BRIEF DESCRIPTION OF THE DRAWINGS

[0094] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0095] Figure 1 This is a flowchart of a multimodal generative generalized zero-shot learning method proposed in the present invention. DETAILED DESCRIPTION

[0096] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, which only illustrate the basic structure of the present invention in a schematic manner, and therefore only show the components related to the present invention.

[0097] refer to Figure 1 , a multimodal generative generalized zero-shot learning method, comprising the following steps:

[0098] S1, collecting image data and text data, and preprocessing the image data and text data respectively;

[0099] S2, using a self-supervised learning strategy to optimize the features of the preprocessed image data, learning the association between local features and global features of the image through the mask reconstruction task, and generating an optimized visual feature vector;

[0100] S3, using the contextual attention mechanism to semantically enhance the preprocessed text data, extracting the global semantic features of the text by modeling the long-distance dependencies in the text sequence, and generating an optimized text feature vector;

[0101] S4. Construct a cross-modal embedding space, map the optimized visual feature vector and text feature vector to the same feature space, align the visual features and text features using a feature transformation network, and construct positive sample pairs of the same category and negative sample pairs of different categories;

[0102] S5. Calculate the feature similarity between the positive sample pairs and the negative sample pairs, adjust the feature mapping of the positive sample pairs and the negative sample pairs, minimize the distance difference between the positive sample pairs and maximize the distance difference between the negative sample pairs, and optimize the feature distribution of the cross-modal embedding space;

[0103] S6. Introducing regularization constraints, using improved Gaussian mixture variational autoencoders to model the feature distribution of known categories, combining Gaussian mixture models to cluster discrete latent variables, and generating new latent features to enhance the feature representation capabilities of unknown categories;

[0104] S7. Based on the enhanced unknown category features, the classification model is trained, category prediction is performed in the optimized cross-modal embedding space, feature extraction is performed on the input data, and category determination is performed based on the multimodal feature distribution.

[0105] In this embodiment, the preprocessing includes discretizing the image data, extracting discrete latent image features and converting them into visual feature vector representations, continuously encoding the text data, extracting continuous text semantic features, and converting them into text feature vector representations of the same dimension as the latent image features.

[0106] In this implementation manner, the S4 specifically includes:

[0107] S41. Construct a cross-modal embedding space to optimize the visual feature vector and text feature vector Normalize and use the feature projection function f v (·) and f t (·) Perform feature projection and convert it into a vector representation of the same dimension:

[0108] V′=f v (V),T′=f t (T);

[0109] in, and are the converted visual features and text features respectively, d v ,d tThey represent the dimensions of the original visual and text features, d is the common feature dimension after projection, and f v (·) and f t (·) Feature projection is achieved through linear transformation:

[0110] f v (V) = W v V+b v ,f t (T) = W t T+b t ;

[0111] in, and is the projection matrix, b v ,b t is the bias vector;

[0112] S42, use the feature transformation network to align the visual features and text features, through the mapping matrix Perform a linear transformation to make the visual features and text features in the same feature space:

[0113] V″=W m V′,T″=W m T′;

[0114] in, and are the converted visual features and text features respectively, W m To map the matrix, ensure that the visual and textual features remain aligned after projection;

[0115] S43. Use the cross-attention mechanism to calculate the mutual influence between cross-modal features and obtain the fused visual features and text features:

[0116]

[0117]

[0118] Among them, softmax(·) is an activation function, which normalizes the numerical vector into a probability distribution vector, and the sum of each probability is 1, W q ,W k A is the query matrix and the key matrix. v and A t are the attention weights of visual features and text features respectively, and are the adjusted visual features and text features;

[0119] S44. Calculate the similarity of the fused features and adjust the feature transformation network parameters to keep the visual features and text features aligned in the cross-modal embedding space:

[0120]

[0121] Among them, α is the balance coefficient, which controls the degree of fusion of the original features and the attention-enhanced features. and Visual features and text features in the final cross-modal embedding space;

[0122] S45, construct positive sample pairs and negative sample pairs, let the visual features V″′ of the same category i and text features T″′ i Composed of positive sample pairs, visual features of different categories V″′ i and text features T″′ j Form a negative sample pair, where i≠j.

[0123] In this implementation manner, S5 specifically includes:

[0124] S51, visual feature vector V″′ i and text feature vector T″′ j Calculate the normalized cosine similarity and construct the positive sample pair set P and the negative sample pair set N:

[0125]

[0126] Among them, sim(V″′ i ,T″′ j ) is the feature similarity score between sample i and sample j, indicating the matching degree between visual features and text features, y i and j Represent the category labels of sample i and sample j respectively;

[0127] S52, calculate the feature similarity of the positive sample pairs, optimize the feature mapping of the positive sample pairs, and reduce the feature distance between the positive sample pairs through optimization, and define the loss function of the positive sample pairs:

[0128]

[0129] Introduce the aggregation loss of the positive sample distribution to minimize the internal variance of the positive sample pairs:

[0130]

[0131] Among them, L pos Optimize the loss for positive pairs, L var Optimize the loss for the positive sample distribution, μ v and μ t They are the means of visual features and text features on positive sample pairs:

[0132]

[0133] S53. Optimize the discrimination of negative sample pairs

[0134] Calculate the similarity between negative sample pairs, optimize the feature mapping of negative sample pairs, and increase the feature distance between negative sample pairs through optimization. Define the loss function of negative sample pairs:

[0135]

[0136] Introduce the minimum distribution interval constraint between negative sample pairs:

[0137]

[0138] Among them, L neg Optimize the loss for negative sample pairs, L margin is the negative sample interval optimization loss, and m is the preset minimum interval hyperparameter;

[0139] S54, comprehensive positive sample optimization loss L pos , negative samples optimize the loss L neg , positive sample distribution optimization loss L var and negative sample interval optimization loss L margin , jointly optimize the cross-modal feature distribution and define the final loss function:

[0140] L=λ 1 L pos +λ 2 L neg +λ 3 L var +λ 4 L margin ;

[0141] Among them, λ 1 ,λ 2 ,λ 3 ,λ 4 is the weight coefficient, which is used to control the optimization impact of different loss terms;

[0142] S55. Use the gradient descent method to quadratically optimize the feature distribution of the cross-modal embedding space and update the parameters of the feature transformation network so that the optimized visual features and text features remain aligned in the cross-modal embedding space and have good distinguishing ability.

[0143] In this implementation manner, S6 specifically includes:

[0144] S61. Extract visual feature vectors based on the optimized cross-modal embedding space and text feature vector Calculate the feature mean vector μ and covariance matrix ∑ of known categories:

[0145]

[0146] Where N is the number of samples of known categories, d is the dimension of the feature, μ is the mean vector of the known category features, and ∑ is the covariance matrix of the known category features;

[0147] In order to improve the consistency of cross-modal features, an inter-modal alignment regularization constraint is introduced to keep the distribution of visual features and text features consistent:

[0148]

[0149] Among them, L align is the cross-modal alignment loss, ||V″′ i -T″′ i || 2 Measure the Euclidean distance between the i-th visual feature and the text feature to ensure that the two are consistent;

[0150] S62. Construct an improved Gaussian mixture variational autoencoder and define the prior distribution of the Gaussian mixture model:

[0151]

[0152] Among them, p(z) is the prior distribution, is a latent variable, representing the implicit feature vector, K is the number of components of the Gaussian mixture distribution, that is, the number of clusters, π k is the mixing coefficient of the kth Gaussian component, satisfying represents the weight of the component, is the kth Gaussian distribution, μ k and∑ k are the mean vector and covariance matrix respectively;

[0153] In the Gaussian mixture latent space, a variational distribution is constructed to approximate the true posterior distribution:

[0154]

[0155] Among them, q(z|x) is the posterior distribution, is the kth Gaussian distribution, r k represents the probability that sample x belongs to the kth cluster, μ′ k and ∑′ k are the mean vector and covariance matrix of the variational distribution respectively;

[0156] Based on the prior distribution and the posterior distribution, the optimization objective adopts a variational lower bound:

[0157]

[0158] Among them, L ELBO is the optimization objective of the variational autoencoder, is the expected value of the probability distribution of log p(x|z) under the distribution of q(z|x), D KL (q(z|x)||p(z)) is the Kullback-Leibler divergence, which is used to measure the deviation of the approximate posterior distribution from the prior distribution;

[0159] S63. Sampling is performed in the optimized Gaussian mixture latent space to generate enhanced latent features z:

[0160]

[0161] The decoder D(·) is used to transform the latent features into new visual features.

[0162]

[0163] in, is the generated visual feature of unknown category, D(·) is the decoding function of Gaussian mixture variational autoencoder;

[0164] S64, introduce the maximum mean difference regularization constraint to optimize the generated feature distribution:

[0165]

[0166] Among them, L MMD is the maximum mean difference regularization loss, k(x,y) is the kernel function, and M is the number of generated samples;

[0167] The regularization term is introduced to constrain the distribution of generated features to make them closer to known categories:

[0168]

[0169] Among them, L reg To generate the feature distribution regularization loss, tr(·) represents the trace operation of the matrix, constraining the covariance matrix of the generated features;

[0170] S65. Introduce adversarial training loss to improve the authenticity of generated features:

[0171]

[0172] Among them, L adv To combat the loss, D adv (·) is the adversarial discriminative network;

[0173] Comprehensive maximum mean difference regularization loss LMMD , Generate feature distribution regularization loss L reg , adversarial loss L adv and the cross-modal alignment loss L align , define the final optimization target L:

[0174] L=λ 1 L ELBO +λ 2 L MMD +λ 3 L reg +λ 4 L adv +λ 5 L align ;

[0175] Among them, λ 1 ,λ 2 ,λ 3 ,λ 4 ,λ 5 is a hyperparameter used to control the weight of each loss;

[0176] S66. Use gradient descent to optimize the Gaussian mixture variational autoencoder so that the generated unknown category features are more consistent with the known category features, thereby enhancing the feature representation capability of the unknown category.

[0177] Embodiment 1:

[0178] In order to verify the feasibility of the present invention in implementation, the present invention is applied to the open world image classification task, and the problem of difficulty in cross-modal feature alignment and insufficient generalization ability of unseen categories in zero-shot learning is solved through a multimodal generative generalized zero-shot learning method. This experiment selected CUB-200-2011 (fine-grained bird classification dataset) and AWA2 (animal attribute dataset) from Kaggle public datasets, and used standard zero-shot learning partitioning to ensure that the training set and test set categories are mutually exclusive, ensuring that the model can only rely on information of known categories for reasoning during testing.

[0179] In this scenario, traditional zero-shot learning methods usually use only visual features for category reasoning, or rely solely on text descriptions for classification. However, the reality is that image and text information are often complementary, and it is difficult to ensure classification accuracy by relying solely on information from a single modality. For example, in the CUB-200-2011 dataset, many birds have similar appearances, which are difficult to distinguish using traditional visual classification methods, while text descriptions can provide additional fine-grained semantic information such as "beak length" and "feather color". In addition, in the AWA2 dataset, some animal categories (such as "sea lions" and "seals") are extremely similar in images, and traditional methods cannot effectively distinguish them, while their attribute features (such as "ear size" and "whether there are long whiskers") are clearly distinguished in text descriptions. Therefore, the present invention uses cross-modal feature alignment and generative enhancement strategies to improve the classification accuracy of the model in these scenarios.

[0180] In this experiment, the dataset is first preprocessed to convert image data into visual feature vectors and text data (category descriptions) into semantic feature vectors. Then, the features of known categories are modeled using an improved Gaussian mixture variational autoencoder, and the alignment of visual features and text features is optimized using a cross-modal contrastive learning strategy. Finally, the generated features of unknown categories are optimized using the maximum mean difference and adversarial training strategies, and these features are used for classification.

[0181] The main evaluation indicators of the experiment include Top-1 classification accuracy, F1 score and feature distribution deviation. The comparative experimental results of the proposed method and the traditional method on CUB-200-2011 and AWA2 datasets are as follows:

[0182] Table 1 Comparative experimental results on CUB-200-2011 and AWA2 datasets

[0183]

[0184]

[0185] As shown in Table 1, on the CUB-200-2011 dataset, the Top-1 classification accuracy of the method of the present invention reached 74.3%, which is 6.5 percentage points higher than the traditional vision-based zero-shot learning method (such as DeepVAE's 67.8%), and 9.1 percentage points higher than the zero-shot learning method that only uses text information (such as SynC's 65.2%). On the AWA2 dataset, the Top-1 classification accuracy of the present invention reached 73.5%, which is 5.2 percentage points higher than the traditional method (such as f-CLSWGAN's 68.3%), and has stronger generalization ability on unseen categories. In addition, the feature distribution deviation of the method of the present invention is reduced by 18.7%, indicating that the unknown category features generated by it are more consistent with the statistical distribution of known categories, which improves the stability of classification.

[0186] In terms of computational efficiency, the method of the present invention is trained on an NVIDIA Tesla V100 GPU, and the time required for training 20 epochs on the CUB-200-2011 dataset is 3.4 hours, which reduces the training time by 34.6% compared with the traditional GAN ​​generation method (such as 5.2 hours of f-CLSWGAN). At the same time, the reasoning time of the method of the present invention is only 38.2ms / image, which is 33.3% less than the traditional zero-sample classification method based on deep network (such as 57.3ms / image of DEM), which improves the reasoning efficiency in practical applications.

[0187] The above experimental results show that the proposed method is superior to the existing methods on both CUB-200-2011 and AWA2 datasets, and shows great advantages in multiple dimensions such as classification accuracy, feature distribution optimization, training time and inference time, which fully verifies its effectiveness and application value in zero-shot learning tasks.

[0188] In summary, in the open-world image classification task, the present invention significantly improves the classification accuracy and generalization ability of the model in zero-shot tasks through multimodal feature alignment, generative feature optimization and cross-modal comparative learning strategies, while reducing the training time and inference delay, making it more suitable for large-scale practical application scenarios.

[0189] The above description is only a preferred specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any technician familiar with the technical field can make equivalent replacements or changes according to the technical scheme and inventive concept of the present invention within the technical scope disclosed by the present invention, which should be covered by the protection scope of the present invention.

Claims

1. A generative generalized zero-shot learning method based on multimodality, characterized in that: The steps include: S1, collecting image data and text data, and preprocessing the image data and text data respectively; S2, using a self-supervised learning strategy to optimize the features of the preprocessed image data, learning the association between local features and global features of the image through the mask reconstruction task, and generating an optimized visual feature vector; S3, using the contextual attention mechanism to semantically enhance the preprocessed text data, extracting the global semantic features of the text by modeling the long-distance dependencies in the text sequence, and generating an optimized text feature vector; S4. Construct a cross-modal embedding space, map the optimized visual feature vector and text feature vector to the same feature space, align the visual features and text features using a feature transformation network, and construct positive sample pairs of the same category and negative sample pairs of different categories; S5. Calculate the feature similarity between the positive sample pairs and the negative sample pairs, adjust the feature mapping of the positive sample pairs and the negative sample pairs, minimize the distance difference between the positive sample pairs and maximize the distance difference between the negative sample pairs, and optimize the feature distribution of the cross-modal embedding space; S6. Introducing regularization constraints, using improved Gaussian mixture variational autoencoders to model the feature distribution of known categories, combining Gaussian mixture models to cluster discrete latent variables, and generating new latent features to enhance the feature representation capabilities of unknown categories; S7. Based on the enhanced unknown category features, the classification model is trained, category prediction is performed in the optimized cross-modal embedding space, feature extraction is performed on the input data, and category determination is performed based on the multimodal feature distribution.

2. The method of generative generalized zero-shot learning based on multimodality according to claim 1, characterized in that: The preprocessing includes discretizing the image data, extracting discrete latent image features and converting them into visual feature vector representations, continuously encoding the text data, extracting continuous text semantic features, and converting them into text feature vector representations with the same dimension as the latent image features.

3. The method of generative generalized zero-shot learning based on multimodality according to claim 1, characterized in that: The S4 specifically includes: S41. Construct a cross-modal embedding space to optimize the visual feature vector and text feature vector Normalize and use the feature projection function f v (·) and f t (·) Perform feature projection and convert it into a vector representation of the same dimension: V′=f v (V),T′=f t (T); in, and are the converted visual features and text features respectively, d v ,d t They represent the dimensions of the original visual and text features, d is the common feature dimension after projection, and f v (·) and f t (·) Feature projection is achieved through linear transformation: <h2 style=";text-align:left;direction:ltr">f<h2 style=";text-align:left;direction:ltr"> v <h2 style=";text-align:left;direction:ltr"> (V)=W<h2 style=";text-align:left;direction:ltr"> v <h2 style=";text-align:left;direction:ltr"> V+b<h2 style=";text-align:left;direction:ltr"> v <h2 style=";text-align:left;direction:ltr"> ,f<h2 style=";text-align:left;direction:ltr"> t <h2 style=";text-align:left;direction:ltr"> 9T)=W<h2 style=";text-align:left;direction:ltr"> t <h2 style=";text-align:left;direction:ltr"> T+b<h2 style=";text-align:left;direction:ltr"> t <h2 style=";text-align:left;direction:ltr"> ; in, and is the projection matrix, b v ,b t is the bias vector; S42, use the feature transformation network to align the visual features and text features, through the mapping matrix Perform a linear transformation to make the visual features and text features in the same feature space: V″=W m V′,T″=W m T′; in, and are the converted visual features and text features respectively, W m To map the matrix, ensure that the visual and textual features remain aligned after projection; S43. Use the cross-attention mechanism to calculate the mutual influence between cross-modal features and obtain the fused visual features and text features: Among them, softmax(·) is an activation function, which normalizes the numerical vector into a probability distribution vector, and the sum of each probability is 1, W q ,W k A is the query matrix and the key matrix. v and A t are the attention weights of visual features and text features respectively, and are the adjusted visual features and text features; S44. Calculate the similarity of the fused features and adjust the feature transformation network parameters to keep the visual features and text features aligned in the cross-modal embedding space: Among them, α is the balance coefficient, which controls the degree of fusion of the original features and the attention-enhanced features. and Visual features and text features in the final cross-modal embedding space; S45, construct positive sample pairs and negative sample pairs, let the visual features V″′ of the same category i and text features T″′ i Composed of positive sample pairs, visual features of different categories V″′ i and text features T″′ j Form a negative sample pair, where i≠j.

4. The method of generative generalized zero-shot learning based on multimodality according to claim 1, characterized in that: The S5 specifically includes: S51, visual feature vector V″′ i and text feature vector T″′ j Calculate the normalized cosine similarity and construct the positive sample pair set P and the negative sample pair set N: P = {(V″′ i ,T″′ i )|y i =y j },N={(V″′ i ,T″′ j )|y i ≠y j }; where sim(V″′ i ,T″′ j ) is the feature similarity score between sample i and sample j, indicating the matching degree between visual features and text features, y i and j Represent the category labels of sample i and sample j respectively; S52, calculate the feature similarity of the positive sample pairs, optimize the feature mapping of the positive sample pairs, and reduce the feature distance between the positive sample pairs through optimization, and define the loss function of the positive sample pairs: Introduce the aggregation loss of the positive sample distribution to minimize the internal variance of the positive sample pairs: Among them, L pos Optimize the loss for positive pairs, L var Optimize the loss for the positive sample distribution, μ v and μ t They are the means of visual features and text features on positive sample pairs: S53. Calculate the similarity between negative sample pairs, optimize the feature mapping of negative sample pairs, and increase the feature distance between negative sample pairs through optimization, and define the loss function of negative sample pairs: Introduce the minimum distribution interval constraint between negative sample pairs: Among them, L neg Optimize the loss for negative sample pairs, L margin is the negative sample interval optimization loss, and m is the preset minimum interval hyperparameter; S54, comprehensive positive sample optimization loss L pos , negative samples optimize the loss L neg , positive sample distribution optimization loss L var and negative sample interval optimization loss L margin , jointly optimize the cross-modal feature distribution and define the final loss function: L=λ1L pos +λ2L neg +λ3L var +λ4L margin ; Among them, λ1, λ2, λ3, λ4 are weight coefficients used to control the optimization impact of different loss terms; S55. Use the gradient descent method to quadratically optimize the feature distribution of the cross-modal embedding space and update the parameters of the feature transformation network so that the optimized visual features and text features remain aligned in the cross-modal embedding space and have good distinguishing ability.

5. The method of generative generalized zero-shot learning based on multimodality according to claim 1, characterized in that: The S6 specifically includes: S61. Extract visual feature vectors based on the optimized cross-modal embedding space and text feature vector Calculate the feature mean vector μ and covariance matrix ∑ of known categories: Where N is the number of samples of known categories, d is the dimension of the feature, μ is the mean vector of the known category features, and ∑ is the covariance matrix of the known category features; In order to improve the consistency of cross-modal features, an inter-modal alignment regularization constraint is introduced to keep the distribution of visual features and text features consistent: Among them, L align is the cross-modal alignment loss, ||V′″ i -T″′ i || 2 Measure the Euclidean distance between the i-th visual feature and the text feature to ensure that the two are consistent; S62. Construct an improved Gaussian mixture variational autoencoder and define the prior distribution of the Gaussian mixture model: Among them, p(z) is the prior distribution, is a latent variable, representing the implicit feature vector, K is the number of components of the Gaussian mixture distribution, that is, the number of clusters, π k is the mixing coefficient of the kth Gaussian component, satisfying represents the weight of the component, is the kth Gaussian distribution, μ k and∑ k are the mean vector and covariance matrix respectively; In the Gaussian mixture latent space, a variational distribution is constructed to approximate the true posterior distribution: Among them, q(z|x) is the posterior distribution, is the kth Gaussian distribution, r k represents the probability that sample x belongs to the kth cluster, μ′ k and ∑′ k are the mean vector and covariance matrix of the variational distribution respectively; Based on the prior distribution and the posterior distribution, the optimization objective adopts a variational lower bound: Among them, L ELBO is the optimization objective of the variational autoencoder, is the expected value of the probability distribution of log p(x|z) under the distribution of q(z|x), D KL (q(z|x)||p(z)) is the Kullback-Leibler divergence, which is used to measure the deviation of the approximate posterior distribution from the prior distribution; S63. Sampling is performed in the optimized Gaussian mixture latent space to generate enhanced latent features z: The decoder D(·) is used to transform the latent features into new visual features. in, is the generated visual feature of unknown category, D(·) is the decoding function of Gaussian mixture variational autoencoder; S64, introduce the maximum mean difference regularization constraint to optimize the generated feature distribution: Among them, L MMD is the maximum mean difference regularization loss, k(x,y) is the kernel function, and M is the number of generated samples; The regularization term is introduced to constrain the distribution of generated features to make them closer to known categories: Among them, L reg To generate the feature distribution regularization loss, tr(·) represents the trace operation of the matrix, constraining the covariance matrix of the generated features; S65. Introduce adversarial training loss to improve the authenticity of generated features: Among them, L adv To combat the loss, D adv (·) is the adversarial discriminative network; Comprehensive maximum mean difference regularization loss L MMD , Generate feature distribution regularization loss L reg , adversarial loss L adv and the cross-modal alignment loss L align , define the final optimization target L: L=λ1L ELBO +λ2L MMD +λ3L reg +λ4L adv +λ5L align ; Among them, λ1,λ2,λ3,λ4,λ5 are hyperparameters used to control the weights of each loss; S66. Use gradient descent to optimize the Gaussian mixture variational autoencoder so that the generated unknown category features are more consistent with the known category features, thereby enhancing the feature representation capability of the unknown category.

Citation Information

Cited By

  • Glioma postoperative radiotherapy prediction method based on iconomics of multi-modal MRI (Magnetic Resonance Imaging)

    CN120413056A

  • Multi-modal feature alignment method based on implicit feature space

    CN120705808A

  • PCIS early risk prediction system driven by multi-modal data

    CN121054265A

  • Multi-modal industrial anomaly detection and segmentation method based on category perception

    CN121304656A

  • Continuous training method of cross-modal retrieval model based on depth evidence correction

    CN122432364A