Generative zero sample learning method based on cross-modal attention and multi-stage generation

Through multi-stage feature generation and cross-modal attention mechanism, the problem of taking into account both the global structure and local details in generative zero-sample learning is solved, the category distinction and recognition accuracy of generated samples are improved, and the generalization ability of the model is enhanced.

CN120277413APending Publication Date: 2025-07-08NORTHEASTERN UNIV CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510436997.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-09
Publication Date
2025-07-08

AI Technical Summary

Technical Problem

The existing generative zero-sample learning method is difficult to take into account the global structure and local details. The fusion method of semantic information and visual features is simple and there is a lack of dynamic interactive modeling, resulting in insufficient inter-class distinction of generated samples, affecting the classification effect.

Method used

A multi-stage feature generation strategy is adopted to generate coarse-grained features through the first stage and fine-grained feature optimization is used in the second stage, combining the Wasserstein GAN framework and the cross-modal attention mechanism to gradually improve the quality of the generated samples.

Benefits of technology

The recognition accuracy of unseen categories and the in-class consistency and inter-class distinction of generated samples are improved, and the generalization ability, stability and separability of generated samples are enhanced in the zero-sample learning task.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120277413A_ABST
    Figure CN120277413A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of computer vision and artificial intelligence, and discloses a generative zero sample learning method based on cross-modal attention and multi-stage generation. Coarse-grained features are generated through a first-stage generator; the coarse-grained features are combined with the semantic attribute vectors to be input into a second-stage generator for semantic enhancement, and fine-grained feature samples are generated; calculating a loss function according to the first-stage discriminator and the second-stage discriminator, and training all generators and discriminators; and the trained generator synthesizes unseen category samples for training a classifier to complete identification of a target category. According to the method, the generated samples are gradually optimized from coarse granularity to fine granularity, and the recognition precision of unseen categories is improved. The cross-modal attention mechanism enhances the matching degree of semantic information and visual features, so that the generated samples are better in intra-class consistency and inter-class discrimination.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision and artificial intelligence, and particularly relates to a generative zero-shot learning method based on cross-modal attention and multi-stage generation. Background Art

[0002] Zero-Shot Learning (ZSL) is a computer vision technology aiming to solve the problem of classifying models for unseen classes. Traditional deep learning methods rely on a large amount of labeled data for training, while zero-shot learning enables the model to recognize classes not present in the training set by leveraging the semantic information of classes. Generalized Zero-Shot Learning (GZSL) further expands the application scope, enabling the model to recognize both seen and unseen classes simultaneously and improving its applicability in real-world applications.

[0003] Existing zero-shot learning methods are mainly divided into two categories. One category is the method based on the embedding space, which maps visual features and semantic information into the same feature space, enabling samples of unseen classes to be classified through semantic information. This type of method relies on the learning of the mapping function, but in practical applications, it is easily affected by data distribution biases, resulting in the model being more inclined to predict seen classes. The other category is the method based on generative models, such as generative adversarial networks (GANs) and variational autoencoders (VAEs). By learning the feature distribution of seen classes and generating synthetic samples of unseen classes, the zero-shot learning problem is transformed into a standard supervised learning task. This method can alleviate the problem of the lack of samples of unseen classes and improve the generalization ability of the model. "Mishra A, Krishna Reddy S, Mittal A, et al. A generative model for zero-shot learning using conditional variational autoencoders [C] / / Proceedings of the IEEE conference on computer vision and pattern recognition workshops. 2018: 2188-2196." proposed a zero-shot learning framework based on conditional variational autoencoders (CVAEs), which uses a probabilistic generative model to generate a sample distribution from class attributes, laying the foundation for generative zero-shot learning. "Xian Y, Lorenz T, Schiele B, et al. Feature generating networks for zero-shot learning [C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2018: 5542-5551." introduced generative adversarial networks (GANs) and proposed a mapping mechanism from semantic information to visual features, enabling the zero-shot learning task to be transformed into a supervised learning problem, thereby enhancing the classification ability of the model.Subsequently, "Xian Y, Sharma S, Schiele B, et al. f-vaegan-d2: A feature generating framework for any-shot learning [C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2019: 10275-10284." further combined the advantages of VAE and GAN, and proposed the F-VAEGAN-D2 method, which stably generated high-quality unseen class features and extended to few-shot learning tasks. In recent years, zero-shot learning methods based on generative models have become a research hotspot, and researchers have continuously improved the stability and semantic consistency of generative models to further enhance the generalization ability of zero-shot learning.

[0004] Generative zero-shot learning converts zero-shot learning into a supervised learning task by synthesizing features, enabling the model to recognize unseen classes. However, its core challenge lies in how to utilize limited semantic attribute information to guide the generator to synthesize unseen class features that are both discriminative and have generalization ability. Most existing methods rely on generative adversarial networks (GANs), where the generator generates features and the discriminator is used for constraint to make the distribution of the generated features as close as possible to that of real samples, while introducing a class prototype alignment mechanism to enhance semantic control ability. However, the current generative zero-shot learning methods still have the following key problems.

[0005] First of all, it is difficult for a single-stage generation framework to balance global structure and local details. Existing methods usually adopt a single-stage generation strategy, directly generating visual features from semantic information in one step, lacking a step-by-step optimization mechanism. This results in insufficient hierarchical representation ability of the generated features, making it difficult to ensure the integrity of the global semantic structure and fine-grained feature expression simultaneously.

[0006] Secondly, the fusion method of semantic information and visual features is relatively simple, lacking dynamic interaction modeling. Existing methods often only splice semantic information with noise and then input it into the generator, without fully modeling the dynamic relationship between semantic attributes and visual features. Since the semantic attributes of different categories play different roles in the visual space, this simple splicing method is difficult to capture the correlation between attributes and local features, resulting in insufficiently delicate feature expressions of the generated samples. For example, "Li J, Jing M, Lu K, et al. Leveraging the invariant side of generative zero-shot learning [C] / / Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. 2019: 7402-7411." directly uses the spliced semantic information to control the generation process and cannot be dynamically adjusted according to the specific feature requirements of different categories, thus limiting the discrimination and semantic expression ability of the generated samples. Due to the lack of an effective local semantic enhancement mechanism in existing methods, the inter-class discrimination of the generated samples is insufficient, resulting in a large overlap in the distribution of samples of different categories in the feature space, further affecting the final classification effect. Summary of the Invention

[0007] The object of the present invention is to propose a generative zero-shot learning method based on cross-modal attention and multi-stage generation to improve the quality of generated samples, enhance the matching of semantic information and visual features, and thus improve the recognition performance of unseen categories. This method uses a multi-stage feature generation strategy to generate coarse-grained features in the first stage to capture the overall structure of the category, and uses a cross-modal attention mechanism to optimize the fine-grained features in the second stage, making the generated samples more conform to the semantic attributes of the real category. This method overcomes the defects of existing generative zero-shot learning methods in terms of insufficient utilization of semantic information, inter-class feature confusion, and insufficient generalization ability of generated samples, thereby improving the practical application value of the model in zero-shot learning tasks.

[0008] The technical solution of the present invention is as follows: A generative zero-shot learning method based on cross-modal attention and multi-stage generation, comprising the following steps:

[0009] Step 1: Generate coarse-grained features through the first-stage generator G1;

[0010] Step 2: Combine the coarse-grained features with the semantic attribute vector and input them into the second-stage generator G2 for semantic enhancement to generate fine-grained feature samples;

[0011] Step 3: Calculate the loss function based on the first-stage discriminator and the second-stage discriminator, and train all the generators and discriminators; the trained generators synthesize unseen class samples for training the classifier to complete the recognition of the target class.

[0012] The first-stage generator G1 consists of a fully connected layer and a LeakyReLU activation function; the noise vector is concatenated with the semantic attribute vector to obtain the vector x (1) = [z, a], which is input into the fully connected layer and undergoes non-linear mapping through the LeakyReLU activation function to obtain the coarse-grained feature

[0013] The first-stage generator G1 adopts the Wasserstein GAN framework, and its loss is shown in formula (1). The loss function is the Wasserstein loss, which promotes the alignment of the generated coarse-grained feature with the distribution of the real samples. Here, in G(z, a), it represents the sample feature generated after the first-stage generator G1 receives the concatenated input of the noise vector z and the semantic information a, and D1(·) represents the first-stage discriminator, which is used to evaluate the authenticity of the generated coarse-grained feature. denotes the mathematical expectation, that is, taking the average over the random variable, then represents the scoring expectation of the discriminator for the generated coarse-grained feature;

[0014]

[0015] The coarse-grained feature and the real feature are input into the first-stage discriminator D1 for training; the first-stage discriminator D1 also adopts the Wasserstein GAN framework; an additional fully connected layer is added to the input end of the first-stage discriminator D1 to map the high-dimensional feature of the real feature to the same dimension as the output feature of the first-stage generator.

[0016] The loss of the first-stage discriminator D1 consists of two parts, as shown in formula (2). The first part is the Wasserstein loss, and the second part is the gradient penalty term to ensure the Lipschitz condition of the Wasserstein distance. Here, represents the scoring expectation of the first-stage discriminator for the real feature. β is the weight hyperparameter of the gradient penalty term, which is used to adjust the contribution of the penalty term to the total loss. is the intermediate sample obtained by interpolating between the real sample and the generated sample, which is used to calculate the gradient. represents the gradient of the first-stage discriminator for the input ;

[0017]

[0018] The coarse-grained features and the semantic attribute vector enter the second-stage generator G2, which includes a cross-modal attention sub-layer and a feed-forward sub-layer; through the cross-modal attention sub-layer, further alignment is performed for fine-grained enhancement and semantic fusion, and the fused intermediate representation generates fine-grained features through the feed-forward sub-layer;

[0019] The cross-modal attention sub-layer includes a multi-head cross-modal attention mechanism multi-head cross-modal attention sub-layer, a residual connection, and layer normalization;

[0020] The specific steps of the cross-modal attention sub-layer are as follows;

[0021] Step 2.1: Generation of semantic tokens; the semantic attribute vector in zero-shot learning only contains one global information; use a fully connected layer to map the semantic attribute vector to a higher-dimensional space and reorganize it into a semantic token sequence A;

[0022] Step 2.2: Consider the coarse-grained features F (1) output by the first-stage generator G1 as the transpose and expansion of the Query of the multi-head cross-modal attention mechanism, denoted as

[0023] Reshape the semantic token sequence A into a matrix with a shape of [L, d attention , as shown in formula (3), and use it as Key and Value; L is the preset number of semantic tokens, and d attention is the feature dimension used in the attention mechanism to unify the representation spaces of Query, Key, and Value.

[0024]

[0025] Step 2.3: Linear mapping and head splitting. First, perform respective linear transformations on Query, Key, and Value to obtain:

[0026]

[0027] W i Q 、W i K and W i V are all learnable linear transformation matrices with a shape of [d, d attention / h], i = 1,..., h represents different attention heads, h is the number of heads, and the dimension of each head is d attention / h;

[0028] Step 2.4: Attention weight calculation; for each attention head, calculate the dot product of Query and Key, divide it by the scaling factor, and then obtain the normalized attention weights through the softmax function, as shown in formula (5):

[0029]

[0030] Step 2.5: Weighted summation; for each attention head, obtain the output of each head by performing a weighted summation on Value:

[0031] O i =W i V i (6)

[0032] Then concatenate the outputs of all attention heads and perform a final linear transformation through a learnable matrix to map back to the original dimension, forming the overall attention output F attention ;

[0033] F attention =Concat(O1,…,Oh)W O (7)

[0034] Step 2.6: Residual connection and normalization; after obtaining the outputs of each attention head through multi-head attention calculation and mapping through concatenation and linear transformation, the overall attention output is obtained; add the overall attention output and the coarse-grained features element-wise to obtain the residual output:

[0035] Fresidual1=Fattention+F (1) (8)

[0036] Perform layer normalization on the residual output to obtain the normalized feature F norm .

[0037] Specifically, step 2.1 is as follows: after passing through the fully connected layer, the semantic attribute vector is mapped into a semantic token sequence A = {a1, a2, …, a attention} with a dimension of L×d L}, where L is the preset number of tokens; a i is regarded as a "local semantic" subrepresentation, representing a fine-grained aspect of the attribute information.

[0038] The feed-forward sublayer is jointly composed of a feed-forward network FF, a residual connection, and a layer normalization module. The specific steps are as follows:

[0039] Step 2.7: Feed-forward network calculation and feature mapping; the feed-forward network FF consists of two fully connected layers and one ReLU activation:

[0040] FF(F norm )=max(0,F norm W1+b1)W2+b2 (9)

[0041] Where W1, W2 and b1, b2 are all learnable parameters, and max(0,·) represents the ReLU activation function;

[0042] Step 2.8: Residual connection and layer normalization module optimize features; add the output features of the feedforward network FF and the normalized features of the input element by element to form a residual:

[0043] Fresidual2=FF+Fnorm (10)

[0044] The residual result is layer normalized to obtain the fine-grained feature F (2) .

[0045] The fine-grained features F are obtained through the second stage generator G2 (2) , the classification loss is introduced in the second stage generator; the second stage generator G2 loss is shown in formula (11), the first part is the Wasserstein loss, and the second part is the classification loss, where λ is the weight hyperparameter of the classification loss term, which is used to balance the relative importance of Wasserstein loss and classification loss in the overall optimization objective, y represents the true category label of the training sample, and P(y|G2(z,a)) represents the predicted probability of the classifier for the category label y under the given generated sample G2(z,a);

[0046]

[0047] The fine-grained features and the real features are input into the second-stage discriminator D2 for training; the second-stage discriminator D2 also follows the Wasserstein GAN adversarial framework; the loss of the second-stage discriminator D1 consists of three parts, as shown in formula (12), the first part is the Wasserstein loss, and the second part is the gradient penalty term, which ensures the Lipschitz condition of the Wasserstein distance;

[0048]

[0049] Compared with the prior art, the beneficial effects of the present invention are as follows: The proposed multi-stage feature generation strategy enables the generated samples to be gradually optimized from coarse-grained to fine-grained, improving the recognition accuracy of unseen categories. The cross-modal attention mechanism enhances the matching degree between semantic information and visual features, making the generated samples perform better in intra-class consistency and inter-class discriminability. Experimental results show that on the standard zero-shot learning dataset, the method of the present invention can generate more discriminative features in the unseen category recognition task, improving the recognition accuracy, and effectively alleviating the category shift problem in the generalized zero-shot learning scenario. At the same time, the training of this method is more stable, the separability and generalization ability of the generated samples are stronger, providing strong support for the practical application of zero-shot learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0050] Figure 1 It is a schematic diagram of a generative zero-shot learning method based on cross-modal attention and multi-stage generation;

[0051] Figure 2 It is a schematic diagram of a semantic enhancement module based on cross-modal attention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0052] Figure 1 It is the main flowchart of the solution of the present invention. As Figure 1 shown, the generative zero-shot learning method based on cross-modal attention and multi-stage generation proposed by the present invention includes the following steps:

[0053] Step 1: Generate coarse-grained features through the first-stage generator G1;

[0054] Step 2: Combine the coarse-grained features with the semantic attribute vector and input them into the second-stage generator G2 for semantic enhancement to generate fine-grained feature samples;

[0055] Step 3: Calculate the loss function according to the first-stage discriminator and the second-stage discriminator, and train all generators and discriminators; the trained generators synthesize unseen category samples for training the classifier to complete the recognition of the target category.

[0056] Step 1.1: The first-stage generator G1 consists of a fully connected layer and a LeakyReLU activation function. First, the vector x obtained by concatenating the noise vector and the semantic attribute vector (1) =[z,a] is input into a fully connected layer (FC), and a non-linear mapping is performed with the LeakyReLU activation function to obtain the coarse-grained feature Since only simple concatenation and one-time mapping are used to integrate noise and semantic information at this time, the generated coarse-grained features mainly reflect the global structure of the category and the preliminary semantic alignment, but may still be insufficient at the local or detailed level. Through the training of this stage, G1 can learn to initially fuse noise and attributes, laying a foundation for subsequent fine-grained enhancement.

[0057] The generator G1 in the first stage adopts the Wasserstein GAN (WGAN) framework, and its loss is shown in formula (1). This loss is the Wasserstein loss, which promotes the alignment of the generated coarse-grained features with the distribution of real samples.

[0058]

[0059] Step 1.2: The generated coarse-grained features and real features are input into the first-stage discriminator D1 to train the first-stage discriminator. The first-stage discriminator D1 also adopts the Wasserstein GAN (WGAN) framework and cooperates with gradient penalty to improve the training stability. In addition, in order to make the real sample features consistent with the features output by the generator G1 after dimensionality reduction, an additional fully connected layer is added at the input end of D1 to map the high-dimensional features of real samples to the same feature dimension as the output of the generator. Its input can be real features or the coarse-grained features F generated by G1. (1) , and the output is used to distinguish between true and false and calculate the adversarial loss. The distribution difference between the two is measured by the adversarial loss.

[0060] The loss of the first-stage discriminator D1 consists of two parts, as shown in formula (2). The first part is the Wasserstein loss, and the second part is the gradient penalty term, which guarantees the Lipschitz condition of the Wasserstein distance.

[0061]

[0062] After obtaining the coarse-grained features F output in the first stage (1) , it is necessary to enter the semantic enhancement module in the second-stage generator G2 for further alignment for fine-grained enhancement and semantic fusion. That is, in step 2, semantic enhancement is performed through cross-modal attention. Here, the idea of multi-head attention in the standard Transformer is borrowed, and its residual connection and layer normalization are combined to enhance the stability and information transmission ability of the model. Its flowchart is as Figure 2 shown, and the specific steps are as follows:

[0063] Step 2.1: Generation of semantic tokens. The semantic attribute vector in zero-shot learning It only contains one piece of global information and is difficult to directly express the details of different local features in the category. Therefore, we use a fully connected layer to map the semantic attribute vector to a higher-dimensional space and reorganize it into a sequence of semantic tokens. Specifically, after passing through the fully connected layer, the semantic attribute vector is mapped to a vector A = {a1, a2, …, a attention} with a dimension of L×d L where L is the preset number of tokens. Here, each a i can be regarded as a "local semantic" subrepresentation, representing a fine-grained aspect of the attribute information. This tokenization process enables the semantic information to be passed to the subsequent attention mechanism in parallel, thereby enhancing the guidance of visual features at the local level.

[0064] Step 2.2: Preparation of data for the cross-modal attention module. Consider the coarse-grained feature F (1) output by the generator G1 in the first stage as the Query of the multi-head attention mechanism. To align with the input shape of the multi-head attention module, we transpose and expand it, denoted as so that it can be directly input into the subsequent attention calculation.

[0065] Reshape the semantic token sequence A into a matrix with a shape of [L, d attention , as shown in formula (3), and use it as the Key and Value.

[0066]

[0067] Step 2.3: Linear mapping and head splitting. For multi-head attention, we first perform separate linear transformations on the Query, Key, and Value to obtain:

[0068]

[0069] W i Q 、W i K and W i V are all learnable linear transformation matrices. Their shape is [d, dattention / h], and different attention heads use different projection matrices, so as to learn different feature patterns in parallel subspaces. Where i = 1, …, h represents different attention heads, h is the number of heads, and the dimension of each head is d attention / h.

[0070] ​Step 2.4: Attention weight calculation. For each head, calculate the dot product of Query and Key, and divide it by the scaling factor. Then obtain the normalized attention weights through the softmax function, as shown in Equation (5):

[0071]

[0072] These weights reflect the importance of each semantic token to the Query.

[0073] Step 2.5: Weighted summation. For each head, obtain the output of each head by performing a weighted sum of the Values:

[0074] O i =W i V i (6)

[0075] Then concatenate the outputs of all heads and perform a final linear transformation through a learnable matrix to map back to the original dimension, forming the overall attention output F attention .

[0076] F attention =Concat(O1,…,O h )W O (7)

[0077] Step 2.6: Residual connection and normalization. After obtaining the outputs of each head through multi-head cross-modal attention calculation and mapping through concatenation and linear transformation, the overall attention output is obtained. Next, we add the attention output and the original Query (coarse-grained feature) element-wise to obtain the residual output:

[0078] F residual1 =F attention +F (1) (8)

[0079] This operation can effectively alleviate the vanishing gradient through the residual connection, while retaining the original feature information and the new information extracted by the attention module. Subsequently, perform layer normalization on the residual output to obtain the normalized feature F norm , making the feature distribution more balanced and the training more stable.

[0080] After the multi-head cross-modal attention sublayer completes the preliminary fusion of the coarse-grained feature and semantic information, we also need to perform further non-linear transformation and feature refinement on the fused intermediate representation. For this purpose, we introduce a feed-forward sublayer (Feed-Forward Network, FFN) in the second-stage generator G2, using a two-layer fully-connected network and a non-linear activation function to enhance the feature representation ability:

[0081] Feedforward network calculation and feature mapping. The feedforward network FF consists of two fully connected layers and one ReLU activation, in the form of:

[0082] FF(F norm ) = max(0, F norm W1 + b1)W2 + b2 (9)

[0083] where W1, W2, b1, and b2 are all learnable parameters, and max(0, ·) represents the ReLU activation function. FF first maps the features to a higher hidden dimension and then maps them back to the original dimension. This process can increase the expressive power of the network while retaining the feature information. Here, we denote the result output by the feedforward network as FF.

[0084] Residual connection and layer normalization optimize features. In the standard Transformer structure, we add FF and the input element-wise to form a residual:

[0085] F residual2 = FF + F norm (10)

[0086] Subsequently, layer normalization is performed on the residual result to obtain the fine-grained feature F (2) .

[0087] Step 4.1: Through the previous steps, the fine-grained feature F (2) has been obtained through the semantic information enhancement module in the generator G2. To better train the generator, a classification loss is introduced into the generator. This loss is the supervised classification loss on the generated samples and the real samples. This loss provides explicit class feedback to the generator by guiding the features generated by the generator to be correctly recognized in the classifier, making the generated features easier to be correctly classified by the discriminator. Its loss is shown in formula (11), where the first part is the Wasserstein loss and the second part is the classification loss.

[0088]

[0089] Step 4.2: The generated fine-grained features and real features are input into the first-stage discriminator D2 to train the discriminator. The second-stage discriminator D2 also follows the Wasserstein GAN (WGAN) adversarial framework and combines gradient penalty to maintain the stability of training. In addition to the adversarial loss, D2 also introduces a classification loss into its loss function to further strengthen F (2)The inter-class discriminability. The loss of the second-stage discriminator D1 consists of three parts, as shown in Equation (12). The first part is the Wasserstein loss, the second part is the gradient penalty term to ensure the Lipschitz condition of the Wasserstein distance, and the third part is the supervised classification loss on the generated samples and the real samples to ensure that the discriminator can identify the correct categories of the generated features, thereby guiding the generator to optimize in a more accurate direction.

[0090]

[0091] Use the trained generator to synthesize unseen-class samples, convert the zero-shot learning into a supervised learning task, and train a classifier to complete the recognition of the target class.

[0092] Through the multi-stage feature generation strategy and the cross-modal attention mechanism, the present invention effectively improves the feature expression ability of generative zero-shot learning and the recognition accuracy of unseen classes. Compared with the traditional single-stage generation method, the multi-stage generation strategy of the present invention enables the generated samples to be gradually optimized at different levels, ensuring that both the global structural information of the class can be maintained and the separability of features can be enhanced in local details.

[0093] The cross-modal attention mechanism of the present invention dynamically regulates visual features using semantic information during the generation process, enabling the generated features to accurately capture category attributes and avoiding the information loss caused by the simple concatenation of semantic information and visual features in traditional methods. This mechanism can highlight key semantic attributes and suppress redundant information, thereby enhancing the class discriminability of the generated samples and making the model have stronger generalization ability for unseen classes.

Claims

1. A generative zero-shot learning method based on cross-modal attention and multi-stage generation, characterized in that, It includes the following steps: Step 1: Generate coarse-grained features through the first-stage generator G1; Step 2: Combine the coarse-grained features with the semantic attribute vector and input them into the second-stage generator G2 for semantic enhancement to generate fine-grained feature samples; Step 3: Calculate the loss function according to the first-stage discriminator and the second-stage discriminator, and train all generators and discriminators; The trained generator synthesizes unseen class samples for training the classifier to complete the recognition of the target class.

2. The generative zero-shot learning method based on cross-modal attention and multi-stage generation according to claim 1, wherein The first-stage generator G1 consists of a fully-connected layer and a LeakyReLU activation function; the noise vector is concatenated with the semantic attribute vector to obtain a vector x (1) = [z, a], which is input into the fully-connected layer and undergoes non-linear mapping through the LeakyReLU activation function to obtain the coarse-grained feature 3. The generative zero-shot learning method based on cross-modal attention and multi-stage generation according to claim 2, wherein The first-stage generator G1 adopts the Wasserstein GAN framework, and its loss is shown in formula (1). The loss function is the Wasserstein loss, which promotes the alignment of the generated coarse-grained features with the distribution of real samples. Among them, in G(z,a), it represents the sample features generated after the first-stage generator G1 receives the concatenated input of the noise vector z and the semantic information a. D1(·) represents the first-stage discriminator, which is used to evaluate the authenticity of the generated coarse-grained features. denotes the mathematical expectation, that is, taking the average over the random variable, then denotes the expected score of the discriminator for the generated coarse-grained features; 4. The generative zero-shot learning method based on cross-modal attention and multi-stage generation according to claim 3, wherein The coarse-grained features and real features are input into the first-stage discriminator D1 for training; the first-stage discriminator D1 also adopts the Wasserstein GAN framework; an additional fully connected layer is added to the input end of the first-stage discriminator D1 to map the high-dimensional features of the real features to the same dimension as the output features of the first-stage generator; The loss of the first-stage discriminator D1 consists of two parts, as shown in Equation (2). The first part is the Wasserstein loss, and the second part is the gradient penalty term, which ensures the Lipschitz condition of the Wasserstein distance. Among them represents the expected score of the first-stage discriminator for the real features. β is the weight hyperparameter of the gradient penalty term, which is used to adjust the contribution of the penalty term to the total loss. is the intermediate sample interpolated between the real sample and the generated sample, which is used to calculate the gradient. represents the gradient of the first-stage discriminator for the input ; 5. The generative zero-shot learning method based on cross-modal attention and multi-stage generation according to claim 1, characterized in that The coarse-grained features and the semantic attribute vector enter the second-stage generator G2, which includes a cross-modal attention sub-layer and a feed-forward sub-layer; First, further alignment is performed through the cross-modal attention sub-layer for fine-grained enhancement and semantic fusion; the fused intermediate representation generates fine-grained features through the feed-forward sub-layer; The cross-modal attention sub-layer includes a multi-head cross-modal attention mechanism, a residual connection, and layer normalization; The specific steps of the cross-modal attention sub-layer are as follows; Step 2.1: Generation of semantic tokens; semantic attribute vectors in zero-shot learning only contain one piece of global information; use a fully connected layer to map the semantic attribute vectors to a higher-dimensional space and reorganize them into a sequence of semantic tokens A; Step 2.2: Treat the coarse-grained features F output by the first-stage generator G1 (1) as the transposed and extended Query of the multi-head cross-modal attention mechanism, denoted as Reshape the semantic token sequence A into a matrix with a shape of [L, d attention , as shown in formula (3), and use it as Key and Value; L is the preset number of semantic tokens, d attention is the feature dimension used in the attention mechanism, which is used to unify the representation spaces of Query, Key, and Value; Step 2.3: Linear mapping and head splitting. First, perform respective linear transformations on Query, Key, and Value to obtain: W i Q , W i K and W i V are all learnable linear transformation matrices, with shapes of [d, d attention / h], where i = 1, …, h represent different attention heads, h is the number of heads, and the dimension of each head is d attention / h; Step 2.4: Calculate the attention weights; for each attention head, calculate the dot product of Query and Key, divide by the scaling factor, and then obtain the normalized attention weights through the softmax function, as shown in formula (5): Step 2.5: Weighted summation; for each attention head, obtain the output of each head by performing weighted summation on Value: O i = W i V i (6) Then, the outputs of all attention heads are concatenated and passed through a learnable matrix for a final linear transformation to map back to the original dimension, forming the overall attention output F attention ; F attention = Concat(O1, …, Oh)W O (7) Step 2.6: Residual connection and normalization; after the multi-head attention calculates the outputs of each attention head and maps them through concatenation and linear transformation, obtain the overall attention output; add the overall attention output and the coarse-grained features element by element to obtain the residual output: Perform layer normalization on the residual output to obtain the normalized feature F norm .

6. The generative zero-shot learning method based on cross-modal attention and multi-stage generation according to claim 5, wherein Specifically, the step 2.1 is as follows: after passing through the fully-connected layer, the semantic attribute vector is mapped into a semantic token sequence with a dimension of L×d attention where L is the preset number of tokens; a is regarded as a "local semantics" subrepresentation, representing a fine-grained aspect of the attribute information. i ​ 7. The generative zero-shot learning method based on cross-modal attention and multi-stage generation according to claim 1, wherein The feed-forward sub-layer is jointly composed of a feed-forward network FF, a residual connection, and a layer normalization module, and the specific steps are as follows: Step 2.7: Feed-forward network calculation and feature mapping; the feed-forward network FF consists of two fully connected layers and one ReLU activation: FF(F norm ) = max(0, F norm W1 + b1)W2 + b2 (9) where W1, W2, b1, and b2 are all learnable parameters, and max(0,·) represents the ReLU activation function; Step 2.8: The residual connection and the layer normalization module optimize the features; Adopt adding the output features of the feed-forward network FF and the normalized input features element by element to form a residual: Perform layer normalization on the residual results to obtain fine-grained feature F (2) .

8. The generative zero-shot learning method based on cross-modal attention and multi-stage generation according to claim 7, wherein The fine-grained feature F is obtained through the second-stage generator G2 (2) , and a classification loss is introduced into the second-stage generator; The loss of the second-stage generator G2 is shown in Equation (11). The first part is the Wasserstein loss, and the second part is the classification loss. Here, λ is the weight hyperparameter of the classification loss term, which is used to balance the relative importance of the Wasserstein loss and the classification loss in the overall optimization objective. y represents the true class label of the training sample, and P(y∣G2(z,a)) represents the predicted probability of the classifier for the class label y under the condition of the given generated sample G2(z,a).

9. The generative zero-shot learning method based on cross-modal attention and multi-stage generation according to claim 8, characterized in that The fine-grained features and the true features are input into the second-stage discriminator D2 for training. The second-stage discriminator D2 also follows the Wasserstein GAN adversarial framework. The loss of the second-stage discriminator D1 consists of three parts, as shown in Equation (12). The first part is the Wasserstein loss, and the second part is the gradient penalty term, which ensures the Lipschitz condition of the Wasserstein distance.