Image Generation Method Using Generative Adversarial Network Based on Mixture of Experts Model

Generative adversarial network based on hybrid expert models generates high-quality image data, which solves the problems of insufficient image data and unbalanced labels, and significantly improves the image classification accuracy and generalization ability of the model.

CN118968258BActive Publication Date: 2025-05-06JIANGNAN UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411423496.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-12
Publication Date
2025-05-06
Estimated Expiration
2044-10-12

AI Technical Summary

Technical Problem

Insufficient image data volume and unbalanced label classification result in low image classification accuracy. Especially when processing weather images, data diversity and quantity are limited, and the category distribution is uneven, making it difficult to generate consistent and correct labels.

Method used

A generative adversarial network based on hybrid expert models is adopted. By building multiple expert networks and a gated network, the contribution of the expert network is dynamically selected and adjusted to generate high-quality and diverse image data, thereby expanding the data set and improving the generalization ability of the model.

Benefits of technology

By generating high-quality image data, the image classification accuracy is significantly improved, the generalization ability and generation efficiency of the model are enhanced, and complex data distribution can be better handled.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118968258B_ABST
    Figure CN118968258B_ABST
Patent Text Reader

Abstract

The present invention discloses an image generation method using a generative adversarial network based on a hybrid expert model, and belongs to the field of image classification technology. The method achieves more refined control of the generation process and more efficient parameter utilization by designing a generative adversarial network based on a hybrid expert model. Thereby improving the quality and diversity of the generated images, and enhancing the model's learning ability for complex data distribution. Furthermore, the present application method also designs a multimodal data fusion network, which can generate new images based on information of data in different modalities such as images, texts, and audios; in addition, the present invention also introduces a new loss function, which not only considers the similarity between the generated image and the real image, but also considers the collaboration and balance between expert networks, further optimizes the training process of the generative adversarial network, thereby improving the image generation quality and the image classification accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to an image generation method using a generative adversarial network based on a hybrid expert model, and belongs to the technical field of image classification. Background Art

[0002] With the development of deep learning, image classification technology has been applied to more and more fields, such as unmanned driving, industrial inspection, medical image analysis, security monitoring, etc. Classification accuracy is an important indicator that reflects the ability of a classification model to correctly classify samples on a given task. Under the premise that the structure of the classification model is determined, the better the quality of the data used for training, the higher the classification accuracy. However, using manual labeling to obtain sufficient and balanced data sets is always expensive, and it may often encounter situations where the number of images is small, the quality is poor, and the categories are unbalanced. For example, weather images are a typical case of insufficient data and uneven distribution of label categories. On the one hand, the diversity and number of the obtained training data sets are limited; on the other hand, in the training set, some categories of weather images are relatively rare compared to other categories, and some images (such as rainy and snowy days) are more subtle. Based on these two reasons, it is difficult to give consistent and correct labels or classifications, which brings great difficulties to image recognition and classification tasks; in addition, the quality and quantity of image data also greatly affect the generalization ability of deep learning models.

[0003] In order to solve the above problems, image generation technology can be used to generate new images based on existing images. An effective image generation method can generate new samples based on existing data to expand the data set and increase data quality. It can overcome the problems of insufficient data and unbalanced label distribution, thereby improving the generalization ability of the model and increasing the diversity of samples; at the same time, using the generated images in the classification model can also improve the classification accuracy.

[0004] Image generation technology usually uses neural networks to generate new images. Existing image generation methods are divided into two categories. One is based on the unsupervised representation learning model VQ-VAE. This model combines vector quantization and variational autoencoders. Through the vector quantization step, the continuous potential representation is converted into a discrete code book index, which reduces the reconstruction error and improves the image generation quality. However, its training process involves additional vector quantization steps and code book learning. Especially when dealing with large-scale data sets, it takes a long time to train for optimization, and the model complexity and computational overhead are large. The other is based on the generative adversarial network. For example, the Chinese patent with publication number CN112598034A discloses a method for generating ore images based on a generative adversarial network and a computer-readable storage medium. The inventor added conditional label information to the original GAN ​​generator structure. The introduction of conditional variables provides additional information for the generator and discriminator, and can generate samples that meet specific conditions or labels, thereby achieving the effect of expanding the data set. In the above method, although the conditional generative adversarial network (CGAN) controls the generation process by introducing conditional variables, the generated images sometimes still have problems such as blurred edges and low resolution. Moreover, if the conditional variables are improperly selected or designed, the quality of the generated results may be reduced or the expected effect may not be achieved. The Chinese patent with publication number CN117252939A discloses a method for generating faces by a deep separable convolutional generative adversarial neural network. The inventors added a deep separable convolutional module to the deep convolutional generative adversarial network (DCGAN) to generate images more efficiently. Although this method can effectively reduce the amount of calculation and the number of parameters of the model, it is accompanied by an increase in the overall complexity of the model, which in turn causes the training time to be extended. This phenomenon is more significant when processing large-scale data sets. The Chinese patent with publication number CN113139916A discloses a method for underwater sonar simulation image generation and data expansion based on a generative adversarial network. The inventor constructed a gradient penalty term in the generative adversarial network to solve the problem of poor image quality caused by gradient vanishing and gradient sudden change during training, and improved image resolution by adding a sampling convolution layer to the network. However, since this model contains multiple hyperparameters (such as gradient penalty coefficient, learning rate, etc.), these hyperparameters need to be found through experiments to find the optimal value, which increases the complexity and difficulty of model parameter adjustment. Summary of the invention

[0005] In order to solve the problem of low image classification accuracy caused by insufficient image data and unbalanced label classification, the present invention provides an image generation method using a generative adversarial network based on a hybrid expert model, which generates high-quality image data to train a classification model, thereby improving image classification accuracy. The method constructs a new deep learning model by designing a generative adversarial network based on a hybrid expert model. This model uses multiple expert networks to generate different parts or features of the image, and dynamically selects the most appropriate expert output through a gating network, thereby generating high-quality and diverse images.

[0006] A method for generating an image using a generative adversarial network based on a hybrid expert model, comprising:

[0007] 1) Constructing a Mixed Expert Model (MoE): This model consists of multiple expert networks and a gating network. Each expert network is an independent convolutional neural network used to generate different parts or features of the image. The gating network is a lightweight neural network structure used to dynamically select the most appropriate expert generator output;

[0008] 2) Integrate MoE into GAN: Use the MoE model as the generator (G) of GAN. During the generation process, the input noise (or conditional information) first passes through the gating network, and dynamically selects and weights different expert networks according to the input features. Each expert network independently generates partial outputs according to the weights assigned to it by the gating network and the input noise received. These partial outputs are weighted and averaged according to their corresponding weights to form the final generated samples;

[0009] 3) Optimize the training process: By alternately training the generator and the discriminator, the parameters of the gated network and the expert network of the MoE model are optimized at the same time. During the training process, the gradient descent optimization algorithm is used so that the generator can gradually approach the real data distribution, while the discriminator can continuously improve its discrimination ability.

[0010] A hybrid expert model generative adversarial network image generation method comprises the following steps:

[0011] S1. Construct a training data sample set;

[0012] S2. Preprocess the sample data;

[0013] S3. Build a hybrid expert model: Design multiple expert networks and a gating mechanism network. Each expert network is responsible for processing different types of features or tasks, and the gating network dynamically selects and adjusts the contribution of each expert network based on the input features or noise;

[0014] S4. Build a generative adversarial network: Design a generator network that is integrated with the hybrid expert model and design a discriminator network at the same time;

[0015] S5. alternately train the generative network and the discriminative network, and during the training process, adjust the weights of the expert network and the parameters of the gating network according to the quality of the generated images and the feedback of the discriminator until a convergence condition or a preset number of iterations is reached;

[0016] S6. After training is complete, use the generator part of the generative adversarial network of the mixture of experts model to generate new images.

[0017] In one implementation, the loss function of the generator network model is:

[0018]

[0019] in, is the total loss function of the generator, is the adversarial loss term, is the expert network loss term, is the gating network loss term, is the expert balance loss term, is the corresponding weight coefficient, which is used to balance the importance of different loss terms.

[0020]

[0021]

[0022]

[0023]

[0024] Adversarial Loss Term middle, is a generator, is the discriminator, is the prior noise distribution The noise vector sampled in ;

[0025] Expert network loss middle, is the number of expert networks, It is Expert network, It is The loss function of the expert network is It is The weight coefficient corresponding to each expert network;

[0026] Gating Network Loss middle, The gating network is The weight coefficients assigned by the expert network;

[0027] Expert Balance Loss Item middle, is the weight coefficient used to balance each expert network, is the number of expert networks, represents the L2 norm.

[0028] In one implementation, the discriminator network model includes an input layer, three convolution blocks, a first Dropout layer, a flattening layer, a first fully connected layer, a Leaky ReLu activation layer, a second Dropout layer, a second fully connected layer, a Sigmoid activation layer and an output layer, and each convolution block consists of a convolution layer, a Leaky ReLu activation layer and a normalization layer.

[0029] In one implementation, the loss function of the discriminator network model is:

[0030]

[0031] in, is the total loss function of the discriminator, is the real sample loss term, is the sample generation loss term.

[0032]

[0033]

[0034] In the real sample loss term, It is real data. represents the probability distribution of real data, is the output of the discriminator for the real sample;

[0035] In the generated sample loss term, is the prior noise distribution The noise vector sampled in ; is the generator according to the noise The generated image, is the output of the discriminator for the generated image.

[0036] In one implementation, if the data sample is multimodal data, step 2 further includes fusing the multimodal data, and the fusing method includes:

[0037] Step S1, performing data cleaning on multimodal data, wherein the multimodal data includes image data, text data, and audio data;

[0038] Step S2, segmenting, encoding, embedding the text data, converting it into a numerical representation, and sampling, filtering, and feature extraction of the audio data;

[0039] Step S3, constructing a multimodal data fusion network to generate multimodal fusion data.

[0040] In one implementation, the multimodal data fusion network includes a lightweight convolutional neural network for extracting image features and two lightweight recurrent neural networks for extracting text and audio features respectively; the lightweight convolutional neural network for extracting image features includes, in sequence, an original image input layer, a convolutional layer, a deep convolutional layer, a global average pooling layer, and a fully connected output layer, wherein the deep convolutional layer includes a point-by-point convolutional layer, a normalization layer, and a ReLU6 activation function; the lightweight recurrent neural network for extracting text features includes, in sequence, an embedding layer, a RNN layer, a bidirectional RNN layer, a merging layer, a first fully connected layer, a second fully connected layer, and an output layer; the lightweight recurrent neural network for extracting audio features includes, in sequence, a convolutional layer, a pooling layer, an RNN layer, a first fully connected layer, a second fully connected layer, and an output layer.

[0041] The second object of the present invention is to provide an image classification method, which uses the above method to generate new image data to expand the training data set.

[0042] The beneficial effects of the present invention are:

[0043] By designing a generative adversarial network based on a hybrid expert model, more refined control of the generation process and more efficient parameter utilization are achieved. In the method of the present invention, the hybrid expert model is composed of multiple expert networks, each of which is responsible for processing different parts or features of the input data, and dynamically adjusts the degree of participation of each expert through a gating mechanism to ensure that the generated data is not only of high quality, but also can make full use of the expertise of each expert network, thereby improving the overall generation efficiency and the generation ability of the model. This integration not only improves the quality and diversity of the generated images, but also enhances the model's learning ability for complex data distribution. In addition, the present invention also introduces a new loss function, which not only takes into account the similarity between the generated image and the real image, but also takes into account the collaboration and balance between the expert networks, further optimizes the training process of the generative adversarial network, thereby improving the image generation quality and image classification accuracy. Furthermore, the method of the present application also designs a multimodal data fusion network, which can generate new images based on information from data of different modalities such as images, text, and audio. Moreover, the combination of multimodal data fusion and a mixed expert model (MoE) can improve the robustness of the model to missing or noisy data, because the architecture can use complementary information between different modalities to compensate each other, give priority to higher quality data by dynamically adjusting the weights of the expert network, and enhance the generalization ability and error tolerance of the model with the help of redundant representation and cross-modal verification, so that accurate predictions and decisions can still be made when faced with incomplete or inaccurate data. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0045] Figure 1 It is a flow chart of an image generation method using a generative adversarial network based on a hybrid expert model provided by the present invention.

[0046] Figure 2 It is a schematic diagram of the structure of a generative adversarial network based on a hybrid expert model in the method of the present invention.

[0047] Figure 3 It is a schematic diagram of the generator network structure in the method of the present invention.

[0048] Figure 4 It is a schematic diagram of the discriminator network structure in the method of the present invention.

[0049] Figure 5It is a partial weather image generated by using the image generation method based on the hybrid expert model generative adversarial network according to the present invention.

[0050] Figure 6 It is a flow chart of an image generation method using a generative adversarial network based on a multimodal hybrid expert model provided by the present invention.

[0051] Figure 7 It is a schematic diagram of the multimodal feature fusion network structure based on the multimodal hybrid expert model in the method of the present invention.

[0052] Fig. 8A It is a schematic diagram of the structure of a lightweight convolutional neural network used to extract image features in the method of the present invention;

[0053] Figure 8B It is a schematic diagram of the structure of a lightweight recurrent neural network used to extract text features in the method of the present invention;

[0054] Figure 8C It is a schematic diagram of the lightweight recurrent neural network structure used to extract audio features in the method of the present invention.

[0055] Fig. 9 It is a schematic diagram of the structure of a generative adversarial network based on a multimodal hybrid expert model in the method of the present invention.

[0056] Fig.10 This is a partial image generated by the image generation method based on the multimodal hybrid expert model of the present invention. DETAILED DESCRIPTION

[0057] In order to make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be further described in detail below with reference to the accompanying drawings.

[0058] Embodiment 1:

[0059] This embodiment provides an image generation method using a generative adversarial network based on a hybrid expert model. Figure 1 , the method comprising:

[0060] Step 101, constructing a training data sample set;

[0061] The data sample set must ensure the diversity and quantity of samples. The data set used in this embodiment is the public data set ImageNet or CIFAR-10, wherein the ImageNet data set contains more than 15 million images from more than 20,000 different categories of real-world objects. Each category has hundreds to thousands of images, with high diversity and complexity. The CIFAR-10 data set is a small data set that is closer to universal objects. A total of 10 categories of RGB color pictures are included: airplane, automobile, bird, cat, deer, dog, frog, horse, ship and truck. The size of each picture in the CIFAR-10 data set is 32x32, each category has 6,000 images, and there are a total of 50,000 training pictures and 10,000 test pictures in the data set.

[0062] Step 102, preprocessing the sample data;

[0063] Data cleaning needs to be performed to remove noise and outliers, feature encoding to convert categorical data, normalization to adjust the data scale, and the dataset is partitioned to separate training and testing data.

[0064] Specifically, first calculate the Z-scores of all data categories (i.e., the difference between the data point and the mean divided by the standard deviation), and treat data categories with Z-scores greater than 3 as noise and outliers; further, use one-hot encoding to create a new binary column for each data category, where one category value is encoded as 1 and the rest are 0, converting categorical data into numerical data that can be processed by machine learning algorithms; finally, use the normalization formula to scale all feature data to the range of [0,1] to facilitate model training and algorithm convergence. The normalization formula is as follows:

[0065]

[0066] in x is the original data, min(x) and max(x) are the minimum and maximum values ​​in the data set respectively. x’ is the normalized data.

[0067] Step 103, constructing a generator network model in a generative adversarial network based on a hybrid expert model;

[0068] like Figure 3As shown in the figure, a hybrid expert model is integrated in the generator. The hybrid expert model consists of multiple expert networks, each of which focuses on different aspects or features of generated data. The weight and contribution of each expert is dynamically adjusted through a gating network, and the outputs of each expert network are fused according to the weight of the gating network to achieve more diverse and realistic data generation.

[0069] Figure 3 In the figure, the hybrid expert model is composed of N expert networks. The outputs of the N expert networks are combined with the weights dynamically determined by the gating network to finally obtain the weighted outputs of each expert network. The weighted outputs of each expert network are fused and then passed through the conversion layer, transposed convolution layer and Tanh activation layer to finally obtain the output of the generator; the conversion layer includes convolution layer, normalization layer, ReLU activation layer and upsampling layer.

[0070] The loss function of the traditional generative adversarial network generator can be expressed as:

[0071]

[0072] in, is the loss of the generative model, is the prior noise distribution, z is the sampled noise vector, is fake data generated by the generative model. It is the judgment result of the discriminant model on the generated data (that is, the probability of believing that the data is true), and E represents the expectation, which represents the average value of all samples under a specific distribution.

[0073] However, for the hybrid expert model constructed in the solution of the present invention, the loss function of the generator network model is a composite loss function that comprehensively considers multiple factors, aiming to optimize the quality and diversity of generated data. Specifically, the loss function also includes the following key parts:

[0074] 1. Expert network loss term: For each expert network responsible for generating specific aspects of the data, such as the edges, textures, or other visual features of the image, an expert loss term is added to the solution of the present invention. This loss ensures that each expert network focuses on specific features of the generated data and is able to improve the quality of relevant aspects of the generated data.

[0075]

[0076] in, is a generator, is the discriminator, is the prior noise distribution The noise vector sampled in , is the number of expert networks, It is Expert network, It is The loss function of the expert network is is the corresponding weight coefficient;

[0077] 2. Gated network loss term: The weights of the gated network reflect the contribution of each expert network to the final output. The design of the gated loss aims to optimize these weights to generate data that is both realistic and diverse. In order to prevent the gated network from always activating the same experts, but to dynamically select experts based on the different characteristics of the input data, the gated network loss term is added to the solution of the present invention. This helps promote the flexibility and adaptability of the model. In this way, the gating network learns how to more efficiently route input data to the most appropriate expert network, thereby improving the quality and diversity of generated data.

[0078]

[0079] in, The gating network is The weight coefficients assigned by experts;

[0080] 3. Expert balance loss term: In order to ensure that all experts can receive balanced training, a balance loss may be introduced. The additional term This encourages the model to distribute input data evenly to all experts, avoiding overloading some experts while leaving others idle.

[0081]

[0082] in, is to encourage the gating network to use the loss terms of each expert in a balanced way, is the weight of the balance loss term. In this embodiment, after multiple experiments, , The gating network is The weight coefficients assigned by experts are is the number of expert networks, represents the L2 norm.

[0083] Therefore, the final generator loss function will be a combination of the traditional adversarial loss and the specially designed expert network loss, gated network loss, and expert balance loss mentioned above to ensure that the generator can not only generate realistic data, but also effectively utilize each component of the hybrid expert model, maintain the stability and efficiency of the model, while avoiding overfitting and ensuring that all expert networks are trained in a balanced manner. Such a loss function helps the generator to fully utilize the advantages of the hybrid expert architecture while maintaining high-quality output.

[0084]

[0085] in, is the corresponding weight coefficient, which is used to balance the importance of different loss items. In this embodiment, after multiple experiments, .

[0086] Step 104, constructing a discriminator network model in a generative adversarial network of a hybrid expert model;

[0087] As attached Figure 4 As shown in the figure, in the discriminator network structure, an image normalization layer is used after each convolution layer for normalization to prevent gradient diffusion, and all activation functions are set to Leaky ReLu functions to ensure that there is a gradient change when x < 0, and two additional Dropout layers are added to regularize the model and reduce overfitting.

[0088] like Figure 4 As shown, the discriminator network includes an input layer, three convolution blocks, a first Dropout layer, a flattening layer, a first fully connected layer, a Leaky ReLu activation layer, a second Dropout layer, a second fully connected layer, a Sigmoid activation layer and an output layer, and each convolution block consists of a convolution layer, a Leaky ReLu activation layer and a normalization layer.

[0089] The loss function of the discriminator D is:

[0090]

[0091] in, is the total loss function of the discriminator, is the real sample loss term, is the sample generation loss term.

[0092] 1. Real sample loss term: This part of the loss function is designed to enable the discriminator to correctly identify real images. For real images, the discriminator should output a probability close to 1, so the loss function can be expressed as:

[0093]

[0094] in, It is real data. represents the probability distribution of real data, is the output of the discriminator for the real sample.

[0095] 2. Generated sample loss term: The loss function of this part is designed to enable the discriminator to correctly identify the fake images generated by the generator. For the generated images, the discriminator should output a probability close to 0, so the loss function can be expressed as:

[0096]

[0097] in, is the prior noise distribution of the generator input, i.e., the distribution of the latent variables, is a random noise vector sampled from the latent variable space, is the generator according to the noise The generated image, is the output of the discriminator for the generated image.

[0098] Step 105, construct a training model.

[0099] Alternate optimization in actual training and The training process of the entire network can be regarded as an iterative optimization process of two loss functions. In each iteration, the generator G is first fixed and optimized. Update the parameters of the discriminator D, and then fix the discriminator D to optimize Update the parameters of the generator G. In the process of training the generative model, the expert layer and gating mechanism in the hybrid expert network model are trained simultaneously to better capture the complex characteristics of the data.

[0100] Step 106, after the training is completed, the generator part in the generative adversarial network of the hybrid expert model is used to generate new images.

[0101] Figure 2 This is a schematic diagram of generative adversarial network training based on a hybrid expert model, such as Figure 2 As shown, the hybrid expert generative adversarial network inputs noise z to the generator, where the generator network structure includes N expert networks E1, E2, ..., E N Each expert network output is fed into the gating network, which assigns weights M1, M2, ..., M to each expert. i The output of the gating network is fed into the data fusion step to form the generated data. The generated data is fed into the discriminator D for real / fake judgment. The output of the discriminator D is used to calculate the discriminator loss and the generator loss The loss function updates the parameters of the discriminator D and the generator G through back propagation. The training process is repeated until the network reaches the predetermined training target in terms of generation quality and discrimination ability.

[0102] The generated images are used as data enhancement methods, some of which are as follows Figure 5As shown in the figure, it can be seen that the generated image has high clarity, and when the rainy and snowy scenes are presented at the same time, the boundary between the two is also very clear. Figure 5 Two snowy and two rainy scene pictures are provided in the figure. It can be seen that the boundary between the rainy and snowy scenes is very clear. This high degree of distinction has a significant promoting effect on improving the accuracy of the classification model, thereby achieving a significant improvement in model performance. This embodiment uses VGG16 as the classification model, and trains it using the original data in the ImageNet and CIFAR-10 datasets and the data after adding the pictures generated by the generative adversarial network of the above-mentioned hybrid expert model. The corresponding training result accuracy is shown in Table 1 below:

[0103] Table 1

[0104]

[0105] As can be seen from Table 1, the accuracy is significantly improved after adding the above-mentioned generative adversarial network. For the types that are missing in the unevenly distributed data set, the generative adversarial network of the hybrid expert model is used to specifically generate minority image types, so as to achieve the purpose of expanding the data set and greatly improve the image classification accuracy.

[0106] By integrating the hybrid expert model into the generative adversarial network to generate images, the GAN generator can handle more complex data distributions, generate more diverse and realistic samples, and improve the generation ability; the parallel computing characteristics of the MoE model make the GAN training process more efficient and reduce the computing cost. It reduces the number of parameters that the model needs to learn, alleviates the overfitting problem of the model, greatly speeds up the training speed, and improves the training efficiency; the expert networks of the MoE model each focus on processing a specific subset of data, thereby improving the generalization ability of the overall model.

[0107] Embodiment 2:

[0108] This embodiment provides an image generation method using a generative adversarial network with a multimodal data fusion network and a hybrid expert model. Figure 6 The combination of multimodal data fusion and the mixture of experts (MoE) model can improve the robustness of the model to missing or noisy data, because the architecture can use the complementary information between different modalities to compensate each other, prioritize high-quality data by dynamically adjusting the weights of the expert network, and enhance the generalization ability and error tolerance of the model with the help of redundant representation and cross-modal verification, so that accurate predictions and decisions can still be made when facing incomplete or inaccurate data. The method includes:

[0109] Step 101, collecting multimodal data samples;

[0110] Collect data from different modalities, such as images, text, audio, etc., and ensure the diversity and representativeness of the data.

[0111] The image data sample set must ensure the diversity and quantity of samples, so the public datasets ImageNet or CIFAR-10 are still used.

[0112] The text data sample set needs to contain data of text and image pairs. The dataset used in this embodiment is the public dataset WIT (Wikipedia-based Image Text). WIT is a large multimodal and multilingual dataset consisting of 37.6 million rich image-text examples, covering more than 100 languages, with at least 12,000 examples for each language, and provides cross-language texts for many images.

[0113] The audio dataset needs to have sufficient resolution and clarity. The dataset used in this embodiment is the A2IR (Audio-to-Image Representation) dataset, which is a dataset for synthetic audio detection. It includes five audio-to-image representations of natural and synthetic audio, such as spectrogram, histogram, scatter plot, bispectral phase map, and bispectral amplitude map.

[0114] Step 102, preprocessing each modal sample data;

[0115] Image data still requires data cleaning to remove noise and outliers, feature encoding to convert categorical data, and normalization to adjust the data scale.

[0116] Specifically, first calculate the Z-scores of all data categories (i.e., the difference between the data point and the mean divided by the standard deviation), and treat data categories with Z-scores greater than 3 as noise and outliers; further, use one-hot encoding to create a new binary column for each data category, where one category value is encoded as 1 and the rest are 0, converting categorical data into numerical data that can be processed by machine learning algorithms; finally, use the normalization formula to scale all feature data to the range of [0,1] to facilitate model training and algorithm convergence. The normalization formula is as follows:

[0117]

[0118] in x is the original data, min(x) and max(x) are the minimum and maximum values ​​in the data set respectively. x’ is the normalized data.

[0119] Text data needs to be segmented, encoded, embedded, and converted into numerical representation.

[0120] Specifically, first, word segmentation is performed to split continuous text strings into meaningful words or phrase units; secondly, the segmented words are encoded to convert the words into digital sequences, which are the positions of the words in the vocabulary; finally, the pre-trained word embedding Word2Vec model is used to convert these digital indexes into vectors in a high-dimensional space.

[0121] Audio data needs to be sampled, filtered, and feature extracted to convert the data into a format suitable for machine learning model processing, such as spectrogram, cepstrum, or feature vector, for subsequent data analysis.

[0122] Specifically, the audio data signal is first sampled and the analog signal is converted into a digital signal to ensure accurate capture of the signal and avoid aliasing; secondly, unnecessary frequency components such as noise or interference are removed through a filter to improve the clarity of the signal; finally, feature extraction is performed.

[0123] Step 103, constructing a multimodal data fusion network to generate multimodal fusion data;

[0124] For data of different modalities (image, text, audio), this embodiment first designs a lightweight convolutional neural network (CNN) for extracting image features, and two lightweight recurrent neural networks (RNN) for extracting text and audio features. Secondly, a module based on the self-attention mechanism is introduced to learn the interaction and importance weights between features of different modalities. Next, combined with the output of the dynamic attention module, features from different modalities are integrated through a weighted fusion strategy. Finally, a multimodal interaction layer is designed to simulate the interaction and information transfer between different modalities. The network structure is as follows: Figure 7 shown.

[0125] The lightweight CNN used to extract image features includes the original image input layer, convolution layer, deep convolution layer, global average pooling layer, and fully connected output layer. The deep convolution layer includes a point-by-point convolution layer, a normalization layer, and a ReLU6 activation function. Fig. 8A As shown in Figure 1, the lightweight RNN used to extract text features includes an embedding layer, an RNN layer, a bidirectional RNN layer, a merging layer, a first fully connected layer, a second fully connected layer, and an output layer. Figure 8B As shown in the figure, the lightweight RNN used to extract audio features includes a convolution layer, a pooling layer, an RNN layer, a first fully connected layer, a second fully connected layer, and an output layer. Figure 8C shown.

[0126] Step 104, constructing a generator network model in a generative adversarial network based on a multimodal hybrid expert model;

[0127] like Fig. 9 As shown in the figure, a multimodal data fusion network and a hybrid expert model are integrated in the generator network. The input multimodal data is first preprocessed and then input into the multimodal data fusion network, which is responsible for integrating the features of different modalities into a coherent and rich data representation. Further, the fused data representation is fed into the hybrid expert model. The model consists of multiple expert networks, each of which focuses on processing specific aspects of the data, thereby achieving refined modeling of different features of the data. Through a gating network, the activation weights of each expert network are dynamically adjusted to ensure that the generator can selectively activate the most relevant expert network according to the specific characteristics of the data. Finally, the gating network performs weighted fusion of the outputs of each expert network, and the generator network produces highly realistic data output.

[0128] Fig. 9 In the paper, the multimodal data fusion network consists of three different neural networks, whose outputs are then fed into a fusion module, which includes a self-attention-based module for learning the interactions and importance weights between features of different modalities. Through the output of the dynamic attention module, the network adopts a weighted fusion strategy to integrate features from different modalities to form a unified multimodal feature representation. Finally, the multimodal interaction layer is designed to simulate the interaction and information transfer between different modalities, enhancing the model's ability to understand and process multimodal data.

[0129] Fig. 9 In the figure, the hybrid expert model is composed of N expert networks. The outputs of the N expert networks are combined with the weights dynamically determined by the gating network to finally obtain the weighted outputs of each expert network. The weighted outputs of each expert network are fused and then passed through the conversion layer, transposed convolution layer and Tanh activation layer to finally obtain the output of the generator; the conversion layer includes convolution layer, normalization layer, ReLU activation layer and upsampling layer.

[0130] The loss function of the traditional generative adversarial network generator can be expressed as:

[0131]

[0132] in, is the loss of the generative model, is the noise sampled from the latent space, is the prior noise distribution, is fake data generated by the generative model. It is the judgment result of the discriminant model on the generated data (that is, the probability of believing that the data is true), and E represents the expectation, which represents the average value of all samples under a specific distribution.

[0133] However, for the hybrid expert model constructed in the present invention, in order to ensure that all expert networks are fully utilized and avoid the situation where some experts are overloaded while others are idle, we introduce the expert load balancing loss This loss function is designed to achieve balanced load distribution by penalizing expert networks whose activation frequencies deviate greatly from the average activation frequency.

[0134]

[0135] in, is assigned to The token ratio (the smallest basic unit of different parts of the input data) of each expert, is the coefficient of variation, which is used to measure the degree of dispersion of a set of data.

[0136] Further,

[0137]

[0138] here, is the number of tokens in the input data, is the number of experts, Is a gated network for each token t The probability distribution of the output, It is Experts were assigned to The probability of a token.

[0139] Furthermore, the proportion of tokens assigned to the i-th expert is usually determined by a gating network. The function of the gating network is to assign a probability distribution to each input token, which indicates which expert the token should be sent to.

[0140] In order to ensure that the generated images are consistent across different modalities, this embodiment introduces a multimodal consistency loss The purpose of this loss function is to ensure that the generated images are not only visually realistic, but also semantically consistent with their corresponding text descriptions and audio content. The multimodal consistency loss is implemented by comparing the similarities between representations of different modalities, and the structural similarity index (SSIM) is used as the metric of similarity:

[0141]

[0142] Among them, the structural similarity index (SSIM) is an indicator to measure the similarity between two images, which takes into account three key features of image brightness, contrast and structure. The value range of SSIM is between -1 and 1. When the two images are exactly the same, the value of SSIM is 1; when the two images are completely different, the value of SSIM is close to -1.

[0143] In order to ensure that the generator can retain the key information and features of the input data when generating images, in this embodiment, the reconstruction loss is introduced The purpose of this loss function is to minimize the difference between the generated image and the original input image, ensuring that the generated image remains highly similar to the original image visually. Here, pixel-level loss (L2 loss) is used:

[0144]

[0145] in, It is real data. is fake data generated by the generative model. represents the L2 norm.

[0146] Therefore, the final generator loss function is It is composed of multiple different loss terms, each of which targets different aspects of model training. The purpose of the total loss function is to optimize the performance of the generator network to generate high-quality, multi-modal consistent images that are consistent with the characteristics of the input data. The total loss function can be expressed as:

[0147]

[0148] in, is the total loss function of the generator, is the adversarial loss term, is the reconstruction loss term, is the multimodal consistency loss term, is the expert load balancing loss term, is the weight used to balance different loss terms. In this embodiment,

[0149] During the training process, by minimizing the total loss function, the generator network is encouraged to generate images that are consistent with the characteristics of the input multimodal data while ensuring the stability and efficiency of the model. The combination of these loss terms enables the generator to maintain image quality while also taking into account the consistency of multimodal data and the load balancing of the model. In this way, the generator network can achieve better performance in multimodal generation tasks.

[0150] Step 105, constructing a discriminator network model in a generative adversarial network of a hybrid expert model;

[0151] As attached Figure 4 As shown in the figure, in the discriminator network structure, an image normalization layer is used after each convolution layer for normalization to prevent gradient diffusion, and all activation functions are set to Leaky ReLu functions to ensure that there is a gradient change when x < 0, and two additional Dropout layers are added to regularize the model and reduce overfitting.

[0152] like Figure 4 As shown, the discriminator network includes an input layer, three convolution blocks, a first Dropout layer, a flattening layer, a first fully connected layer, a Leaky ReLu activation layer, a second Dropout layer, a second fully connected layer, a Sigmoid activation layer and an output layer, and each convolution block consists of a convolution layer, a Leaky ReLu activation layer and a normalization layer.

[0153] The loss function of the discriminator D is:

[0154]

[0155] in, is the loss of the discriminative model, x is the real data, represents the probability distribution of real data, is the prior noise distribution of the generator input, i.e., the distribution of the latent variables; is the prior noise distribution The noise vector sampled in , is the output of the discriminator for the real sample, which should ideally be close to 1. is the output of the discriminator for the generated samples, and ideally should be close to 0. The first term encourages the discriminator D to correctly classify real data as real, and the second term encourages the discriminator D to classify fake data generated by the generator G as fake.

[0156] Step 106, constructing a training model;

[0157] Alternate optimization in actual training and The training process of the entire network can be regarded as an iterative optimization process of two loss functions. In each iteration, the generator G is first fixed and optimized. Update the parameters of the discriminator D, and then fix the discriminator D to optimize Update the parameters of the generator G. In the process of training the generative model, the expert layer and gating mechanism in the hybrid expert network model are trained simultaneously to better capture the complex characteristics of the data.

[0158] Step 107, after the training is completed, the generator part in the generative adversarial network of the multimodal hybrid expert model is used to generate new images.

[0159] This embodiment still uses VGG16 as the classification model, and uses the original images, texts, audio data in the ImageNet and CIFAR-10, WIT and A2IR data sets and the data after adding the pictures generated by the generative adversarial network of the hybrid expert model with the above multimodal data fusion mechanism to train it. The corresponding training result accuracy is shown in Table 2 below:

[0160] Table 2

[0161]

[0162] Experimental results show that the classification accuracy of the model is significantly improved after training with original images, text, audio data, and images generated by multimodal data fusion mechanism and hybrid expert model. Fig.10 It can be seen that by introducing the data enhancement technology combining the two networks, the model also shows excellent performance and robustness when generating images of other extreme weather conditions such as fog, sandstorms, and storms. These results further verify the effectiveness of the hybrid expert model in improving the performance of generative adversarial networks.

[0163] Some steps in the embodiments of the present invention may be implemented using software, and the corresponding software program may be stored in a readable storage medium, such as a CD or a hard disk.

[0164] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. An image generation method using a generative adversarial network based on a hybrid expert model, characterized in that: The method comprises: Step 1, construct a training data sample set; Step 2, preprocessing the data in the sample set; Step 3, constructing a generator network model in a generative adversarial network based on a hybrid expert model, wherein the loss function of the generator network model includes an adversarial loss term, an expert network loss term, a gating network loss term, and an expert balance loss term; Step 4, constructing a discriminator network model in the generative adversarial network of the hybrid expert model; Step 5, iteratively optimize and train the generator network model and the discriminator network model; Step 6: After the training is completed, the generator part of the generative adversarial network of the hybrid expert model is used to generate new images for specific application scenarios; The generator network model includes An expert network and a gating network, where each expert network focuses on different aspects or features of generated data, and the gating network is used to dynamically adjust the weight and contribution of each expert; the loss function of the generator network model is: in, is the total loss function of the generator, is the adversarial loss term, is the expert network loss term, is the gating network loss term, is the expert balance loss term, is the corresponding weight coefficient of each loss term, which is used to balance the importance of different loss terms; Adversarial Loss Term middle, is a generator, is the discriminator, is the prior noise distribution The noise vector sampled in ; Expert network loss middle, is the number of expert networks, It is Expert network, It is The loss function of the expert network is It is The weight coefficient corresponding to each expert network; Gating Network Loss middle, The gating network is The weight coefficients assigned by the expert network; Expert Balance Loss Item middle, is the weight coefficient used to balance each expert network, is the number of expert networks, represents the L2 norm; The loss function of the discriminator network model is: in, is the total loss function of the discriminator, is the real sample loss term, is the sample generation loss term; In the real sample loss term, It is real data. represents the probability distribution of real data, is the output of the discriminator for the real sample; In the generated sample loss term, is the prior noise distribution The noise vector sampled in ; is the generator according to the noise The generated image, is the output of the discriminator for the generated image.

2. The image generation method using a generative adversarial network based on a hybrid expert model according to claim 1, characterized in that: The discriminator network model includes an input layer, three convolution blocks, a first Dropout layer, a flattening layer, a first fully connected layer, a Leaky ReLu activation layer, a second Dropout layer, a second fully connected layer, a Sigmoid activation layer and an output layer, and each convolution block consists of a convolution layer, a Leaky ReLu activation layer and a normalization layer.

3. The image generation method using a generative adversarial network based on a hybrid expert model according to claim 2, characterized in that: The loss function of the generator network model is alternately optimized in step 5. And the loss function of the discriminator network model The training is iteratively optimized by first fixing the generator network model and optimizing the loss function of the discriminator network model. Update the parameters of the discriminator network model, and then fix the discriminator network model to optimize the loss function of the generator network model Update the parameters of the generator network model.

4. The image generation method using a generative adversarial network based on a hybrid expert model according to claim 3, characterized in that: If the data sample is multimodal data, step 2 further includes fusing the multimodal data, and the fusing method includes: Step S1, performing data cleaning on multimodal data, wherein the multimodal data includes image data, text data, and audio data; Step S2, segmenting, encoding, embedding the text data, converting it into numerical representation, and sampling, filtering, and feature extraction of the audio data; Step S3, constructing a multimodal data fusion network to generate multimodal fusion data.

5. The image generation method using a generative adversarial network based on a hybrid expert model according to claim 4, characterized in that: The multimodal data fusion network includes a lightweight convolutional neural network for extracting image features and two lightweight recurrent neural networks for extracting text and audio features respectively; the lightweight convolutional neural network for extracting image features includes, in sequence, an original image input layer, a convolutional layer, a deep convolutional layer, a global average pooling layer, and a fully connected output layer, wherein the deep convolutional layer includes a point-by-point convolutional layer, a normalization layer, and a ReLU6 activation function; the lightweight recurrent neural network for extracting text features includes, in sequence, an embedding layer, a RNN layer, a bidirectional RNN layer, a merging layer, a first fully connected layer, a second fully connected layer, and an output layer; the lightweight recurrent neural network for extracting audio features includes, in sequence, a convolutional layer, a pooling layer, an RNN layer, a first fully connected layer, a second fully connected layer, and an output layer.

6. An image classification method, characterized in that: The method generates new image data based on the method described in any one of claims 1 to 5 to expand the training data set.

Citation Information

Patent Citations

  • Ore image generation method based on generative adversarial network and computer readable storage medium

    CN112598034A

  • Underwater sonar simulation image generation and data expansion method based on generative adversarial network

    CN113139916A

  • Method for generating face through deep separable convolutional generative adversarial neural network

    CN117252939A

  • Structural recognition model training method and device, model structure recognition method and device and medium

    CN117609870A

  • Method and system for improving structure of language model based on hybrid expert model

    CN118194917A