Image generation method and system of frequency division quantization variational auto-encoder based on 2D-FFT
By introducing 2D-FFT and graph convolution theory in VQ-VAE, the problem of degraded generation quality and impaired feature global relationships when processing complex images is solved, and higher quality and diverse image generation is achieved.
Patent Information
- Application Number
- CN202510003356.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-06
AI Technical Summary
When processing complex images, existing VQ-VAEs may experience problems such as degradation in generation quality and damage to global characteristics, resulting in the lack of details of the generated image, affecting the sense of reality and diversity.
Using a frequency-dividing quantized variational autoencoder based on 2D-FFT, the image is encoded into frequency domain feature components through two-dimensional fast Fourier transform, quantized, and dynamically updated the codebook, and introduced graph convolution theory to correct the quantized features.
Significantly improves the quality and diversity of generated images, reduces information loss, and enhances the expressiveness and visual appeal of images, especially when generating high-definition face images and natural scenes.
Smart Images

Figure CN119941894A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image generation, and in particular to an image generation method and system of a frequency-divided quantized variational autoencoder based on 2D-FFT. Background Art
[0002] Image generation is a key task in the field of computer vision. Its core lies in using algorithms to generate new images, which can be completely innovative or variants of existing images. The significant advantage of image generation technology is that it can automatically generate high-quality images, which plays an important role in many fields. In the entertainment industry, it helps to create virtual characters and scenes; in architectural design, it quickly generates a variety of architectural exterior design schemes; in the field of medical imaging, it is used to generate synthetic medical images to assist doctors in diagnosis and research. In addition, image generation technology has also shown great application potential in many fields such as artistic creation, fashion design, and advertising production.
[0003] With the rapid development of deep learning technology, the field of image generation has become particularly active, and many innovative generation models have emerged. Among them, the vector quantized variational autoencoder (VQ-VAE) has attracted much attention due to its excellent performance. VQ-VAE simplifies the optimization process of the model by mapping continuous features to discrete space and selecting the optimal discrete vector using the nearest neighbor algorithm, enabling it to efficiently capture and represent the complex structure and characteristics of the data. However, despite its excellent performance in image generation tasks, VQ-VAE still faces many technical difficulties in processing complex images. For example, when generating highly complex images such as faces and buildings, problems such as decreased generation quality and damaged global feature relationships may occur, resulting in the loss of details in the generated images, affecting the realism and diversity of the images. In order to meet these challenges, existing studies have attempted to introduce methods such as attention mechanisms and multi-scale feature extraction to maintain the global structure of the generated images. However, these methods often increase the complexity of the model and may have an adverse effect on the training efficiency and generation speed of the model.
[0004] In addition, the codebook collapse problem remains a serious challenge for VQ-VAE. In some cases, the model may only optimize a few discrete vectors, resulting in a large number of code vectors not being fully utilized, which in turn affects the diversity and expressiveness of the model. Summary of the invention
[0005] The purpose of the present invention is to propose an image generation method and system based on a 2D-FFT-based frequency-division quantized variational autoencoder, which retains the overall relationship of the image from a frequency domain perspective and significantly improves the generation quality; reduces information loss during image processing and enhances the expressiveness and diversity of the generated image.
[0006] According to a first aspect of an embodiment of the present disclosure, a method for generating an image using a frequency-divided quantized variational autoencoder based on 2D-FFT is provided, comprising the following steps:
[0007] Obtain public image datasets and perform preprocessing;
[0008] The common image is encoded into frequency domain feature components using 2D-FFT, and the frequency domain information is extracted and quantified by linear combination.
[0009] Dynamically update the codebook based on the usage frequency of code vectors and the dependencies between them;
[0010] The graph convolution theory is introduced to quantize the correction coefficients of the features, thereby repairing the losses in the quantization process and improving the quality and diversity of the generated images.
[0011] In one embodiment, a common image is encoded into frequency domain feature components using a two-dimensional fast Fourier transform (2D-FFT), and frequency domain information is extracted and quantized by a linear combination method, specifically:
[0012] Sampling is performed through convolution, so that the public image is mapped into the latent space and the channel dimension is advanced to obtain the latent feature z(x)∈R D×h×w For each channel, the feature decomposition and extraction of different frequency domains are achieved through the transformation of different frequencies. The specific implementation is:
[0013] z e (x)=K☉F(x) (1)
[0014] Where ⊙ represents the element-wise product operation; is a learnable filter; F(x) represents the feature after 2D-FFT, z e (x) is the final latent feature of the encoder, and each element in it will be quantized as the minimum unit.
[0015] In one embodiment, a set of N filter coefficients m is generated by a multilayer perceptron (MLP) i , (where i∈0,1,...N), and define a set of filter basis For different frequency domain characteristic filters, the following definitions are given:
[0016]
[0017] Finally, the combination of learnable filters is:
[0018]
[0019] The encoder extracts frequency domain features in the channel dimension of the potential feature z(x), and finally quantizes the image features with different frequency features of different channels as the smallest unit, thus retaining the global relationship of the image.
[0020] In one embodiment, the codebook is dynamically updated according to the usage frequency of the code vectors and the dependencies between them, and the updated code vectors as follows:
[0021]
[0022] in, is the resampled anchor vector, which is used to provide the correct update direction for the code vector. This vector is obtained by finding the nearest neighbor potential vector; is the attenuation value:
[0023]
[0024] Where M is the number of anchor vectors, ∈ is a constant; represents the average number of times the kth entry is used after t steps, and its update formula is:
[0025]
[0026] in is the usage count of the kth entry in the mini-batch after t steps, is the average number of times the kth entry is used after t steps, γ is the decay hyperparameter, B is the batch size; Bhw is the set of vectors for each batch.
[0027] In one embodiment, the Pearson correlation coefficient δ is introduced when dynamically updating the codebook to evaluate the linear correlation strength between the vector in the codebook and the other vectors; for any code vector e in the codebook k , Pearson correlation coefficient δ k for:
[0028]
[0029] Among them, j represents the number of bits in the codebook except e k The remaining code vectors outside; cov(e k ,e j ) is e k and e j The covariance of is their standard deviation; covariance cov(e k ,e β ) is expressed as:
[0030]
[0031] Where n represents the dimension of the code vector.
[0032] In one embodiment, the latent feature z can be restored approximately losslessly by using the correction coefficient w of the quantized feature. e (x), specifically expressed as:
[0033] z e (x)≈w1e1+w2e2+…+w n e n (9)
[0034] Graph convolution theory is introduced to solve the correction coefficient w. Specifically, the feature map of the common image is first converted to perform a global average pooling operation (GAP), and its dimension is converted to U×1×1, where U=N×D is the product of the number of filter bases and the number of channels, that is, the number of final frequency features, representing the number of feature vertices in the graph; then a 1×1 convolution operation is performed to further enhance the feature representation capability, and thus a set of preliminary feature mapping groups are obtained; then, the graph convolution module is passed to obtain the correction coefficient.
[0035] In one embodiment, let the mapping of each frequency feature be a graph node, and the graph convolution module is represented as:
[0036]
[0037] Where W is the weight of different network layers; f and They represent the input and output feature maps respectively; A is initialized to a U×U unit matrix, representing the relationship between the feature vertices themselves; B is a U×U diagonal matrix, representing the weights of the feature vertices, obtained by performing one-dimensional convolution and softmax operations on f, so that the network can adaptively emphasize or suppress features; C is a learnable U×U adjacency matrix, which is optimized by back-propagation and represents the relationship between feature vertices.
[0038] The weights obtained by the graph convolution operation are used as the weights of the frequency features, and then the frequency level response is recalibrated; the frequency features after weighted summation are converted into the spatial domain, and finally the potential features in the spatial domain are restored to the original space through the convolution operation to achieve high-quality image generation.
[0039] According to a second aspect of an embodiment of the present disclosure, an image generation system of a frequency-divided quantized variational autoencoder based on 2D-FFT is provided, comprising:
[0040] The preprocessing module obtains the public image dataset and performs preprocessing;
[0041] The encoding module uses two-dimensional fast Fourier transform 2D-FFT to encode the public image into frequency domain feature components and extracts the frequency domain information through linear combination;
[0042] The quantization module dynamically updates the codebook according to the usage frequency of the code vectors and the dependencies between them, and quantizes the frequency domain features in the continuous space to the discrete space represented by the codebook;
[0043] The decoding module introduces graph convolution theory to generate correction coefficients of discrete features, thereby repairing the loss in the quantization process and improving the quality and diversity of the generated images.
[0044] According to a third aspect of an embodiment of the present disclosure, an electronic device is provided, comprising a memory, a processor, and a computer program stored and running on the memory, wherein when the processor executes the program, the image generation method of the frequency-divided quantized variational autoencoder based on 2D-FFT is implemented.
[0045] According to a fourth aspect of an embodiment of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the program is executed by a processor, the image generation method of the frequency-divided quantized variational autoencoder based on 2D-FFT is implemented.
[0046] Compared with the prior art, the above technical solution adopted by the present invention has the following advantages: the present invention makes full use of the frequency domain features and can more comprehensively capture the complex structures and details in the image, thereby showing higher accuracy in the practical application of image generation. In particular, when generating high-definition face images, natural scenes, and performing artistic style conversion, the present invention can provide higher-quality image output and significantly improve the user's visual experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] The drawings in the specification, which constitute a part of the present application, are used to provide further understanding of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute improper limitations on the present application.
[0048] Figure 1 This is a model architecture diagram of the frequency-division quantized variational autoencoder of the present invention;
[0049] Figure 2 This is a diagram of the graph convolution module architecture of the present invention;
[0050] Figure 3 This is a visualization diagram of the codebook usage rate of the present invention on the MNIST (top) and CIFAR 10 (bottom) validation sets;
[0051] Figure 4This is a visualization diagram comparing the information content per unit of the codebook of the present invention on the CelebA (top) and LSUN Church (bottom) validation sets;
[0052] Figure 5 This is an example diagram of the qualitative comparison results of the present invention in the LSUN Church validation set. DETAILED DESCRIPTION
[0054] The present disclosure is further described below in conjunction with the accompanying drawings and embodiments.
[0055] It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.
[0056] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0057] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the methods and systems according to various embodiments of the present disclosure. It should be noted that each box in the flowchart or block diagram can represent a module, a program segment, or a part of a code, and the module, program segment, or a part of a code may include one or more executable instructions for implementing the logical functions specified in each embodiment. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, or they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the flowchart and / or block diagram, and the combination of boxes in the flowchart and / or block diagram can be implemented using a dedicated hardware-based system that performs a specified function or operation, or can be implemented using a combination of dedicated hardware and computer instructions.
[0058] Embodiment 1:
[0059] The original common image is encoded into feature components in the frequency domain and quantized by two-dimensional fast Fourier transform 2D-FFT to solve the problem of local distortion in the image generation process caused by the loss of global relationships due to spatial domain segmentation. In addition, the present invention comprehensively considers the frequency and dependency of code vectors to improve the codebook in view of the insufficient generalization performance of image generation caused by the low frequency of use and high dependency of the codebook, and adopts a dynamic update method. At the same time, the graph convolution theory is introduced to prevent the loss of important information in the generated image caused by quantization errors, and the correction coefficient of the quantization feature is designed to repair the quantization feature and further reduce the quantization loss. This innovative solution provides a new idea for solving the problems of reduced generation quality and insufficient feature expression of the model when processing complex images. In order to make the purpose, features and advantages of the present invention clearer and easier to understand, it will be further described in detail below in conjunction with the accompanying drawings and specific implementation details.
[0060] This embodiment provides an image generation method of a frequency-division quantized variational autoencoder based on 2D-FFT, comprising the following steps:
[0061] S1. Obtain a public image dataset and perform preprocessing;
[0062] Specifically, public image datasets include MNIST, CIFAR-10, CelebA, and LSUN Church. MNIST is a grayscale image dataset containing 70,000 handwritten digits; CIFAR-10 contains 60,000 32×32 color images divided into 10 categories, such as airplanes, cars, animals, etc.; CelebA contains 200,000 facial images of celebrities, annotated with 40 different attributes; LSUN Church is a large-scale image dataset that mainly contains images of buildings and indoor scenes, focusing on scene understanding and generation tasks.
[0063] The acquired images were preprocessed to improve the stability during training. For the MNIST and CIFAR10 datasets, the experiments used their original resolutions, i.e., 28×28 pixels and 32×32 pixels. The CelebA dataset was first center-cropped to a size of 140×140 pixels, and then resized to 64×64 pixels. The LSUN Church dataset used the official version and was divided into training, validation, and test sets in a ratio of 8:1:1, and the image size was resized to 128×128 pixels.
[0064] S2. Use 2D-FFT to encode the common image into frequency domain feature components, extract the frequency domain information through linear combination and quantize it;
[0065] like Figure 1As shown in the figure, in order to deal with the global relationship loss and local distortion problems caused by spatial domain segmentation during image generation, the spatial domain feature map is converted into an important feature combination in the frequency domain through 2D-FFT, so as to effectively and adaptively extract more useful feature information to improve the overall quality and consistency of the generated image. The details are as follows:
[0066] Traditional methods usually map the original data into a set of regular image blocks in a high-dimensional space domain and quantize them in units of these image blocks. However, this grid-based segmentation operation often destroys the global relationship of the image, resulting in information loss. In contrast, the method proposed in the present invention always retains the global characteristics of the image during the encoding process. The encoder encodes the original data into a linear combination of these features in units of different frequency feature components in the frequency domain.
[0067] For a given two-dimensional signal x(h,w), the two-dimensional discrete Fourier transform is defined as:
[0068]
[0069] in
[0070] 2D-DFT maps data from the spatial domain to the frequency domain, where The space to which it belongs is the frequency domain. The 2D-DFT generates a full complex matrix, half of which is redundant due to the conjugate symmetry of the real input, which results in its direct implementation having 2D-FFT reduces the time complexity by taking advantage of the symmetry and periodicity in DFT calculations to reduce the number of multiplications and additions required.
[0071] This method first follows the general operation and samples through convolution, mapping the original data into the latent space and advancing the channel dimension to obtain the latent feature z(x)∈R D×h×w , where each channel is actually a standard two-dimensional signal. For each channel, different frequency transformations are used to achieve feature decomposition and extraction in different frequency domains. The specific implementation is as follows:
[0072] Z e (x) = K⊙F(x) (2)
[0073] Where ⊙ represents the element-wise product operation; is a learnable filter; F(x) represents the feature after 2D-FFT, z e(x) is the final potential feature of the encoder, in which each element will be quantized as the smallest unit. The purpose of K is to dynamically generate filters according to data features and capture important frequency components in different data, so it has dynamic and adaptive properties. The present invention generates a set of N filter coefficients m through a multilayer perceptron (MLP). i , (where i∈0,1,...N), and define a set of filter basis For different frequency domain characteristic filters, the following definitions are given:
[0074]
[0075] Finally, the combination of learnable filters is:
[0076]
[0077] The encoder extracts frequency domain features in the channel dimension of the potential feature z(x), and finally quantizes the image features with different frequency features of different channels as the minimum unit. This method does not perform segmentation in the spatial domain, and retains the global relationship of the image. At the same time, the feature extraction scheme based on the global dynamic filter of the FFT used in the present invention has a smaller time complexity than other global feature extractors (such as Transformer):
[0078] S3. Dynamically update the codebook according to the usage frequency of the code vectors and the dependencies between them;
[0079] Specifically, in response to the problem of insufficient generalization performance of image generation caused by low frequency of use and high dependency of the codebook, the present invention proposes a method for dynamically updating the codebook, which comprehensively considers the frequency of use and dependency of the code vector. In terms of frequency of use, this method regularly updates all code vectors to prevent codebook collapse, thereby improving the adaptability of the generation model to diverse images. In terms of dependency, special attention is paid to code vectors with high dependencies to ensure that they can obtain more updates, which can reduce the coupling and sharing of information. This dynamic update mechanism effectively improves the diversity and quality of generated images, enabling the generation model to better adapt to complex and diverse image generation tasks. The details are as follows:
[0080] In order to optimize the inactive code vectors in the code book, the present invention proposes an improved dynamic code book update strategy. The strategy dynamically updates the code book according to the usage frequency of the code vectors and the dependencies between them. The updated code vector as follows:
[0081]
[0082] in, is the resampled anchor vector, which is used to provide the correct update direction for the code vector. This vector is obtained by finding the nearest neighbor potential vector; is the attenuation value:
[0083]
[0084] Where M is the number of anchor vectors and ∈ is a constant (can be set to 0.001). represents the average number of times the kth entry is used after t steps, and its update formula is:
[0085]
[0086] in is the usage count of the kth entry in the mini-batch after t steps, is the average number of times the k-th entry is used after t steps, γ is the decay hyperparameter (can be set to 0.99), and B is the batch size.
[0087] In addition, in order to reduce the dependency between code vectors in the codebook, so as to more accurately linearly represent the quantized data and avoid excessive information coupling, the Pearson correlation coefficient δ is introduced, which is used to evaluate the linear correlation strength between the code vector in the codebook and the other vectors. k , Pearson correlation coefficient δ k for:
[0088]
[0089] Among them, j represents the number of bits in the codebook except e k The remaining code vectors. k ,e j ) is e k and e j The covariance of is their standard deviation. The covariance can be expressed as:
[0090]
[0091] Where n represents the dimension of the code vector. k The size of reflects the linear correlation between the code vector and other code vectors. When the correlation is higher, the code vector is updated more. This dynamic learning process helps the code vector better capture the essential characteristics of the data and improve linear independence.
[0092] S4. Introduce graph convolution theory and quantize the correction coefficients of features, thereby repairing the loss in the quantization process and improving the quality and diversity of generated images.
[0093] In order to prevent the loss of key detail information in the generated image due to quantization error, the present invention introduces graph convolution theory, designs and solves the correction coefficient of the quantized features to repair the loss generated during the quantization process. This method significantly reduces the quantization loss, thereby improving the quality and diversity of the generated images. By introducing the graph convolutional model (GCM), the present invention can assign correction weights to features of different frequencies, thereby further reducing the information loss caused by quantization. This correction mechanism ensures that the generative model can effectively retain the details and features in the image, enhance the effect of image generation, and improve the richness and visual appeal of the generated image. The details are as follows:
[0094] Traditional quantization methods inevitably lead to information loss, mainly because there is a distance deviation between the code vector and the latent vector. In the method of the present invention, the input of the decoder is a linear combination of different frequency domain features. At the same time, due to the optimization of the correlation strength during the codebook training process, the codebook entries tend to be linearly independent groups. According to the theorem in linear algebra, if there is a set of N-dimensional linearly independent vector groups, and the number is greater than N, then any N-dimensional vector can be linearly represented. This provides a theoretical basis for the correction of quantized features. By solving the correction coefficient w of the quantized feature group, the potential feature z can be approximately restored losslessly e (x). Specifically expressed as:
[0095] Z e (x)≈w1e1+w2e2+…+w n e n #(10)
[0096] This method introduces graph convolution theory to solve the correction w. Specifically, the feature map of the common image is first converted to global average pooling (GAP) and its dimension is converted to U×1×1, where U=N×D is the product of the number of filter bases and the number of channels, that is, the number of final frequency features, which represents the number of feature vertices in the graph. After that, a 1×1 convolution operation is performed to further enhance the feature representation capability, and thus a set of preliminary feature mapping groups is obtained. It then passes through the graph convolution module to obtain the feature weights. The architecture of the graph convolution module is as follows: Figure 2 As shown in Figure 1, the input feature map is mapped into feature weights containing the relationship between features through the graph convolution network. Let the mapping of each frequency feature be a graph node, and the graph convolution module can be expressed as:
[0097]
[0098] Where W is the weight of different network layers. f and Represent the input and output feature maps respectively. A is initialized as a U×U identity matrix, representing the relationship between the feature vertices themselves. B is a U×U diagonal matrix, representing the weights of the feature vertices, obtained by performing a one-dimensional convolution and softmax operation on f, which allows the network to adaptively emphasize or suppress features. C is a learnable U×U adjacency matrix, which is optimized by back-propagation and represents the relationship between feature vertices.
[0099] The feature weights calculated by the graph convolution operation will be used as the weights of the frequency features, and then the frequency level response will be recalibrated. After the weighted sum of the frequency features, the features in the frequency domain are converted to the spatial domain. The subsequent operations are the same as those of the traditional VQ-VAE. Finally, the potential features in the spatial domain are restored to the original space through the convolution operation to achieve high-quality image generation.
[0100] The architecture of the present invention includes an encoder, a quantizer, and a decoder. The encoder first transfers the features of the input image to the frequency domain through Fourier transform, and then maps the frequency domain features into a set of multiple continuous components through different filters; the quantizer mainly maintains a codebook with a dimension of N, which is mutually optimized with the above-mentioned continuous components after training, and then quantizes the continuous components into discrete components through nearest neighbor replacement; the decoder generates a set of corrected weights of discrete components through a graph convolution module, which are transferred to the spatial domain through inverse Fourier transform after correction, and finally restores the features to the original space through continuous upsampling to complete image generation.
[0101] Embodiment 2:
[0102] This embodiment provides an image generation system of a frequency-division quantized variational autoencoder based on 2D-FFT, including:
[0103] The preprocessing module obtains the public image dataset and performs preprocessing;
[0104] The encoding module uses two-dimensional fast Fourier transform 2D-FFT to encode the public image into frequency domain feature components and extracts the frequency domain information through linear combination;
[0105] The quantization module dynamically updates the codebook according to the usage frequency of the code vectors and the dependencies between them, and quantizes the frequency domain features in the continuous space to the discrete space represented by the codebook;
[0106] The decoding module introduces graph convolution theory to generate correction coefficients of discrete features, thereby repairing the loss in the quantization process and improving the quality and diversity of the generated images.
[0107] Embodiment three:
[0108] An electronic device includes a memory, a processor, and a computer program stored and running on the memory, wherein the processor implements the above-mentioned image generation method based on 2D-FFT frequency division quantization variational autoencoder when executing the program, including:
[0109] Obtain public image datasets and perform preprocessing;
[0110] The common image is encoded into frequency domain feature components using 2D-FFT, and the frequency domain information is extracted and quantified by linear combination.
[0111] Dynamically update the codebook based on the usage frequency of code vectors and the dependencies between them;
[0112] The graph convolution theory is introduced to quantize the correction coefficients of the features, thereby repairing the losses in the quantization process and improving the quality and diversity of the generated images.
[0113] Embodiment 4:
[0114] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the above-mentioned image generation method of the frequency-divided quantized variational autoencoder based on 2D-FFT, comprising:
[0115] Obtain public image datasets and perform preprocessing;
[0116] The common image is encoded into frequency domain feature components using 2D-FFT, and the frequency domain information is extracted and quantified by linear combination.
[0117] Dynamically update the codebook based on the usage frequency of code vectors and the dependencies between them;
[0118] The graph convolution theory is introduced to quantize the correction coefficients of the features, thereby repairing the losses in the quantization process and improving the quality and diversity of the generated images.
[0119] The present invention uses the following indicators for performance evaluation: Structural Similarity (SSIM) at the image block level, Learning Perceptual Image Patch Similarity (LPIPS) at the feature level, and Reconstructed Fréchet Inception Distance (rFID) at the dataset level.
[0120] This embodiment conducts experiments on image generation tasks. To ensure fairness, all experiments use the same downsampling rate and hyperparameter settings. In the generation experiment, the superiority of the method of the present invention is verified by comparing the image generation quality, codebook utilization, and codebook unit information volume in the image generation task. The ablation experiment further confirms the impact of all components on the model performance and its practicality, and determines the optimal combination of the number of channels and the number of filters. The details are as follows:
[0121] This example first maps the features through a convolutional layer. For the MNIST, CIFAR10, and CelebA datasets, only the channel dimension is changed without changing the resolution size after resizing the image; for the LSUN Church dataset, j and w are adjusted to H / 2 and W / 2, respectively.
[0122] For the learning rate, except for MNIST, it is set to 2×10 -3 , the generation tasks on the other datasets are all set to 3×10 -3 The weight hyperparameter β in the loss is set to 0.25, and the decay hyperparameter β is set to 0.99. All experiments are performed on an NVIDIA A800 PCIe with 80GB of video memory.
[0123] The present invention performs image generation tasks on four data sets. In the experiment, all data are uniformly standardized: the mean is 0.5 and the variance is 0.5. In order to ensure fairness, the number of encoded features of the method adopted in this embodiment is consistent with the number of latent vectors encoded by the comparison method, which is 1 / 4 of the original resolution, that is, the downsampling rate of f=4, which is applied to all data sets.
[0124] By measuring and visualizing the usage rate of the codebook on the MNIST and CIFAR10 validation sets, the present invention verifies its effective improvement on the codebook. Specifically, the present invention counts the total number of times all entries in the codebook are used in the validation set. Figure 3 As shown, the curve in the figure is the fitting curve of the codebook entries sorted by usage rate. It can be clearly seen that the codebook utilization rate of the method of the present invention is higher than that of the basic model VQ-VAE and the most advanced CVQ-VAE method. In VQ-VAE, unused or rarely used entries account for a large proportion, and only a small number of entries are used very frequently, which causes the phenomenon of codebook collapse. Although CVQ-VAE optimizes the usage distribution of active entries, there is still unevenness. In contrast, the method of the present invention can achieve better codebook utilization, making the usage distribution of code vectors more uniform, thereby significantly improving the performance and diversity of image generation.
[0125] Furthermore, previous studies have not considered the calculation of the amount of information per unit of the codebook. This embodiment calculates the amount of information of the codebook vector by removing a codebook vector and measuring the mean square error (MSE) loss difference between the generated image and the original image before and after the removal. Specifically, the amount of information per unit of the code vector is evaluated by comparing the MSE loss difference before and after the removal on the scale of the data set. This indicator effectively reflects the effectiveness and compactness of the codebook information. According to the measurement results on the CelebA and LSUN Church data sets, Figure 4As shown in the figure, the average unit information content of the codebook generated by the method of the present invention is higher than that of the previous methods. This result shows that compared with the traditional image features obtained by spatial domain segmentation, the frequency domain features proposed by the present invention contain more effective information, and the structure of the codebook is more compact and efficient, thereby significantly improving the quality of the generated image.
[0126] Furthermore, in order to verify the advantages of the network model proposed in the present invention, this embodiment quantitatively compares image generation in four public datasets: MNIST, CIFAR10, CelebA, and LSUN Church, and compares it with the most advanced quantized autoencoder models VQ-VAE, HVQ-VAE, SQ-VAE, and CVQ-VAE. The comparison results are shown in Table 1.
[0127] Table 1 Generation results on CIFAR-10, MNIST, CelebA and LSUNChurch validation sets
[0128]
[0129]
[0130] From the experimental results, it can be seen that the method of the present invention effectively improves the generation of images by quantizing frequency domain features, optimizing codebook structure, and correcting quantized features. This indicates that the model has learned more effective information of the original image. In all performance metrics on all datasets, except for the SSIM index in the MNIST and CelebA datasets, the method of the present invention outperforms the four SOTA baseline models.
[0131] Figure 5 Some example generation results trained on the LSUN Church dataset are shown. It can be seen that the proposed method achieves better visual effects and also shows good performance in challenging details.
[0132] In order to more comprehensively evaluate the impact of the improvement measures proposed in the present invention on the model performance, this embodiment conducts an ablation experiment, by gradually introducing each core component to verify the contribution of each part to the improvement of model performance. The experimental results are shown in Table 2.
[0133] Table 2 Evaluation results of core components on CIFAR-10, MNIST, CelebA and LSUNChurch validation sets
[0134]
[0135] It can be seen from the experimental results that all components designed by the present invention effectively and greatly improve the quality of image generation, starting with the baseline configuration (A) that implements VQ-VAE. In configuration (B), frequency domain feature quantization replaces the spatial domain vector quantization method, which significantly improves the performance of the model and also shows that frequency domain feature quantization is the main influencing factor of the performance improvement of the present invention. Configuration (C) uses selected anchor points to reinitialize the code vectors with low usage frequency. Although it shows good gain on the small data sets of MNIST and CIFAR10, it performs poorly on CelebA and LSUN Church due to the more serious information coupling of large data sets. Configuration (D) adds the consideration of codebook dependency on the basis of (C), reduces the information coupling of the codebook, and the performance of all indicators in various data sets is significantly improved. Finally, in configuration (E), the graph convolutional network generates correction coefficients for the quantized features. The corrected features reduce the quantization error and further improve the results. This method effectively solves the global relationship loss caused by spatial domain segmentation and avoids local distortion that may occur during the generation process. This breakthrough directly enhances the coherence and realism of image generation. In addition, in order to address the insufficient generalization performance caused by the low frequency of codebook use and high dependency, a dynamic update mechanism was introduced to significantly improve the adaptability of the generation model, thereby expanding the diversity of generated images. At the same time, the quantization error was reduced through the graph convolutional network, ensuring that the key information in the generated image was retained, and improving the overall effect and quality of image generation.
[0136] Consistent with the expected results, the image generated by the present invention is clearer and more realistic than the method based on spatial domain vector quantization.
[0137] Those skilled in the art should understand that the modules or steps of the present disclosure can be implemented by a general-purpose computer device, or alternatively, they can be implemented by a program code executable by a computing device, so that they can be stored in a storage device and executed by the computing device, or they can be made into individual integrated circuit modules, or multiple modules or steps therein can be made into a single integrated circuit module for implementation. The present disclosure is not limited to any specific combination of hardware and software.
[0138] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.
[0139] Although the above describes the specific implementation methods of the present disclosure in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present disclosure. Technical personnel in the relevant field should understand that on the basis of the technical solution of the present disclosure, various modifications or variations that can be made by those skilled in the art without creative work are still within the scope of protection of the present disclosure.
Claims
1. An image generation method based on a 2D-FFT frequency-division quantized variational autoencoder, characterized in that: The following steps are involved: Obtain public image datasets and perform preprocessing; The common image is encoded into frequency domain feature components using two-dimensional fast Fourier transform (2D-FFT), and the frequency domain information is extracted through linear combination. The codebook is dynamically updated according to the usage frequency of code vectors and the dependencies between them, and the frequency domain features in the continuous space are quantized to the discrete space represented by the codebook; Graph convolution theory is introduced to generate correction coefficients of discrete features, thereby repairing the loss in the quantization process and improving the quality and diversity of generated images.
2. The image generation method of the frequency-division quantized variational autoencoder based on 2D-FFT according to claim 1 is characterized in that: The public image is encoded into frequency domain feature components using two-dimensional fast Fourier transform (2D-FFT), and the frequency domain information is extracted and quantized by linear combination, specifically: Sampling is performed through convolution, so that the public image is mapped into the latent space and the channel dimension is advanced to obtain the latent feature z(x)∈R D×h×w For each channel, the feature decomposition and extraction of different frequency domains are achieved through the transformation of different frequencies. The specific implementation is: z e (x)=K⊙F(x) (1) Where ⊙ represents the element-wise multiplication operation; is a learnable filter; F(x) represents the features after 2D-FFT, z e (x) is the final latent feature of the encoder, and each element in it will be quantized as the minimum unit.
3. The image generation method of the frequency-division quantized variational autoencoder based on 2D-FFT according to claim 2 is characterized in that: Generate a set of N filter coefficients m through a multi-layer perceptron i , (where i∈0,1,...N), and define a set of filter basis For different frequency domain characteristic filters, the following definitions are given: Finally, the combination of learnable filters is: The encoder extracts frequency domain features in the channel dimension of the potential feature z(x), and finally quantizes the image features with different frequency features of different channels as the smallest unit, preserving the global relationship of the image.
4. The image generation method based on 2D-FFT frequency division quantization variational autoencoder according to claim 1, characterized in that: The codebook is dynamically updated based on the usage frequency of the code vectors and the dependencies between them. The updated code vectors as follows: in, is the resampled anchor vector, which is used to provide the correct update direction for the code vector. This vector is obtained by finding the nearest neighbor potential vector; is the attenuation value: Where M is the number of anchor vectors, ∈ is a constant; represents the average number of times the kth entry is used after t steps, and its update formula is: in is the usage count of the kth entry in the mini-batch after t steps, is the average number of times the kth entry is used after t steps, γ is the decay hyperparameter, B is the batch size; Bhw is the set of vectors for each batch.
5. The image generation method of the frequency-division quantized variational autoencoder based on 2D-FFT according to claim 1 or 4, characterized in that: When dynamically updating the codebook, the Pearson correlation coefficient δ is introduced to evaluate the linear correlation strength between the vector in the codebook and the other vectors; for any code vector e in the codebook k , Pearson correlation coefficient δ k for: Among them, j represents the number of bits in the codebook except e k The remaining code vectors outside; cov(e k ,e j ) is e k and e j The covariance of is their standard deviation; covariance cov(e k ,e β ) is expressed as: Where n represents the dimension of the code vector.
6. The image generation method based on 2D-FFT frequency division quantization variational autoencoder according to claim 1, characterized in that: By quantizing the correction coefficient w of the feature, the potential feature z can be restored approximately losslessly e (x), specifically expressed as: z e (x)≈w1e1+w2e2+…+w n e n (9) Graph convolution theory is introduced to solve the correction coefficient w. Specifically, the feature map of the common image is first converted to a global average pooling operation, and its dimension is converted to U×1×1, where U=N×D is the product of the number of filter bases and the number of channels, that is, the number of final frequency features, representing the number of feature vertices in the graph; then, a convolution operation is performed to further enhance the feature representation capability, and thus a set of preliminary feature mapping groups are obtained; then, the graph convolution module is passed to obtain the correction coefficient.
7. The image generation method based on 2D-FFT frequency division quantization variational autoencoder according to claim 6 is characterized in that: Let the mapping of each frequency feature be a graph node, and the graph convolution module is expressed as: Where W is the weight of different network layers; f and Represent the input and output feature maps respectively; A is initialized as a U×U unit matrix, representing the relationship between the feature vertices themselves; B is a U×U diagonal matrix, representing the weights of the feature vertices, obtained by performing a one-dimensional convolution and softmax operation on f; C is a learnable U×U adjacency matrix, which is optimized by back-propagation and represents the relationship between feature vertices; The weights obtained by the graph convolution operation are used as the weights of the frequency features, and then the frequency level response is recalibrated; the frequency features after weighted summation are converted into the spatial domain, and finally the potential features in the spatial domain are restored to the original space through the convolution operation to achieve high-quality image generation.
8. An image generation system based on 2D-FFT frequency-division quantized variational autoencoder, characterized in that: include: The preprocessing module obtains the public image dataset and performs preprocessing; The encoding module uses two-dimensional fast Fourier transform 2D-FFT to encode the public image into frequency domain feature components and extracts the frequency domain information through linear combination; The quantization module dynamically updates the codebook according to the usage frequency of the code vectors and the dependencies between them, and quantizes the frequency domain features in the continuous space to the discrete space represented by the codebook; The decoding module introduces graph convolution theory to generate correction coefficients of discrete features, thereby repairing the loss in the quantization process and improving the quality and diversity of the generated images.
9. An electronic device comprising a memory, a processor and a computer program stored and running on the memory, characterized in that: When the processor executes the program, the image generation method of the frequency-division quantized variational autoencoder based on 2D-FFT is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the image generation method of the frequency-divided quantized variational autoencoder based on 2D-FFT is implemented.
Citation Information
Cited By
Hyperspectral image reconstruction method, device and system and storage medium
CN120707732A
Remote sensing image registration method and system based on frequency domain transfer learning and feature extraction
CN121414798A