Feature clustering and region self-adaption multi-discriminator generative adversarial network enhancement method
Through the multi-discriminator generation adversarial network enhancement method of feature clustering and region adaptation, the limitations of the traditional single discriminator GAN architecture when processing complex texture images are solved, and the image detail recovery quality and authenticity are improved.
Patent Information
- Application Number
- CN202510560810.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-04-30
AI Technical Summary
The traditional single discriminator GAN architecture has limitations when processing images with complex textures, making it difficult to balance the generation quality of multimodal features, and is unable to effectively integrate cross-scale features, resulting in artifacts or excessive smoothing problems in image details reconstruction.
The multi-discriminator generation adversarial network enhancement method is adopted for feature clustering and region adaptation. By inputting the low-resolution image input generator to generate super-resolution images and inputting them with the high-resolution real image for discrimination, the outputs of each discriminator are weighted and fused to optimize the multi-discriminator to generate adversarial networks to improve image quality.
It improves the generalization and stability of the network, enhances feature extraction and recognition capabilities, solves the problem that traditional models are difficult to take into account both global and local features, and improves the detail recovery quality and authenticity of super-resolution images.
Smart Images

Figure CN120087422A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and in particular, to a feature clustering and region - adaptive multi - discriminator generative adversarial network enhancement method (Feature Clustering and Region - Adaptive Multi - Discriminator Generative Adversarial Network Enhancement Method, FCR - ESRGAN). Background Art
[0002] The super - resolution (SISR) technology of images has been widely applied in multiple fields such as medical imaging, remote sensing images, video surveillance, etc. By restoring the details of low - resolution images, the clarity and readability of images are increased. Currently, due to the insufficient learning ability of the network model in the super - resolution technology, the restored high - resolution images often have noise and artifacts.
[0003] In the prior art, image super - resolution mainly relies on interpolation and regularization methods, such as bicubic interpolation and bilinear interpolation, to restore image details by introducing priors and constraints. In recent years, deep - learning - based methods, such as convolutional neural networks (CNNs) and generative adversarial networks (GANs), have become more effective solutions. The super - resolution generative adversarial network (SRGAN) significantly improves the visual quality through perceptual loss; the image restoration model SwinIR combines the advantages of CNNs and Transformers, improving the image restoration efficiency. However, the traditional single - discriminator GAN architecture has limitations in processing images with complex textures, such as limited feature extraction range, single optimization direction, difficulty in balancing the generation quality of multi - modal features, and inability to effectively integrate cross - scale features, resulting in artifacts or over - smoothing problems during image detail reconstruction. Summary of the Invention
[0004] In view of this, the present invention provides a feature clustering and region - adaptive multi - discriminator generative adversarial network enhancement method to solve the limitations of the single discriminator in traditional models and improve the generalization and stability of the network.
[0005] In a first aspect, the present invention provides a feature clustering and region - adaptive multi - discriminator generative adversarial network enhancement method, and the method includes: S1. Input a low - resolution image into a generator to generate a super - resolution image ; S2. The super - resolution image and the high - resolution real image Multiple discriminators are used for joint discrimination. The outputs of each discriminator are weighted and fused respectively through weights to obtain a comprehensive discrimination result; S3. According to the comprehensive discrimination result, use a loss function to calculate the loss to obtain a loss value; S4. Update the parameters according to the loss value and the gradient descent algorithm to optimize the multi-discriminator generative adversarial network; S5. Repeat the above steps until the iteration is completed or the loss value reaches the minimum, obtain the optimal weights, and apply them to the multi-discriminator generative adversarial network. At this time, input the low-resolution image into the generator to obtain the final super-resolution image.
[0006] Optionally, the S1 includes: Input the low-resolution image into the generator. First, extract initial features through a convolutional layer; then perform deep feature extraction through multiple basic blocks; then adopt skip connections to enable effective transmission of shallow information to the deep layer; perform upsampling at the backend of the multi-discriminator generative adversarial network to restore high-resolution features, and perform detail optimization through multiple convolutional layers, and finally output a super-resolution image .
[0007] Optionally, the S2 includes: Improve the single discriminator of the generative adversarial network ESRGAN into a multi-discriminator, which includes a basic discriminator Discriminator1, a feature clustering discriminator Discriminator2, and a feature block discriminator Discriminator3; the basic discriminator is the discriminator in ESRGAN, and the feature clustering discriminator includes feature clustering, a routing mechanism, and an expert network; the feature block discriminator includes regional adaptive processing, a routing mechanism, and an expert network; The process of the basic discriminator includes adopting an architecture based on the convolutional neural network CNN, consisting of multiple convolutional layers and the activation function Leaky ReLU, gradually extracting image features, and discriminating whether the input image is a super-resolution image or a high-resolution real image ; The processes of the feature clustering discriminator and the feature block discriminator include a feature extraction module FEM, a spectral clustering module ISKM, and an image block divider ITS, which respectively process the input super-resolution image and high-resolution real image ; after being fused by the feature routing module FRM, a feature map to is obtained, and passed through n expert sub-networks to Further processing is performed, and finally, the comprehensive discrimination result is obtained through the feature selection module CFS.
[0008] Optionally, the spectral clustering module ISKM includes: ISKM adopts a container filling-based strategy to cluster features and dynamically adjusts the center value of the cluster to make the clustering result more stable and robust. Its expression is: ; Each cluster acts as a container for storing the allocated pixel point features , where represents the number of filled pixels, is the center of the cluster, ; During the clustering process, according to the distance from each cluster center, the new pixel point features are assigned to the nearest cluster and the mean value is updated; when the container is not full, they are directly added and the mean value is updated; when the container is full, the existing pixels are replaced to keep the cluster size stable.
[0009] Optionally, the process of the image block divider ITS includes: ITS extracts local features through a sliding window and reorganizes the data using an extended Unfold operation to enhance the feature expression ability and adapt to subsequent clustering analysis; the input feature map has a shape of , where is the batch size, is the number of channels, is the feature map size; ITS extracts local regions through the window size and the stride ; if then the windows do not overlap, if then there is overlap, which helps to enhance the local feature capture ability; then, ITS adopts an extended Unfold operation to flatten each region into a -dimensional feature vector and converts the entire feature map into a tensor with the shape , where is the total number of windows, and the calculation method is .
[0010] Optionally, the process of the feature routing module FRM includes: FRM assigns the optimal expert sub-network ESN to the feature map processed by ISKM or ITS through a routing matrix to improve the feature expression ability and calculation efficiency; first, the input feature map passes through a learnable weight matrix The routing matrix R is calculated, and the allocation weights of each sub-network are obtained through Softmax normalization; subsequently, according to the maximum matching degree , each feature map is sent to the most suitable expert sub-network for processing; the processed features are aggregated through the feature selection module CFS in the final fusion stage to form global features; among them, the weight matrix is input into the routing mechanism to obtain the routing matrix R, and its expression is: , where NULL represents invalid.
[0011] Optionally, the S3 includes: Using reconstruction loss and perceptual loss to optimize the mapping between the low-resolution image and the super-resolution image to improve the image quality; calculating the perceptual loss between the super-resolution image and the high-resolution real image to ensure the perceptual quality of the generated image; calculating the intra-class feature compactness loss to reduce the intra-class feature dispersion; calculating the inter-class feature dispersion loss to increase the distance between different class features; calculating the expert load consistency loss to balance the task allocation of the expert network; by fusing the previous adversarial loss , constructing the discriminator loss function , and its expression is: ; Among them, , and are weight coefficients; is the feature density loss, including the intra-class feature compactness loss and the inter-class feature dispersion loss , and its expression is: ; and are weight coefficients used to adjust the influence degree of the two parts of the loss.
[0012] Optionally, the intra-class feature compactness loss is used to reduce the dispersion degree of intra-class features, so that the feature points of the same class are more concentrated in the feature space, and its expression is: ; Among them, and represent the feature vectors of the same-class samples, and the Gaussian kernel function is used to measure the similarity of each pair of samples in the loss calculation; the smaller the Euclidean distance , the closer the value of the exponential function is to 1, indicating that the intra-class samples are closer; The inter-class feature dispersion loss is used to increase the feature distance between different classes, so that the feature distributions of different classes are far away from each other, thereby improving the separability of classification. Its expression is: ; Wherein, and represent the normalized feature vectors of different classes. The inter-class feature dispersion loss calculates the cosine similarity of feature points between all classes and takes the average value.
[0013] Optionally, the expert load consistency loss is designed to optimize the load balance of different expert sub-networks ESN during the network allocation process, ensuring that all expert sub-networks can evenly share tasks when processing features. Its expression is: ; Wherein, represents the number of expert sub-networks, represents the total number of input feature blocks, reflects the probability that the th feature block is assigned to the th expert sub-network. is used as a binary indicator to indicate whether the feature block is actually assigned to the expert sub-network; the calculation method of the expert load consistency loss is to first obtain the average allocation weight and the actual allocation situation , then take their product and perform weighted normalization to measure the overall load balance.
[0014] In a second aspect, an embodiment of the present invention provides a computer-readable storage medium. The computer-readable storage medium includes a stored program. Wherein, when the program runs, it controls the device where the computer-readable storage medium is located to execute the feature clustering and region-adaptive multi-discriminator generative adversarial network enhancement method in the first aspect or any possible implementation manner of the first aspect.
[0015] In a third aspect, an embodiment of the present invention provides a device, including: one or more processors; a memory; and one or more computer programs. Wherein, the one or more computer programs are stored in the memory, and the one or more computer programs include instructions. When the instructions are executed by the device, the device is caused to execute the feature clustering and region-adaptive multi-discriminator generative adversarial network enhancement method in the first aspect or any possible implementation manner of the first aspect.
[0016] In the technical solution provided by the present invention, the method includes inputting a low-resolution image into a generator to generate a super-resolution image; jointly inputting the super-resolution image and a high-resolution real image into multiple discriminators for discrimination, and the outputs of each discriminator are respectively weighted and fused through weights to obtain a comprehensive discrimination result; according to the comprehensive discrimination result, using a loss function to calculate the loss to obtain a loss value; updating the parameters according to the loss value and the gradient descent algorithm to optimize the multi-discriminator generative adversarial network; repeating the above steps until the iteration is completed or the loss value reaches the minimum, obtaining the optimal weights, and applying them to the multi-discriminator generative adversarial network. At this time, inputting the low-resolution image into the generator to obtain the final super-resolution image. This method solves the limitation problem of a single discriminator in the traditional model and improves the generalization and stability of the network. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0018] Figure 1 It is a flowchart of the feature clustering and region adaptive multi-discriminator generative adversarial network enhancement method provided by the embodiment of the present invention; Figure 2 It is a structural diagram of the generator provided by the embodiment of the present invention; Figure 3 It is an architecture diagram of the multi-discriminator generative adversarial network provided by the embodiment of the present invention; Figure 4 It is an architecture diagram of the discriminator provided by the embodiment of the present invention; Figure 5 It is a visual comparison example diagram provided by the embodiment of the present invention; Figure 6 It is a schematic diagram of an electronic device provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0019] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.
[0020] It should be clear that the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of protection of the present invention.
[0021] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "the", and "said" used in the embodiments of the present invention are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0022] It should be understood that the term "and / or" used herein is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " in this text generally represents an "or" relationship between the associated objects before and after.
[0023] Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining" or "in response to detecting". Similarly, depending on the context, the phrase "if determined" or "if detected (stated condition or event)" can be interpreted as "when determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)".
[0024] Figure 1 The flowchart of the feature clustering and region adaptive multi-discriminator generative adversarial network enhancement method provided for the embodiments of the present invention is as Figure 1 shown, and the method includes: S1. Input a low-resolution image into the generator to generate a super-resolution image .
[0025] In the embodiments of the present invention, the generator adopts the generator of ESRGAN, which is composed of a basic block (Residual-in-Residual Dense Block, RRDB) structure, and the batch normalization (BN) layer is removed to retain image details and reduce unnecessary computational overhead. RRDB is composed of multiple dense residual blocks (Dense Residual Blocks, DRB) to enhance the feature extraction and reuse capabilities, and at the same time alleviate the problem of gradient disappearance. The generator performs upsampling through sub-pixel convolution or transposed convolution to achieve super-resolution reconstruction. The overall architecture combines deep residual learning to improve the detail quality and perceptual realism of the generated images, making ESRGAN perform excellently in the image super-resolution task.
[0026] In the embodiments of the present invention, as Figure 2 shown, S1 includes: Input the low-resolution image into the generator. First, extract the initial features through the convolutional layer; then perform deep feature extraction through multiple basic blocks to enhance the expression ability of the network; then adopt skip connections to enable the effective transmission of shallow information to the deep layer and improve the training stability; perform upsampling at the backend of the multi-discriminator generative adversarial network to restore the high-resolution features, and perform detail optimization through multiple convolutional layers, and finally output the super-resolution image .
[0027] S2. Input the super-resolution image and the high-resolution real image into the multi-discriminator for discrimination. The outputs of each discriminator are respectively weighted and fused through weights to obtain the comprehensive discrimination result.
[0028] In the embodiments of the present invention, as Figure 3 and Figure 4 shown, S2 includes: Improve the single discriminator of the generative adversarial network ESRGAN into a multi-discriminator, which includes a basic discriminator Discriminator1, a feature clustering discriminator Discriminator2, and a feature block discriminator Discriminator3; the basic discriminator is the discriminator in ESRGAN, and the feature clustering discriminator includes feature clustering, a routing mechanism, and an expert network; the feature block discriminator includes regional adaptive processing, a routing mechanism, and an expert network; In the embodiments of the present invention, the basic discriminator is used for global image quality discrimination to ensure overall visual consistency; the feature clustering discriminator combines MiniBatch K-Means clustering to classify the input features and enhance the feature distinctiveness; the feature block discriminator adopts a sliding window strategy to divide the local area and improve the edge detail restoration ability. The feature clustering discriminator and the feature block discriminator adopt an expert network (Expert Network) and a routing mechanism (Feature Routing Mechanism) for feature optimization. Each expert network focuses on different features (such as texture, edge, color) to improve the multi-modal processing ability. The feature clustering discriminator preprocesses the input data through feature clustering (ISKM) to optimize the training of the expert network, ensure the feature matching degree, and avoid feature mixing. The feature block discriminator adopts regional adaptive processing (ITS), divides the image into small blocks, and independently models the local features to improve the detail restoration ability of complex structure areas and reduce artifacts and over-smoothing problems.
[0029] The process of the basic discriminator includes adopting an architecture based on the convolutional neural network (CNN), which consists of multiple convolutional layers and the activation function Leaky ReLU, gradually extracting image features, and discriminating whether the input image is a super-resolution image through a fully connected layer and the activation function Sigmoid or a high-resolution real image ; The processes of the feature clustering discriminator and the feature block discriminator include a feature extraction module (FEM), a spectral clustering module (ISKM), and an image tiler (ITS), which respectively process the input super-resolution image and the high-resolution real image ; after being fused by the feature routing module (FRM), feature maps to are obtained, and are further processed by n expert sub-networks to , and finally, a comprehensive discrimination result is obtained through the feature selection module (CFS).
[0030] In the embodiment of the present invention, the spectral clustering module (ISKM) includes: ISKM adopts a strategy based on container filling to cluster features, and dynamically adjusts the center value of the cluster to make the clustering result more stable and robust. Its expression is: ; Each cluster acts as a container for storing the allocated pixel point features , where represents the number of filled pixels, is the center of the cluster, ; during the clustering process, according to the distance from each cluster center, the new pixel point features are assigned to the nearest cluster, and the mean value is updated; when the container is not full, it is directly added and the mean value is updated; when the container is full, the existing pixels are replaced to keep the cluster size stable.
[0031] In the embodiment of the present invention, the process of the image tiler module (ITS) includes: ITS extracts local features through a sliding window and reorganizes the data using an extended Unfold operation to enhance the feature expression ability and adapt to subsequent clustering analysis; the input feature map has a shape of , where is the batch size, is the number of channels, is the feature map size; ITS extracts local regions through the window size and the stride , if then the windows do not overlap, if There is overlap, which helps to enhance the ability to capture local features; then, ITS adopts an extended Unfold operation to flatten each region into -dimensional feature vectors and converts the entire feature map into a tensor with the shape , where is the total number of windows, calculated as . This conversion method not only enhances the local feature expression ability but also optimizes the data format, making it more suitable for subsequent spectral clustering modules such as ISKM, improving the robustness and adaptability of clustering, and can be flexibly adapted to different tasks such as super-resolution and object detection by adjusting P and S.
[0032] In the embodiments of the present invention, the process of the feature routing module FRM includes: FRM assigns the optimal expert sub-network ESN to the feature map processed by ISKM or ITS through a routing matrix to improve the feature expression ability and computational efficiency; first, the input feature map passes through a learnable weight matrix to calculate the routing matrix R, and the allocation weights of each sub-network are obtained through Softmax normalization; subsequently, according to the maximum matching degree , each feature map is sent to the most suitable expert sub-network for processing, and the expert sub-network is specifically optimized for different feature types (such as texture, edge, or structural information); the processed features are aggregated through the feature selection module CFS in the final fusion stage to form global features; among them, the weight matrix is input into the routing mechanism to obtain the routing matrix R, and its expression is: , where NULL represents invalid.
[0033] In the embodiments of the present invention, feature clustering includes: The K-means++ initialization algorithm is used to select the initial cluster centers to improve the clustering convergence; the assignment of each feature vector gives priority to its distance from the cluster center and is adjusted according to feature similarity; a loss function is set to make the intra-cluster distance as small as possible and the inter-cluster distance as large as possible, expanding the independence between different features.
[0034] In the embodiments of the present invention, region adaptive processing includes: The input feature map is divided into multiple small blocks according to a predetermined size; each small block includes a local area in the feature map and is sent as an independent input to the subsequent network module for processing.
[0035] In the embodiments of the present invention, the expert network consists of multiple sub-networks and a routing weight. Each sub-network processes a specific type of data, and the routing network is responsible for allocating the input data to the expert sub-networks according to the routing weight.
[0036] In the embodiments of the present invention, in order to constrain the expert sub-networks and two data preprocessing methods, a feature density loss function and a load consistency loss function are set during the training of the network, and the optimal solution to the problem is found by minimizing the loss value.
[0037] S3. According to the comprehensive discrimination result, use the loss function to calculate the loss to obtain the loss value.
[0038] In the embodiments of the present invention, the loss of the generator includes pixel loss , reconstruction loss , adversarial loss , etc., and the loss of the discriminator includes feature density loss , expert load consistency loss , etc.
[0039] In the embodiments of the present invention, S3 includes: Adopt reconstruction loss and perceptual loss to optimize the mapping between the low-resolution image and the super-resolution image to improve the image quality; calculate the perceptual loss between the super-resolution image and the high-resolution real image to ensure the perceptual quality of the generated image; calculate the intra-class feature compactness loss to reduce the intra-class feature dispersion and improve the feature discrimination; calculate the inter-class feature dispersion loss to increase the distance between different class features; calculate the expert load consistency loss to balance the task allocation of the expert network and improve the network stability; by fusing the previous adversarial loss , construct the discriminator loss function , and its expression is: ; Among them, , and are weight coefficients; is the feature density loss, which is used to optimize the feature distribution, make the same-class features more compact and different-class features more dispersed, so as to improve the discrimination ability of the model. It includes the intra-class feature compactness loss and the inter-class feature dispersion loss , and its expression is: ; and is the weight coefficient, which is used to adjust the influence degree of the two parts of losses.
[0040] In the embodiments of the present invention, the intra-class feature compactness loss is used to reduce the dispersion degree of intra-class features, so that the feature points of the same category are more concentrated in the feature space. Its expression is: ; where and represent the feature vectors of samples in the same class. The loss calculation uses the Gaussian kernel function to measure the similarity of each pair of samples; the Euclidean distance is smaller, the value of the exponential function is closer to 1, indicating that the intra-class samples are closer; since this term takes a negative value, the model will tend to minimize the loss during the optimization process, so that the intra-class features are more compact. Intuitively, this loss term encourages the sample features of the same category to be close to the center, forming a more compact intra-class distribution, reducing the discreteness of features, and avoiding the impact of feature dispersion on the classification performance.
[0041] The inter-class feature dispersion loss is used to increase the feature distance between different classes, so that the feature distributions of different classes are far away from each other, thereby improving the separability of classification and avoiding the impact of feature overlap on the classification performance. Its expression is: ; where and represent the normalized feature vectors of different classes. The inter-class feature dispersion loss calculates the cosine similarity of feature points between all classes and takes the average value. Since it is desired that the features of different classes are as different as possible, the optimization objective is to make the value of as small as possible, that is, the included angle between features of different classes is as large as possible. Intuitively, this loss term encourages the features of different classes to be as far away from each other as possible in the feature space, reducing the overlap of inter-class features, thereby improving the discriminant ability of the model.
[0042] In the embodiments of the present invention, the expert load consistency loss (Load Equilibrium Loss, abbreviated as ) aims to optimize the load balance of different expert sub-networks ESN during the network allocation process, ensure that all expert sub-networks can evenly share tasks when processing features, and avoid over-concentration or waste of computing resources. This loss adjusts the feature allocation strategy by measuring the task allocation situation of different expert networks, so that the task loads of all expert sub-networks are balanced, improving the computing efficiency and stability of the network. Its expression is: ; where represents the number of expert sub-networks, represents the total number of input feature blocks, reflects the The probability that a feature block is assigned to the th expert sub-network, as a binary indicator to indicate whether the feature block is actually assigned to the expert sub-network; the calculation method of the expert load consistency loss is to first obtain the average assignment weight of each expert and the actual assignment situation , then take their product, and perform weighted normalization to measure the overall load balance. The optimization goal of this loss is to minimize it, so that each expert sub-network can obtain an equal task assignment, avoiding some experts being overloaded while others are not fully utilized. The introduction of the load consistency loss helps to improve the stability of the network, improve the computational efficiency of the model, and enhance the synergy of different sub-networks.
[0043] S4. Update the parameters according to the loss value and the gradient descent algorithm to optimize the multi-discriminator generative adversarial network.
[0044] In the embodiments of the present invention, the generator optimizes the parameters according to the adversarial loss feedback by the discriminator, which is used to improve the quality of the super-resolution image ; the discriminator is updated simultaneously to improve the discrimination ability for the super-resolution image and the high-resolution real image for image reconstruction.
[0045] In the embodiments of the present invention, the update, that is, the training process, includes: inputting the low-resolution image into the generator to generate a super-resolution image ; inputting the super-resolution image and the high-resolution real image into the multi-discriminator for discrimination at the same time. The outputs of each discriminator are weighted and fused respectively through weights to obtain a comprehensive discrimination result; according to the comprehensive discrimination result, the loss is calculated using the loss function, and the weights of the multi-discriminator generative adversarial network are updated through the gradient descent algorithm; repeat the above steps until the number of training rounds of the network reaches the threshold or the loss value reaches the minimum.
[0046] S5. Repeat the above steps until the iteration is completed or the loss value reaches the minimum, obtain the optimal weights, and apply them to the multi-discriminator generative adversarial network. At this time, input the low-resolution image into the generator to obtain the final super-resolution image.
[0047] In the embodiments of the present invention, the final super-resolution image is output as the final result, realizing high-quality image magnification and detail restoration.
[0048] In the embodiments of the present invention, the training set uses the DIV2K dataset (800 high-quality images), and low-resolution images (LR) are generated by bicubic downsampling (×4). During training, it is randomly cropped to 128×128 and data augmentation is performed. The test sets include Set5, Set14, BSD100, and Urban100. PSNR and SSIM are used as pixel-level metrics for model evaluation, and LPIPS and DISTS are used to evaluate the perceptual quality to comprehensively measure the reconstruction effect and adaptability of the model.
[0049] To prevent the discriminator from being too strong and causing the generator to collapse, pre-trained weights are first used and then formal training is carried out. The balance weights of LG and LD are set to (0.01, 1, 0.05) and (1, 10, 0.05) respectively, and the hyperparameters are set to (1, 0.1, 0.1). The model uses the Adam optimizer, and the learning rate starts from and trains for 60,000 iterations. Using a 4060 GPU, it takes about 47 hours.
[0050] In the embodiments of the present invention, the performance of 12 mainstream models is compared (SRCNN and FSRCNN based on the CNN benchmark framework, SRGAN, USRGAN, BSRGAN, ESRGAN based on the GAN benchmark framework, and models based on SwinIR and RRBD) on 4 datasets. As shown in Table 1, the peak signal-to-noise ratio (PSNR) and structural similarity (SSIM) are used as evaluation metrics. The results in Table 1 show that there are significant differences among different methods. The traditional super-resolution models SRCNN and FSRCNN use the MSE loss. Although they have higher scores in PSNR and SSIM, the reconstructed images are overly smooth and lack details. In contrast, GAN methods such as SRGAN and ESRGAN have slightly lower scores, but the visual perception effect is better. The SwirlIR and RRDB series of models perform excellently in PSNR / SSIM metrics through a multi-scale residual dense architecture and Transformer-CNN design, and also take into account the visual quality. The proposed FCR-ESRGAN in the present invention has better PSNR and SSIM than other methods on the Set5, Set14, Mangal109, and Urban100 datasets, thanks to the multi-discriminator architecture and the optimization of complex features by the expert network.
[0051] Table 1 Comparison results of PSNR and SSIM of different models 。
[0052] In the embodiments of the present invention, the visual comparison of the existing methods (HR, SwinlR, BSRGAN, ESRGAN, USRGAN) and FCR-ESRGAN in the classical image super-resolution task is as follows Figure 5 As shown, the images restored by other methods have false textures, such as the distorted stone bricks in the first row, the false textures on the building walls in the second row, and the unnatural stone brick textures in the last row. Through these comparisons, it can be observed that the images restored by other methods may contain unrealistic textures, while the method of the present invention can generate more natural and detailed results, eliminating artifacts and enhancing the realism of texture details, which proves that the multi-discriminator using the expert network has a stronger ability to distinguish fine textures.
[0053] The present invention uses a multi-discriminator to replace the single discriminator of the traditional model, and adds an expert network, a routing mechanism, and two data preprocessing mechanisms to the discriminator, aiming to solve the problem that the single discriminator of the traditional model cannot cope with the complexity and diversity of image features in the real world, and improve the generalization and stability of the model.
[0054] Compared with the prior art, the present invention has the following beneficial effects: (1) The multi-discriminator architecture of the present invention has stronger feature extraction and recognition capabilities, solves the problem that it is difficult for the traditional single discriminator to balance global and local features, and improves the detail restoration quality and authenticity of super-resolution images.
[0055] (2) The feature clustering discriminator (Discriminator2) of the present invention has more accurate feature classification and processing capabilities, solves the problem of the decline in generation quality caused by the mixing of different types of image features, and improves the generalization ability of the model and its adaptability to complex textures.
[0056] (3) The feature block discriminator (Discriminator3) of the present invention has more accurate local area feature modeling capabilities, solves the problems of blurred edge details and loss of high-frequency information, and improves the detail restoration ability of image reconstruction.
[0057] (4) The expert network and routing mechanism of the present invention have the ability to perform adaptive optimization for different feature categories, solve the problem of the single optimization direction of the traditional discriminator when dealing with multi-modal features, and improve the quality and diversity of image super-resolution reconstruction.
[0058] (5) The feature clustering (ISKM) preprocessing of the present invention has the effect of enhancing the independence of input data and optimizing discriminative learning, solves the problem of unstable training caused by inaccurate feature classification, and improves the stability of adversarial training and the consistency of image generation.
[0059] (6) The regional adaptive processing (ITS) of the present invention has a more accurate local region feature restoration ability, solves the problems of local detail loss and insufficient texture reconstruction, and improves the restoration quality of complex image regions (such as edges and high-frequency texture regions).
[0060] In the technical solution provided by the present invention, the method includes inputting a low-resolution image into a generator to generate a super-resolution image; jointly inputting the super-resolution image and a high-resolution real image into multiple discriminators for discrimination, and the outputs of each discriminator are respectively weighted and fused through weights to obtain a comprehensive discrimination result; according to the comprehensive discrimination result, using a loss function to calculate the loss to obtain a loss value; updating the parameters according to the loss value and the gradient descent algorithm to optimize the multi-discriminator generative adversarial network; repeating the above steps until the iteration is completed or the loss value reaches the minimum, obtaining the optimal weights, and applying them to the multi-discriminator generative adversarial network. At this time, input the low-resolution image into the generator to obtain the final super-resolution image. This method solves the limitation problem of a single discriminator in the traditional model and improves the generalization and stability of the network.
[0061] Each step of the embodiments of the present invention can be executed by an electronic device. Among them, the electronic device includes, but is not limited to, mobile phones, tablet computers, portable PCs, desktop computers, etc.
[0062] The embodiments of the present invention provide a computer-readable storage medium. The computer-readable storage medium includes a stored program. Among them, when the program runs, it controls the electronic device where the computer-readable storage medium is located to execute the embodiments of the above-mentioned feature clustering and region adaptive multi-discriminator generative adversarial network enhancement method.
[0063] Figure 6 It is a schematic diagram of an electronic device provided by an embodiment of the present invention. As Figure 6 shown, the electronic device 21 includes: a processor 211, a memory 212, and a computer program 213 stored in the memory 212 and executable on the processor 211. When the computer program 213 is executed by the processor 211, it implements the feature clustering and region adaptive multi-discriminator generative adversarial network enhancement method in the embodiment. To avoid repetition, it will not be elaborated here one by one.
[0064] The electronic device 21 includes, but is not limited to, a processor 211 and a memory 212. Those skilled in the art can understand that Figure 6 this is only an example of the electronic device 21 and does not constitute a limitation on the electronic device 21. It may include more or fewer components than shown in the figure, or combine some components, or different components. For example, the electronic device may also include input / output devices, network access devices, buses, etc.
[0065] The so-called processor 211 may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.
[0066] The memory 212 may be an internal storage unit of the electronic device 21, such as the hard disk or memory of the electronic device 21. The memory 212 may also be an external storage device of the electronic device 21, such as a plug-in hard disk equipped on the electronic device 21, a Smart Media Card (SMC), a Secure Digital (SD) card, a FlashCard, etc. Further, the memory 212 may also include both the internal storage unit and the external storage device of the electronic device 21. The memory 212 is used to store computer programs and other programs and data required by the network device. The memory 212 may also be used to temporarily store data that has been output or is to be output.
[0067] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems, devices, and units can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0068] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.
Claims
1. A feature clustering and region-adaptive multi-discriminator generative adversarial network enhancement method, characterized in that: The method comprises: S1, low resolution image Input to the generator to generate super-resolution images ; S2, super-resolution image and high-resolution real images The multiple discriminators are input together for discrimination, and the outputs of each discriminator are weighted and fused to obtain a comprehensive discrimination result; S3. According to the comprehensive discrimination result, the loss is calculated using the loss function to obtain the loss value; S4. Update parameters according to the loss value and gradient descent algorithm to optimize the multi-discriminator generative adversarial network; S5. Repeat the above steps until the iteration is completed or the loss value reaches the minimum, and the optimal weight is obtained, and it is applied to the multi-discriminator generative adversarial network. At this time, the low-resolution image Input to the generator to get the final super-resolution image.
2. The method according to claim 1, characterized in that The S1 includes: Convert low-resolution images The input is sent to the generator, where the initial features are first extracted through the convolution layer. Then, deep features are extracted through multiple basic blocks. Then, skip connections are used to effectively transfer shallow layer information to the deep layer. Upsampling is performed at the back end of the multi-discriminator generative adversarial network to restore high-resolution features, and details are optimized through multiple convolution layers to finally output a super-resolution image. .
3. The method according to claim 1, characterized in that The S2 includes: The single discriminator of the generative adversarial network ESRGAN is improved into multiple discriminators, which include basic discriminator Discriminator1, feature clustering discriminator Discriminator2, and feature block discriminator Discriminator3; the basic discriminator is the discriminator in ESRGAN, the feature clustering discriminator includes feature clustering, routing mechanism, and expert network; the feature block discriminator includes regional adaptive processing, routing mechanism, and expert network; The basic discriminator process includes adopting a convolutional neural network (CNN)-based architecture, which consists of multiple convolutional layers and activation functions Leaky ReLU, gradually extracting image features, and using fully connected layers and activation functions Sigmoid to discriminate whether the input image is a super-resolution image. or high-resolution real images ; The process of feature clustering discriminator and feature block discriminator includes feature extraction module FEM, spectral clustering module ISKM and image blocker ITS, which respectively perform super-resolution image segmentation on the input image. and high-resolution real images After being fused by the feature routing module FRM, the feature map is obtained. to , and through n expert sub-networks to After further processing, the comprehensive discrimination result is finally obtained through the feature selection module CFS.
4. The method according to claim 3, characterized in that The spectral clustering module ISKM comprises: ISKM uses a container filling-based strategy to cluster features and dynamically adjusts the center value of the cluster to make the clustering result more stable and robust. Its expression is: ; Each cluster As a container, it is used to store the assigned pixel features. ,in Indicates the number of pixels filled. is the center of the cluster, ; During the clustering process, the new pixel features are assigned to the nearest cluster according to the distance from the center of each cluster, and the mean is updated; when the container is not full, it is directly added and the mean is updated; when the container is full, the existing pixels are replaced to keep the cluster size stable.
5. The method according to claim 3, characterized in that: The process of the image blocker ITS includes: ITS extracts local features through sliding windows and reorganizes data using extended Unfold operations to improve feature expression capabilities and adapt to subsequent clustering analysis; input feature map Shape ,in is the batch size, is the number of channels, is the feature map size; ITS is based on the window size and step length Extract local area, if The windows do not overlap if There is overlap, which helps to enhance the local feature capture capability; then, ITS uses the extended Unfold operation to convert each The region is flattened to dimensional feature vector and convert the entire feature map to shape A tensor of is the total number of windows, calculated as .
6. The method according to claim 3, characterized in that The process of the feature routing module FRM includes: FRM uses the routing matrix to assign the optimal expert subnetwork ESN to the feature graph processed by ISKM or ITS to improve the feature expression ability and computational efficiency. First, the input feature graph Through the learnable weight matrix The routing matrix R is calculated, and the allocation weights of each sub-network are obtained through Softmax normalization; then, according to the maximum matching degree , each feature map The features are sent to the most suitable expert sub-network for processing; the processed features are aggregated through the feature selection module CFS in the final fusion stage to form global features; among them, the weight matrix Input into the routing mechanism to get the routing matrix R, which is expressed as: , NULL means invalid.
7. The method according to claim 1, characterized in that The S3 includes: Adopt reconstruction loss and perceptual loss to optimize low-resolution images With super-resolution images Mapping between them is used to improve image quality; calculate super-resolution images With high-resolution real images The perceptual loss between is used to ensure the perceptual quality of the generated image; the compactness loss of the intra-class features is calculated , used to reduce the dispersion of intra-class features; calculate the dispersion loss of inter-class features , used to increase the spacing between features of different categories; calculate expert load consistency loss , used to balance the task allocation of expert networks; by integrating the previous adversarial loss , construct the discriminator loss function , whose expression is: ; in, , and is the weight coefficient; is the feature density loss, including the intra-class feature compactness loss and inter-class feature dispersion loss , whose expression is: ; and is the weight coefficient, which is used to adjust the influence of the two parts of loss.
8. The method according to claim 7, characterized in that The intra-class feature compactness loss is used to reduce the dispersion of intra-class features, so that the feature points of the same category are more concentrated in the feature space. Its expression is: ; in, and The feature vector represents the same type of samples. The loss calculation uses the Gaussian kernel function to measure the similarity of each pair of samples; Euclidean distance The smaller it is, the closer the value of the exponential function is to 1, indicating that the samples within the class are closer; The inter-class feature dispersion loss is used to increase the feature spacing of different categories, so that the feature distribution of different categories is far away, thereby improving the separability of classification. Its expression is: ; in, and Represents the normalized feature vectors of different categories. The inter-class feature dispersion loss calculates the cosine similarity of feature points between all categories and takes the average value.
9. The method according to claim 7, characterized in that: The expert load consistency loss It aims to optimize the load balancing of different expert sub-networks ESN in the network allocation process to ensure that all expert sub-networks can evenly share the tasks during feature processing. The expression is: ; in, represents the number of expert sub-networks, represents the total number of input feature blocks, Reflect the The feature blocks are assigned to The probability of an expert sub-network, As a binary indicator, it is used to indicate whether the feature block is actually assigned to the expert sub-network; the calculation method of expert load consistency loss is to first calculate the average allocation weight of each expert and actual distribution , and then take their product and weight normalize them to measure the overall load balance.
Citation Information
Patent Citations
Image super-resolution method based on generative adversarial network
CN111583109A
Medical image segmentation method, electronic equipment and storage medium
CN117808831A
Port container number identification method and system based on machine learning
CN117809310A
Digital human live broadcast method based on GAN neural network system
CN118658098A
Remote sensing image super-resolution recovery method based on GAN neural network
CN119359540A