Image processing method and system based on generative adversarial network

By constructing a grid-based high-frequency information evaluation and quantification model to screen high-frequency information and combining it with a generative adversarial network for image processing, the problem of high-frequency information loss in the generative adversarial network in super-resolution reconstruction is solved, and the accurate restoration of image details is achieved.

CN120725877AActive Publication Date: 2025-09-30SICHUAN UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511146337.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-15
Publication Date
2025-09-30
Estimated Expiration
2045-08-15

AI Technical Summary

Technical Problem

Existing generative adversarial networks are prone to causing high-frequency information loss during image super-resolution reconstruction, resulting in loss of image details.

Method used

By calculating the high-frequency energy ratio of the image to construct a grid, the complexity of the high-frequency information is quantitatively evaluated, and the high-frequency information grid is screened using the information loss probability model. The high-frequency information is extracted for compressed transmission, and super-resolution reconstruction and fusion are performed in combination with a generative adversarial network.

Benefits of technology

While ensuring the image transmission speed, it accurately restores image details, overcoming the problem of high-frequency detail loss in existing technologies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120725877A_ABST
    Figure CN120725877A_ABST
Patent Text Reader

Abstract

The invention relates to an image processing method and system based on a generative adversarial network, and the method comprises the steps: carrying out the meshing of an image through high-frequency information before the image is compressed and transmitted, and carrying out the quantitative evaluation of the complexity of the high-frequency information of each grid, secondly, judging the loss probability of high-frequency information of each grid after compression and super-resolution reconstruction by using a pre-constructed information loss probability model; if the probability is large, high-frequency information of the corresponding grid is extracted, and during sending, the high-frequency information of the corresponding grid and the compressed image are sent to the target terminal together. And super-resolution reconstruction based on the generative adversarial network is carried out on the target terminal, and high-frequency information in the original image and the reconstructed image are fused together to obtain a fused image. According to the method and the device, the details of the reconstructed image can be accurately restored while the transmission speed of a large batch of file images is ensured, and the defect that high-frequency details are easy to lose during super-resolution reconstruction of the existing generative adversarial network is overcome.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of image processing technology, and in particular to an image processing method and system based on a generative adversarial network. Background Art

[0002] Generative Adversarial Networks (GANs) are a model that generates high-quality data through adversarial training of two neural networks. Their core concept is derived from the zero-sum game theory. They consist of a generator (G) and a discriminator (D), which compete and optimize against each other to ultimately generate realistic data.

[0003] Generative adversarial networks (GANs) are widely used in image super-resolution processing. Low-resolution images have a clear advantage during transmission. Using GANs for super-resolution reconstruction after transmission enables the rapid transmission of large quantities of high-quality images. However, existing GANs for super-resolution reconstruction are prone to losing high-frequency information in the image, resulting in image information loss. Summary of the Invention

[0004] In view of this, the purpose of the present invention is to provide an image processing method and system based on a generative adversarial network to solve the above technical problems.

[0005] In order to achieve the above object, the present invention adopts the following technical solutions: The image processing method based on a generative adversarial network of the present invention comprises the steps of: Acquire an image to be processed and size information of the image to be processed; and preprocess the image to be processed to obtain a preprocessed image; Calculating the high-frequency energy ratio of the pre-processed image, and constructing a grid covering the pre-processed image based on the high-frequency energy ratio and the size information of the image to be processed, wherein the higher the high-frequency energy ratio, the smaller the grid size; Compressing the preprocessed image to obtain a compressed image; and extracting high-frequency information from the preprocessed image to obtain a high-frequency information image, wherein the high-frequency information includes contour features and noise points; Dividing the high-frequency information image based on the grid to obtain a plurality of high-frequency information grids; and quantitatively evaluating the complexity of the high-frequency information in each high-frequency information grid to obtain a quantitative evaluation vector; Determining the information loss probability corresponding to the quantized evaluation vector of each high-frequency information grid based on a pre-constructed information loss probability model, wherein the information loss probability model characterizes the high-frequency information loss probability corresponding to the quantized evaluation vector of the high-frequency information; Taking a high-frequency information grid with an information loss probability higher than a preset probability threshold as a target grid, and extracting the high-frequency information of the target grid; sending the high-frequency information of the target grid and the compressed image to a target terminal; Based on a pre-built generative adversarial network, super-resolution reconstruction is performed on the compressed image in the target terminal to obtain a reconstructed image, and high-frequency information of the target grid is fused with the reconstructed image to obtain a fused image.

[0006] In one embodiment of the present application, preprocessing the image to be processed to obtain a preprocessed image includes: Performing grayscale conversion on the image to be processed to obtain a grayscale image; Performing high-pass filtering on the grayscale image to obtain a filtered image; Contrast enhancement is performed on the filtered image to obtain a preprocessed image.

[0007] In one embodiment of the present application, calculating the high-frequency energy ratio of the pre-processed image and constructing a grid covering the pre-processed image based on the high-frequency energy ratio and size information of the image to be processed includes: performing normalization processing on the preprocessed image to obtain a normalized image; Performing a fast Fourier transform on the normalized image to obtain a normalized frequency domain image; Performing centering processing on the normalized frequency domain image to obtain a centering processed image, wherein a low-frequency component of the centering processed image is located at a center position and a high-frequency component is located at an edge position; Determining a center point of the centralized image, and constructing a high- and low-frequency dividing circle based on a set radius threshold and the center point, wherein the interior of the high- and low-frequency dividing circle is the low-frequency component, and the exterior of the high- and low-frequency dividing circle is the high-frequency component; Calculate the total energy of the centralized image and the high-frequency energy outside the high- and low-frequency dividing circle , and based on the total energy and the high frequency energy Calculate the proportion of high-frequency energy , ; The high frequency energy ratio Compare with the preset ratio threshold and calculate the ratio of high frequency energy When the ratio is greater than the preset threshold, based on the high frequency energy ratio Determine the number of height divisions of the image grid and width division quantity , the number of height divisions and width division quantity The calculation formula is: ; in, It is the preset reference value of the number of divisions in the height direction. is the preset reference value of the number of divisions in the width direction, is the preset benchmark value; Divide the quantity based on the height , the width is divided into a number of And the size information of the image to be processed is used to construct a grid covering the pre-processed image.

[0008] In one embodiment of the present application, extracting high-frequency information from the pre-processed image to obtain a high-frequency information image includes: Performing fast Fourier transform on the preprocessed image to obtain a frequency domain image; Performing high-pass filtering on the frequency domain image based on a pre-built high-pass filter to obtain a filtered image; Perform inverse Fourier transform on the filtered image to obtain a high-frequency information image.

[0009] In one embodiment of the present application, the method for constructing the information loss probability model includes: Acquiring a plurality of image samples, and compressing the plurality of image samples to obtain a plurality of compressed image samples; Inputting the compressed image sample into a pre-built generative adversarial network to obtain a reconstructed sample; Extracting high-frequency information image samples of the image samples and high-frequency information image samples of the reconstructed samples; performing gridding processing on the high-frequency information image samples of the image samples to obtain gridded samples; and performing gridding processing on the high-frequency information image samples of the reconstructed samples to obtain gridded reconstructed samples; and constructing sample pairs based on the gridded samples and the gridded reconstructed samples; Perform quantitative evaluation on multiple grids in the gridded sample to obtain quantitative evaluation vectors of the multiple grids ,in, Represents grid coordinates; For any sample pair, normalize and subtract the two grids with the same grid coordinates to obtain a difference grid image, and calculate the sum of the pixel values ​​of the difference grid image to obtain the difference value representing the difference between the two. ; The quantitative evaluation vector and the difference value Constructing analysis basis vectors ; Taking the quantitative evaluation vector as the clustering benchmark, multiple analysis basis vectors are clustered to obtain multiple vector clusters; Calculate the reference range of multiple quantitative evaluation vectors in each vector cluster, and calculate the probability value of abnormal analysis basic vectors with difference values ​​greater than the difference threshold in each vector cluster accounting for the total number of analysis basic vectors in the vector cluster ; Reference ranges based on multiple quantitative evaluation vectors and probability values ​​corresponding to the reference ranges of multiple quantitative evaluation vectors Construct an information loss probability model.

[0010] In one embodiment of the present application, the complexity of the high-frequency information in each high-frequency information grid is quantitatively evaluated to obtain a quantitative evaluation vector, including: Segment the high-frequency information in each high-frequency information grid to obtain high-frequency information and its types, wherein the types of high-frequency information include contours, noise, textures, and mutation points; Convert the high-frequency information into the mid-frequency domain and calculate the energy of each high-frequency information from the frequency , information entropy , non-zero coefficient ratio and variance , where energy , information entropy The calculation formulas are: ; Where, Represents pixel points in high-frequency information The frequency components of The frequency component of the pixel in the high-frequency information is probability; The energy , the information entropy , the non-zero coefficient ratio and the variance Normalize and get the normalized energy , normalized information entropy , normalized non-zero coefficient ratio and normalized variance ; The normalized energy , the normalized information entropy , the normalized non-zero coefficient ratio and the normalized variance Perform weighted summation to obtain the quantized value of high-frequency information; For any high-frequency information grid, a quantitative evaluation vector is constructed based on the quantized value of the contour, the quantized value of the noise, the quantized value of the texture and the quantized value of the mutation point.

[0011] In one embodiment of the present application, determining the information loss probability corresponding to the quantized evaluation vector of each high-frequency information grid based on a pre-built information loss probability model includes: Substituting the quantized evaluation vector of the high-frequency information in each high-frequency information grid into the information loss probability model; When the quantitative evaluation vector falls into any target reference range in the information loss probability model, the probability corresponding to the target reference range is used as the information loss probability corresponding to the quantitative evaluation vector.

[0012] In one embodiment of the present application, a generative adversarial network is pre-built, including: Obtain a training dataset containing pairs of low-resolution and high-resolution images; Constructing a generator and a discriminator, and defining a loss function; training the generator and the discriminator based on the training data set and the loss function to obtain a generative adversarial network, wherein the loss function includes adversarial loss, content loss, and pixel-level loss.

[0013] In one embodiment of the present application, fusing the high-frequency information of the target grid with the reconstructed image to obtain a fused image includes: The high-frequency information of the target grid is superimposed on the corresponding position of the reconstructed image, and Gaussian blur is performed on the superimposed edge to obtain a fused image.

[0014] This application provides an image processing system based on a generative adversarial network, including: An acquisition module is used to acquire an image to be processed and size information of the image to be processed; and preprocess the image to be processed to obtain a preprocessed image; a gridding module, configured to calculate the high-frequency energy ratio of the pre-processed image and construct a grid covering the pre-processed image based on the high-frequency energy ratio and the size information of the image to be processed, wherein the higher the high-frequency energy ratio, the smaller the grid size; A compression and feature extraction module, configured to compress the pre-processed image to obtain a compressed image; and extract high-frequency information from the pre-processed image to obtain a high-frequency information image, wherein the high-frequency information includes contour features and noise points; a quantization module, configured to divide the high-frequency information image based on the grid to obtain a plurality of high-frequency information grids; and to quantitatively evaluate the complexity of the high-frequency information in each high-frequency information grid to obtain a quantitative evaluation vector; A probability evaluation module is used to determine the information loss probability corresponding to the quantized evaluation vector of each high-frequency information grid based on a pre-built information loss probability model, wherein the information loss probability model characterizes the high-frequency information loss probability corresponding to the quantized evaluation vector of the high-frequency information; a transmission module, configured to take a high-frequency information grid whose information loss probability is higher than a preset probability threshold as a target grid, extract the high-frequency information of the target grid, and transmit the high-frequency information of the target grid and the compressed image to a target terminal; The restoration module is used to perform super-resolution reconstruction on the compressed image in the target terminal based on a pre-built generative adversarial network to obtain a reconstructed image, and to fuse the high-frequency information of the target grid with the reconstructed image to obtain a fused image.

[0015] The beneficial effects of the present invention are as follows: the image processing method and system based on the generative adversarial network of the present invention, before compressing and sending the image, the present application first evaluates the proportion of high-frequency information of the image. If the proportion of high-frequency information is large, it means that there are more contours or details. Therefore, the image is gridded using high-frequency information, and the complexity of the high-frequency information of each grid is quantitatively evaluated. Then, a pre-built information loss probability model is used to judge the probability of loss of high-frequency information of each grid after compression and super-resolution reconstruction. If the probability is large, the high-frequency information of the corresponding grid is extracted, and when sending, the high-frequency information of the corresponding grid is sent to the target terminal together with the compressed image. Super-resolution reconstruction based on the generative adversarial network is performed at the target terminal, and the high-frequency information in the original image is fused with the reconstructed image to obtain a fused image. The present application can simultaneously ensure the transmission speed of large batches of file images while accurately restoring the details of the reconstructed image, overcoming the shortcomings of the existing generative adversarial network super-resolution reconstruction that easily loses high-frequency details. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The present invention will be further described below in conjunction with the accompanying drawings and embodiments: Figure 1 This is a structural diagram of a file transmission system in one embodiment of the present application; Figure 2 is a flowchart of an image processing method based on a generative adversarial network shown in an embodiment of the present application; Figure 3 A schematic diagram showing the comparison of an original image and a high-frequency image in an embodiment of the present application; Figure 4 This is a schematic diagram of information extraction of a target grid in one embodiment of the present application; Figure 5 is a structural diagram of an image processing system based on a generative adversarial network shown in one embodiment of the present application; Figure 6A schematic diagram of the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application is shown. DETAILED DESCRIPTION

[0017] The following describes the embodiments of the present invention through specific examples. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments. The details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the following embodiments and features in the embodiments can be combined with each other unless they conflict.

[0018] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention. Therefore, the drawings only show the layers related to the present invention and are not drawn according to the number, shape and size ratio of the layers in actual implementation. In actual implementation, the type and number of each layer can be changed at will, and the layer layout may also be more complicated.

[0019] In the following description, numerous details are set forth to provide a more thorough explanation of the embodiments of the present invention; however, it is apparent to one skilled in the art that the embodiments of the present invention may be practiced without these specific details.

[0020] Figure 1 This is a structural diagram of a file transmission system in one embodiment of the present application. Figure 1 As shown, the file transfer system includes a compression module, an image analysis and extraction module, and a target terminal. In this application, the input image is analyzed by the image analysis and extraction module to determine whether the input image needs to contain high-frequency information that is easily lost in the subsequent processing process. If so, the high-frequency information is extracted accordingly. The image compression module compresses the input image and sends the extracted high-frequency information and the compressed image together to the target terminal. The generative adversarial network deployed in the target terminal performs super-resolution reconstruction on the compressed image and fuses the reconstructed image with the high-frequency information to obtain a fused image that can accurately restore the image details.

[0021] Figure 2 is a flowchart of an image processing method based on a generative adversarial network shown in one embodiment of the present application. Figure 2 As shown in FIG. 1 , the image processing method based on the generative adversarial network of this embodiment may include steps S210 to S270 : S210, obtaining an image to be processed and size information of the image to be processed; and preprocessing the image to be processed to obtain a preprocessed image; The image to be processed is the original image, the size information is H*W, and the color channel is RGB.

[0022] Before analysis and processing, the original image needs to be preprocessed to reduce processing information and optimize image details to facilitate subsequent analysis and processing. The preprocessing process includes: S211, performing grayscale conversion on the image to be processed to obtain a grayscale image ; Convert a color image to grayscale, removing color information and retaining only brightness information. This is commonly achieved by weighting the RGB channels. Removing color interference creates a clearer brightness distribution, facilitating subsequent edge detection, texture analysis, and other operations.

[0023] S212, the grayscale image Perform high-pass filtering to obtain a filtered image; High-pass filtering can retain high-frequency information (such as edges and textures) and suppress low-frequency information (such as smooth areas).

[0024] S213, performing contrast enhancement on the filtered image to obtain a pre-processed image .

[0025] By adjusting the dynamic range of pixel values, the image is made clearer. This is achieved using histogram equalization.

[0026] S220, calculating the high-frequency energy ratio of the pre-processed image, and constructing a grid covering the pre-processed image based on the high-frequency energy ratio and the size information of the image to be processed, wherein the higher the high-frequency energy ratio, the smaller the grid size; High-frequency information, such as contours, textures, noise, etc., are easily lost in subsequent super-resolution reconstruction. Therefore, the higher the proportion of high-frequency energy in an image, the greater the probability of loss. Therefore, this application first calculates the pre-processed image The proportion of medium and high frequency energy is used to divide the image with a higher proportion of high frequency energy into more detailed parts, thereby increasing the granularity. The grid division process includes: S221, performing normalization processing on the pre-processed image to obtain a normalized image; Scale the pixel values ​​of the preprocessed image to a uniform range (e.g., [0, 1] or [-1, 1]) to eliminate brightness differences between different images. This application uses the maximum-minimum normalization method for normalization.

[0027] S222, performing fast Fourier transform on the normalized image to obtain a normalized frequency domain image; Convert the normalized image from the spatial domain to the frequency domain to separate low-frequency (smooth areas) and high-frequency (edges, textures) components. Decompose the image into different frequency components to facilitate the subsequent separation of high- and low-frequency information.

[0028] S223, performing centralization processing on the normalized frequency domain image to obtain a centralized image, wherein a low-frequency component of the centralized image is located at a central position, and a high-frequency component is located at an edge position; The centering process moves the low-frequency components of the frequency domain image to the center and the high-frequency components to the edge.

[0029] S224, determining a center point of the centralized image, and constructing a high- and low-frequency dividing circle based on a set radius threshold and the center point, wherein the interior of the high- and low-frequency dividing circle is the low-frequency component, and the exterior of the high- and low-frequency dividing circle is the high-frequency component; In a centralized frequency domain image, a circle is constructed with a radius set around the center point. The radius can be adjusted based on task requirements, flexibly adapting to different image characteristics.

[0030] S225, calculating the total energy of the centralized processed image and the high-frequency energy outside the high- and low-frequency dividing circle , and based on the total energy and the high frequency energy Calculate the proportion of high-frequency energy , ; The calculation formula for total energy is: ; Where, Indicates the pixel points in the central processing image The frequency components of High-frequency energy The calculation formula is: ; Represents pixel points in high-frequency information The frequency components of S226, the high frequency energy ratio Compare with the preset ratio threshold and calculate the ratio of high frequency energy When the ratio is greater than the preset threshold, based on the high frequency energy ratio Determine the number of height divisions of the image grid and width division quantity , the number of height divisions and width division quantity The calculation formula is: ; in, It is the preset reference value of the number of divisions in the height direction. is the preset reference value of the number of divisions in the width direction, is the preset benchmark value; High-frequency energy ratio Reflects the richness of image details (such as texture density and edge complexity). When the proportion of high-frequency energy is high, increase the number of divisions to capture more details; otherwise, reduce the number of divisions to reduce computational costs.

[0031] S227, dividing the quantity based on the height , the width is divided into a number of And the size information of the image to be processed is used to construct a grid covering the pre-processed image.

[0032] Finally, a grid is generated to evenly cover the entire image, ensuring that no areas are missed. The grid density is dynamically adjusted according to the image content, taking into account both efficiency and accuracy.

[0033] In this process, frequency domain analysis and energy calculations are used to quickly locate high-frequency areas in the image, avoiding global over-processing. The number of grid divisions is dynamically adjusted based on image content, balancing computing resources and processing performance. Normalization and centering processes reduce noise interference and ensure accurate energy calculations.

[0034] S230, compressing the preprocessed image to obtain a compressed image; and extracting high-frequency information from the preprocessed image to obtain a high-frequency information image, wherein the high-frequency information includes contour features and noise points; The image compression method is JPEG compression based on discrete cosine transform (DCT). The compressed image has lower resolution and data volume for faster transmission.

[0035] The process of extracting high-frequency information includes: S231, performing fast Fourier transform on the preprocessed image to obtain a frequency domain image; The preprocessed image is converted from the spatial domain to the frequency domain to separate the low-frequency and high-frequency components.

[0036] S232, performing high-pass filtering on the frequency domain image based on a pre-built high-pass filter to obtain a filtered image; This application uses a Butterworth high-pass filter (HPF) to filter the signal and ensure smooth transitions and reduce ringing effects.

[0037] S233, performing inverse Fourier transform on the filtered image to obtain a high-frequency information image.

[0038] The filtered frequency domain image is converted back to the spatial domain to obtain an image containing only high-frequency information. Figure 3 is a schematic diagram comparing the original image and the high-frequency image in one embodiment of the present application, as shown in FIG. Figure 3As shown, the upper part is a high-frequency image (an image containing only high-frequency information, including contours, noise, and other information), and the lower part is a grayscale image.

[0039] S240, dividing the high-frequency information image based on the grid to obtain a plurality of high-frequency information grids; and quantitatively evaluating the complexity of the high-frequency information in each high-frequency information grid to obtain a quantitative evaluation vector; The high-frequency information image itself contains a certain amount of data. If all of it is used for transmission, the transmission rate will be slowed down. Therefore, in order to extract the part of the high-frequency information image that is prone to information loss, this application performs grid processing on the high-frequency information and then quantifies the complexity of the high-frequency information in each grid. Generally speaking, the more complex the high-frequency information, the more likely it is to be lost in the subsequent super-resolution reconstruction. Therefore, the process of quantifying the complexity of the high-frequency information includes: S241, segmenting the high-frequency information in each high-frequency information grid to obtain high-frequency information and types of high-frequency information, wherein the types of high-frequency information include contours, noise points, textures, and mutation points; In this embodiment, an edge extraction operator (such as the Canny operator) is used to segment the high-frequency information grid to obtain edge features. A texture segmenter (such as a Gabor filter) is then used to segment the high-frequency information grid to obtain texture features. Statistical analysis or morphological operations are then used to separate the noise regions, such as by smoothing the image with a Gaussian filter to separate the high-frequency noise. The remaining unsegmented high-frequency information may be mutation points. Based on this process, the high-frequency information can be classified.

[0040] Since different types of high-frequency information have different loss probabilities when being reconstructed, this application quantizes each type of high-frequency information in the high-frequency information grid to form a quantized evaluation vector. The specific process is as follows.

[0041] S242, converting the high frequency information into a mid-frequency domain, and calculating the energy of each high frequency information from the frequency , information entropy , non-zero coefficient ratio and variance , where energy , information entropy The calculation formulas are: ; Where, is the total number of pixels in the high-frequency information grid, is the average frequency component of the pixels in the high-frequency information grid, Represents pixel points in high-frequency information The frequency components of The frequency component of the pixel in the high-frequency information is The probability of is determined by statistically analyzing the high-frequency coefficient histogram.

[0042] S243, respectively, the energy , the information entropy , the non-zero coefficient ratio and the variance Normalize and get the normalized energy , normalized information entropy , normalized non-zero coefficient ratio and normalized variance ; Normalization is performed using maximum-minimum normalization to obtain a quantized value in the range of 0-1.

[0043] S244, the normalized energy , the normalized information entropy , the normalized non-zero coefficient ratio and the normalized variance Perform weighted summation to obtain the quantized value of high-frequency information , Indicates the type of high-frequency information.

[0044] Final quantized value for: ; in, is the first weight, is the second weight, is the third weight, The fourth weight.

[0045] S245 , for any high-frequency information grid, construct a quantitative evaluation vector based on the quantized value of the contour, the quantized value of the noise point, the quantized value of the texture, and the quantized value of the mutation point.

[0046] Finally, the grid The corresponding quantitative evaluation vector is , represents the quantized value of the contour, Represents the quantized value of noise, Represents the quantized value of the texture, Indicates the quantitative value of the mutation point.

[0047] S250, determining an information loss probability corresponding to a quantized evaluation vector of each high-frequency information grid based on a pre-constructed information loss probability model, wherein the information loss probability model characterizes a high-frequency information loss probability corresponding to a quantized evaluation vector of the high-frequency information; In this application, the information loss probability is constructed in advance based on the actual performance of the generative adversarial network to characterize the information loss probability corresponding to different evaluation vectors, specifically including: (1) obtaining a plurality of image samples and compressing the plurality of image samples to obtain a plurality of compressed image samples; The compression method here is the same as the compression method described above and will not be repeated here.

[0048] (2) Inputting the compressed image sample into a pre-built generative adversarial network to obtain a reconstructed sample; Generate adversarial network pre-training, the training process may include: Obtain a training dataset containing pairs of low-resolution images [LR] and high-resolution images [HR]; Constructing a generator and a discriminator, and defining a loss function; training the generator and the discriminator based on the training data set and the loss function to obtain a generative adversarial network, wherein the loss function includes adversarial loss, content loss, and pixel-level loss.

[0049] The goal of a generator (G) is to generate samples (e.g., images, text) similar to real data from random noise. The input is a random noise vector (e.g., a normal or uniform distribution). The output is fake data, with the goal of making it as close to the real data distribution as possible.

[0050] The goal of the Discriminator (D) is to distinguish whether the input data comes from a real dataset or fake data generated by the Generator. It takes in real data or generated data and outputs a probability value, representing the probability that the input data is "real."

[0051] After training, the generator can output high-resolution images based on low-resolution images, achieving super-resolution reconstruction of the image.

[0052] (3) extracting high-frequency information image samples of the image samples and high-frequency information image samples of the reconstructed samples; performing gridding processing on the high-frequency information image samples of the image samples to obtain gridded samples; and performing gridding processing on the high-frequency information image samples of the reconstructed samples to obtain gridded reconstructed samples; and constructing sample pairs based on the gridded samples and the gridded reconstructed samples; (4) Quantitatively evaluate multiple grids in the gridded sample to obtain quantitative evaluation vectors of the multiple grids ,in, Represents grid coordinates; Please refer to the previous article for gridding and quantization processing, which will not be repeated here.

[0053] (5) For any sample pair, normalize and subtract the two grids with the same grid coordinates to obtain a difference grid image, and calculate the sum of the pixel values ​​of the difference grid image to obtain a difference value representing the difference between the two. ; For two grids with the same coordinates, normalization and difference (absolute value of the difference) are performed respectively to obtain the difference image between the two. The pixel values ​​of the difference image are summed up, and the total difference obtained is used to measure the difference between the two. The larger the value, the greater the difference between the two. It is a quantitative value reflecting the difference between the two.

[0054] (6) Using the quantitative evaluation vector and the difference value Constructing analysis basis vectors ; Then the grid-based quantitative evaluation vector and difference value Build basic data to analyze variances The relationship between the quantitative evaluation vector and the grid.

[0055] (7) Using the quantitative evaluation vector as the clustering benchmark, cluster multiple analysis basis vectors to obtain multiple vector clusters; In this embodiment, density clustering is used to cluster analysis basic vectors with similar quantitative evaluation vectors into one cluster, thereby obtaining vector clusters corresponding to various typical quantitative evaluation vectors.

[0056] (8) Calculate the reference range of multiple quantitative evaluation vectors in each vector cluster, and calculate the probability value of the abnormal analysis basis vectors in each vector cluster whose difference value is greater than the difference threshold to the total number of analysis basis vectors in the vector cluster. ; The reference range is a reference range that conforms to three times the standard deviation and is constructed based on the mean and standard deviation of each parameter in the quantitative evaluation vector. The quantitative evaluation vector includes four parameters, so the reference range includes four corresponding three times the standard deviation ranges.

[0057] Then, based on the difference threshold, all the data in the cluster are screened to obtain the basic vector for abnormal analysis. The ratio of this part of the vector to the total vector is the corresponding probability. That is to say, when the value of the quantitative evaluation vector falls into one of the reference ranges, after super-resolution reconstruction, the probability of a large difference from the high-frequency information of the original image is the corresponding probability value. .

[0058] (9) Based on the reference ranges of multiple quantitative evaluation vectors and the probability values ​​corresponding to the reference ranges of multiple quantitative evaluation vectors Construct an information loss probability model.

[0059] After obtaining the information loss probability model, the quantitative evaluation vector of the high-frequency information in each high-frequency information grid is substituted into the information loss probability model; when the quantitative evaluation vector falls into any target reference range in the information loss probability model, the probability corresponding to the target reference range is used as the information loss probability corresponding to the quantitative evaluation vector.

[0060] S260, taking a high-frequency information grid with an information loss probability higher than a preset probability threshold as a target grid, extracting high-frequency information of the target grid; and sending the high-frequency information of the target grid and the compressed image to a target terminal; For grids with a high probability of loss, high-frequency information is extracted, and then the high-frequency information is packaged with the compressed image and sent to the target terminal. Figure 4 This is a schematic diagram of information extraction of a target grid in one embodiment of the present application. Figure 4 As shown, the grid to be filled is the grid with a higher probability of loss, and the high-frequency information of the grid with a higher probability of loss is extracted to obtain the positions and high-frequency information of multiple target high-frequency grids.

[0061] S270, performing super-resolution reconstruction on the compressed image in the target terminal based on a pre-built generative adversarial network to obtain a reconstructed image, and fusing the high-frequency information of the target grid with the reconstructed image to obtain a fused image.

[0062] When the target terminal receives the data packet, it parses the compressed image and high-frequency information and uses a generative adversarial network deployed locally or in the cloud to perform super-resolution reconstruction of the compressed image. Because the high-frequency information retains positional information, it can be aligned with the reconstructed image based on this positional information. Gaussian blurring is then performed on the overlaid edges to create a fused image. This avoids noticeable stitching artifacts and creates a more natural-looking stitching.

[0063] The image processing method based on the generative adversarial network of the present invention, before compressing and sending the image, the present application first evaluates the proportion of high-frequency information of the image. If the proportion of high-frequency information is large, it means that there are more contours or details. Therefore, the image is gridded using the high-frequency information, and the complexity of the high-frequency information of each grid is quantitatively evaluated. Then, the pre-built information loss probability model is used to judge the probability of loss of high-frequency information of each grid after compression and super-resolution reconstruction. If the probability is large, the high-frequency information of the corresponding grid is extracted, and when sending, the high-frequency information of the corresponding grid is sent to the target terminal together with the compressed image. And super-resolution reconstruction based on the generative adversarial network is performed at the target terminal, and the high-frequency information in the original image is fused with the reconstructed image to obtain a fused image. The present application can simultaneously ensure the transmission speed of large batches of file images while accurately restoring the details of the reconstructed image, overcoming the shortcomings of the existing generative adversarial network super-resolution reconstruction, which easily loses high-frequency details.

[0064] like Figure 5 As shown, the present application provides an image processing system based on a generative adversarial network, comprising: An acquisition module is used to acquire an image to be processed and size information of the image to be processed; and preprocess the image to be processed to obtain a preprocessed image; a gridding module, configured to calculate the high-frequency energy ratio of the pre-processed image and construct a grid covering the pre-processed image based on the high-frequency energy ratio and the size information of the image to be processed, wherein the higher the high-frequency energy ratio, the smaller the grid size; A compression and feature extraction module, configured to compress the pre-processed image to obtain a compressed image; and extract high-frequency information from the pre-processed image to obtain a high-frequency information image, wherein the high-frequency information includes contour features and noise points; a quantization module, configured to divide the high-frequency information image based on the grid to obtain a plurality of high-frequency information grids; and to quantitatively evaluate the complexity of the high-frequency information in each high-frequency information grid to obtain a quantitative evaluation vector; A probability evaluation module is used to determine the information loss probability corresponding to the quantized evaluation vector of each high-frequency information grid based on a pre-built information loss probability model, wherein the information loss probability model characterizes the high-frequency information loss probability corresponding to the quantized evaluation vector of the high-frequency information; a transmission module, configured to take a high-frequency information grid whose information loss probability is higher than a preset probability threshold as a target grid, extract the high-frequency information of the target grid, and transmit the high-frequency information of the target grid and the compressed image to a target terminal; The restoration module is used to perform super-resolution reconstruction on the compressed image in the target terminal based on a pre-built generative adversarial network to obtain a reconstructed image, and to fuse the high-frequency information of the target grid with the reconstructed image to obtain a fused image.

[0065] The image processing method and system based on the generative adversarial network of the present invention, before compressing and sending the image, the present application first evaluates the proportion of high-frequency information of the image. If the proportion of high-frequency information is large, it means that there are more contours or details. Therefore, the image is gridded using high-frequency information, and the complexity of the high-frequency information of each grid is quantitatively evaluated. Then, a pre-built information loss probability model is used to judge the probability of loss of high-frequency information of each grid after compression and super-resolution reconstruction. If the probability is large, the high-frequency information of the corresponding grid is extracted, and when sending, the high-frequency information of the corresponding grid is sent to the target terminal together with the compressed image. Super-resolution reconstruction based on the generative adversarial network is performed at the target terminal, and the high-frequency information in the original image is fused with the reconstructed image to obtain a fused image. The present application can simultaneously ensure the transmission speed of large batches of file images while accurately restoring the details of the reconstructed image, overcoming the shortcomings of the existing generative adversarial network super-resolution reconstruction that easily loses high-frequency details.

[0066] Figure 6 The following is a schematic diagram showing the structure of a computer system suitable for implementing an electronic device according to an embodiment of the present application. Figure 6 The computer system of the electronic device shown is only an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0067] like Figure 6 As shown, the computer system includes a central processing unit (CPU) 601, which can perform various appropriate actions and processes, such as executing the methods described in the above embodiments, based on programs stored in a read-only memory (ROM) 602 or programs loaded from a storage unit 608 into a random access memory (RAM) 603. RAM 603 also stores various programs and data required for system operation. CPU 601, ROM 602, and RAM 603 are interconnected via a bus 604. An input / output (I / O) interface 605 is also connected to bus 604.

[0068] The following components are connected to the I / O interface 605: an input section 606 including a keyboard, mouse, and the like; an output section 607 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and speakers; a storage section 608 including a hard disk; and a communication section 609 including a network interface card such as a LAN (Local Area Network) card or a modem. The communication section 609 performs communication processing via a network such as the Internet. A drive 610 is also connected to the I / O interface 605 as needed. Removable media 611, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 610 as needed, so that computer programs read from the removable media can be installed in the storage section 608 as needed.

[0069] In particular, according to embodiments of the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program including a computer program for executing the methods illustrated in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication section 609 and / or installed from removable media 611. When executed by the central processing unit (CPU) 601, the computer program performs the various functions defined in the system of the present application.

[0070] It should be noted that the computer-readable medium described in the embodiments of this application may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may, for example, be an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to, an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM), flash memory, optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof. In this application, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying a computer-readable computer program. This propagated data signal may take a variety of forms, including, but not limited to, an electromagnetic signal, an optical signal, or any suitable combination thereof. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. A computer program embodied on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0071] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions and operations of the systems, methods and computer program products according to various embodiments of the present application. Each box in the flowchart or block diagram can represent a module, program segment or part of the code, and the above-mentioned module, program segment or part of the code contains one or more executable instructions for implementing the specified logical functions. It should also be noted that in some alternative implementations, the functions marked in the boxes can also occur in an order different from that marked in the accompanying drawings. For example, two boxes shown in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart and the combination of boxes in the block diagram or flowchart can be implemented using a dedicated hardware-based system that performs the specified functions or operations, or can be implemented using a combination of dedicated hardware and computer instructions.

[0072] The units involved in the embodiments described in this application may be implemented by software or hardware, and the units described may also be set in a processor. In some cases, the names of these units do not constitute limitations on the units themselves.

[0073] Another aspect of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a computer processor, the computer executes the aforementioned method. The computer-readable storage medium may be included in the electronic device described in the above embodiments, or may exist independently and not be incorporated into the electronic device.

[0074] Another aspect of the present application provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the above embodiments.

[0075] The above embodiments are only preferred embodiments for fully illustrating the present application, and the protection scope of the present application is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art based on the present application are within the protection scope of the present application.

Claims

1. An image processing method based on a generative adversarial network, characterized in that: Including steps: Acquire an image to be processed and size information of the image to be processed; and preprocess the image to be processed to obtain a preprocessed image; Calculating the high-frequency energy ratio of the pre-processed image, and constructing a grid covering the pre-processed image based on the high-frequency energy ratio and the size information of the image to be processed, wherein the higher the high-frequency energy ratio, the smaller the grid size; Compressing the preprocessed image to obtain a compressed image; and extracting high-frequency information from the preprocessed image to obtain a high-frequency information image, wherein the high-frequency information includes contour features and noise points; Dividing the high-frequency information image based on the grid to obtain a plurality of high-frequency information grids; and quantitatively evaluating the complexity of the high-frequency information in each high-frequency information grid to obtain a quantitative evaluation vector; Determining the information loss probability corresponding to the quantized evaluation vector of each high-frequency information grid based on a pre-constructed information loss probability model, wherein the information loss probability model characterizes the high-frequency information loss probability corresponding to the quantized evaluation vector of the high-frequency information; Taking a high-frequency information grid with an information loss probability higher than a preset probability threshold as a target grid, and extracting the high-frequency information of the target grid; sending the high-frequency information of the target grid and the compressed image to a target terminal; Based on a pre-built generative adversarial network, super-resolution reconstruction is performed on the compressed image in the target terminal to obtain a reconstructed image, and high-frequency information of the target grid is fused with the reconstructed image to obtain a fused image.

2. The image processing method based on generative adversarial network according to claim 1, characterized in that Preprocessing the image to be processed to obtain a preprocessed image includes: Performing grayscale conversion on the image to be processed to obtain a grayscale image; Performing high-pass filtering on the grayscale image to obtain a filtered image; Contrast enhancement is performed on the filtered image to obtain a preprocessed image.

3. The image processing method based on generative adversarial network according to claim 1, characterized in that: Calculating the high-frequency energy ratio of the pre-processed image, and constructing a grid covering the pre-processed image based on the high-frequency energy ratio and size information of the image to be processed, including: performing normalization processing on the preprocessed image to obtain a normalized image; Performing a fast Fourier transform on the normalized image to obtain a normalized frequency domain image; Performing centering processing on the normalized frequency domain image to obtain a centering processed image, wherein a low-frequency component of the centering processed image is located at a center position and a high-frequency component is located at an edge position; Determining a center point of the centralized image, and constructing a high- and low-frequency dividing circle based on a set radius threshold and the center point, wherein the interior of the high- and low-frequency dividing circle is the low-frequency component, and the exterior of the high- and low-frequency dividing circle is the high-frequency component; Calculate the total energy of the centralized image and the high-frequency energy outside the high- and low-frequency dividing circle , and based on the total energy and the high frequency energy Calculate the proportion of high-frequency energy , ; The high frequency energy ratio Compare with the preset ratio threshold and calculate the ratio of high frequency energy When the ratio is greater than the preset threshold, based on the high frequency energy ratio Determine the number of height divisions of the image grid and width division quantity , the number of height divisions and width division quantity The calculation formula is: ; in, It is the preset reference value of the number of divisions in the height direction. is the preset reference value of the number of divisions in the width direction, is the preset benchmark value; Divide the quantity based on the height , the width is divided into a number of And the size information of the image to be processed is used to construct a grid covering the pre-processed image.

4. The image processing method based on generative adversarial network according to claim 1, characterized in that Extracting high-frequency information from the preprocessed image to obtain a high-frequency information image includes: Performing fast Fourier transform on the preprocessed image to obtain a frequency domain image; Performing high-pass filtering on the frequency domain image based on a pre-built high-pass filter to obtain a filtered image; Perform inverse Fourier transform on the filtered image to obtain a high-frequency information image.

5. The image processing method based on generative adversarial network according to claim 1, characterized in that: The method for constructing the information loss probability model includes: Acquiring a plurality of image samples, and compressing the plurality of image samples to obtain a plurality of compressed image samples; Inputting the compressed image sample into a pre-built generative adversarial network to obtain a reconstructed sample; Extracting high-frequency information image samples of the image samples and high-frequency information image samples of the reconstructed samples; performing gridding processing on the high-frequency information image samples of the image samples to obtain gridded samples; and performing gridding processing on the high-frequency information image samples of the reconstructed samples to obtain gridded reconstructed samples; and constructing sample pairs based on the gridded samples and the gridded reconstructed samples; Perform quantitative evaluation on multiple grids in the gridded sample to obtain quantitative evaluation vectors of the multiple grids ,in, Represents grid coordinates; For any sample pair, normalize and subtract the two grids with the same grid coordinates to obtain a difference grid image, and calculate the sum of the pixel values ​​of the difference grid image to obtain the difference value representing the difference between the two. ; The quantitative evaluation vector and the difference value Constructing analysis basis vectors ; Taking the quantitative evaluation vector as the clustering benchmark, multiple analysis basis vectors are clustered to obtain multiple vector clusters; Calculate the reference range of multiple quantitative evaluation vectors in each vector cluster, and calculate the probability value of abnormal analysis basic vectors with difference values ​​greater than the difference threshold in each vector cluster accounting for the total number of analysis basic vectors in the vector cluster ; Reference ranges based on multiple quantitative evaluation vectors and probability values ​​corresponding to the reference ranges of multiple quantitative evaluation vectors Construct an information loss probability model.

6. The image processing method based on generative adversarial network according to claim 1 or 5, characterized in that: The complexity of the high-frequency information in each high-frequency information grid is quantitatively evaluated to obtain a quantitative evaluation vector, including: Segment the high-frequency information in each high-frequency information grid to obtain high-frequency information and its types, wherein the types of high-frequency information include contours, noise, textures, and mutation points; Convert the high-frequency information into the mid-frequency domain and calculate the energy of each high-frequency information from the frequency , information entropy , non-zero coefficient ratio and variance , where energy , information entropy The calculation formulas are: ; Where, Represents pixel points in high-frequency information The frequency components of The frequency component of the pixel in the high-frequency information is probability; The energy , the information entropy , the non-zero coefficient ratio and the variance Normalize and get the normalized energy , normalized information entropy , the normalized non-zero coefficient ratio and the normalized variance Perform weighted summation to obtain the quantized value of high-frequency information; For any high-frequency information grid, a quantitative evaluation vector is constructed based on the quantized value of the contour, the quantized value of the noise, the quantized value of the texture and the quantized value of the mutation point.

7. The image processing method based on generative adversarial network according to claim 1, characterized in that: The information loss probability corresponding to the quantitative evaluation vector of each high-frequency information grid is determined based on a pre-built information loss probability model, including: Substituting the quantized evaluation vector of the high-frequency information in each high-frequency information grid into the information loss probability model; When the quantitative evaluation vector falls into any target reference range in the information loss probability model, the probability corresponding to the target reference range is used as the information loss probability corresponding to the quantitative evaluation vector.

8. The image processing method based on generative adversarial network according to claim 2, characterized in that: Pre-built Generative Adversarial Networks, including: Obtain a training dataset containing pairs of low-resolution and high-resolution images; Constructing a generator and a discriminator, and defining a loss function; training the generator and the discriminator based on the training data set and the loss function to obtain a generative adversarial network, wherein the loss function includes adversarial loss, content loss, and pixel-level loss.

9. The image processing method based on generative adversarial network according to claim 1, characterized in that: Fusing the high-frequency information of the target grid with the reconstructed image to obtain a fused image, including: The high-frequency information of the target grid is superimposed on the corresponding position of the reconstructed image, and Gaussian blur is performed on the superimposed edge to obtain a fused image.

10. An image processing system based on a generative adversarial network, characterized in that: include: An acquisition module, used to acquire the image to be processed and the size information of the image to be processed; and preprocessing the image to be processed to obtain a preprocessed image; a gridding module, configured to calculate the high-frequency energy ratio of the pre-processed image and construct a grid covering the pre-processed image based on the high-frequency energy ratio and the size information of the image to be processed, wherein the higher the high-frequency energy ratio, the smaller the grid size; A compression and feature extraction module, configured to compress the pre-processed image to obtain a compressed image; and extract high-frequency information from the pre-processed image to obtain a high-frequency information image, wherein the high-frequency information includes contour features and noise points; a quantization module, configured to divide the high-frequency information image based on the grid to obtain a plurality of high-frequency information grids; and to quantitatively evaluate the complexity of the high-frequency information in each high-frequency information grid to obtain a quantitative evaluation vector; A probability evaluation module is used to determine the information loss probability corresponding to the quantized evaluation vector of each high-frequency information grid based on a pre-built information loss probability model, wherein the information loss probability model characterizes the high-frequency information loss probability corresponding to the quantized evaluation vector of the high-frequency information; a transmission module, configured to take a high-frequency information grid whose information loss probability is higher than a preset probability threshold as a target grid, extract the high-frequency information of the target grid, and transmit the high-frequency information of the target grid and the compressed image to a target terminal; The restoration module is used to perform super-resolution reconstruction on the compressed image in the target terminal based on a pre-built generative adversarial network to obtain a reconstructed image, and to fuse the high-frequency information of the target grid with the reconstructed image to obtain a fused image.

Citation Information

Patent Citations

  • Medical ultrasonic image super-resolution reconstruction method based on multi-image fusion

    CN114792287A

  • Terahertz image super-resolution reconstruction method based on generative adversarial network

    CN115358922A

  • Image super-resolution reconstructing

    US20230206396A1