Image processing method and device, electronic equipment and storage medium

By performing wavelet decomposition and upsampling of the original feature map, combined with the generation adversarial network improvement model, the problem of lack of realism in the generated image is solved, and a more stable image processing effect is achieved.

CN120339080APending Publication Date: 2025-07-18SHENZHEN XUMI YUNTU SPACE TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510182933.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-19
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

In the prior art, the generated image results lack realism, especially in portrait training data, there is image blurring and detail loss caused by beauty, and the effect of the existing methods is unstable.

Method used

By performing wavelet decomposition of the original feature map, point sampling sets are generated and upsampled, combined with the generation of adversarial network improvement models, the realism and detail retention of the image are enhanced.

Benefits of technology

It improves the stability and realism of image processing results, reduces information loss, and enhances the real texture and detailed performance of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120339080A_ABST
    Figure CN120339080A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computers, and provides an image processing method and device, electronic equipment and a storage medium. The method comprises the following steps: acquiring an original feature map; performing wavelet decomposition on the original feature map to obtain a plurality of sub-feature maps, and splicing the plurality of sub-feature maps to obtain a down-sampling feature map; a sampling point generator is used for processing the down-sampling feature map, a point sampling set is generated, and the point sampling set is used for defining the coordinate position of each point on the target feature map; and performing up-sampling on the down-sampling feature map according to the point sampling set to obtain a target feature map. The problem that in the prior art, the generated image result is poor in effect is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to an image processing method, apparatus, electronic device, and storage medium. Background Art

[0002] With the rapid development of image generation technology, many currently generated images lack a sense of realism in some aspects or do not conform to natural phenomena in daily experience. For example, in portrait training data, there are situations such as blurred images, beauty filters, and certain detail losses in the diffusion model itself, resulting in the lack of real textures in the images generated by the diffusion model. In the prior art, a realistic LoRA is often superimposed to generate realistic images, but the effect of this method is unstable when generating image results. Summary of the Invention

[0003] In view of this, embodiments of this application provide an image processing method, apparatus, electronic device, and storage medium to solve the problem of poor effect in generating image results in the prior art.

[0004] In the first aspect of the embodiments of this application, an image processing method is provided, including: obtaining an original feature map; performing wavelet decomposition on the original feature map to obtain multiple sub-feature maps, splicing the multiple sub-feature maps to obtain a downsampled feature map; using a sampling point generator to process the downsampled feature map to generate a point sampling set, where the point sampling set is used to define the coordinate positions of each point on the target feature map; and performing upsampling on the downsampled feature map according to the point sampling set to obtain the target feature map.

[0005] In the second aspect of the embodiments of this application, an image processing apparatus is provided, including: an obtaining module configured to obtain an original feature map; a downsampling module configured to perform wavelet decomposition on the original feature map to obtain multiple sub-feature maps, and splice the multiple sub-feature maps to obtain a downsampled feature map; a sampling point generation module configured to use a sampling point generator to process the downsampled feature map to generate a point sampling set, where the point sampling set is used to define the coordinate positions of each point on the target feature map; and an upsampling module configured to perform upsampling on the downsampled feature map according to the point sampling set to obtain the target feature map.

[0006] In the third aspect of the embodiments of this application, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, where the processor implements the steps of the above method when executing the computer program.

[0007] In the fourth aspect of the embodiments of this application, a storage medium is provided, where the storage medium stores a computer program, and the computer program implements the steps of the above method when executed by a processor.

[0008] The beneficial effects of the embodiments of the present application compared with the prior art are as follows: By obtaining the original feature map from the image to be processed, performing wavelet decomposition on the original feature map, sub-feature maps corresponding to various decomposition methods can be obtained. Stitching multiple sub-feature maps together to obtain the downsampled feature map, using the sampling point generator to process the downsampled feature map to generate the point sampling set, and performing upsampling on the downsampled feature map based on this sampling set, the target feature map can be obtained. In this way, by performing wavelet downsampling and dynamic upsampling on the original feature map to obtain the target feature map that meets the target requirements, the stability of the image processing result is ensured, and a realistic image processing result can be obtained. Description of the Drawings

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0010] Figure 1 It is a schematic flowchart of an image processing method provided by an embodiment of the present application;

[0011] Figure 2 It is a schematic flowchart of another image processing method provided by an embodiment of the present application;

[0012] Figure 3 It is a schematic flowchart of yet another image processing method provided by an embodiment of the present application;

[0013] Figure 4 It is a schematic flowchart of still another image processing method provided by an embodiment of the present application;

[0014] Figure 5 It is a schematic structural diagram of an image processing device provided by an embodiment of the present application;

[0015] Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed Embodiments

[0016] In the following description, specific details such as specific system structures and technologies are presented for the purpose of illustration rather than limitation, so as to thoroughly understand the embodiments of the present application. However, those skilled in the art should clearly understand that the present application can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present application.

[0017] Due to reasons such as beautification in portrait training data and certain loss of detailed information in the diffusion model itself, the generated faces by the diffusion model have problems such as being too smooth, lacking skin texture, and insufficient realism of the characters. The general method is often to generate realistic faces by overlaying realistic LoRA. However, when generating specific characters, this method is prone to unstable character generation. Therefore, this application uses the Generative Adversarial Network (GAN) and improves and designs a stable network model based on the Generative Facial Prior - Generative Adversarial Network (GFPGAN).

[0018] It should be understood that in the image processing process of the existing GFPGAN network, a low - quality image can be input and degraded and removed by the degradation removal module for the input low - quality image. Then, the loss between the generated image and the real image is calculated through various loss functions to ensure that the generated high - quality image removes the degradation effect while maintaining the characteristics of the original image. However, the target image obtained by the existing GFPGAN network in processing the realism of human images is too smooth and lacks skin texture. Therefore, this application improves the Unet of the degradation removal module (Degradation Removal) in the existing GFPGAN to increase the realism of the image processing results, such as removing beautification and skin smoothing and enhancing the texture of the face image.

[0019] Next, a method and device for image processing according to an embodiment of the present application will be described in detail with reference to the accompanying drawings.

[0020] Figure 1 It is a schematic flowchart of a method for image processing provided by an embodiment of the present application. As Figure 1 shown, the method for image processing includes:

[0021] S101, obtaining an original feature map.

[0022] Specifically, the original feature map can be obtained from the image to be processed for subsequent processing to obtain a realistic image. Here, the image to be processed can be an image lacking realism, such as a portrait image with added beautification.

[0023] The original feature map can contain basic information of the image to be processed, such as texture, edges, etc. This lays a foundation for subsequent processing, can retain the spatial information and texture characteristics of the image to be processed, and helps to achieve more accurate processing.

[0024] S102. Perform wavelet decomposition on the original feature map to obtain multiple sub-feature maps, and splice the multiple sub-feature maps to obtain a downsampled feature map.

[0025] Specifically, the wavelet transform can be used to decompose the original feature map into multiple sub-feature maps of different scales and frequencies. Next, these sub-feature maps are spliced according to the feature channels to generate a downsampled feature map. In this way, not only the resolution of the original feature map is reduced through downsampling, but also the high-frequency and low-frequency information is retained as much as possible, avoiding information loss caused by direct convolutional downsampling.

[0026] Among them, the wavelet transform can apply the Haar wavelet transform, that is, a transformation method based on a simple rectangular wave, which decomposes the image into low-frequency and high-frequency information of multiple scales. Through the low-frequency and high-frequency components after the Haar wavelet transform, the image can be reconstructed to a certain extent, the detailed information of the image can be reduced, and the overall and smooth structural features can be retained, reducing information loss in the process of processing the realism of portraits and improving the effect of skin retouching removal.

[0027] It should be noted that this step can be applied to the downsampling of the degradation removal module in the GFPGAN network. Through the downsampling of the wavelet transform, multi-level semantic features of the original feature map are extracted, reducing information loss and improving the effect of skin retouching removal in the process of image beauty removal.

[0028] S103. Use a sampling point generator to process the downsampled feature map to generate a point sampling set, which is used to define the coordinate positions of each point on the target feature map.

[0029] Specifically, in order to make up for the high-frequency part that may be lost during the downsampling process, the downsampled feature map can be upsampled. First, a point sampling set can be generated by the sampling point generator, and the point sampling set is used for upsampling from the downsampled feature map, thereby reducing detail loss and over-smoothing, improving the effect of beauty removal, and obtaining a more realistic image. When using the sampling point generator to generate the point sampling set, a constraint factor can be introduced to ensure the reasonable distribution of the sampling points, and the constraint factor can be dynamically adjusted based on the feature information of the input original feature map or the downsampled feature map to obtain a reasonable point sampling set, avoiding information loss caused by unnecessary sampling. Of course, in the application process of the embodiment, a fixed value can also be set based on the specific situation to constrain the moving range of the point sampling set offset, which is not limited here.

[0030] S104. Upsample the downsampled feature map according to the point sampling set to obtain the target feature map.

[0031] Specifically, the downsampled feature map is upsampled using the coordinates in the point sampling set to generate the target feature map. The downsampled feature map obtained during the downsampling process in GFPGAN is input into the upsampling stage through skip connections and fused with the upsampled features. At each step of upsampling, the downsampled feature map is combined to obtain the target feature map. This process is well-known to those skilled in the art and will not be elaborated here. In this way, the spatial resolution of the image can be gradually restored through step-by-step upsampling, and the features extracted in the downsampling stage are combined to generate the final high-resolution output.

[0032] It should also be noted that based on the existing GFPGAN, after sampling the original feature map, local feature enhancement processing is required. The target feature map obtained in this application can be understood as the target feature map for further detail enhancement after wavelet downsampling and upsampling in the improved GFPGAN network. Of course, the solution of this application is not limited to the existing GFPGAN settings, and subsequent processing of the target feature map can be adaptively selected according to the development of network technology, such as directly outputting or undergoing other feature processing, which will not be elaborated here.

[0033] According to the technical solution provided in the above embodiment, by obtaining the original feature map from the image to be processed, performing wavelet decomposition on the original feature map to obtain multiple sub-feature maps, splicing the multiple sub-feature maps to obtain the downsampled feature map, using the sampling point generator to process the downsampled feature map to generate the point sampling set, and performing upsampling on the downsampled feature map according to the point sampling set to obtain the target feature map, the details and structure of the input data can be captured through upsampling, and the downsampled feature map can be processed through upsampling, avoiding the loss of details or over-smoothing of the feature map, and improving the accuracy of the algorithm and the effect of image realism processing.

[0034] In some embodiments, performing wavelet decomposition on the original feature map to obtain multiple sub-feature maps includes: performing convolution on the original feature map in the horizontal and vertical directions using the first filter to obtain the first sub-feature map; performing convolution on the original feature map in the horizontal direction using the second filter and performing convolution on the original feature map in the vertical direction using the first filter to obtain the second sub-feature map; performing convolution on the original feature map in the horizontal direction using the first filter and performing convolution on the original feature map in the vertical direction using the second filter to obtain the third sub-feature map; performing convolution on the original feature map in the horizontal and vertical directions using the second filter to obtain the fourth sub-feature map, and the sub-feature maps include the first sub-feature map, the second sub-feature map, the third sub-feature map, and the fourth sub-feature map.

[0035] Among them, the first filter can represent a low-pass wavelet filter, and the second filter can represent a high-pass wavelet filter.

[0036] Figure 2 It is a schematic flowchart of another image processing method provided by an embodiment of the present application. The following will describe this embodiment in conjunction with Figure 2 this embodiment.

[0037] Please refer to Figure 2 , the downsampling module of Haar wavelet can be used to decompose the original feature map by applying wavelet transform, obtain the corresponding sub-feature maps, splice each sub-feature map to obtain a lossless coding of downsampling, and then learn frequency domain features through 1×1 convolution to obtain a downsampled feature map.

[0038] Furthermore, performing a two-dimensional discrete wavelet first-order decomposition on the input original feature map can obtain an LL sub-map (i.e., the first sub-feature map), an LH sub-map (i.e., the second sub-feature map), an HL sub-map (i.e., the third sub-feature map), and an HH sub-map (i.e., the fourth sub-feature map). The input original feature map can be a feature map of size h×w×c, where h represents the height of the original feature map, w represents the width of the original feature map, and c can represent the number of channels of the original feature map.

[0039] Among them, the first sub-feature map is an approximate representation of the original feature map and is the wavelet coefficients obtained by convolving with the first filter in the horizontal and vertical directions.

[0040] The second sub-feature map represents the horizontal direction detail sub-map of the original image and can be used to highlight the singular characteristics in the horizontal direction of the image. It is the wavelet coefficients obtained by convolving with the second filter in the horizontal direction and then convolving with the first filter in the vertical direction;

[0041] The third sub-feature map can represent the vertical direction detail sub-map of the original image and is used to highlight the singular characteristics in the vertical direction of the image. It is the wavelet coefficients obtained by convolving with the first filter in the horizontal direction and then convolving with the second filter in the vertical direction;

[0042] The fourth sub-feature map can represent the diagonal direction detail sub-map of the original image and is used to highlight the diagonal edge characteristics of the image. It is the wavelet coefficients obtained by convolving with the second filter in both the horizontal and vertical directions.

[0043] According to the technical solution provided by the above embodiment, wavelet transform can be applied to decompose and downsample the original feature map, which reduces the detail loss that may be caused by direct convolution downsampling. Splicing the decomposed sub-feature maps to obtain a downsampled feature map reduces the resolution of the original feature map and retains high and low frequency information as much as possible, avoiding information loss caused by direct convolution downsampling, enhancing the sensitivity of the model to details, and providing more effective information for restoring details in the upsampling stage.

[0044] In some embodiments, a downsampled feature map is processed by a sampling point generator to generate a point sampling set, including: dynamically determining a constraint factor according to the downsampled feature map, and determining a sampling set offset corresponding to the downsampled feature map based on the constraint factor, where the constraint factor is used to constrain the sampling set offset; adding the sampling set offset and an original sampling grid to obtain a point sampling set, and the original sampling grid represents a basic sampling coordinate.

[0045] First, it should be noted that this embodiment can be understood as improving the Unet upsampling in the degradation removal module of GFPGAN to dynamic upsampling, that is, introducing a dynamic constraint factor in the sampling point generator so that the network can adaptively adjust the upsampling strategy according to the data features of the input image to be processed, in order to better capture the details and structure of the input data.

[0046] Specifically, the constraint factor can be calculated from the downsampled feature map. After determining the constraint factor, the sampling set offset of the downsampled feature map is constrained based on this constraint factor. The sampling set offset can represent the moving range of the offset. By constraining with the constraint factor, the disorder of the sampling points can be avoided, and the flexibility of the sampling set offset is increased by dynamically determining the constraint factor, thereby increasing the flexibility of subsequent upsampling. It should be noted that the constraint factor can be adaptively adjusted with the change of the input features, that is, the original feature map and the downsampled feature map. For example, features such as the shape, edge, texture, and set structure of the object represented by the input features may all affect the size of the constraint factor. That is, different input features will generate different offset constraints. It should be understood that the constraint factor corresponds to a fixed value in different situations, but generally the constraint factor can be controlled within a range, for example, normalizing the constraint factor to the range of 0-1.

[0047] Furthermore, adding the sampling set offset and the original sampling grid can obtain a point sampling set for upsampling.

[0048] According to the technical solution provided by the above embodiment, the constraint factor can be dynamically determined according to the input feature map, and the sampling set offset is based on the constraint factor, ensuring the rationality and orderliness of the sampling point distribution. The point sampling set can dynamically adapt to the distribution of the input features, so as to generate a higher-quality target feature map in the subsequent upsampling process, improving the detail retention and image realism.

[0049] In some embodiments, dynamically determining a constraint factor according to a downsampled feature map includes: performing a convolution process on the downsampled feature map to obtain a basic offset; normalizing the basic offset to obtain a normalized basic offset, and multiplying the normalized basic offset by a preset value to obtain a constraint factor corresponding to the downsampled feature map.

[0050] Specifically, a convolution operation can be performed on the downsampled feature map to calculate the basic offset. Through the convolution operation, local information in the downsampled feature map can be extracted to generate the initial offset, and these initial offsets can represent the offset relationship from the original sampling grid to the final sampling points.

[0051] Among them, the basic offset, as the initial displacement information, provides input for subsequent normalization and dynamic constraints to ensure that the offset adjustment can be adaptively generated based on the input features. Normalization can limit the value of the basic offset within a fixed range, which can be achieved through the sigmoid activation function to limit the value within the range of [0, 1]. In this way, a more stable numerical range can be obtained, and unstable sampling caused by excessive offsets can be avoided.

[0052] Furthermore, a preset value (such as 0.5) is used as a scaling factor to further adjust the normalized basic offset and limit it within the range of [0, preset value], thereby obtaining the amplitude for controlling the offset and limiting the displacement range of the sampling points.

[0053] According to the technical solution provided in the above embodiment, by performing convolution on the downsampled feature map to extract local information, initial displacement data is provided for subsequent sampling point adjustment. By normalizing the basic offset, the range of the offset is limited, the stability and controllability of the offset are enhanced, and the normalized basic offset is further scaled to dynamically control the amplitude of the sampling point offset, improving the rationality and flexibility of the sampling point distribution. By adaptively adjusting the upsampling strategy through the input feature map, the details and structure of the input can be better captured.

[0054] In some embodiments, determining the sampling set offset corresponding to the downsampled feature map based on the constraint factor includes: multiplying the constraint factor by the basic offset to obtain the initial offset; performing pixel shuffling on the initial offset to obtain the sampling set offset.

[0055] Specifically, after determining the constraint factor, the basic offset can be restricted by multiplication to obtain the initial offset, so as to dynamically adjust the amplitude of the offset and ensure that the offset value is within a reasonable range to obtain the initial offset. In this way, the initial offset can be dynamically generated according to the input feature map, which can adapt to the characteristics of different inputs and further improve the ability to model local details.

[0056] Among them, the determination of the constraint factor and the initial offset can be achieved through two branches. That is, the basic offset is obtained by performing convolution on the downsampled feature map through the first branch network, and the constraint factor is obtained by performing convolution on the downsampled feature map through the second branch network and then multiplying the obtained basic offset by the preset value. Further, the basic offset of the first branch network can be multiplied by the constraint factor of the second branch network to obtain the initial offset.

[0057] Further, the initial offset can be pixel-shuffled, that is, the initial offset can be redistributed into a coordinate grid with a higher spatial resolution to generate a final sampling set offset.

[0058] According to the technical solution provided by the embodiments of the present application, by multiplying the constraint factor by the base offset, an initial offset is obtained, and pixel-shuffling the initial offset can map the initial offset to a grid with a higher resolution, generating a finer sampling point distribution, which helps to more reasonably distribute the sampling points to ensure that the generated sampling set offset can accurately describe the position coordinates of the target feature map and avoid information loss caused by resolution conversion.

[0059] Figure 3 is a schematic flowchart of another image processing method provided by the embodiments of the present application. The following will be further described in conjunction with Figure 3 to further illustrate the above embodiments.

[0060] Please refer to Figure 3 , which can be understood as the generation process of the point sampling set.

[0061] First, input the downsampled feature map X of size h×w×c, and process X through two branches, namely the first branch network linear1 and the second branch network linear2, respectively, to obtain a base offset of size (h×w×2gs 2 ). Normalize and scale the base offset in the second branch network to map the convolution output to the range [0, 0.5] to obtain a constraint factor, and multiply the constraint factor by the base offset of the first branch network to obtain an initial offset of (h×w×2gs 2 ), where 2gs 2 represents the offsets of g sampling points in their horizontal and vertical directions. Pixel-shuffle the initial offset to convert the offset size from (h×w×2gs 2 ) to (sh×sw×2g) to obtain the sampling set offset O, where sh = s×s and sw = w×s: representing the upsampled spatial size. Here, G represents the original sampling grid, that is, an initial regular grid representing the positions of regular sampling points during the upsampling process.

[0062] Further, adding the sampling set offset obtained after pixel-shuffling (piexl_shuffle) to the original grid can generate a point sampling set sampling_set, which can contain the sampling coordinates at each position after upsampling.

[0063] The generation process of the point sampling set can be referred to the following formula:

[0064] sampling_set = piexl_shuffle(0.5sigmoid(linear2(X))linear1(X)) + G

[0065] Among them, 0.5sigmoid(linear2(X)) is a dynamic constraint factor, which is used to constrain the moving range of the offset, avoid the disorder of sampling points, and avoid poor sampling effects. And this dynamic constraint factor is dynamic, which can increase the flexibility of the offset and the flexibility of upsampling. linear1 and linear2 are 1×1 convolutions. O is the sampling set offset, G is the original sampling grid (providing the initial position of the sampling point is the reference coordinate), and pixel shuffle is pixel shuffling, which changes the feature map with a size of h×w×2gs 2 to sh×sw×2g, making the feature map larger and the number of channels smaller.

[0066] In this way, the sampling set offset can be dynamically generated, the range of the offset can be constrained by the dynamically determined constraint factor, the sampling set offset can be dynamically generated, ensuring a reasonable distribution of sampling points. By pixel shuffling, the spatial resolution of the feature map is adjusted, enabling the network to better capture the geometric information in the downsampled feature map.

[0067] In some embodiments, upsampling the downsampled feature map according to the point sampling set to obtain the target feature map includes: performing upsampling on the downsampled feature map according to the point sampling set through bilinear interpolation to obtain the target feature map.

[0068] Specifically, after obtaining the point sampling set, the downsampled feature map can be upsampled based on this point sampling set. During the sampling process, the coordinates in the point sampling set are corresponded to the downsampled feature map, and the corresponding feature values are extracted. In this way, through the redistribution of sampling points, the target feature map can be obtained.

[0069] During the process of sampling point redistribution, the sampling points in the point sampling set may not be completely aligned with the discrete network points of the downsampled feature map. For example, the sampling points in the point sampling set may fall at a certain position between two grids. Therefore, in this application, the bilinear interpolation method can be used to calculate the value of this sampling point to obtain the final value of the sampling point, thereby improving the target feature map.

[0070] According to the technical solutions provided in the above embodiments, through difference calculation, feature values can be accurately extracted at non-discrete grid point positions. Through the interpolation method, the smoothness during the sampling process is ensured, the accuracy, continuity, and smoothness of the target feature map are improved, the errors caused by sampling point offset and grid misalignment are avoided, and at the same time, the effect of feature map reconstruction is improved.

[0071] Figure 4It is a schematic flowchart of another image processing method provided by an embodiment of the present application. The following will further explain the present application in conjunction with Figure 4 to make a further description of the present application.

[0072] Figure 4 It can be regarded as the overall process of dynamic upsampling (DySample), that is, the input upsampled feature map X is processed by a sampling point generator (i.e., Sampling point generator) to obtain a point sampling set sampling_set of size h×w×2g, and then the generation of the target feature map Xup is realized through the grid_sample function in the Pytorch framework, and the target feature map with an input size of sh×sw×c is input.

[0073] This upsampling process can be expressed by the following formula:

[0074] Xup = grid_sample(X, sampling_set)

[0075] In this article, the same symbols represent the same meanings, and the formula symbols will not be explained here. Through such an upsampling method, the image processing process can be achieved. Compared with the upsampling methods in the prior art, the Dysample upsampling has the following advantages:

[0076] 1. Adaptability: Dysample sampling can adaptively adjust the upsampling strategy according to the characteristics of the input data. This means that it can apply different upsampling methods in different regions to better capture the details and structure of the input data.

[0077] 2. Higher precision: Since Dysample can make different processing decisions for different input data, it can usually provide higher reconstruction precision in complex scenarios. For regions with more details such as edges and textures, dynamic upsampling can choose a more refined processing method to improve the quality of the feature map.

[0078] 3. Reducing over-smoothing: Dysample avoids the loss of details or over-smoothing of the feature map and improves the algorithm precision.

[0079] 4. Trainability: In a deep learning model, Dysample is implemented through a trainable module, which means that it can optimize its upsampling process in a data-driven manner to further improve the performance and effect.

[0080] In this way, the face prior knowledge can be effectively utilized by GFPGAN, and a deringing training method can be adopted to enhance the realism of the face. And the Unet downsampling of the degradation removal module is improved based on the wavelet transform method, and the Unet upsampling of the degradation removal module is improved based on the Dysample method.

[0081] In some embodiments, after upsampling the downsampled feature map according to the point sampling set to obtain the target feature map, the following steps are further included: mapping the target feature map to the target channel format to generate an image processing result, presenting the image processing result on a visualization interface; receiving an interaction operation for the image processing result, and adjusting the image processing result based on the interaction operation.

[0082] Specifically, the target feature map can represent a feature map containing required feature information. After obtaining the high-resolution target feature map, the target feature map can be visualized to obtain an image processing result, and the image processing result is presented on the visualization interface.

[0083] In the actual application process, the target feature map may have multiple channels. Therefore, the target feature map can be first converted into a 3-channel (RGB) or 1-channel (grayscale) map. Further, the feature values are normalized to the image display range, and then the image processing result can be obtained. Furthermore, the generated image file can be saved or directly output in combination with the user interface tool or the terminal environment. After output, the processing effect of the model can be intuitively presented through the visualization interface, which is convenient for the user to further optimize the model or modify the processing flow according to the image processing result.

[0084] Receiving the interaction operation of the user for the target processed image can be understood as the user's manual intervention or adjustment of the image processing result, such as modifying the brightness, contrast, color saturation, etc., to further correct the detail problems of the image.

[0085] According to the technical solution provided in the above embodiments, by mapping the target feature map to the target channel, it can be ensured that the image processing result can be intuitively viewed, and accepting the interaction operation can further optimize the image effect. This not only ensures that the output result of the model can be intuitively perceived, but also improves the application value of the image through interactive adjustment, providing greater flexibility for the image processing tasks in the real scenario.

[0086] All the above optional technical solutions can be combined arbitrarily to form the optional embodiments of the present application, which will not be elaborated herein one by one.

[0087] The following is the device embodiment of the present application, which can be used to execute the method embodiment of the present application. For the details not disclosed in the device embodiment of the present application, please refer to the method embodiment of the present application.

[0088] Figure 5 is a schematic diagram of an image processing device provided by an embodiment of the present application. As Figure 5 shown, the image processing device includes:

[0089] An acquisition module 501, configured to acquire an original feature map;

[0090] The downsampling module 502 is configured to perform wavelet decomposition on the original feature map to obtain multiple sub-feature maps, and splice the multiple sub-feature maps to obtain a downsampled feature map;

[0091] The sampling point generation module 503 is configured to process the downsampled feature map using a sampling point generator to generate a point sampling set, and the point sampling set is used to define the coordinate positions of each point on the target feature map;

[0092] The upsampling module 504 is configured to upsample the downsampled feature map according to the point sampling set to obtain the target feature map.

[0093] In some embodiments, the downsampling module 502 is specifically configured to perform convolution on the original feature map in the horizontal and vertical directions using a first filter to obtain a first sub-feature map; perform convolution on the original feature map in the horizontal direction using a second filter and perform convolution on the original feature map in the vertical direction using the first filter to obtain a second sub-feature map; perform convolution on the original feature map in the horizontal direction using the first filter and perform convolution on the original feature map in the vertical direction using the second filter to obtain a third sub-feature map; perform convolution on the original feature map in the horizontal and vertical directions using the second filter to obtain a fourth sub-feature map, and the sub-feature maps include the first sub-feature map, the second sub-feature map, the third sub-feature map, and the fourth sub-feature map.

[0094] In some embodiments, the sampling point generation module 503 is specifically configured to dynamically determine a constraint factor according to the downsampled feature map, and determine a sampling set offset corresponding to the downsampled feature map based on the constraint factor, and the constraint factor is used to constrain the sampling set offset; add the sampling set offset and the original sampling grid to obtain the point sampling set, and the original sampling grid represents the basic sampling coordinates.

[0095] In some embodiments, the sampling point generation module 503 is specifically configured to perform convolution processing on the downsampled feature map to obtain a basic offset; normalize the basic offset to obtain a normalized basic offset, and multiply the normalized basic offset by a preset value to obtain a constraint factor corresponding to the downsampled feature map.

[0096] In some embodiments, the sampling point generation module 503 is specifically configured to multiply the constraint factor by the basic offset to obtain an initial offset; perform pixel shuffling on the initial offset to obtain the sampling set offset.

[0097] In some embodiments, the upsampling module 504 is specifically configured to perform bilinear interpolation to upsample the downsampled feature map according to the point sampling set to obtain the target feature map.

[0098] In some embodiments, the upsampling module 504 is specifically configured to map the target feature map to the target channel format, generate an image processing result, display the image processing result on the visualization interface; receive an interaction operation for the image processing result, and adjust the image processing result based on the interaction operation.

[0099] It should be understood that the sequence numbers of the steps in the above embodiments do not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0100] Figure 6 is a schematic diagram of the electronic device 6 provided by the embodiments of the present application. As Figure 6 shown, the electronic device 6 in this embodiment includes: a processor 601, a memory 602, and a computer program 603 stored in the memory 602 and executable on the processor 601. When the processor 601 executes the computer program 603, the steps in the above-mentioned method embodiments are implemented. Alternatively, when the processor 601 executes the computer program 603, the functions of each module / unit in the above-mentioned device embodiments are implemented.

[0101] The electronic device 6 may be a desktop computer, a notebook, a palm computer, a cloud server, or other electronic devices. The electronic device 6 may include, but is not limited to, the processor 601 and the memory 602. Those skilled in the art can understand that Figure 6 merely examples of the electronic device 6, and do not constitute a limitation to the electronic device 6. It may include more or fewer components than shown in the figure, or different components.

[0102] The processor 601 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc.

[0103] The memory 602 can be an internal storage unit of the electronic device 6. For example, it can be the hard disk or memory of the electronic device 6. The memory 602 can also be an external storage device of the electronic device 6. For example, it can be a plug-in hard disk equipped on the electronic device 6, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. The memory 602 can also include both the internal storage unit and the external storage device of the electronic device 6. The memory 602 is used to store computer programs and other programs and data required by the electronic device.

[0104] Those skilled in the art can clearly understand that for the convenience and simplicity of description, only the above division of each functional unit and module is used as an example. In actual applications, the above functions can be allocated to different functional units and modules according to needs, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0105] If the integrated module / unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, to implement all or part of the processes in the above embodiment methods of this application, it can also be completed by instructing related hardware through a computer program. The computer program can be stored in the readable storage medium. When the computer program is executed by a processor, it can implement the steps of each of the above method embodiments. The computer program can include computer program code, and the computer program code can be in the form of source code, object code, executable file, or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a Read-Only Memory (ROM), a Random Access Memory (RAM), an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0106] The above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the present application in various embodiments, and should all be included within the protection scope of the present application.

Claims

1. An image processing method, characterized in that, Including: Obtain the original feature map; Perform wavelet decomposition on the original feature map to obtain multiple sub-feature maps, and splice the multiple sub-feature maps to obtain a downsampled feature map; Use a sampling point generator to process the downsampled feature map to generate a point sampling set, and the point sampling set is used to define the coordinate positions of each point on the target feature map; Upsample the downsampled feature map according to the point sampling set to obtain the target feature map.

2. The method according to claim 1, characterized in that, The performing wavelet decomposition on the original feature map to obtain multiple sub-feature maps includes: Use a first filter to perform convolution in the horizontal and vertical directions of the original feature map to obtain a first sub-feature map; Use a second filter to perform convolution in the horizontal direction of the original feature map, and perform convolution in the vertical direction of the original feature map through the first filter to obtain a second sub-feature map; Use a first filter to perform convolution in the horizontal direction of the original feature map, and perform convolution in the vertical direction of the original feature map through the second filter to obtain a third sub-feature map; Use a second filter to perform convolution in the horizontal and vertical directions of the original feature map to obtain a fourth sub-feature map, and the sub-feature maps include the first sub-feature map, the second sub-feature map, the third sub-feature map, and the fourth sub-feature map.

3. The method according to claim 1, wherein The using a sampling point generator to process the downsampled feature map to generate a point sampling set includes: Dynamically determine a constraint factor according to the downsampled feature map, and determine a sampling set offset corresponding to the downsampled feature map based on the constraint factor, and the constraint factor is used to constrain the sampling set offset; Add the sampling set offset and the original sampling grid to obtain the point sampling set, and the original sampling grid represents the basic sampling coordinates.

4. The method according to claim 3, wherein The dynamically determining a constraint factor according to the downsampled feature map includes: Perform convolution processing on the downsampled feature map to obtain a basic offset; Normalize the basic offset to obtain a normalized basic offset, and multiply the normalized basic offset by a preset value to obtain a constraint factor corresponding to the downsampled feature map.

5. The method according to claim 4, wherein The determining a sampling set offset corresponding to the downsampled feature map based on the constraint factor includes: Multiply the constraint factor by the basic offset to obtain an initial offset; Perform pixel shuffling on the initial offset to obtain the sampling set offset.

6. The method according to claim 1, characterized in that, The upsampling the downsampled feature map according to the point sampling set to obtain the target feature map includes: Perform upsampling on the downsampled feature map according to the point sampling set through bilinear interpolation to obtain the target feature map.

7. The method according to claim 1, characterized in that, After the upsampling the downsampled feature map according to the point sampling set to obtain the target feature map, it further includes: Map the target feature map to a target channel format, generate an image processing result, and display the image processing result on a visualization interface; Receive an interaction operation for the image processing result, and adjust the image processing result based on the interaction operation.

8. An image processing apparatus, characterized in that, Including: An acquisition module configured to acquire an original feature map; The downsampling module is configured to perform wavelet decomposition on the original feature map to obtain a plurality of sub-feature maps, and splice the plurality of sub-feature maps to obtain a downsampled feature map; The sampling point generation module is configured to process the downsampled feature map by using a sampling point generator to generate a point sampling set, and the point sampling set is used to define the coordinate positions of each point on the target feature map; The upsampling module is configured to perform upsampling on the downsampled feature map according to the point sampling set to obtain the target feature map.

9. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.

10. A storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the method according to any one of claims 1 to 7.