Image Harmonization System Based on Style Transfer
Through the image harmony system based on style transfer, the incongruity problem of the foreground and background fusion in image synthesis is solved, real and natural harmonious image effects are achieved, and manual processing costs are reduced.
Patent Information
- Application Number
- CN202210589150.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-27
- Publication Date
- 2025-06-17
- Estimated Expiration
- 2042-05-27
AI Technical Summary
In the prior art, in image synthesis, foreground objects are prone to incongruence when fusing with background images, and traditional manual image processing methods are time-consuming and costly.
A kind of image harmony system based on style transfer is proposed. Multi-scale feature maps are extracted through feature extraction modules, and the area segmentation module scales the mask segmentation map to match the feature maps. The generator uses this information to generate the final harmonious image.
A more continuous visual style of foreground and background images is achieved, making the composite image look real and without any sense of incongruity, and reducing the cost of manual processing.
Smart Images

Figure CN115100024B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of computer vision technology, and in particular to an image harmonization system based on style transfer. Background Art
[0002] Image synthesis refers to cutting out part of the content of a picture as the foreground image and pasting it onto another background picture to obtain a composite image. However, there are many problems with the composite images obtained by segmentation or cutout algorithms. Among them, the disharmony of the composite image caused by the difference in the shooting environment of the foreground image and the background image is one of the representative problems.
[0003] Traditional manual image processing methods use image editing software (such as Photoshop) to manually adjust the differences between images. Since the operation involves multi-dimensional parameters, including brightness, contrast, saturation, etc., and they are interdependent, it takes professionals a lot of time to process, which increases labor costs. To address the above problem, an image harmonization scheme is proposed that aims to automatically adjust the foreground area to blend the background.
[0004] Early image harmonization schemes were mainly based on traditional image processing techniques to design better feature statistics and matching methods, including color, texture, or Poisson fusion, etc. However, faced with massive and diverse natural images, it is difficult for the above schemes to accurately estimate statistical features and thus the harmonization effect is limited. Summary of the invention
[0005] The present application aims to solve one of the technical problems in the related art at least to some extent.
[0006] To this end, the first purpose of the present application is to propose an image harmonization system based on style transfer, which solves the technical problem of the existing method that there is a sense of incongruity when the foreground object and the background image are merged, and realizes that the foreground and background images have a more continuous visual style, so that it can generate a real and natural harmonized image result for a given simple copy-and-paste synthetic image, that is, the synthetic image to be adjusted.
[0007] To achieve the above-mentioned purpose, the first aspect of the present application proposes an image harmonization system based on style transfer, including: a feature extraction module, used to extract a multi-scale feature map from a synthetic image through standard residual blocks of different scales, and send the feature maps of different scales to the corresponding level of the generator; a region segmentation module, used to scale the foreground and background mask segmentation maps so that the scaled foreground and background mask segmentation maps match the scale of the multi-scale feature map; a generator, used to generate a final harmonized synthetic image based on feature maps of different scales, the scaled foreground and background mask segmentation maps, and the background area.
[0008] The image harmonization system based on style transfer in the embodiments of the present application extracts multi-scale features from the synthesized image through multiple standard residual blocks; obtains the foreground and background images in the synthesized image by using the scaled foreground and background mask segmentation maps; anti-modulates the statistical features of the background region into the generated feature map through the region-aware adaptive instance normalization layer, so that the generated feature map has similar statistical characteristics to the background region; further, the region-aware adaptive denormalization layer learns the affine transformation parameters from the extracted foreground region and applies them to the generated feature map through multiplication and addition operations, retaining the semantic structure information similar to the foreground region; combines the newly generated foreground image with the extracted background image to generate the finally harmonized image, thereby making the harmonized synthesized image have better style consistency and look real and non-intrusive.
[0009] Optionally, in an embodiment of the present application, the region segmentation module is specifically used for:
[0010] Invert the foreground mask segmentation map to obtain the background mask segmentation map;
[0011] Scale the foreground mask segmentation map and the background mask segmentation map according to the sizes of the feature maps output by different scale residual blocks, so that the masks match the multi-scale feature maps.
[0012] Optionally, in an embodiment of the present application, the generator includes multiple levels, each level includes a region-aware adaptive instance normalization layer and a region-aware adaptive denormalization residual block, and the scales of all levels match the scales of the feature maps.
[0013] Optionally, in an embodiment of the present application, the region-aware adaptive instance normalization layer is used to obtain the background region feature map by using the corresponding scaled background mask segmentation map, and then extract the statistical features from the background region feature map and apply the statistical features to the output feature map of the previous layer through anti-modulation;
[0014] The region-aware adaptive denormalization residual block is used to extract the foreground region feature map by using the corresponding scaled foreground mask segmentation map, learn the affine transformation parameters from the foreground region feature map through the convolutional layer, and apply the affine transformation parameters to the feature map output by the region-aware adaptive instance normalization layer in the same layer through multiplication and addition operations;
[0015] Among them, the region-aware adaptive denormalization residual block has a structure opposite to that of the standard residual block.
[0016] Optionally, in an embodiment of the present application, the region-aware adaptive instance normalization layer is specifically used for:
[0017] Obtain the background region feature map according to the corresponding scaled background mask segmentation map, and perform instance normalization on the features obtained in the previous layer to obtain the normalized features of the previous layer;
[0018] Extract the corresponding mean statistical feature and standard deviation statistical feature from the background region feature map in the channel manner;
[0019] Calculate according to the standard deviation statistical feature, mean statistical feature, and normalized features of the previous layer, and output a new feature map.
[0020] Optionally, in an embodiment of the present application, the region-aware adaptive denormalization residual block is specifically used for:
[0021] Obtain the foreground region feature map according to the scaled foreground mask segmentation map, and perform batch normalization on the feature map output by the region-aware adaptive instance normalization layer;
[0022] Learn the affine transformation parameters from the foreground region feature map in the pixel manner through a convolutional layer, where the affine transformation parameters include the scaling scale parameter and bias parameter at the channel level;
[0023] Multiply the extracted scaling scale parameter by the normalized feature output by the adaptive instance normalization layer, and then add the learned bias parameter as the final feature and output it to the next layer.
[0024] Optionally, in an embodiment of the present application, the generator is specifically used for:
[0025] Output the adjusted foreground region image through multiple scales of region-aware adaptive instance normalization layers and region-aware adaptive denormalization residual blocks;
[0026] Obtain the background mask segmentation map of the same size as the synthesized image by inverting the original foreground mask segmentation area, and extract the background of the synthesized image using the background mask segmentation map;
[0027] Combine the adjusted foreground region image and the background region image to generate the finally harmonized synthesized image.
[0028] Additional aspects and advantages of the present application will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0029] The above and / or additional aspects and advantages of the present application will become apparent and easy to understand from the following description of the embodiments in conjunction with the drawings, where:
[0030] Figure 1Schematic diagram of a style transfer-based image harmonization system provided by an embodiment of the present application;
[0031] Figure 2 Structural diagram of the region-aware adaptive instance normalization layer according to an embodiment of the present application;
[0032] Figure 3 Structural diagram of the region-aware adaptive denormalization layer according to an embodiment of the present application;
[0033] Figure 4 Network architecture diagram of the style transfer-based image harmonization according to an embodiment of the present application. Detailed implementation manners
[0034] The embodiments of the present application will be described in detail below. The examples of the embodiments are shown in the accompanying drawings, in which the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present application and should not be construed as limiting the present application.
[0035] The style transfer-based image harmonization system according to an embodiment of the present application will be described below with reference to the accompanying drawings.
[0036] In order to implement the above embodiments, the present application proposes a style transfer-based image harmonization system.
[0037] Figure 1 Schematic diagram of a style transfer-based image harmonization system provided by an embodiment of the present application.
[0038] As Figure 1 shown, the style transfer-based image harmonization system includes a feature extraction module, a region segmentation module, and a generator, where:
[0039] The feature extraction module is configured to extract multi-scale feature maps from the synthetic image through standard residual blocks of different scales and send the feature maps of different scales to the corresponding levels of the generator;
[0040] The region segmentation module is configured to scale the foreground and background mask segmentation map so that the scaled foreground and background mask segmentation map matches the scale of the multi-scale feature maps;
[0041] The generator is configured to generate a final harmonized synthetic image according to the feature maps of different scales, the scaled foreground and background mask segmentation map, and the background region.
[0042] The image harmonization system based on style transfer in the embodiments of the present application extracts multi-scale features from the synthesized image through multiple standard residual blocks; obtains the foreground and background images in the synthesized image by using the scaled foreground and background mask segmentation maps; anti-modulates the statistical features of the background region into the generated feature map through the region-aware adaptive instance normalization layer, so that the generated feature map has similar statistical characteristics to the background region; further, the region-aware adaptive denormalization layer learns the affine transformation parameters from the extracted foreground region and applies them to the generated feature map through multiplication and addition operations, retaining the semantic structure information similar to the foreground region; combines the newly generated foreground image with the extracted background image to generate the final harmonized image, thereby making the harmonized synthesized image have better style consistency and look real and non-intrusive.
[0043] The network for capturing the features of the synthesized picture in the present application is composed of multiple standard residual blocks. Each residual block specifically includes three convolutional layers, where one convolutional layer is used to learn the skip connection. The instance normalization layer is used in the standard residual block, and Leaky ReLU is used as the activation function;
[0044] Before inputting the image feature map into the residual block, it is first downsampled by a factor of 2 according to the nearest neighbor method.
[0045] Optionally, in an embodiment of the present application, the region segmentation module is specifically used for:
[0046] Distinguish the foreground region and the background region in the synthesized image by using the foreground mask segmentation map. Specifically, invert the foreground mask map to obtain the background mask segmentation map, and scale the foreground mask segmentation map and the background mask segmentation map according to the size of the feature map output by different-scale residual blocks, so that the mask and the feature map can be perfectly matched.
[0047] Optionally, in an embodiment of the present application, the generator includes multiple levels, each level includes a region-aware adaptive instance normalization layer and a region-aware adaptive denormalization residual block, and the scales of all levels match the scale of the feature map.
[0048] Optionally, in an embodiment of the present application, the region-aware adaptive instance normalization layer is as Figure 2 shown, and is used to obtain the background region feature map by using the corresponding scaled background mask segmentation map, further extract the statistical features from the background region feature map, and apply the statistical features to the output feature map of the previous layer through anti-modulation;
[0049] The region-aware adaptive denormalization residual block is used to extract the foreground region feature map using the corresponding scaled foreground mask segmentation map, learn the affine transformation parameters from the foreground region feature map through a convolutional layer, and apply the affine transformation parameters to the feature map output by the region-aware adaptive instance normalization layer in the same layer through multiplication and addition operations;
[0050] Among them, the region-aware adaptive denormalization residual block is the inverse of the standard residual block structure.
[0051] Optionally, in an embodiment of the present application, the region-aware adaptive instance normalization layer is specifically used for:
[0052] Under the indication of the scaled background mask segmentation map, obtain the feature map of the background region, and perform instance normalization on the features obtained from the previous layer;
[0053] Extract its statistical features from the background region in a channel-wise manner, specifically including the mean and standard deviation at the channel level. The statistical feature information obtained here is not affected by the foreground region;
[0054] Multiply the extracted standard deviation statistical feature by the normalized features obtained from the previous layer, and then add the obtained mean statistical feature, which can be summarized as:
[0055]
[0056] Among them,
[0057]
[0058]
[0059] Here, represents the input feature of channel c of the i-th region-aware adaptive instance normalization layer, H and W respectively represent the height and width of the feature map, so the subscript h, w, c represents the feature value at the corresponding position, and ε represents a very small value, which is set to 10 -7 ;
[0060] And,
[0061]
[0062]
[0063] Here, represents the scaled background mask segmentation map of the i-th layer, which can be calculated by the formula Calculated, sum{M i =1} represents the number of pixels with a pixel value equal to 1 in the scaled foreground mask segmentation map of the i-th layer, and ο represents the Hadamard product;
[0064] The features after demodulation are fed as output to the subsequent region-aware adaptive denormalization residual block.
[0065] Optionally, in an embodiment of the present application, in the generator, a plurality of region-aware adaptive denormalization residual blocks are deployed, and the output of each residual block is upsampled by a factor of 2 according to the nearest neighbor method. Each residual block is deployed after the region-aware adaptive instance normalization layer at the same level;
[0066] The region-aware adaptive denormalization residual block adopts a structure inverse to the standard residual block used for feature capture, and uses the region-aware adaptive denormalization layer to replace the batch normalization layer, with Leaky ReLU as the activation function;
[0067] The region-aware adaptive denormalization layer is as Figure 3 shown. Under the indication of the scaled foreground mask segmentation map, the feature map of the foreground region is obtained, and batch normalization is performed on the features output by the region-aware adaptive instance normalization layer;
[0068] The affine transformation parameters are learned pixel by pixel from the foreground region through a convolutional layer, specifically including the scaling scale parameters and bias parameters at the channel level;
[0069] Multiplying the extracted scaling scale parameters by the normalized features output by the adaptive instance normalization layer and then adding the learned bias parameters can be expressed as:
[0070]
[0071] where,
[0072]
[0073]
[0074] Here, represents the input feature of channel c of the i-th region-aware adaptive denormalization layer. N, H, and W respectively represent the batch size and the height and width of the feature map. Thus, the subscripts n, h, w, c represent the feature values at the corresponding positions, and ε represents a very small value, which is set to 10 -7 ; and are the demodulated scaling scale parameters and bias parameters learned through one-layer convolution of the features of the obtained foreground region.
[0075] Optionally, in an embodiment of the present application, the generator is specifically used for:
[0076] The foreground image after style transfer is finally obtained through the region-aware adaptive instance normalization layer and the region-aware adaptive denormalization residual block at multiple scales.
[0077] By inverting the original foreground mask segmentation area, a background mask segmentation map of the same size as the synthesized image is obtained, and the background of the synthesized image is extracted using the background mask segmentation map;
[0078] The newly generated foreground area image is combined with the background area image to generate the final harmonized composite image, which has a consistent style and looks real and harmonious.
[0079] In the generator, adversarial loss, perceptual loss and feature matching loss are used. The loss function of the generator can be expressed as:
[0080]
[0081] Here, G represents the generator, z0 represents the Gaussian noise map of the initial input of the generator, M is the foreground mask segmentation map, and I c represents the original composite image of the input, is a perceptual loss function that aims to minimize the difference between the feature representations extracted by the VGG-19 network. It represents the feature matching loss function, which is used to match the intermediate features of different layers of the multi-scale discriminator. f represents the foreground image, I b represents the background image, λ adv ,λ P and FM is the corresponding weight parameter;
[0082] Based on the multi-scale discriminator, there are three regular discriminators with the same structure to discriminate images of different scales, respectively operating on the real image and the synthetic image output by the harmonized network. The loss function is expressed as:
[0083] L D =E[max(0,1-D(I))]+E[max(0,1+D(G(z0,I) c ,M)))]
[0084] Among them, D represents the discriminator and I represents the real image.
[0085] In the embodiment, the image harmonization network architecture based on style transfer is as follows: Figure 4As shown, the number of network layers of the actual image harmonization network does not necessarily have to be as shown in the figure, and any number can be set. The image harmonization network specifically includes four parts, namely, a feature extraction stream, a region-aware adaptive instance normalization layer, a region-aware adaptive denormalization residual block, and a discriminator, where the discriminator is not shown in Figure 4 . The foreground image after style transfer obtained by the image harmonization network is recombined with the background region of the input composite image, thus ensuring the consistency of the background image before and after processing.
[0086] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.
[0087] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present application, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically and clearly defined.
[0088] Any process or method description shown in the flowchart or described in other ways herein can be understood to represent a module, segment, or part of code including one or more executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of the present application includes additional implementations, where the functions can be executed in a substantially simultaneous manner or in a reverse order according to the functions involved, rather than in the order shown or discussed, which should be understood by those skilled in the art of the embodiments of the present application.
[0089] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a definite sequence list of executable instructions for implementing logical functions, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or used in combination with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection portion having one or more wirings (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable media on which the program can be printed, because the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or other appropriate processing as necessary, and then stored in a computer memory.
[0090] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.
[0091] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried by the method of implementing the above embodiments can be completed by a program instructing relevant hardware, and the program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.
[0092] In addition, each functional unit in various embodiments of the present application may be integrated into a processing module, may exist separately physically for each unit, or two or more units may be integrated into one module. The above-mentioned integrated module may be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it may also be stored in a computer-readable storage medium.
[0093] The above-mentioned storage medium may be a read-only memory, a magnetic disk, an optical disc, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.
Claims
1. An image harmonization system based on style transfer, characterized in that, It includes a feature extraction module, a region segmentation module, and a generator, where: The feature extraction module is used to extract multi-scale feature maps from the synthetic image through standard residual blocks of different scales, and send the feature maps of different scales to the corresponding levels of the generator; The region segmentation module is used to scale the foreground and background mask segmentation map so that the scaled foreground and background mask segmentation map matches the scale of the multi-scale feature map; The generator is used to generate the final harmonized synthetic image according to the feature maps of different scales, the scaled foreground and background mask segmentation map, and the background region; Among them, the generator includes multiple levels, each level contains a region-aware adaptive instance normalization layer and a region-aware adaptive denormalization residual block, and the scales of all levels match the scale of the feature map; The region-aware adaptive instance normalization layer is used to obtain the background region feature map by using the corresponding scaled background mask segmentation map, then extract the statistical features from the background region feature map, and apply the statistical features to the output feature map of the previous layer through inverse modulation; The region-aware adaptive denormalization residual block is used to extract the foreground region feature map by using the corresponding scaled foreground mask segmentation map, learn the affine transformation parameters from the foreground region feature map through the convolutional layer, and apply the affine transformation parameters to the feature map output by the region-aware adaptive instance normalization layer in the same layer through multiplication and addition operations; Among them, the structure of the region-aware adaptive denormalization residual block is inverse to that of the standard residual block; The region-aware adaptive denormalization residual block is specifically used for: Obtain the foreground region feature map according to the scaled foreground mask segmentation map, and perform batch normalization on the feature map output by the region-aware adaptive instance normalization layer; Learn the affine transformation parameters pixel by pixel from the foreground region feature map through the convolutional layer, where the affine transformation parameters include the scaling scale parameters and bias parameters at the channel level; Multiply the extracted scaling scale parameters by the normalized output feature of the adaptive instance normalization layer, and then add the learned bias parameters as the final feature and output it to the next layer.
2. The system according to claim 1, characterized in that, The region segmentation module is specifically used for: Invert the foreground mask segmentation map to obtain the background mask segmentation map; Scale the foreground mask segmentation map and the background mask segmentation map according to the size of the feature map output by the different-scale residual blocks, so that the mask matches the multi-scale feature map.
3. The system according to claim 1, characterized in that, The region-aware adaptive instance normalization layer is specifically used for: Obtain the background region feature map according to the corresponding scaled background mask segmentation map, and perform instance normalization on the feature obtained from the previous layer to obtain the normalized feature of the previous layer; Extract the corresponding mean statistical feature and standard deviation statistical feature from the background region feature map in the channel manner; Calculate according to the standard deviation statistical feature, the mean statistical feature, and the normalized feature of the previous layer to obtain a new feature map and output it.
4. The system according to claim 1, characterized in that, The generator is specifically used for: Output the adjusted foreground region image through the region-aware adaptive instance normalization layer and the region-aware adaptive denormalization residual block at multiple scales; Obtain a background mask segmentation map of the same size as the synthesized image by inverting the original foreground mask segmentation area, and extract the background of the synthesized image using the background mask segmentation map; Combine the adjusted foreground region image and the background region image to generate the final harmonized synthesized image.
Citation Information
Patent Citations
Examiner identity appraising system based on bionic and biological characteristic recognition
CN101246543A
Interactive three-dimensional cartoon human face generating method and device
CN101493953A