Position-sensitive illumination consistency rendering method and device
Through the illumination consistency rendering method based on soft contrast learning, using position-aware segmentation and comprehensive content optimization strategies, the problems of unnatural and discordantness of the foreground objects in the background image are solved, and a more natural and harmonious image synthesis is achieved.
Patent Information
- Application Number
- CN202510814364.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-06-18
AI Technical Summary
Prior art After inserting a realistic foreground object into another background image, the synthetic image may appear unnatural and discordant, mainly due to the failure to effectively consider local style differences.
The light consistency rendering method based on soft contrast learning is adopted, and the generator model is trained through position-aware segmentation and soft contrast learning, combining the generator and discriminator to achieve local style consistency and global style alignment, and the image harmony is maintained using a comprehensive content optimization strategy.
The problem of unnatural and discordant foreground objects in the background image was successfully solved, and a more natural and harmonious synthetic image was generated to meet the visual effects that meet users' expectations.
Smart Images

Figure CN120339491A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of illumination consistency rendering based on deep learning technology, and particularly relates to a position-sensitive illumination consistency rendering method and device. Background Art
[0002] In image processing, image synthesis is an important operation. After inserting a realistic foreground object into another background image, the synthesized image may look unnatural and discordant. The purpose of illumination consistency rendering is to transfer the style of the background to the foreground object, which is a challenging task because there is a large domain gap between the foreground and the background. Currently, mainstream methods (PHDNet, ProPIH) generally choose to align the foreground style with the overall style of the background. However, the style differences in different regions of the image may be large, and this method does not consider local styles, resulting in disharmonious and unnatural situations locally. Therefore, the present invention proposes a position-sensitive illumination consistency rendering method. Summary of the Invention
[0003] Aiming at the deficiencies in the prior art, the present invention provides a position-sensitive illumination consistency rendering method and device. The specific technical solutions are as follows:
[0004] In a first aspect, an embodiment of the present application provides a position-sensitive illumination consistency rendering method, including the following steps:
[0005] First, make a dataset to establish an image dataset required for network training.
[0006] Then, construct a basic illumination consistency rendering model based on soft contrast learning.
[0007] (1) Theoretical modeling of the basic illumination consistency rendering model based on soft contrast learning.
[0008] (2) Construct a training network for illumination consistency rendering based on soft contrast learning.
[0009] (3) Train to generate a network model with a position-sensitive image coordination effect, that is, obtain a basic illumination consistency rendering model based on soft contrast learning.
[0010] The trained neural network model receives the picture that needs to be subjected to illumination consistency rendering, and outputs the picture after completing the harmonious processing.
[0011] In a possible implementation, when making the dataset, the public datasets COCO and WikiArt dataset are downloaded. The images from the WikiaArt dataset are used as background images, and the instance segmentation masks provided by the COCO dataset are used to extract photographic foreground objects from the COCO dataset. The foreground objects are randomly inserted into the background images using the instance segmentation masks to form composite pictures. .
[0012] In a possible implementation, the process of the theoretical modeling includes:
[0013] Using a generator including a style conversion module to take the composite picture, the background image, and the foreground mask as inputs and output a harmonious image. Secondly, a discriminator is proposed to guide the generator to produce more realistic and coordinated images.
[0014] The generator is trained by continuously optimizing the loss. The total loss of the generator includes: a local soft contrast style loss for achieving local style consistency, an overall style loss for comprehensive content optimization, an adversarial loss, a structural similarity loss, and a detail similarity loss.
[0015] In a possible implementation, the training network for light consistency rendering based on soft contrast learning includes a generator and a discriminator. The generator consists of an encoder, a style conversion module, and a decoder. The encoder uses a pre-trained VGG-19 network, and the decoder is designed to reflect the architecture of the encoder by using upsampling operations. The style conversion module is used to capture the style information of the background and apply it to the foreground objects. The discriminator is used to evaluate whether there are discordant regions in the harmonious image.
[0016] In a possible implementation, the discriminator includes a lightweight spatial encoder and a lightweight autoencoder. Specifically, six downsampling blocks are applied in the lightweight spatial encoder, where each encoding block contains a convolutional layer with a kernel size of 4 and a stride of 2, followed by a batch normalization layer and an activation function layer; the lightweight autoencoder includes two downsampling blocks and two upsampling blocks; each downsampling block in the lightweight autoencoder sequentially contains a convolutional layer with a kernel size of 3 and a stride of 1, a batch normalization layer, and an activation function layer; while each upsampling block has the same structure as the downsampling block of the lightweight autoencoder, except for the activation function; the downsampling blocks use LeakyReLU activation and the upsampling blocks use ReLU activation.
[0017] In a possible implementation, the training process is supervised by soft contrast learning. First, a position-aware segmentation module is used to obtain candidate patches that need to be coordinated from the composite picture, and then soft contrast learning is used to transform the style of the foreground in each patch to achieve local consistency.
[0018] To preserve the details of the transferred foreground, a content retention process (i.e., an integrated content optimization strategy) is proposed, which includes two parts: global style alignment and content retention.
[0019] The global style alignment initially aligns the styles of the foreground region and the background region using an adversarial loss, and further enhances the style alignment effect through a distribution-based style alignment loss (i.e., the overall style loss):
[0020] Content retention maintains the consistency of the structure before and after transformation through a structural similarity loss. In addition, a detail similarity loss based on contrastive learning is used to better retain details.
[0021] In one possible implementation, the process of performing harmonious processing by a trained neural network model includes:
[0022] First, load the weights of the trained illumination consistency rendering network model and update the parameters in the model. Second, use the synthetic image to be harmonized as input data and input it into the generator network model. The input data passes through the generator to obtain the harmonized output image.
[0023] In a second aspect, an embodiment of the present application provides a location-sensitive illumination consistency rendering device, including the following modules:
[0024] A dataset production module for establishing an image dataset required for network training.
[0025] A model construction and training module for constructing a basic illumination consistency rendering model based on soft contrastive learning, including:
[0026] (1) Theoretical modeling of the basic illumination consistency rendering model based on soft contrastive learning.
[0027] (2) Construct a training network for illumination consistency rendering based on soft contrastive learning.
[0028] (3) Train to generate a network model with a location-sensitive image coordination effect, that is, obtain a basic illumination consistency rendering model based on soft contrastive learning.
[0029] A picture harmonization module uses the neural network model trained by the model construction and training module to receive the picture that needs to perform illumination consistency rendering, and outputs the picture after completing the harmonization process.
[0030] In a third aspect, an embodiment of the present application provides an electronic device, including a processor and a memory;
[0031] The memory is used to store a computer program;
[0032] When the processor is used to execute the program stored in the memory, it implements any one of the light consistency rendering methods described in this application.
[0033] In a fourth aspect, an embodiment of this application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements any one of the light consistency rendering methods described in this application.
[0034] In a fifth aspect, an embodiment of this application provides a computer program product containing instructions. When it runs on a computer, it causes the computer to execute any one of the light consistency rendering methods described in this application.
[0035] The beneficial effects of the present invention are as follows:
[0036] The method of the present invention successfully trains the generator model through position-aware segmentation and soft contrast learning to enable it to have position-sensitive light consistency rendering ability, which can effectively solve the problem that the synthesized image may look unnatural and disharmonious after inserting a realistic foreground object into another background image. In addition, the present invention realizes local style consistency through the method of soft contrast learning, and at the same time realizes global style alignment and content retention through a comprehensive content optimization strategy, and finally realizes the natural harmony of the synthesized image, so as to achieve the purpose of light consistency rendering. And the results based on user research show that our method can produce more satisfactory and harmonious images. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] In order to more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of this application, and those of ordinary skill in the art can also obtain other embodiments according to these drawings.
[0038] Figure 1 It is a schematic flowchart of the method according to the embodiment of the present invention.
[0039] Figure 2 It is a structural diagram of the generator model according to the embodiment of the present invention.
[0040] Figure 3 It is a comparison diagram of experimental effects according to the embodiment of the present invention.
[0041] Figure 4 It is a diagram of the results of user research according to the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0042] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art based on the present application belong to the scope of protection of the present application.
[0043] The present invention proposes a method based on soft contrast learning to achieve local style consistency. First, adversarial learning is used to bridge the domain gap between the foreground feature map and the background feature map, and then soft contrast learning is used to achieve position-sensitive illumination consistency rendering. At the same time, a comprehensive content optimization strategy is designed to achieve global style alignment and content retention, and finally position-sensitive adaptive illumination consistency rendering is achieved. The generator consists of an encoder, a style conversion module, and a decoder. The encoder uses a pre-trained VGG-19 network, and the decoder is designed to reflect the architecture of the encoder by using upsampling operations instead of downsampling operations. The style conversion module is used to capture the style information of the background and apply it to the foreground object. The discriminator is used to evaluate whether there are discordant regions in the harmonized image. The discriminator enhances the generator through confrontation, prompting the improved foreground feature map generated by the generator to be indistinguishable from the background feature map, so that the synthesized image looks more natural. To achieve position-sensitive illumination consistency rendering, the present invention first uses position-aware segmentation to obtain candidate patches that need to be coordinated from the synthesized image, and then uses soft contrast learning to transform the style of the foreground in each patch to achieve local consistency.
[0044] Through this method, the present invention can effectively solve the problem that the synthesized image may look unnatural and discordant after inserting a realistic foreground object into another background image. In addition, the present invention achieves local style consistency through the method of soft contrast learning, and at the same time achieves global style alignment and content retention through a comprehensive content optimization strategy, and finally achieves the natural harmony of the synthesized image, so as to achieve the purpose of illumination consistency rendering.
[0045] In a possible implementation manner, as Figure 1 shown, the method of the present invention specifically includes the following steps:
[0046] Step 1: Dataset production.
[0047] Build the image dataset required for network training in Step 2. Download the publicly available datasets COCO and WikiArt dataset. The COCO dataset is a large-scale dataset containing 123,287 images, which has instance segmentation annotations for objects in 80 categories. The Wikiart dataset is a large digital art dataset containing 81,444 images from 27 different styles. In the present invention, the images from the WikiaArt dataset are used as background images, and the foreground photographic objects are extracted from the COCO dataset using the instance segmentation masks provided by the COCO dataset. The foreground objects are randomly inserted into the background images using the instance segmentation masks to form composite pictures .
[0048] Step 2: Construct a basic illumination consistency rendering model based on soft contrast learning.
[0049] Step 2.1: Theoretical modeling of the basic illumination consistency rendering model based on soft contrast learning.
[0050] Use a generator including a style transfer module to take the composite picture , the background image and the foreground mask M as inputs, and output a harmonious image. Secondly, a discriminator is proposed to guide the generator to produce more realistic and harmonious images.
[0051] The theoretical model of the generator is expressed by the formula:[[]]
[0052]
[0053] where represents the composite picture, represents the background picture, the foreground mask M represents the area to be coordinated, and G represents the generator network.
[0054] The generator is trained by continuously optimizing the loss, and the total loss of the generator can be expressed as:[[]]
[0055]
[0056] where is the local soft contrast style loss used to achieve local style consistency, , , and are the overall style loss, adversarial loss, structural similarity loss and detail similarity loss respectively used for comprehensive content optimization. Among them ( = 1,..., 5) are hyperparameters, which are set to 0.5, 3, 6, 4, 10 in the present invention.
[0057] Step 2.2: Construct a training network for illumination consistency rendering based on soft contrast learning.
[0058] The network model includes a generator and a discriminator. The generator takes a synthetic image ( ), a background image ( ), and a foreground mask (M) as inputs. As Figure 2 shown, the generator consists of an encoder, a style transfer module, and a decoder. The encoder Em contains the first 31 layers of the pre-trained VGG-19 (up to ReLU4_1), and the decoder structure is symmetric to the encoder. When training the network, skip connections are added to the encoder Em at ReLU1_1, ReLU2_1, and Relu3_1 to preserve the content details in the low-level feature maps. First, the encoder Em extracts L = 4 layers of feature maps from the background image and the synthetic image , respectively from the four encoder layers ReLU1_1, ReLU2_1, ReLU3_1, and ReLU4_1, obtaining the background image features { } and the synthetic image features { }. For the th layer, the feature maps and are input into the style transfer module (AdaIN layer) together with the resized foreground mask , which adjusts the statistical information in the foreground region of to be consistent with the statistical information of , thus obtaining the stylized feature map The decoder is a symmetric structure of the encoder, which takes the coordinated encoder features as the input of the decoder or connects with the decoder features through skip connections to obtain the final output of the generator.
[0059] The input of the discriminator is the harmonious image , The discriminator includes a lightweight spatial encoder Ds and a lightweight autoencoder Da, which are built on top of downsampling (DS) blocks and upsampling (US) blocks. Specifically, six DS blocks are applied within the lightweight spatial encoder Ds, where each DS block contains a convolutional layer with a kernel size of 4 and a stride of 2, followed by a batch normalization (BN) layer and a LeakyReLU activation. Da is a small autoencoder that includes two DS blocks and two US blocks. Each DS block in Da sequentially contains a Conv layer with a kernel size of 3 and a stride of 1, a BN layer, and a LeakyReLU activation. Each US block has the same structure as the DS block in Da, except for the ReLU activation. The downsampling DS block uses LeakyReLU activation while the upsampling US block uses ReLU activation. In this way, the discriminator can identify whether there are discordant parts in the harmonious images.
[0060] Step 2.3: Train the training network for soft contrast learning-based illumination consistency rendering to form a network model that produces a position-sensitive illumination consistency rendering effect, that is, obtain the basic illumination consistency rendering model based on soft contrast learning.
[0061] To train the generator network to achieve position-sensitive illumination consistency rendering, the training process needs to be supervised by soft contrast learning. First, use the position-aware segmentation module to obtain the candidate patches that need to be coordinated from the synthetic image , and then use soft contrast learning to transform the style of the foreground in each patch to achieve local consistency. Specifically, the position-aware segmentation module first divides the synthetic image into k non-overlapping patches of size p*p, which are named { , ,... ,... }, and similarly divides the background image ( ) and the foreground mask (M) into { , ,... ,... } and { , ,... ,... }.
[0062] Take the part of in as an index to describe the style consistency degree of the i-th patch. This index is obtained through and is expressed as:
[0063]
[0064] where ∈[0,1] represents the style consistency degree of the i-th patch. The larger the value of, the better the consistency.
[0065] Considering that relatively large patches have good style consistency, which means that the information they provide is less and cannot effectively help us learn good style transformation, thereby reducing style inconsistency. Therefore, the present invention only selects relatively small patches as candidate patches for subsequent processing, thus obtaining candidate patches:
[0066]
[0067] where is a preset threshold, taking the value of 0.7 in the embodiment, represents the candidate patch set. The candidate patch set is obtained in the same way.
[0068] Traditional contrast loss assigns the same weight to all negative samples. However, this setting is somewhat unreasonable because different negative patches obviously show different importance. Although this different importance may be perceived when the network has learned to distinguish style differences, it is still challenging in the initial stage of training. Considering that contains the degree of style consistency, it is appropriate to introduce as a soft label to assist contrast learning. Therefore, the present invention takes ∈ as negative samples and ∈ as positive samples for soft contrast learning, and the formula is expressed as:
[0069]
[0070] where is the local soft contrast style loss; is the corresponding patch extracted from according to formula (4), is the deep feature extracted by the patch through the encoder ReLU4_1 and then obtained through the projection network. N represents the number of candidate patches in each set. Here, τ and are two hyperparameters, τ is the temperature coefficient set to 0.03, is the total number of negative samples, set to 1024. Among them, the negative samples are a queue of size 1024, which is slowly updated by momentum, and N negative samples are updated for each picture.
[0071] The above patch-based soft contrast loss learns local style information by maximizing the mutual information between corresponding patches in the input and output images. However, this patch-based process may overemphasize the consistency of local styles while neglecting the global style alignment between the foreground and the background, thus potentially introducing local style inconsistencies, especially in backgrounds with uneven style distributions. In addition, since no real labels are provided during the training phase, it is also a challenge to preserve the details of the transferred foreground. Therefore, the present invention proposes a content retention process (i.e., an integrated content optimization strategy), which consists of two parts: global style alignment and content retention.
[0072] Global style alignment aims to eliminate potential local style inconsistencies by globally aligning the style of the foreground with the background. First, an adversarial loss is used to preliminarily align the styles of the foreground region and the background region, and the expression is:
[0073]
[0074] where D represents the discriminator. Next, another distribution-based style alignment loss (overall style loss) is introduced to further enhance the style alignment effect:
[0075]
[0076] where represents the downsampled foreground mask M. represents the output of the i-1th layer of the encoder ReLU, represents the average value, represents the standard deviation.
[0077] Content retention aims to maintain the consistency of the details of the transferred foreground region with the original content. To this end, first, a structural similarity loss is introduced to maintain the consistency of the structure before and after the transformation, and the expression is:
[0078]
[0079] where represents the feature tensor extracted by the encoder ReLU4_1. In addition, a detail similarity loss based on contrastive learning is further formulated to better retain details.
[0080] Specifically, for each and , first, a background is randomly selected and denoted as , and then is used to generate with the same foreground as and An input generator is used to obtain a harmonized output Then, use the bounding box to crop to obtain the anchor points The detail similarity loss can be expressed as:
[0081]
[0082] where represents the negative samples. τ is the temperature coefficient set to 0.03, is the total number of negative samples, set to 1024. Among them, the negative samples are a queue of size 1024, which is slowly updated by momentum, and 1 negative sample is updated for each picture.
[0083] Network structure and implementation details. Adam optimizer is used for training, where β1 = 0 and β2 = 0.9. By default, the learning rate of the discriminator is 0.0004, and the learning rate of the generator is 0.0001. During training and testing, the size of the input image is adjusted to 256×256.
[0084] Step 3: The trained neural network model receives the picture that needs to be rendered with lighting consistency, and outputs the picture after completing the harmonization process.
[0085] First, load the weights of the lighting consistency rendering network model trained in Step 2 and update the parameters in the model. Secondly, use the synthetic picture to be harmonized as the input data and input it into the generator network model. The input data passes through the generator to obtain the harmonized output picture.
[0086] The embodiment of the present application also provides a position-sensitive lighting consistency rendering device, including the following modules:
[0087] A dataset production module, which is used to establish the image dataset required for network training.
[0088] A model construction and training module, which is used to construct a basic lighting consistency rendering model based on soft contrast learning, including:
[0089] (1) Theoretical modeling of the basic lighting consistency rendering model based on soft contrast learning.
[0090] (2) Construct a training network for lighting consistency rendering based on soft contrast learning.
[0091] (3) Train to generate a network model with a position-sensitive image coordination effect, that is, obtain the basic lighting consistency rendering model based on soft contrast learning.
[0092] The picture harmony module uses the neural network model trained by the model construction and training module to receive the picture that needs to be rendered with lighting consistency, and outputs the picture after completing the harmony processing.
[0093] In a possible implementation, the dataset production module uses the images from the WikiaArt dataset as background images, and extracts photographic foreground objects from the COCO dataset using the instance segmentation masks provided by the COCO dataset. The foreground objects are randomly inserted into the background images using the instance segmentation masks to form composite pictures.
[0094] In a possible implementation, the process of theoretical modeling of the model construction and training module includes:
[0095] Using a generator including a style transfer module to take the composite picture, background image and foreground mask as inputs and output a harmonious image. Secondly, a discriminator is proposed to guide the generator to produce more realistic and harmonious images.
[0096] The generator is trained by continuously optimizing the loss. The total loss of the generator includes: a local soft contrast style loss for achieving local style consistency, an overall style loss for comprehensive content optimization, an adversarial loss, a structural similarity loss, and a detail similarity loss.
[0097] In a possible implementation, the training network for lighting consistency rendering based on soft contrast learning of the model construction and training module includes a generator and a discriminator. The generator consists of an encoder, a style transfer module and a decoder. The encoder uses a pre-trained VGG-19 network, and the decoder is designed to reflect the architecture of the encoder by using upsampling operations. The style transfer module is used to capture the style information of the background and apply it to the foreground objects. The discriminator is used to evaluate whether there are disharmonious regions in the image after harmony.
[0098] In a possible implementation, the discriminator includes a light spatial encoder Ds and a lightweight autoencoder Da. Specifically, six DS blocks are applied in the light spatial encoder Ds, where each DS block contains a convolutional layer with a kernel size of 4 and a stride of 2, followed by a batch normalization (BN) layer and a LeakyReLU activation. Da is a small autoencoder, including two DS blocks and two US blocks. Each DS block in Da sequentially contains a Conv layer with a kernel size of 3 and a stride of 1, a BN layer and a LeakyReLU activation. And each US block has the same structure as the DS block of Da, except for the ReLU activation. The downsampling DS block uses LeakyReLU activation while the upsampling US block uses ReLU activation.
[0099] In a possible implementation, the training process of the model construction and training module is supervised by soft contrastive learning. First, the position-aware segmentation module is used to obtain candidate patches that need to be coordinated from the synthetic image, and then soft contrastive learning is used to transform the style of the foreground in each patch to achieve local consistency.
[0100] To preserve the details of the transferred foreground, a content retention process (i.e., comprehensive content optimization strategy) is proposed, which includes two parts: global style alignment and content retention.
[0101] Global style alignment uses adversarial loss to initially align the styles of the foreground region and the background region, and further enhances the style alignment effect through a distribution-based style alignment loss (i.e., overall style loss):
[0102] Content retention maintains the consistency of the structure before and after transformation through structural similarity loss. In addition, a detail similarity loss based on contrastive learning is used to better retain details.
[0103] In a possible implementation, the process of the picture harmony module for harmony processing includes:
[0104] First, load the weights of the light consistency rendering network model trained by the model construction and training module, and update the parameters in the model. Second, take the synthetic image to be harmonized as the input data and input it into the generator network model. The input data passes through the generator to obtain the harmonized output image.
[0105] Experimental comparison: The method of the present invention successfully realizes position-sensitive light consistency rendering, while the existing advanced methods (PHDNet, ProPIH) do not have the ability to sense styles according to positions. The results are as Figure 3 shown. In addition, we also conducted a user study to objectively and quantitatively evaluate 12 comparison methods. We first randomly selected background images, foreground images, and corresponding masks to generate 100 synthetic images, and obtained their harmony results through all comparison methods. Therefore, we obtained 100 groups of harmony images, each group containing 12 harmony images of the same synthetic image generated by different methods. For each group of images, 66 different combinations can be obtained to form a pair of different images. Therefore, we can obtain 6600 pairs of harmony images for the user study. We organized 15 participants to rate these 6600 pairs of harmony images and rank each pair of images according to visual quality. Then, we used the Bradley-Terry (B-T) model (Bradley and Terry, 1952) to calculate the user study scores of the comparison methods. The results are as Figure 4 shown.
[0106] An embodiment of the present application also provides an electronic device, which includes a processor and a memory;
[0107] The memory is used to store a computer program;
[0108] When the processor is used to execute the program stored on the memory, the method described in any one of the present application is implemented.
[0109] In a possible implementation manner, the electronic device of the embodiment of the present application further includes a communication interface and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus.
[0110] The communication bus mentioned in the above electronic device may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0111] The communication interface is used for communication between the above electronic device and other devices.
[0112] The memory may include a Random Access Memory (RAM), or may also include a Non-Volatile Memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0113] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0114] In another embodiment provided by the present application, a computer-readable storage medium is further provided. A computer program is stored in the computer-readable storage medium, and when the computer program is executed by a processor, the method described in any one of the present application is implemented.
[0115] In another embodiment provided by the present application, a computer program product including instructions is further provided. When it runs on a computer, the computer is caused to execute the method described in any one of the present application.
[0116] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions may be transmitted from a website, a computer, a server, or a data center to another website, a computer, a server, or a data center in a wired manner (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer, or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (for example, a floppy disk, a hard disk, a magnetic tape), an optical medium (for example, a DVD), or a semiconductor medium (for example, a Solid State Disk (SSD)).
[0117] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0118] Each embodiment in this specification is described in a related manner. Each embodiment focuses on the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other.
[0119] The above are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are all included in the protection scope of the present application.
Claims
1. A position-sensitive light consistency rendering method, characterized in that It includes the following steps: First, make a dataset to establish the image dataset required for network training; Then, construct a basic illumination consistency rendering model based on soft contrast learning: (1) Theoretical modeling of the basic illumination consistency rendering model based on soft contrast learning; (2) Construct a training network for illumination consistency rendering based on soft contrast learning; (3) Train to generate a network model with a position-sensitive image coordination effect, that is, obtain the basic illumination consistency rendering model based on soft contrast learning; The trained neural network model receives the picture that needs to be rendered with illumination consistency, and outputs the picture after completing the harmonious processing.
2. The position-sensitive light consistency rendering method according to claim 1, characterized in that When making the dataset, use the images from the WikiaArt dataset as the background images, and use the instance segmentation masks provided by the COCO dataset to extract the photographic foreground objects from the COCO dataset; randomly use the instance segmentation masks to insert the foreground objects into the background images to form synthetic pictures.
3. A position-sensitive light consistency rendering method according to claim 1, characterized in that The process of the theoretical modeling includes: Use a generator including a style transfer module to take the synthetic picture, the background image, and the foreground mask as inputs and output a harmonious image; secondly, propose a discriminator to guide the generator to generate more realistic and coordinated images; The generator is trained by continuously optimizing the loss. The total loss of the generator includes: a local soft contrast style loss for achieving local style consistency, an overall style loss for comprehensive content optimization, an adversarial loss, a structural similarity loss, and a detail similarity loss.
4. A position-sensitive light consistency rendering method according to claim 1, characterized in that The training network for illumination consistency rendering based on soft contrast learning includes a generator and a discriminator; The generator consists of an encoder, a style transfer module, and a decoder; the encoder uses a pre-trained VGG-19 network, and the decoder is designed to reflect the architecture of the encoder by using upsampling operations; The style transfer module is used to capture the style information of the background and apply it to the foreground object; the discriminator is used to evaluate whether there are discordant regions in the harmonious image.
5. A position-sensitive light consistency rendering method according to claim 4, wherein The discriminator includes a light spatial encoder and a lightweight autoencoder; specifically, six downsampling blocks are applied in the light spatial encoder, and each encoding block contains a convolutional layer with a kernel size of 4 and a stride of 2, followed by a batch normalization layer and an activation function layer; the lightweight autoencoder includes two downsampling blocks and two upsampling blocks; each downsampling block in the lightweight autoencoder sequentially contains a convolutional layer with a kernel size of 3 and a stride of 1, a batch normalization layer, and an activation function layer; while each upsampling block has the same structure as the downsampling block, only the activation function is different; the downsampling block uses the LeakyReLU activation and the upsampling block uses the ReLU activation.
6. A position-sensitive light consistency rendering method according to claim 1, characterized in that The training process is supervised by soft contrast learning; first, use a position-aware segmentation module to obtain the candidate patches that need to be coordinated from the synthetic picture, and then use soft contrast learning to transform the style of the foreground in each patch to achieve local consistency; To maintain the details of the transferred foreground, a content retention process is proposed, which includes two parts: global style alignment and content retention; Global style alignment initially aligns the styles of the foreground and background regions using an adversarial loss, and further enhances the style alignment effect through a distribution-based style alignment loss: Content preservation maintains the consistency of the structure before and after transformation through a structural similarity loss. Additionally, a detail similarity loss based on contrastive learning is used to better preserve details.
7. A position-sensitive light consistency rendering method according to claim 1, characterized in that The process of harmonious processing by the trained neural network model includes: First, load the weights of the trained illumination consistency rendering network model and update the parameters in the model. Second, take the synthetic image to be harmonized as the input data and input it into the generator network model. The input data passes through the generator to obtain the harmonized output image.
8. A position-sensitive light consistency rendering device, characterized in that, It includes the following modules: The dataset production module is used to establish the image dataset required for network training; The model construction and training module is used to construct a basic illumination consistency rendering model based on soft contrastive learning, including: (1) Theoretical modeling of the basic illumination consistency rendering model based on soft contrastive learning; (2) Construct a training network for illumination consistency rendering based on soft contrastive learning; (3) Train to generate a network model with a position-sensitive image coordination effect, that is, obtain the basic illumination consistency rendering model based on soft contrastive learning; The image harmonization module uses the neural network model trained by the model construction and training module to receive the image that needs to be rendered for illumination consistency, and outputs the image after completing the harmonization process.
9. An electronic device, characterized in that, It includes a processor and a memory; The memory is used to store computer programs; When the processor is used to execute the program stored on the memory, it implements the illumination consistency rendering method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, it implements the illumination consistency rendering method as described in any one of claims 1-7.
Citation Information
Patent Citations
Rendering image illumination method based on generative adversarial network
CN107862734A
Illumination remapping method and device for image synthesis, storage medium and processor
CN110288512A
Method and system for harmonizing synthetic image based on foreground reference image
CN115205544A
Synthetic image harmonization model training method, harmonization method and device
CN115456921A
Image harmony method based on attention generative adversarial network
CN117408900A