Composite image processing method, device, electronic device and storage medium

By extracting and dividing the foreground area of the synthetic image, and using the blocked foreground processing model to perform adaptive color transformation, the problem of inaccurate color transformation in the prior art is solved, and the harmony and authenticity of the synthetic image is improved.

CN114612357BActive Publication Date: 2025-08-08TENCENT TECH SHANGHAI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210217122.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-03-07
Publication Date
2025-08-08
Estimated Expiration
2042-03-07

AI Technical Summary

Technical Problem

In the prior art, the harmony algorithm of synthetic images is not accurate enough to process the color transformation of the foreground area, resulting in poor harmony of some color blocks.

Method used

By extracting the foreground area of the composite image to be processed, performing global color transformation processing, it is divided into multiple sub-regions, and using the chunked foreground processing model to adapt to different degrees of color transformation processing for each sub-region.

Benefits of technology

It improves the accuracy of color transformation processing, coordinates the foreground area with the background area, and enhances the authenticity of the synthetic image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114612357B_ABST
    Figure CN114612357B_ABST
Patent Text Reader

Abstract

The present application discloses a composite image processing method, apparatus, electronic device, and storage medium. The method comprises: performing global color transformation processing on the foreground area of the composite image to be processed to obtain an initial composite image, and dividing the foreground area of the composite image to be processed into multiple sub-areas. The feature information of the block areas corresponding to the initial composite image and the multiple sub-areas is input into a block foreground processing model, and the multiple sub-areas are adaptively subjected to different degrees of color transformation processing to obtain a target composite image. The method can perform different degrees of color transformation processing on different sub-areas in the foreground area of the composite image to be processed, so that each sub-area in the foreground area can be coordinated with the background area, thereby improving the accuracy of the color transformation processing and improving the authenticity of the target composite image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of synthetic computer technology, and in particular to synthetic image processing methods, devices, electronic devices and storage media. Background Art

[0002] The goal of an image harmonization algorithm is to transform the colors of the implanted foreground portion of a composite image, ensuring that the overall composite image appears realistic and harmonious while preserving its content. In existing technologies, harmonization algorithms typically only apply the same color transformation to the entire foreground region of the composite image to produce the harmonized image. However, because the foreground often contains different color blocks, the same color transformation may result in poor harmonization of some color blocks, leading to low color transformation accuracy. Summary of the Invention

[0003] The present application provides a composite image processing method, device, electronic device and storage medium, which can improve the accuracy of color conversion processing.

[0004] In one aspect, the present application provides a composite image processing method, the method comprising:

[0005] performing a global color transformation process on the synthesized image to be processed based on feature information of a foreground region of the synthesized image to be processed, to obtain an initial synthesized image; wherein the feature information of the foreground region is obtained by extracting features from the foreground region of the synthesized image to be processed;

[0006] Dividing the foreground area into multiple sub-areas;

[0007] The initial composite image and the block region feature information corresponding to each of the multiple sub-regions are input into the block foreground processing model for color transformation processing to obtain a target composite image, wherein each sub-region in the target composite image corresponds to a color transformation processing result of different degrees.

[0008] Another aspect provides a composite image processing device, the device comprising:

[0009] a global processing module configured to perform global color transformation processing on the composite image to be processed based on feature information of a foreground region of the composite image to be processed, to obtain an initial composite image, wherein the feature information of the foreground region is obtained by extracting features from the foreground region of the composite image to be processed;

[0010] A sub-region division module, configured to divide the foreground region into multiple sub-regions;

[0011] The block processing module is used to input the block area feature information corresponding to the initial composite image and the multiple sub-areas into the block foreground processing model for color transformation processing to obtain a target composite image, wherein each sub-area in the target composite image corresponds to a color transformation processing result of different degrees.

[0012] On the other hand, an electronic device is provided, which includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement a synthetic image processing method as described above.

[0013] On the other hand, a computer-readable storage medium is provided, which includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement a synthetic image processing method as described above.

[0014] On the other hand, a computer program product is provided, comprising a computer program, which implements the above-mentioned composite image processing method when executed by a processor.

[0015] The present application provides a synthetic image processing method, device, electronic device, and storage medium. The method can perform global color transformation processing on the foreground area of the synthetic image to be processed to obtain an initial synthetic image, and divide the foreground area of the synthetic image to be processed into multiple sub-areas. The feature information of the block areas corresponding to the initial synthetic image and the multiple sub-areas is input into a block foreground processing model, and the multiple sub-areas are adaptively subjected to different degrees of color transformation processing to obtain a target synthetic image. The method can perform different degrees of color transformation processing on different sub-areas in the foreground area of the synthetic image to be processed, so that each sub-area in the foreground area can be coordinated with the background area, thereby improving the accuracy of the color transformation processing and improving the authenticity of the target synthetic image. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0017] Figure 1 A schematic diagram of an application scenario of a composite image processing method provided in an embodiment of the present application;

[0018] Figure 2 A flowchart of a composite image processing method provided in an embodiment of the present application;

[0019] Figure 3 A flowchart of obtaining an initial composite image by using a global foreground processing model in a composite image processing method provided in an embodiment of the present application;

[0020] Figure 4 A flowchart of region division in a composite image processing method provided in an embodiment of the present application;

[0021] Figure 5 A schematic diagram of a composite image to be processed, feature information of a foreground region, and feature information of sub-regions corresponding to the sub-regions in a composite image processing method provided in an embodiment of the present application;

[0022] Figure 6 A flowchart of adding global feature processing results to a block foreground processing model in a synthetic image processing method provided in an embodiment of the present application;

[0023] Figure 7 A flowchart of obtaining a target composite image by using a block foreground processing model in a composite image processing method provided in an embodiment of the present application;

[0024] Figure 8 A schematic diagram of feature fusion performed by a feature fusion layer in a block foreground processing model of a synthetic image processing method provided in an embodiment of the present application;

[0025] Figure 9 A schematic diagram of model training in a synthetic image processing method provided in an embodiment of the present application;

[0026] Figure 10 A schematic diagram of a model structure corresponding to a synthetic image processing method provided in an embodiment of the present application;

[0027] Figure 11 A comparison chart showing the processing results of a synthetic image processing method provided in an embodiment of the present application and the processing results of a real image, a synthetic image, and a basic algorithm;

[0028] Figure 12 A schematic structural diagram of a composite image processing device provided in an embodiment of the present application;

[0029] Figure 13 A schematic diagram of the hardware structure of a device provided in an embodiment of the present application for implementing the method provided in an embodiment of the present application. DETAILED DESCRIPTION

[0030] To make the objectives, technical solutions, and advantages of this application more clear, this application will be further described in detail below with reference to the accompanying drawings. It is clear that the embodiments described are only some of the embodiments of this application, and not all of them. All other embodiments obtained by persons of ordinary skill in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.

[0031] In the description of the present application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the features. Moreover, the terms "first", "second", etc. are applicable to distinguishing similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way are interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein.

[0032] It is understandable that in the specific implementation of this application, related data such as user information is involved. When the above embodiments of this application are applied to specific products or technologies, user permission or consent is required, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards of relevant countries and regions.

[0033] See Figure 1 , which shows a schematic diagram of an application scenario of the composite image processing method provided by an embodiment of the present application. The application scenario includes a client 110 and a server 120. Server 120 receives a composite image to be processed from client 110, performs global color transformation on the composite image, then divides the foreground area of the composite image into regions. The global color transformation results and the region division results are subjected to block-by-block color transformation using a block-by-block foreground processing model to obtain a target composite image. Server 120 then sends the target composite image to client 110 for display.

[0034] In the embodiment of the present application, the client 110 includes but is not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle-mounted terminals, aircraft, etc.

[0035] In the embodiment of the present application, the server 120 may include an independent server, a distributed server, or a server cluster composed of multiple servers. The server 120 may include a network communication unit, a processor, a memory, and the like.

[0036] In an embodiment of the present application, the server 120 can perform color conversion processing on the processed composite image through a machine learning algorithm. Machine Learning (ML) is a multi-disciplinary interdisciplinary subject involving probability theory, statistics, approximation theory, convex analysis, algorithm complexity theory and other disciplines. It specializes in studying how computers simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their own performance. Machine learning is the core of artificial intelligence and the fundamental way to make computers intelligent. Its applications are spread across all areas of artificial intelligence. Machine learning and deep learning generally include artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and formulaic learning.

[0037] See Figure 2 , which shows a synthetic image processing method that can be applied on the server side, the method includes:

[0038] S210. Based on the foreground region feature information of the composite image to be processed, the composite image to be processed is subjected to global color transformation processing to obtain an initial composite image, where the foreground region feature information is obtained by feature extraction of the foreground region of the composite image to be processed;

[0039] In some embodiments, feature extraction is performed on the foreground region of the composite image to be processed to obtain foreground region feature information. The foreground region feature information may be mask feature information of the foreground region, which may be in various forms such as a binary image or a grayscale image. The foreground region feature information is used to distinguish the foreground region from the background region in the composite image to be processed. The composite image to be processed is an image in which the foreground region and the background region are synthesized and the foreground region and the background region do not match.

[0040] In some embodiments, performing global color transformation on the synthesized image to be processed based on feature information of the foreground region of the synthesized image to be processed to obtain an initial synthesized image includes:

[0041] The synthesized image to be processed and the feature information of the foreground area are input into the global foreground processing model for global color transformation processing to obtain the initial synthesized image.

[0042] In some embodiments, global color transformation processing can be performed using a global foreground processing model, which can include models such as Residual Network (ResNet), VGG16, and U-Net. When applying the global foreground processing model, if partial information from the global foreground processing model needs to be input into the block foreground processing model, the network structure of the global foreground processing model must correspond one-to-one with the network structure of the block foreground processing model.

[0043] In some embodiments, the input of the global foreground processing model is the composite image to be processed and foreground region feature information. In the global foreground processing model, based on the foreground region feature information of the composite image to be processed, global color transformation processing is performed on the composite image to be processed. That is, the foreground and background regions in the composite image to be processed are subjected to overall color transformation processing to obtain an initial composite image. The initial composite image is the result of the coarse-grained color transformation processing.

[0044] In other words, in the global foreground processing model, the foreground and background regions are harmonized at a coarse granularity level so that the image feature information corresponding to the foreground region matches the image feature information corresponding to the background region, thereby obtaining an initial composite image. Matching here means that the image feature information corresponding to the foreground region and the image feature information corresponding to the background region belong to the same environment as a whole. The image feature information can be pixel information.

[0045] For example, the pixel information corresponding to the background area belongs to an environment with high light brightness and bright color tones, while the pixel information corresponding to the foreground area belongs to an environment with low light brightness and grayish color tones. After performing color transformation processing on the pixel information corresponding to the foreground area, the pixel information corresponding to the foreground area also belongs to an environment with high light brightness and bright color tones. At this time, the image feature information corresponding to the foreground area matches the image feature information corresponding to the background area.

[0046] Performing global color transformation processing through the global foreground processing model can improve the efficiency of global processing and provide the original feature information of the synthesized image to be processed for the block foreground processing model in the subsequent step.

[0047] In some embodiments, see Figure 3 The global foreground processing model includes multiple sequentially arranged global encoding layers and a global decoding layer corresponding to each global encoding layer. The synthesized image to be processed and the foreground region feature information are input into the global foreground processing model for global color transformation processing. The initial synthesized image includes:

[0048] S310. Inputting the synthesized image to be processed and the foreground region feature information into a plurality of sequentially arranged global coding layers for feature processing, obtaining initial global feature information output by each global coding layer;

[0049] S320. When the current global decoding layer is the first global decoding layer, input the third initial global feature information into the first global decoding layer for feature processing to obtain target global feature information output by the first global decoding layer, where the third initial global feature information is information output by the last global coding layer;

[0050] S330. When the current global decoding layer is not the first global decoding layer, obtaining the previous target global feature information corresponding to the current global decoding layer;

[0051] S340. Input the previous target global feature information and the fourth initial global feature information into the current global decoding layer for feature processing to obtain the target global feature information output by the current global decoding layer, and the fourth initial global feature information is the information output by the global encoding layer corresponding to the current global decoding layer;

[0052] S350. Obtain an initial synthesized image based on the target global feature information output by the last global decoding layer.

[0053] In some embodiments, the global foreground processing model may include a plurality of global encoding layers arranged in sequence and a global decoding layer corresponding to each global encoding layer.

[0054] When the global coding layer currently performing feature processing is the first global coding layer, the input information of the first global coding layer is the synthesized image to be processed and the foreground region feature information. Before being input into the first global coding layer, the synthesized image to be processed and the foreground region feature information may be concatenated to obtain first concatenated feature information. The first concatenated feature information is then input into the first global coding layer for feature processing to obtain initial global feature information output by the first global coding layer.

[0055] When the global coding layer currently performing feature processing is not the first global coding layer, the initial global feature information output by each global coding layer is used as the input information of the next global coding layer to obtain the initial global feature information output by the global coding layer currently performing feature processing.

[0056] When the current global decoding layer currently performing feature processing is the first global decoding layer, the initial global feature information output by the last global coding layer among the multiple global coding layers arranged in sequence, that is, the third initial global feature information, is used as the input information of the first global decoding layer, and is input into the first global decoding layer for feature processing to obtain the target global feature information corresponding to the first global decoding layer.

[0057] When the current global decoding layer currently undergoing feature processing is not the first global decoding layer, the previous target global feature information corresponding to the current global decoding layer is obtained. The previous target global feature information is the target global feature information output by the global decoding layer immediately preceding the current global decoding layer. The initial global feature information output by the global coding layer corresponding to the current global decoding layer, i.e., the fourth initial global feature information, is also obtained. The previous target global feature information and the fourth initial global feature information are input into the current global decoding layer for feature processing to obtain the target global feature information output by the current global decoding layer.

[0058] The correspondence between the global decoding layer and the global coding layer indicates that the global coding layer and the corresponding global coding layer have the same dimension, which can be the feature resolution. For example, if a global decoding layer outputs 64x64 target global feature information, the corresponding global coding layer outputs 64x64 initial global feature information.

[0059] When performing feature processing in the global coding layer, the feature resolution corresponding to each global coding layer gradually decreases, while the number of channels gradually increases. Convolution operations can be used to abstract high-level features from the first concatenated feature information, resulting in the initial global feature information corresponding to each global coding layer. When performing feature processing in the global decoder, the feature resolution corresponding to each global decoding layer gradually increases, while the number of channels gradually decreases. Deconvolution operations can be used to obtain an initial composite image at the target feature resolution.

[0060] When performing feature processing through the decoding layer, inputting the output information of the previous layer of each decoding layer and the output information of the encoding layer corresponding to each decoding layer can reduce the information loss in the input to-be-processed synthetic image and the feature information of the foreground area, thereby improving the accuracy of the global color transformation processing.

[0061] S220. Divide the foreground area into multiple sub-areas;

[0062] In some embodiments, the foreground area is divided into multiple sub-areas. The number of sub-areas can be a preset number, which is the result of multiple experiments. In actual applications, there is no limit on the number of sub-areas. For example, the sub-areas can be determined randomly.

[0063] In some embodiments, see Figure 4 , divide the foreground area into multiple sub-areas including:

[0064] S410. Determine multiple initial block center information based on the image feature information corresponding to the foreground area;

[0065] S420. Using multiple initial block center information as current block center information;

[0066] S430. Based on the foreground area and the current block center information, determine the current block area corresponding to each current block center information;

[0067] S440. Based on the current block area, the current block center information corresponding to the current block area is updated;

[0068] S450. Repeat the steps of determining the current block area corresponding to each current block center information based on the foreground area and the current block center information, and updating the current block center information corresponding to the current block area based on the current block area, until the current block center information before and after the update matches, and using the current block center information as the target block center information;

[0069] S460. Use the current block area corresponding to the target block center information as a sub-area.

[0070] In some embodiments, the foreground area can be divided into multiple sub-areas based on a preset clustering algorithm. The clustering algorithm may include a K-Means Clustering Algorithm (KMeans), an improved K-Means Clustering Algorithm (KMeans++), a density-based spatial clustering algorithm (Density-Based Spatial Clustering of Applications with Noise, DBSCAN), etc.

[0071] In some embodiments, multiple initial block center information may be randomly determined based on the image feature information corresponding to the foreground area. The initial block center information is distributed in the foreground area. In the process of determining the initial block center information, multiple target feature point information may be randomly obtained from the image feature information corresponding to the foreground area as the initial block center information.

[0072] The multiple initial block center information is used as the current block center information. Based on the foreground area and the current block center information, the feature point information other than the current block center information in the image feature information corresponding to the foreground area can be determined, and the distance between each other feature point information and the current block center information can be determined. Each other feature point information can be classified based on the distance, and each other feature point information is associated with the current block center information. After the classification is completed, the area corresponding to the other feature point information associated with the current block center information is obtained. This area is the current block area corresponding to each current block center information.

[0073] Based on the feature point information in the current block area, the cluster center information of the feature point information in the current block area can be re-determined, and based on the cluster center information, the current block center information corresponding to the current block area is updated to obtain updated block center information. The cluster center information is the centroid of the feature point information in the current block area.

[0074] The updated block center information is used as the current block center information. The above steps of determining the current block area corresponding to the current block center information and updating the current block center information based on each current block area are repeated until the current block center information before and after the update does not change. In other words, when the current block center information before and after the update matches, the current block center information is determined as the target block center information. The current block area corresponding to the target block center information is the sub-area.

[0075] In some embodiments, the foreground area in the image to be processed can be regarded as a collection of pixel information, and the image feature information corresponding to the foreground area, the initial block center information, and other feature point information in the foreground area except the initial block center information can all be pixel information.

[0076] When determining the current block area, the set of feature point information associated with the current block center information is the pixel information set. The area corresponding to this pixel information set in the foreground area can be determined to obtain the current block area. When the current block center information corresponding to the current block area is updated, the pixel information serving as the current block center information is updated. When obtaining the sub-area corresponding to the target block center information, the colors corresponding to the pixel information in each sub-area are within the same color value range, and different sub-areas correspond to different colors.

[0077] In some embodiments, see Figure 5 ,like Figure 5 Shown is a schematic diagram of a synthesized image to be processed, feature information of a foreground region, and feature information of a sub-region corresponding to a block region. Figure 5 The first one from the left in the first row is the composite image to be processed, the first one from the left in the second row is the feature information of the foreground area, and the remaining four images are the feature information of the block areas, which correspond to the sub-area of the first color, the sub-area of the second color, the sub-area of the third color, and the sub-area of the fourth color in the composite image to be processed respectively.

[0078] By using a preset clustering algorithm to divide the foreground area into regions, sub-regions corresponding to different colors are obtained, which can improve the rationality of the region division and facilitate the adaptive color transformation processing of each sub-region in subsequent steps.

[0079] S230. Input the initial composite image and the block region feature information corresponding to each of the multiple sub-regions into the block foreground processing model for color transformation processing to obtain a target composite image, where each sub-region in the target composite image corresponds to a color transformation processing result of different degrees.

[0080] In some embodiments, sub-regions of different colors require different degrees of harmonization. The block foreground processing model can adaptively perform different degrees of color transformation processing on each sub-region, so that each sub-region in the target composite image corresponds to a different degree of color transformation processing result.

[0081] In some embodiments, see Figure 6 , the initial composite image and the feature information of the sub-regions corresponding to each of the multiple sub-regions are input into the sub-foreground processing model for color conversion processing, and the target composite image is obtained including:

[0082] S610. Obtaining a global feature processing result output by the target network structure in the global foreground processing model, where the network structure in the global foreground processing model corresponds to the network structure in the block foreground processing model;

[0083] S620. Input the initial synthetic image, the block region feature information corresponding to each of the multiple sub-regions, and the global feature processing result into the block foreground processing model for color conversion processing to obtain the target synthetic image.

[0084] In some embodiments, the block foreground processing model may include a residual network (ResNet), VGG16, U-net, etc. The network structure in the global foreground processing model corresponds to the network structure in the block foreground processing model. For example, when the global foreground processing model is a UNet model, the block foreground processing model is also a UNet model, and the number of network layers in the two UNet models is the same.

[0085] The global feature processing result corresponding to the target network structure in the global foreground processing model is obtained. The global feature processing result contains the original feature information of the synthesized image to be processed. Therefore, the global feature processing result is also input into the block foreground processing model, so that the block foreground processing model can utilize the original feature information of the synthesized image to be processed. In a specific embodiment, the target network structure can be an encoding layer, and the global feature processing result can be the output result corresponding to the encoding layer. The global feature processing result can be input into the decoding layer of the block foreground processing model to assist the decoding layer in feature processing.

[0086] When both the global foreground processing model and the block foreground processing model include multi-layer network structures, a correspondence is established between the network structure of the global foreground processing model and the network structure of the block foreground processing model having the same feature dimension. The encoding layer in the global foreground processing model corresponds to the encoding layer in the block foreground processing model, and the decoding layer in the global foreground processing model corresponds to the decoding layer in the block foreground processing model. Since the encoding layer and the decoding layer in the block foreground processing model also have a correspondence, the global feature processing result output by the encoding layer in the global foreground processing model can be input into the corresponding decoding layer in the block foreground processing model.

[0087] The global feature processing results output by the target network structure in the global processing stage are input into the block foreground processing model, so that the block foreground processing model can make full use of the original feature information of the image to be processed, thereby improving the accuracy of the block foreground processing model.

[0088] In some embodiments, see Figure 7 The block foreground processing model includes multiple sequentially arranged block coding layers, a block decoding layer corresponding to each block coding layer, and a feature fusion layer. The initial global feature information output by each global coding layer in the global foreground processing model of the dimension corresponding to each block coding layer is used as the global feature processing result. The initial synthetic image, the block region feature information corresponding to each of the multiple sub-regions, and the global feature processing result are input into the block foreground processing model for color conversion processing. The target synthetic image includes:

[0089] S710. Input the initial synthesized image and the block region feature information into a plurality of sequentially arranged block coding layers for feature processing to obtain the initial block feature information output by each block coding layer;

[0090] S720. When the current block decoding layer is the first block decoding layer, input the first initial block feature information and the first initial global feature information into the first block decoding layer for feature processing to obtain target block feature information output by the first block decoding layer, the first initial block feature information being information output by the last block coding layer, and the first initial global feature information being information output by a global coding layer of the same dimension as the last block coding layer;

[0091] S730. When the current block decoding layer is not the first block decoding layer, obtain the previous target block feature information corresponding to the current block decoding layer;

[0092] S740. Input the previous target block feature information, the current initial block feature information, and the second initial global feature information into the current block decoding layer for feature processing, and obtain the target block feature information output by the current block decoding layer. The current initial block feature information is the information output by the current block encoding layer corresponding to the current block decoding layer, and the second initial global feature information is the information output by the global encoding layer of the same dimension as the current block encoding layer;

[0093] S750. Input the target block feature information output by each block decoding layer into the feature fusion layer for feature fusion to obtain the target composite image.

[0094] In some embodiments, before inputting the initial composite image and the block region feature information corresponding to each of the plurality of subregions into the first block coding layer, the initial composite image and the block region feature information corresponding to each of the plurality of subregions may be spliced together to obtain second spliced feature information. When the block coding layer currently undergoing feature processing is the first block coding layer, the second spliced feature information is input into the first block coding layer for feature processing.

[0095] When the block coding layer currently performing feature processing is not the first block coding layer, the initial global feature information output by each block coding layer is used as the input information of the next block coding layer to obtain the initial block feature information output by the block coding layer currently performing feature processing.

[0096] When the current block decoding layer is the first block decoding layer, the initial global feature information output by the global coding layer of the same dimension as the last block coding layer is used as the first initial global feature information. The initial block feature information output by the last block coding layer and the first initial global feature information are input into the first block decoding layer for feature processing. The target block feature information output by the first block decoding layer can be obtained. The initial block feature information output by the last block coding layer is the first initial block feature information. The same dimension can refer to having the same feature resolution.

[0097] When the current block decoding layer is not the first block decoding layer, obtain the previous target block feature information corresponding to the current block decoding layer, and obtain the current initial block feature information output by the current block coding layer corresponding to the current block decoding layer, and the initial global feature information output by the global coding layer of the same dimension as the current block coding layer as the second initial global feature information. The previous target block feature information, the current initial block feature information, and the second initial global feature information are input into the current block decoding layer for feature processing, and the target block feature information output by the current block decoding layer can be obtained. The same dimension can refer to having the same feature resolution.

[0098] The target block feature information output by each block decoding layer is input into the feature fusion layer for feature fusion to obtain the target composite image. Figure 8 ,like Figure 8 Figure 2 shows a schematic diagram of feature fusion performed by the feature fusion layer in the block foreground processing model. Feature fusion uses the concat function to combine the target block feature information output by each block decoding layer into a single target feature information. Feature processing is then performed on this target feature information to produce a target composite image.

[0099] When performing feature processing in the block coding layer, the feature resolution corresponding to each block coding layer gradually decreases, while the number of channels gradually increases. A convolution operation can be used to abstract the high-level features in the second spliced feature information, obtaining the initial block feature information corresponding to each block coding layer. When performing feature processing in the block decoder, the feature resolution corresponding to each block decoding layer gradually increases, while the number of channels gradually decreases. A deconvolution operation can be used to obtain the target composite image with the target feature resolution.

[0100] The block foreground processing model is used to perform feature processing on the initial synthetic image and the block region feature information corresponding to each of the multiple sub-regions, and the target block feature information output by each block decoding layer is fused. This allows the block foreground processing model to adaptively perform different degrees of color transformation processing on sub-regions corresponding to different colors, thereby improving the authenticity of the target synthetic image.

[0101] In some embodiments, see Figure 9 , the method further comprises:

[0102] S910 determines the foreground difference image corresponding to the sample image, the foreground difference image foreground training feature information and the foreground region feature information corresponding to the sample image differ, the foreground training feature information is obtained by feature extraction of the foreground region of the foreground difference image;

[0103] S920. Based on the foreground training feature information, perform color transformation on the foreground difference image to obtain an initial training composite image;

[0104] S930. Based on the foreground training feature information, the foreground area of the foreground difference image is divided into regions to obtain multiple training sub-regions;

[0105] S940. Input the initial training synthetic image and the block region feature information corresponding to each of the multiple training sub-regions into the to-be-trained model for color conversion processing to obtain a target training synthetic image;

[0106] S950. Based on the target training synthetic image and the sample image, the to-be-trained model is trained to obtain a block foreground processing model.

[0107] In some embodiments, a random color transformation is performed on the foreground region of the sample image to obtain a foreground difference image. The random color transformation may be a simple color transformation of brightness, saturation, contrast, hue, color temperature, etc. The foreground region in the sample image matches the background region, but the foreground region in the foreground difference image obtained after the random color transformation does not match the background region. Therefore, there is a difference between the foreground training feature information of the foreground difference image and the foreground feature information corresponding to the sample image. Each sample image may correspond to at least one foreground difference image.

[0108] Feature extraction is performed on the foreground difference image to obtain foreground training feature information, which can be foreground mask feature information in the form of a binary image. Based on the foreground training feature information, global color transformation processing is performed on the foreground difference image to obtain an initial training composite image. In the case of model processing, the foreground training feature information and the foreground difference image are input into the first to-be-trained model for global color transformation processing to obtain the initial training composite image.

[0109] Using a preset clustering algorithm, the foreground region of the foreground difference image is partitioned based on the foreground training feature information to obtain multiple training sub-regions. The initial training composite image and the corresponding block feature information for each of the multiple training sub-regions are input into the second to-be-trained model for color conversion, resulting in the target training composite image.

[0110] Based on the image difference information between the initial training synthetic image and the sample image, first loss data is determined. The first loss data may be mean square error loss data. Based on the image difference information between the target training synthetic image and the sample image, second loss data is determined. The second loss data may be mean square error loss data. The first loss data and the second loss data are combined to obtain target loss data. Based on the target loss data, the first and second models to be trained are trained to obtain a global foreground processing model and a block foreground processing model.

[0111] During the model training process, a foreground difference image can be generated by performing random color transformation on the sample image, and then the foreground difference image can be used for model training. A large amount of training data can be generated with a relatively low labor cost, thereby improving the efficiency of training data generation and further improving the efficiency of model training.

[0112] In some embodiments, see Figure 10 ,like Figure 10Figure 2 shows the model structure corresponding to the synthetic image processing method. The network structure of the global foreground processing model corresponds to that of the block foreground processing model, consisting of four 4x4 convolutional layers (conv) and 4x4 deconvolutional layers (deconv). The convolutional layers are used for encoding, and the deconvolutional layers are used for decoding. In the global foreground processing model, the initial global feature information output by each convolutional layer is connected to the corresponding deconvolutional layer and serves as the input information for the corresponding deconvolutional layer, along with the target global feature information output by the previous deconvolutional layer. Similarly, in the block foreground processing model, the initial block feature information output by each convolutional layer is connected to the corresponding deconvolutional layer and serves as the input information for the corresponding deconvolutional layer, along with the target block feature information output by the previous deconvolutional layer. The global feature processing results output by the convolutional layer in the global foreground processing model can also be input to the deconvolutional layer of the corresponding block foreground processing model after a single 3x3 convolution.

[0113] Input the synthesized image to be processed and the foreground region feature information into the global foreground processing model for color conversion processing, and output the initial synthesized image. Also obtain the block region feature information corresponding to each sub-region in the foreground region of the image to be processed. Obtain the global feature processing results of the convolutional layer in the global foreground processing model, input the block region feature information, the initial synthesized image, and the global feature processing results of the convolutional layer into the block foreground processing model for color conversion processing, and output the target synthesized image. See Figure 11 ,like Figure 11 The figure shows the comparison between the processing results of the synthetic image processing method and the processing results of the real image, synthetic image and basic algorithm. Figure 11 In the example images in the first row, the foreground region is a person, consisting of clothing and hair subregions. Existing algorithms handle the hair well, but not the clothing well. However, this synthetic image processing method handles both well. In the example images in the second row, the foreground region is a stone pillar and a wreath, consisting of the stone pillar and wreath subregions. Existing algorithms handle the stone pillar well, but not the wreath well. However, this synthetic image processing method handles both well.

[0114] An embodiment of the present application provides a synthetic image processing method, the method comprising: performing global color transformation processing on the foreground area of the synthetic image to be processed to obtain an initial synthetic image, and dividing the foreground area of the synthetic image to be processed into regions to obtain multiple sub-regions. Inputting the block region feature information corresponding to the initial synthetic image and the multiple sub-regions into a block foreground processing model, adaptively performing color transformation processing of different degrees on the multiple sub-regions to obtain a target synthetic image. The method can perform color transformation processing of different degrees on different sub-regions in the foreground area of the synthetic image to be processed, so that each sub-region in the foreground area can be coordinated with the background area, thereby improving the accuracy of the color transformation processing and improving the authenticity of the target synthetic image.

[0115] The present application also provides a synthetic image processing device. Figure 12 , the device comprises:

[0116] A global processing module 1210 is configured to perform global color transformation processing on the composite image to be processed based on feature information of the foreground region of the composite image to be processed, thereby obtaining an initial composite image. The feature information of the foreground region is obtained by extracting features from the foreground region of the composite image to be processed.

[0117] A sub-region division module 1220 is used to divide the foreground region into multiple sub-regions;

[0118] The block processing module 1230 is used to input the initial composite image and the block area feature information corresponding to each of the multiple sub-areas into the block foreground processing model for color transformation processing to obtain a target composite image. Each sub-area in the target composite image corresponds to a different degree of color transformation processing result.

[0119] In some embodiments, the global processing module 1210 includes:

[0120] The first model processing unit is used to input the synthesized image to be processed and the foreground area feature information into the global foreground processing model to perform global color transformation processing to obtain an initial synthesized image.

[0121] In some embodiments, the block processing module includes:

[0122] A global information acquisition unit is used to obtain a global feature processing result output by a target network structure in a global foreground processing model, wherein the network structure in the global foreground processing model corresponds to the network structure in the block foreground processing model;

[0123] The second model processing unit is used to input the initial synthetic image, the block region feature information corresponding to each of the multiple sub-regions and the global feature processing result into the block foreground processing model for color conversion processing to obtain the target synthetic image.

[0124] In some embodiments, the block foreground processing model includes a plurality of sequentially arranged block coding layers, a block decoding layer corresponding to each block coding layer, and a feature fusion layer. The initial global feature information output by each global coding layer in the global foreground processing model of the dimension corresponding to each block coding layer is used as the global feature processing result. The second model processing unit includes:

[0125] A block coding processing unit is used to input the initial synthesized image and the block region feature information into a plurality of sequentially arranged block coding layers for feature processing, thereby obtaining initial block feature information output by each block coding layer;

[0126] A first block decoding processing unit is configured to input the first initial block feature information and the first initial global feature information into the first block decoding layer for feature processing when the current block decoding layer is the first block decoding layer, to obtain target block feature information output by the first block decoding layer, where the first initial block feature information is information output by the last block coding layer, and the first initial global feature information is information output by the global coding layer of the same dimension as the last block coding layer;

[0127] The previous block information acquisition unit is used to acquire the previous target block feature information corresponding to the current block decoding layer when the current block decoding layer is not the first block decoding layer;

[0128] A second block decoding processing unit is configured to input the previous target block feature information, the current initial block feature information, and the second initial global feature information into the current block decoding layer for feature processing, thereby obtaining the target block feature information output by the current block decoding layer, wherein the current initial block feature information is information output by the current block coding layer corresponding to the current block decoding layer, and the second initial global feature information is information output by the global coding layer of the same dimension as the current block coding layer;

[0129] The feature fusion unit is used to input the target block feature information output by each block decoding layer into the feature fusion layer for feature fusion to obtain the target composite image.

[0130] In some embodiments, the sub-region division module 1220 includes:

[0131] An initial center determination unit, configured to determine a plurality of initial block center information based on image feature information corresponding to a foreground area;

[0132] a current center determination unit, configured to use the multiple initial block center information as the current block center information;

[0133] A current block region determining unit, configured to determine a current block region corresponding to each current block center information based on the foreground region and the current block center information;

[0134] A current center updating unit, configured to update the current block center information corresponding to the current block area based on the current block area;

[0135] a repeating execution unit, configured to repeat the steps of determining, based on the foreground area and the current block center information, the current block area corresponding to each current block center information, and updating, based on the current block area, the current block center information corresponding to the current block area, until the current block center information before and after the update matches, and using the current block center information as the target block center information;

[0136] The sub-region determining unit is configured to use the current block region corresponding to the target block center information as a sub-region.

[0137] In some embodiments, the global foreground processing model includes a plurality of sequentially arranged global encoding layers and a global decoding layer corresponding to each global encoding layer, and the first model processing unit includes:

[0138] A global coding processing unit, configured to input the synthesized image to be processed and the feature information of the foreground area into a plurality of sequentially arranged global coding layers for feature processing, thereby obtaining initial global feature information output by each global coding layer;

[0139] A first global decoding processing unit is configured to, when the current global decoding layer is the first global decoding layer, input the third initial global feature information into the first global decoding layer for feature processing to obtain target global feature information output by the first global decoding layer, where the third initial global feature information is information output by the last global encoding layer;

[0140] A previous global information acquisition unit, configured to acquire the previous target global feature information corresponding to the current global decoding layer when the current global decoding layer is not the first global decoding layer;

[0141] A second global decoding processing unit is configured to input the previous target global feature information and the fourth initial global feature information into the current global decoding layer for feature processing, thereby obtaining the target global feature information output by the current global decoding layer, where the fourth initial global feature information is information output by the global encoding layer corresponding to the current global decoding layer;

[0142] The initial synthesized image determining unit is used to obtain an initial synthesized image based on the target global feature information output by the last global decoding layer.

[0143] In some embodiments, the apparatus further comprises:

[0144] a foreground difference image acquisition module, configured to determine a foreground difference image corresponding to the sample image, wherein there is a difference between foreground training feature information of the foreground difference image and feature information of the foreground region corresponding to the sample image, and the foreground training feature information is obtained by extracting features from the foreground region of the foreground difference image;

[0145] A global training processing module is used to perform global color transformation processing on the foreground difference image based on the foreground training feature information to obtain an initial training synthetic image;

[0146] A training sub-region division module is used to divide the foreground region of the foreground difference image into regions based on the foreground training feature information to obtain multiple training sub-regions;

[0147] The block training processing module is used to input the initial training synthetic image and the block region feature information corresponding to each of the multiple training sub-regions into the to-be-trained model for color conversion processing to obtain the target training synthetic image;

[0148] The model training module is used to train the to-be-trained model based on the target training synthetic image and the sample image to obtain a block foreground processing model.

[0149] The apparatus provided in the above embodiments can execute the method provided in any embodiment of the present application, and has the corresponding functional modules and beneficial effects of executing the method. For technical details not fully described in the above embodiments, please refer to a synthetic image processing method provided in any embodiment of the present application.

[0150] This embodiment further provides a computer-readable storage medium, in which computer-executable instructions are stored. The computer-executable instructions are loaded by a processor and execute the above-mentioned synthetic image processing method of this embodiment.

[0151] This embodiment also provides a computer program product or computer program, which includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations of the aforementioned composite image processing.

[0152] This embodiment further provides an electronic device, which includes a processor and a memory, wherein the memory stores a computer program, and the computer program is suitable for being loaded by the processor and executing the above-mentioned synthetic image processing method of this embodiment.

[0153] The device may be a computer terminal, a mobile terminal or a server, and the device may also participate in constituting the apparatus or system provided in the embodiments of the present application. Figure 13 As shown, the server 13 may include one or more (illustrated as 1302a, 1302b, ..., 1302n in the figure) processors 1302 (the processor 1302 may include but is not limited to a microprocessor MCU or a programmable logic device FPGA, etc.), a memory 1304 for storing data, and a transmission device 1306 for communication functions. In addition, it may also include: an input / output interface (I / O interface) and a network interface. It will be understood by those skilled in the art that Figure 13 The structure shown is only for illustration and does not limit the structure of the above electronic device. Figure 13 More or fewer components than shown, or with Figure 13 Different configurations shown.

[0154] It should be noted that the one or more processors 1302 and / or other data processing circuitry described above may generally be referred to herein as "data processing circuitry." The data processing circuitry may be embodied in whole or in part as software, hardware, firmware, or any other combination thereof. Furthermore, the data processing circuitry may be a single, standalone processing module, or may be fully or partially integrated into any of the other components of the server 13.

[0155] The memory 1304 can be used to store software programs and modules of application software, such as program instructions / data storage devices corresponding to the methods described in the embodiments of the present application. The processor 1302 executes various functional applications and data processing by running the software programs and modules stored in the memory 1304, thereby realizing the above-mentioned method for generating a temporal behavior capture frame based on a self-attention network. The memory 1304 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1304 may further include a memory remotely located relative to the processor 1302, and these remote memories may be connected to the mobile device 13 via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0156] Transmission device 1306 is used to receive or send data via a network. Specific examples of the aforementioned network may include a wireless network provided by the communication provider of server 13. In one embodiment, transmission device 1306 includes a network interface controller (NIC), which can be connected to other network devices via a base station to communicate with the Internet.

[0157] This specification provides method operation steps as described in the embodiments or flowcharts, but more or fewer operation steps may be included based on routine or non-creative work. The steps and order listed in the embodiments are only one way of executing the order of many steps and do not represent the only execution order. When an actual system or interrupt product is executed, it can be executed sequentially or in parallel according to the method shown in the embodiments or the drawings (for example, in a parallel processor or multi-threaded processing environment).

[0158] The structure shown in this embodiment is only a partial structure related to the scheme of the present application, and does not constitute a limitation on the device to which the scheme of the present application is applied. The specific device may include more or fewer components than shown, or combine certain components, or have different arrangements of components. It should be understood that the methods, devices, etc. disclosed in this embodiment can be implemented in other ways. For example, the device embodiment described above is only schematic. For example, the division of the modules is only a division of logical functions. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or unit modules.

[0159] Based on this understanding, the technical solution of this application, or the portion that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for causing a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of this application. The aforementioned storage medium includes: a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., various media that can store program code.

[0160] Those skilled in the art may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed in this specification can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the above description has generally described the composition and steps of each example according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0161] As described above, the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A synthetic image processing method, characterized in that: The method comprises: performing a global color transformation process on the synthesized image to be processed based on feature information of a foreground region of the synthesized image to be processed to obtain an initial synthesized image, wherein the feature information of the foreground region is obtained by extracting features from the foreground region of the synthesized image to be processed; the global color transformation process is performed using a global foreground processing model; Divide the foreground area into multiple sub-areas; the colors corresponding to the pixel information in each sub-area are within the same color value range, and different sub-areas correspond to different color value ranges; Obtaining a global feature processing result output by a target network structure in the global foreground processing model, wherein the network structure in the global foreground processing model corresponds to the network structure in the block foreground processing model; The initial composite image, the block region feature information corresponding to each of the multiple sub-regions, and the global feature processing result are input into the block foreground processing model for color transformation processing to obtain a target composite image; each sub-region in the target composite image corresponds to a color transformation processing result of different degrees.

2. The synthetic image processing method according to claim 1, wherein: The performing global color transformation processing on the synthesized image to be processed based on the foreground region feature information of the synthesized image to be processed to obtain an initial synthesized image includes: The synthesized image to be processed and the foreground region feature information are input into a global foreground processing model for global color transformation processing to obtain the initial synthesized image.

3. The synthetic image processing method according to claim 1, wherein: The block foreground processing model includes a plurality of block coding layers arranged in sequence, a block decoding layer and a feature fusion layer corresponding to each block coding layer, using the initial global feature information output by each global coding layer in the global foreground processing model of the dimension corresponding to each block coding layer as the global feature processing result, and inputting the initial synthesized image, the block region feature information corresponding to each of the plurality of sub-regions, and the global feature processing result into the block foreground processing model for color conversion processing to obtain the target synthesized image, comprising: Inputting the initial synthesized image and the block region feature information into the plurality of sequentially arranged block coding layers for feature processing, thereby obtaining initial block feature information output by each block coding layer; When the current block decoding layer is the first block decoding layer, inputting the first initial block feature information and the first initial global feature information into the first block decoding layer for feature processing to obtain target block feature information output by the first block decoding layer, where the first initial block feature information is information output by the last block coding layer, and the first initial global feature information is information output by a global coding layer of the same dimension as the last block coding layer; When the current block decoding layer is not the first block decoding layer, obtaining the previous target block feature information corresponding to the current block decoding layer; Inputting the previous target block feature information, the current initial block feature information, and the second initial global feature information into the current block decoding layer for feature processing, thereby obtaining the target block feature information output by the current block decoding layer, wherein the current initial block feature information is information output by the current block encoding layer corresponding to the current block decoding layer, and the second initial global feature information is information output by the global encoding layer of the same dimension as the current block encoding layer; The target block feature information output by each block decoding layer is input into the feature fusion layer for feature fusion to obtain the target composite image.

4. The synthetic image processing method according to claim 1, wherein: The foreground area is divided into multiple sub-areas, including: Determining a plurality of initial block center information based on image feature information corresponding to the foreground area; Using the multiple initial block center information as current block center information; Determine, based on the foreground area and the current block center information, a current block area corresponding to each current block center information; Based on the current block area, updating the current block center information corresponding to the current block area; Repeating the steps of determining the current block area corresponding to each current block center information based on the foreground area and the current block center information and updating the current block center information corresponding to the current block area based on the current block area until the current block center information before and after the update matches, and using the current block center information as the target block center information; The current block area corresponding to the target block center information is used as a sub-area.

5. The synthetic image processing method according to claim 2, wherein: The global foreground processing model includes a plurality of sequentially arranged global coding layers and a global decoding layer corresponding to each global coding layer. Inputting the synthesized image to be processed and the foreground region feature information into the global foreground processing model for performing global color transformation processing to obtain the initial synthesized image includes: Inputting the to-be-processed synthesized image and the foreground region feature information into the plurality of sequentially arranged global coding layers for feature processing, thereby obtaining initial global feature information output by each global coding layer; When the current global decoding layer is the first global decoding layer, inputting the third initial global feature information into the first global decoding layer for feature processing to obtain target global feature information output by the first global decoding layer, where the third initial global feature information is information output by the last global encoding layer; When the current global decoding layer is not the first global decoding layer, obtaining the previous target global feature information corresponding to the current global decoding layer; Inputting the previous target global feature information and the fourth initial global feature information into the current global decoding layer for feature processing to obtain the target global feature information output by the current global decoding layer, wherein the fourth initial global feature information is information output by the global encoding layer corresponding to the current global decoding layer; The initial synthesized image is obtained based on the target global feature information output by the last global decoding layer.

6. The synthetic image processing method according to claim 1, wherein: The method further comprises: Determining a foreground difference image corresponding to the sample image, wherein there is a difference between foreground training feature information of the foreground difference image and feature information of a foreground region corresponding to the sample image, the foreground training feature information being obtained by extracting features from the foreground region of the foreground difference image; Based on the foreground training feature information, performing global color transformation processing on the foreground difference image to obtain an initial training synthetic image; Based on the foreground training feature information, the foreground area of the foreground difference image is divided into regions to obtain a plurality of training sub-regions; Inputting the initial training synthetic image and the block region feature information corresponding to each of the plurality of training sub-regions into the to-be-trained model for color conversion processing to obtain a target training synthetic image; The model to be trained is trained based on the target training synthetic image and the sample image to obtain the block foreground processing model.

7. A synthetic image processing device, characterized in that: The device comprises: a global processing module configured to perform a global color transformation process on the composite image to be processed based on feature information of a foreground region of the composite image to be processed, to obtain an initial composite image, wherein the feature information of the foreground region is obtained by extracting features from the foreground region of the composite image to be processed; and the global color transformation process is performed using a global foreground processing model; A sub-region division module is used to divide the foreground region into multiple sub-regions; the colors corresponding to the pixel information in each sub-region are within the same color value range, and different sub-regions correspond to different color value ranges; a block processing module for inputting the initial composite image and the feature information of the block regions corresponding to the plurality of subregions into a block foreground processing model to perform color transformation processing, thereby obtaining a target composite image, wherein each subregion in the target composite image corresponds to a color transformation processing result of a different degree; The block processing module includes: a global information acquisition unit, configured to acquire a global feature processing result output by a target network structure in the global foreground processing model, wherein the network structure in the global foreground processing model corresponds to the network structure in the block foreground processing model; The second model processing unit is used to input the initial synthetic image, the block region feature information corresponding to each of the multiple sub-regions and the global feature processing result into the block foreground processing model for color conversion processing to obtain the target synthetic image.

8. The device according to claim 7, characterized in that The global processing module includes: The first model processing unit is configured to input the synthesized image to be processed and the foreground region feature information into a global foreground processing model to perform global color transformation processing to obtain the initial synthesized image.

9. The device according to claim 7, characterized in that The block foreground processing model includes a plurality of sequentially arranged block coding layers, a block decoding layer corresponding to each block coding layer, and a feature fusion layer. The initial global feature information output by each global coding layer in the global foreground processing model of the dimension corresponding to each block coding layer is used as the global feature processing result. The second model processing unit includes: A block coding processing unit, configured to input the initial synthesized image and the block region feature information into the plurality of sequentially arranged block coding layers for feature processing, thereby obtaining initial block feature information output by each block coding layer; A first block decoding processing unit is configured to, when the current block decoding layer is the first block decoding layer, input first initial block feature information and first initial global feature information into the first block decoding layer for feature processing to obtain target block feature information output by the first block decoding layer, where the first initial block feature information is information output by the last block coding layer, and the first initial global feature information is information output by a global coding layer of the same dimension as the last block coding layer; A previous block information acquisition unit, configured to acquire, when the current block decoding layer is not the first block decoding layer, previous target block feature information corresponding to the current block decoding layer; A second block decoding processing unit is configured to input the previous target block feature information, the current initial block feature information, and the second initial global feature information into the current block decoding layer for feature processing to obtain the target block feature information output by the current block decoding layer, wherein the current initial block feature information is information output by the current block coding layer corresponding to the current block decoding layer, and the second initial global feature information is information output by the global coding layer of the same dimension as the current block coding layer; The feature fusion unit is used to input the target block feature information output by each block decoding layer into the feature fusion layer for feature fusion to obtain the target composite image.

10. The device according to claim 7, characterized in that The sub-area division module includes: an initial center determination unit, configured to determine a plurality of initial block center information based on image feature information corresponding to the foreground area; a current center determining unit, configured to use the plurality of initial block center information as current block center information; a current block region determining unit, configured to determine a current block region corresponding to each current block center information based on the foreground region and the current block center information; A current center updating unit, configured to update current block center information corresponding to the current block area based on the current block area; a repeating execution unit, configured to repeat the steps of determining the current block area corresponding to each current block center information based on the foreground area and the current block center information, and updating the current block center information corresponding to the current block area based on the current block area, until the current block center information before and after the update matches, and use the current block center information as the target block center information; The sub-region determining unit is configured to use the current block region corresponding to the target block center information as a sub-region.

11. The device according to claim 8, characterized in that The global foreground processing model includes a plurality of sequentially arranged global coding layers and a global decoding layer corresponding to each global coding layer. The first model processing unit includes: a global coding processing unit, configured to input the synthesized image to be processed and the foreground region feature information into the plurality of sequentially arranged global coding layers for feature processing, thereby obtaining initial global feature information output by each global coding layer; a first global decoding processing unit, configured to, when the current global decoding layer is the first global decoding layer, input the third initial global feature information into the first global decoding layer for feature processing to obtain target global feature information output by the first global decoding layer, where the third initial global feature information is information output by the last global encoding layer; a previous global information acquisition unit, configured to acquire the previous target global feature information corresponding to the current global decoding layer when the current global decoding layer is not the first global decoding layer; a second global decoding processing unit, configured to input the previous target global feature information and the fourth initial global feature information into the current global decoding layer for feature processing, to obtain the target global feature information output by the current global decoding layer, where the fourth initial global feature information is information output by the global encoding layer corresponding to the current global decoding layer; The initial synthesized image determining unit is configured to obtain the initial synthesized image based on the target global feature information output by the last global decoding layer.

12. The device according to claim 7, characterized in that The device further comprises: a foreground difference image acquisition module, configured to determine a foreground difference image corresponding to a sample image, wherein there is a difference between foreground training feature information of the foreground difference image and feature information of a foreground region corresponding to the sample image, the foreground training feature information being obtained by extracting features from the foreground region of the foreground difference image; A global training processing module, configured to perform a global color transformation process on the foreground difference image based on the foreground training feature information to obtain an initial training composite image; a training sub-region division module, configured to divide the foreground region of the foreground difference image into regions based on the foreground training feature information to obtain a plurality of training sub-regions; A block training processing module is used to input the initial training synthetic image and the block region feature information corresponding to each of the multiple training sub-regions into the to-be-trained model for color conversion processing to obtain a target training synthetic image; The model training module is used to train the model to be trained based on the target training synthetic image and the sample image to obtain the block foreground processing model.

13. An electronic device, characterized in that: The electronic device includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement a synthetic image processing method as described in any one of claims 1-6.

14. A computer-readable storage medium, characterized in that The storage medium includes a processor and a memory, wherein the memory stores at least one instruction or at least one program, and the at least one instruction or the at least one program is loaded and executed by the processor to implement a synthetic image processing method as described in any one of claims 1-6.

15. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the composite image processing method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Image harmonious synthesis method based on color constancy

    CN113222875A

  • Image processing method and device, and storage medium

    US20210118112A1