Image processing methods, apparatus, terminal devices, and readable storage media

By using the identity features of the image to be processed to correct the replaced image during the image style replacement process, the problem of changes in the identity attributes of the target object is solved, and the identity attributes of the target object remain unchanged after style replacement, thus improving the accuracy and quality of image processing.

CN113850715BActive Publication Date: 2025-10-31BEIJING QIYI CENTURY SCI & TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111154576.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-09-29
Publication Date
2025-10-31
Estimated Expiration
2041-09-29

AI Technical Summary

Technical Problem

Existing technologies cannot accurately extract local image features during image style replacement, leading to changes in the identity attributes of the target object and affecting the recognition effect.

Method used

The style of the image to be processed is replaced based on the target reference image, and the identity features of the target object in the image to be processed are used to correct the replaced image, so as to maximize the preservation of the identity attributes of the target object. The feature fusion and supplementation are performed by mask guidance.

Benefits of technology

It effectively preserves the identity attributes of the target object, improves the accuracy and quality of image processing, and ensures that the recognition effect of the target object is not affected before and after replacement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113850715B_ABST
    Figure CN113850715B_ABST
Patent Text Reader

Abstract

This application provides an image processing method, apparatus, terminal device, and readable storage medium. The method includes: performing style replacement on an image to be processed based on a target reference image to obtain a replaced image; and, with the goal of maximizing the similarity between the identity attributes of a target object in the replaced image and the identity attributes of the target object in the image to be processed, correcting the replaced image using the identity features of the target object in the image to be processed to obtain a target image. Through this method, after style replacement of a target object in any image, the identity features of the target object in the image before replacement can be used to supplement the lost identity features in the replaced image, thereby ensuring that the identity attributes of the target object remain unchanged before and after replacement.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image processing technology, and in particular to an image processing method, apparatus, terminal device, and readable storage medium. Background Technology

[0002] In the field of image processing, style transformation has always received widespread attention. For an image containing a target object, there are many types of styles, such as the shape, color, size, background, clothing, hairstyle, etc. of the target object. In related technologies, when replacing a certain style in the image to be processed with the corresponding style in a reference image, the image features of the corresponding region in the reference image are directly extracted and replaced in the image to be processed. However, there is no style extractor in related technologies that can completely and accurately extract the image features of local regions. This means that in the image after feature replacement, other regions that do not need style replacement will also be affected (e.g., distortion or color difference).

[0003] In some application scenarios, it is necessary to preserve the identity attributes of the target object after style replacement, and the above methods are clearly insufficient. For example, replacing a person's hairstyle may result in facial distortion, or replacing a vehicle's background may alter the vehicle's shape. In these cases, the attributes of the person or vehicle itself change, making it difficult for the person or device to be accurately recognized. Therefore, how to preserve the identity attributes of the target object during style replacement has become a pressing problem to be solved. Summary of the Invention

[0004] This application provides an image processing method, apparatus, terminal device, and readable storage medium, aiming to maintain the identity attributes of the target object unchanged before and after style replacement. The specific technical solution is as follows:

[0005] In a first aspect of this application, an image processing method is provided, the method comprising:

[0006] The style of the image to be processed is replaced based on the target reference image to obtain the replaced image;

[0007] With the goal of maximizing the similarity between the identity attributes of the target object in the replaced image and the identity attributes of the target object in the image to be processed, the replaced image is modified using the identity features of the target object in the image to be processed to obtain the target image.

[0008] Optionally, the replaced image is corrected using the identity features of the target object in the image to be processed to obtain the target image, including:

[0009] The image features of the first identity region of the target object are used as the identity features of the target object, and the first identity region includes features that can uniquely identify the target object;

[0010] By using a mask-guided method, the image features of the first identity region of the target object in the image to be processed and the image features of the first identity region of the target object in the replaced image are fused to obtain the target image.

[0011] Optionally, the replaced image is corrected using the identity features of the target object in the image to be processed to obtain the target image, including:

[0012] The image features of the first identity region of the target object and the image features of the second identity region of the target object are used as the identity features of the target object, wherein the second identity region includes the background region of the target object;

[0013] By using a mask-guided method, image features of the first identity region of the target object in the image to be processed, the second identity region of the target object in the image to be processed, the first identity region of the target object in the replaced image, and the second identity region of the target object in the replaced image are fused to obtain the target image.

[0014] Optionally, style replacement is performed on the image to be processed based on the target reference image to obtain a replaced image, including:

[0015] The first network in the image generation model is used to extract features from the target reference image to obtain the image features of each preset style region in the target reference image;

[0016] The image generation model uses a generative network to perform feature fusion on the image features of each preset style region of the target reference image and the image features of each preset style region of the image to be processed, according to a preset replacement target, to obtain a replaced image. The preset replacement target represents the weight of the image features of each preset style region when performing feature fusion.

[0017] Optionally, the replaced image is corrected using the identity features of the target object in the image to be processed to obtain the target image, including:

[0018] The image to be processed is used to extract features from the image to be processed by the second network in the image generation model to obtain the identity features of the target object in the image to be processed. The second network is different from the first network.

[0019] The replacement image is corrected using the identity features of the target object in the image to be processed through the generation network to obtain the target image.

[0020] Optionally, the generating network includes an encoder and a decoder. The encoder is used to acquire image features of each preset style region of the image to be processed. The decoder takes the output of the first network, the output of the encoder, and the output of the second network as input. The second network includes a first branch network and a second branch network.

[0021] Correcting the replaced image using the identity features of the target object in the image to be processed includes:

[0022] The decoder uses image features of the first identity region of the target object in the image to be processed, output by the first branch network, to supplement the features of the replaced image; or

[0023] The decoder uses the image features of the first identity region of the target object in the image to be processed output by the first branch network and the image features of the second identity region of the target object in the image to be processed output by the second branch network to supplement the features of the replaced image.

[0024] Optionally, the method further includes:

[0025] Obtain a first image after style replacement based on a first target reference image, and a second image after style replacement based on a second target reference image;

[0026] The loss function for the first identity region is obtained based on the absolute value error between the image of the first identity region of the target object in the first image and the image of the first identity region of the target object in the second image.

[0027] The image generation model is optimized using the loss function of the first identity region.

[0028] Optionally, the method further includes:

[0029] Based on the error in the Lab color space between the image of the second identity region of the target object in the target image and the image of the second identity region of the target object in the image to be processed, a loss function for the second identity region is obtained.

[0030] The image generation model is optimized using the loss functions of the first identity region and the second identity region.

[0031] In a second aspect of this application, an image processing apparatus is also provided, the apparatus comprising:

[0032] The replacement module is used to perform style replacement on the image to be processed based on the target reference image, so as to obtain the replaced image.

[0033] The correction module is used to correct the replaced image by utilizing the identity features of the target object in the image to be processed, with the goal of maximizing the similarity between the identity attributes of the target object in the replaced image and the identity attributes of the target object in the image to be processed, so as to obtain the target image.

[0034] Optionally, the correction module includes:

[0035] The first determining submodule is used to use the image features of the first identity region of the target object as the identity features of the target object, wherein the first identity region includes the key feature identifier of the target object;

[0036] The first acquisition submodule is used to perform feature fusion on the image features of the first identity region of the target object in the image to be processed and the image features of the first identity region of the target object in the replaced image through a mask-guided method to obtain the target image.

[0037] Optionally, the correction module includes:

[0038] The second determining submodule is used to use the image features of the first identity region of the target object and the image features of the second identity region of the target object as the identity features of the target object, wherein the second identity region includes the background region of the target object;

[0039] The second acquisition submodule is used to perform feature fusion on the image features of the first identity region of the target object in the image to be processed, the image features of the second identity region of the target object in the image to be processed, the image features of the first identity region of the target object in the replaced image, and the image features of the second identity region of the target object in the replaced image through a mask-guided method to obtain a target image.

[0040] Optionally, the replacement module includes:

[0041] The first extraction submodule is used to extract features from the target reference image through the first network in the image generation model to obtain image features of each preset style region in the target reference image;

[0042] The third submodule is used to perform feature fusion on the image features of each preset style region of the target reference image and the image features of each preset style region of the image to be processed, according to the preset replacement target, through the generative network in the image generation model, to obtain the replaced image. The preset replacement target represents the weight of the image features of each preset style region when performing feature fusion.

[0043] Optionally, the correction module includes:

[0044] The second extraction submodule is used to extract features from the image to be processed through the second network in the image generation model to obtain the identity features of the target object in the image to be processed. The second network is different from the first network.

[0045] The fourth submodule is used to modify the replaced image by using the identity features of the target object in the image to be processed through the generation network, so as to obtain the target image.

[0046] Optionally, the generating network includes an encoder and a decoder. The encoder is used to acquire image features of various preset style regions of the image to be processed. The decoder takes the output of the first network, the output of the encoder, and the output of the second network as input. The second network includes a first branch network and a second branch network. The correction module includes:

[0047] The first correction submodule is used to supplement the replaced image with features from the first identity region of the target object in the image to be processed, using the decoder and the image features output by the first branch network; or

[0048] The second correction submodule is used to supplement the replaced image with features by using the image features of the first identity region of the target object in the image to be processed output by the first branch network and the image features of the second identity region of the target object in the image to be processed output by the second branch network through the decoder.

[0049] Optionally, the device further includes:

[0050] The first acquisition module is used to acquire a first image after style replacement based on a first target reference image, and a second image after style replacement based on a second target reference image;

[0051] The second obtaining module is used to obtain a loss function for the first identity region based on the absolute value error between the image of the first identity region of the target object in the first image and the image of the first identity region of the target object in the second image;

[0052] The first optimization module is used to utilize the first identity area. domain The loss function is used to optimize the image generation model.

[0053] Optionally, the device further includes:

[0054] The third obtaining module is used to obtain a loss function for the second identity region based on the error between the image of the second identity region of the target object in the target image and the image of the second identity region of the target object in the image to be processed in the Lab color space.

[0055] The second optimization module is used to optimize the image generation model using the loss function of the first identity region and the loss function of the second identity region.

[0056] In a third aspect of the embodiments of this application, a terminal device is also provided, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;

[0057] Memory, used to store computer programs;

[0058] When a processor executes a program stored in memory, it implements the steps of the image processing method described in the first aspect of the embodiments of this application.

[0059] In a fourth aspect of the embodiments of this application, a computer-readable storage medium is also provided, wherein instructions are stored therein, which, when executed on a computer, cause the computer to perform the steps in any of the image processing methods described above.

[0060] The image processing method of this application first performs style replacement on the image to be processed based on the target reference image to obtain the replaced image. Then, with the goal of maximizing the similarity between the identity attributes of the target object in the replaced image and the identity attributes of the target object in the image to be processed, the replaced image is corrected using the identity features of the target object in the image to be processed, resulting in the target image. Through this method, after style replacement of a target object in any image, the identity features of the target object in the image before replacement can be used to supplement the lost identity features in the replaced image, thereby ensuring that the identity attributes of the target object remain unchanged before and after replacement. Attached Figure Description

[0061] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below.

[0062] Figure 1 This is a flowchart illustrating an image processing method according to an embodiment of this application;

[0063] Figure 2 This is a schematic diagram of the structure of an image generation model according to an embodiment of this application;

[0064] Figure 3 This is a structural block diagram of an image processing apparatus according to an embodiment of this application;

[0065] Figure 4 This is a schematic diagram of the structure of a terminal device shown in one embodiment of this application. Detailed Implementation

[0066] The technical solutions in the embodiments of this application will now be described with reference to the accompanying drawings.

[0067] Figure 1 This is a flowchart illustrating an image processing method according to an embodiment of this application. (In conjunction with...) Figure 1 The image processing method of this application may specifically include the following steps:

[0068] Step S11: Perform style replacement on the image to be processed based on the target reference image to obtain the replaced image.

[0069] In this embodiment, the image to be processed contains a target object, which can be a human figure or a non-human object (such as an animal, vehicle, etc.). When the target object is a human figure, the style can refer to the human's hairstyle, face, background, etc. When the target object is a non-human figure, the style can refer to the object's face, background, shape, color, hair, etc. Of course, styles can also include other types, and this embodiment does not impose specific limitations on them.

[0070] Style replacement refers to replacing one or more styles of a target object in an image with the corresponding style in another image (target reference image). For example, replacing the hairstyle of a person in image A with the hairstyle of a person in image B, or replacing the background of a vehicle in image A with the background of a vehicle in image B.

[0071] In this embodiment, the style of the target object in the image to be processed can be replaced with the corresponding style of the target object in the target reference image in any way. This embodiment does not impose specific restrictions on the replacement method.

[0072] Step S12: With the goal of maximizing the similarity between the identity attributes of the target object in the replaced image and the identity attributes of the target object in the image to be processed, the replaced image is modified using the identity features of the target object in the image to be processed to obtain the target image.

[0073] In this embodiment, if the task is simply to replace a certain style, such as simply replacing the hairstyle of a person in the image to be processed with the hairstyle of a person in the target reference image, or simply replacing the background of a vehicle in the image to be processed with the background of a vehicle in the target reference image, after the replacement, image areas that do not need to be replaced may be affected, such as by deformation or color difference, causing the user or device to be unable to distinguish whether the target object in the replaced image is still the target object before the replacement.

[0074] Therefore, if we want to ensure that the identity attributes of the target object remain unchanged after replacement (i.e., it is still the original target object), we need to optimize the image after replacement. In this embodiment, we can use the features in the image to be processed before replacement that can characterize the identity attributes of the target object to supplement the missing features in the image after replacement that can characterize the identity attributes of the target object.

[0075] Different target objects have different identity attributes, which can be reflected using features that uniquely represent the target object. For example, user A and user B have different identity attributes, and facial features can represent the uniqueness of a person (the facial features of user A and user B are clearly different). Therefore, user A's facial features can be used as user A's identity feature, and user B's facial features can be used as user B's identity feature. As another example, vehicle A's shape, color, license plate, etc., can be used as vehicle A's identity feature, and vehicle B's shape, color, license plate, etc., can be used as vehicle B's identity feature.

[0076] In this embodiment, to ensure that the identity attributes of the target object remain unchanged after style replacement, the following approach can be adopted: The replaced image is modified using the identity features of the target object in the image to be processed until the similarity between the identity attributes of the target object in the replaced image and the identity attributes of the target object in the image to be processed is maximized. In practice, maximization can be achieved using a preset value; when the similarity of the target object's identity attributes between the two images exceeds the preset value, it is determined that maximization has been achieved.

[0077] For example, taking the replacement of a person's hairstyle, where facial features are used as identity features, after replacing the hairstyle of the person in the image to be processed with the hairstyle of the person in the target reference image, a replaced image is obtained. In this replaced image, the person's face is distorted, so the face shape needs to be corrected. At this point, the facial features of the person in the image to be processed before replacement can be obtained, and these facial features can be used to supplement the facial features of the person in the replaced image until the similarity between the facial features of the person before and after replacement is maximized.

[0078] To illustrate further, consider vehicle replacement where the external surface features are used as identity features. After replacing the background of the vehicle in the image to be processed with the background of the vehicle in the target reference image, the replaced image is obtained. The shape of the vehicle in the replaced image is distorted, so the shape of the vehicle needs to be corrected. At this point, the external surface features (shape, color, license plate, etc.) of the vehicle in the image to be processed before replacement can be obtained. These external surface features are then used to supplement the external surface features of the vehicle in the replaced image until the similarity between the external surface features of the target vehicle before and after replacement is maximized.

[0079] In practice, the image after replacement can be modified by using the identity features of the target object in the image to be processed in any way, and the identity attributes of the target object in the modified image can be determined in any way to maximize the similarity between the target object and the target object in the image to be processed. This embodiment does not impose specific restrictions on the modification method and the method of determining whether the similarity is maximized.

[0080] The method in this embodiment first performs style replacement on the image to be processed based on the target reference image to obtain a replaced image. Then, with the goal of maximizing the similarity between the identity attributes of the target object in the replaced image and the identity attributes of the target object in the image to be processed, the replaced image is corrected using the identity features of the target object in the image to be processed, resulting in the target image. Through this method, after replacing the style of a target object in any image, the identity features of the target object in the image before replacement can be used to supplement the lost identity features in the replaced image, thereby ensuring that the identity attributes of the target object remain unchanged before and after replacement.

[0081] In conjunction with the above embodiments, in one implementation, this embodiment provides a first method for correcting a replaced image. Specifically, the replaced image is corrected using the identity features of the target object in the image to be processed to obtain a target image, including:

[0082] The image features of the first identity region of the target object are used as the identity features of the target object, and the first identity region includes features that can uniquely identify the target object;

[0083] By using a mask-guided method, the image features of the first identity region of the target object in the image to be processed and the image features of the first identity region of the target object in the replaced image are fused to obtain the target image.

[0084] In this embodiment, the image features of the first identity region of the target object can be used as the identity features of the target object. The first identity region includes features that can uniquely identify the target object.

[0085] The solution in this embodiment is applicable to style replacement scenarios where the first identity region is not replaced. In such style replacement scenarios, the image features of the first identity region of the target object in the image to be processed are used to supplement the image features of the lost first identity region of the target object in the replaced image, which can effectively preserve the identity attributes of the target object before and after replacement.

[0086] For example, when the target object is a person, the first identity region can be the facial region, and the image features of the facial region can be the person's face shape, facial features, skin color, texture, etc. When the target object is a vehicle, the first identity region can be the vehicle's outer surface region, and the image features of the outer surface region can be the vehicle's shape, color, license plate, etc. When the target object is an animal, the first identity region can be the animal's facial region, such as face shape, facial features, fur, color, texture, etc. In specific implementations, the image features of the first identity region of the target object can be set according to actual needs; this embodiment does not impose specific limitations on this.

[0087] Mask-guided fusion can fuse image features according to a given task objective. Specifically, it can fuse the image features of the first identity region of the target object in the image to be processed with those of the first identity region of the target object in the replaced image. The task objective can be to assign a larger weight to the image features of the first identity region of the target object in the image to be processed, thereby supplementing the lost first identity region image features of the target object in the replaced image using those features.

[0088] In this embodiment, the image features of the first identity region of the target object can be directly used as the identity features of the target object, and this feature can be used to supplement the image features of the first identity region of the target object that were lost in the replaced image, which reduces the difficulty of correcting the replaced image and improves the image processing efficiency.

[0089] In conjunction with the above embodiments, in one implementation, this embodiment provides a second method for correcting a replaced image. Specifically, correcting the replaced image using the identity features of the target object in the image to be processed to obtain a target image may include:

[0090] The image features of the first identity region of the target object and the image features of the second identity region of the target object are used as the identity features of the target object, wherein the second identity region includes the background region of the target object;

[0091] By using a mask-guided method, image features of the first identity region of the target object in the image to be processed, the second identity region of the target object in the image to be processed, the first identity region of the target object in the replaced image, and the second identity region of the target object in the replaced image are fused to obtain the target image.

[0092] The solution in this embodiment is applicable to style replacement scenarios where neither the first identity region nor the second identity region is replaced. In such style replacement scenarios where neither the first identity region nor the second identity region is replaced, the image features of the first identity region and the second identity region of the target object in the image to be processed are used to supplement the lost image features of the first identity region and the second identity region of the target object in the replaced image. This can effectively preserve the identity attributes of the target object before and after replacement and improve the quality of the target image.

[0093] In this embodiment, the second identity region can be a background region, and the background region of the target object will also affect the identity characteristics of the target object. For example, if the target object is a person and the first identity region is the face region, the color, lighting, edges, etc. in the background region will all affect the facial features of the target object. If the target object is a person, the background region can refer to the entire background environment of the person object, or it can refer to other areas of the person object other than the first identity region, such as the clothing area, hand area, etc.

[0094] Therefore, in order to further improve the quality of image correction and ensure that the identity attributes of the target object remain unchanged before and after replacement, the image features of the first identity region and the image features of the second identity region can be combined to supplement the lost identity features in the replaced image.

[0095] In practical implementation, on the one hand, mask-guided feature fusion can be performed between the image features of the first identity region of the target object in the image to be processed and the image features of the first identity region of the target object in the replaced image. This allows the missing image features of the first identity region of the target object in the replaced image to be supplemented. On the other hand, mask-guided feature fusion can also be performed between the image features of the second identity region of the target object in the image to be processed and the image features of the second identity region of the target object in the replaced image. This allows the missing image features of the second identity region of the target object in the replaced image to be supplemented. Finally, the supplemented image features of the first and second identity regions are combined to obtain the corrected target image.

[0096] For example, taking a target object as a person, a first identity region as the face, and a second identity region as the background region of the entire person as an example, during feature fusion, on the one hand, mask guidance can be used to fuse the features of the face region of the person in the image to be processed with the features of the face region of the person in the replaced image, using the features of the face region of the person in the image to supplement the missing features of the face region of the person in the replaced image. On the other hand, mask guidance can also be used to fuse the features of the background region of the person in the image to be processed with the features of the background region of the person in the replaced image, using the features of the background region of the person in the image to supplement the missing features of the background region of the person in the replaced image. Finally, the supplemented features of the face region and the background region are combined to obtain the corrected target image.

[0097] For example, taking a vehicle as the vehicle object, the first identity region as the outer surface region, and the second identity region as the background region of the entire vehicle object, feature fusion can be performed in two ways. First, mask-guided feature fusion can be used to fuse the features of the vehicle's outer surface region in the image to be processed (e.g., body shape, color, license plate features) with the features of the vehicle's outer surface region in the replaced image. This supplements the missing features of the vehicle's outer surface region in the replaced image. Second, mask-guided feature fusion can be used to fuse the features of the vehicle's background region in the image to be processed with the features of the vehicle's background region in the replaced image. This supplements the missing features of the vehicle's background region in the replaced image, preventing changes in vehicle identity attributes due to color differences at the vehicle's edges. Finally, the supplemented features of the outer surface region and the background region are combined to obtain the corrected target image.

[0098] In this embodiment, the image features of the first identity region and the second identity region of the target object are combined to supplement the image features of the first identity region and the second identity region of the target object that were lost in the replaced image, thereby further improving the quality of image correction and ensuring that the identity attributes of the target object remain unchanged before and after the replacement.

[0099] In conjunction with the above embodiments, in one implementation, the image processing method of this application can be implemented through an image generation model. After the image to be processed and the target reference image are input into the image generation model, the model can output the target image. The image generation model includes at least a first network and a generation network.

[0100] Based on this image generation model, the replacement image is corrected using the identity features of the target object in the image to be processed to obtain the target image, which may include:

[0101] The first network in the image generation model is used to extract features from the target reference image to obtain the image features of each preset style region in the target reference image;

[0102] The image generation model uses a generative network to perform feature fusion on the image features of each preset style region of the target reference image and the image features of each preset style region of the image to be processed, according to a preset replacement target, to obtain a replaced image. The preset replacement target represents the weight of the image features of each preset style region when performing feature fusion.

[0103] In this embodiment, after the target reference image and the image to be processed are input into the image generation model, the image generation model extracts features from the target reference image through a first network to obtain image features of each preset style region in the target reference image. Simultaneously, the image generation model extracts features from the image to be processed through a generation network to obtain image features of each preset style region in the image to be processed.

[0104] The preset style area is a pre-defined style area. For example, when the target object is a human figure, the preset style area can be set to the face area, hairstyle area, background area, etc. Similarly, when the target object is an animal, the preset style area can be set to the face area, fur area, background area, etc. The preset style area can be divided according to the actual situation of the target object; this embodiment does not impose specific limitations on the division method.

[0105] After obtaining the image features of each preset style region in the target reference image, the first network inputs the image features of each preset style region in the target reference image into the generator network, so that the generator network performs feature fusion on the image features of each preset style region in the target reference image and the image features of each preset style region in the image to be processed according to the preset replacement target.

[0106] The preset replacement target represents the weight of image features in each preset style region during feature fusion. For example, if preset style region A in the image to be processed is to be replaced with preset style region A in the target reference image, then in the preset replacement target, the image features of preset style region A in the target reference image have the largest weight. Preset style region A can refer to one preset style region or multiple preset style regions.

[0107] In this embodiment, each replacement can replace only one preset style region or multiple preset style regions. For example, if the region to be replaced is the hairstyle region of a person, then when replacing the hairstyle region, the hairstyle of the person in the image to be processed can be replaced with the hairstyle of the person in the target reference image. As another example, if the regions to be replaced are the hairstyle region and the hat region of a person, then when replacing the hairstyle region and the hat region, the hairstyle of the person in the image to be processed can be replaced with the hairstyle of the person in the target reference image, and the hat of the person in the image to be processed can be replaced with the hat of the person in the target reference image. In this embodiment, the Adain operation can be used to fuse the image features of each preset style region of the target reference image and the image features of each preset style region of the image to be processed to obtain the replaced image.

[0108] In conjunction with the above embodiments, in one implementation, the image generation model further includes a second network for extracting the identity features of the target object in the image to be processed. Based on this, the replaced image is corrected using the identity features of the target object in the image to be processed to obtain the target image, specifically including:

[0109] The image to be processed is used to extract features from the image to be processed by the second network in the image generation model to obtain the identity features of the target object in the image to be processed. The second network is different from the first network.

[0110] The replacement image is corrected using the identity features of the target object in the image to be processed through the generation network to obtain the target image.

[0111] In this embodiment, the image generation model can pre-obtain image regions that can characterize the identity features of the target object from the image to be processed. Then, the obtained image regions are input into a second network, which extracts the identity features of the target object. These identity features are then input into a generation network, which uses these features to correct the replaced image, thus obtaining the target image. The correction process can be referred to the previous description and will not be repeated here.

[0112] In this embodiment, the image generation model includes a first network, a generator network, and a second network. After extracting image features from each preset style region of the image to be processed, the generator network can fuse these features with the output of the first network (image features of each preset style region in the target reference image) and the output of the second network (identity features of the target object) to obtain the final target image. The output of the first network can be fused into the generator using an Adain operation, and the output of the second network can be fused into the generator using a mask-guided method. The specific fusion process is described above and will not be repeated here.

[0113] In this embodiment, a second network is added to the image generation model to extract the identity features of the target object, enabling the image generation model to generate the target image in a multi-parallel manner (first network, generation network, and second network), thereby improving image processing efficiency.

[0114] In one embodiment, this application also provides an internal structure design for a second network, specifically including a first branch network and a second branch network. The first branch network is used to extract image features of a first identity region of the target object, and the second branch network is used to extract image features of a second identity region of the target object. Part of the structure of the image generation model in this application scenario can be as follows: Figure 2 As shown. Figure 2 This is a schematic diagram of the structure of an image generation model shown in one embodiment of this application.

[0115] exist Figure 2 In this model, the image generation model includes a first network, a generator network, and a second network. The generator network further includes an encoder and a decoder, which can exchange information via skip connections. The encoder is used to acquire image features of various preset style regions of the image to be processed. The second network further includes a first branch network and a second branch network. The decoder takes the outputs of the first network, the encoder, and the second network as inputs and generates the target image based on the multiple inputs.

[0116] Based on this, the replaced image is corrected using the identity features of the target object in the image to be processed, specifically including:

[0117] The decoder uses image features of the first identity region of the target object in the image to be processed, output by the first branch network, to supplement the features of the replaced image; or

[0118] The decoder uses the image features of the first identity region of the target object in the image to be processed output by the first branch network and the image features of the second identity region of the target object in the image to be processed output by the second branch network to supplement the features of the replaced image.

[0119] Without the first network, the decoder can generate an image based on the image features extracted by the encoder; this image is the original image to be processed. With the first network, the image features of each preset style region in the target reference image can be applied to the decoder through an adain operation. The decoder then performs feature fusion on the image features extracted by the encoder and the image features of each preset style region in the target reference image output by the first network to obtain the style-replaced image.

[0120] However, the image quality after style replacement obtained directly from the first network and decoder is poor, and the identity features of the target object may change. Therefore, this application further designs a second network to extract the identity features of the target object and apply it to the decoder, which can replenish the lost identity feature information in the style-replaced image, enabling the decoder to generate an image with the same identity attributes as the target object before replacement.

[0121] In practical implementation, if only the image features of the first identity region are used as identity features, the image of the first identity region of the target object in the image to be processed can be input into the first branch network for feature extraction to obtain the image features of the first identity region. Then, the features are applied to the decoder through a mask-guided method and fused with the features of the corresponding resolution layer of the decoder (when the encoder extracts image features, the number of channels increases and the resolution decreases, while the decoder is the opposite of the encoder, when generating images, the number of channels decreases and the resolution increases).

[0122] If image features of both the first and second identity regions are used as identity features simultaneously, on one hand, the image of the first identity region of the target object in the image to be processed can be input into the first branch network for feature extraction to obtain the image features of the first identity region. These features are then applied to the decoder using a mask-guided method and fused with the features of the corresponding resolution layer of the decoder. On the other hand, the image of the second identity region of the target object in the image to be processed can be input into the second branch network for feature extraction to obtain the image features of the second identity region. These features are then applied to the decoder using a mask-guided method and fused with the features of the corresponding resolution layer of the decoder.

[0123] In this embodiment, the first branch network can be composed of four residual network modules, which can extract high-level semantic features of the first identity region. For example, when the first identity region is a facial region, the high-level semantic features of the facial region can include skin color, texture, overall topological structure information, etc.

[0124] Considering the diversity and complexity of the second identity region, the second branch network can be composed of six residual network modules, which is two more residual modules than the first branch network, allowing for better feature extraction.

[0125] In this embodiment, an internal structure design of the second network is provided, enabling the decoder to correct the replaced image based on the features of the first identity region of the target object, and also to combine the image features of the first and second identity regions of the target object to correct the replaced image, thus enriching the methods for obtaining the target image. Secondly, the decoder can generate the target image using a parallel processing method with multiple inputs (output of the first network, output of the encoder, output of the first branch network, and output of the second branch network), improving image processing efficiency.

[0126] In conjunction with the above embodiments, when using an image generation model to obtain a target image, the image to be processed and the target reference image can be preprocessed first, and then the preprocessed image to be processed and the target reference image can be used as input to the image generation model.

[0127] The same preprocessing method can be used for both the image to be processed and the target reference image. The specific steps are as follows:

[0128] Align objects in an image by angle;

[0129] Crop the aligned image to the preset size.

[0130] Angle alignment can be understood as correcting angled objects. When the object is a person, angle alignment involves facial detection and key point alignment to straighten an angled face. When the object is a vehicle, angle alignment can straighten an angled vehicle.

[0131] Generally, the preset image size is matched to the computational power of the network in the image generation model. To reduce the computational load on the network, the preset size is not too large; for example, based on experience, it can be 256*256 pixels. In actual implementation, the preset size can be set according to the actual computational power of the network.

[0132] In conjunction with the above embodiments, in one implementation, after preprocessing the image to be processed, masks of each preset style region (e.g., face region, hairstyle region, and background region) in the image to be processed can be obtained, and then the preprocessed image to be processed, the masks of each preset style region in the image to be processed, and the target reference image are used as inputs to the image generation model.

[0133] In this context, the mask is a two-dimensional matrix array, with elements consisting of either 0 or 1. The mask is used to extract a specified region from an image to be processed. Taking a portrait image as an example, if the preset style regions include the face, hairstyle, and background, then multiplying the face mask by the image to be processed yields the face image; multiplying the background mask by the image to be processed yields the background image, and so on. The concept of a mask can be found in existing technologies, and will not be elaborated upon in this embodiment.

[0134] In practice, any algorithm can be used to obtain the mask corresponding to each preset style region.

[0135] In this embodiment, pre-obtaining masks for each preset style region in the image to be processed enables the image generation model to quickly determine the images of each preset style region in the image to be processed and perform feature extraction. For example, the encoder can quickly extract the image features corresponding to each preset style region based on the mask of each preset style region. As another example, the first branch network can quickly extract the image features of the first identity region based on the mask of the first identity region, and the second branch network can quickly extract the image features of the second identity region based on the mask of the second identity region.

[0136] In this embodiment, pre-obtaining masks for each preset style region of the image to be processed can reduce the computational load of the network and further improve the processing efficiency of the image generation model.

[0137] In conjunction with the above embodiments, in one implementation, this application further designs a loss function for the first identity region. Specifically, the loss function for the first identity region can be obtained in the following manner:

[0138] The loss function for the first identity region is obtained based on the error between the image of the first identity region of the target object in the target image and the image of the first identity region of the target object in the image to be processed.

[0139] In this method, the error between the image of the first identity region of the obtained target image and the image of the first identity region of the target object in the image to be processed is used as the loss function of the first identity region, which makes it easier and faster to obtain the loss function of the first identity region.

[0140] In one implementation, to further improve model performance, this application also provides another method for obtaining the loss function of the first identity region, which may specifically include:

[0141] Obtain a first image after style replacement based on a first target reference image, and a second image after style replacement based on a second target reference image;

[0142] The loss function for the first identity region is obtained based on the absolute value error between the image of the first identity region of the target object in the first image and the image of the first identity region of the target object in the second image.

[0143] The image generation model is optimized using the loss function of the first identity region.

[0144] In this embodiment, the loss function for the first identity region is shown as dom1 loss. Here, dom1 represents the first identity region.

[0145]

[0146] Where G(.) represents the decoder of the generator network, and x represents the input image to be processed (or the preprocessed image to be processed). Indicates from reference Figure 1 Features of each preset style region extracted from it. Indicates from reference Figure 2 The features of each preset style region extracted are shown, and doml mask1 and doml mask2 represent the corresponding generated images, respectively. and The mask of each of their first identity regions, ||·|| represents the calculation of absolute value error.

[0147] In practical implementation, for the same image to be processed, if style replacement is performed based on different target reference images, resulting in two different target images, and if the absolute error between the first identity region images of the target object in these two different target images is very small, then it can be said that the image generation model has a strong ability to preserve identity attributes during style replacement. Therefore, according to dom1 loss, if the loss function for obtaining the first identity region is obtained by using two style-replaced images based on different target reference images, the ability of the image generation model to preserve identity attributes during style replacement can be further improved.

[0148] In this embodiment, the image generation model can be optimized using the dom1 loss function of the first identity region, which can significantly improve the image processing quality of the image generation model and ensure that the attributes of the target object remain unchanged before and after style replacement.

[0149] In conjunction with the above embodiments, in one implementation, if the image features of both the first identity region and the second identity region of the target object are used simultaneously as the identity features of the target object, in addition to the loss function dom1 loss for the first identity region, this application also designs a loss function dom2 loss for the second identity region. Specifically, the method of this application may further include:

[0150] Based on the error in the Lab color space between the image of the second identity region of the target object in the target image and the image of the second identity region of the target object in the image to be processed, a loss function for the second identity region is obtained.

[0151] The image generation model is optimized using the loss functions of the first identity region and the second identity region.

[0152] In this embodiment, the loss function for the second identity region is shown as dom2 loss:

[0153] dom2 loss=labloss(x_fake*fake_dom2_mask,x_real*real_dom2_mask)

[0154] Where x_fake represents the target image, x_real represents the input image to be processed (or the preprocessed image to be processed), fake dom2 mask represents the mask of the second identity region of the target image, real_dom2_mask represents the mask of the second identity region of the image to be processed, and labloss(.) represents the error in the Lab color space.

[0155] In one implementation, the image generation model can be optimized using the loss function of the first identity region and the loss function of the second identity region, respectively. For example, the image generation model can be optimized first using the loss function of the first identity region, and then optimized again using the loss function of the second identity region. Alternatively, the image generation model can be optimized first using the loss function of the second identity region, and then optimized again using the loss function of the first identity region.

[0156] In another implementation, the expression for the sum of the loss functions of the first identity region and the second identity region can be used as the total loss function, and then the image generation model can be optimized using the total loss function.

[0157] In this embodiment, by combining the loss function of the first identity region and the loss function of the second identity region to optimize the image generation model, the image processing quality of the image generation model can be further improved, ensuring that the attributes of the target object remain unchanged before and after style replacement.

[0158] In conjunction with the above embodiments, in one implementation, the image generation model of this application can be constructed based on GAN (Generative Adversarial Network) and StarGANv2 technology. Specifically, the generation network can be implemented using the generation network in a GAN network, the first network can be implemented based on the style encoder in StarGANv2 technology, and the image discriminator can be implemented using the discriminator in a GAN network.

[0159] Because the style encoder lacks decoupling during feature extraction and cannot perform local feature extraction, it operates on the entire target reference image (extracting features from various pre-defined style regions). This results in even areas that don't require style replacement undergoing some degree of change, failing to preserve the object's identity attributes before and after the replacement. For example, when only the target person's hairstyle is replaced, because the style encoder operates on the entire target reference image, both the target person's facial area and the background area will be affected after the hairstyle replacement.

[0160] To overcome the aforementioned problems, this application improves upon GAN (Generative Adversarial Network) and StarGANv2 technologies by adding a second network. This second network extracts the identity features before replacement and replenishes any lost identity features after replacement (as described above), thereby preserving the identity attributes of the target object before and after replacement. Furthermore, to further optimize the image generation model, this application designs loss functions for the first and second identity regions. These loss functions are then used to optimize the image generation model, thereby further improving the image processing quality and achieving diverse style transformations while maintaining the identity attributes of the target object.

[0161] The image processing method of this application will be described in general with reference to a specific embodiment below. In this embodiment, the task is to replace the hairstyle in the user portrait image i with the hairstyle in the hairstyle reference image ref, where the first identity region is the face region and the second identity region is the background region.

[0162] Phase 1: Preprocessing

[0163] Step 1: Using a face alignment algorithm, perform face alignment processing on the user portrait image i and the hairstyle reference image ref respectively.

[0164] Step 2: Crop the aligned user portrait image i and hairstyle reference image ref to a size of 256*256. Only when the user portrait image i and hairstyle reference image ref are the same size can the hairstyle be replaced.

[0165] Step 3: Use a face parsing algorithm to process the user portrait image i obtained in Step 2 to obtain the facial region mask, hairstyle region mask, and background region mask.

[0166] After steps 1-3, the input information of the image generation model is obtained: preprocessed user portrait image i, preprocessed hairstyle reference image ref, facial region mask, hairstyle region mask, and background region mask.

[0167] Phase Two: Image Generation

[0168] Step 1: Input the preprocessed user portrait image i, hairstyle reference image ref, facial region mask, hairstyle region mask, and background region mask into the encoder for feature extraction to obtain the features of each region.

[0169] Step 2: Input the preprocessed hairstyle reference image (ref) into the first network for feature extraction to obtain the features of each region, and then apply them to the decoder through the adain operation.

[0170] Step 3: Input the preprocessed user portrait image i and the facial region mask into the first branch network to obtain the facial region features, and apply them to the decoder in a mask-guided manner.

[0171] Step 4: Input the preprocessed user portrait image i and the background region mask into the second branch network to obtain the features of the background region, and apply them to the decoder in a mask-guided manner.

[0172] Steps 1 through 4 above can be performed simultaneously.

[0173] Step 5: The decoder performs feature fusion on the features output by the encoder, the features output by the first network, the features output by the first branch network, and the features output by the second branch network to obtain the target image.

[0174] In terms of network structure, the first network of this application still performs overall feature extraction on the hairstyle reference image ref, and then applies it to the decoder through an adain operation. To better preserve the facial and background information of the input user portrait image i, information exchange between the encoder and decoder is introduced, which is implemented through skip connections. In specific implementation, based on the input user portrait image i, the facial region mask, and the background region mask, two additional parts can be obtained: the facial region and the background region (the clothing region also belongs to the background region). For the facial region, this application designs a separate facial feature extraction branch network, and for the background region, this application also designs a separate background feature extraction branch network. The facial feature extraction branch network can be composed of four residual network modules, aiming to extract high-level semantic features of the facial region, such as skin color, texture, and overall topological structure information. Then, it will be fused with the corresponding resolution layer of the decoder through a mask-guided method to ensure that the facial information lost under the action of the first network is supplemented. Considering the diversity and complexity of the background region, the background feature extraction branch network has two more residual modules than the facial feature extraction branch network, aiming to extract features better. Feature fusion is still performed using a mask-guided method.

[0175] Phase 3: Model Optimization

[0176] Method 1: Obtain hairstyle reference Figure 1 Target portrait after style replacement Figure 1 and based on hairstyle reference Figure 2 Target portrait after style replacement Figure 2 According to the target portrait Figure 1 Images of the facial region, compared with the target portrait Figure 2 The absolute value error between images of the facial region is used to obtain the loss function for the facial region, Face loss; Face loss is then used to optimize the image generation model.

[0177] Method 2: Based on the error in the Lab color space between the background region of the target portrait image and the background region of the user portrait image i before replacement, obtain the Lab loss function for the background region; optimize the image generation model using Face loss and Lab loss.

[0178] The image processing method of this application will now be described in general using another specific embodiment. In this embodiment, the task is to replace the background in vehicle image i with the background in background reference image ref. The first identity region is the outer surface region of the vehicle, and there is no need to use a second identity region.

[0179] Phase 1: Preprocessing

[0180] Step 1: Perform angle alignment processing on the vehicle image i and the background reference image ref respectively.

[0181] Step 2: Crop the aligned vehicle image i and background reference image ref to a size of 256*256. Only when the vehicle image i and the background reference image ref have the same size can the background replacement be performed.

[0182] Step 3: Process the vehicle image i obtained in Step 2 to obtain the outer surface area mask and the background area mask.

[0183] After steps 1-3, the input information of the image generation model is obtained: the preprocessed vehicle image i, the preprocessed background reference image ref, the outer surface region mask, and the background region mask.

[0184] Phase Two: Image Generation

[0185] Step 1: Input the preprocessed vehicle image i, background reference image ref, outer surface region mask, and background region mask into the encoder for feature extraction to obtain the features of each region.

[0186] Step 2: Input the preprocessed background reference image ref into the first network for feature extraction to obtain the features of each region, and then apply them to the decoder through the adain operation to achieve background replacement.

[0187] Step 3: Input the preprocessed vehicle image i and the outer surface region mask into the first branch network to obtain the features of the outer surface region, and apply them to the decoder in a mask-guided manner.

[0188] Steps 1 through 3 above can be performed simultaneously.

[0189] Step 4: The decoder performs feature fusion on the features output by the encoder, the features output by the first network, and the features output by the first branch network to obtain the target image. Specifically, the features output by the first branch network are used to supplement the features of the image generated by the first network.

[0190] Phase 3: Model Optimization

[0191] Get background-based reference Figure 1 Target vehicle after background replacement Figure 1 and based on background reference Figure 2 Target vehicle after background replacement Figure 2 According to the target vehicle Figure 1 Images of the outer surface area, and the target vehicle Figure 2The loss function for the outer surface region is obtained by calculating the absolute error between images of the outer surface region, and then the image generation model is optimized using the obtained loss function for the outer surface region.

[0192] This application proposes a novel network structure that divides the input into multiple parallel paths. By adding feature extraction branches for the first and second identity regions, it obtains higher-level abstract semantic features and employs a mask-guided feature fusion method, which better preserves the identity attributes of the target object before and after replacement. Furthermore, this application proposes a loss function based on the consistency of color features between the first and second identity regions to avoid changes in lighting and color in areas that do not need to be replaced due to the style features of the target reference image, thus ensuring the quality of image style replacement.

[0193] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of this application are not limited to the described order of actions, because according to the embodiments of this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of this application.

[0194] Based on the same inventive concept, one embodiment of this application provides an image processing apparatus 300. (Reference) Figure 3 , Figure 3 This is a structural block diagram of an image processing apparatus according to an embodiment of this application. Figure 3 As shown, the device 300 includes:

[0195] The replacement module 301 is used to perform style replacement on the image to be processed based on the target reference image to obtain the replaced image;

[0196] The correction module 302 is used to correct the replaced image by using the identity features of the target object in the image to be processed, with the goal of maximizing the similarity between the identity attributes of the target object in the replaced image and the identity attributes of the target object in the image to be processed, so as to obtain the target image.

[0197] Optionally, the correction module 302 includes:

[0198] The first determining submodule is used to use the image features of the first identity region of the target object as the identity features of the target object, wherein the first identity region includes features that can uniquely identify the target object.

[0199] The first acquisition submodule is used to perform feature fusion on the image features of the first identity region of the target object in the image to be processed and the image features of the first identity region of the target object in the replaced image through a mask-guided method to obtain the target image.

[0200] Optionally, the correction module 302 includes:

[0201] The second determining submodule is used to use the image features of the first identity region of the target object and the image features of the second identity region of the target object as the identity features of the target object, wherein the second identity region includes the background region of the target object;

[0202] The second acquisition submodule is used to perform feature fusion on the image features of the first identity region of the target object in the image to be processed, the image features of the second identity region of the target object in the image to be processed, the image features of the first identity region of the target object in the replaced image, and the image features of the second identity region of the target object in the replaced image through a mask-guided method to obtain a target image.

[0203] Optionally, the replacement module 301 includes:

[0204] The first extraction submodule is used to extract features from the target reference image through the first network in the image generation model to obtain image features of each preset style region in the target reference image;

[0205] The third submodule is used to perform feature fusion on the image features of each preset style region of the target reference image and the image features of each preset style region of the image to be processed, according to the preset replacement target, through the generative network in the image generation model, to obtain the replaced image. The preset replacement target represents the weight of the image features of each preset style region when performing feature fusion.

[0206] Optionally, the correction module 302 includes:

[0207] The second extraction submodule is used to extract features from the image to be processed through the second network in the image generation model to obtain the identity features of the target object in the image to be processed. The second network is different from the first network.

[0208] The fourth submodule is used to modify the replaced image by using the identity features of the target object in the image to be processed through the generation network, so as to obtain the target image.

[0209] Optionally, the generating network includes an encoder and a decoder. The encoder is used to acquire image features of various preset style regions of the image to be processed. The decoder takes the output of the first network, the output of the encoder, and the output of the second network as input. The second network includes a first branch network and a second branch network. The correction module 302 includes:

[0210] The first correction submodule is used to supplement the replaced image with features from the first identity region of the target object in the image to be processed, using the decoder and the image features output by the first branch network; or

[0211] The second correction submodule is used to supplement the replaced image with features by using the image features of the first identity region of the target object in the image to be processed output by the first branch network and the image features of the second identity region of the target object in the image to be processed output by the second branch network through the decoder.

[0212] Optionally, the device 300 further includes:

[0213] The first acquisition module is used to acquire a first image after style replacement based on a first target reference image, and a second image after style replacement based on a second target reference image;

[0214] The second obtaining module is used to obtain a loss function for the first identity region based on the absolute value error between the image of the first identity region of the target object in the first image and the image of the first identity region of the target object in the second image;

[0215] The first optimization module is used to utilize the first identity area. domain The loss function is used to optimize the image generation model.

[0216] Optionally, the device 300 further includes:

[0217] The third obtaining module is used to obtain a loss function for the second identity region based on the error between the image of the second identity region of the target object in the target image and the image of the second identity region of the target object in the image to be processed in the Lab color space.

[0218] The second optimization module is used to optimize the image generation model using the loss function of the first identity region and the loss function of the second identity region.

[0219] This application also provides a terminal device, such as... Figure 4 As shown. Figure 4 This is a schematic diagram illustrating the structure of a terminal device according to an embodiment of this application. (Refer to...) Figure 4The terminal device includes a processor 41, a communication interface 42, a memory 43, and a communication bus 44, wherein the processor 41, the communication interface 42, and the memory 43 communicate with each other through the communication bus 44.

[0220] Memory 43 is used to store computer programs;

[0221] When processor 41 executes the program stored in memory 43, it performs the following steps:

[0222] The style of the image to be processed is replaced based on the target reference image to obtain the replaced image;

[0223] With the goal of maximizing the similarity between the identity attributes of the target object in the replaced image and the identity attributes of the target object in the image to be processed, the replaced image is modified using the identity features of the target object in the image to be processed to obtain the target image.

[0224] Alternatively, when the processor 41 executes the program stored in the memory 43, it implements the steps described in the other method embodiments.

[0225] The communication bus mentioned above can be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into address bus, data bus, control bus, etc. For ease of illustration, only one thick line is used to represent it in the diagram, but this does not mean that there is only one bus or one type of bus.

[0226] The communication interface is used for communication between the aforementioned terminal and other devices.

[0227] The memory may include random access memory (RAM) or non-volatile memory, such as at least one disk storage device. Optionally, the memory may also be at least one storage device located remotely from the aforementioned processor.

[0228] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0229] In another embodiment provided in this application, a computer-readable storage medium is also provided, which stores instructions that, when executed on a computer, cause the computer to perform any of the webpage display methods described in the above embodiments.

[0230] In another embodiment provided in this application, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform any of the webpage display methods described in the above embodiments.

[0231] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0232] It should be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0233] The various embodiments in this specification are described in a related manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, the system embodiments are basically similar to the method embodiments, so the description is relatively simple; relevant parts can be referred to the descriptions of the method embodiments.

[0234] The above description is merely a preferred embodiment of this application and is not intended to limit the scope of protection of this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application are included within the scope of protection of this application.

Claims

1. An image processing method, characterized in that, include: The style of the image to be processed is replaced based on the target reference image to obtain the replaced image; With the goal of maximizing the similarity between the identity attributes of the target object in the replaced image and the identity attributes of the target object in the image to be processed, the replaced image is modified using the identity features of the target object in the image to be processed to obtain the target image; The step of correcting the replaced image using the identity features of the target object in the image to be processed to obtain the target image includes: The image features of the first identity region of the target object and the image features of the second identity region of the target object are used as the identity features of the target object, wherein the second identity region includes the background region of the target object; By using a mask-guided method, image features of the first identity region of the target object in the image to be processed, the second identity region of the target object in the image to be processed, the first identity region of the target object in the replaced image, and the second identity region of the target object in the replaced image are fused to obtain the target image.

2. The method according to claim 1, characterized in that, Style replacement is performed on the image to be processed based on the target reference image, resulting in the replaced image, including: The first network in the image generation model is used to extract features from the target reference image to obtain the image features of each preset style region in the target reference image; The image generation model uses a generative network to perform feature fusion on the image features of each preset style region of the target reference image and the image features of each preset style region of the image to be processed, according to a preset replacement target, to obtain a replaced image. The preset replacement target represents the weight of the image features of each preset style region when performing feature fusion.

3. The method according to claim 2, characterized in that, The replacement image is corrected using the identity features of the target object in the image to be processed to obtain the target image, including: The image to be processed is used to extract features from the image to be processed by the second network in the image generation model to obtain the identity features of the target object in the image to be processed. The second network is different from the first network. The replacement image is corrected using the identity features of the target object in the image to be processed through the generation network to obtain the target image.

4. The method according to claim 3, characterized in that, The generating network includes an encoder and a decoder. The encoder is used to acquire image features of each preset style region of the image to be processed. The decoder takes the output of the first network, the output of the encoder, and the output of the second network as input. The second network includes: a first branch network and a second branch network; Correcting the replaced image using the identity features of the target object in the image to be processed includes: The decoder uses image features of the first identity region of the target object in the image to be processed, output by the first branch network, to supplement the features of the replaced image; or The decoder uses the image features of the first identity region of the target object in the image to be processed output by the first branch network and the image features of the second identity region of the target object in the image to be processed output by the second branch network to supplement the features of the replaced image.

5. The method according to claim 2, characterized in that, The method further includes: Obtain a first image after style replacement based on a first target reference image, and a second image after style replacement based on a second target reference image; The loss function for the first identity region is obtained based on the absolute value error between the image of the first identity region of the target object in the first image and the image of the first identity region of the target object in the second image. The image generation model is optimized using the loss function of the first identity region.

6. The method according to claim 5, characterized in that, The method further includes: Based on the error in the Lab color space between the image of the second identity region of the target object in the target image and the image of the second identity region of the target object in the image to be processed, a loss function for the second identity region is obtained. The image generation model is optimized using the loss functions of the first identity region and the second identity region.

7. An image processing apparatus, characterized in that, include: The replacement module is used to perform style replacement on the image to be processed based on the target reference image, so as to obtain the replaced image. The correction module is used to correct the replaced image by using the identity features of the target object in the image to be processed, with the goal of maximizing the similarity between the identity attributes of the target object in the replaced image and the identity attributes of the target object in the image to be processed, so as to obtain the target image. The replacement module is specifically used to use the image features of the first identity region of the target object and the image features of the second identity region of the target object as the identity features of the target object, wherein the second identity region includes the background region of the target object; By using a mask-guided method, image features of the first identity region of the target object in the image to be processed, the second identity region of the target object in the image to be processed, the first identity region of the target object in the replaced image, and the second identity region of the target object in the replaced image are fused to obtain the target image.

8. A terminal device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus; Memory, used to store computer programs; A processor, when executing a program stored in memory, implements the steps of the image processing method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the image processing method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Image generation method and device, network training method and device, computer equipment and storage medium

    CN113361490A