Image processing method and device and electronic equipment

By using the target mask image and image content processing model in the image content replacement technology, the replacement content of the adversarial network and the target extension network is generated, and the problem of poor replacement effect in the prior art is solved, and better image content replacement effect and image quality are achieved.

CN120147477APending Publication Date: 2025-06-13GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311711441.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-12
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The existing image content replacement technology has the problem of poor replacement effect.

Method used

Generation of the target image is achieved by determining the target object from the to-process image, obtaining the target mask image, and generating replacement content generated by the adversarial network and the target extension network based on the to-process image, the target mask image, and the image content processing model.

Benefits of technology

The effect of image content replacement is improved, so that the generated target image has better picture quality and more controllable content.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120147477A_ABST
    Figure CN120147477A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses an image processing method and device and electronic equipment. The method comprises the following steps: determining a target object from a to-be-processed image; obtaining a target mask image corresponding to the to-be-processed image, wherein the target mask image comprises a target mask area corresponding to the target object; based on the to-be-processed image, the target mask image and an image content processing model, first replacement content corresponding to the target object and output by the image content processing model is obtained, and the image content processing model generates the first replacement content through a generative adversarial network and a target expansion network; the target expansion network is obtained by training through training data of image categories; and changing the image content corresponding to the target object in the to-be-processed image into the first replacement content to obtain a target image. Therefore, the obtained target image has better image quality, and the replacement effect of the image content is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and more particularly, to an image processing method, apparatus, and electronic device. Background Art

[0002] In the process of image content replacement, specific elements in an image can be replaced with other elements (for example, replacing a person in an image with other content) to achieve the modification and beautification of the image content. Image content replacement technology can have wide applications in multiple fields, such as advertising, film and television, games, and social media. However, the related image content replacement methods still have the problem of poor replacement effects. Summary of the Invention

[0003] In view of the above problems, this application proposes an image processing method, apparatus, and electronic device to improve the above problems.

[0004] In a first aspect, this application provides an image processing method, the method comprising: determining a target object from a to-be-processed image; obtaining a target mask image corresponding to the to-be-processed image, the target mask image including a target mask region corresponding to the target object; based on the to-be-processed image, the target mask image, and an image content processing model, obtaining a first replacement content corresponding to the target object output by the image content processing model, wherein the image content processing model generates the first replacement content through a generative adversarial network and a target expansion network, and the target expansion network is trained through training data of image categories; changing the image content corresponding to the target object in the to-be-processed image to the first replacement content to obtain a target image.

[0005] In a second aspect, this application provides an image processing apparatus, the apparatus comprising: a target object determination unit for determining a target object from a to-be-processed image; a mask obtaining unit for obtaining a target mask image corresponding to the to-be-processed image, the target mask image including a target mask region corresponding to the target object; a replacement content obtaining unit for obtaining a first replacement content corresponding to the target object output by the image content processing model based on the to-be-processed image, the target mask image, and the image content processing model, wherein the image content processing model generates the first replacement content through a generative adversarial network and a target expansion network, and the target expansion network is trained through training data of image categories; and a picture fusion unit for changing the image content corresponding to the target object in the to-be-processed image to the first replacement content to obtain a target image.

[0006] In a third aspect, the present application provides an electronic device, which at least includes a processor and a memory; one or more programs are stored in the memory and configured to be executed by the processor to implement the above method.

[0007] In a fourth aspect, the present application provides a computer-readable storage medium, in which program code is stored, and when the program code is run by a processor, the above method is executed.

[0008] A method, apparatus, and electronic device for image processing provided by the present application can, after determining a target object from an image to be processed, obtain a target mask image corresponding to the image to be processed, and then, based on the image to be processed, the target mask image, and an image content processing model, obtain first replacement content output by the image content processing model corresponding to the target object, and change the image content corresponding to the target object in the image to be processed to the first replacement content to obtain a target image. Thus, when the image content processing model generates the first replacement content through a generative adversarial network and a target expansion network, and the target expansion network is trained with training data of an image category, the image content (for example, the first replacement content) generated by the target expansion network is more controllable and has better image quality. Therefore, after replacing the image content corresponding to the target object in the image to be processed based on the first replacement content, the obtained target image has better image quality, thereby improving the replacement effect of the image content. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application, and those skilled in the art can obtain other drawings without creative efforts based on these drawings.

[0010] Figure 1 A schematic diagram showing an application scenario of the image processing method proposed in an embodiment of the present application;

[0011] Figure 2 A schematic diagram showing another application scenario of the image processing method proposed in an embodiment of the present application;

[0012] Figure 3 A flowchart showing an image processing method proposed in an embodiment of the present application;

[0013] Figure 4 A schematic diagram showing candidate objects in an image to be processed in an embodiment of the present application;

[0014] Figure 5A schematic diagram showing identification of candidate objects in an embodiment of the present application is shown;

[0015] Figure 6 A schematic diagram of a target mask image in an embodiment of the present application is shown;

[0016] Figure 7 A schematic diagram showing replacement of an image of a target object in an embodiment of the present application is shown;

[0017] Figure 8 A flowchart of an image processing method proposed in another embodiment of the present application is shown;

[0018] Figure 9 A schematic diagram of intercepting a first partial image and a second partial image in the present application is shown;

[0019] Figure 10 A schematic diagram showing replacement of a replacement result image in the present application is shown;

[0020] Figure 11 A flowchart of an image processing method proposed in another embodiment of the present application is shown;

[0021] Figure 12 A schematic diagram showing the removal of intersecting areas in the present application is shown;

[0022] Figure 13 A flowchart of an image processing method proposed in an embodiment of the present application is shown;

[0023] Figure 14 A flowchart of an image processing method proposed in another embodiment of the present application is shown;

[0024] Figure 15 A schematic diagram showing a processing method using an image content processing model in the present application is shown;

[0025] Figure 16 A schematic diagram of obtaining a target training mask image in the present application is shown;

[0026] Figure 17 A schematic diagram of a reference mask image region in the present application is shown;

[0027] Figure 18 A schematic diagram of a reference mask image region in the present application is shown;

[0028] Figure 19 A structural block diagram of an image processing device proposed in an embodiment of the present application is shown;

[0029] Figure 20 A structural block diagram of an image processing device proposed in an embodiment of the present application is shown;

[0030] Figure 21 The block diagram of another electronic device for executing the image processing method according to the embodiments of the present application is shown;

[0031] Figure 22 It is a storage unit for storing or carrying the program code for implementing the image processing method according to the embodiments of the present application. Detailed implementation manners

[0032] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0033] Image processing is a technology that involves various operations and analyses of images. It is widely applied in various fields, including computer vision, medical image analysis, digital image processing, remote sensing image processing, etc. By processing images, we can achieve various tasks, such as image restoration, image enhancement, feature extraction, object detection and recognition, etc.

[0034] In some cases, the effect of image processing can be achieved by replacing some content in the image. For example, during the process of replacing the content of a picture, specific elements in the picture can be replaced with other elements (for example, replacing the people in the picture with other content) to achieve the modification and beautification of the picture content.

[0035] However, the inventors found in the research that there are still problems with the poor replacement effect in the related image content replacement methods. Therefore, after discovering the above problems in the research, the inventors proposed the image processing method, device, and electronic device in the present application that can improve the above problems. In this method, after determining the target object from the image to be processed, the target mask image corresponding to the image to be processed can be obtained. Then, based on the image to be processed, the target mask image, and the image content processing model, the first replacement content corresponding to the target object output by the image content processing model is obtained, and the image content corresponding to the target object in the image to be processed is changed to the first replacement content to obtain the target image.

[0036] Thus, when the image content processing model generates the first replacement content through a generative adversarial network and a target expansion network, and the target expansion network is trained with training data of image categories, the image content (e.g., the first replacement content) generated by the target expansion network is more controllable and has better image quality. As a result, after replacing the image content corresponding to the target object in the image to be processed with the first replacement content, the obtained target image has better image quality, thereby improving the replacement effect of the image content.

[0037] Before further elaborating on the embodiments of the present application, an application environment involved in the embodiments of the present application is introduced.

[0038] First, the application scenarios involved in the embodiments of the present application are introduced below.

[0039] In the embodiments of the present application, the provided image processing method can be executed by an electronic device. In this manner of being executed by an electronic device, all steps in the image processing method provided in the embodiments of the present application can be executed by the electronic device. For example, as Figure 1 shown, all steps in the image processing method provided in the embodiments of the present application can be executed by the processor of the electronic device 100.

[0040] Alternatively, the image processing method provided in the embodiments of the present application can also be executed by a server. Correspondingly, in this manner of being executed by the server, the server can start executing the steps in the image processing method provided in the embodiments of the present application in response to a trigger instruction. Among them, the trigger instruction can be sent by the electronic device used by the user, or can be locally triggered by the server in response to some automated events.

[0041] In addition, the image processing method provided in the embodiments of the present application can also be executed collaboratively by an electronic device and a server. In this manner of being executed collaboratively by an electronic device and a server, some steps in the image processing method provided in the embodiments of the present application are executed by the electronic device, while other steps are executed by the server. Exemplarily, as Figure 2 shown, the electronic device 100 can execute the following steps included in the image processing method: determining a target object from the image to be processed and obtaining a target mask image corresponding to the image to be processed. Then, the electronic device 100 transmits the image to be processed and the target mask image to the server 200. After receiving the image to be processed and the target mask image, the server 200 can execute subsequent steps to obtain a target image, and then transmit the target image to the electronic device 100. After receiving the target image, the electronic device 100 can store, display, and share the target image.

[0042] It should be noted that in this way of collaborative execution by the electronic device and the server, the steps respectively executed by the electronic device and the server are not limited to the ways introduced in the above examples. In actual applications, the steps respectively executed by the electronic device and the server can be dynamically adjusted according to the actual situation.

[0043] It should be noted that in addition to being Figure 1 and Figure 2 the smart phone shown in, the electronic device 100 can also be devices such as a tablet computer, a smart watch, a smart voice assistant, etc. The server 200 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers. It can also be a cloud server that provides basic cloud computing services such as cloud services, cloud computing, cloud storage, network services, cloud communication, middleware services, CDN (Content Delivery Network), and artificial intelligence platforms. Among them, when the image processing method provided in the embodiment of the present application is executed by a server cluster or a distributed system composed of multiple physical servers, different steps in the image processing method can be respectively executed by different physical servers, or can be executed by a server based on a distributed system in a distributed manner.

[0044] Next, each embodiment of the present application will be specifically described with reference to the accompanying drawings.

[0045] Please refer to Figure 3 , an image processing method provided by an embodiment of the present application, the method includes:

[0046] S110: Determine a target object from the image to be processed.

[0047] In the embodiment of the present application, the image to be processed can be understood as an image to be subjected to content replacement. Among them, there are various ways to obtain the image to be processed.

[0048] As a way, the preview image currently displayed on the electronic device can be used as the image to be processed. In this way, the electronic device can collect images in real time through the camera and display the collected images in the preview area. Furthermore, the user can operate the electronic device to use the image (a single frame of picture) displayed in the preview area as the image to be processed. As another way, the electronic device can obtain the image to be processed from the album. In this way, when the album is launched, the electronic device can display the pictures in the album and, in response to the user's selection, select a picture from the album as the image to be processed. As yet another way, the electronic device can use the picture transmitted from other devices as the image to be processed. For example, when the user of the electronic device chats with other friends through an instant messaging program, other friends can send a picture through the instant messaging program. In this case, the electronic device can use the sent picture as the image to be processed.

[0049] Among them, the target object in the image to be processed can be understood as the object that needs to be replaced. In the embodiments of the present application, there are various ways to determine the target object.

[0050] As a way, the user can determine the target object in the image to be processed. For example, when the image to be processed includes multiple objects (such as multiple foreground objects), these multiple objects can all be used as candidate objects. Exemplarily, as Figure 4 shown, by recognizing Figure 4 , the person 10 and the person 20 in Figure 4 can be used as candidate objects. Furthermore, the user can select one person from the person 10 and the person 20 as the target object. Among them, in the embodiments of the present application, the way for the user to select the target object from the candidate objects is not specifically limited. For example, as Figure 5 shown, the candidate objects can be marked with a dotted box in the image to be processed. If it is detected that a dotted box is clicked, the candidate object marked by the clicked dotted box can be used as the target object.

[0051] As a way, the target object in the image to be processed can be determined through the current processing scenario. Optionally, the electronic device can enter the specified processing scenario in response to the user's operation. Then, in the specified processing scenario, the object corresponding to the image to be processed can be used as the target object. For example, the specified processing scenario can be the scenario of removing passers-by. In this scenario of removing passers-by, the electronic device can use the foreground objects (such as other people) in the image to be processed obtained, except for the specified person, as the target objects. Among them, the electronic device can identify the specified person from the image to be processed through face recognition. The specified person can be the user of the electronic device or other persons specified by the electronic device. For another example, the specified processing scenario can be the scenario of removing people. In this case, all the people in the image to be processed can be used as the target objects.

[0052] S120: Obtain the target mask image corresponding to the image to be processed, where the target mask image includes the target mask region corresponding to the target object.

[0053] In the embodiments of the present application, the mask image can be understood as a binary image with the same size as the original image (such as the image to be processed). In this binary image, only pixels of two colors, black and white, can be included. Among them, the mask image can include a mask region, and this mask region can be used to identify the region in the original image that needs to be processed. For example, the pixels in the mask region in the mask image can all be black, and the colors of the pixels in the non-mask region can all be white. Among them, the non-mask region can be understood as all regions outside the mask region.

[0054] In this case, in the target mask image corresponding to the image to be processed, there can be a target mask region corresponding to the target object. This target mask region is used to identify the position of the target object in the image to be processed, and thus also identifies the region in the image to be processed that needs to be processed. Among them, the position of the target mask region in the target mask image is the same as the position of the target object in the image to be processed, and the contour shape of the target mask region can also be the same as the contour shape of the target object.

[0055] As Figure 6 shown, when it is determined that the object 20 in the image to be processed is the target object, the target mask image corresponding to the image to be processed can be as Figure 6 shown in the right image. In the Figure 6 shown target mask image, there are a region with the color of black and a region with the color of white. Among them, the region with the color of black can be understood as the target mask region, and the region with the color of white can be understood as the non-mask region.

[0056] S130: Based on the image to be processed, the target mask image, and the image content processing model, obtain the first replacement content corresponding to the target object output by the image content processing model. Among them, the image content processing model generates the first replacement content through a generative adversarial network and a target extension network, and the target extension network is trained with training data of image categories.

[0057] In the embodiment of the present application, the image content processing model can determine the area where the target object is located according to the identifier of the target mask area in the target mask image, and then generate the first replacement content corresponding to the target object in the image to be processed.

[0058] Among them, the first replacement content generated by the image content processing model can be determined by the training data of the image content processing model. Optionally, during the training process of the image content processing model, the images in the training data used carry data labels, and most of the data labels are labels about background content. Then, when the trained image content processing model generates content, it will more likely generate background content and will no longer generate foreground content (for example, people in the image).

[0059] Among them, the first replacement content can be generated by the target extension network (Unet) in the image content processing model. In this process, the target extension network can be controlled by the generative adversarial network (GAN inpaint model) to generate the first replacement content under the control of the generative adversarial network. The target extension network can be trained with training data of image categories (for example, a large number of pictures as training data). The images generated by the target extension network are relatively good in terms of semantics and image quality, which can make up for the defects of the generative adversarial network. At the same time, the target extension network is controlled by the output result of the generative adversarial network, and thus the generated image content is controllable and semantically reasonable. Among them, semantically reasonable can be understood as the generated content is image content that can be understood by users.

[0060] S140: Change the image content corresponding to the target object in the image to be processed to the first replacement content to obtain the target image.

[0061] After obtaining the first replacement content, the first replacement content can be fused with the image to be processed. This fusion process can be understood as replacing the image content corresponding to the target object in the image to be processed with the first replacement content to obtain the target image.

[0062] Exemplarily, as Figure 7 shown, based on Figure 4 and Figure 6 shown, in the case where object 20 is determined as the target object and the generated first replacement content is a tree, thenFigure 7 The target image shown. As Figure 7 shown in the target image, the image content corresponding to the object 20 is replaced with the trees generated by the image content processing model.

[0063] It should be noted that, in the embodiments of the present application, the image content processing model can randomly generate the first replacement content, or can generate the first replacement content according to the background content in the image to be processed, so that after the generated first replacement content is fused into the image to be processed, the first replacement content is more coordinated with the original content in the image to be processed. For example, when it is recognized that the background content in the image to be processed is a plant, the generated first replacement content can also be content related to plants. When it is recognized that the background content in the image to be processed is a street, the generated first replacement content can be content that appears in the street.

[0064] An image processing method provided in this embodiment, so that when the image content processing model generates the first replacement content through a generative adversarial network and a target expansion network, and the target expansion network is trained by the training data of the image category, the image content (for example, the first replacement content) generated by the target expansion network is more controllable and has better image quality, so that after replacing the image content corresponding to the target object in the image to be processed based on the first replacement content, the obtained target image has better image quality, thereby improving the replacement effect of the image content.

[0065] Please refer to Figure 8 , an image processing method provided in the embodiments of the present application, the method includes:

[0066] S210: Determine the target object from the image to be processed.

[0067] S220: Obtain the target mask image corresponding to the image to be processed, where the target mask image includes the target mask area corresponding to the target object.

[0068] S230: Obtain the first partial image from the image to be processed, and obtain the second partial image corresponding to the first partial image from the target mask image, where the first partial image includes the target object, and the second partial image includes the target mask area.

[0069] It should be noted that in some cases, the image content processing model may compress the input image, which may result in the final generated target image not being clear enough. Moreover, when the image to be processed is relatively large, directly inputting the image to be processed into the image content processing model will also cause a high data calculation volume. In this case, to further ensure the clarity of the subsequent obtained target image and to a certain extent reduce the data calculation volume during the image processing process, a part of the image content can be intercepted from the image to be processed for processing. Correspondingly, a part of the image content will also be intercepted from the target mask image. Among them, a part of the image content intercepted from the target mask image corresponds to a part of the image content intercepted from the image to be processed. It can be understood that intercepting a part of the image content from the image to be processed will include the target object, and intercepting a part of the image content from the target mask image will include the target mask area corresponding to the target object.

[0070] In the embodiment of the present application, a part of the image content obtained from the image to be processed can be used as the first part of the image, and a part of the image content obtained from the target mask image can be used as the second part of the image.

[0071] As a way, in order to make the shapes of the obtained first part of the image and the second part of the image more regular, during the process of obtaining the first part of the image from the image to be processed, the image corresponding to the first circumscribed rectangle of the target object in the image to be processed can be obtained as the first part of the image. Correspondingly, during the process of obtaining the second part of the image corresponding to the first part of the image from the target mask image, the image corresponding to the second circumscribed rectangle of the target mask area in the target mask image can be obtained as the second part of the image, and the second circumscribed rectangle corresponds to the first circumscribed rectangle. The correspondence between the first circumscribed rectangle and the second circumscribed rectangle can be understood as that the shapes and sizes of the first circumscribed rectangle and the second circumscribed rectangle are the same, and the area covered by the first circumscribed rectangle in the image to be processed is the same as the area covered by the second circumscribed rectangle in the target mask image.

[0072] It should be noted that the circumscribed rectangle of an object or region in an image can be understood as a rectangle that completely encloses the object or region. Correspondingly, the first circumscribed rectangle of the target object can be a rectangle that completely encloses the target object, and the second circumscribed rectangle of the target mask region can be understood as a rectangle that completely encloses the target mask region. Among them, for an object or region, there may be multiple corresponding circumscribed rectangles. In this case, the smallest circumscribed rectangle among the multiple circumscribed rectangles corresponding to the target object can be used as the first circumscribed rectangle corresponding to the target object. The smallest circumscribed rectangle among the multiple circumscribed rectangles corresponding to the target mask region can be used as the second circumscribed rectangle corresponding to the target mask region.

[0073] Exemplarily, as Figure 9 shown, after the first circumscribed rectangle 21 and the second circumscribed rectangle 22 are used to intercept the image to be processed and the corresponding target mask image respectively, the first partial image and the second partial image obtained. As can be seen from Figure 9 it, the first partial image includes the target object, and the second partial image includes the target mask region corresponding to the target object. Moreover, the position of the first partial image in the image to be processed is the same as the position of the second partial image in the target mask image.

[0074] S240: Input the first partial image and the second partial image into the image content processing model, and obtain the first replacement content corresponding to the target object output by the image content processing model. Among them, the image content processing model generates the first replacement content through a generative adversarial network and a target expansion network, and the target expansion network is trained through training data of image categories.

[0075] S250: Change the image content corresponding to the target object in the image to be processed to the first replacement content to obtain the target image.

[0076] It should be noted that the image content output by the image processing model (for example, the first replacement content) may have a color difference from the original image (for example, the image to be processed). Then, in order to correct this color difference, the image content output by the image content processing model can be color-corrected.

[0077] Optionally, after the first partial image and the second partial image are input into the image content processing model, the image content processing model can output a replacement result image, where the replacement result image is an image obtained by changing the image content corresponding to the target object in the first partial image to the first replacement content. Exemplarily, based on Figure 9 the situation shown, the obtained replacement result image can be the image shown in Figure 10 As shown. As Figure 10As shown, in the replacement result image, the target object in the original image (e.g., person 20) has been replaced with the first replacement content (tree).

[0078] During the color correction process, the color of the first replacement content can be corrected based on the color information of the image content other than the first replacement content in the replacement result image, and the color information of the image content other than the target object in the first part of the image. Among them, the color information can include the mean value of the color values and the standard deviation of the color values. Optionally, the color correction can be performed through the following formula, and the formula is:

[0079] image_new = (sc - s_mean) * (t_std / s_std) + t_mean

[0080] Among them, image_new represents the replacement result image after color correction, sc represents the replacement result image that has not been color-corrected, s_mean represents the mean value of the non-masked area of the replacement result image that has not been color-corrected (that is, the area other than the first replacement content in the replacement result image that has not been color-corrected), s_std represents the standard deviation of the non-masked area of the replacement result image that has not been color-corrected, t_mean represents the mean value of the non-masked area of the first part of the image (that is, the area other than the target object in the first part of the image), and t_std represents the standard deviation of the non-masked area of the first part of the image. Among them, the mean value can be understood as the mean value of the color values, and the standard deviation is the standard deviation of the color values.

[0081] After obtaining the first replacement content with the corrected color, the image content corresponding to the target object in the image to be processed can be changed to the first replacement content with the corrected color to obtain the target image.

[0082] An image processing method provided in this embodiment improves the replacement effect of the image content. Moreover, in this embodiment, after obtaining the image to be processed and the target mask image, the first part of the image can be obtained from the image to be processed, and the second part of the image corresponding to the first part of the image can be obtained from the target mask image. Furthermore, by inputting the first part of the image and the second part of the image into the image content processing model for processing, it is possible to reduce the amount of data required for processing while obtaining the first replacement content. Moreover, in this embodiment, after obtaining the first replacement content, the color of the first replacement content can be corrected to obtain the first replacement content with the corrected color, so that subsequently, the first replacement content with the corrected color can be fused into the image to be processed to further improve the content replacement effect.

[0083] Please refer to Figure 11 , an image processing method provided in an embodiment of the present application, the method includes:

[0084] S310: Identify the image to be processed to obtain candidate objects in the image to be processed and an initial mask image corresponding to the image to be processed. The initial mask image includes an initial mask region corresponding to the candidate object.

[0085] As a method, the image to be processed can be input into an image segmentation model to identify the image to be processed through the image segmentation model, so as to obtain candidate objects in the image to be processed and an initial mask image corresponding to the image to be processed.

[0086] Optionally, the image segmentation model can identify all foreground objects in the image to be processed as candidate objects. Among them, the foreground and background can be distinguished by the depth information of each region in the image to be processed. Among them, the region with relatively smaller depth information can be identified as the foreground, and the region with relatively larger depth information can be identified as the background. Among them, the objects in the region identified as the foreground region will be identified as foreground objects.

[0087] S320: Determine the target object from the image to be processed.

[0088] S330: Remove the initial mask regions of the remaining candidate objects in the initial mask image to obtain a target mask image corresponding to the image to be processed. The remaining candidate objects are candidate objects other than the target object.

[0089] Among them, removing the initial mask regions of the remaining candidate objects in the initial mask image can be understood as changing the color values of the initial mask regions of the remaining candidate objects to be the same as the color values of the original non-mask regions. Exemplarily, the color values of the initial mask regions of multiple candidate objects in the initial mask image can all be black, and the color of the non-mask region can all be white. After determining the target object, the color values of the initial mask regions of the remaining candidate objects can be changed to white, so that only the target mask region corresponding to the target object exists in the obtained target mask image. In this case, the target mask region corresponding to the target object is the same as the initial mask region corresponding to the target object.

[0090] It should be noted that in the case where the object in the image includes relatively small parts, the mask area of the object may not be able to cover these relatively small parts well. For example, when the target object is a person, the initial mask area of the target object may not cover parts such as the person's hair. As an improvement method, before removing the mask areas of candidate objects other than the target object in the initial mask image to obtain the target mask image corresponding to the image to be processed, the initial mask area of the target object can also be enlarged to obtain an enlarged mask area. And if the enlarged mask area intersects with the initial mask areas of other candidate objects, the intersecting area is obtained, and the intersecting area is removed from the enlarged mask area to obtain the target mask area of the target object.

[0091] Exemplarily, as Figure 12 shown, where the mask area 31 can be understood as the enlarged mask area, and the mask area 30 can be understood as the initial mask areas of other candidate objects. As Figure 12 shown, there is an intersecting area between the mask area 31 and the mask area 30 ( Figure 12 the shaded part in), in this case, the intersecting area can be removed from the mask area 31, so as to obtain Figure 12 the target mask area shown in the right image. It should be noted that Figure 12 the shapes of the mask areas in are only for explaining the solution and are not used to limit the actual shapes of the mask areas.

[0092] As an enlargement method, after obtaining the initial mask area of the target object, the area of the initial mask area can be obtained, then the square root of the area is taken, then the result of taking the square root is rounded, and then the rounded result is multiplied by a preset coefficient to obtain the number of pixels included in the enlarged initial mask area.

[0093] S340: Based on the image to be processed, the target mask image, and the image content processing model, obtain the first replacement content corresponding to the target object output by the image content processing model. Among them, the image content processing model generates the first replacement content through a generative adversarial network and a target expansion network, and the target expansion network is trained through training data of image categories.

[0094] S350: Change the image content corresponding to the target object in the image to be processed to the first replacement content to obtain the target image.

[0095] As a way, after obtaining the first part of the image and the second part of the image, the first part of the image and the second part of the image can be downsampled respectively, and then the downsampled first part of the image and the second part of the image are input into the image content processing model for processing. In this case, after obtaining the first replacement content output by the image content processing model, the first replacement content can be upsampled first, and then the image content corresponding to the target object in the image to be processed is changed to the upsampled first replacement content to obtain the target image. Among them, in the process of downsampling, downsampling can be performed according to the resolution adapted by the image content processing model. For example, if the resolution adapted by the image content processing model is 512×512, then both the first part of the image and the second part of the image can be downsampled to 512×512. In the process of upsampling the first replacement content, the first replacement content can be upsampled to the resolution ratio of the image to be processed.

[0096] Please refer to Figure 13 , and then the process of an image processing method involved in this embodiment will be described below through Figure 13 an example. Figure 13 The scenario shown can be understood as a processing scenario for replacing passers-by in the image to be processed. In this case, the passers-by can be understood as the target object. Among them, after the user confirms the passers-by (target object), the initial mask area of the passers-by can be dilated. Dilating the initial mask area of the passers-by can be understood as increasing the initial mask area of the passers-by to obtain an enlarged initial mask area. The enlarged initial mask area may cover the mask areas of other objects, and then the operation of removing the main body occlusion can be performed to obtain the target mask image of the passers-by. Among them, in the process of performing the operation of removing the main body occlusion, the intersecting area in the enlarged initial mask area can be removed, and the intersecting area is the area where the enlarged initial mask area intersects with the mask areas of other objects.

[0097] After obtaining the target mask image of the passers-by, the original image and the target mask image of the passers-by can be cropped respectively to obtain the first part of the image and the second part of the image. In Figure 13 the case shown, the first part of the image will include passers-by, and the second part of the image will include the target mask area corresponding to the passers-by. Then, the first part of the image and the second part of the image will be downsampled respectively, and then the downsampled first part of the image and the downsampled second part of the image are input into the passer-by removal model to obtain the replacement result. In Figure 13In the scene shown, the passerby removal model can be understood as an image content processing model, and the replacement result can be understood as the replacement result image. Then, in the step of color transfer, color correction as shown in the foregoing content is performed. After undergoing color transfer, the first replacement content in the replacement result image is upsampled, and then the upsampled first replacement content is fused with the original image to obtain the target image.

[0098] An image processing method provided in this embodiment improves the replacement effect of image content. Moreover, in this embodiment, foreground objects in the image to be processed can be recognized to obtain multiple candidate objects, so that the user can select the target object to be processed from the candidate objects according to their own preferences or inclinations, thereby improving the flexibility of image processing. Furthermore, in this embodiment, after obtaining the initial mask region of the target object, the initial mask region can be enlarged so that the enlarged mask region can cover all parts included in the target object. And, in order to avoid the enlarged mask region from obscuring other foreground objects, the region of the enlarged mask region that intersects with the initial mask regions of other objects will also be removed, thereby avoiding replacing part of the image of other foreground objects during the subsequent content replacement process.

[0099] Please refer to Figure 14 , an image processing method provided in an embodiment of the present application, the method includes:

[0100] S410: Determine a target object from the image to be processed.

[0101] S420: Obtain a target mask image corresponding to the image to be processed, where the target mask image includes a target mask region corresponding to the target object.

[0102] S430: Based on the image to be processed, the target mask image, and the adversarial generation network, obtain a second replacement content.

[0103] As a method, based on the target mask image, the color values of the image corresponding to the target object in the image to be processed are processed into target values to obtain the image to be processed with processed colors, and the target values are used to mark the positions where content replacement needs to be performed. Based on the image to be processed with processed colors and the adversarial generation network, a second replacement content is obtained. Among them, the target mask region in the target mask image can be used to identify the position of the target object in the image to be processed. And, as shown in the foregoing content, the target mask image is a binary image. Furthermore, by multiplying the target mask image with the image to be processed, the color values of the image corresponding to the target object in the image to be processed can be processed into target values to obtain the image to be processed with processed colors. For example, when the target mask region in the binary image is black, the target value is black.

[0104] Optionally, when the color value of the target mask region in the target mask image is black, after multiplying the target mask image with the image to be processed, in the processed image with the processed color, the color of the region where the target object is located will be configured to black. In this case, after inputting the processed image with the processed color into the adversarial generation network, the adversarial generation network can identify the black region as the region that needs to be replaced, and then generate the corresponding second replacement content for the black region.

[0105] S440: Obtain the first replacement content based on the target expansion network and the second replacement content.

[0106] As a way, under the control of the target vector, the target expansion network obtains the first replacement content based on the processed image to be processed and the target mask image, and the target vector is obtained by extracting features from the second replacement content.

[0107] S450: Change the image content corresponding to the target object in the image to be processed to the first replacement content to obtain the target image.

[0108] Next, through Figure 15 to illustrate the content involved in this embodiment.

[0109] As Figure 15 shown, after obtaining the processed image to be processed (masked image), the processed image to be processed can be input into the adversarial generation network (GAN, Generative Adversarial Network) to obtain the replacement result image output by the adversarial generation network. The replacement result image output by the adversarial generation network may include the second replacement content. For the replacement result image output by the adversarial generation network, its corresponding image feature (for example, an image feature in vector form) can be obtained, and then the image feature is input into the target expansion network (Unet) to control the target expansion network to output the replacement result image. The replacement result image output by the target expansion network may include the first replacement content.

[0110] Among them, the image to be processed after color processing will also be processed by an encoder (VAE encoder) to obtain a corresponding first vector representation, and the target mask image will also be downsampled (resized), and then a second vector representation of the downsampled target mask image will be obtained. Then, the first vector representation and the second vector representation will be concatenated to obtain a concatenated vector. Then, the concatenated vector will be input into the target expansion network, so that the target expansion network, under the control of the aforementioned image features, generates a corresponding replacement result image based on the concatenated vector. Among them, the image content output by the target expansion network is also in the form of a vector. Furthermore, the decoder (VAE decoder) can be used to decode the vector-form image content output by the target expansion network to obtain a corresponding replacement result image. Among them, the VAE (Variational AutoEncoder) encoder can be used to compress an image into a latent space to obtain a corresponding latent variable, and the VAE decoder can be used to convert the latent variable into an image.

[0111] Among them, since the target expansion network is pre-trained with a large number of pictures, the generated images are relatively good in terms of semantics and image quality, which can make up for the shortcomings of the solution based on the generative adversarial network. At the same time, it is controlled by the output result of the generative adversarial network, so that the content generated by the target expansion network is controllable and semantically reasonable.

[0112] It should be noted that in the embodiments of the present application, a part of the image can be intercepted from the image to be processed and the target mask image for subsequent processing. In this case, after obtaining the image to be processed and the target mask image, a first part of the image can be obtained from the image to be processed, and a second part of the image corresponding to the first part of the image can be obtained from the target mask image. The first part of the image includes the target object, and the second part of the image includes the target mask area. In this case, a second replacement content can be obtained based on the first part of the image, the second part of the image, and the object generation network, and then the first replacement content can be obtained based on the target expansion network and the second replacement content.

[0113] Among them, in the case of intercepting a part of the image from the image to be processed and the target mask image for subsequent processing, it is also applicable to Figure 15 the processing flow shown.

[0114] An image processing method provided in this embodiment obtains a second replacement content based on an image to be processed, a target mask image, and a generative adversarial network, and obtains a first replacement content based on a target expansion network and the second replacement content. Moreover, when the target expansion network is trained using training data of an image category, the image content generated by the target expansion network is more controllable and has better image quality. As a result, after replacing the image content corresponding to the target object in the image to be processed based on the first replacement content, the obtained target image has better image quality, thereby improving the replacement effect of the image content.

[0115] It should be noted that for an image content processing model, during the process of generating the first replacement content corresponding to the target object, the specific category of the generated first replacement content can be determined by training the image content processing model. For example, among multiple images used to train the image content processing model, data labels can be carried, and these data labels are used to distinguish different contents (contents in the foreground area or background area) in the images used for training.

[0116] The specific training process may include:

[0117] Identify the training images to obtain the initial training mask images corresponding to the training images. The initial mask images include the first mask regions corresponding to the objects in the training images.

[0118] If the ratio of the intersection region between the reference mask region randomly generated in the initial training mask image and the first mask region to the first mask region is greater than or equal to the specified threshold, then remove the intersection region in the first mask region to obtain the second mask region.

[0119] Among them, the randomly generated reference mask can be a mask of any irregular shape, a mask of a regular shape (for example, a rectangle), or a mask generated locally in history. As Figure 16 shown, a mask region can be randomly selected from the random mask region, the rectangular mask region, and the local mask region as the reference mask region and generated in the initial mask image to perform the operation of removing the target occlusion. The process of the operation of removing the target occlusion can be as Figure 17 and Figure 18 shown.

[0120] Exemplarily, as Figure 17 shown, Figure 17 the left image shown is the initial training mask image, and a triangular first mask region is included in the initial mask image. As Figure 17 shown in the middle image, the reference mask region randomly generated in the initial mask image is Figure 17If the reference mask region of the rectangle shown in the middle image, which intersects with the first mask region and the ratio of the intersection region to the first mask region is greater than or equal to the specified threshold, the intersection region will be removed from the first mask region to obtain Figure 17 the second mask region shown in the right image, and then obtain the target training mask image.

[0121] Based on the second mask region, obtain the target training mask image corresponding to the training image.

[0122] If the ratio of the intersection region of the reference mask region randomly generated in the initial training mask image to the first mask region is less than the specified threshold, the initial training mask image will be used as the target training mask image corresponding to the training image.

[0123] Exemplarily, as Figure 18 shown, as Figure 18 shown in the middle image of Figure 18 the reference mask region of the rectangle shown in the middle image of the initial mask image, which intersects with the first mask region and the ratio of the intersection region to the first mask region is less than the specified threshold, then as Figure 18 shown in the right image of , directly use the initial training mask image as the target training mask image.

[0124] In the embodiments of the present application, the specific value of the specified threshold is not specifically limited. For example, the specified threshold can be 40%, 50% or 60%.

[0125] Based on multiple training images and the corresponding target training mask images, train the initial image content processing model to obtain the image content processing model.

[0126] Through the above training method, it can be made that in the process of generating the first replacement content or the second replacement content by the trained image content processing model, more background content is generated.

[0127] Please refer to Figure 19 , an image processing apparatus 500 provided in the embodiments of the present application, the apparatus 500 includes:

[0128] A target object determination unit 610, configured to determine a target object from the image to be processed.

[0129] A mask acquisition unit 620, configured to acquire a target mask image corresponding to the image to be processed, where the target mask image includes a target mask region corresponding to the target object.

[0130] A replacement content acquisition unit 630 is configured to obtain a first replacement content corresponding to a target object output by an image content processing model based on an image to be processed, a target mask image, and the image content processing model. The image content processing model generates the first replacement content through a generative adversarial network and a target expansion network, and the target expansion network is trained with training data of image categories.

[0131] An image fusion unit 640 is configured to change the image content corresponding to the target object in the image to be processed to the first replacement content to obtain a target image.

[0132] As a way, the replacement content acquisition unit 630 is specifically configured to obtain a first part of the image from the image to be processed and a second part of the image corresponding to the first part of the image from the target mask image. The first part of the image includes the target object, and the second part of the image includes the target mask region. The first part of the image and the second part of the image are input into the image content processing model, and the first replacement content corresponding to the target object output by the image content processing model is obtained.

[0133] Optionally, the replacement content acquisition unit 630 is specifically configured to obtain the second part of the image corresponding to the first part of the image from the target mask image, including: obtaining the image corresponding to the second circumscribed rectangle of the target mask region in the target mask image as the second part of the image, and the second circumscribed rectangle corresponds to the first circumscribed rectangle.

[0134] Optionally, the replacement content acquisition unit 630 is specifically configured to input the first part of the image and the second part of the image into the image content processing model to output a replacement result image through the image content processing model. The replacement result image is an image obtained by changing the image content corresponding to the target object in the first part of the image to the first replacement content. Based on the color information of the image content other than the first replacement content in the replacement result image and the color information of the image content other than the target object in the first part of the image, the color of the first replacement content is corrected to obtain the first replacement content with the corrected color. In this way, the image fusion unit 640 is specifically configured to change the image content corresponding to the target object in the image to be processed to the first replacement content with the corrected color to obtain a target image.

[0135] As Figure 20 shown, the apparatus 500 further includes:

[0136] An image recognition unit 650 is configured to recognize a to-be-processed image to obtain candidate objects in the to-be-processed image and an initial mask image corresponding to the to-be-processed image. The initial mask image includes initial mask regions corresponding to the candidate objects. A mask acquisition unit 620 is specifically configured to remove the initial mask regions of the remaining candidate objects in the initial mask image to obtain a target mask image corresponding to the to-be-processed image. The remaining candidate objects are candidate objects other than the target object.

[0137] Optionally, the mask acquisition unit 620 is specifically configured to enlarge the initial mask region of the target object to obtain an enlarged mask region; if the enlarged mask region intersects with the initial mask regions of other candidate objects, obtain the intersecting regions; and remove the intersecting regions from the enlarged mask region to obtain the target mask region of the target object.

[0138] As a way, a replacement content acquisition unit 630 is specifically configured to obtain a second replacement content based on the to-be-processed image, the target mask image, and a generative adversarial network; and obtain a first replacement content based on a target expansion network and the second replacement content.

[0139] Optionally, the replacement content acquisition unit 630 is specifically configured to process the color values of the image corresponding to the target object in the to-be-processed image to a target value based on the target mask image to obtain a to-be-processed image with processed colors, where the target value is used to mark positions where content replacement is required; and obtain a second replacement content based on the to-be-processed image with processed colors and a generative adversarial network.

[0140] Optionally, the replacement content acquisition unit 630 is specifically configured to, under the control of a target vector, obtain a first replacement content based on the to-be-processed image with processed colors and the target mask image by a target expansion network, where the target vector is obtained by performing feature extraction on the second replacement content.

[0141] It should be noted that the apparatus embodiments in this application correspond to the foregoing method embodiments. The specific principles in the apparatus embodiments can refer to the content in the foregoing method embodiments and will not be elaborated here.

[0142] Next, Figure 21 an electronic device provided by this application will be described.

[0143] Please refer to Figure 21, based on the above image processing method and apparatus, another electronic device 100 capable of executing the foregoing image processing method is further provided in an embodiment of the present application. The electronic device 100 includes one or more (only one is shown in the figure) processors 102, a memory 104, and a network module 106 that are coupled to each other. Among them, the memory 104 stores a program that can execute the content in the foregoing embodiments, and the processor 102 can execute the program stored in the memory 104.

[0144] Among them, the processor 102 may include one or more processing cores. The processor 102 connects various parts within the entire electronic device 100 through various interfaces and lines, and executes various functions of the electronic device 100 and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 104, and by calling data stored in the memory 104. Optionally, the processor 102 may be implemented in at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). The processor 102 may integrate a combination of one or more of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. Among them, the CPU mainly processes the operating system, user interface, application programs, etc.; the GPU is responsible for rendering and drawing display content; the modem is used to process wireless communication. It can be understood that the above modem may not be integrated into the processor 102 and may be implemented separately by a communication chip.

[0145] The memory 104 may include random access memory (RAM) and may also include read-only memory. The memory 104 can be used to store instructions, programs, codes, code sets, or instruction sets. The memory 104 may include a program storage area and a data storage area. Among them, the program storage area may store instructions for implementing the operating system, instructions for implementing at least one function (such as touch function, sound playback function, image playback function, etc.), instructions for implementing the following various method embodiments, etc. The data storage area may also store data created during the use of the terminal 100 (such as phone book, audio and video data, chat record data, etc.).

[0146] The network module 106 is used to receive and transmit electromagnetic waves, realizing the mutual conversion between electromagnetic waves and electrical signals, so as to communicate with a communication network or other devices, such as communicating with an audio playback device. The network module 106 may include various existing circuit elements for performing these functions, such as antennas, radio frequency transceivers, digital signal processors, encryption / decryption chips, subscriber identity module (SIM) cards, memories, and so on. The network module 106 can communicate with various networks such as the Internet, enterprise intranets, wireless networks or communicate with other devices through a wireless network. The above-mentioned wireless network may include a cellular phone network, a wireless local area network or a metropolitan area network. For example, the network module 106 can interact with a base station.

[0147] Please refer to Figure 22 , which shows a structural block diagram of a computer-readable storage medium provided by an embodiment of the present application. Program code is stored in the computer-readable medium 800, and the program code can be called by a processor to execute the method described in the above method embodiment.

[0148] The computer-readable storage medium 800 can be an electronic memory such as a flash memory, an EEPROM (electrically erasable programmable read-only memory), an EPROM, a hard disk or a ROM. Optionally, the computer-readable storage medium 800 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 800 has a storage space for the program code 810 that executes any method step in the above method. These program codes can be read out from or written into one or more computer program products. The program code 810 can be compressed in an appropriate form, for example.

[0149] In summary, for an image processing method, apparatus, and electronic device provided by the present application, after determining a target object from an image to be processed, a target mask image corresponding to the image to be processed can be obtained. Then, based on the image to be processed, the target mask image, and an image content processing model, a first replacement content output by the image content processing model corresponding to the target object is obtained, and the image content corresponding to the target object in the image to be processed is changed to the first replacement content to obtain a target image. Therefore, when the image content processing model generates the first replacement content through a generative adversarial network and a target extension network, and the target extension network is trained with training data of image categories, the image content (e.g., the first replacement content) generated by the target extension network is more controllable and has better image quality. As a result, after replacing the image content corresponding to the target object in the image to be processed with the first replacement content, the obtained target image has better image quality, thereby improving the replacement effect of the image content.

[0150] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.

Claims

1. An image processing method, characterized in that, the method includes: determining a target object from the image to be processed; obtaining a target mask image corresponding to the image to be processed, the target mask image including a target mask region corresponding to the target object; based on the image to be processed, the target mask image, and an image content processing model, obtaining a first replacement content corresponding to the target object output by the image content processing model, wherein the image content processing model generates the first replacement content through a generative adversarial network and a target expansion network, and the target expansion network is trained with training data of image categories; changing the image content corresponding to the target object in the image to be processed to the first replacement content to obtain a target image.

2. The method according to claim 1, characterized in that, the obtaining the first replacement content corresponding to the target object output by the image content processing model based on the image to be processed, the target mask image, and the image content processing model includes: obtaining a first partial image from the image to be processed, and obtaining a second partial image corresponding to the first partial image from the target mask image, the first partial image including the target object, and the second partial image including the target mask region; inputting the first partial image and the second partial image into the image content processing model, and obtaining the first replacement content corresponding to the target object output by the image content processing model.

3. The method according to claim 2, characterized in that, the obtaining the first partial image from the image to be processed includes: obtaining an image corresponding to the first circumscribed rectangle of the target object in the image to be processed as the first partial image; the obtaining the second partial image corresponding to the first partial image from the target mask image includes: obtaining an image corresponding to the second circumscribed rectangle of the target mask region in the target mask image as the second partial image, and the second circumscribed rectangle corresponds to the first circumscribed rectangle.

4. The method according to claim 2, characterized in that, the inputting the first partial image and the second partial image into the image content processing model, and obtaining the first replacement content corresponding to the target object output by the image content processing model includes: inputting the first partial image and the second partial image into the image content processing model to output a replacement result image through the image content processing model, the replacement result image being an image obtained by changing the image content corresponding to the target object in the first partial image to the first replacement content; correcting the color of the first replacement content based on the color information of the image content other than the first replacement content in the replacement result image and the color information of the image content other than the target object in the first partial image to obtain a first replacement content with corrected color; Changing the image content corresponding to the target object in the to-be-processed image to the first replacement content to obtain a target image includes: Changing the image content corresponding to the target object in the to-be-processed image to the first replacement content with the corrected color to obtain a target image.

5. The method according to claim 1, wherein, before determining the target object from the to-be-processed image, it further includes: Performing recognition on the to-be-processed image to obtain candidate objects in the to-be-processed image and an initial mask image corresponding to the to-be-processed image, and the initial mask image includes initial mask regions corresponding to the candidate objects; Obtaining the target mask image corresponding to the to-be-processed image includes: Removing the initial mask regions of the remaining candidate objects in the initial mask image to obtain the target mask image corresponding to the to-be-processed image, and the remaining candidate objects are candidate objects other than the target object.

6. The method according to claim 5, wherein, before removing the mask regions of the candidate objects other than the target object in the initial mask image to obtain the target mask image corresponding to the to-be-processed image, it further includes: Enlarging the initial mask region of the target object to obtain an enlarged mask region; If the enlarged mask region intersects with the initial mask regions of other candidate objects, obtaining the intersecting regions; Removing the intersecting regions from the enlarged mask region to obtain the target mask region of the target object.

7. The method according to claim 1, wherein, Based on the to-be-processed image, the target mask image, and the image content processing model, obtaining the first replacement content corresponding to the target object output by the image content processing model includes: Based on the to-be-processed image, the target mask image, and the adversarial generation network, obtaining a second replacement content; Based on the target extension network and the second replacement content, obtaining the first replacement content.

8. The method according to claim 7, wherein, Based on the to-be-processed image, the target mask image, and the adversarial generation network, obtaining the second replacement content includes: Based on the target mask image, processing the color value of the image corresponding to the target object in the to-be-processed image into a target value to obtain a to-be-processed image with processed color, and the target value is used to mark the position where content replacement is required; Based on the to-be-processed image with processed color and the adversarial generation network, obtaining the second replacement content.

9. The method according to claim 8, wherein, Based on the target extension network and the second replacement content, obtaining the first replacement content includes: Under the control of a target vector, the target extension network obtains the first replacement content based on the to-be-processed image with processed color and the target mask image, and the target vector is obtained by performing feature extraction on the second replacement content.

10. The method according to claim 1, wherein, The method further includes: Identify the training image to obtain the initial training mask image corresponding to the training image, where the initial mask image includes a first mask region of the object in the training image; If the ratio of the intersection region of the reference mask region randomly generated in the initial training mask image to the first mask region is greater than or equal to a specified threshold, then remove the intersection region in the first mask region to obtain a second mask region; Based on the second mask region, obtain the target training mask image corresponding to the training image; If the ratio of the intersection region of the reference mask region randomly generated in the initial training mask image to the first mask region is less than the specified threshold, then use the initial training mask image as the target training mask image corresponding to the training image; Based on multiple training images and the corresponding target training mask images, train the initial image content processing model to obtain an image content processing model.

11. An image processing device, Characterized in that, The device includes: A target object determination unit for determining a target object from the image to be processed; A mask acquisition unit for acquiring a target mask image corresponding to the image to be processed, where the target mask image includes a target mask region corresponding to the target object; A replacement content acquisition unit for obtaining a first replacement content corresponding to the target object output by the image content processing model based on the image to be processed, the target mask image, and the image content processing model, where the image content processing model generates the first replacement content through a generative adversarial network and a target extension network, and the target extension network is trained through training data of image categories; A picture fusion unit for changing the image content corresponding to the target object in the image to be processed to the first replacement content to obtain a target image.

12. An electronic device, Characterized in that, It includes a processor and a memory; one or more programs are stored in the memory and configured to be executed by the processor to implement the method according to any one of claims 1-10.

13. A computer-readable storage medium, Characterized in that, Program code is stored in the computer-readable storage medium, where the method according to any one of claims 1-10 is executed when the program code is run by a processor.