Image elimination and repair method, electronic device, chip system and storage medium

CN121280281BActive Publication Date: 2026-09-11HONOR DEVICE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411128646.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2024-06-27
Filing Date
2024-08-15
Publication Date
2026-09-11
Estimated Expiration
2044-08-15

AI Technical Summary

Technical Problem

[0004]本申请提供一种图像消除及修复方法、电子设备、芯片系统及存储介质,解决了多人场景下消除某个人像又重新生成畸形人像的问题,提升了图像消除修复体验

Benefits of technology

[0049] For example, when the first and second portraits intersect, or the distance between the two portraits is less than a certain value, eliminating the first portrait may also eliminate the second portrait, generating a deformed portrait in the elimination area. The solution in this application determines the positional relationship between the object to be eliminated and other portraits, and determines whether it is a crowd background based on the positional relationship. Then, if it is determined to be a crowd background, the first image, the second mask image, and a powerful image elimination model are used to avoid generating deformed portraits.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121280281B_ABST
    Figure CN121280281B_ABST
Patent Text Reader

Abstract

The application provides an image erasing and repairing method, an electronic device, a chip system and a storage medium, and relates to the field of image processing. In response to a user operation on a first image, a to-be-erased object in a selected region is identified and a background image is identified. In a case where the to-be-erased object is a portrait and the background image is a crowd background image, according to the first image, a first mask image (including a mask region corresponding to the to-be-erased object) and a second mask image (including a mask region of each portrait in the first image), the to-be-erased object is subjected to image erasing and background repairing to generate a second image, and the selected region of the second image does not include a portrait. Through the scheme, even if there are multiple portraits in a photo, the portrait can be accurately erased after a certain portrait is selected, and no image of other people or objects, for example, no deformed portrait, is generated after the repairing processing, thereby improving the effect of image erasing and repairing.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to Chinese Patent Application No. 202410858256.X, filed on June 27, 2024, entitled "Image Processing Method, Electronic Device, Chip System and Storage Medium", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of image processing technology, and in particular to an image removal and restoration method, electronic device, chip system and storage medium. Background Technology

[0003] When users take photos outdoors with their mobile phones, passersby may appear in the photos. In such cases, users usually want to remove the images of passersby from the photos. However, in some situations, the area of ​​the passerby image to be removed in the photo may be obscured or very close to other people's images. In these cases, if the user uses the removal and repair function to remove the passerby image, it is very likely that images of other people or objects (such as deformed human figures) will be generated at the removal location. As a result, the removed and repaired photo will not meet the user's needs and will reduce the user experience. Summary of the Invention

[0004] This application provides an image removal and restoration method, electronic device, chip system, and storage medium, which solves the problem of removing a person's image and then regenerating a deformed image in a multi-person scene, thus improving the image removal and restoration experience.

[0005] In a first aspect, this application provides an image removal and repair method, the method comprising: activating an image removal function in response to a first user operation; highlighting a selected area in the first image in response to a second user operation on a first image; identifying whether the object to be removed in the selected area is a human portrait; identifying whether the background image of the object to be removed is a crowd background image; when the object to be removed is a human portrait and the background image of the object to be removed is a crowd background image, performing image removal and repair processing on the object to be removed according to the first image, a first mask image, and a second mask image to obtain a second image; wherein, the mask area in the first mask image is the mask area corresponding to the object to be removed, and the second mask image includes the mask area of ​​each human portrait in the first image; displaying the second image; wherein, the selected area of ​​the second image does not include human portraits.

[0006] It should be noted that the first mask image and the second mask image are obtained based on the first image.

[0007] The background image of the object to be eliminated is the image surrounding the object.

[0008] The image removal and restoration method provided in this application, in response to a user's selection of a region on a first image, can identify the object to be removed in the selected region and its background image, so as to adopt an appropriate image erasure and restoration processing strategy to obtain a better image erasure and restoration effect. When the object to be removed is a human portrait and the background image of the object to be removed is a crowd background image, image removal and background restoration can be performed on the object to be removed based on the first image, the first mask image (including the mask area corresponding to the object to be removed), and the second mask image (including the mask area of ​​each human portrait in the first image), to obtain a second image. The selected area of ​​the second image does not include human portraits. With this solution, in practical use, even if there are multiple human portraits in a photo, it is possible to accurately remove a human portrait after selecting it, and perform background restoration in the removed area without generating images of other people or objects, such as deformed human portraits, thereby improving the removal and restoration effect.

[0009] It's important to note that the object to be removed in the selected area may be occluded by other images. In this case, after the object is removed, other images (such as deformed human figures) may be generated in the removed area. This is because the removal and restoration mechanism typically uses model inference based on the images surrounding the removal area to restore it. Since the object to be removed is in a crowd background with multiple human figures around the removal area, model inference based on the surrounding images is likely to generate a single human figure. Consequently, a new human figure may be generated in the removed area, and it is very likely to be deformed, resulting in poor restoration. Therefore, it is essential to employ more effective image removal and restoration methods to accurately remove selected human figures without generating images of other people or objects.

[0010] The proposed solution can determine the positional relationship between the object to be removed and other human figures, and determine whether it is a crowd background based on the positional relationship. Then, if it is determined to be a crowd background, the solution uses a first image, a first mask image, a second mask image, and a powerful image removal model to perform image processing, thereby avoiding the generation of deformed human figures.

[0011] In this context, "eliminate" can also be replaced with "erase," "wipe," "remove," or other words that are the same as or similar to "eliminate."

[0012] Among them, restoration refers to the restoration of the background image after the image has been removed.

[0013] The mask image can also be called a mask image, a masking image, or a masking region.

[0014] The selected area can be a regular shape such as a rectangle or an ellipse, or it can be an irregular shape, depending on the actual use case.

[0015] For example, the first operation is clicking the smart elimination control. The second operation is a selection operation or a smearing operation. For instance, the second operation could be the user sliding their finger around the image to be eliminated or dragging the elimination cursor on the first image, thereby forming a selected area of ​​regular or irregular shape.

[0016] In some possible implementations, highlighting a selected area in the first image in response to a second user operation on the first image includes: determining the closed area formed by the movement trajectory of the circle operation as the selected area in response to a user's circle operation on the first image; or, determining the smeared area as the selected area in response to a user's smear operation on the first image.

[0017] With this solution, users can directly select an area on a photo and remove the image of a person or object in that area with one click, while restoring the background image in the removed area, thus improving the removal and repair effect.

[0018] This application provides the following two user scenarios applicable to the solution proposed in this application:

[0019] User Scenario 1: In response to user input, the electronic device enters the gallery or photo album interface. The gallery interface displays the first image. After the user clicks the edit option in the gallery interface, the electronic device displays options such as doodle and AI removal. When the AI ​​removal option is selected, in response to the user's selection or smearing action on a certain area of ​​the first image, the electronic device removes the object (e.g., a person or object) in the selected area and performs restoration processing on the removed area to ensure that the removed area is consistent with the background image, thus improving the image removal effect.

[0020] User Scenario 2: With the camera app enabled, in response to the user clicking the shutter button on the shooting interface, the electronic device captures a first image and displays a thumbnail of the first image on the shooting interface. In response to the user's interaction with the thumbnail, the electronic device navigates to the gallery interface, where the first image is displayed. After the user clicks the edit option in the gallery interface, the electronic device displays options such as drawing and AI removal. When the AI ​​removal option is selected, in response to the user's selection or drawing action on a specific area of ​​the first image, the electronic device removes objects (such as people or objects) from the selected area and performs retouching processing on the removed area to ensure that the removed area remains consistent with the background image, improving the image removal effect.

[0021] The proposed solution can eliminate human figures or objects in the user-selected area, meeting various image elimination needs in terms of computational load, image processing performance, and image processing effect. In particular, it can avoid regenerating other images such as deformed human figures in elimination scenarios with crowd backgrounds.

[0022] Specifically, in response to the user's operation of selecting a region on the first image, the object to be removed in the selected region can be identified, as well as the background image of the object to be removed, so as to adopt an appropriate image erasing and repair processing strategy and obtain a better image erasing and repair effect.

[0023] It should be noted that the solution in this application can identify whether the image removal scenario involves a crowd. If there is a crowd, a powerful image removal model can be used to avoid generating distorted human figures after image removal. If there is no crowd, an image removal model that meets the actual usage requirements can be used.

[0024] The following are the image erasure and restoration strategies selected under different conditions:

[0025] Strategy 1: Image Removal and Restoration Strategies for Crowd Backgrounds

[0026] When the object to be removed is a human portrait and the background image of the object to be removed is a crowd background image, an image removal and repair strategy based on the crowd background can be adopted.

[0027] In some possible implementations, the method further includes: inputting a first image into a first image segmentation model to obtain a second mask image; wherein the first image segmentation model is a portrait instance segmentation network model or a panoptic segmentation network model. Based on the second mask image, all human figures in the image can be eliminated, thereby avoiding the phenomenon of regenerating distorted human figures based on multiple human figures.

[0028] In some possible implementations, the step of performing image removal and repair processing on the object to be removed based on the first image, the first mask image, and the second mask image to obtain a second image includes: performing image processing based on the first image and the second mask image to obtain a third image; wherein the third image does not include a human image; and performing image processing based on the first image, the third image, and the first mask image to obtain the second image.

[0029] The first image and the second mask image can be input into the first image elimination model to obtain the third image.

[0030] The first image removal model is used to remove the image from the masked area and perform background restoration on the removed area.

[0031] By using the proposed solution, and by determining whether the background of the object to be removed is a crowd background, appropriate image removal models can be used to remove the image and repair the background, thereby improving the image removal effect.

[0032] Among some possible implementations, the first image removal model can be a stable diffusion (SD) based image removal model. For removal scenarios with crowd backgrounds, the requirements for removal and restoration are high, necessitating a high-performance network model. Therefore, using a higher-performance image removal model can achieve better removal and restoration results.

[0033] In some possible implementations, the step of performing image processing based on the first image, the third image, and the first mask image to obtain the second image includes: performing an AND operation on the third image and the first mask image to obtain a fourth image; performing an AND operation on the first image and the fifth image to obtain a sixth image; the fifth image is an image obtained by inverting the pixel values ​​of the first mask image; and performing image fusion on the fourth image and the sixth image to obtain the second image.

[0034] The proposed solution achieves better image restoration results by performing post-processing on the first image, the third image, and the first mask image.

[0035] Using the above scheme, when the selected area includes multiple portraits, the first image and the second mask image (including the mask area of ​​each portrait in the first image) can be input into the first image removal model to obtain a third image, which does not contain portraits. Then, image processing is performed based on the first image, the third image, and the first mask image (including the mask area corresponding to the object to be removed) to generate a second image. The selected area of ​​the second image does not contain portraits, that is, it does not contain the portrait to be removed, nor does it contain other regenerated portraits. Through the embodiments of this application, image removal and restoration in multi-person scenes are improved, solving the problem of regenerating distorted portraits after removing a certain portrait in a multi-person scene, thus improving the image removal and restoration experience.

[0036] In some possible implementations, performing a bitwise AND operation between the third image and the first mask image includes performing a bitwise AND operation between the third image and the dilated first mask image. The fifth image is obtained by inverting the pixel values ​​of the dilated first mask image.

[0037] The proposed solution achieves better image restoration results through processes such as mask dilation and image fusion.

[0038] In some possible implementations, the method further includes: when the object to be eliminated is determined to be a human image, determining whether the elimination operation meets the elimination conditions based on the second mask image and the third mask image; wherein the mask region of the third mask image is the mask region corresponding to the selected region.

[0039] Accordingly, identifying whether the background image of the object to be eliminated is a crowd background image includes: if it is determined that the elimination operation meets the elimination conditions, identifying whether the background image of the object to be eliminated is a crowd background image.

[0040] In some possible implementations, when the object to be eliminated is identified as a human image and the selected area covers part of the human body, it is determined whether the elimination operation meets the elimination conditions based on the second mask image and the third mask image.

[0041] The solution of this application can determine whether the area selected by the user (e.g., circled or painted) meets the elimination conditions in response to the user's selection operation on the first image. If the circled or painted area meets the elimination conditions, the object to be eliminated is eliminated. If the circled or painted area does not meet the elimination conditions, the object to be eliminated is not eliminated, so as to prevent accidental elimination, such as preventing accidental elimination of human body parts in the image.

[0042] In some possible implementations, before identifying whether the background image of the object to be eliminated is a crowd background image, the method further includes: determining a first area ratio based on the ratio of the number of pixels in the mask region of the first mask image to the number of pixels in the first mask image.

[0043] Accordingly, identifying whether the background image of the object to be eliminated is a crowd background image includes: identifying whether the background image of the object to be eliminated is a crowd background image when the first area ratio is greater than or equal to a first threshold and less than a second threshold.

[0044] For example, the first threshold is 10%, and the second threshold is 25%. The first and second thresholds are illustrative examples, and can be adjusted according to actual usage requirements in actual implementation.

[0045] In some possible implementations, identifying whether the background image of the object to be eliminated is a crowd background image includes: if the number of human figures in the second mask image is less than a third threshold, then determining that the background image of the object to be eliminated is not a crowd background image; if the number of human figures in the second mask image is greater than or equal to the third threshold, then determining a first human figure mask region in the second mask image based on the first mask image and the second mask image; if the first human figure mask region intersects with at least one other human figure mask region in the second mask image or the distance is less than a fourth threshold, then determining that the background image of the object to be eliminated is a crowd background image; if the minimum distance between the first human figure mask region and each human figure mask region in the second mask image is greater than or equal to the fourth threshold, then determining that the background image of the object to be eliminated is not a crowd background image.

[0046] For example, the third threshold can be 3, and the fourth threshold can be the product of the width of the smallest bounding rectangle of the determined first portrait mask area and a certain set value, such as 0.5.

[0047] In some possible implementations, the second mask image includes Y portrait mask regions. In this case, determining the first portrait mask region in the second mask image based on the first mask image and the second mask image includes: obtaining the i-th intersection area and the i-th union area between the first mask region and the i-th portrait mask region in the second mask image; if the ratio of the i-th intersection area and the i-th union area is greater than a threshold (e.g., 0.8), recording the i-th portrait mask region (i.e., the first portrait mask region). Wherein, 1≤i≤Y.

[0048] In some possible implementations, the first portrait mask region intersects with at least one other portrait mask region in the second mask image or the distance between them is less than a fourth threshold, including: if the minimum bounding rectangle of the i-th portrait mask region intersects with the minimum bounding rectangle of the j-th portrait mask region in the second mask image, then the portrait mask regions are determined to intersect; if the minimum distance between the minimum bounding rectangle of the i-th portrait mask region and the minimum bounding rectangle of the j-th portrait mask region in the second mask image is less than the fourth threshold, then the distance is determined to be less than the fourth threshold; wherein, 1≤j≤Y.

[0049] For example, when the first and second portraits intersect, or the distance between the two portraits is less than a certain value, eliminating the first portrait may also eliminate the second portrait, generating a deformed portrait in the elimination area. The solution in this application determines the positional relationship between the object to be eliminated and other portraits, and determines whether it is a crowd background based on the positional relationship. Then, if it is determined to be a crowd background, the first image, the second mask image, and a powerful image elimination model are used to avoid generating deformed portraits.

[0050] Through the embodiments of this application, in actual use, even if there are multiple portraits in a photo, it is possible to accurately remove a portrait after selecting a certain portrait, and restore the background image in the removed area, without generating images of other people or objects (such as deformed portraits), thus improving the effect of image removal and restoration.

[0051] Strategy 2: Image Removal and Restoration Strategies for Non-Crowded Backgrounds

[0052] When the background image of the object to be removed is not a crowd background image, an image removal and repair strategy that does not involve a crowd background can be adopted.

[0053] It should be noted that, as mentioned above, if the number of human figures in the second mask image is less than the third threshold (e.g., 3), then the background image of the object to be eliminated is determined to be a non-crowd background image (i.e., a non-crowd background).

[0054] Alternatively, if the number of human figures in the second mask image is greater than or equal to the third threshold, the positional relationship between the object to be eliminated and other human figures is further determined, and whether it is a crowd background is determined based on the positional relationship. If the minimum distance between the first human figure mask area and each human figure mask area in the second mask image is greater than or equal to the fourth threshold, the background image of the object to be eliminated is determined to be a crowd background image (i.e., non-crowd background).

[0055] In some possible implementations, the method further includes: when it is determined that the background image of the object to be eliminated is not a crowd background image, performing image elimination and repair processing on the object to be eliminated according to the first image, the first mask image, and the second image elimination model to obtain a repaired image; wherein, the second image elimination model is used to eliminate the image of the mask region and perform background repair on the eliminated region.

[0056] In some possible implementations, the step of performing image removal and restoration processing on the object to be removed based on the first image, the first mask image, and the second image removal model to obtain a restored image includes: cropping a first cropped image including the object to be removed from the first image; cropping a second cropped image including the first mask region from the first mask image, and performing mask dilation processing on the second cropped image; inputting the first cropped image and the dilated second cropped image into the second image removal model to obtain a model-generated image; and filling the cropped region of the first image with the model-generated image to obtain the restored image.

[0057] The proposed solution allows for image cropping of the first image, ensuring that the object to be removed is centered, in a suitable position, or at a suitable proportion in the cropped image, thereby improving the image removal and restoration effect.

[0058] In some possible implementations, for image removal scenarios without crowds in the background, a relatively low-computation image removal model can be used to meet the basic requirements of image removal and minimal color difference. The second image removal model is based on a Generative Adversarial Network (GAN).

[0059] Alternatively, the second image removal model can also adopt an SD-based image removal model.

[0060] In practical implementation, for regions with a small area in the image, the requirements for elimination and restoration are relatively low, and a GAN-based image elimination model with low computational cost can be used, and the image processing effect can meet basic requirements such as small color difference. For regions with a large area, the requirements for elimination and restoration are higher, so a SD-based image elimination model with better performance can be used to obtain better elimination and restoration effects.

[0061] Secondly, this application provides an image restoration and processing apparatus, which includes units for performing the method described in the first aspect above. This apparatus can correspond to performing the method described in the first aspect above. For a detailed description of the units within this apparatus, please refer to the description in the first aspect above; for brevity, it will not be repeated here.

[0062] The method described in the first aspect above can be implemented in hardware or by hardware executing corresponding software. The hardware or software includes one or more modules or units corresponding to the above functions. For example, a processing module or unit, a display module or unit, etc.

[0063] Thirdly, this application provides an electronic device, which includes a processor, a computer program or instructions stored in a memory, wherein the processor is used to execute the computer program or instructions to cause the method in the first aspect to be performed.

[0064] Fourthly, this application provides a computer-readable storage medium having a computer program (also referred to as instructions or code) stored thereon for implementing the method of the first aspect. For example, when the computer program is executed by a computer, it enables the computer to perform the method of the first aspect.

[0065] Fifthly, this application provides a chip including a processor. The processor is used to read and execute a computer program stored in a memory to perform the methods in the first aspect and any possible implementation thereof. Optionally, the chip further includes a memory connected to the processor via a circuit or wire.

[0066] Sixthly, this application provides a chip system including a processor. The processor is used to read and execute a computer program stored in a memory to perform the methods in the first aspect and any possible implementation thereof. Optionally, the chip system further includes a memory connected to the processor via a circuit or wire.

[0067] In a seventh aspect, this application provides a computer program product comprising a computer program (also referred to as instructions or code), which, when executed by an electronic device, causes the electronic device to implement the method in the first aspect.

[0068] It is understood that the beneficial effects of the second to seventh aspects mentioned above can be found in the relevant descriptions in the first aspect mentioned above, and will not be repeated here. Attached Figure Description

[0069] Figure 1 A schematic diagram of the interface for generating deformed human images after image removal and restoration in related technologies;

[0070] Figure 2A This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application;

[0071] Figure 2B A schematic diagram of the software architecture of an electronic device provided in an embodiment of this application;

[0072] Figure 3A schematic diagram of a user interface scenario for the image removal and restoration method provided in the embodiments of this application;

[0073] Figures 4A to 4C These are schematic diagrams of user interface scenarios for the image removal and restoration methods provided in the embodiments of this application;

[0074] Figure 5 A schematic flowchart illustrating the image removal and restoration method provided in this application embodiment;

[0075] Figure 6 This is a schematic diagram illustrating an application scenario of the image removal and restoration method provided in the embodiments of this application;

[0076] Figure 7 This is a schematic diagram of image cropping processing in the image removal and restoration method provided in the embodiments of this application;

[0077] Figure 8 This is a schematic diagram illustrating the cropping, scaling, and mask dilation processes in the image removal and restoration method provided in the embodiments of this application.

[0078] Figure 9 This is another schematic flowchart of the image removal and restoration method provided in the embodiments of this application;

[0079] Figure 10 This is a schematic diagram illustrating a scenario where the image removal and restoration method provided in this application removes images of objects that account for a relatively small percentage of the total image.

[0080] Figure 11 This is a flowchart illustrating the process of determining whether an image removal and restoration method has a complex background, as provided in the embodiments of this application.

[0081] Figure 12 This diagram illustrates the complex background judgment of objects with a slightly larger proportion and the execution of different image processing strategies in the image removal and restoration method provided in the embodiments of this application.

[0082] Figure 13A This diagram illustrates the image removal and restoration method provided in this application, which involves judging the background of a large proportion of objects and implementing different image processing strategies accordingly.

[0083] Figure 13B This illustration shows a non-human-background image removal scenario for the image removal and restoration method provided in the embodiments of this application. Figure 3 ;

[0084] Figure 13C Schematic diagram four illustrating the image removal scenario of a crowd background for the image removal and restoration method provided in this application embodiment;

[0085] Figure 13DAn improved schematic diagram of an image removal and restoration method for a crowd background scenario provided in the embodiments of this application;

[0086] Figure 14 This is a flowchart illustrating the process of determining whether the image removal and restoration method provided in this application is a scene with a crowd background.

[0087] Figure 15 This is another schematic diagram of the image removal and restoration method provided in the embodiments of this application;

[0088] Figure 16 This is a schematic diagram illustrating the acquisition of the second mask image in the image removal and restoration method provided in the embodiments of this application;

[0089] Figure 17 This is a flowchart illustrating the process of determining whether a selected or painted area meets the elimination conditions in the image elimination and repair method provided in this application embodiment.

[0090] Figure 18 This is a schematic diagram illustrating a scenario in the image removal and restoration method provided in this application, where the selected or painted area is determined to meet the removal conditions.

[0091] Figure 19 This is a schematic diagram illustrating a scenario in the image removal and repair method provided in this application, where the selected or painted area does not meet the removal conditions.

[0092] Figure 20 This is a schematic diagram illustrating a scenario in the image removal and repair method provided in this application, where the selected or painted area is determined to meet the removal conditions.

[0093] Figure 21 This is a flowchart illustrating the process of determining whether a selected or painted area meets the elimination conditions in the image elimination and repair method provided in this application embodiment.

[0094] Figure 22A and Figure 22B A schematic diagram of the user interface for the image removal and restoration method provided in the embodiments of this application;

[0095] Figure 23 This is a schematic diagram of the structure of an image processing device provided in an embodiment of this application. Detailed Implementation

[0096] In the embodiments of this application, the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature. In the description of this embodiment, unless otherwise stated, "multiple" means two or more.

[0097] To facilitate understanding of the embodiments of this application, the relevant concepts involved in the embodiments of this application will be briefly explained first.

[0098] 1. GAN-based image inpainting network model

[0099] Image inpainting network models based on GANs can be simply referred to as GAN image removal models. GAN image removal models are a technique for image inpainting using generative adversarial networks (GANs), which can repair and complete specified regions (masked regions) in an image. GAN image removal models can be deep learning network models trained through iterative adversarial training of a generator network and a discriminator network. The training samples consist of a large number of images (millions of images) and randomly generated removal regions (masked regions), with the training objective being to make the inpainted image as close as possible to restoring the original image.

[0100] 2. SD-based image inpainting network model

[0101] The image inpainting network model based on SD (Stable Diffusion) can be simply referred to as the SD image removal model. The SD image removal model is a method based on a stable diffusion model. This method has a high number of parameters (approximately 1GB) and a large computational cost. Because this method is trained on large-scale image-text data (5 billion image-text pairs), it has stronger image-text matching and background generation capabilities, thus enabling the SD image removal model to handle large-area and complex background removal problems. The SD image removal model can include an encoder network, a denoising UNet network, and a decoder network.

[0102] Currently, when users take photos outdoors with their mobile phones, passersby may appear in the photos. In such cases, users usually want to remove the images of passersby from the photos. Figure 1As shown, in some cases, the area of ​​the passerby image to be removed in the photo is obscured by other people's image areas or is relatively close to each other. In this case, if the user uses the removal and repair function to remove the passerby image, a deformed human image may be generated at the removal location after the passerby image is removed. As a result, the removed and repaired photo cannot meet the user's needs and reduces the user experience.

[0103] It should be noted that after removing a human figure, other images (such as deformed human figures) may be generated in the removed area. This is because the removal and restoration mechanism usually performs model inference based on the images around the removed area and then restores the removed area. Since the object to be removed is in a crowd background and there are multiple human figures around the area to be removed, the model inference based on the images around the area to be removed is likely to generate a human figure. Therefore, a human figure will be regenerated in the removed area, and it is likely to be a deformed human figure, thus failing to achieve a good restoration effect.

[0104] To address the aforementioned problems, this application provides an image removal and restoration method and electronic device. In an embodiment of this application, in response to a user's operation on a first image, the object to be removed in the selected area is identified, and the background image is also identified. When the object to be removed is a human portrait and the background image is a crowd background image, image removal and background restoration are performed on the object to be removed based on the first image, a first mask image (including the mask area corresponding to the object to be removed), and a second mask image (including the mask area of ​​each portrait in the first image), generating a second image. The selected area of ​​this second image does not include human portraits. Through this solution, even if there are multiple portraits in a photo, it is possible to accurately remove a specific portrait after selection, and no other images or objects are generated after restoration processing, such as deformed portraits, thereby improving the effect of image removal and restoration.

[0105] The image removal and repair method and electronic device in the embodiments of this application will now be described with reference to the accompanying drawings.

[0106] Figure 2A A hardware system for an electronic device applicable to this application is shown.

[0107] Electronic device 100 can be a mobile phone, smart screen, tablet computer, wearable electronic device, in-vehicle electronic device, augmented reality (AR) device, virtual reality (VR) device, laptop computer, ultra-mobile personal computer (UMPC), netbook, personal digital assistant (PDA), projector, etc. This application embodiment does not limit the specific type of electronic device 100.

[0108] Electronic device 100 may include processor 110, external memory interface 120, internal memory 121, universal serial bus (USB) interface 130, charging management module 140, power management module 141, battery 142, antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, audio module 170, speaker 170A, receiver 170B, microphone 170C, headphone jack 170D, sensor module 180, button 190, motor 191, indicator 192, camera 193, display screen 194, and subscriber identification module (SIM) card interface 195, etc. The sensor module 180 may include a pressure sensor 180A, a gyroscope sensor 180B, a barometric pressure sensor 180C, a magnetic sensor 180D, an accelerometer sensor 180E, a distance sensor 180F, a proximity sensor 180G, a fingerprint sensor 180H, a temperature sensor 180J, a touch sensor 180K, an ambient light sensor 180L, a bone conduction sensor 180M, etc.

[0109] It should be noted that, Figure 2A The structure shown does not constitute a specific limitation on the electronic device 100. In other embodiments of this application, the electronic device 100 may include... Figure 2A The components shown may include more or fewer components, or the electronic device 100 may include... Figure 1 The components shown may be a combination of certain components, or the electronic device 100 may include... Figure 2A Sub-components of some of the components shown. Figure 2A The components shown can be implemented in hardware, software, or a combination of both.

[0110] Processor 110 may include one or more processing units. For example, processor 110 may include at least one of the following processing units: application processor (AP), modem processor, graphics processing unit (GPU), image signal processor (ISP), controller, video codec, digital signal processor (DSP), baseband processor, and neural network processing unit (NPU). These different processing units may be independent devices or integrated devices. The controller can generate operation control signals based on instruction opcodes and timing signals to control instruction fetching and execution.

[0111] The processor 110 may also include a memory for storing instructions and data. In some embodiments, the memory in the processor 110 is a cache memory. This memory can store instructions or data that the processor 110 has just used or that are used repeatedly. If the processor 110 needs to use the instruction or data again, it can retrieve it directly from the memory. This avoids repeated accesses, reduces the waiting time of the processor 110, and thus improves the efficiency of the system.

[0112] For example, processor 110 can be used to execute the image removal and repair method of the embodiments of this application: in response to a first operation by a user, enabling the image removal function; in response to a second operation by a user on a first image, highlighting a selected area in the first image; identifying whether the object to be removed in the selected area is a human portrait; identifying whether the background image of the object to be removed is a crowd background image; when the object to be removed is a human portrait and the background image of the object to be removed is a crowd background image, performing image removal and repair processing on the object to be removed according to the first image, a first mask image, and a second mask image to obtain a second image; wherein, the mask area in the first mask image is the mask area corresponding to the object to be removed, and the second mask image includes the mask area of ​​each human portrait in the first image; displaying the second image; wherein, the selected area of ​​the second image does not include human portraits.

[0113] Figure 2A The connection relationships between the modules shown are merely illustrative and do not constitute a limitation on the connection relationships between the modules of the electronic device 100. Optionally, the modules of the electronic device 100 may also adopt a combination of various connection methods described in the above embodiments.

[0114] The wireless communication function of electronic device 100 can be realized through devices such as antenna 1, antenna 2, mobile communication module 150, wireless communication module 160, modem processor and baseband processor.

[0115] Electronic device 100 can implement display functions through a GPU, a display screen 194, and an application processor. The GPU is a microprocessor for image processing, connected to the display screen 194 and the application processor. The GPU is used to perform mathematical and geometric calculations and for graphics rendering. Processor 110 may include one or more GPUs, which execute program instructions to generate or modify display information.

[0116] The display screen 194 can be used to display images or videos. Exemplarily, in an embodiment of this application, the display screen 194 can be used to display a first captured image (the image to be removed and repaired), and to display a second image (the image after removal and repair).

[0117] Electronic device 100 can perform shooting functions through ISP, camera 193, video codec, GPU, display screen 194 and application processor.

[0118] The ISP is used to process data fed back from the camera 193. For example, when taking a picture, the shutter is opened, and light is transmitted through the camera to the camera's photosensitive element. The light signal is converted into an electrical signal, and the camera's photosensitive element transmits the electrical signal to the ISP for processing, transforming it into an image visible to the naked eye. The ISP can perform algorithmic optimization of image noise, brightness, and color. The ISP can also optimize parameters such as exposure and color temperature of the shooting scene. In some embodiments, the ISP can be set in the camera 193.

[0119] The camera 193 (also known as a lens) is used to capture still images or videos. It can be activated via application commands to enable photo-taking, such as capturing images of any scene. The camera may include components such as an imaging lens, filters, and an image sensor. Light emitted or reflected by an object enters the imaging lens, passes through the filter, and is finally focused onto the image sensor. The imaging lens is primarily used to focus and image the light emitted or reflected by all objects within the shooting field of view (also known as the scene to be captured, the target scene, or the scene image the user expects to capture). The filter is primarily used to filter out excess light waves (such as infrared light waves other than visible light). The image sensor can be a charge-coupled device (CCD) or a complementary metal-oxide-semiconductor (CMOS) phototransistor. The image sensor is primarily used to perform photoelectric conversion on the received light signal, converting it into an electrical signal, which is then transmitted to the ISP for conversion into a digital image signal. The ISP outputs the digital image signal to the DSP for processing. The DSP converts the digital image signal into standard RGB, YUV, and other image signal formats.

[0120] In some embodiments, the electronic device 100 may include one or N cameras 193, where N is a positive integer greater than 1.

[0121] The camera 193 can be located in front of the electronic device 100 or on the back of the electronic device 100. The specific number and arrangement of the cameras can be set according to the requirements, and this application does not impose any restrictions.

[0122] A touch sensor 180K, also known as a touch device, can be disposed on a display screen 194. The touch sensor 180K and the display screen 194 together form a touchscreen, also known as a touch display. The touch sensor 180K is used to detect touch operations applied to or near it. The touch sensor 180K can transmit the detected touch operation to an application processor to determine the type of touch event. Visual output related to the touch operation can be provided through the display screen 194. In other embodiments, the touch sensor 180K may also be disposed on the surface of the electronic device 100, and in a different location from the display screen 194. Exemplarily, in an embodiment of this application, the touch sensor 180K can be used to detect user-triggered image deletion and restoration operations.

[0123] The hardware system of electronic device 100 has been described in detail above. The software system of electronic device 100 will be introduced below.

[0124] Figure 2BThis is a schematic diagram of the software system of the electronic device 100 provided in this application embodiment. The layered architecture divides the system into several layers, each with a clear role and division of labor. The layers communicate with each other through software interfaces. In some embodiments, the Android system is divided into five layers, from top to bottom: applications 201, application framework 202, hardware abstraction layer (HAL) 203, kernel 204, and hardware layer 205.

[0125] Application layer 201 may include a series of application packages. In this embodiment, the application packages may include applications such as camera and gallery.

[0126] The application framework layer 202 provides an application programming interface (API) and programming framework for applications in the application layer. The application framework layer 202 includes some predefined functions. In this embodiment, the application framework layer may include a camera access interface, wherein the image processing interface may include image removal and restoration services. The image processing interface is used to provide an API and programming framework for the gallery application.

[0127] The hardware abstraction layer 203 is an interface layer located between the application framework layer 202 and the kernel layer 204, providing a virtual hardware platform for the operating system. In this embodiment, the hardware abstraction layer 203 may include a camera hardware abstraction layer and an image processing algorithm library.

[0128] The camera hardware abstraction layer (HAL) can provide virtual hardware for camera devices such as the front and rear cameras (referred to as camera HAL). The image processing algorithm library may include image removal and restoration algorithms; in other words, the image processing algorithm library contains runtime code and data that implement the image removal and restoration methods provided in the embodiments of this application.

[0129] Kernel layer 204 is the layer between hardware and software. It includes drivers for various hardware components. Kernel layer 204 may include camera device drivers, digital signal processor (DSP) drivers, and image processor (IP) drivers, among others. Specifically, the camera device driver drives the camera sensor to acquire images and the image signal processor to preprocess them. The DSP driver drives the DSP to process images. The IIP driver drives the graphics processor to process images.

[0130] Hardware layer 205 includes camera sensors, image signal processors, digital signal processors, and image processors, etc.

[0131] It should be noted that although the embodiments of this application are described using the Android system as an example, the basic principles are also applicable to electronic devices based on operating systems such as iOS or Windows.

[0132] The image removal and repair method provided in this application can be implemented by the aforementioned electronic device (such as a mobile phone), or by a functional module and / or entity within the electronic device capable of implementing the image removal and repair method. Furthermore, the solution in this application can be implemented through hardware and / or software, the specific implementation depending on actual usage requirements, and this application does not impose any limitations. The following description uses an electronic device as an example, along with the accompanying drawings, to exemplarily illustrate the image removal and repair method provided in this application.

[0133] To better understand the embodiments of this application, the "image removal and repair" function provided in the embodiments of this application will be briefly described below with reference to a mobile phone interface diagram:

[0134] In some embodiments, the electronic device adds an "AI Removal" option to the editing interface of a gallery application (or photo album application). This "AI Removal" option serves as an entry point to trigger the display of image removal and restoration functions. Exemplarily, the image removal and restoration functions may include features such as intelligent selection and manual smearing. It should be noted that the various functions included in the image removal and restoration functions and their names are illustrative examples and are not limited in scope by this application.

[0135] Intelligent selection refers to the process by which electronic devices, in response to a user's swiping action on an image, select targets based on the swiping trajectory.

[0136] Manual smearing refers to the electronic device responding to the user's smearing action on an image and selecting the target based on the smeared area.

[0137] Figure 3 This illustration shows a user interface diagram of the "AI elimination" option in the image elimination and restoration method provided in this application embodiment. Figure 3 As shown in (a), the electronic device displays an icon for the Gallery app on desktop 10. When the user clicks the Gallery app icon, as shown in (a), the user sees the following: Figure 3 As shown in (b), the electronic device updates from displaying the desktop 10 to displaying the all-photos interface 11. When the user clicks on a photo in the all-photos interface 11, as shown in (b), the electronic device updates from displaying the desktop 10 to displaying the all-photos interface 11. Figure 3 As shown in (c), the electronic device updates from displaying all photos interface 11 to displaying a single photo interface 12. The single photo interface includes editing options 13. Figure 3As shown in (d), when editing option 13 is selected, the electronic device updates from displaying a single photo interface 12 to displaying an editing interface 14. Editing interface 14 includes the photo selected by the user and the "AI Removal" option 15. When the "AI Removal" option 15 is selected in editing interface 14, the electronic device displays the functions corresponding to the "AI Removal" option 15, such as intelligent selection and manual smoothing.

[0138] It should be noted that electronic devices can identify the selected area and the object to be removed through functions such as intelligent selection or manual smearing.

[0139] In some embodiments, the object to be eliminated can be a human body, such as a passerby in a photograph.

[0140] In other embodiments, the object to be eliminated can be an item, such as a trash can or clutter in a photograph.

[0141] According to the solution of this application, electronic devices can respond to the user's selection operation on the image, remove the object to be removed in the user-specified area (the specified area) of the image, and repair the removed area to generate a more aesthetically pleasing image.

[0142] Figure 4A This illustration shows a user interface diagram illustrating the image removal and restoration method provided in this application, where image removal and restoration are performed. For example... Figure 4A As shown in (a), the electronic device displays the user-selected photo (i.e., the original image) 16 in the editing interface 14 and displays the functions corresponding to the "AI Removal" option 15 in the function option area, such as intelligent selection and manual smearing. Figure 4A As shown in (a) and (b), when the user selects the "Smart Selection" function option 17, the electronic device can display a trajectory cursor 18 for smart selection in the original image 16. The user can drag the trajectory cursor 18 to select the objects to be eliminated in the original image 16. The user drags the trajectory cursor 18 from the starting point 19 to the ending point 20 and releases it; the selection trajectory is as follows. Figure 4A As shown by the dashed lines in (b) and (c), the selected trajectory can be closed manually by the user or automatically extended and closed by the electronic device. Figure 4A As shown in (d), the selection trajectory is closed, and the area enclosed by this trajectory is the area within the dashed box 21. The object to be eliminated is a human body, and the background is flowers and plants. Figure 4A As shown in (e), the image elimination and repair method provided in this application embodiment is used to eliminate and repair the object to be eliminated. The human image in the dashed frame 21 in the original image 16 is eliminated, while the background such as flowers and plants in the dashed frame 21 is still retained. The image after elimination and repair is shown in 22.

[0143] In some embodiments, after the user drags the trajectory cursor 18 from the starting point 19 to the ending point 20 and releases it, the object to be eliminated can be displayed using a dynamic flashing display method.

[0144] Figure 4B This illustration shows another user interface diagram illustrating the image removal and restoration method provided in this application, where image removal and restoration are performed. For example... Figure 4B As shown in (a) to (d), in response to the user's selection operation 23 on the original image, the electronic device recognizes that the object to be eliminated is a human image and meets the elimination conditions. The object to be eliminated 24 is highlighted. The electronic device performs image elimination and repair processing on the object to be eliminated 24 in the background, and then updates the original image. The updated image 25 does not include the object to be eliminated 24, and the eliminated area is repaired as the background image.

[0145] Figure 4C This illustration shows another user interface diagram illustrating the image removal and restoration method provided in this application, where image removal and restoration are performed. For example... Figure 4C As shown in (a) to (f), in response to the user's operation on the manual smear option 26, the electronic device starts the manual smear function; in response to the user's operation on the smear tool 27, the electronic device enables the smear tool 27; in response to the user's smear operation on the object to be eliminated 28 (trash can) in the original image, the electronic device recognizes that the object to be eliminated is not a human image and meets the elimination conditions. The electronic device performs image elimination and repair processing on the smeared area of ​​the object to be eliminated 28 in the background, and then updates and displays the original image. The updated image 29 does not include the object to be eliminated 28, and the eliminated area is repaired to the background image.

[0146] The image removal and repair method provided in this application will be described in detail below with reference to specific embodiments and the accompanying drawings.

[0147] Figure 5 This is a schematic flowchart illustrating an image removal and restoration method provided in an embodiment of this application. The method can be... Figure 1 The electronic device shown executes the method, which includes steps S301 to S304.

[0148] S301. In response to the user's selection operation (circle or smudge operation), identify the object to be eliminated in the first image.

[0149] According to the solution of this application, an electronic device can respond to the user's selection operation in an image (or first image, original image or image to be repaired), remove the portrait or object image selected by the user through circle or smear operation in the image, and repair the removed area to generate a more beautiful image.

[0150] The area circled or drawn by the user is called the selected area, which can include people or objects selected by the user in the image by circling or drawing.

[0151] For example, a selection operation can be an action where a user selects a region in an image by sliding their finger or dragging a cursor; this selection operation can be called smart selection. Accordingly, the smart-selected region can be defined as the selected area. Objects within the selected area are the objects to be eliminated.

[0152] As another example, the selection operation can be a user clicking on a specific area of ​​an image, which can trigger the electronic device to automatically select the object to be eliminated at the clicked location. Accordingly, the automatically selected area can be defined as the selected region.

[0153] For example, the selection operation can be a user's smearing action on a certain area of ​​an image. Accordingly, the area covered by the smearing can be defined as the selected area.

[0154] For ease of explanation, this application uses the example of a smart selection operation for illustration.

[0155] The objects in the selected area can be human bodies or items, etc.

[0156] In some embodiments, in response to a user's intelligent selection operation, a selection trajectory can be marked in the first image, and the selected area can be determined based on the selection trajectory. For example, the area enclosed by the selection trajectory can be determined as the selected area. Optionally, the selection trajectory can be a line trajectory, or it can be a closed rectangle, ellipse, or irregular shape. For ease of explanation, this application embodiment uses a line trajectory as an example for illustrative purposes.

[0157] In some embodiments, when the selected trajectory is closed, the area enclosed by the selected trajectory is defined as the selected area.

[0158] In other embodiments, if the distance between the start and end points of the selected trajectory meets a preset distance condition when the selected trajectory is not closed, the electronic device can automatically connect the start and end points to form a closed area of ​​the selected trajectory, and then use the closed area as the selected area.

[0159] In some embodiments, if the distance between the start and end points of the selected trajectory does not meet a preset distance condition (e.g., the distance between the start and end points is too large) when the selected trajectory is not closed, the electronic device determines that the selected trajectory does not meet the condition and displays a prompt message. For example, the prompt message may be: This selection operation is invalid; please select again.

[0160] S302. Obtain a first mask image based on the first image, the first mask image including a first mask region corresponding to the object to be eliminated.

[0161] The first mask image can be referred to as a mask image or a binary image. For example, the first mask image includes a mask region where pixels are marked as 1, and other regions where pixels are marked as 0.

[0162] Specifically, pixels in the first mask region are marked as 1 and displayed as white. Pixels in other regions of the first mask image, excluding the first mask region, are marked as 0 and displayed as black.

[0163] In some embodiments, the electronic device can determine a selected area based on a selection trajectory and determine the object to be eliminated based on the selected area, and then generate a first mask image. The shape of the mask area in the first mask image is equivalent to the shape of the object to be eliminated.

[0164] In other embodiments, the electronic device can obtain a first mask image by inputting an image marked with a circled trajectory into a mask generation model. The masked region in the first mask image corresponds to the selected region enclosed by the circled trajectory. The shape of the first masked region is similar to the shape of the region enclosed by the circled trajectory.

[0165] S303. Calculate the first area ratio (denoted as P) of the first mask region in the first mask image.

[0166] The area ratio of the first mask region in the first mask image is equal to the area ratio of the object to be removed in the first image.

[0167] In some embodiments, the first area ratio P can be obtained by acquiring the number of pixels in the first mask image and the number of pixels in the first mask region, and by calculating the ratio of the number of pixels in the first mask region to the number of pixels in the first mask image.

[0168] The number of pixels in the first mask image is the sum of the number of all 1 pixels and the number of all 0 pixels.

[0169] The number of pixels in the first mask region is equal to the total number of pixels 1 in the first mask image.

[0170] It can be understood that the area ratio of the first mask region in the first mask image is equivalent to the area ratio of the object to be removed in the first image.

[0171] Where 0 < P ≤ 100%.

[0172] It should be noted that when P is small (e.g., P < 2%), meaning the proportion of objects to be eliminated is small, elimination and repair are relatively easy; when P is large (e.g., 2% ≤ P < 25%), meaning the proportion of objects to be eliminated is large, elimination and repair are more complex and require higher standards.

[0173] S304. The image processing strategy corresponding to the first area ratio P is used to eliminate and repair the objects to be eliminated in the first image.

[0174] For example, when P < X1, the first image processing strategy is used to eliminate and repair the object to be eliminated.

[0175] When X1≤P<X2, the second image processing strategy is used to eliminate and repair the object to be eliminated.

[0176] When X2≤P<X3, the third image processing strategy is used to eliminate and repair the object to be eliminated.

[0177] When P≥X3, no elimination or repair processing is performed on the object to be eliminated.

[0178] Where 0 ≤ X1 < X2 < X3.

[0179] It should be noted that in actual implementation, the number of thresholds can be adjusted (increased or decreased) according to actual usage needs. For example, four thresholds X1, X2, X3, and X4 can be set, where 0 ≤ X1 < X2 < X3 < X4, making image processing more refined and greatly improving the removal and restoration effects. For ease of explanation, X1, X2, and X3 are used as examples in the following embodiments.

[0180] It should be noted that in actual implementation, the values ​​of X1, X2, and X3 can be set according to actual usage requirements, and this application does not limit them. For ease of explanation, the following embodiments use X1 as 2%, X2 as 10%, and X3 as 25% as examples for illustrative purposes.

[0181] With the solution proposed in this application, for objects to be eliminated with different area proportions, an image processing strategy corresponding to the area proportion can be selected to eliminate and repair the objects, resulting in different image processing effects.

[0182] It should be noted that the embodiments of this application do not limit the image removal models used by the three image processing strategies. The image removal models used by the three image processing strategies may be the same or different.

[0183] In some embodiments, the first image processing strategy, the second image processing strategy, and the third image processing strategy may employ the same or different image elimination models under different circumstances; the image processing procedures may be the same or different, for example, some perform image cropping processing while others do not; the computational load of image processing increases sequentially, and the image processing effect is enhanced sequentially.

[0184] In some embodiments, for objects to be eliminated that have a very small area ratio, a first image processing strategy is adopted to perform image cropping processing, and elimination and repair processing is performed based on the first mask image, the first image, and the GAN image elimination model.

[0185] In other embodiments, for objects to be eliminated that have a slightly larger area, a second image processing strategy is adopted to perform image cropping and determine whether the background of the object to be eliminated is a complex background. If it is a simple background (e.g., sky or solid color background), then elimination and repair processing is performed based on the first mask image, the first image, and the GAN image elimination model. If it is a complex background (e.g., crowd background), then elimination and repair processing is performed based on the first mask image, the first image, and the SD image elimination model.

[0186] In some embodiments, for objects to be eliminated that occupy a large area, a third image processing strategy is adopted. Instead of image cropping, it determines whether the elimination scenario has a crowd background. If it does not, elimination and repair processing is performed based on the first mask image and an image processing network model based on SD or GAN. If it does have a crowd background, elimination and repair processing is performed based on the first mask image, the second mask image, and the SD image elimination model. The first mask image includes the mask region corresponding to the object to be eliminated in the first image, and the second mask image includes the mask region corresponding to each portrait in the first image.

[0187] It should be noted that, compared to GAN image removal models, SD image removal models have higher computational requirements but superior image processing performance. SD image removal models can handle large-area and complex background removal problems.

[0188] Figure 6 The diagram illustrates three scenarios where different image processing strategies are used to eliminate and repair objects with different area proportions.

[0189] Scene 1:

[0190] like Figure 6As shown in (a), the first image includes the object to be eliminated, 31, indicated by the dashed line. The object to be eliminated is a male, and it corresponds to the first mask region 32 in the first mask image. The area ratio P of the first mask region in the first mask image is calculated, where P < 2%. Since the area ratio of the object to be eliminated is very small, the first image processing strategy can be adopted: based on the first mask image and the first image, the GAN image elimination model is used to eliminate and repair the object 31. This approach has a small computational load and the image processing effect meets basic requirements such as small color difference.

[0191] exist Figure 6 In the image shown in (a) after removal and restoration, the image to be removed (the male image) selected by the circle or smear operation in the first image is removed. It should be noted that for image removal scenarios where the object to be removed accounts for a small proportion and removal and restoration are relatively easy to achieve, the GAN image removal model can basically meet the removal and restoration requirements and save computational resources.

[0192] Scene Two:

[0193] like Figure 6 As shown in (b), the first image includes the object to be removed, 33, indicated by the dashed line. The object to be removed, 33, is a trash can and corresponds to the first mask region 34 in the first mask image. The area percentage P is calculated; 2% ≤ P < 10%. Since the area percentage of the object to be removed is slightly larger, a second image processing strategy can be adopted.

[0194] Case 1: If the background of the object to be eliminated 33 is a simple background (such as sky or solid color background), then the elimination and repair process is performed according to the first mask image and the first image and GAN image elimination model. The amount of computation is small and the basic requirements such as small color difference are met.

[0195] Case 2: If the background of the object to be eliminated 33 is a complex background (e.g., a crowd background), then the elimination and repair process is performed according to the first mask image and the first image and SD image elimination model. The elimination effect is good and meets the basic requirements such as small color difference.

[0196] exist Figure 6 In the image shown in (b) after removal and restoration, the image to be removed (the trash can image) selected by the circle or smear operation in the first image is eliminated. By adopting the second image processing strategy, the removal and restoration requirements can be met, while saving computational resources.

[0197] Third Scene:

[0198] like Figure 6As shown in (c), the first image includes the object to be removed, 35, indicated by the dashed line. The object to be removed 35 is a woman, and it corresponds to the first mask region 36 in the first mask image. The area percentage P is calculated; 10% ≤ P < 25%. Since the area percentage of the object to be removed is large, a third image processing strategy with high computational complexity and good image processing effect can be used to remove and repair the object 35.

[0199] Case 1: If the background of the object to be eliminated 35 is a non-human scene, then the elimination and repair process is performed according to the first mask image and the first image and SD image elimination model. The elimination effect is good and meets the basic requirements such as small color difference.

[0200] Scenario 2: If the background of the object to be removed 35 is a crowd scene, then the removal and repair processing is performed based on the first image, the first mask image, the second mask image, and the SD image removal model. The removal effect is good and meets basic requirements such as small color difference. The second mask image contains the masked areas of all human figures in the first image.

[0201] exist Figure 6 In the image after removal and restoration shown in (c), the image to be removed (the woman's image) selected by circling or smearing operations in the first image is removed. It should be noted that for image removal scenarios where the object to be removed occupies a large proportion and the requirements for removal and restoration are high, performing removal and restoration processing based on the first mask image, the first image, the second mask image, and the SD image removal model can well meet the removal and restoration requirements and achieve good removal and restoration results.

[0202] It should be noted that the embodiments of this application do not limit the image removal models used in the above three scenarios. For example, Figure 6 In (a), either a GAN-based image removal model or an SD-based image removal model can be used instead; similarly, in Figure 6 In (b) and (c), either an SD-based image removal model or a GAN-based image removal model can be used instead. The specific image removal model used can be determined based on the actual usage requirements.

[0203] The solution of this application allows for the selection of objects to be eliminated in the original image with different area proportions chosen by the user. An image processing strategy corresponding to the area proportion can be selected to eliminate and repair the objects, resulting in different image processing effects. The image elimination and repair method provided by this application can eliminate people or objects within selected or painted areas, meeting various image elimination needs in terms of computational load, image processing performance, and image processing effects.

[0204] The image removal and repair method provided in this application is applicable not only to image removal scenarios with simple backgrounds, but also to image removal scenarios with complex backgrounds and image removal scenarios with crowds in the background.

[0205] The method provided in this application selects an appropriate image processing strategy to eliminate and repair the object to be removed based on the area ratio of the selected or painted area in the original image. This results in more refined image processing and improved elimination and repair effects.

[0206] Image cropping:

[0207] In some embodiments, in a scenario where the object to be eliminated is selected by circling or smearing in the first image for elimination and repair processing, if the area ratio P meets certain conditions, the first image and the first mask image can be cropped first, and then the elimination and repair processing can be performed based on the cropped first image and the first mask image. This can optimize the effect of elimination and repair processing.

[0208] For example, when the area percentage P is less than X2 (e.g., 10%), the first image and the first mask image can be cropped separately. The cropped mask image contains the first mask region, and the cropped first image contains the object to be eliminated. Then, based on the cropped mask image and the cropped first image, the object to be eliminated selected by circling or smearing in the first image can be eliminated and repaired. The following describes the possible implementation methods for determining the cropping region and specific cropping provided by the embodiments of this application.

[0209] 1) Calculate the number of pixels in the first mask region (denoted as S).

[0210] 2) Multiply S by the coefficient corresponding to P to calculate the number of pixels in the cropped region (denoted as T).

[0211] In some embodiments, when P < 2%, the coefficient corresponding to P is N6, and the number of pixels in the cropped region T = S * N6 is calculated. Optionally, N6 can be 17. Multiplying S by 17 can be understood as cropping the image by 6%.

[0212] In some embodiments, when 2% ≤ P < 10%, the coefficient corresponding to P is N7, and the number of pixels in the cropped region T = S * N7 is calculated. Optionally, N7 can be 10. Multiplying S by 10 can be understood as cropping the image by 10%.

[0213] 3) Calculate the width and height (H1) of the cropped area (denoted as W1) based on the number of pixels T in the cropped area. For example, you can find the square root of T and determine the width W1 and height H1 of the cropped area. For instance, if T is 360000, then the width W1 and height H1 of the cropped area are both 600.

[0214] 4) such as Figure 7 As shown in (a), the width (denoted as W2) and height (denoted as H2) of the minimum bounding rectangle 17 of the first mask region are calculated.

[0215] 5) Adjust the width W1 and height H1 of the clipping region based on the width L2 and height H2 of the minimum bounding rectangle 17.

[0216] For example, if the width W1 of the clipping region is less than the width W2 of the minimum bounding rectangle 17, then W1 = W2 + N8.

[0217] For example, if the height H1 of the clipping region is less than the height H2 of the minimum bounding rectangle 17, then H1 = H2 + N8.

[0218] For example, if the width W1 or height H1 of the cropping region is less than N9, then the width W1 or height H1 of the cropping region is set to N9. Here, N9*N9 is the input image size defined by the inpainting network.

[0219] Optionally, N8 can be 6, and N9 can be 768.

[0220] 6) such as Figure 7 As shown in (a), the clipping region 42 is determined with the midpoint of the minimum bounding rectangle 41 of the first mask region as the center, according to the adjusted width L1 and height H1.

[0221] Here, it is assumed that the size of both the first image and the first mask image is a first dimension. For example, the first dimension is 1600*1200.

[0222] like Figure 7 As shown in (a) and (b) above, a first mask image of a second size is cropped from a first mask image, and a first image of a second size is cropped from a first image, following steps 1) to 6) above. For example, the second size is 1200*800.

[0223] Mask region dilation processing:

[0224] In some embodiments, in a scenario where the object to be eliminated is selected by circling or smearing in the first image for elimination and repair processing, the first mask region in the first mask image can be dilated first, and then the elimination and repair processing can be performed based on the dilated mask image, which can optimize the effect of elimination and repair processing.

[0225] In this embodiment of the application, the first mask image can be subjected to mask region dilation processing to increase the mask region, thereby increasing the scope of elimination and repair, so that the object to be eliminated can be eliminated and repaired more comprehensively, which can improve the elimination and repair effect.

[0226] In some embodiments, a mask dilation ratio N10 and an initial dilation kernel size N11 can be set. The dilation kernel size is increased by 1 each time, and the first mask region is cyclically dilated to obtain the dilated first mask image (denoted as dilatedMask) until the area (number of pixels) of the dilated first mask image is greater than or equal to R, where R = S * N10. S represents the area (number of pixels) of the first mask region.

[0227] Optionally, the expansion ratio N10 can be 1.2, and the initial expansion kernel size N11 can be 3.

[0228] Figure 8 This diagram illustrates cropping, scaling, and dilation of the first mask image. (See attached diagram.) Figure 8 As shown, the first mask image has a first size, the cropped first mask image has a second size, the dilated first mask image has a second size, and the dilated first mask area is larger than the undilated first mask area.

[0229] It is understandable that when N10 is greater than 1, the outline of the first mask region after expansion is larger than the outline of the first mask region before expansion. In this way, by increasing the mask region, the scope of elimination and repair is increased, thereby improving the elimination and repair effect.

[0230] Image removal and restoration process

[0231] The proposed solution allows for the selection of appropriate image processing methods for image removal and background restoration based on the area ratio of the object to be removed, resulting in more refined image processing and improved removal and restoration effects.

[0232] It should be noted that electronic devices can employ various possible methods to implement the above-mentioned S304, which uses an image processing strategy corresponding to the first area ratio P to eliminate and repair the objects to be eliminated in the first image (original image) by circling or smearing.

[0233] In this application embodiment, to facilitate the explanation of a specific scenario, the concept of a crowd background is introduced. This crowd background indicates different image features from the image background. The crowd background refers to the human portrait features relative to the object to be eliminated; a large number of human portraits exist around the object to be eliminated, forming a crowd background. The image background refers to image features in the image other than human portraits (such as trees, flowers, roads, bicycles, etc.).

[0234] It should also be noted that in the following examples, the process of image removal and background restoration involved in the solution of this application is represented by spots in order to better illustrate the effect of image removal and restoration.

[0235] The following details the possible implementation methods for image removal and background restoration by selecting an appropriate image processing method based on the area ratio of the object to be removed.

[0236] For example, combined Figure 5 ,like Figure 9 As shown, S304 may include the following S1 to S25.

[0237] S1. Determine whether P is less than 2%.

[0238] If P is determined to be less than 2%, continue with steps S2-S8, which involve using the first image processing strategy to perform image removal and repair.

[0239] If P is determined to be no less than 2%, proceed to S9, which further determines whether P is less than 10%.

[0240] When P is less than 2%, the first image processing strategy is adopted.

[0241] S2. If P is determined to be less than 2%, calculate the area to be clipped using a 6% mask ratio.

[0242] For example, calculating the cropping region with a 6% mask ratio involves multiplying the number of pixels in the first mask region by 17 to obtain the number of pixels T. Then, the square root of the number of pixels T is taken to obtain the width W1 and height H1. Next, based on the width W2 and height H2 of the minimum bounding rectangle of the first mask region, the width W1 and height H1 are adjusted (increased) to obtain the width W3 and height H3. The width W3 and height H3 are then used as the width and height of the cropping region. The specific calculation process for the cropping region can be found in the detailed description above regarding the calculation of the cropping region, and will not be repeated here.

[0243] S3. Crop and scale the first image and the first mask image to obtain the cropped and scaled first image and the first mask image.

[0244] First, the first image and the first mask image are cropped according to the calculated cropping area. The cropped first image includes the object to be removed, and the cropped first mask image includes the first mask area. For details on the image cropping process, please refer to [link to documentation]. Figure 7 As shown.

[0245] Then, both the cropped first image and the first mask image are scaled to N9*N9 (e.g., 768*768) to meet the requirements of the elimination and repair network for the input image size.

[0246] The process of cropping and scaling the first mask image can be found in [reference needed]. Figure 16 It should be noted that the process of cropping and scaling the first image is similar to the process of cropping and scaling the first mask image.

[0247] The dimensions of the first image and the first mask image can be a first size, and the dimensions of the cropped area can be a second size. The second size is smaller than the first size.

[0248] S4. Dilate the mask region of the scaled first mask image.

[0249] In this application embodiment, return to reference Figure 8 As shown, the mask region of the first mask image is dilated to increase the mask region, thereby increasing the range of elimination and repair, so that the object to be eliminated is eliminated and repaired more comprehensively, thus improving the elimination and repair effect.

[0250] S5. Input the cropped and scaled first image and the cropped, scaled and dilated first mask image into the GAN image elimination model for model inference to obtain the seventh image.

[0251] The size of the seventh image is specified by the GAN image elimination model. For example, the size of the seventh image is 768*768.

[0252] In some embodiments, the GAN image removal model can be used to repair the large mask LAMA model.

[0253] In some embodiments, the GAN image removal model can be an aggregated contextual-transformation (AOT) model.

[0254] In some embodiments, the GAN image removal model can be a mask-interactive generative adversarial network (MI-GAN) model.

[0255] S6. After scaling the seventh image to the second size, fill it into the cropped area of ​​the first image to obtain the third image of the first size.

[0256] In this embodiment of the application, the seventh image is scaled (e.g., enlarged) from the third size to the second size, i.e., size restoration.

[0257] In some embodiments, an image super-resolution method may be used to scale the second image to be eliminated. Super-resolution processing refers to the process of recovering a high-resolution image from a low-resolution image.

[0258] It is understandable that by scaling the seventh image to the same size as the cropped area (the second size), the scaled seventh image can be seamlessly filled into the cropped area of ​​the first image.

[0259] S7. After scaling the cropped, scaled, and dilated first mask image to the second size, fill it into the cropped area of ​​the first mask image to obtain the dilated and filled first mask image.

[0260] This application does not limit the execution order of S6 and S7. For example, S6 can be executed first and then S7; or S7 can be executed first and then S6; or S6 and S7 can be executed simultaneously.

[0261] S8. Based on the first image, the third image, and the first mask image after dilation and filling, obtain the second image.

[0262] In this embodiment of the application, the second image is obtained by the following image matrix calculation formula.

[0263] Pi = Po × (1 – Pm) + Pe × Pm.

[0264] Where Po represents the first image, Pe represents the third image, Pm represents the first mask image after dilation and filling, and Pi represents the second image.

[0265] Where (1–Pm) represents inverting the pixel values ​​of the first mask image.

[0266] Figure 10 The diagram illustrates the process of image removal and restoration using the GAN image removal model when P is less than 2%.

[0267] like Figure 10 As shown, in response to a user's selection operation (e.g., circling or smearing) on ​​the first image (original image), the electronic device determines the object to be removed and calculates the area percentage P of the object to be removed in the first image, which is less than 2%.

[0268] A first mask image is obtained based on the first image. The mask region in the first mask image is the mask region of the object to be eliminated. The size of both the first image and the first mask image is the first size.

[0269] Then, based on the mask region in the first mask image, a cropped region is calculated with a first mask percentage (e.g., 6%). The cropped region includes and is larger than the mask region, and its size is a second size. The first image and the first mask image are cropped based on the cropped region to obtain a first image of the second size and a first mask image of the second size, respectively.

[0270] Then, the first image of the second size and the first mask image of the second size are scaled to obtain the first image of the third size and the first mask image of the third size.

[0271] Then, the mask region of the first mask image with the third size is dilated to obtain the first mask image with the third size dilated.

[0272] Then, the first image at the third size and the first masked image at the third size are input into the GAN image elimination model to obtain the seventh image at the third size. The seventh image at the third size is then scaled to obtain the seventh image at the second size.

[0273] Then, the seventh image of the second size is filled into the cropped area of ​​the first image to obtain the third image of the first size.

[0274] On the other hand, the first mask image, after being expanded to a third size, is filled into the cropped area of ​​the first mask image of the first size to obtain the first mask image after being expanded and filled to the first size. Then, the pixel values ​​of the first mask image after being expanded and filled are inverted to obtain the fifth image. Then, the fifth image is multiplied by the first image to obtain the sixth image.

[0275] The third image and the first masked image after dilation and filling are ANDed (i.e. multiplied) to obtain the fourth image.

[0276] Then, the fourth and sixth images are fused together to obtain the second image.

[0277] like Figure 10 As shown, in the second image, the image to be eliminated (portrait) selected by circling or smearing is eliminated, and the background image of the eliminated area has been repaired.

[0278] With the proposed solution, for objects with a small area to be eliminated, the requirements for elimination and repair are relatively low. An image elimination model with low computational load can be used, and the image processing effect only needs to meet basic requirements such as small color difference.

[0279] The above steps S1-S8 illustrate the image removal and repair process when the area percentage P of the object to be removed is very small (e.g., P < 2%). The following steps will describe the image removal and repair process when the area percentage P of the object to be removed is slightly larger (e.g., 2% ≤ P < 10% or P < 25%).

[0280] S9. If P is determined to be no less than 2%, determine whether P is less than 10%.

[0281] If P is determined to be less than 10%, proceed with S10.

[0282] If P is determined to be no less than 10%, proceed to S13, which is to further determine whether P is less than 25%.

[0283] The second image processing strategy is adopted when 2% ≤ P < 10%.

[0284] S10. If P is determined to be less than 10%, calculate the area to be clipped using the second mask percentage (e.g., 10%).

[0285] The process of calculating the area to be clipped using a 10% mask ratio can refer to the process described above for calculating the area to be clipped using a 6% mask ratio.

[0286] S11. Crop the first image and the first mask image to obtain the cropped first image and the first mask image.

[0287] Both the first image and the first mask image are of the first size, and after cropping, they are of the second size.

[0288] The implementation process of S11 is similar to the implementation process of cropping the first image and the first mask image in S3 above, and will not be repeated here.

[0289] S12. Based on the cropped first image and the first mask image, determine whether the background of the object to be eliminated is a complex background.

[0290] Figure 11 This diagram illustrates how to determine whether the background of the object to be removed is a complex background. Figure 11 As shown, the first mask region is dilated, and then the difference is calculated to obtain the outer annular region. The feature point density in the outer annular region can be used to determine whether the background of the object to be removed is a complex background.

[0291] For example, the first mask region is expanded by an expansion ratio N12 to obtain dialtedMask1. Furthermore, the first mask region is expanded by an expansion ratio N13 to obtain dialtedMask2. Optionally, N12 is 15 and N13 is 75.

[0292] Then, determine the outer ring region of the first mask region: AroundMask = dialtedMask2 - dialtedMask1.

[0293] Then, the number of feature points FeaPtNum and the number of pixels PAM in the AroundMask annular region of the first image are calculated.

[0294] Then, the feature point density FeaPtDensity in the AroundMask annular region of the first image is calculated:

[0295] FeaPtDensity=FeaPtNum / PAM.

[0296] If FeaPtDensity > N14, then the background of the object to be eliminated is determined to be a complex background.

[0297] If FeaPtDensity≤N14, then the background of the object to be eliminated is determined to be a simple background.

[0298] Optionally, N14 is 1%.

[0299] Optionally, the number of feature points FeaPtNum can be calculated using the FAST (features from accelerated segment test) feature point extraction method, the ORB (oriented FAST and rotated BRIEF) feature point extraction method, or the SURF (speeded uprobust features) feature point extraction method.

[0300] refer to Figure 11 If it is determined that the background of the object to be removed is not a complex background, then proceed to steps S4-S8, which involves using a GAN image removal model that meets the actual usage requirements to perform image removal and restoration processing. If it is determined that the background of the object to be removed is a complex background in the image removal scenario, then the SD image removal model, which has stronger image processing performance, can be used.

[0301] For example, Figure 12 The flowchart illustrates the process of image removal and restoration for non-complex backgrounds when 2% ≤ P < 10%.

[0302] like Figure 12 As shown, in response to the user's selection operation on the first image (original image), the electronic device determines the selected area and calculates the area percentage P of the object to be eliminated in the first image, where 2% ≤ P < 10%.

[0303] A first mask image is obtained based on the first image. The mask region in the first mask image is the mask region of the object to be eliminated. The size of both the first image and the first mask image is the first size.

[0304] Then, based on the mask region in the first mask image, a cropped region is calculated with a mask ratio of 10%. The cropped region includes the mask region and is larger than the mask region, and the size of the cropped region is the second size. The first image and the first mask image are cropped based on the cropped region to obtain a first image of the second size and a first mask image of the second size, respectively.

[0305] Unlike Figure 10 The point is, in Figure 12 In the process, after image cropping, the first cropped image and the first mask image are used to determine whether the background of the object to be removed is a complex background.

[0306] If the background of the object to be eliminated is determined to be a complex background, the cropped first image and the first mask image are scaled, and then the mask region of the scaled first mask image is dilated. Then, the cropped and scaled first image and the cropped, scaled and dilated first mask image are input into the SD image elimination model to obtain the seventh image.

[0307] Alternatively, if it is determined that the background of the object to be eliminated is not a complex background, the cropped first image and the first mask image are scaled, and then the mask region of the scaled first mask image is dilated. Then, the cropped and scaled first image and the cropped, scaled and dilated first mask image are input into the GAN image elimination model to obtain the seventh image.

[0308] Then the seventh image is filled into the cropped area of ​​the first image to obtain the third image of the first size.

[0309] On the other hand, the first mask image, after being expanded to a third size, is filled into the cropped area of ​​the first image of the first size to obtain the first mask image after being expanded and filled to the first size. Then, the pixel values ​​of the first mask image after being expanded and filled are inverted to obtain the fifth image. Then, the fifth image is multiplied by the first image to obtain the sixth image.

[0310] The third image and the first masked image after dilation and filling are ANDed (i.e. multiplied) to obtain the fourth image.

[0311] Then, the fourth and sixth images are fused together to obtain the second image.

[0312] like Figure 12 As shown, in the second image, the portraits selected by circling or smearing are eliminated, and the background image of the eliminated area has been restored.

[0313] The solution proposed in this application allows for the use of a GAN image removal model that meets practical application requirements for image removal scenarios with non-complex backgrounds. For image removal scenarios with complex backgrounds, a SD image removal model with stronger image processing performance can be used. The image removal and restoration method provided in this application can remove images of people or objects from the selected area, and can meet various image removal needs in terms of computational load, image processing performance, and image processing effect.

[0314] Optionally, if the background of the object to be removed is determined to be a complex background, it can be determined whether the removal scenario involves a crowd. It should be noted that for complex backgrounds, it is possible to further identify whether the background is a crowd. If it is not a crowd background, an image removal model that meets practical requirements can be used. If it is a crowd background, to avoid generating distorted human images after image removal, a powerful image removal model (such as the SD image removal model) can be used based on the first image and the second mask image to avoid generating distorted human images. The specific implementation process will be explained below.

[0315] It should also be noted that, in the case of 2% ≤ P < 10%, since image cropping was performed in the preprocessing of image elimination and restoration, the cropped image filling process, as in S5 and S6, needs to be completed in the postprocessing of image elimination and restoration.

[0316] When 10% ≤ When P < 25%, the third image processing strategy is adopted.

[0317] The above describes the image removal and repair process when the area percentage P of the object to be removed is very small (e.g., P < 2%), and when the area percentage P of the object to be removed is slightly large (e.g., 2% ≤ P < 10%). The following describes the image removal and repair process when the area percentage P of the object to be removed is even larger (e.g., 10% ≤ P < 25%).

[0318] S13. If P is not less than 10%, determine whether P is less than 25%.

[0319] If P is determined to be less than 25%, proceed to S14, which determines whether the current elimination scene is a scene with a crowd background. If P is determined to be not less than 25%, proceed to S25.

[0320] S14. If P is less than 25%, determine whether the current elimination scene is an elimination scene with a crowd background.

[0321] The scenario of eliminating crowd backgrounds refers to a situation where the object to be eliminated is a portrait or object, and the background of the object to be eliminated is a crowd background (with people overlapping or obscuring each other). The specific judgment method will be explained below.

[0322] If it is determined that the current elimination scene is not an elimination scene with a crowd background, continue to execute S15-S19, or if it is determined that the current elimination scene is an elimination scene with a crowd background, continue to execute S20-S24.

[0323] The following section first explains the elimination and repair process for elimination scenarios where the current elimination scene is not a crowd background.

[0324] S15. In cases where the current elimination scene is not a scene with a crowd background, scale the first image and the first mask image to the required size of the SD image elimination model.

[0325] S16. Perform mask region dilation processing on the scaled first mask image to obtain the dilated first mask image.

[0326] S17. Input the scaled first image and the scaled and dilated first mask image into the SD image elimination model for model inference to obtain the third image.

[0327] S18. Restore the size of the dilated first mask image to the first size to obtain the dilated first mask image.

[0328] S19. Based on the first image, the third image, and the first masked image after dilation, obtain the second image.

[0329] In this embodiment of the application, the second image is obtained by the following image matrix calculation formula.

[0330] Pi = Po × (1 – Pm) + Pe × Pm.

[0331] Where Po represents the first image, Pe represents the third image, Pm represents the first masked image after dilation, and Pi represents the second image.

[0332] Where (1–Pm) represents the inversion of pixel values ​​in the first mask image after dilation.

[0333] Figure 13A This diagram illustrates the image removal and restoration method provided in this application, which involves judging the background of a large proportion of objects and implementing different image processing strategies.

[0334] If the background image of the object to be removed is not a crowd background image, then the first image and the first mask image (containing the mask region corresponding to the object to be removed) are input into the second image removal model, and then the image output by the model is post-processed to obtain the second image. The specific implementation process can be found below. Figure 13B .

[0335] If the background image of the object to be removed is a crowd background image, then the first image and the second mask image (including the mask region of each portrait in the first image) are input into the first image removal model, and then the image output by the model is post-processed to obtain the second image. The specific implementation process can be found below. Figure 13C .

[0336] Figure 13B The diagram illustrates the process of image removal and restoration for scenes where the current scene to be removed is not a crowd background, provided that 10% ≤ P < 25%.

[0337] like Figure 13B As shown, in response to the user's selection operation on the first image (original image), the electronic device determines the selected area and calculates the area percentage P of the object to be eliminated in the first image, where 10% ≤ P < 25%.

[0338] A first mask image is obtained based on the first image. The mask region in the first mask image is the mask region of the object to be eliminated. The size of both the first image and the first mask image is the first size.

[0339] Unlike Figure 10 and Figure 12 The point is, in Figure 13B In this process, neither image cropping nor image filling is performed.

[0340] Unlike Figure 10 and Figure 12 The point is also that, in Figure 13B In the process, it is necessary to determine whether the current scene is a scene with a crowd background, that is, to determine whether the object to be eliminated is a human figure and whether the background of the object to be eliminated is a crowd background.

[0341] If it is determined that the current scene is not a scene where a crowd background is eliminated, the first image and the first mask image are scaled, and then the mask region of the scaled first mask image is dilated. Then, the scaled first image and the scaled and dilated first mask image are input into the SD image elimination model to obtain the third image.

[0342] The third image and the dilated first mask image are ANDed (i.e. multiplied) to obtain the fourth image.

[0343] Invert the pixel values ​​of the first masked image after dilation to obtain the fifth image. Then multiply the fifth image by the first image to obtain the sixth image.

[0344] Then, the fourth and sixth images are fused together to obtain the second image.

[0345] like Figure 13B As shown, the first object image (such as the first portrait) in the second image is eliminated, and the background image of the eliminated area has been repaired.

[0346] The proposed solution addresses the challenges of eliminating and repairing objects with large areas requiring high performance. Therefore, employing a superior image elimination model yields better elimination and repair results.

[0347] The above explains the elimination and repair process for elimination scenarios where the current elimination scene is not a crowd background. The following explains the elimination and repair process for elimination scenarios where the current elimination scene is a crowd background.

[0348] In this embodiment, the removal is performed on a first object image (which may be a human body or an object) with a crowd background. This image removal scenario has higher requirements because removing the first object image with a crowd background may generate deformed human figures in the area where the first object image was removed. To avoid generating deformed human figures, this application adopts the following scheme to perform image removal and restoration processing.

[0349] S20. In the scenario of eliminating the background of the crowd, obtain a multi-person portrait instance mask image (i.e., the second mask image) based on the first image.

[0350] S21. Scale the first image and the multi-person portrait instance mask image to the required size (third size) of the SD image elimination model.

[0351] For example, the SD image removal model requires a size of 768*768.

[0352] S22. Input the scaled first image and the multi-person portrait instance mask image into the SD image elimination model for model inference to obtain the third image.

[0353] It should be noted that the size of the third image is the third dimension, for example, the size of the third image can be 768*768.

[0354] S23. Dilate the mask region of the first mask image to obtain the dilated first mask image.

[0355] This application does not limit the execution order of S22 and S23. For example, S22 can be executed first and then S23; or S23 can be executed first and then S22; or S22 and S23 can be executed simultaneously.

[0356] S24. Based on the first image, the third image, and the first masked image after dilation, obtain the second image.

[0357] In this embodiment of the application, the second image is obtained by the following image matrix calculation formula.

[0358] Pi = Po × (1 – Pm) + Pe × Pm.

[0359] Where Po represents the first image, Pe represents the third image, Pm represents the first masked image after dilation, and Pi represents the second image.

[0360] Where (1–Pm) represents the inversion of pixel values ​​in the first mask image after dilation.

[0361] Figure 13C This diagram illustrates the process of image removal and restoration in a scene where the current removal scene is a crowd background, provided that 10% ≤ P < 25%.

[0362] like Figure 13C As shown, in response to the user's selection operation on the first image (original image), the electronic device determines the selected area or the smeared area and determines the object to be eliminated, and calculates the area percentage P of the object to be eliminated in the first image, where 10% ≤ P < 25%.

[0363] It should be noted that the judgment logic of this application can determine that: the first image contains multiple portraits and the portraits in the selected area are occluded (i.e. closely adjacent) to other portraits. This elimination scenario is a crowd background elimination scenario, that is, to eliminate and repair a certain portrait in the crowd, which has relatively high requirements for image elimination and repair.

[0364] A first mask image is obtained based on the first image. The mask region in the first mask image is the mask region of the object to be eliminated. The size of both the first image and the first mask image is the first size.

[0365] Figure 13C Unlike Figure 13B The point is, in Figure 13C In this process, a multi-portal instance mask image (i.e., a second mask image) is obtained based on the first image. This multi-portal instance mask image includes the mask region for each portrait in the first image.

[0366] Figure 13C Unlike Figure 13B The point is also that, in Figure 13C In the process, the first image and the second mask image are scaled, and then the scaled first image and the scaled second mask image are input into the SD image elimination model to obtain the third image.

[0367] exist Figure 13C In the process, the mask region of the first mask image is dilated to obtain the dilated first mask image.

[0368] The third image and the dilated first mask image are ANDed (i.e. multiplied) to obtain the fourth image.

[0369] Invert the pixel values ​​of the first masked image after dilation to obtain the fifth image. Then multiply the fifth image by the first image to obtain the sixth image.

[0370] Then, the fourth and sixth images are fused together to obtain the second image.

[0371] like Figure 13C As shown, the human figures in the selected area of ​​the second image are removed, and the background image of the removed area is restored.

[0372] like Figure 13D As shown in the embodiments of this application, the image removal and restoration in multi-person scenes has been improved. After performing removal and restoration processing on the image area specified by the user in the photo, new portraits are avoided. Therefore, the solution of this application solves the problem of regenerating deformed portraits after removing a portrait in a multi-person scene, thus improving the image removal and restoration experience.

[0373] It should be noted that, in this embodiment of the application, the example of eliminating a certain human figure in the background of a crowd to avoid the generation of deformed human figures is used for illustrative purposes. In actual implementation, this embodiment of the application is also applicable to the scenario of eliminating a certain object (such as a trash can) in the background of a crowd. In other words, a certain object in the background of a crowd can be eliminated, and the generation of deformed human figures can also be avoided.

[0374] S25. If P is determined to be greater than or equal to 25%, prompt that the selection operation is invalid.

[0375] If the selected area is too large, to avoid poor image removal and repair results, a message can be displayed indicating that the selection operation is invalid and suggesting that the operation be repeated. This will ensure better image removal and repair results and improve the user experience.

[0376] The method provided in this application selects an appropriate image processing strategy to eliminate and repair the object to be eliminated based on the area ratio of the object to be eliminated in the original image. The image processing is more refined, which improves the elimination and repair effect.

[0377] It should be noted that the thresholds of 2%, 10%, and 25% mentioned above are illustrative examples. In actual implementation, the thresholds can be adjusted according to actual usage needs. For example, the thresholds can be adjusted to 3%, 12%, 50%, or 1%, 15%, and 75%.

[0378] In practice, the number of thresholds can be increased or decreased according to actual usage needs.

[0379] For example, a threshold can be set, such as 25%. When the area of ​​the object to be removed in the first image is less than the threshold of 25%, one image processing strategy can be used for image removal and restoration. When the area of ​​the object to be removed in the first image is greater than or equal to the threshold of 25%, another image processing strategy can be used for image removal and restoration.

[0380] As another example, two thresholds can be set, such as 10% and 25%. This reduces the computational load while ensuring the effectiveness of image removal and modification.

[0381] For example, four thresholds can be set, such as 2%, 10%, 15%, and 25%, which can make image processing more refined and improve the effect of image removal and restoration.

[0382] The image removal and repair method provided in this application can be applied to image removal scenarios with simple backgrounds or complex backgrounds.

[0383] In this application, the elimination and restoration method based on the GAN image elimination model and the elimination and restoration method based on the SD image elimination model are effectively combined, which can improve the effect and performance of elimination and restoration.

[0384] Image removal scenario to determine if it is a crowd background

[0385] The above embodiments illustrate the scenario of eliminating crowd backgrounds. The following details the possible implementation methods of image elimination scenarios for determining whether an image has a crowd background, as provided in this application.

[0386] Figure 14 An exemplary flowchart illustrating an image removal scenario for determining whether it is a crowd background, as provided in an embodiment of this application, is shown.

[0387] S401, the first image is input into a human image instance segmentation network or a panoramic segmentation network model to obtain a second mask image. This second mask image includes the masked regions of the Y human images. The second mask image can also be called a human image instance segmentation mask image.

[0388] S402, determine whether Y is greater than the threshold N15.

[0389] Optionally, N15 is 3.

[0390] If the second mask image is empty (Y=0), meaning there is no mask region or no region marked as 1, then it indicates that there is no human figure in the first image. Therefore, it can be determined that this scene does not belong to the scene where crowd backgrounds are eliminated.

[0391] If the second mask image is not empty, that is, if there is a mask region or a region marked as 1, then it means that there is at least one human figure in the first image.

[0392] In some embodiments, if the second mask image is not empty (i.e., there are human figures), and the number of human figure instances Y is less than N15, then it can be determined that the scene does not belong to the scene where the crowd background is eliminated.

[0393] In some embodiments, if the second mask image is not empty (i.e., there is a human image) and the number of human image instances is greater than or equal to N15, then S403 continues to be executed.

[0394] S403. Iterate through and calculate the number of pixels in the intersection area (denoted as Q1) and the number of pixels in the union area (denoted as Q2) between the first mask area (i.e., the mask area of ​​the object to be removed) and the i-th portrait mask area in the second mask image, as well as the intersection-union ratio (denoted as IoU). i starts from 1 and can take values ​​up to Y.

[0395] Where IoU = Q1 / Q2.

[0396] S404. Determine whether Q1 / Q2 is greater than the threshold N16.

[0397] Optionally, N16 = 0.8.

[0398] S405. If Q1 / Q2≥N16, then the first mask region is determined to be a single person region, and the traversal calculation is stopped and the mask image of the i-th person instance is recorded.

[0399] If no human face instance mask meets the conditions after traversing all the data, it is determined that the first mask area does not contain human faces, and therefore it can be determined that the scene does not belong to the scene where the crowd background is eliminated.

[0400] S406. Iterate through and calculate the minimum distance L between the i-th portrait mask region and the j-th portrait mask region in the second mask image. j starts from 1 and can take a value up to Y.

[0401] After recording the mask image of a single person's portrait instance, it is possible to determine whether the scene is a scene where the background of a crowd is to be eliminated based on the close adjacency of the recorded single person's portrait mask with the other portrait masks.

[0402] S407. Determine whether the minimum distance L is less than D, where D = W × N17.

[0403] Optionally, N17 can be 0.5.

[0404] Where W represents the width of the minimum bounding rectangle of the single-person portrait mask.

[0405] Specifically, it determines whether the minimum bounding rectangle of the recorded i-th person's image mask intersects with the minimum bounding rectangle of any other person's image mask, or whether the minimum distance between them is less than D.

[0406] S408. If the minimum bounding rectangle of the recorded i-th person's mask intersects with the minimum bounding rectangle of any other person's mask, or the minimum distance between them is less than W×N17, then it is determined that the i-th person's mask and the other person's mask are in a close adjacency state, and the traversal calculation stops. In this case, it can be determined that the scene belongs to the scene of eliminating crowd background.

[0407] S409. If no human face instance mask meets the conditions after traversal, it is determined that the scene is not tightly connected and therefore does not belong to the elimination scene with a crowd background (i.e., the elimination scene without a crowd background).

[0408] In this embodiment of the application, in the scenario of eliminating and repairing objects selected by circling or smearing in the first image, it is determined whether the current scene is an elimination scenario with a crowd background based on the first mask image and the mask image of multiple portrait instances (i.e., the second mask image). If it is determined that the current scene is an elimination scenario with a crowd background, an elimination and repair method suitable for crowd backgrounds is adopted to perform image elimination and repair, so as to avoid generating deformed portraits. This can optimize the effect of elimination and repair processing.

[0409] It should be noted that various possible methods can be used to implement the "image removal scenario of determining whether there is a crowd background," and this application does not limit this. In some embodiments, a large image-text model can also be used to determine whether there is a crowd background. For example, an image-text question-and-answer model can be used to ask whether the object to be removed is in a crowd, and the image removal scenario of determining whether there is a crowd background can be determined based on the output or answer of the image-text question-and-answer model.

[0410] Determine whether the elimination condition is met.

[0411] In some embodiments, in scenarios where objects to be eliminated are selected by circling or smearing in the original image, in order to avoid the elimination of part of the human body image due to the circling or smearing area being located on the human body, resulting in a fragmented human body image and affecting the elimination effect, the solution of this application can first determine whether the circling or smearing area meets the elimination conditions. If the circling or smearing area meets the elimination conditions, then the elimination and repair process is performed. If the circling or smearing area does not meet the elimination conditions, then the elimination and repair process is not performed. This can optimize the effect of elimination and repair processing.

[0412] For example, combined Figure 5 ,like Figure 15 As shown, after S302 and before S303, the method also includes S305 and S306.

[0413] S305. Obtain a second mask image and a third mask image based on the first image. The second mask image includes mask regions of multiple human figures, and the third mask image includes mask regions corresponding to the selected regions.

[0414] The second mask image includes the mask region corresponding to each portrait in the first image.

[0415] In some embodiments, the first image can be input into a human instance segmentation network to obtain a second mask image. The second mask image can also be called a human instance segmentation mask image (HumanInstanceMask). The human instance segmentation network is a network model used to identify human figures or human bodies in the first image and generate a mask image for each human figure.

[0416] In this process, after inputting the first image into the human image instance segmentation network model, a mask image of each human image instance in the first image can be obtained. By combining the mask images of all human image instances in the first image, a second mask image is obtained.

[0417] In other embodiments, the first image can be input into a panoramic segmentation network model to obtain a human image instance segmentation mask image. The panoramic segmentation network model is a network model used to identify people or objects in the first image and then generate a mask image for each person or object.

[0418] In this process, after inputting the first image into the panoramic segmentation network model, a mask image of each person or object in the first image can be obtained. By combining the mask images of all people in the first image, a second mask image is obtained.

[0419] Figure 16 A schematic diagram illustrating the acquisition of a second mask image based on a first image is shown. For example... Figure 16As shown, the first image includes three portraits: portrait 1, portrait 2, and portrait 3; correspondingly, the second mask image generated based on the first image includes three portrait mask regions: mask1, mask2, and mask3, which are the mask regions corresponding to the three portraits respectively.

[0420] It should be noted that if the second mask image is empty, that is, there is no mask area or no area marked as 1, then it means that there is no human image in the first image. If the second mask image is not empty, that is, there is a mask area or an area marked as 1, then it means that there is at least one human image in the first image.

[0421] S306. Based on the third mask image and the second mask image, determine whether the selected area or the painted area meets the elimination conditions.

[0422] It should be noted that the purpose of determining whether the selected or painted area meets the elimination criteria is to avoid the removal of parts of the human body image due to the selected or painted area being located on the human body, resulting in incomplete human body images and affecting the elimination effect.

[0423] Figure 17 The flowchart illustrates the process of determining whether a selected or smeared area meets the elimination criteria based on a third mask image and a second mask image. It is understood that further selection of elimination and repair methods is only permitted when the selected or smeared area is determined to meet the elimination criteria, allowing for image elimination and repair processing of the object to be eliminated. If the selected or smeared area does not meet the elimination criteria, an invalid selection operation can be indicated. This avoids the elimination of parts of the human body image due to the selected area being located within the human body, thus improving the elimination and repair effect for portraits.

[0424] In some embodiments, if the second mask image is empty, that is, there is no human image in the first image, then the object to be eliminated will not be a human body. In other words, the object to be eliminated may be an item such as a trash can or debris, which meets the elimination conditions. Therefore, elimination and repair processing can be performed on the object to be eliminated in the selected area of ​​the first image.

[0425] If the second mask image is not empty, meaning there is a human figure in the first image, then obtain the following parameters:

[0426] (1) Calculate the area S of the mask region of the third mask image (which corresponds to the selected region in the first image). The area of ​​the first mask region can be determined by counting the number of pixels in the first mask region. For example, refer to... Figure 10 As shown, the white area in the third mask image is the mask area. Calculate the number of pixels in this mask area to obtain the number of pixels S in the selected or smeared area.

[0427] (2) Calculate the area Si of each portrait mask region in the second mask image. For example, refer to... Figure 10 As shown, the second mask image includes three portrait mask regions: mask1, mask2, and mask3. The number of pixels in mask1 is denoted as S1, the number of pixels in mask2 is denoted as S2, and the number of pixels in mask3 is denoted as S3. Si includes S1, S2, and S3.

[0428] (3) Calculate the area S′ of the intersection region between the first mask region and each portrait mask region in the second mask image. The number of pixels in the intersection region between the first mask region and mask1 is denoted as S1′, the number of pixels in the intersection region between the first mask region and mask2 is denoted as S2′, and the number of pixels in the intersection region between the first mask region and mask3 is denoted as S3′. Si′ includes S1′, S2′, and S3′.

[0429] Among them, the first mask region and mask2 have an overlapping region (such as... Figure 10 (As shown by the diagonal lines in the diagram), the number of pixels in the intersection region S2′ is greater than zero. The first mask region has no intersection with either mask1 or mask3, and the number of pixels in the intersection region S2′=0 and S3′=0.

[0430] In the embodiments of this application, after calculating parameters such as S, Si (including S1, S1 and S3) and Si′ (including S1′, S2′ and S3′), it is determined whether the selected area or the smeared area meets the elimination condition in the following manner. i is taken as 1, 2, and 3 in sequence.

[0431] If Si′ / Si∈(N1,N2) and Si′ / S>N3, then the selected or painted area does not meet the elimination conditions, and the elimination process is not performed on the object to be eliminated.

[0432] Optionally, N1 can be 1%, N2 can be 50%, and N3 can be 80%. For ease of explanation, the following embodiments are illustrated with N1 as 1%, N2 as 50%, and N3 as 80%. That is, when Si′ / Si∈(1%, 50%) and Si′ / S>80%, it is determined that the selected area or the smeared area does not meet the elimination condition.

[0433] It should be noted that Si′ / S can be understood as the area ratio of the intersection region Si′ to the selected or smeared region S. If this area ratio is less than or equal to 80%, it means that part of the selected or smeared region falls within the portrait area, and the object to be removed could be a portrait or other objects. If this area ratio is greater than 80%, it means that most or all of the selected area falls within the portrait area, allowing for a more accurate identification of the object to be removed as a portrait.

[0434] Si′ / Si can be understood as the area ratio of the intersection region Si′ to the human image Si. If this area ratio is within the range of (1%, 50%), it means that the selected area or the painted area covers part of the human image. In this case, the object to be removed should not be removed to avoid accidental removal of parts of the human body, resulting in incomplete human image and affecting the image removal effect.

[0435] In other words, if Si′ / S>80% and Si′ / Si∈(1%,50%), then it can be determined that the object to be eliminated is a human body and the selected or smeared area is part of the human body. Therefore, it is considered that the elimination conditions are not met and the object to be eliminated will not be eliminated.

[0436] If neither Si′ nor Si belongs to (1%, 50%), then the elimination condition is met, and the elimination process can be performed on the object to be eliminated.

[0437] It should be noted that the calculation can be performed on each portrait mask area. If the elimination conditions are not met, the calculation can be stopped and an invalid selection operation will be displayed. If the conditions are met after traversing and calculating all portrait mask areas, the selected or painted area is determined to meet the elimination conditions, and the elimination process can be performed on the object to be eliminated. By performing traversal calculations, the amount of computation can be reduced.

[0438] The solution proposed in this application can determine whether the selected or painted area meets the elimination conditions in response to the user's selection operation on the first image. If the selected or painted area meets the elimination conditions, the object to be eliminated is eliminated. If the selected or painted area does not meet the elimination conditions, the object to be eliminated is not eliminated, so as to prevent accidental elimination, such as preventing accidental elimination of human body parts in the image.

[0439] The following illustrations, with reference to the accompanying drawings, illustrate scenarios where the selected or painted area meets the elimination criteria and scenarios where the selected or painted area does not meet the elimination criteria.

[0440] Scenes where the selected area (circled area or painted area) meets the elimination criteria.

[0441] The following is for reference. Figure 18This diagram illustrates how to determine if a selected or painted area meets the elimination criteria. For example... Figure 18 As shown, the selected area in the first image is indicated by the dashed line. The number of pixels S of the object to be eliminated is determined by the mask area in the third mask image; the number of pixels S1, S2, and S3 of each portrait is determined by the mask area of ​​each portrait in the second mask image; the number of pixels S1′, S2′, and S3′ of the intersection area of ​​the first mask area and each portrait mask area is calculated. S1′ = 0, S3′ = 0.

[0442] like Figure 18 As shown, starting with i = 1, the traversal begins. Since S1′ = 0, S1′ / S1 = 0, and S1′ / S = 0. The traversal continues. Then, i = 2, S2′ / S2 > 50%, and S2′ / S > 80%. The traversal continues. Then, i = 3, since S3′ = 0, S3′ / S3 = 0, and S3′ / S = 0. The traversal calculation ends. Si′ / Si does not belong to (1%, 50%), satisfying the elimination condition. Therefore, the elimination process can be performed on the object to be eliminated.

[0443] Scenes where the selected area (circled area or painted area) does not meet the elimination criteria.

[0444] The following is for reference. Figure 19 This diagram illustrates how to determine if a selected or painted area does not meet the elimination criteria. For example... Figure 19 As shown, the selected area in the first image is indicated by the dashed line. The number of pixels S of the object to be eliminated can be determined by the mask area in the third mask image. The number of pixels S1, S2, and S3 of each portrait can be determined by the mask area of ​​each portrait in the second mask image. The number of pixels S1′, S2′, and S3′ of the intersection area of ​​the first mask area and each portrait mask area can be calculated.

[0445] like Figure 19 As shown, the calculation starts from i=1. Since the number of pixels in portrait 1, S1′=0, therefore S1′ / S1=0, S1′ / S=0. The calculation continues. Then, i=2, S2′ / S2∈(1%, 50%) and S2′ / S>80%, indicating that the object to be eliminated is a human body and the selected area is part of the human body. Therefore, the elimination condition is not met, and the calculation stops. In cases where the elimination condition is not met, no elimination processing is performed on the object to be eliminated. This avoids the situation where part of the human body image is eliminated because the selected area is located within the human body.

[0446] Optionally, if a certain portrait mask region has an area ratio of less than N4 in all mask regions of the second mask image, meaning the portrait mask region is too small, then it can be excluded from the above traversal calculation, saving computational resources and speeding up the calculation. Optionally, N4 can be 1%. The following example illustrates this with N4 set to 1%.

[0447] For example, refer to the above Figure 18 If the electronic device determines that the area of ​​mask3 in the second mask image is less than 1%, then the electronic device does not need to perform traversal calculations on mask3.

[0448] Specifically, starting with i=1, the calculation iterates through the data. Since the number of pixels in portrait 1, S1′=0, therefore S1′ / S1=0 and S1′ / S=0, satisfying the elimination condition. Then, when i=2, S2′ / S2>50% and S2′ / S>80%, satisfying the elimination condition again. Since mask3's area is too small (less than 1%), it is not calculated. The calculation ends. If all elimination conditions are met after the calculation, the object to be eliminated can be processed.

[0449] It should be noted that various possible methods can be used to "determine whether the selected area or the painted area meets the elimination conditions," and this application does not limit this. For example, embodiments of this application also provide the following implementation method to determine whether the selected area or the painted area meets the elimination conditions.

[0450] For example, after calculating parameters such as the number of pixels S of the object to be eliminated, the number of pixels Si of each portrait (including S1, S2, and S3), and the number of pixels Si′ of the intersection region (including S1′, S2′, and S3′), the selected area or the painted area is judged to meet the elimination conditions in the following manner. i is taken as 1, 2, and 3 in sequence.

[0451] In some embodiments, if the ratio of the number of pixels Si′ in the intersection region to the number of pixels Si in the portrait region is within the range of (N1, N2) and the BBOX ratio is greater than N5, then the selected region or the smeared region is determined not to meet the elimination conditions. This situation usually occurs when the selected region or the smeared region contains a human body and would cause an elimination effect on the human body area, therefore, the elimination process is not performed on the object to be eliminated.

[0452] BBOX is the width of the bounding rectangle of the intersection area of ​​portrait i, divided by the width of the bounding rectangle of the mask of portrait i.

[0453] The intersection area of ​​portrait i refers to the intersection area between the mask area of ​​portrait i and the first mask area.

[0454] Let Y(i) denote the BBOX ratio. Y(i) = BBOX(i′).width / BBOX(i).width.

[0455] Here, BBOX represents the outer rectangle. BBOX(i).width is the width of the outer rectangle of the mask for each portrait. BBOX(i′).width represents the width of the outer rectangle of the intersection area of ​​the portraits.

[0456] Figure 20 A diagram illustrating the percentage of BBOX is shown, such as... Figure 20 As shown, it can be determined that mask2 and the first mask region have an intersection area, and the proportion of the intersection area is Y(2)=BBOX(2′).width / BBOX(2).width.

[0457] Optionally, N5 can be 30%.

[0458] It's understandable that when the BBOX percentage is greater than 30%, it means that more than 30% of the selected or smeared area falls within the area containing the human image. If the number of pixels in the intersection area Si′ / the number of pixels in the human image area Si is within the range of (1%, 50%), then it can be determined that the object to be removed is a human body and the selected area is part of the human body. Therefore, it can be determined that the selected or smeared area does not meet the removal conditions. This situation usually occurs when the selected or smeared area contains a human body and will cause a removal effect on the human body area; therefore, removal processing is not performed on the object to be removed.

[0459] In other embodiments, if the elimination conditions are met after the traversal calculation is completed, this usually means that the selected area or the smeared area does not contain a human body (i.e., the object to be eliminated is not a human body), or the selected area or the smeared area contains a human body but will not have an elimination effect on the human body area. Therefore, the object to be eliminated can be eliminated.

[0460] Figure 21 A comparative diagram is shown to determine whether the selected area meets the elimination criteria.

[0461] like Figure 21 As shown in (a), the selected region 21 surrounds the human body image. Through the above judgment process, it can be determined that the selected region meets the elimination conditions, so elimination and repair processing is performed.

[0462] like Figure 21 As shown in (b) above, the dashed box 22 is the selected area in the first image. The selected area does not completely surround the human figure, but only covers a part of the human body. Through the above judgment process, it can be determined that the selected area does not meet the elimination condition. In this case, the electronic device can prompt that the selection operation is invalid.

[0463] In this embodiment of the application, if the selected area (e.g., the circled area or the painted area) does not meet the elimination conditions, the electronic device may display a prompt message, such as "This photo does not support elimination".

[0464] For example, such as Figure 22A As shown in (a) and (b) of the image, in the scenario of AI elimination through intelligent selection, in response to the user's selection operation in the first image, the electronic device displays the selection trajectory 51 in the first image. Through the solution of this application, the electronic device determines that the selected area corresponding to the selection trajectory 51 is a human body area, which does not meet the elimination conditions, as shown in (a) and (b). Figure 22A As shown in (c), the electronic device displays the message 52 "Prompt: This photo does not support deletion".

[0465] For example, such as Figure 22B As shown in (a) and (b) of the image, in a scenario where AI removal is performed by manual smearing, in response to the user's manual smearing operation in the first image, the electronic device displays smear mark 53 in the first image. Using the solution of this application, the electronic device determines that the area corresponding to smear mark 53 is a human body area and does not meet the removal conditions, as shown in (a) and (b). Figure 22B As shown in (c), the electronic device displays the message 54 "Tip: This photo does not support deletion".

[0466] The proposed solution responds to a user's selection or smearing operation on the first image, and if the selected or smeared area meets the elimination conditions, image elimination and repair are then performed.

[0467] It should be noted that in the embodiments of this application, "greater than" can be replaced with "greater than or equal to", "less than or equal to" can be replaced with "less than", or "greater than or equal to" can be replaced with "greater than", and "less than" can be replaced with "less than or equal to".

[0468] The various embodiments described herein can be independent solutions or combinations thereof based on their inherent logic, and all such solutions fall within the protection scope of this application.

[0469] The foregoing mainly describes the solutions provided by the embodiments of this application from the perspective of method steps. It is understood that, in order to achieve the above functions, the electronic device implementing this method includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should recognize that, in conjunction with the units and algorithm steps of the various examples described in the embodiments disclosed herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed by hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of protection of this application.

[0470] This application embodiment can divide an electronic device into functional modules based on the above method example. For example, each function can be divided into its own functional modules, or two or more functions can be integrated into one processing module. The integrated modules can be implemented in hardware or as software functional modules. It should be noted that the module division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, other feasible division methods may exist. The following description uses the division of functional modules according to each function as an example.

[0471] Figure 23 This is a schematic block diagram of an image removal and restoration apparatus 500 provided in an embodiment of this application. The apparatus 500 can be used to perform the actions performed by the electronic device in the above method embodiments. The apparatus 500 includes an image processing unit 510 and an image display unit 520.

[0472] The image processing unit 510 is configured to enable the image removal function in response to a first operation by the user; and to highlight the selected area in the first image in response to a second operation by the user on the first image.

[0473] The image processing unit 510 is also used to identify whether the object to be eliminated in the selected area is a human image; and to identify whether the background image of the object to be eliminated is a crowd background image.

[0474] The image processing unit 510 is further configured to perform image elimination and repair processing on the object to be eliminated based on the first image, the first mask image, and the second mask image when the object to be eliminated is a human portrait and the background image of the object to be eliminated is a crowd background image, to obtain a second image; wherein, the mask area in the first mask image is the mask area corresponding to the object to be eliminated, and the second mask image includes the mask area of ​​each human portrait in the first image.

[0475] The image display unit 520 is used to display a second image; wherein the selected area of ​​the second image does not include a human figure.

[0476] The image processing unit 510 is specifically used for: performing image processing based on the first image and the second mask image to obtain a third image; wherein the third image does not include a human figure; and performing image processing based on the first image, the third image and the first mask image to obtain a second image.

[0477] The image processing unit 510 is further configured to: input the first image and the second mask image into the first image elimination model to obtain the third image; wherein the first image elimination model is used to eliminate the image of the mask region and perform background restoration on the eliminated region.

[0478] The first image removal model can be an SD-based image removal model.

[0479] The image processing apparatus provided in this application embodiment, in response to a user's operation of selecting a region on a first image, can identify the object to be eliminated in the selected region and identify the background image of the object to be eliminated, so as to adopt an appropriate image erasure and repair processing strategy and obtain a better image erasure and repair effect. When the object to be eliminated is a human portrait and the background image of the object to be eliminated is a crowd background image, the object to be eliminated can be image-eliminating and background-repaired based on the first image, the first mask image (including the mask area corresponding to the object to be eliminated), and the second mask image (including the mask area of ​​each human portrait in the first image), to obtain a second image, where the selected region of the second image does not include human portraits. With this application solution, in practical use, even if there are multiple human portraits in a photo, it is possible to accurately eliminate a human portrait after selecting it, and perform background repair in the eliminated area without generating images of other people or objects, such as deformed human portraits, thereby improving the elimination and repair effect.

[0480] The apparatus 500 according to the embodiments of this application can correspond to the execution of the method described in the embodiments of this application, and the above and other operations and / or functions of the units in the apparatus 500 are respectively for implementing the corresponding process of the method, which will not be repeated here for the sake of brevity.

[0481] This application also provides a chip coupled to a memory, which is used to read and execute computer programs or instructions stored in the memory to perform the methods in the above embodiments.

[0482] This application also provides an electronic device including a chip for reading and executing computer programs or instructions stored in a memory, causing the methods in the various embodiments to be performed.

[0483] This embodiment also provides a computer-readable storage medium storing computer instructions. When the computer instructions are executed on an electronic device, the electronic device performs the aforementioned method steps to implement the image removal and repair method described in the above embodiment.

[0484] This embodiment also provides a computer program product. The computer-readable storage medium stores program code. When the computer program product is run on a computer, it causes the computer to perform the above-mentioned related steps to realize the image elimination and repair method in the above embodiment.

[0485] In this embodiment, the electronic device, computer-readable storage medium, computer program product or chip are all used to execute the corresponding methods provided above. Therefore, the beneficial effects that can be achieved can be referred to the beneficial effects of the corresponding methods provided above, and will not be repeated here.

[0486] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another apparatus, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0487] In this article, the term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. The symbol " / " in this article indicates that the related objects are in an "or" relationship; for example, A / B means A or B.

[0488] The terms "first" and "second," etc., used in the specification and claims herein are used to distinguish different objects, not to describe a specific order of objects. In the description of embodiments in this application, unless otherwise stated, "multiple" means two or more; for example, multiple processing units refer to two or more processing units, etc.; multiple elements refer to two or more elements, etc.

[0489] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. An image removal and restoration method, characterized in that, include: The image removal function is enabled in response to the user's first action; In response to a second user action on the first image, the selected area is highlighted in the first image; Identify whether the object to be eliminated in the selected area is a human figure; Identify whether the background image of the object to be eliminated is a crowd background image; When the object to be eliminated is a human portrait and the background image of the object to be eliminated is a crowd background image, the object to be eliminated is processed by image elimination and repair according to the first image, the first mask image and the second mask image to obtain the second image; wherein, the mask area in the first mask image is the mask area corresponding to the object to be eliminated, and the second mask image includes the mask area of ​​each human portrait in the first image. The second image is displayed; wherein the selected area of ​​the second image does not include human figures; The step of performing image removal and repair processing on the object to be removed based on the first image, the first mask image, and the second mask image to obtain the second image includes: Image processing is performed based on the first image and the second mask image to obtain a third image; wherein, the third image does not include a human figure; The third image and the first mask image are ANDed to obtain the fourth image; The first image and the fifth image are ANDed to obtain the sixth image; the fifth image is obtained by inverting the pixel values ​​of the first mask image. The fourth image and the sixth image are fused together to obtain the second image.

2. The method according to claim 1, characterized in that, The step of performing a bitwise AND operation on the third image and the first mask image includes: performing a bitwise AND operation on the third image and the dilated first mask image; The fifth image is obtained by inverting the pixel values ​​of the first mask image after it has been dilated.

3. The method according to claim 1, characterized in that, The step of performing image processing based on the first image and the second mask image to obtain the third image includes: The first image and the second mask image are input into the first image elimination model to obtain the third image; The first image removal model is used to remove the image from the masked area and perform background restoration on the removed area.

4. The method according to claim 3, characterized in that, The first image removal model is an image removal model based on stable diffusion (SD).

5. The method according to claim 1, characterized in that, The method further includes: The first image is input into the first image segmentation model to obtain the second mask image; The first image segmentation model is either a portrait instance segmentation network model or a panoramic segmentation network model.

6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: If the object to be eliminated is determined to be a human image, the elimination operation is determined to meet the elimination conditions based on the second mask image and the third mask image; wherein, the mask area of ​​the third mask image is the mask area corresponding to the selected area; The step of identifying whether the background image of the object to be eliminated is a crowd background image includes: If the elimination operation meets the elimination conditions, identify whether the background image of the object to be eliminated is a crowd background image.

7. The method according to any one of claims 1 to 5, characterized in that, Before identifying whether the background image of the object to be eliminated is a crowd background image, the method further includes: The first area ratio is determined based on the ratio of the number of pixels in the masked region of the first mask image to the total number of pixels in the first mask image. The step of identifying whether the background image of the object to be eliminated is a crowd background image includes: If the first area ratio is greater than or equal to the first threshold and less than the second threshold, identify whether the background image of the object to be eliminated is a crowd background image.

8. The method according to any one of claims 1 to 5, characterized in that, The step of identifying whether the background image of the object to be eliminated is a crowd background image includes: If the number of human figures in the second mask image is less than the third threshold, then it is determined that the background image of the object to be eliminated is not a crowd background image. If the number of human figures in the second mask image is greater than or equal to the third threshold, then the first human figure mask region in the second mask image is determined based on the first mask image and the second mask image. If the first portrait mask region intersects with at least one other portrait mask region in the second mask image or the distance is less than the fourth threshold, then the background image of the object to be eliminated is determined to be a crowd background image. If the minimum distance between the first portrait mask region and each portrait mask region in the second mask image is greater than or equal to the fourth threshold, then it is determined that the background image of the object to be eliminated is not a crowd background image.

9. The method according to any one of claims 1 to 5, characterized in that, The method further includes: If it is determined that the background image of the object to be eliminated is not a crowd background image, the object to be eliminated is subjected to image elimination and repair processing based on the first image, the first mask image and the second image elimination model to obtain a repaired image. The second image removal model is used to remove the image from the masked area and perform background restoration on the removed area.

10. The method according to claim 9, characterized in that, The step of performing image removal and restoration processing on the object to be removed based on the first image, the first mask image, and the second image removal model to obtain a restored image includes: A first cropped image including the object to be eliminated is obtained by cropping from the first image; A second cropped image including the first mask region is obtained by cropping from the first mask image, and the second cropped image is subjected to mask dilation processing; The first cropped image and the second cropped image after dilation are input into the second image elimination model to obtain the model-generated image. The model-generated image is filled into the cropped area of ​​the first image to obtain the repaired image.

11. The method according to claim 9, characterized in that, The second image removal model is either a generative adversarial network (GAN) based image removal model or an SD-based image removal model.

12. The method according to any one of claims 1 to 5, characterized in that, The step of highlighting a selected area in the first image in response to a second user action on the first image includes: In response to a user's selection operation on the first image, the closed region formed by the movement trajectory of the selection operation is determined as the selected region; or, In response to the user's smearing action on the first image, the smeared area is determined as the selected area.

13. An electronic device, characterized in that, The electronic device includes: one or more processors, and memory; The memory is coupled to the one or more processors, the memory being used to store computer program code, the computer program code including computer instructions, the one or more processors invoking the computer instructions to cause the electronic device to perform the method as described in any one of claims 1 to 12.

14. A chip system, characterized in that, The chip system is applied to an electronic device, the chip system including one or more processors, the one or more processors being used to invoke computer instructions to cause the electronic device to perform the method as described in any one of claims 1 to 12.

15. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes instructions that, when executed on an electronic device, cause the electronic device to perform the method as described in any one of claims 1 to 12.

Citation Information

Patent Citations

  • Image processing method and device and computer equipment

    CN117011417A