Image processing method, device and product

By identifying the region of the object to be deleted in the image and using a pre-trained model to process and fuse high-frequency features, the problem of low efficiency in manual image editing in existing technologies is solved, achieving automated and efficient image processing results.

CN121767495APending Publication Date: 2026-03-31BEIJING DUYOU INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-22
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

Existing image editing methods rely on specialized software to manually process objects to be deleted from images, which is inefficient and complex, making it difficult to meet the rapidly growing demand for image editing.

Method used

By identifying the region of the object to be deleted in the original image, a pre-trained image processing model is used to delete the object and fuse the high-frequency component image of high-frequency features to generate a processed image, thereby improving the image processing effect.

Benefits of technology

It achieves automated image processing, improving the efficiency and effectiveness of image editing, especially when deleting facial hair, stubble, and other unwanted features from face images, resulting in more natural and detailed images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121767495A_ABST
    Figure CN121767495A_ABST
Patent Text Reader

Abstract

The invention provides an image processing method and device, electronic equipment, a storage medium and a computer program product, relates to the technical field of computers, in particular to the technical fields of computer vision, artificial intelligence, image editing and the like, and can be applied to an image processing scene. According to the specific implementation scheme, the method comprises the steps of determining a region image representing a to-be-deleted object in an original image; deleting the to-be-deleted object in the area image to obtain an object deletion image; fusing the object deletion image and the high-frequency component image representing the high-frequency features in the region image to obtain a fused image; and generating a processed image based on the original image and the fused image. On the basis of deleting the to-be-deleted object in the original image, the image processing effect is improved by fusing the high-frequency component images representing the high-frequency features.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer technology, specifically to the fields of computer vision, artificial intelligence, and image editing, and particularly to an image processing method, apparatus, electronic device, storage medium, and computer program product that can be applied to image processing scenarios. Background Technology

[0002] With the rapid development of digital content creation, social entertainment, and online office work, the demand for image editing is increasing, such as removing stubble and gray hairs from facial images. To achieve this, traditional image editing methods rely on professional image editing software such as Photoshop and GIMP (GNU Image Manipulation Program) for manual image processing. Summary of the Invention

[0003] This disclosure provides an image processing method, apparatus, electronic device, storage medium, and computer program product.

[0004] According to the first aspect, an image processing method is provided, comprising: determining a region image representing an object to be deleted in an original image; deleting the object to be deleted from the region image to obtain an object deletion image; fusing the object deletion image and a high-frequency component image of high-frequency features in the region image to obtain a fused image; and generating a processed image based on the original image and the fused image.

[0005] According to a second aspect, an image processing apparatus is provided, comprising: an image determining unit configured to determine a region image representing an object to be deleted in an original image; an object deleting unit configured to delete the object to be deleted in the region image to obtain an object-deleted image; an image fusion unit configured to fuse the object-deleted image and a high-frequency component image representing high-frequency features in the region image to obtain a fused image; and an image generating unit configured to generate a processed image based on the original image and the fused image.

[0006] According to a third aspect, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform a method as described in any implementation of the first aspect.

[0007] According to a fourth aspect, a non-transitory computer-readable storage medium is provided that stores computer instructions for causing a computer to perform the method described in any implementation of the first aspect.

[0008] According to a fifth aspect, a computer program product is provided, comprising: a computer program that, when executed by a processor, implements the method as described in any implementation of the first aspect.

[0009] According to the technology disclosed herein, an image processing method and apparatus are provided, which involves determining a region image representing an object to be deleted in an original image; deleting the object to be deleted from the region image to obtain an object deletion image; fusing the object deletion image and a high-frequency component image representing high-frequency features in the region image to obtain a fused image; and generating a processed image based on the original image and the fused image. By deleting the object to be deleted from the original image and fusing the high-frequency component image representing high-frequency features, the image processing effect is improved.

[0010] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description

[0011] The accompanying drawings are provided to better understand this solution and do not constitute a limitation of this disclosure. Wherein: Figure 1 This is an exemplary system architecture diagram that can be applied to an embodiment of this disclosure; Figure 2 This is a flowchart of an embodiment of the image processing method according to the present disclosure; Figure 3 This is a schematic diagram illustrating an application scenario of the image processing method according to this embodiment; Figure 4 This is a flowchart of yet another embodiment of the image processing method according to the present disclosure; Figure 5 This is a structural diagram of an embodiment of the image processing apparatus according to the present disclosure; Figure 6 This is a schematic diagram of the structure of a computer system suitable for implementing the embodiments of the present disclosure. Detailed Implementation

[0012] The exemplary embodiments of this disclosure are described below with reference to the accompanying drawings, including various details of the embodiments to aid understanding, and should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.

[0013] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in the technical solution disclosed herein comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0014] Figure 1 An exemplary architecture 100 is shown that the image processing methods and apparatus of this disclosure can be applied.

[0015] like Figure 1 As shown, the system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. The communication connections between terminal devices 101, 102, and 103 form a network topology. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links or fiber optic cables, etc.

[0016] Terminal devices 101, 102, and 103 can be hardware or software that supports network connectivity for data interaction and processing. When terminal devices 101, 102, and 103 are hardware, they can be various electronic devices that support network connectivity, information acquisition, interaction, display, and processing functions, including but not limited to smartphones, tablets, e-book readers, laptops, and desktop computers. When terminal devices 101, 102, and 103 are software, they can be installed in the aforementioned electronic devices. They can be implemented as, for example, multiple software programs or software modules to provide distributed services, or as a single software program or software module. No specific limitations are imposed here.

[0017] Server 105 can be a server that provides various services, such as a background processing server that retrieves original images uploaded by users through terminal devices 101, 102, and 103 and deletes objects to be deleted from the original images. As an example, server 105 can be a cloud server.

[0018] It should be noted that a server can be either hardware or software. When the server is hardware, it can be implemented as a distributed server cluster consisting of multiple servers, or as a single server. When the server is software, it can be implemented as multiple software programs or software modules (such as software programs or software modules used to provide distributed services), or as a single software program or software module. No specific limitations are made here.

[0019] It should also be noted that the image processing method provided in the embodiments of this disclosure is generally executed by a server, but the possibility of it being executed by a terminal device, or by the server and the terminal device cooperating with each other, is not excluded. Accordingly, the various parts (e.g., various units) included in the image processing apparatus can be all located in the server, all located in the terminal device, or separately located in the server and the terminal device.

[0020] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Depending on implementation needs, any number of terminal devices, networks, and servers can be included. When the electronic devices on which the image processing method runs do not require data transmission with other electronic devices, the system architecture may consist only of the electronic devices on which the image processing method runs (e.g., servers or terminal devices).

[0021] Please refer to Figure 2 , Figure 2 A flowchart of an image processing method provided in this disclosure embodiment. Flowchart 200 includes the following steps: Step 201: Determine the region in the original image that represents the object to be deleted.

[0022] In this embodiment, the entity executing the image processing method (e.g., Figure 1 The server can acquire the full image remotely or locally via a wired or wireless network connection and determine the region of the original image that represents the object to be deleted.

[0023] First, identify the object to be deleted in the original image and determine the area of ​​the object in the original image; then, divide the original image into regions according to the determined area.

[0024] The objects to be deleted in the original image are generally flaws or redundant elements in the original image, such as flaws on body parts like beards, spots, and scars, or flaws on attachments like stains and wrinkles.

[0025] Depending on the application scenario and the content of the original image, targeted object recognition can be performed on the original image. In a scenario where object deletion is targeted at a specific type, only the objects to be deleted under the target type in the original image need to be identified; in a general object deletion scenario, all objects to be deleted in the original image need to be identified.

[0026] As an example, based on the semantic segmentation model, the input is the original image and the category description data of the object to be deleted (such as facial spots, fine hairs). The image processing model calls the pre-trained object feature library (including texture, shape, and color distribution features) and identifies the region range of the object to be deleted in the original image through pixel-level feature matching. Then, the region image is segmented from the original image according to the region range.

[0027] As another example, an edge detection and color clustering fusion method is adopted. First, the contours of pixel changes (such as the boundary between the object and the background) in the original image are extracted by the edge detection algorithm. Then, color clustering is performed on the original image to filter out the color range corresponding to the object to be deleted. Then, the intersection operation of the edge contour and the target color range is performed to remove the background interference area. The final continuous or discrete region is the range of the object to be deleted, which is suitable for the location of objects with distinct color features.

[0028] In a specific example, the objects to be deleted are the mustache, stubble, and beard in a face image (original image). Based on facial hair features and visual representation, the following definitions are given for mustache, stubble, and beard: Hu Qing specifically refers to the bluish pigmentation or hue in the area where facial hair grows. It is a color feature without obvious physical hair and belongs to the low-frequency component in the image frequency domain.

[0029] Stubble refers to the hair growing in the beard area of ​​a person's face. It has a discrete and fine structure and belongs to the high-frequency components in the image frequency domain.

[0030] Beard, as a general term for facial hair, encompasses the hair and related visual features of areas such as the lips, chin, and cheeks. It includes both the bluish-green base color corresponding to beard hair and the short, fine hairs like stubble, representing a combination of both.

[0031] In some optional implementations of this embodiment, the execution entity can perform step 201 as follows: The first step is to determine the location image representing the part of the object to be deleted from the original image.

[0032] Continuing with the example of the object to be deleted being facial hair, stubble, or beard, the object to be deleted is located on the face. The original image is input into a part detection model (e.g., a face detection model) to determine the region of the object to be deleted within the original image, such as the upper left corner coordinates (x1, y1) and the lower right corner coordinates (x2, y2). The part image is then segmented from the original image according to the region range.

[0033] The part detection model can be a single-stage or two-stage detection model, such as the YOLO (You Only Look Once) model.

[0034] The second step is to determine the region in the original image that represents the object to be deleted, based on the key points in the part image.

[0035] Identify key points in the image of the part; based on the key points corresponding to the object to be deleted, determine the region image in the original image that represents the object to be deleted.

[0036] Continuing with the example of a face image, the original image is cropped according to the coordinates of the top-left corner (x1, y1) and bottom-right corner (x2, y2) of the face detection bounding box to obtain the face image. The face image is then input into a facial landmark detection model, which generates multiple facial landmarks and outputs the coordinates of each landmark. Landmarks near the beard region are selected from these landmarks, and the top-left corner (x3, y3) and bottom-right corner (x4, y4) of the beard region are determined accordingly to crop the region image.

[0037] In this implementation, a two-stage determination method based on part detection and key point detection is used to determine the region image, which improves the accuracy of the region image.

[0038] Step 202: Delete the objects to be deleted from the region image to obtain the object deletion image.

[0039] In this embodiment, the aforementioned execution entity can delete the object to be deleted in the region image to obtain the object deletion image.

[0040] As an example, firstly, the pixel values ​​of the pixels within the object region of the object to be deleted in the image are set to 0, for example; then, normal textures (such as skin textures and background textures) within a specified pixel range (e.g., 10-20) around the object region of the object to be deleted are extracted, and texture blocks with the highest matching degree to the edge of the area to be deleted are selected by texture similarity calculation; finally, the selected texture blocks are filled into the area to be deleted according to the pixel arrangement pattern, and then Gaussian smoothing is performed on the boundary between the filled area and the original image to eliminate edge color difference and discontinuity, and finally an image without the object to be deleted and with coherent texture is generated, which is suitable for deleting objects with uniform texture such as spots and small stains.

[0041] As another example, firstly, the region image and the region mask of the object to be deleted are input into a pre-trained image inpainting model. Based on the image context information outside the mask (such as color distribution and lighting angle), the model learns to generate filling content that matches the style of the surrounding environment. During the generation process, a constraint loss function is used to ensure that the details of the filling region (such as skin pores and background texture) match the original image, thus completing the removal of the object to be deleted. This method is suitable for scenarios where complex details need to be restored, such as hair and small occlusions.

[0042] In some optional implementations of this embodiment, the execution entity can perform step 202 as follows: by using a pre-trained image processing model, delete the object to be deleted in the region image and fill the object region of the object to be deleted in the region image to obtain the object deletion image.

[0043] Pre-trained image processing models include, for example, neural network models with image editing capabilities and large-scale artificial intelligence models. Large-scale artificial intelligence models (or simply large models) refer to a class of artificial intelligence models with a large number of parameters constructed from artificial neural networks, such as large language models, large vision models, multimodal large models, and large basic science models. This embodiment can specifically employ a multimodal large language model.

[0044] A multimodal large language model can be fine-tuned in the following ways: First, a training sample set is obtained, which includes images containing and without the objects to be deleted. For the training image dataset containing the objects to be deleted (such as spots, hair, stains, etc.), multi-dimensional data augmentation operations are performed: random horizontal flipping adapts to different left-right orientations of the objects to be deleted in the image; fine-tuning brightness and contrast simulates the visual appearance of the objects to be deleted under different lighting conditions; and slight Gaussian blur enhances the model's ability to recognize the objects to be deleted in slightly blurred scenes. These operations enrich the data diversity and improve the model's generalization ability to the objects to be deleted in different scenarios from the source.

[0045] Subsequently, based on the Flux.Kontext image editing pre-trained model, a parameter-efficient LoRA (Low-Rank Adaptation) fine-tuning technique was employed to train the removal function targeting the objects to be deleted. This technique does not require a full update of the original model's massive backbone parameters; it only inserts a small number of trainable low-rank matrices into the key layers of the model, significantly reducing training resource consumption while accurately focusing on feature learning of the objects to be deleted.

[0046] During training, forward diffusion noise is first added to the image containing the object to be deleted: random noise is gradually added to the image according to the preset diffusion steps to generate noisy images with different noise intensities, simulating the various states of the object to be deleted being disturbed by noise in real scenes; then, the noisy image is input into the model equipped with the LoRA adapter, and the model outputs the prediction result of adding noise to the image based on the learned features of the object to be deleted.

[0047] Then, the difference between the predicted noise output by the model and the actual noise added to the image is calculated using the MSE (mean squared error) loss function. This loss value is used as a supervision signal to backpropagate to the model: only the parameters of the LoRA low-rank matrix are updated, and the backbone parameters of the original model are frozen to retain its general image editing capabilities, avoiding overfitting or loss of basic performance due to full training.

[0048] Finally, once the training loss has stabilized and converged, a LoRA weight file adapted to the Flux model ecosystem and tailored to the target objects to be deleted is saved. This weight file can then be directly loaded to quickly remove similar objects from the model.

[0049] Continuing with the example of objects to be deleted being stubble, beard, and mustache, you can fine-tune the large model for each of the three objects, or you can fine-tune the large model for only some of the objects (such as stubble and mustache). After fine-tuning, the large model can be adapted to the objects to be deleted, including stubble, beard, and mustache.

[0050] In this implementation, a pre-trained image processing model is used to delete the object to be deleted from the region image and fill the object region of the object to be deleted in the region image to obtain the object deletion image, which improves the generation efficiency and generation effect of the object deletion image.

[0051] Step 203: The high-frequency component images of high-frequency features in the image of the object to be fused and the image representing the region are fused to obtain the fused image.

[0052] In this embodiment, the aforementioned execution entity can fuse the object deletion image and the high-frequency component image of the high-frequency features in the characterizing region image to obtain a fused image.

[0053] High-frequency features characterize components in an image where pixel grayscale or color changes abruptly; these are typically texture or edge features. For example, a Gaussian high-pass filter is first used to process the region image. This filter, through preset filtering parameters, directly suppresses low-frequency components with gradual grayscale / color changes, allowing only high-frequency components with abrupt pixel value changes to pass through. During processing, the filter performs weighted calculations on the neighborhood of each pixel, strengthening areas with significant differences between pixels. The final output is a high-frequency component image that retains only high-frequency features, suitable for scenarios requiring rapid separation of high and low frequencies while preserving the complete high-frequency region.

[0054] For example, the Laplacian operator can be applied directly. This operator calculates the grayscale / color difference between each pixel in the region image and its neighboring pixels, highlighting areas with abrupt pixel value changes while suppressing areas with gradual changes. After the operation, the numerical range of the result is adjusted to eliminate negative numerical interference, ultimately generating a high-frequency component image containing only high-frequency features, suitable for scenarios requiring precise capture of local pixel abrupt changes.

[0055] After identifying the high-frequency component images representing the high-frequency features in the region image, the object deletion image and the high-frequency component images can be fused to obtain a fused image. For example, first, the region in the object deletion image where the object to be deleted once existed is located. For this deletion region, the pixel values ​​of the corresponding region in the high-frequency component image are superimposed with the pixel values ​​of the corresponding region in the object deletion image at a higher proportion (e.g., 70%-80%); for the non-deletion region, the proportion of the high-frequency component is reduced (e.g., 20%-30%) to avoid interfering with the details of the original image. After superposition, the overall image is normalized for brightness to eliminate pixel value overflow, ultimately generating a fused image that retains high-frequency features and makes the deleted region appear natural.

[0056] For example, firstly, both the object deletion image and the high-frequency component image are decomposed into a base layer and a detail layer. The base layer preserves the overall tone of the image, while the detail layer corresponds to subtle changes in the image. In the base layer, the base layer of the object deletion image is used as the primary layer; in the detail layer, the detail layer of the high-frequency component image completely replaces the detail layer of the object deletion image, and the boundary areas are smoothed by pixel transitions to eliminate layering marks. Finally, the fused base layer and detail layer are reconstructed to obtain a fused image that incorporates high-frequency features.

[0057] In some optional implementations of this embodiment, the execution entity can perform step 203 as follows: The first step is to remove low-frequency features from the region image to obtain the high-frequency component image.

[0058] For example, the Sobel gradient operator is used to calculate the gray-level difference between adjacent pixels in the horizontal and vertical directions of the image. By superimposing the difference results in the two directions, areas with abrupt changes in pixel values ​​are enhanced, while low-frequency components with gradual gray-level changes are weakened, directly obtaining a high-frequency component image that retains only high-frequency features, which is suitable for scenarios where pixel abrupt changes need to be highlighted.

[0059] For example, the region image is divided into small image blocks, and the pixel gray-level variance of each block is calculated. For regions with variances less than a threshold (determined to be low-frequency blocks), the mean gray-level value within the block is subtracted from the original block pixel value; regions with variances meeting the threshold retain their original pixel values. Finally, all blocks are integrated to generate a high-frequency component image, thus achieving low-frequency feature removal.

[0060] The second step involves removing the image and high-frequency component image from the fusion object to obtain the fused image.

[0061] In this implementation, image fusion can be performed using the above-described method of deleting images and high-frequency component images from the fusion object, which will not be elaborated upon here.

[0062] In this implementation, a high-frequency component image is obtained by deleting low-frequency features from the regional image. This can filter out redundant information such as smooth background colors and accurately retain detailed features. Then, by fusing the deleted object image and the high-frequency component image, the details of the filled area after the object is deleted can be filled, avoiding harsh blurring at the repaired area and making the fused image richer in detail and more realistic.

[0063] In some optional implementations of this embodiment, the execution entity can perform the first step as follows: First, a Gaussian low-pass filter is used to filter high-frequency features in the region image to obtain a low-frequency component image.

[0064] Set the parameters of the Gaussian low-pass filter: choose an odd-sized filter kernel (e.g., 5×5) to ensure symmetrical pixel calculation, and adjust the standard deviation according to the detail density of the region image (a slightly larger standard deviation is needed for less detail). Then, use this filter to process the image pixel by pixel. Each pixel value is a weighted sum of its neighboring pixels according to Gaussian weights, with smaller weights for neighbors farther from the center pixel. This process smooths out the high-frequency components where pixel values ​​change abruptly, ultimately outputting a low-frequency component image that preserves the overall tone and general outline.

[0065] Then, based on the pixel-level difference between the region image and the low-frequency component image, the high-frequency component image is generated.

[0066] The region image and the low-frequency component image are identical in properties such as resolution and size. For each pixel in the region image, the pixel value of that pixel in the region image is subtracted from the pixel value of the corresponding pixel in the low-frequency component image to determine the pixel value of the corresponding pixel in the high-frequency component image, thus generating the high-frequency component image.

[0067] In this implementation, Gaussian low-pass filtering can smoothly filter high frequencies and accurately preserve low-frequency features such as the overall tone and contour of the image; then, through pixel-level interpolation, high-frequency details of the original image can be completely extracted, and the separation of high and low frequencies is thorough without loss of additional information, providing clean high-frequency and low-frequency materials for subsequent processing.

[0068] In some optional implementations of this embodiment, the execution entity can perform the pixel-level interpolation operation to generate a high-frequency component image in the following manner: First, in response to the fact that the resolution of the region image and the low-frequency component image are the same, multiple color channels of the region image and multiple color channels of the low-frequency component image with the same color attribute are combined to obtain multiple channel pairs.

[0069] The aforementioned execution entity needs to verify the consistency between the region image and the low-frequency component image in terms of resolution and number of channels; in response to consistency, it combines multiple color channels of the region image and multiple color channels of the low-frequency component image that have the same color attribute to obtain multiple channel pairs.

[0070] Taking an RGB (Red, Green, Blue) image as an example, both the regional image and the low-frequency component image include three color channels: red, green, and blue. The red channel of the regional image and the red channel of the low-frequency component image are combined into one channel pair; the green channel of the regional image and the green channel of the low-frequency component image are combined into one channel pair; and the blue channel of the regional image and the blue channel of the low-frequency component image are combined into one channel pair, resulting in a total of three channel pairs.

[0071] Then, based on the pixel-level difference between two color channels in multiple channel pairs, a high-frequency component image is generated.

[0072] For each pair of color channels, the pixel value of the corresponding pixel is subtracted from the pixel value of the low-frequency component image from the pixel value of the region image. This determines the pixel value of the corresponding pixel in the color channel of the high-frequency component image, and finally generates the high-frequency component image.

[0073] For example, for each pixel in the red channel of the region image, the pixel value of the corresponding pixel in the red channel of the low-frequency component image is subtracted from the pixel value of the red channel of the region image to determine the pixel value of the corresponding pixel in the red channel of the high-frequency component image, thus generating the red channel of the high-frequency component image. Each color channel is processed in turn to finally obtain the high-frequency component image.

[0074] In this implementation, consistent resolution ensures channel alignment without misalignment, pairing channels with the same color attributes avoids color mixing interference, and pixel-level difference calculations can accurately extract the unique high-frequency features of each color channel, ensuring the color fidelity of high-frequency component images.

[0075] In some optional implementations of this embodiment, the execution entity can perform the second step described above to obtain the fused image in the following manner: First, the fusion weights of the high-frequency component images are determined based on the objects to be deleted; then, the objects to be deleted and the high-frequency component images are fused according to the fusion weights to obtain the fused image.

[0076] First, the type and characteristics of the object to be deleted are determined, and the fusion weights of the high-frequency component images are determined based on the type and characteristics. For example, if the object to be deleted is fine hair (such as stubble or lip hair), more high-frequency details need to be added after deletion to restore the natural texture of the skin. The fusion weight of the high-frequency component images is set to 60%-70%, and the weight of the corresponding region in the object deletion image is 30%-40%. If it is light-colored spots or slight stains, the need for high-frequency details after deletion is lower. The weight of the high-frequency component images is set to 20%-30%, and the weight of the object deletion image is 70%-80%.

[0077] Next, the object region in the object deletion image where the object to be deleted once existed is located, ensuring that this region is perfectly aligned with the corresponding region in the high-frequency component image in terms of resolution and position. For the aligned region, pixel-level fusion calculation is performed according to the set weights (i.e., the final value of a pixel = pixel value of the object deletion image × object deletion weight + pixel value of the high-frequency component image × high-frequency weight), while the non-object regions retain the original pixels of the object deletion image. Finally, the boundary edges between the fused and non-fused regions are slightly smoothed to eliminate weight transition traces, resulting in a fused image with natural details and no fusion discontinuities.

[0078] In this implementation, the weights are precisely determined according to the characteristics of the object to be deleted, adapting to the repair needs of different objects, which makes the details of the fused image natural and the repair effect more accurate.

[0079] Step 204: Generate the processed image based on the original image and the fused image.

[0080] In this embodiment, the aforementioned execution entity can generate a processed image based on the original image and the fused image.

[0081] As an example, firstly, the object region of the object to be deleted in the original image is located. The part of the fused image that completely corresponds to this represented region is directly used to replace the pixels of the object region in the original image; the pixels of the non-object regions in the original image remain unchanged. After replacement, the boundary edges between the object region and the non-object region are smoothed using the mean of neighboring pixels to eliminate replacement traces, and finally, a processed image with the object to be deleted is generated.

[0082] As another example, a global pixel-by-pixel weighted fusion is performed on the original image and the fused image. Weights are set according to region attributes: non-object regions in the original image are assigned the first weight (e.g., 80%), and the corresponding regions in the fused image are assigned the second weight (e.g., 20%), prioritizing the preservation of the original image's natural texture; object regions in the original image are assigned the third weight (e.g., 20%), and the corresponding regions in the fused image are assigned the fourth weight (e.g., 80%), highlighting the restoration effect of the fused image. After weighted calculation, the image pixel values ​​are uniformly adjusted to the normal range, generating a harmoniously processed image.

[0083] In some optional implementations of this embodiment, the execution entity may also perform the following operation: determine a region marker map representing the region range of the region image in the original image.

[0084] A labeled image uses a digital image as its medium and visual symbols such as pixel values, colors, and lines to specifically mark the area of ​​a region within the original image. Figure 1 Generally, the size is consistent with the corresponding labeled image. For example, it can be a mask image, a mask layer image, a contour marker image, a grayscale region marker image, etc.

[0085] Taking the region marker map as a mask image as an example, the mask image is generally a binary image with the same size and resolution as the original image. The pixel value within the region range in the mask image is the first pixel value, such as 255 or 1, and the pixel value outside the region range is the second pixel value, such as 0.

[0086] For example, first, generate an initial mask image with the same resolution and size as the original image, where the pixel value of each pixel is 0; then, set the pixel values ​​within the region corresponding to the region image in the initial mask image to the marker value 255 or 1 to obtain the mask image.

[0087] For example, based on the determination of the region of the object to be deleted in the original image by the part detection model, a region marker map representing the region of the region image in the original image is output.

[0088] In this implementation, the above-mentioned execution entity can perform step 204 as follows: fuse the original image and the fused image according to the region label map to obtain the processed image.

[0089] In the region labeling map, the pixel values ​​within the region range corresponding to the region image (which is also the region range of the fused image in the original image) are labeled with a value of 255 or 1, while the pixel values ​​in other regions are 0. During the fusion process, the weight of the pixel values ​​within the region range corresponding to the region image in the original image is set to 0, and the weight of the pixel values ​​outside the region range is set to 1. The weight of the pixel values ​​in the fused image is set to 1, thereby performing pixel-by-pixel fusion of the original image and the fused image to obtain the processed image.

[0090] In this implementation, the region-labeled map improves the accuracy of image fusion while ensuring the efficiency of image fusion.

[0091] In some optional implementations of this embodiment, the execution entity may also perform the following operations: First, determine the location marker map that represents the area of ​​the location in the original image.

[0092] For example, a pre-trained part detection model can be used to determine the region of the part to be deleted in the original image; and a part marker map representing the region of the part in the original image can be generated.

[0093] Then, a skin marker map is determined to represent the area of ​​skin corresponding to the location in the original image.

[0094] By using a pre-trained skin segmentation model, the region of skin in the original image corresponding to the area of ​​the image to be deleted is determined; a skin marker map representing the region of skin in the original image corresponding to the area is generated.

[0095] The part marker map, skin marker map, and region marker map use the same type of marker map, such as a mask image.

[0096] Finally, by combining the site marker map and the skin marker map, a fused marker map is obtained.

[0097] A fused marker map is obtained by performing a union operation between the body part marker map and the skin marker map. Taking a marker value of 1 in the marker map as an example, all pixels with a marker value of 1 in both the body part marker map and the skin marker map are retained and combined to obtain the fused marker map.

[0098] In this implementation, the aforementioned execution entity can determine the region marker map representing the region range of the region image in the original image in the following way: determine the region marker map representing the region range of the region image in the original image from the fused marker map.

[0099] Taking the top-left corner coordinates (x3, y3) and bottom-right corner coordinates (x4, y4) of the region as an example, the region marker map within the region with the top-left corner coordinates (x3, y3) and bottom-right corner coordinates (x4, y4) is determined from the fused marker map.

[0100] In this implementation, by combining the part marking map and the skin marking map, a more accurate fused marking map is obtained, from which the region marking map is determined, further improving the accuracy and comprehensiveness of the region marking map.

[0101] In some optional implementations of this embodiment, the execution entity can obtain the region marking map in the following manner: First, an initial marker map representing the region extent of the region image in the original image is determined from the fused marker map.

[0102] Taking the top-left corner coordinates (x3, y3) and bottom-right corner coordinates (x4, y4) of the region as an example, the initial marker map within the region with the top-left corner coordinates (x3, y3) and bottom-right corner coordinates (x4, y4) is determined from the fused marker map.

[0103] Then, the label values ​​of the edge regions of the initial label map are incremented from the outside to the inside to generate a region label map.

[0104] In the initial labeled image, the pixel values ​​of the pixels in the labeled regions (corresponding to the region range of the region image in the original image) are all labeled values, such as 255 or 1. The labeled values ​​of the edge regions of the labeled regions increase from the outside in until they reach the labeled value, thus obtaining the region labeled image. The edge regions are, for example, regions composed of a preset number of pixels (e.g., 5).

[0105] In this implementation, by setting gradient marker values ​​in the edge region of the region marker map, abrupt changes in boundary pixel values ​​are avoided, eliminating harsh breaks in subsequent image processing, and ensuring that the area to be processed covered by the marker map is naturally connected with the surrounding area, thus guaranteeing the visual smoothness of the processed image.

[0106] See also Figure 3 , Figure 3 This is a schematic diagram of an application scenario 300 of the image processing method according to this embodiment. The target user 301 sends the original image 304 to the server 303 via a terminal device 302. The original image is a face image, and the face region includes the beard of the object to be deleted. The server first determines the region image representing the object to be deleted in the original image; then, it deletes the object to be deleted from the region image, obtaining an object deletion image; then, it fuses the object deletion image and the high-frequency component image of the high-frequency features in the region image, obtaining a fused image; finally, based on the original image and the fused image, it generates the processed image 305.

[0107] In this embodiment, an image processing method is provided, which involves determining a region image representing an object to be deleted in the original image; deleting the object to be deleted from the region image to obtain an object deletion image; fusing the object deletion image and a high-frequency component image representing high-frequency features in the region image to obtain a fused image; and generating a processed image based on the original image and the fused image. By deleting the object to be deleted from the original image and fusing the high-frequency component image representing high-frequency features, the image processing effect is improved.

[0108] Continue to refer to Figure 4 The illustration shows a schematic flow 400 of another embodiment of the image processing method according to the present disclosure. Flow 400 includes the following steps: Step 401: Determine the location image representing the part of the object to be deleted from the original image.

[0109] Step 402: Based on the key points in the part image, determine the region image representing the object to be deleted in the original image.

[0110] Step 403: Using a pre-trained image processing model, delete the objects to be deleted from the region image and fill the object regions of the objects to be deleted in the region image to obtain the object deletion image.

[0111] Step 404: Use a Gaussian low-pass filter to filter high-frequency features in the region image to obtain a low-frequency component image.

[0112] Step 405: Generate a high-frequency component image based on the pixel-level difference between the region image and the low-frequency component image.

[0113] Step 406: Determine the fusion weights of the high-frequency component images based on the objects to be deleted.

[0114] Step 407: Delete the image and high-frequency component image according to the fusion weight to obtain the fused image.

[0115] Step 408: Determine the location marker map that represents the area range of the location in the original image.

[0116] Step 409: Determine the skin marker map representing the area of ​​skin corresponding to the location in the original image.

[0117] Step 410: Combine the site marker map and the skin marker map to obtain the fused marker map.

[0118] Step 411: Determine an initial marker map from the fused marker map to represent the region extent of the region image in the original image.

[0119] Step 412: Increment the label values ​​of the edge regions of the initial label map from the outside to the inside to generate a region label map.

[0120] Step 413: Merge the original image and the merged image based on the region label map to obtain the processed image.

[0121] The image processing method in this embodiment, in process 400, compared to process 200 above, specifically describes the generation method of the region image, the generation method of the object deletion image, the generation method of the fused image, and the generation method of the processed image, which further improves the processing efficiency and effect of the object deletion operation and can meet the needs of large-scale image processing.

[0122] Continue to refer to Figure 5 As an implementation of the methods shown in the above figures, this disclosure provides an embodiment of an image processing apparatus, which is similar to... Figure 2 Corresponding to the method embodiments shown, the system can be specifically applied to various electronic devices.

[0123] like Figure 5As shown, the image processing apparatus 500 includes: an image determination unit 501 configured to determine a region image representing an object to be deleted in the original image; an object deletion unit 502 configured to delete the object to be deleted in the region image to obtain an object deletion image; an image fusion unit 503 configured to fuse the object deletion image and a high-frequency component image representing high-frequency features in the region image to obtain a fused image; and an image generation unit 504 configured to generate a processed image based on the original image and the fused image.

[0124] In some optional implementations of this embodiment, the image fusion unit 503 is further configured to: delete low-frequency features in the region image to obtain a high-frequency component image; and merge the deleted image and the high-frequency component image to obtain a fused image.

[0125] In some optional implementations of this embodiment, the image fusion unit 503 is further configured to: filter high-frequency features in the region image through a Gaussian low-pass filter to obtain a low-frequency component image; and generate a high-frequency component image based on pixel-level difference calculation between the region image and the low-frequency component image.

[0126] In some optional implementations of this embodiment, the image fusion unit 503 is further configured to: in response to the fact that the resolution of the regional image and the resolution of the low-frequency component image are the same, combine multiple color channels of the regional image and multiple color channels of the low-frequency component image with color channels of the same color attribute to obtain multiple channel pairs; and generate a high-frequency component image based on the pixel-level difference operation between two color channels in the multiple channel pairs.

[0127] In some optional implementations of this embodiment, the image fusion unit 503 is further configured to: determine the fusion weight of the high-frequency component image based on the object to be deleted; and fuse the object to be deleted and the high-frequency component image according to the fusion weight to obtain a fused image.

[0128] In some optional implementations of this embodiment, the image determination unit 501 is further configured to: determine a part image representing the location of the object to be deleted from the original image; and determine a region image representing the object to be deleted in the original image based on key points in the part image.

[0129] In some optional implementations of this embodiment, the above apparatus further includes: a marker map generation unit (not shown in the figure), configured to determine a region marker map representing the region range of the region image in the original image; and an image generation unit 504 further configured to: fuse the original image and the fused image according to the region marker map to obtain a processed image.

[0130] In some optional implementations of this embodiment, the marker map generation unit is further configured to: determine a part marker map representing the region range of the part in the original image; determine a skin marker map representing the region range of the skin corresponding to the part in the original image; combine the part marker map and the skin marker map to obtain a fused marker map; and the marker map generation unit is further configured to: determine a region marker map representing the region range of the region image in the original image from the fused marker map.

[0131] In some optional implementations of this embodiment, the marker map generation unit is further configured to: determine an initial marker map from the fused marker map to characterize the region range of the region image in the original image; and increment the marker values ​​of the edge regions of the marker regions in the initial marker map from the outside to the inside to generate a region marker map.

[0132] In some optional implementations of this embodiment, the object deletion unit 502 is further configured to: delete the object to be deleted in the region image through a pre-trained image processing model, and fill the object region of the object to be deleted in the region image to obtain the object deletion image.

[0133] In this embodiment, an image processing apparatus is provided. The image determination unit in the image processing apparatus determines the region image representing the object to be deleted in the original image; the object deletion unit deletes the object to be deleted in the region image to obtain the object deletion image; the image fusion unit fuses the object deletion image and the high-frequency component image of the high-frequency features in the region image to obtain the fused image; and the image generation unit generates a processed image based on the original image and the fused image. Thus, by deleting the object to be deleted in the original image and fusing the high-frequency component image representing the high-frequency features, the image processing effect is improved.

[0134] According to embodiments of the present disclosure, the present disclosure also provides an electronic device, the electronic device comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to implement the image processing method described in any of the above embodiments.

[0135] According to embodiments of the present disclosure, the present disclosure also provides a readable storage medium storing computer instructions that enable a computer to perform the image processing method described in any of the above embodiments.

[0136] This disclosure provides a computer program product that, when executed by a processor, can implement the image processing method described in any of the above embodiments.

[0137] Figure 6 A schematic block diagram of an example electronic device 600 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0138] like Figure 6 As shown, device 600 includes a computing unit 601, which can perform various appropriate actions and processes based on a computer program stored in read-only memory (ROM) 602 or a computer program loaded into random access memory (RAM) 603 from storage unit 608. RAM 603 may also store various programs and data required for the operation of device 600. The computing unit 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.

[0139] Multiple components in device 600 are connected to I / O interface 605, including: input unit 606, such as keyboard, mouse, etc.; output unit 607, such as various types of monitors, speakers, etc.; storage unit 608, such as disk, optical disk, etc.; and communication unit 609, such as network card, modem, wireless transceiver, etc. Communication unit 609 allows device 600 to exchange information / data with other devices through computer networks such as the Internet and / or various telecommunications networks.

[0140] The computing unit 601 can be a variety of general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 601 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various special-purpose artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 601 performs the various methods and processes described above, such as image processing methods. For example, in some embodiments, the image processing method may be implemented as a computer software program tangibly contained in a machine-readable medium, such as storage unit 608. In some embodiments, part or all of the computer program may be loaded and / or installed on device 600 via ROM 602 and / or communication unit 609. When the computer program is loaded into RAM 603 and executed by the computing unit 601, one or more steps of the image processing method described above may be performed. Alternatively, in other embodiments, the computing unit 601 may be configured to perform image processing methods by any other suitable means (e.g., by means of firmware).

[0141] Various embodiments of the systems and techniques described above herein can be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems-on-a-chip (SoCs), payload-programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transmitting data and instructions to the storage system, the at least one input device, and the at least one output device.

[0142] The program code used to implement the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, special-purpose computer, or other programmable image processing apparatus, such that when executed by the processor or controller, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a machine, partially on a machine, as a standalone software package partially on a machine and partially on a remote machine, or entirely on a remote machine or server.

[0143] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0144] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).

[0145] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.

[0146] Computer systems can include clients and servers. Clients and servers are generally geographically separated and typically interact via communication networks. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. Servers can be cloud servers, also known as cloud computing servers or cloud hosts, which are hosting products within the cloud computing service system to address the management difficulties and weak business scalability inherent in traditional physical hosts and Virtual Private Servers (VPS) services; they can also be servers for distributed systems or servers incorporating blockchain technology.

[0147] According to the technical solution of the embodiments of this disclosure, an image processing method and apparatus are provided. The method involves determining a region image representing an object to be deleted in the original image; deleting the object to be deleted from the region image to obtain an object deletion image; fusing the object deletion image and a high-frequency component image representing high-frequency features in the region image to obtain a fused image; and generating a processed image based on the original image and the fused image. By deleting the object to be deleted from the original image and fusing the high-frequency component image representing high-frequency features, the image processing effect is improved.

[0148] It should be understood that the various forms of processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution provided in this disclosure can be achieved, and this is not limited herein.

[0149] The specific embodiments described above do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. An image processing method, comprising: determining a region image in an original image representing a to-be-deleted object; deleting the to-be-deleted object in the region image to obtain an object-deleted image; fusing the object-deleted image and a high-frequency component image representing high-frequency features in the region image to obtain a fused image; generating a processed image based on the original image and the fused image.

2. The method of claim 1, wherein, The fusing the object-deleted image and the high-frequency component image representing high-frequency features in the region image to obtain a fused image comprises: deleting low-frequency features in the region image to obtain the high-frequency component image; fusing the object-deleted image and the high-frequency component image to obtain the fused image.

3. The method of claim 2, wherein, The deleting low-frequency features in the region image to obtain the high-frequency component image comprises: filtering high-frequency features in the region image through a Gaussian low-pass filter to obtain a low-frequency component image; generating the high-frequency component image based on pixel-level difference operations between the region image and the low-frequency component image.

4. The method of claim 3, wherein, The generating the high-frequency component image based on pixel-level difference operations between the region image and the low-frequency component image comprises: in response to the resolution of the region image being consistent with the resolution of the low-frequency component image, combining color channels of the same color attribute in multiple color channels of the region image and multiple color channels of the low-frequency component image to obtain multiple channel pairs; generating the high-frequency component image based on pixel-level difference operations between two color channels in the multiple channel pairs.

5. The method of claim 2, wherein, The fusing the object-deleted image and the high-frequency component image to obtain the fused image comprises: determining a fusion weight of the high-frequency component image according to the to-be-deleted object; fusing the object-deleted image and the high-frequency component image according to the fusion weight to obtain the fused image.

6. The method of claim 1, wherein, The determining a region image in an original image representing a to-be-deleted object comprises: determining a part image representing a part of the to-be-deleted object from the original image; determining a region image representing the to-be-deleted object in the original image based on key points in the part image.

7. The method of claim 1 or 6, wherein, Further comprising: determining a region marker image representing a region range of the region image in the original image; and The generating a processed image based on the original image and the fused image comprises: fusing the original image and the fused image according to the region marker image to obtain the processed image. Further comprising:

8. The method of claim 6, wherein, determining a part marker image representing a region range of the part of the to-be-deleted object in the original image; determining a skin marker image representing a region range of skin corresponding to the part of the to-be-deleted object in the original image; combining the part marker image and the skin marker image to obtain a fused marker image; and The determining a region marker image representing a region range of the region image in the original image comprises: determining a region marker image representing a region range of the region image in the original image from the fused marker image. The determining a region marker image representing a region range of the region image in the original image from the fused marker image comprises: ​ 9. The method of claim 8, wherein, ​ determine an initial marker map representing a region range of the region image in the original image from the fusion marker map; incrementally increase marker values of edge regions of the marker regions in the initial marker map from outside to inside, to generate the region marker map.

10. The method of claim 1, wherein, The deleting the to-be-deleted object in the region image to obtain an object-deleted image comprises: deleting the to-be-deleted object in the region image and filling an object region of the to-be-deleted object in the region image by using a pre-trained image processing model, to obtain the object-deleted image. 11.An image processing apparatus, comprising: an image determining unit configured to determine a region image representing a to-be-deleted object in an original image; an object deleting unit configured to delete the to-be-deleted object in the region image to obtain an object-deleted image; an image fusing unit configured to fuse the object-deleted image and a high-frequency component image representing high-frequency features in the region image to obtain a fusion image; an image generating unit configured to generate a processed image based on the original image and the fusion image.

12. An electronic device, comprising: comprise: at least one processor; and a memory connected with the at least one processor in communication; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1-10.

13. A non-transitory computer-readable storage medium having stored thereon computer instructions, wherein, The computer instructions are used to enable the computer to perform the method of any one of claims 1-10.

14. A computer program product, comprising: A computer program, when executed by a processor, implements the method of any one of claims 1-10.