Image processing method and device, electronic equipment, chip and medium
By identifying the first region in the original image and using an image processing model for foreground removal and background completion, the problem of poor background completion quality is solved, achieving higher quality image processing results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-03
- Publication Date
- 2026-04-14
AI Technical Summary
When performing foreground removal and background completion on region A in the original image, the existing technology results in poor background completion quality, which can easily generate unrealistic illusionary content.
By identifying the first region to be processed from the original image, and performing foreground removal and background completion based on the image within the target region, interference from the second region is blocked, and image processing models such as GANs or VAEs are used for processing.
It improves the quality of background completion, avoids the generation of foreground elements related to content, and optimizes user experience and image processing efficiency.
Smart Images

Figure CN121860891A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of image processing technology, and in particular to an image processing method, apparatus, electronic device, chip, and storage medium. Background Technology
[0002] Currently, when performing foreground removal and background completion on region A in the original image, most methods rely on images outside region A in the original image to perform foreground removal and background completion on region A in the original image, resulting in poor background completion quality. Summary of the Invention
[0003] This disclosure provides an image processing method, apparatus, electronic device, chip, and storage medium to at least solve the problem of poor background completion quality in related technologies. The technical solution of this disclosure is as follows: According to a first aspect of the present disclosure, an image processing method is provided, comprising: in response to a first operation, determining a first region to be foreground removal from an original image; if the original image includes a second region, performing foreground removal and background completion processing on the first region in the original image based on an image within a target region in the original image to obtain a target image; wherein the second region is associated with the content of the first region, and the target region includes image regions in the original image other than the first region and the second region.
[0004] According to a second aspect of the present disclosure, an image processing apparatus is provided, comprising: a determining module configured to, in response to a first operation, determine a first region to be foreground removal from an original image; and a processing module configured to, if the original image includes a second region, perform foreground removal and background completion processing on the first region in the original image based on an image within a target region in the original image to obtain a target image; wherein the second region is associated with the content of the first region, and the target region includes image regions in the original image other than the first region and the second region.
[0005] According to a third aspect of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the steps of the image processing method described in the first aspect of the present disclosure.
[0006] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided that stores computer program instructions thereon, which, when executed by a processor, implement the steps of the image processing method described in the first aspect of the present disclosure.
[0007] According to a fifth aspect of the present disclosure, a chip is provided, the chip including an interface circuit and a processing circuit coupled to each other, the interface circuit being used to input or output signals, and the processing circuit being configured to implement the steps of the image processing method described in the first aspect of the present disclosure.
[0008] According to a sixth aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements the steps of the image processing method described in the first aspect of the present disclosure.
[0009] The technical solution provided by the embodiments of this disclosure brings at least the following beneficial effects: In response to a first operation, a first region to be foreground removal is determined from the original image. If the original image includes a second region, foreground removal and background completion processing are performed on the first region in the original image based on the image within the target region in the original image to obtain a target image. The second region is content-related to the first region, and the target region includes image regions in the original image other than the first and second regions. Therefore, when the original image includes a second region, the image within the target region in the original image can be considered when performing foreground removal and background completion processing on the first region in the original image to obtain a target image. That is, the image within the target region in the original image is used as a reference for generating the first region to perform foreground removal and background completion processing on the first region. This can block the interference of the image within the second region on the background completion of the first region, thereby avoiding the generation of content-related foreground elements in the first region, improving the background completion quality, and optimizing the user experience.
[0010] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0011] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0012] Figure 1 This is a flowchart illustrating an image processing method according to an exemplary embodiment.
[0013] Figure 2 This is a flowchart illustrating an image processing method according to another exemplary embodiment.
[0014] Figure 3 This is a flowchart illustrating an image processing method according to another exemplary embodiment.
[0015] Figure 4This is a flowchart illustrating a model training method according to an exemplary embodiment.
[0016] Figure 5 This is a schematic diagram of the structure of an image processing apparatus according to an exemplary embodiment.
[0017] Figure 6 This is a schematic diagram of the structure of an electronic device according to an exemplary embodiment.
[0018] Figure 7 This is a schematic diagram of the structure of a chip according to an exemplary embodiment. Detailed Implementation
[0019] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0020] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0021] The image processing method, apparatus, electronic device, chip, and storage medium of the present disclosure are described below with reference to the accompanying drawings.
[0022] Figure 1 This is a flowchart illustrating an image processing method according to an exemplary embodiment, such as... Figure 1 As shown, the image processing method of this disclosure includes the following steps.
[0023] S101, in response to the first operation, determine the first region to be foreground removal from the original image.
[0024] It should be noted that the execution subject of the image processing method in this embodiment is an electronic device, such as a terminal device, vehicle, camera, server, chip, etc. The terminal device may include a mobile phone, wearable device (such as a smartwatch, smart glasses), laptop computer, etc. The vehicle may include an in-vehicle terminal, in-vehicle controller, etc. The chip may include an ISP (Image Signal Processor), etc.
[0025] The image processing method of this disclosure embodiment can be executed by the image processing device of this disclosure embodiment. The image processing device of this disclosure embodiment can be configured in any electronic device to execute the image processing method of this disclosure embodiment.
[0026] There are no strict limitations on the original images. For example, in an image editing scenario, the original image can include images captured by the terminal device, such as facial images, landscape images, and building images. In a driving scenario, the original image can include images of the vehicle's external environment (such as images of the road in front of the vehicle), images of the vehicle's interior (such as images of the driver's area), and road surveillance images. In a live streaming scenario, the original image can include at least one frame from the live video. In a smart home scenario, the original image can include at least one frame from home surveillance video.
[0027] The original image may include RGB, HSV, HSL, YCbCr, Lab, YUV images, etc.
[0028] There are no strict restrictions on the first operation, such as including mechanical operation, touch operation, voice interaction operation, etc.
[0029] For example, in response to a first operation, a first region to be foreground removal is determined from the original image, including, in response to a selection operation, using a selected image region in the original image as the first region. The selection operation is not overly limited; it can include smearing, selecting by bounding, clicking, polygon drawing, etc.
[0030] For example, in response to a first operation, determining a first region to be foreground removal from the original image includes, in response to a voice interaction operation, acquiring voice information and performing voice recognition on the voice information to determine the first region from the original image.
[0031] Optionally, before determining the first region to be foreground removal from the original image in response to the first operation, the method further includes determining the original image to be foreground removal in response to the second operation.
[0032] It should be noted that there are no strict limitations on the second operation, such as including mechanical operation, touch operation, voice interaction operation, etc.
[0033] For example, in response to a second operation, determining the original image to be foreground removal includes, in response to a selection operation, using a selected candidate image as the original image.
[0034] For example, in response to a second operation, determining the original image to be foreground removal includes, in response to a voice interaction operation, acquiring voice information and performing voice recognition on the voice information to determine the original image from an image library.
[0035] S102, if the original image includes a second region, based on the image within the target region in the original image, perform foreground removal and background completion processing on the first region in the original image to obtain the target image.
[0036] It should be noted that the content of the second region is related to that of the first region, and the target region includes image regions in the original image other than the first and second regions. The second region and the first region are different image regions.
[0037] The association between the content of the second region and the first region refers to the association between the image content in the second region of the original image and the image content in the first region of the original image.
[0038] Optionally, the content of the second region is associated with that of the first region, including at least one of the following: The content category of the second region is consistent with that of the first region, which means that the image content in the second region of the original image is of the same category as the image content in the first region of the original image. In other words, the instance corresponding to the second region of the original image is the same type of instance as the instance corresponding to the first region of the original image. The content of the second region and the first region belong to the same object. This means that the image content in the second region of the original image and the image content in the first region of the original image belong to the same object. In other words, the second region and the first region of the original image show different parts of the same object. The object belonging to the second region is created by the object belonging to the first region blocking light. It refers to the object belonging to the second region in the original image. It is a dark area formed on the projection surface because the object belonging to the first region in the original image blocks light. In other words, the object belonging to the second region is the "shadow" of the object belonging to the first region.
[0039] The target region can be any image region other than the first and second regions in the original image, or it can be a local image region other than the first and second regions in the original image; there are no strict limitations here. For example, the target region can be the area surrounding the first region, or the distance between the target region and the first region can be less than a set threshold.
[0040] For example, if the original image includes regions A, B, C, D, and E, and the image content in region A is face 1, the image content in region B is face 2, the image content in region C is text 1, the image content in region D is text 2, and the image content in region E is the background, then the target region must meet the following conditions: the distance between the target region and the first region is less than a set threshold.
[0041] If the first region is region A, the distances between regions C and E and region A are all less than a set threshold, and the distance between region D and region A is greater than a set threshold, then the second region includes region B, and the target region includes regions C and E.
[0042] If the first region is region C, and the distances between regions A, B, and E and region C are all less than a set threshold, then the second region includes region D, and the target region includes regions A, B, and E.
[0043] Currently, when performing foreground removal and background completion on region A in the original image, most methods rely on images outside region A in the original image to perform foreground removal and background completion on region A in the original image, resulting in poor background completion quality.
[0044] For example, if the image content in region A of the original image is face 1, and the image content outside region A of the original image includes face 2, then region A may be processed by foreground removal and background completion based on the image of face 2. This would result in the image content of region A having face 1 removed after processing, but unexpectedly having face 2 added, or even having a fake face based on face 2, i.e. generating unreal "illusion content" in region A, which would seriously affect the quality of background completion.
[0045] For example, if the image content in region A of the original image is text 1, and the image content outside region A of the original image includes text 2, then region A may be processed by foreground removal and background completion based on the image of text 2. This may result in the image content of region A having text 1 removed, but unexpectedly having text 2 added, or even text forged based on text 2, i.e. generating unreal "illusion content" in region A, which seriously affects the quality of background completion.
[0046] For example, if the image content in region A of the original image is the left half of building 1, and the image content outside region A of the original image includes the right half of building 1, then region A may be processed by foreground removal and background completion based on the right half of building 1. This would result in region A having the left half of building 1 removed in the processed image, but unexpectedly having the right half of building 1 added. It might even result in a fake building based on the right half of building 1, i.e., generating unreal "illusion content" in region A, which seriously affects the quality of background completion.
[0047] For example, if the image content within region A of the original image is human body 1, and the image content outside region A of the original image includes the shadow of human body 1, then region A may be processed by foreground removal and background completion based on the shadow of human body 1. This would result in the image content of region A having human body 1 removed after processing, but unexpectedly having the shadow of human body 1 added. It is even possible that a fake human body based on the shadow of human body 1 will appear, that is, an unreal "illusion content" will be generated in region A, which seriously affects the quality of background completion.
[0048] In this disclosure, when the original image includes a second region, the image within the target region of the original image can be taken into account, and foreground removal and background completion processing can be performed on the first region of the original image to obtain the target image. That is, the image within the target region of the original image is used as a reference for generating the first region, and foreground removal and background completion processing can be performed on the first region. This can block the interference of the image within the second region on the background completion of the first region, thereby avoiding the generation of content-related foreground elements in the first region, improving the background completion quality, and optimizing the user experience.
[0049] For example, it can avoid generating foreground elements with the same content category in the first area, avoid generating foreground elements of the rest of the same object in the first area, and avoid generating shadows of foreground elements to be eliminated in the first area.
[0050] Optionally, based on the image within the target region in the original image, foreground removal and background completion processing are performed on the first region in the original image to obtain the target image. This includes using an image processing model to perform foreground removal and background completion processing on the first region in the original image based on the image within the target region in the original image to obtain the target image.
[0051] It should be noted that the image processing models are not limited in many ways, and include diffusion models, GANs (Generative Adversarial Networks), VAEs (Variational Autoencoders), etc. It should be clarified that VAEs include both encoders and decoders.
[0052] Optionally, based on the image within the target region of the original image, a foreground removal and background completion process is performed on the first region of the original image using an image processing model to obtain the target image. This includes performing foreground removal and background completion on the first region of the original image based on the image within the target region of the original image and the prompt text using an image processing model to obtain the target image.
[0053] It should be noted that the prompt text is not subject to many restrictions, such as including "Please remove the faces in the selected area".
[0054] Optionally, the first region in the original image is processed by foreground removal and background completion based on the image within the target region in the original image using an image processing model to obtain the target image. This includes, if the image processing model includes a diffusion model, denoising the pure noise image based on the image within the target region in the original image using a diffusion model to obtain the target image.
[0055] Optionally, based on the image within the target region of the original image, the first region in the original image is subjected to foreground removal and background completion processing by an image processing model to obtain the target image. This includes, in the case that the image processing model includes an encoder, a diffusion model, and a decoder, encoding the image within the target region of the original image by the encoder to obtain a first image code, denoising the pure noise image based on the first image code by the diffusion model to obtain a second image code, and decoding the second image code by the decoder to obtain the target image.
[0056] It should be noted that this disclosure does not impose any restrictions on the execution sequence of steps S101 to S102. Figure 1 The example is performed by executing steps S101 to S102 in sequence only.
[0057] The image processing method provided in the embodiments of this disclosure, in response to a first operation, determines a first region to be foreground removal from the original image. If the original image includes a second region, foreground removal and background completion processing are performed on the first region in the original image based on the image within the target region in the original image to obtain a target image. The second region is content-related to the first region, and the target region includes image regions in the original image other than the first and second regions. Therefore, when the original image includes a second region, the image within the target region in the original image can be considered when performing foreground removal and background completion processing on the first region to obtain the target image. That is, the image within the target region in the original image is used as a reference for generating the first region to perform foreground removal and background completion processing on the first region. This can prevent interference from the image within the second region on the background completion of the first region, thereby avoiding the generation of content-related foreground elements in the first region, improving the background completion quality, and optimizing the user experience.
[0058] Based on any of the above embodiments, the method further includes determining candidate regions associated with the content of the first region from the original image, and using the candidate regions as the second region; or, if the target region satisfies a set condition including that the target region is located in the surrounding area of the first region, determining candidate regions located in the surrounding area of the first region, and using them as the second region. Thus, candidate regions associated with the content of the first region can be determined from the original image, and directly used as the second region. Alternatively, if the target region satisfies a set condition including that the target region is located in the surrounding area of the first region, determining candidate regions located in the surrounding area of the first region, and using them as the second region.
[0059] Optionally, a candidate region located in the surrounding area of the first region is determined as the second region. This includes obtaining the distance between the second region and the first region when the distance between the target region and the first region is less than a set threshold, and selecting the candidate region whose distance is less than the set threshold as the second region.
[0060] For example, if the original image includes regions A, B, C, D, and E, the image content in region A is face 1, the image content in region B is face 2, the image content in region C is face 3, the image content in region D is face 4, and the image content in region E is the background.
[0061] If the first region is region A, and the distances between regions B, C, and E and region A are all less than a set threshold, and the distance between region D and region A is greater than a set threshold, then the second region includes regions B and C, and the target region includes region E.
[0062] Optionally, the method further includes cropping the image within the first region of the original image and the image within the surrounding region of the first region to obtain a cropped image; performing masking processing on the first region in the cropped image to obtain a mask for the first region; performing instance segmentation on the cropped image to determine the category and segmentation mask of each candidate instance in the cropped image; obtaining the overlap parameter between the segmentation mask and the mask of the first region; using the segmentation mask corresponding to the largest overlap parameter as the target mask; and obtaining the category of the target instance corresponding to the target mask as the content category of the first region.
[0063] It should be noted that there are no strict limitations on the surrounding areas of the first region, such as including regions whose distance from the first region is less than a set threshold. The cropped image includes the image within the first region of the original image, as well as the images within the surrounding areas of the first region of the original image. There are no strict limitations on the overlap parameters, such as similarity, the size of the overlapping region, the ratio of the overlapping region, and IOU (Intersection over Union).
[0064] The processing area identified by the mask of the first region is the first region. No restrictions are placed on the types of masks, such as arrays, multi-valued images (e.g., binary images), grayscale images, etc. For example, taking a binary image as the mask, pixels with a value of 1 belong to the processing area, and pixels with a value of 0 belong to the non-processing area; that is, white pixels in a binary image belong to the processing area, and black pixels belong to the non-processing area. The mask can be obtained using any mask generation method from relevant technologies; no further restrictions are placed here.
[0065] It should be noted that in this embodiment, if the mask of the first region is an image, the size of the mask of the first region is the same as the size of the cropped image.
[0066] For example, taking a binary image as the mask of the first region, the first region in the cropped image is masked to obtain the mask of the first region. This includes setting the pixel value of the pixels in the first region of the cropped image to 1 and setting the pixel value of the pixels in the image region outside the first region of the cropped image to 0, so as to obtain the mask of the first region.
[0067] Based on any of the above embodiments, the method further includes merging the mask of the first region and the mask of the second region to obtain a merged mask, and determining the image within the target region in the original image based on the merged mask; wherein the merged mask is used to identify the target region.
[0068] It should be noted that the processing area identified by the mask of the second region is the second region itself, while the processing area identified by the merge mask includes both the first and second regions. Merging the masks of the first and second regions can be achieved using any mask merging method from relevant technologies; no further limitations are imposed here.
[0069] For example, the masks of the first region and the second region are merged to obtain a merged mask, which includes using the processing areas identified by the masks of the first region and the processing areas identified by the masks of the second region as the processing areas identified by the merged mask.
[0070] For example, taking a binary image as an example where both the mask for the first region and the mask for the second region are used, the masks for the first region and the second region are merged to obtain a merged mask. This includes adding the positions of pixels with a value of 1 in the mask for the first region and the positions of pixels with a value of 1 in the mask for the second region to a first position set, determining the first pixel and the second pixel in the initial mask, where the position of the first pixel exists in the first position set and the position of the second pixel does not exist in the first position set, setting the pixel value of the first pixel in the initial mask to 1, and setting the pixel value of the second pixel in the initial mask to 0, to obtain the merged mask.
[0071] Figure 2 This is a flowchart illustrating an image processing method according to another exemplary embodiment, such as... Figure 2 As shown, the image processing method of this disclosure includes the following steps.
[0072] S201, in response to the first operation, determine the first region to be foreground removal from the original image.
[0073] S202, if the original image includes a second region, crop the image in the first region of the original image and the image in the surrounding region of the first region to obtain a cropped image.
[0074] S203, perform masking on the first region in the cropped image to obtain the mask of the first region.
[0075] The details of steps S201-S203 can be found in the above embodiments and will not be repeated here.
[0076] S204, perform masking on the second region in the cropped image to obtain the mask of the second region.
[0077] It should be noted that in this embodiment, if the mask of the second region is an image, the size of the mask of the second region is the same as the size of the cropped image.
[0078] For example, taking the second region mask as a binary image, the second region in the cropped image is masked to obtain the mask of the second region. This includes setting the pixel value of the pixels in the second region of the cropped image to 1 and setting the pixel value of the pixels in the image region outside the second region of the cropped image to 0, so as to obtain the mask of the second region.
[0079] S205, merge the mask of the first region and the mask of the second region to obtain the merged mask.
[0080] It should be noted that in this embodiment, if the merge mask is an image, the size of the merge mask is the same as the size of the cropped image.
[0081] The details of step S205 can be found in the above embodiments and will not be repeated here.
[0082] S206, based on the merging mask, the first region and the second region in the cropped image are masked to obtain the first reference image, which serves as the image within the target region of the cropped image.
[0083] In this embodiment, the target region includes the image region other than the first and second regions in the cropped image. Compared with directly using all image regions other than the first and second regions in the original image as the target region, the target region in this embodiment is only the surrounding region of the first region. Only the surrounding region of the first region needs to be used as the generation reference for the first region to perform background completion processing on the first region. This makes the difference between the processed image of the first region and its surrounding images smaller, and the transition between the processed image of the first region and its surrounding images smoother, thus improving the visual consistency and naturalness of the target image.
[0084] The first reference image has the same size as the cropped image. The first reference image retains the image within the target region of the cropped image and removes the images within the first and second regions of the cropped image. Therefore, the first reference image can be regarded as the image within the target region of the cropped image.
[0085] It should be noted that, based on the merging mask, the first and second regions in the cropped image are masked to obtain the first reference image. This can be achieved using any image masking method in related technologies, and no further restrictions are imposed here.
[0086] For example, taking a binary image with a merged mask as an example, based on the merged mask, the first and second regions in the cropped image are masked to obtain a first reference image. This includes adding the positions of pixels with a value of 1 in the merged mask to a second position set, determining the third pixel in the cropped image (where the position of the third pixel exists in the second position set), and setting the pixel value of the third pixel in the cropped image to 0. The pixels with a value of 1 in the merged mask are located in both the first and second regions.
[0087] S207, based on the first reference image, perform foreground removal and background completion processing on the first region in the cropped image to obtain the first intermediate image.
[0088] In this disclosure, the cropped image is used as the image to be processed, and foreground removal and background completion are performed only on the first region therein. It is not necessary to use the entire original image as the image to be processed, which can significantly reduce the amount of computation and improve the efficiency of image processing. At the same time, it can focus more on the first region, avoid unnecessary impact on image regions that do not need to be processed, and improve the quality of image processing.
[0089] For example, the size of the first intermediate image is the same as the size of the cropped image.
[0090] Optionally, based on the first reference image, foreground removal and background completion processing are performed on the first region in the cropped image to obtain a first intermediate image, including performing foreground removal and background completion processing on the first region in the cropped image based on the first reference image using an image processing model to obtain the first intermediate image.
[0091] Optionally, based on the first reference image, an image processing model is used to perform foreground removal and background completion processing on the first region of the cropped image to obtain a first intermediate image. This includes performing foreground removal and background completion processing on the first region of the cropped image based on the first reference image and the prompt text using an image processing model to obtain a first intermediate image.
[0092] Optionally, a first intermediate image is obtained by performing foreground removal and background completion processing on a first region in the cropped image based on a first reference image using an image processing model. This includes, if the image processing model includes a diffusion model, performing denoising processing on a purely noisy image based on the first reference image using a diffusion model to obtain the first intermediate image.
[0093] Optionally, based on the first reference image, the first region in the cropped image is subjected to foreground removal and background completion processing by an image processing model to obtain a first intermediate image. This includes encoding the first reference image by the encoder to obtain a first image code when the image processing model includes an encoder, a diffusion model, and a decoder; performing denoising processing on the pure noise image based on the first image code by the diffusion model to obtain a second image code; and decoding the second image code by the decoder to obtain the target image.
[0094] S208, crop the image in the first region of the first intermediate image to obtain the first completed image.
[0095] S209, the first completed image is overlaid on the image in the first region of the original image to obtain the target image.
[0096] It should be noted that the first completed image includes the image within the first region of the first intermediate image.
[0097] In this disclosure, the first completed image can cover the image in the first region of the original image to obtain the target image. This allows for greater focus on the first region while keeping the image content of the image regions outside the first region in the original image unchanged, thereby enhancing the local controllability of foreground elimination and background completion processing and improving image processing quality.
[0098] Optionally, after covering the image in the first region of the original image with the first completed image to obtain the target image, the process further includes at least one of color correction, edge smoothing, and brightness consistency adjustment of the stitching area of the target image to ensure that the first completed image and the original image are visually seamlessly connected, avoiding artifacts such as edge breaks and brightness inconsistencies, and improving the overall naturalness.
[0099] It should be noted that this disclosure does not impose any restrictions on the execution sequence of steps S201 to S209. Figure 2 The example only demonstrates the sequential execution of steps S201 to S209.
[0100] The image processing method provided in the embodiments of this disclosure includes an image region other than the first and second regions in the cropped image as the target region. Compared with directly using all image regions other than the first and second regions in the original image as the target region, the target region in this embodiment is only the surrounding area of the first region. It is only necessary to use the surrounding area of the first region as the generation reference of the first region to perform background completion processing on the first region, so that the difference between the processed image of the first region and its surrounding images is small, the transition between the processed image of the first region and its surrounding images is smoother, and the visual consistency and naturalness of the target image are improved.
[0101] Furthermore, by using the cropped image as the image to be processed and performing foreground removal and background completion only on the first region within it, the entire original image can be used as the image to be processed, which can significantly reduce the amount of computation and improve image processing efficiency. At the same time, it can focus more on the first region, avoiding unnecessary impact on image regions that do not need to be processed, thus improving image processing quality.
[0102] In addition, the first completed image can be overlaid on the image in the first region of the original image to obtain the target image. This allows for a greater focus on the first region while keeping the image content of the image regions outside the first region in the original image unchanged. This enhances the local controllability of foreground elimination and background completion processing and improves the image processing quality.
[0103] Figure 3 This is a flowchart illustrating an image processing method according to another exemplary embodiment, such as... Figure 3 As shown, the image processing method of this disclosure includes the following steps.
[0104] S301, in response to the first operation, determine the first region to be foreground removal from the original image.
[0105] The details of step S301 can be found in the above embodiments and will not be repeated here.
[0106] S302, if the original image includes a second region, perform masking processing on the first region in the original image to obtain a mask for the first region.
[0107] It should be noted that in this embodiment, if the mask of the first region is an image, the size of the mask of the first region is the same as the size of the original image.
[0108] For example, taking the first region's mask as a binary image, the first region in the original image is masked to obtain the mask of the first region. This includes setting the pixel value of the pixels in the first region of the original image to 1 and setting the pixel value of the pixels in the image regions outside the first region of the original image to 0, so as to obtain the mask of the first region.
[0109] S303, perform masking on the second region in the original image to obtain the mask for the second region.
[0110] It should be noted that in this embodiment, if the mask of the second region is an image, the size of the mask of the second region is the same as the size of the original image.
[0111] For example, taking the second region mask as a binary image, the second region in the original image is masked to obtain the mask of the second region. This includes setting the pixel value of the pixels in the second region of the original image to 1 and setting the pixel value of the pixels in the image region outside the second region of the original image to 0, so as to obtain the mask of the second region.
[0112] S304, merge the mask of the first region and the mask of the second region to obtain the merged mask.
[0113] It should be noted that in this embodiment, if the merged mask is an image, the size of the merged mask is the same as the size of the original image.
[0114] The details of step S205 can be found in the above embodiments and will not be repeated here.
[0115] S305, based on the merging mask, masking processing is performed on the first and second regions in the original image to obtain a second reference image, which serves as the image within the target region in the original image.
[0116] In this embodiment, the target region is all image regions other than the first and second regions in the original image. Images within the target region can provide richer image information to perform foreground removal and background completion processing on the first region.
[0117] The second reference image has the same dimensions as the original image. The second reference image retains the image within the target region of the original image while removing the images from the first and second regions of the original image. Therefore, the second reference image can be considered as the image within the target region of the original image.
[0118] It should be noted that, based on the merging mask, the first and second regions in the original image are masked to obtain the second reference image. This can be achieved using any image masking method in the relevant technologies, and no further restrictions are imposed here.
[0119] For example, taking a binary image with a merged mask as an example, based on the merged mask, the first and second regions in the original image are masked to obtain a second reference image. This includes adding the positions of pixels with a value of 1 in the merged mask to a third location set, determining the fourth pixel in the original image (where the position of the fourth pixel exists in the third location set), and setting the pixel value of the fourth pixel in the original image to 0. The pixels with a value of 1 in the merged mask are located in both the first and second regions.
[0120] S306, Based on the second reference image, perform foreground removal and background completion processing on the first region in the original image to obtain the target image.
[0121] Optionally, based on the second reference image, foreground removal and background completion processing are performed on the first region of the original image to obtain the target image. This includes performing foreground removal and background completion processing on the first region of the original image based on the second reference image to obtain a second intermediate image, cropping the image within the first region of the second intermediate image to obtain a second completed image, and covering the image within the first region of the original image with the second completed image to obtain the target image. Therefore, by covering the image within the first region of the original image with the second completed image to obtain the target image, the focus can be more on the first region while maintaining the image content of the image regions outside the first region in the original image unchanged. This enhances the local controllability of foreground removal and background completion processing and improves image processing quality.
[0122] For example, the size of the second intermediate image is the same as the size of the original image. The second completed image includes the image within the first region of the second intermediate image. Based on the second reference image, foreground removal and background completion processing are performed on the first region of the original image to obtain the relevant content of the second intermediate image. This process can be similar to performing foreground removal and background completion processing on the first region of the original image based on the second reference image to obtain the relevant content of the target image, which will not be elaborated here.
[0123] Optionally, after covering the image in the first region of the original image with the second completed image to obtain the target image, the process further includes at least one of color correction, edge smoothing, and brightness consistency adjustment of the stitching area of the target image to ensure that the second completed image and the original image are visually seamlessly connected, avoid artifacts such as edge breaks and brightness inconsistencies, and improve the overall naturalness.
[0124] Optionally, based on the second reference image, foreground removal and background completion processing are performed on the first region in the original image to obtain the target image, including performing foreground removal and background completion processing on the first region in the original image based on the second reference image using an image processing model to obtain the target image.
[0125] Optionally, the first region in the original image is processed by foreground removal and background completion based on the second reference image using an image processing model to obtain the target image. This includes processing the first region in the original image by foreground removal and background completion based on the second reference image and the prompt text using an image processing model to obtain the target image.
[0126] Optionally, the first region in the original image is subjected to foreground removal and background completion processing based on the second reference image using an image processing model to obtain the target image. This includes, if the image processing model includes a diffusion model, denoising the pure noise image based on the second reference image using a diffusion model to obtain the target image.
[0127] Optionally, based on the second reference image, the first region in the original image is subjected to foreground removal and background completion processing by an image processing model to obtain a target image. This includes encoding the second reference image by the encoder to obtain a first image code when the image processing model includes an encoder, a diffusion model, and a decoder; denoising the pure noise image based on the first image code by the diffusion model to obtain a second image code; and decoding the second image code by the decoder to obtain the target image.
[0128] It should be noted that this disclosure does not impose any restrictions on the execution sequence of steps S301 to S306. Figure 3 The example only demonstrates the sequential execution of steps S301 to S306.
[0129] The image processing method provided in the embodiments of this disclosure performs masking processing on a first region in the original image to obtain a mask for the first region, performs masking processing on a second region in the original image to obtain a mask for the second region, and performs occlusion processing on the first and second regions in the original image based on the merged mask to obtain a second reference image, which serves as the image within the target region in the original image. Based on the second reference image, foreground removal and background completion processing are performed on the first region in the original image to obtain the target image. Thus, the target region is all image regions other than the first and second regions in the original image, and the image within the target region can provide richer image information for performing foreground removal and background completion processing on the first region.
[0130] The training process of the image processing model will be described below.
[0131] Figure 4 This is a flowchart illustrating a model training method according to an exemplary embodiment, such as... Figure 4 As shown, the model training method of this disclosure includes the following steps.
[0132] S401, Obtain a first sample image; wherein, the first sample image includes the sample region to be eliminated.
[0133] S402, based on the mask of the sample region, the sample region in the first sample image is masked to obtain the sample reference image.
[0134] It should be noted that the execution subject of the model training method in this embodiment is an electronic device, such as a terminal device, vehicle, camera, server, chip, etc. The terminal device may include a mobile phone, wearable device (such as a smartwatch, smart glasses), laptop computer, etc. The vehicle may include an in-vehicle terminal, in-vehicle controller, etc., and the chip may include an ISP, etc.
[0135] The model training method of this disclosure embodiment can be executed by the image processing device of this disclosure embodiment. The image processing device of this disclosure embodiment can be configured in any electronic device to execute the model training method of this disclosure embodiment.
[0136] Optionally, the method further includes masking the sample regions in the first sample image to obtain a mask for the sample regions.
[0137] If the mask of the sample region is an image, the size of the mask of the sample region is the same as the size of the first sample image.
[0138] For example, taking a binary image as the mask of the sample region, the sample region in the first sample image is masked to obtain the mask of the sample region. This includes setting the pixel value of the pixel in the sample region of the first sample image to 1 and setting the pixel value of the pixel in the image region outside the sample region of the first sample image to 0, so as to obtain the mask of the sample region.
[0139] The size of the sample reference image is the same as that of the first sample image. The sample reference image retains the image region outside the sample region in the first sample image and removes the image region within the sample region in the first sample image. Therefore, the sample reference image can be regarded as the image region outside the sample region in the first sample image.
[0140] It should be noted that, based on the mask of the sample region, the sample region in the first sample image is masked to obtain the sample reference image. This can be achieved using any image masking method in the relevant technology, and no further restrictions are imposed here.
[0141] For example, taking a binary image as the mask for the sample region, the sample region in the first sample image is masked based on the mask to obtain a sample reference image. This includes adding the positions of pixels with a value of 1 in the mask of the sample region to a fourth position set, determining the fifth pixel in the first sample image (where the position of the fifth pixel exists in the fourth position set), and setting the pixel value of the fifth pixel in the first sample image to 0. Pixels with a value of 1 in the mask of the sample region are located within the sample region.
[0142] S403, based on the sample reference image, the image processing model performs foreground removal and background completion processing on the sample region in the first sample image to obtain the predicted image.
[0143] Optionally, based on the sample reference image, an image processing model is used to perform foreground removal and background completion processing on the sample region in the first sample image to obtain a predicted image. This includes performing foreground removal and background completion processing on the sample region in the first sample image based on the sample reference image and sample prompt text using an image processing model to obtain a predicted image.
[0144] Optionally, a foreground removal and background completion process is performed on the sample region in the first sample image based on the sample reference image using an image processing model to obtain a predicted image. This includes, if the image processing model includes a diffusion model, denoising the pure noise image based on the sample reference image using a diffusion model to obtain a predicted image.
[0145] Optionally, based on the sample reference image, the image processing model performs foreground removal and background completion processing on the sample region in the first sample image to obtain a predicted image. This includes, in the case where the image processing model includes an encoder, a diffusion model, and a decoder, encoding the sample reference image by the encoder to obtain a first predicted image encoding, denoising the pure noise image based on the first predicted image encoding by the diffusion model to obtain a second predicted image encoding, and decoding the second predicted image encoding by the decoder to obtain the predicted image.
[0146] S404, perform quality evaluation on the predicted image to obtain a quality score for the predicted image.
[0147] S405 trains the image processing model based on quality scores.
[0148] In this disclosure, the image processing model can be trained with regard to the quality score of the predicted image. This can capture and correct semantic errors such as accidental deletion and accidental generation in the image processing model, significantly enhancing the image processing model's ability to identify and suppress "illusory content". This helps to improve the quality of the images generated by the image processing model, especially the credibility and quality performance in the subjective evaluation dimension.
[0149] Optionally, the predicted image is quality evaluated to obtain a quality score, including evaluating the quality of the predicted image using a multimodal large model to obtain a quality score.
[0150] The dimensions for quality assessment of predicted images include at least one of the following dimensions: Foreground elimination of quality dimension; Background completion quality dimensions; Aesthetic dimension.
[0151] It should be noted that the foreground removal quality dimension includes whether the foreground removal is thorough, the background completion quality dimension includes whether similar foreground elements are mistakenly generated and whether the generation is reasonable, and the aesthetic dimension includes whether the spliced area blends naturally, clarity, composition, etc.
[0152] Optionally, the image processing model is trained based on the quality score, including determining the loss function of the image processing model based on the difference information between the set score and the quality score, and training the image processing model based on the loss function.
[0153] Optionally, the image processing model is trained based on the quality score, including obtaining a second sample image after performing foreground removal and background completion processing on the first sample image, and training the image processing model based on the predicted image, the second sample image, and the quality score. Thus, the image processing model can be trained by comprehensively considering the predicted image, the second sample image, and the quality score.
[0154] Optionally, the image processing model is trained based on the predicted image, the second sample image, and the quality score, including determining a first loss function of the image processing model based on the difference information between the set score and the quality score, determining a second loss function of the image processing model based on the difference information between the predicted image and the second sample image, determining a total loss function of the image processing model based on the first loss function and the second loss function, and training the image processing model based on the total loss function.
[0155] The model training method provided in this disclosure involves obtaining a first sample image, wherein the first sample image includes sample regions to be eliminated. Based on a mask of the sample regions, the sample regions in the first sample image are masked to obtain a sample reference image. An image processing model then performs foreground elimination and background completion processing on the sample regions in the first sample image based on the sample reference image to obtain a predicted image. The predicted image is then evaluated for quality to obtain a quality score. Based on the quality score, the image processing model is trained. Therefore, by considering the quality score of the predicted image when training the image processing model, semantic-level errors such as accidental deletion and generation in the image processing model can be captured and corrected. This significantly enhances the image processing model's ability to identify and suppress "illusory content," helping to improve the quality of images generated by the image processing model, especially in terms of credibility and quality performance in the subjective evaluation dimension.
[0156] Figure 5 This is a schematic diagram of the structure of an image processing apparatus according to an exemplary embodiment.
[0157] Reference Figure 5 The image processing apparatus 500 of this embodiment includes: a determination module 501 and a processing module 502.
[0158] The determination module 501 is configured to determine a first region to be foreground removal from the original image in response to the first operation; The processing module 502 is configured to, when the original image includes a second region, perform foreground removal and background completion processing on the first region in the original image based on the image within the target region in the original image, so as to obtain a target image; The second region is associated with the content of the first region, and the target region includes image regions in the original image other than the first region and the second region.
[0159] In some possible implementations, the processing module 502 is further configured to: merge the mask of the first region and the mask of the second region to obtain a merged mask; and determine the image within the target region in the original image based on the merged mask; wherein the merged mask is used to identify the target region.
[0160] In some possible implementations, the processing module 502 is further configured to: crop the image in the first region of the original image and the image in the surrounding region of the first region to obtain a cropped image; perform masking processing on the first region in the cropped image to obtain a mask of the first region; and / or perform masking processing on the second region in the cropped image to obtain a mask of the second region.
[0161] In some possible implementations, the target region includes image regions other than the first and second regions in the cropped image.
[0162] In some possible implementations, the processing module 502 is further configured to: perform masking processing on the first region and the second region in the cropped image based on the merging mask to obtain a first reference image as the image within the target region in the cropped image; perform foreground removal and background completion processing on the first region in the cropped image based on the first reference image to obtain a first intermediate image; crop the image within the first region in the first intermediate image to obtain a first completed image; and cover the image within the first region in the original image with the first completed image to obtain the target image.
[0163] In some possible implementations, the processing module 502 is further configured to: perform foreground removal and background completion processing on the first region in the cropped image based on the first reference image using an image processing model to obtain the first intermediate image; In some possible implementations, the processing module 502 is further configured to: when the image processing model includes a diffusion model, perform denoising processing on the pure noise image based on the first reference image using the diffusion model to obtain the first intermediate image.
[0164] In some possible implementations, the processing module 502 is further configured to: perform masking processing on the first region in the original image to obtain a mask for the first region; and / or perform masking processing on the second region in the original image to obtain a mask for the second region.
[0165] In some possible implementations, the processing module 502 is further configured to: perform masking processing on the first region and the second region in the original image based on the merging mask to obtain a second reference image as the image within the target region in the original image; and perform foreground removal and background completion processing on the first region in the original image based on the second reference image to obtain the target image.
[0166] In some possible implementations, the processing module 502 is further configured to: perform foreground removal and background completion processing on the first region in the original image based on the second reference image to obtain a second intermediate image; crop the image in the first region of the second intermediate image to obtain a second completed image; and cover the image in the first region of the original image with the second completed image to obtain the target image.
[0167] In some possible implementations, the processing module 502 is further configured to: perform foreground removal and background completion processing on the first region in the original image based on the second reference image using an image processing model to obtain the target image; wherein, In some possible implementations, the processing module 502 is further configured to: when the image processing model includes a diffusion model, perform denoising processing on the pure noise image based on the second reference image using the diffusion model to obtain the target image.
[0168] In some possible implementations, the processing module 502 is further configured to: determine a candidate region from the original image that is associated with the content of the first region; use the candidate region as the second region; or, if the target region satisfies a set condition including that the target region is located in the surrounding area of the first region, determine a candidate region located in the surrounding area of the first region as the second region.
[0169] In some possible implementations, the processing module 502 is further configured to: when the set condition includes that the distance between the target region and the first region is less than a set threshold, obtain the distance between the second region and the first region; and select the candidate region whose distance is less than the set threshold as the second region.
[0170] In some possible implementations, the apparatus further includes: a training module configured to: acquire a first sample image; wherein the first sample image includes sample regions to be eliminated; perform masking processing on the sample regions in the first sample image based on a mask of the sample regions to obtain a sample reference image; perform foreground elimination and background completion processing on the sample regions in the first sample image based on the sample reference image using the image processing model to obtain a predicted image; perform quality evaluation on the predicted image to obtain a quality score for the predicted image; and train the image processing model based on the quality score.
[0171] In some possible implementations, the training module is further configured to: perform quality evaluation on the predicted image using a multimodal large model to obtain the quality score; The dimensions for quality evaluation of the predicted image include at least one of the following dimensions: Foreground elimination of quality dimension; Background completion quality dimensions; Aesthetic dimension.
[0172] In some possible implementations, the training module is further configured to: acquire a second sample image after performing foreground removal and background completion processing on the first sample image; and train the image processing model based on the predicted image, the second sample image, and the quality score.
[0173] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.
[0174] The image processing apparatus provided in the embodiments of this disclosure, in response to a first operation, determines a first region to be foreground removal from an original image. If the original image includes a second region, foreground removal and background completion processing are performed on the first region in the original image based on an image within a target region in the original image to obtain a target image. The second region is content-related to the first region, and the target region includes image regions in the original image other than the first and second regions. Therefore, when the original image includes a second region, foreground removal and background completion processing can be performed on the first region in the original image, taking into account the image within the target region, to obtain the target image. That is, the image within the target region in the original image is used as a reference for generating the first region to perform foreground removal and background completion processing on the first region. This can prevent interference from images within the second region on background completion of the first region, thereby avoiding the generation of content-related foreground elements in the first region, improving background completion quality, and optimizing the user experience.
[0175] To implement the above embodiments, this disclosure also proposes an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the steps of the image processing method provided in this disclosure.
[0176] Figure 6 This is a schematic diagram illustrating the structure of an electronic device according to an exemplary embodiment. For example, the electronic device 600 may be a vehicle, mobile phone, computer, digital broadcasting terminal, messaging device, game console, tablet device, medical device, fitness equipment, personal digital assistant, etc.
[0177] Reference Figure 6 The electronic device 600 may include one or more of the following components: processing component 602, memory 604, power component 606, multimedia component 608, audio component 610, input / output (I / O) interface 612, sensor component 614, and communication component 616.
[0178] Processing component 602 typically controls the overall operation of electronic device 600, such as operations associated with display, telephone calls, data communication, camera operation, and recording. Processing component 602 may include one or more processors 620 to execute instructions to complete all or part of the steps of the image processing method described above. Furthermore, processing component 602 may include one or more modules to facilitate interaction between processing component 602 and other components. For example, processing component 602 may include a multimedia module to facilitate interaction between multimedia component 608 and processing component 602.
[0179] Memory 604 is configured to store various types of data to support the operation of electronic device 600. Examples of this data include instructions for any application or method operating on electronic device 600, contact data, phonebook data, messages, pictures, videos, etc. Memory 604 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0180] Power component 606 provides power to various components of electronic device 600. Power component 606 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to electronic device 600.
[0181] Multimedia component 608 includes a screen that provides an output interface between electronic device 600 and user. In some embodiments, the screen may include a Liquid Crystal Display (LCD) and a Touch Panel (TP). If the screen includes a Touch Panel, the screen may be implemented as a touchscreen to receive input signals from the user. The Touch Panel includes one or more touch sensors to sense touches, swipes, and gestures on the Touch Panel. The touch sensors may sense not only the boundaries of touch or swipe actions but also the duration and pressure associated with the touch or swipe operation. In some embodiments, multimedia component 608 includes a front-facing camera and / or a rear-facing camera. When electronic device 600 is in an operating mode, such as a shooting mode or video mode, the front-facing camera and / or rear-facing camera may receive external multimedia data. Each front-facing camera and rear-facing camera may be a fixed optical lens system or have focal length and optical zoom capabilities.
[0182] Audio component 610 is configured to output and / or input audio signals. For example, audio component 610 includes a microphone (MIC) configured to receive external audio signals when electronic device 600 is in an operating mode, such as call mode, recording mode, and voice recognition mode. The received audio signals may be further stored in memory 604 or transmitted via communication component 616. In some embodiments, audio component 610 also includes a speaker for outputting audio signals.
[0183] I / O interface 612 provides an interface between processing component 602 and peripheral interface modules, such as keyboards, click wheels, buttons, etc. These buttons may include, but are not limited to, home buttons, volume buttons, start buttons, and lock buttons.
[0184] Sensor assembly 614 includes one or more sensors for providing state assessments of various aspects of electronic device 600. For example, sensor assembly 614 may detect the on / off state of electronic device 600, the relative positioning of components such as the display and keypad of electronic device 600, changes in position of electronic device 600 or a component of electronic device 600, the presence or absence of user contact with electronic device 600, orientation or acceleration / deceleration of electronic device 600, and temperature changes of electronic device 600. Sensor assembly 614 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. Sensor assembly 614 may also include an optical sensor, such as a complementary metal-oxide-semiconductor (CMOS) or charge-coupled device (CCD) image sensor, for use in imaging applications. In some embodiments, sensor assembly 614 may also include an accelerometer, gyroscope, magnetometer, pressure sensor, or temperature sensor.
[0185] Communication component 616 is configured to facilitate wired or wireless communication between electronic device 600 and other devices. Electronic device 600 can access wireless networks based on communication standards, such as WiFi, 4G, or 5G, or combinations thereof. In one exemplary embodiment, communication component 616 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In one exemplary embodiment, communication component 616 also includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module may be implemented based on Radio Frequency Identification (RFID), Infrared Data Association (IrDA), Ultra-Wideband (UWB), Bluetooth, and other technologies.
[0186] In an exemplary embodiment, the electronic device 600 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to perform the steps of the image processing method described above.
[0187] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 604 including instructions, which can be executed by a processor 620 of an electronic device 600 to complete the image processing method described above. For example, the non-transitory computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, a floppy disk, and an optical data storage device, etc.
[0188] To implement the above embodiments, this disclosure also proposes a computer-readable storage medium storing computer program instructions thereon, which, when executed by a processor, implement the steps of the image processing method provided in this disclosure.
[0189] Alternatively, the computer-readable storage medium may be ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device, etc.
[0190] To implement the above embodiments, this disclosure also proposes a chip including an interface circuit and a processing circuit coupled to each other. The interface circuit is used to input or output signals, and the processing circuit is configured to implement the steps of the image processing method provided in this disclosure.
[0191] Figure 7 This is a schematic diagram illustrating the structure of a chip according to an exemplary embodiment. See also... Figure 7 The diagram shown is a schematic representation of the structure of chip 700, but it is not limited to this.
[0192] Chip 700 includes processing circuit 701, which is configured to perform the steps of any of the above image processing methods.
[0193] In some embodiments, chip 700 further includes one or more interface circuits 702. Optionally, interface circuit 702 is connected to memory 703, and interface circuit 702 can be used to receive signals from memory 703 or other devices, and interface circuit 702 can be used to send signals to memory 703 or other devices. For example, interface circuit 702 can read instructions stored in memory 703 and send the instructions to processing circuit 701.
[0194] In some embodiments, the interface circuit 702 performs at least one of the communication steps such as sending and / or receiving in the above method, and the processing circuit 701 performs other steps.
[0195] In some embodiments, the terms interface circuit, interface, transceiver pin, transceiver, etc., can be used interchangeably.
[0196] In some embodiments, chip 700 further includes one or more memories 703 for storing instructions. Optionally, all or part of the memories 703 may be located outside of chip 700.
[0197] To implement the above embodiments, this disclosure also proposes a computer program product, including a computer program that, when executed by a processor, implements the steps of the image processing method provided in this disclosure.
[0198] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0199] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. An image processing method, characterized in that, include: In response to the first operation, a first region to be foreground removal is determined from the original image; If the original image includes a second region, based on the image within the target region in the original image, foreground removal and background completion processing are performed on the first region in the original image to obtain the target image; The second region is associated with the content of the first region, and the target region includes image regions in the original image other than the first region and the second region.
2. The method according to claim 1, characterized in that, The method further includes: The masks of the first region and the second region are merged to obtain a merged mask; Based on the merge mask, the image within the target region in the original image is determined; wherein, the merge mask is used to identify the target region.
3. The method according to claim 2, characterized in that, The method further includes: The image within the first region and the image within the surrounding region of the first region in the original image are cropped to obtain a cropped image; The first region in the cropped image is masked to obtain a mask for the first region; and / or the second region in the cropped image is masked to obtain a mask for the second region.
4. The method according to claim 3, characterized in that, The target region includes the image region other than the first region and the second region in the cropped image.
5. The method according to claim 3, characterized in that, The step of determining the image within the target region in the original image based on the merged mask includes: Based on the merging mask, the first region and the second region in the cropped image are masked to obtain a first reference image, which serves as the image within the target region of the cropped image. The step of performing foreground removal and background completion processing on the first region in the original image based on the image within the target region in the original image to obtain the target image includes: Based on the first reference image, foreground removal and background completion are performed on the first region in the cropped image to obtain a first intermediate image; The image within the first region of the first intermediate image is cropped to obtain the first completed image; The first completed image is overlaid on the image within the first region of the original image to obtain the target image.
6. The method according to claim 5, characterized in that, The step of performing foreground removal and background completion processing on the first region in the cropped image based on the first reference image to obtain a first intermediate image includes: Based on the first reference image, an image processing model is used to perform foreground removal and background completion processing on the first region in the cropped image to obtain the first intermediate image; wherein, The step of performing foreground removal and background completion processing on the first region in the cropped image based on the first reference image using an image processing model to obtain the first intermediate image includes: When the image processing model includes a diffusion model, the pure noise image is denoised based on the first reference image using the diffusion model to obtain the first intermediate image.
7. The method according to claim 2, characterized in that, The method further includes: The first region in the original image is masked to obtain a mask for the first region; and / or, The second region in the original image is masked to obtain a mask for the second region.
8. The method according to claim 7, characterized in that, The step of determining the image within the target region of the original image based on the merged mask includes: Based on the merging mask, the first region and the second region in the original image are masked to obtain a second reference image, which serves as the image within the target region in the original image; The step of performing foreground removal and background completion processing on the first region in the original image based on the image within the target region in the original image to obtain the target image includes: Based on the second reference image, foreground removal and background completion are performed on the first region in the original image to obtain the target image.
9. The method according to claim 8, characterized in that, The step of performing foreground removal and background completion processing on the first region in the original image based on the second reference image to obtain the target image includes: Based on the second reference image, foreground removal and background completion are performed on the first region in the original image to obtain a second intermediate image; The image in the first region of the second intermediate image is cropped to obtain the second completed image; The second completed image is used to overlay the image within the first region of the original image to obtain the target image.
10. The method according to claim 8, characterized in that, The step of performing foreground removal and background completion processing on the first region in the original image based on the second reference image to obtain the target image includes: Based on the second reference image, an image processing model is used to perform foreground removal and background completion processing on the first region in the original image to obtain the target image; wherein, The step of performing foreground removal and background completion processing on the first region in the original image based on the second reference image using an image processing model to obtain the target image includes: When the image processing model includes a diffusion model, the pure noise image is denoised based on the second reference image using the diffusion model to obtain the target image.
11. The method according to any one of claims 1-10, characterized in that, The method further includes: Determine candidate regions from the original image that are associated with the content of the first region; The candidate region is used as the second region; or, if the target region satisfies a set condition including that the target region is located in the area surrounding the first region, a candidate region located in the area surrounding the first region is determined as the second region.
12. The method according to claim 11, characterized in that, The step of determining a candidate region located in the surrounding area of the first region as the second region includes: If the set condition includes that the distance between the target area and the first area is less than a set threshold, the distance between the second area and the first area is obtained; Candidate regions whose distance is less than the set threshold are designated as the second region.
13. The method according to claim 6 or 10, characterized in that, The image processing model was trained using the following method: Obtain a first sample image; wherein the first sample image includes the sample region to be eliminated; Based on the mask of the sample region, the sample region in the first sample image is masked to obtain a sample reference image; Based on the sample reference image, the image processing model performs foreground removal and background completion processing on the sample region in the first sample image to obtain a predicted image. The predicted image is evaluated for quality to obtain a quality score. The image processing model is trained based on the quality score.
14. The method according to claim 13, characterized in that, The process of evaluating the quality of the predicted image to obtain a quality score includes: The quality score is obtained by evaluating the quality of the predicted image using a multimodal large model. The dimensions for quality evaluation of the predicted image include at least one of the following dimensions: Foreground elimination of quality dimension; Background completion quality dimensions; Aesthetic dimension.
15. The method according to claim 13, characterized in that, The step of training the image processing model based on the quality score includes: Obtain a second sample image after performing foreground removal and background completion processing on the first sample image; The image processing model is trained based on the predicted image, the second sample image, and the quality score.
16. An image processing apparatus, characterized in that, include: The determination module is configured to determine a first region to be foreground removal from the original image in response to a first operation; The processing module is configured to, when the original image includes a second region, perform foreground removal and background completion processing on the first region in the original image based on the image within the target region in the original image, to obtain a target image; The second region is associated with the content of the first region, and the target region includes image regions in the original image other than the first region and the second region.
17. The apparatus according to claim 16, characterized in that, The processing module is further configured to: The masks of the first region and the second region are merged to obtain a merged mask; Based on the merge mask, the image within the target region in the original image is determined; wherein, the merge mask is used to identify the target region.
18. The apparatus according to claim 17, characterized in that, The processing module is further configured to: The image within the first region and the image within the surrounding region of the first region in the original image are cropped to obtain a cropped image; The first region in the cropped image is masked to obtain the mask of the first region; And / or, mask the second region in the cropped image to obtain a mask for the second region.
19. An electronic device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the steps of the method according to any one of claims 1-15.
20. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When executed by a processor, the program instructions implement the steps of the method described in any one of claims 1-15.
21. A chip, characterized in that, The chip includes an interface circuit and a processing circuit coupled to each other. The interface circuit is used to input or output signals, and the processing circuit is configured to implement the steps of the method according to any one of claims 1-15.
22. A computer program product, characterized in that, Includes a computer program, which, when executed by a processor, implements the steps of the method according to any one of claims 1-15.