Image processing method and apparatus, electronic device and storage medium

By segmenting the initial image and using the pre-trained image processing model to generate matching backgrounds, the problem of poor segmentation traces and consistency during background replacement in the prior art is solved, and a higher quality image synthesis is achieved.

WO2025108419A1PCT designated stage expired Publication Date: 2025-05-30BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/133814
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-23
Filing Date
2024-11-22
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

When replacing the pictures in the prior art, commonly used image segmentation methods lead to obvious segmentation traces and poor background consistency in the synthetic image, which affects the image quality.

Method used

By obtaining the initial image and performing segmentation processing, the segmented image is obtained to indicate the target object, and then using the pre-trained image processing model combined with the segmented image to generate a matching image background, thereby creating a synthetic image to ensure that the target object does not interfere with the background visually.

Benefits of technology

It effectively reduces segmentation traces in the composite image, improves the consistency between the background and the target object, thereby improving the authenticity and visual perception of the image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024133814_30052025_PF_FP_ABST
    Figure CN2024133814_30052025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure provide an image processing method and apparatus, an electronic device and a storage medium. The method comprises: acquiring an initial image; processing the initial image to obtain a segmented image, wherein the segmented image is used for indicating a target object in the initial image; and obtaining a synthesized image by means of a pre-trained image processing model and the segmented image, wherein the image processing model is used for generating a matching image background for the target object, so that no visual physical interference occurs between the target object and the image background in the output synthesized image. By using the image generation capability of the image processing model, no visual physical interference occurs between the target object and the image background in the synthesized image, thereby improving the authenticity of the synthesized image, reducing the segmentation trace between the background and the target object, and improving the consistency between the background and the target object.
Need to check novelty before this filing date? Find Prior Art

Description

Image processing method, device, electronic device and storage medium

[0001] This application claims priority to Chinese Patent Application No. 202311575601.0 filed on November 23, 2023, and the contents of the above-mentioned Chinese patent application disclosure are hereby incorporated by reference in their entirety as a part of this application. Technical Field

[0002] Embodiments of the present disclosure relate to an image processing method, apparatus, electronic device, and storage medium. Background Art

[0003] Currently, through the special effects functions provided by various applications and platforms, the visual effect of replacing the background in a picture or video can be achieved, which increases the interest and visual expressiveness of the picture or video.

[0004] However, the existing solutions for background replacement of images are usually based on simple image segmentation, which results in obvious segmentation marks and poor background consistency in the generated synthetic images, affecting the image quality. Summary of the Invention

[0005] The embodiments of the present disclosure provide an image processing method, apparatus, electronic device, and storage medium to overcome the problems of obvious segmentation traces and poor background consistency.

[0006] In a first aspect, an embodiment of the present disclosure provides an image processing method, comprising:

[0007] An initial image is acquired and processed to obtain a segmented image, wherein the segmented image is used to indicate a target object in the initial image; a composite image is obtained by using a pre-trained image processing model and the segmented image, wherein the image processing model is used to generate a matching image background for the target object so that the target object in the output composite image does not visually physically interfere with the image background.

[0008] In a second aspect, an embodiment of the present disclosure provides an image processing device, including:

[0009] an acquisition module, configured to acquire an initial image and process the initial image to obtain a segmented image, wherein the segmented image is used to indicate a target object in the initial image;

[0010] A processing module is used to obtain a composite image using a pre-trained image processing model and the segmented image, wherein the image processing model is used to generate a matching image background for the target object so that the target object in the output composite image does not visually physically interfere with the image background.

[0011] In a third aspect, an embodiment of the present disclosure provides an electronic device, including: a processor and a memory;

[0012] The memory stores computer-executable instructions;

[0013] The processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the image processing method described in the first aspect and various possible designs of the first aspect.

[0014] In a fourth aspect, an embodiment of the present disclosure provides a computer-readable storage medium, in which computer execution instructions are stored. When a processor executes the computer execution instructions, the image processing method described in the first aspect and various possible designs of the first aspect is implemented.

[0015] In a fifth aspect, an embodiment of the present disclosure provides a computer program product, including a computer program, which, when executed by a processor, implements the image processing method described in the first aspect and various possible designs of the first aspect. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] In order to more clearly illustrate the embodiments of the present disclosure or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0017] FIG1 is a diagram of an application scenario of the image processing method provided by an embodiment of the present disclosure;

[0018] FIG2 is a flowchart of an image processing method according to an embodiment of the present disclosure;

[0019] FIG3 is a flowchart of a specific implementation of step S103 in the embodiment shown in FIG2 ;

[0020] FIG4 is a flowchart of a specific implementation method of step S1031 in the embodiment shown in FIG3 ;

[0021] FIG5 is a schematic diagram of a process for generating prompt word information provided by an embodiment of the present disclosure;

[0022] FIG6 is a second flow chart of the image processing method provided by an embodiment of the present disclosure;

[0023] FIG7 is a schematic diagram of an image processing model provided by an embodiment of the present disclosure;

[0024] FIG8 is a flowchart of a specific implementation of step S205 in the embodiment shown in FIG6 ;

[0025] FIG9 is a schematic diagram of a directional noise removal process provided by this embodiment;

[0026] FIG10 is a flowchart of a specific implementation of step S2051 in the embodiment shown in FIG8 ;

[0027] FIG11 is a structural block diagram of an image processing device provided by an embodiment of the present disclosure;

[0028] FIG12 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure;

[0029] FIG13 is a schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION

[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present disclosure without making any creative efforts shall fall within the scope of protection of the present disclosure.

[0031] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0032] The following explains the application scenarios of the embodiments of the present disclosure:

[0033] FIG1 is a diagram of an application scenario of the image processing method provided by an embodiment of the present disclosure. The image processing method provided by an embodiment of the present disclosure can be applied to applications with image editing and image generation functions. More specifically, it can be applied to application scenarios in which background replacement of videos and pictures is performed using special effects props within the application. The execution subject of this embodiment can be a terminal device running the above-mentioned application with image synthesis function, or a server that deploys the server corresponding to the above-mentioned application, or other electronic devices that perform similar functions. Referring to FIG1 , taking a terminal device as an example, after loading an initial image (such as a photo or video), the terminal device displays the image data in the interactive interface of the application. Thereafter, by triggering a special effects prop (component) set in the interactive interface, shown as component #1 in the figure, the terminal device processes the initial image through the method provided by the embodiment of the present disclosure and generates a composite image. The composite image retains the main object 10 in the initial image, such as the portrait shown in the figure, but the image background in the composite image is a second image background 12, which is different from the first image background 11 in the initial image, that is, the first image background 11 in the initial image is replaced by the second image background 12.

[0034] In the prior art, image background replacement for pictures and videos is achieved through simple image segmentation. For example, after segmenting the main object from the initial image, a target background image that meets the user's needs is obtained from a library. This segmented image is directly overlaid on the target background image to generate a composite image, completing the image background replacement. However, because the background image is a pre-generated fixed-style image, there are obvious segmentation traces between the main object segmented using this scheme and the background image, and there is visual physical interference. This results in poor authenticity of the generated composite image, affecting the visual perception of the image.

[0035] The embodiments of the present disclosure provide an image processing method to solve the above problems.

[0036] Referring to FIG2 , FIG2 is a flow chart of an image processing method according to an embodiment of the present disclosure. The method of this embodiment can be applied in a terminal device, and the image processing method includes:

[0037] Step S101: Acquire an initial image.

[0038] Step S102: Processing the initial image to obtain a segmented image, where the segmented image is used to indicate the target object in the initial image.

[0039] For example, referring to the application scenario diagram shown in FIG1 , a terminal device first loads an image or video through an application and, in response to user input through the application, loads the image or video as the initial image. Specifically, if the loaded image is a video, the initial image can be a frame from the video. Furthermore, the image or video can be stored locally on the terminal device or on a network external to the terminal device. The terminal device can obtain the initial image directly locally or by accessing an external server, as required. The terminal device then performs image segmentation on the obtained initial image, segmenting the target object therein to obtain a segmented image. The segmented image can be a transparent background image containing only the target object, or a mask image used to locate the target object, or a combination of the mask image and the initial image. The target object refers to an element object in the initial image, such as an object or a person. The initial image can include multiple element objects. The target object can be determined based on user instructions or pre-marked in the initial image. Using the marking information in the initial image, the terminal device can determine the target object from the multiple element objects and then perform segmentation to obtain a segmented image.

[0040] Among them, the specific implementation method of identifying and segmenting image elements in the image is the existing technology, which can be achieved through a segmentation model, a diffusion model, or by identifying the contours of element objects and then performing segmentation based on a script, which will not be described in detail here.

[0041] Step S103: Obtain a composite image through a pre-trained image processing model and the segmented image, wherein the image processing model is used to generate a matching image background for the target object so that there is no visual physical interference between the target object and the image background in the output composite image.

[0042] Exemplarily, after obtaining a segmented image, the segmented image is further processed using a pre-trained image processing model to generate a composite image. The segmented image is an image used to indicate the target object in the initial image, equivalent to an image containing only the target object without the image background. The segmented image can be a transparent background image, a mask image, or a combination of a mask image and the initial image. Pixel information about the target object can be obtained from the segmented image. Subsequently, the image processing model is used to generate an image background that matches the pixel information corresponding to the target object, thereby obtaining a composite image. Because the composite image is dynamically generated by the image processing model based on the target object, the trained image processing model can control the generation process of the composite image during the generation process, thereby ensuring that the composite image satisfies certain constraints, namely, that there is no visual physical interference between the target object in the composite image and the image background. More specifically, for example, by taking pictures in a real environment and marking the element objects in the pictures, training samples can be obtained, and the image processing model can be trained based on the training samples. After the image processing model is fully trained, the synthetic image generated can be the same as the picture in the real environment, and there is no visual physical interference between the target object and the image background. The specific training process will not be repeated here.

[0043] Specifically, multiple image processing models are trained on an image restoration (inpaiting) model using sample images of different image scenes and image styles, and each image processing model is used to generate an image background of a different image style and / or image scene. In one possible implementation, before obtaining a composite image through a pre-trained image processing model and a segmented image, the terminal device first obtains model scene information based on a user instruction. The model scene information is used to characterize the image style and / or image scene of the image background generated by the image processing model. The model scene information includes, for example, a scene type identifier. A corresponding scene model is obtained through the model scene information and determined as a target image processing model. Afterwards, subsequent processing steps are performed based on the target image processing model to generate a composite image, thereby achieving content control of the image background in the composite image.

[0044] Furthermore, exemplarily, the image processing model is a pre-trained generative model. In one possible implementation, the image processing model is a diffusion model, more specifically, it is trained based on an image restoration (inpainting) model. The synthetic image generated by the image processing model includes two parts: the target object and the image background, wherein the content of the image background is controlled by the image processing model. In one possible implementation, multiple image processing models are first trained on the inpainting model using sample images of different image scenes and image styles. Each scene model is used to generate an image background of a different image style and / or image scene. Afterwards, the terminal device obtains model scene information based on user instructions. The model scene information is used to characterize the image style and / or image scene of the image background generated by the image processing model. The model scene information includes, for example, a scene type identifier. Afterwards, the terminal device obtains a corresponding image processing model through the model scene information and determines it as the target image processing model. Then, subsequent processing steps are performed based on the target image processing model to generate a synthetic image, thereby achieving content control of the image background in the synthetic image.

[0045] Specifically, for example, a terminal device receives an input instruction and obtains model scene information based on the input instruction. For example, if the model scene information represents a "snow scene," the terminal device will obtain a target image processing model M1 corresponding to the "snow scene" to process the segmented image and obtain a corresponding composite image P1. The image content of the image background in the composite image is "snow." Similarly, if the model scene information represents a "starry sky scene," the terminal device will obtain a target image processing model M2 corresponding to the "starry sky scene" to process the segmented image and obtain a corresponding composite image P2. The image content of the image background in the composite image is "starry sky."

[0046] In another possible implementation, the terminal device can guide the image content of the composite image output by the image processing model through the prompt word information, thereby achieving control of the image background in the composite image. For example, as shown in FIG3 , the specific implementation of step S103 includes:

[0047] Step S1031: Obtain prompt word information, which is used to guide the image processing model to generate an image background with target image content.

[0048] Step S1032: Input the initial image, prompt word information and segmented image into the image processing model to generate a composite image.

[0049] For example, prompt information is used to guide the output content of a generative model. It typically consists of one or more text segments, each of which indicates one or more restriction requirements. The specific implementation of prompt information is related to the image processing model and is not limited here. It can be set according to specific circumstances.

[0050] Furthermore, the prompt word information can be input by the user through a terminal device. For example, within an application running on the terminal device, the prompt word is generated by receiving user-entered text via a text box. Alternatively, the prompt word information can be generated by the terminal device based on a preset configuration file, combined with user-entered information regarding image background requirements. The segmented image can be a mask image. The mask image and the initial image can be used to locate the target object in the initial image. The terminal device inputs the initial image, the prompt word information, and the segmented image into an image processing model. The image processing model controls the process of generating an image background that matches the target object indicated by the segmented image based on the prompt word information, thereby generating a composite image whose image background content meets the requirements of the prompt word information.

[0051] In this embodiment, the prompt word information is used to guide the output of the image processing model, so that the synthetic image generated by the image processing model is more accurate, better meets the personalized needs of users, and improves the diversity and visual expression of the synthetic image.

[0052] Furthermore, in another possible implementation, the prompt word information may be automatically generated based on the content of the segmented image. As shown in FIG4 , the specific implementation of step S1031 includes:

[0053] Step S1031A: Acquire shooting parameters of the segmented image, where the shooting parameters represent shooting angle characteristics and / or illumination characteristics of the target object in the segmented image.

[0054] Step S1031B: Obtain corresponding description text according to the shooting parameters.

[0055] Step S1031C: Generate prompt word information based on the description text.

[0056] For example, in one possible implementation, the terminal device can obtain the shooting parameters corresponding to the segmented image by performing image recognition on the segmented image, and the shooting parameters represent the shooting angle characteristics and / or lighting characteristics of the target object in the segmented image. The recognition of the shooting parameters can be achieved through a preset feature recognition model, and the process is not described here. In another possible implementation, the terminal device obtains the above-mentioned shooting parameters by reading the attribute information of the initial image corresponding to the segmented image. The shooting parameters can be identifiers, key-value pairs or matrices that represent the shooting angle characteristics and / or lighting characteristics. Furthermore, after obtaining the shooting parameters, they are converted into descriptive text according to the meaning they represent, and then prompt word information is generated based on the preset prompt word text. FIG5 is a schematic diagram of a process for generating prompt word information provided by an embodiment of the present disclosure. As shown in FIG5 , first, feature extraction is performed based on the segmented image to obtain shooting parameters, wherein the shooting parameters include: parameter F_1 = [1, 30], representing that the illumination angle of the target object is 30 degrees; parameter F_2 = [2, 0], representing that the illumination intensity of the target object in the segmented image is level 0 (for example, representing the highest level, that is, the surface of the target object is well-lit and the target object is bright). Afterwards, according to the meaning represented by the above shooting parameters, it is mapped to a description text. For example, parameter F_1 is mapped to description text #1: "The illumination angle is 30 degrees to the left"; parameter F_2 is mapped to description text #2: "The surface is well-lit". Finally, based on the above description text #1 and description text #2, combined with the preset prompt word generation model, prompt word information that limits the generated content of the image processing model is generated. In the synthetic image subsequently generated based on the prompt word information, the image background has similar shooting parameters to those of the target object. For example, the illumination angle corresponding to the image background is also 30 degrees to the left, and the surface is well illuminated, thus making the image background and the target object have consistent lighting characteristics. The process of controlling the shooting angle characteristics of the image background through shooting parameters is similar and will not be repeated here.

[0057] In the steps of this embodiment, by obtaining the shooting parameters of the segmented image, and forming prompt word information according to the shooting parameters, and then generating a synthetic image based on the prompt word information, the shooting angle characteristics and / or lighting characteristics of the image background in the synthetic image are controlled, so that the image background and the target object in the synthetic image have the same or similar shadow effects and visual angles, thereby improving the consistency between the image background and the target object and improving the image quality of the synthetic image.

[0058] In this embodiment, an initial image is acquired and processed to obtain a segmented image, which is used to indicate a target object in the initial image. A pre-trained image processing model and the segmented image are then used to obtain a composite image, wherein the image processing model is used to generate a matching image background for the target object so that there is no visual physical interference between the target object and the image background in the output composite image. The initial image is segmented to obtain a segmented image indicating the target object in the initial image. The image processing model is then used to generate a corresponding image background based on the segmented image, thereby obtaining a composite image containing the target object and the image background. The image generation capabilities of the image processing model are utilized to eliminate visual physical interference between the target object and the image background in the composite image, thereby improving the authenticity of the composite image, reducing segmentation artifacts between the background and the target object, and improving the consistency between the background and the target object.

[0059] Referring to FIG6 , FIG6 is a second flow chart of the image processing method provided by an embodiment of the present disclosure. Based on the embodiment shown in FIG2 , this embodiment further refines step S103 , and the image processing model includes an encoder unit, a decoder unit, and a diffusion unit. The image processing method includes:

[0060] Step S201: Acquire an initial image.

[0061] Step S202: Processing the initial image to obtain a segmented image, where the segmented image is used to indicate the target object in the initial image.

[0062] Step S203: Perform feature extraction on the initial image through the encoder unit to obtain a first feature map.

[0063] Step S204: adding noise to the first feature map through a diffusion unit to generate a noisy feature map.

[0064] Step S205: performing directional denoising on the noisy feature map through a diffusion unit, and in the directional denoising process, weightedly superimposing a target area of ​​the first feature map to obtain a denoised feature map, wherein the target area is an image area determined based on the segmented image.

[0065] Step S206: Decode the denoised feature map through a decoder unit to generate a composite image.

[0066] Exemplarily, Figure 7 is a schematic diagram of an image processing model provided by an embodiment of the present disclosure. The above process is introduced below in conjunction with Figure 7. The image processing model includes an encoder unit, a decoder unit and a diffusion unit, wherein the encoder unit includes at least one image encoder (Encoder). More specifically, the image encoder is a variational autoencoder (VAE Encoder), and the image encoder is used to extract features (embedding) on ​​the initial image to obtain a first feature map.

[0067] Exemplarily, based on different implementations of segmenting the image, when the segmented image is a mask image, the terminal device performs feature extraction on the initial image to obtain a first feature map representing the image features of the initial image. When the segmented image is a transparent image, or a combination of a mask image and the initial image, since the segmented image contains input information required for subsequent calculation steps, another implementation of step S203 is to perform feature extraction on the segmented image using an encoder unit to obtain the first feature map. The specific implementation process is similar and will not be repeated here.

[0068] Then, noise is added to the first feature map using a diffusion unit, wherein the diffusion unit can be a diffusion model including one or more orderly arranged noise adding units for adding Gaussian noise to the first feature map once or step by step, and after adding the Gaussian noise, a (nearly) completely random noise map is finally obtained, i.e., the noisy feature map. In one possible implementation, the noise intensity in the noisy feature map can be set based on a preset configuration parameter or user instruction, or based on the image content of the initial image, thereby achieving dynamic adjustment of the noise intensity in the noisy feature map. Afterwards, the noisy feature map after adding noise is subjected to directional denoising by the diffusion unit. In the directional denoising process, the target area of ​​the first feature map, that is, the area where the target object indicated by the segmented image is located, is weightedly superimposed to obtain a denoised feature map. In this case, directional denoising of the noisy feature map refers to removing the noise components in the noisy feature map according to a certain rule and retaining some content, thereby achieving the expression of the image background. The process of performing step-by-step denoising on the noisy feature map is the process of generating the image background. On the other hand, weighted superposition of the target area of ​​the first feature map refers to superimposing the information of the target object in the initial image onto the denoised feature map and affecting the generation process of the image background, thereby achieving the fusion of the target object and the image background. Finally, after one or more levels of directional denoising and superimposing the target area of ​​the first feature map, a denoised feature map containing both the target object and the image background is obtained. In this case, directional denoising refers to controlling the image content represented by the denoised feature map, such as controlling the image style, content, and composition of the image background. This can be achieved by inputting additional prompt words or images into the model corresponding to the diffusion unit (the denoising unit within the diffusion unit). Finally, the denoised feature map is decoded by the decoder unit to restore the corresponding image, i.e., the synthesized image.

[0069] Furthermore, in one possible implementation, step S205, i.e., the process of performing directional denoising on the noisy feature map, can be implemented by a denoising unit within the diffusion unit. For example, the diffusion unit includes at least one denoising unit arranged in an orderly manner, as shown in FIG8 . The specific implementation process of step S205 includes:

[0070] Step S2051: Obtain the sequence value of the current de-noising unit.

[0071] Step S2052: If the sequence value is less than the sequence threshold, the current de-noising unit is used to perform directionally de-noising on the current noisy feature map to obtain an intermediate de-noised feature map output by the current de-noising unit, and then the current de-noising unit is updated to the next de-noising unit, the current noisy feature map is updated to the intermediate de-noised feature map, and the process returns to step S2051.

[0072] Step S2053: If the sequence value is greater than or equal to the sequence threshold and less than the last sequence corresponding to the last denoising unit, then the target region of the first feature map is weighted and superimposed on the current noisy feature map to obtain the superimposed feature map corresponding to the current denoising unit. Then, directional denoising is performed on the superimposed feature map to obtain the intermediate denoised feature map output by the current denoising unit. Next, the current denoising unit is updated to the next denoising unit, the current noisy feature map is updated to the intermediate denoised feature map, and return to step S2051.

[0073] Step S2054: If the sequence value is equal to the last sequence corresponding to the last denoising unit, then the target region of the first feature map is weighted and superimposed on the current noisy feature map to obtain the superimposed feature map corresponding to the current denoising unit. Then, directional denoising is performed on the superimposed feature map to obtain the denoised feature map.

[0074] FIG. 9 is a schematic diagram of a directional denoising process provided by this embodiment. The above steps will be introduced in detail below with reference to FIG. 9. Exemplarily, as shown in FIG. 9, the number of denoising units in the diffusion unit is N, and N is an integer greater than 1. Starting from the first denoising unit (Node_1), first obtain the sequence value of the current denoising unit, that is, Node_1, for example, it is 1. Then, determine whether the sequence value of Node_1 is less than the sequence threshold. Here, the sequence threshold is, for example, half of the number N of all denoising units, for example, it is M (M is greater than or equal to 1). Since the sequence value of Node_1 is less than the sequence threshold (1 < M), then Node_1 is used to perform directional denoising on the input input_1 of Node_1, that is, the current noisy feature map, to obtain the intermediate denoised feature map output by Node_1, that is, output_1. For the first denoising unit Node_1, its corresponding current noisy feature map (that is, input_1) is the noisy feature map F0 generated in the previous step. Then, the current denoising unit is set as the next denoising unit, and the current noisy feature map is set as the intermediate denoised feature map. That is, the current denoising unit is set as Node_2 (current denoising unit = Node_2), and the input input_2 of Node_2 is set as the output output_1 of Node_1 (input_2 = output_1), and return to step S2051, repeating the above steps.

[0075] In a possible implementation manner, when the sequence value is less than i, set the target weighting coefficient = 0 (shown as w_coef = 0 in the figure), that is, use 0 as the weighting coefficient corresponding to the target region of the first feature map (shown as feature map F in the figure), and calculate the superposition of the target region of the first feature map and the current noisy feature map, and perform directional denoising on the superposition result (superimposed feature map) to obtain the intermediate denoised feature map. Since this superimposed feature map is the same as the current noisy feature map, it is also equivalent to directly performing directional denoising on the superimposed feature map.

[0076] Furthermore, when the above steps are executed until the current denoising unit is Node_i, the sequence value of Node_i is i. When i is greater than or equal to M and less than N, the target area of ​​the first feature map (shown as F in the figure) is first weightedly superimposed (with a weighting coefficient of w_coef=c1) on the current noisy feature map, that is, the input input_i of the current denoising unit Node_i, to obtain the superimposed feature map P_i corresponding to the current denoising unit, and then the superimposed feature map P_i is directionally de-noised to obtain the intermediate denoised feature map output by the current denoising unit, that is, output_i. Afterwards, similar to the previous step, the current denoising unit is set to the next denoising unit, the current noisy feature map is set to the intermediate denoising feature map, that is, the current denoising unit is set to Node_i+1 (current denoising unit=Node_i+1), the input input_i+1 of Node_i+1 is set to output_i output by Node_i (input_i+1=output_i), and the process returns to step S2051.

[0077] When the above steps are executed to the last de-noising unit, that is, the last de-noising unit in the one or more de-noising units arranged in order in the diffusion unit, that is, the current de-noising unit is Node_N, first the target area of ​​the first feature map (shown as F in the figure) is weighted superimposed (the weighting coefficient is w_coef = cN) to the current noisy feature map, that is, the input input_N of the current de-noising unit Node_N, to obtain the superimposed feature map P_N corresponding to the current de-noising unit, and then the superimposed feature map P_N is directionally de-noised to obtain the intermediate de-noising feature map output by the current de-noising unit, that is, output_N. After that, output_N is determined as the final de-noising feature map.

[0078] Among them, the above-mentioned denoising unit before Node_i directly performs directional denoising, and for the denoising unit after Node_i, the target area of ​​the first feature map is weighted and superimposed on the current noisy feature map to obtain the superimposed feature map and then perform denoising. In the specific implementation process, the two different implementation methods mentioned above can be achieved by adjusting the weighting coefficient during weighted superposition. Specifically, when the sequence value of the denoising unit is less than i, the target weighting coefficient is set to 0, that is, 0 is used as the weighting coefficient, the target area of ​​the first feature map is weighted, and it is superimposed on the current noisy feature map to obtain a superimposed feature map, which is the same as the current noisy feature map.

[0079] By performing cyclic denoising through the above-mentioned multiple denoising units, the diffusion unit will eventually output a denoised feature map. At the same time, in the above-mentioned denoising process, when the denoising process reaches a certain level, the target area of ​​the first feature map is weighted and superimposed on the current denoised feature map, and the next denoising output is generated through the current denoised feature map, thereby achieving control over the denoising result (i.e., the image background). In the final denoised feature map, the target object and the image background can match each other, achieving an effect of no visual physical interference, and at the same time, the target object can be strengthened. In the final denoised feature map, the target object can be included in a complete and accurate manner. This avoids the problem of target object deformation in the process of generating a synthetic image and achieves accurate restoration of the target object in the synthetic image.

[0080] Furthermore, in the step of weighting and superimposing the target area of ​​the first feature map onto the current noisy feature map to obtain the superimposed feature map corresponding to the current de-noising unit, the target area of ​​the first feature map and the current noisy feature map are superimposed using a target weighting coefficient. The target weighting coefficient can be manually set by the user or determined by the terminal device based on the segmentation confidence corresponding to the segmented image. In one possible implementation, as shown in FIG10 , the specific implementation method of obtaining the target weighting coefficient includes:

[0081] Step S2051A: Obtain the segmentation confidence corresponding to the segmented image, where the segmentation confidence represents the accuracy of segmentation of the target object.

[0082] Step S2051B: Obtain a target weighting coefficient according to the segmentation confidence, wherein the segmentation confidence is positively correlated with the target weighting coefficient.

[0083] For example, the target weighting coefficient represents the weighted weight of the first feature map when the first feature map and the noisy feature map are weighted and superimposed. The segmentation confidence is used to represent the accuracy of segmenting the target object. When the image content of the initial image is complex and the elements in the image are numerous and overlapping, the difficulty of segmenting the target object in the initial image increases, thereby reducing the accuracy of the target object segmentation. The segmentation confidence can be used to evaluate the accuracy of the above segmentation results. The segmentation confidence can be information output by the segmentation model. The specific implementation method is not repeated here. In this embodiment, the corresponding target weighting information is obtained through the segmentation confidence. The greater the segmentation confidence, the more accurate the target object (contour) in the segmented image. In this case, the target weighting coefficient is increased, that is, the weighted weight of the first feature map is increased, thereby improving the accuracy of the target object in the synthesized image; conversely, the smaller the segmentation confidence, the less accurate the target object (contour) in the segmented image. In this case, the target weighting coefficient is reduced, thereby reducing the influence of the blank area in the segmented image (that is, the image background that has not been deleted in the initial image), reducing the prominence of the blank area in the synthesized image, and improving the image quality.

[0084] Among them, in one possible implementation, the denoising unit is implemented based on the Unet network. The denoising unit in the trained diffusion unit can realize the denoising process of the noisy feature map. The specific implementation principle will not be repeated here.

[0085] Step S207: Based on the segmented image, the image texture corresponding to the target object is superimposed on the synthesized image to obtain a first post-processed image.

[0086] For example, after obtaining a composite image, in order to further improve the clarity and accuracy of the target object in the composite image, the image texture corresponding to the target object can be superimposed on the composite image by segmenting the image. For example, the segmented image includes a mask image. The pixel data corresponding to (the area where) the target object is located is determined by using the mask image and the initial image. The pixel data is superimposed on the corresponding position in the composite image, thereby achieving the superposition of the image texture corresponding to the target object, so that the target object in the composite image has an image texture similar to or identical to that of the target object in the initial image, thereby improving the clarity and visual effect of the target image. In one possible implementation, when superimposing the image texture corresponding to the target object on the composite image, the center point (e.g., the geometric center point or the center point of the graphic) of the target object can be determined based on the segmented image. Then, based on the distance between the pixel point of the target object and the center point (hereinafter referred to as the center distance), a fuzzy coefficient related to the center distance is determined, wherein the smaller the center distance of the pixel point (i.e., the closer to the center point), the smaller the fuzzy coefficient, and the larger the center distance of the pixel point (i.e., the farther away from the center point), the larger the fuzzy coefficient. This achieves the effect of clear image texture in the center area of ​​the target object and blurred image texture in the edge area of ​​the target object, reduces the sense of separation between the target object and the image background in the optimized synthetic image (i.e., the first post-processed image), and improves visual authenticity.

[0087] The mapping relationship between the fuzzy coefficients related to the center distance can be set based on needs, and can be a linear mapping or a nonlinear mapping, which is not specifically limited here.

[0088] Step S208: Processing the synthesized image or the first post-processed image through a generative adversarial network model to obtain a super-resolution synthesized image, wherein the image resolution of the super-resolution synthesized image is greater than the image resolution of the synthesized image or the first post-processed image.

[0089] Furthermore, after obtaining the above-mentioned composite image or the first post-processed image, the composite image or the first post-processed image can be further subjected to super-resolution processing, for example, by processing the composite image or the first post-processed image through a generative adversarial network (GAN) model, thereby obtaining a super-resolution composite image with higher image resolution, thereby further improving the resolution of the target object or the image background. Since the image background in the contract image obtained in this embodiment is based on the content generated by the model, there may be a problem that the image resolution of the image background part is inconsistent with the image resolution of the target object part. In the steps of this embodiment, the composite image or the first post-processed image is processed through a generative adversarial network model to obtain a super-resolution composite image, so as to improve the image resolution of the image background and the target object in the composite image or the first post-processed image to the same level, thereby avoiding the sense of fragmentation caused by the inconsistent resolution of the two, and improving the authenticity and visual expression of the image.

[0090] In this embodiment, the implementation of step S201-step S202 is the same as the implementation of step S101-step S102 in the embodiment shown in FIG. 2 of the present disclosure, and will not be described in detail here.

[0091] Corresponding to the image processing method of the above embodiment, FIG11 is a block diagram of the structure of the image processing device provided by the embodiment of the present disclosure. For the sake of convenience, only the parts related to the embodiment of the present disclosure are shown. Referring to FIG11, the image processing device 3 includes:

[0092] An acquisition module 31 is used to acquire an initial image and process the initial image to obtain a segmented image, where the segmented image is used to indicate a target object in the initial image;

[0093] The processing module 32 is used to obtain a composite image through a pre-trained image processing model and a segmented image, wherein the image processing model is used to generate a matching image background for the target object so that the target object in the output composite image does not have visual physical interference with the image background.

[0094] In one embodiment of the present disclosure, the image processing model includes an encoder unit, a decoder unit and a diffusion unit; the processing module 32 is specifically used to: extract features from the initial image through the encoder unit to obtain a first feature map; add noise to the first feature map through the diffusion unit to generate a noisy feature map; perform directionally denoising on the noisy feature map through the diffusion unit, and in the process of directionally denoising, weightedly superimpose the target area of ​​the first feature map to obtain a denoised feature map, wherein the target area is an image area determined based on the segmented image; and decode the denoised feature map through the decoder unit to generate a composite image.

[0095] In one embodiment of the present disclosure, the diffusion unit includes at least one de-noising unit arranged in an orderly manner. When the processing module 32 performs directional de-noising on the noisy feature map through the diffusion unit and, in the directional de-noising process, weightedly superimposes the target area of ​​the first feature map to obtain the de-noised feature map, it is specifically used to: for each de-noising unit, sequentially perform the following steps: obtain the sequence value of the current de-noising unit; if the sequence value is less than the sequence threshold, use the current de-noising unit to perform directional de-noising on the current noisy feature map to obtain the intermediate de-noised feature map output by the current de-noising unit, then set the current de-noising unit as the next de-noising unit, set the current noisy feature map as the intermediate de-noised feature map, and return to the step of obtaining the sequence value of the current de-noising unit; if the sequence value is less than the sequence threshold, If the value is greater than or equal to the sequence threshold and less than the last sequence corresponding to the last de-noising unit, the target area of ​​the first feature map is weightedly superimposed on the current noisy feature map to obtain the superimposed feature map corresponding to the current de-noising unit, and then the superimposed feature map is directionally de-noised to obtain the intermediate de-noised feature map output by the current de-noising unit. The current de-noising unit is then set as the next de-noising unit, the current noisy feature map is set as the intermediate de-noised feature map, and the step of obtaining the sequence value of the current de-noising unit is returned to execute; if the sequence value is equal to the last sequence corresponding to the last de-noising unit, the target area of ​​the first feature map is weightedly superimposed on the current noisy feature map to obtain the superimposed feature map corresponding to the current de-noising unit, and then the superimposed feature map is directionally de-noised to obtain the de-noised feature map.

[0096] In one embodiment of the present disclosure, the processing module 32 is further used to: obtain the segmentation confidence corresponding to the segmented image, the segmentation confidence characterizing the accuracy of the segmentation of the target object; obtain the target weighting coefficient according to the segmentation confidence; when the processing module 32 weightedly superimposes the target area of ​​the first feature map on the current noisy feature map to obtain the superimposed feature map corresponding to the current de-noising unit, it is specifically used to: based on the target weighting coefficient, weightedly superimpose the target area of ​​the first feature map on the current noisy feature map to obtain the superimposed feature map corresponding to the current de-noising unit.

[0097] In one embodiment of the present disclosure, the processing module 32 is specifically used to: obtain prompt word information, which is used to guide the image processing model to generate an image background with target image content; input the initial image, prompt word information and segmented image into the image processing model to generate a composite image.

[0098] In one embodiment of the present disclosure, when obtaining prompt word information, the processing module 32 is specifically used to: obtain the shooting parameters of the segmented image, the shooting parameters representing the shooting angle characteristics and / or lighting characteristics of the target object in the segmented image; obtain the corresponding description text based on the shooting parameters; and generate prompt word information based on the description text.

[0099] In one embodiment of the present disclosure, the acquisition module 31 is further used to: acquire model scene information, which is used to characterize the image style and / or image scene of the image background generated by the image processing model; obtain the target image processing model based on the model scene information; the processing module 32 is specifically used to: process the segmented image through the target image processing model to obtain a composite image.

[0100] In one embodiment of the present disclosure, the segmented image includes a first segmented image or a second segmented image, wherein the first segmented image is a transparent background image containing the target object; and the second segmented image is a mask image for locating the target object.

[0101] In one embodiment of the present disclosure, after obtaining the composite image, the processing module 32 is further used for at least one of the following: based on the segmented image, superimposing the image texture corresponding to the target object into the composite image; processing the composite image through a generative adversarial network model to obtain a super-resolution composite image, wherein the image resolution of the super-resolution composite image is greater than the image resolution of the composite image.

[0102] The acquisition module 31 and the processing module 32 are connected. The image processing device 3 provided in this embodiment can implement the technical solution of the above method embodiment, and its implementation principle and technical effect are similar, which will not be described in detail in this embodiment.

[0103] FIG12 is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. As shown in FIG12 , the electronic device 4 includes:

[0104] A processor 41, and a memory 42 communicatively connected to the processor 41;

[0105] Memory 42 stores computer-executable instructions;

[0106] The processor 41 executes the computer-executable instructions stored in the memory 42 to implement the image processing method in the embodiments shown in FIG. 2 to FIG. 10 .

[0107] Optionally, the processor 41 and the memory 42 are connected via a bus 43 .

[0108] The relevant explanations can be understood by referring to the relevant descriptions and effects corresponding to the steps in the embodiments corresponding to Figures 2 to 10, and no further details will be given here.

[0109] An embodiment of the present disclosure provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, they are used to implement the image processing method provided in any of the embodiments corresponding to Figures 2 to 10 of the present disclosure.

[0110] In order to implement the above embodiment, the embodiment of the present disclosure further provides an electronic device.

[0111] Referring to FIG13 , a schematic diagram of the structure of an electronic device 900 suitable for implementing an embodiment of the present disclosure is shown. The electronic device 900 may be a terminal device or a server. The terminal device may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, personal digital assistants (PDAs), tablet computers (Portable Android Devices, PADs), portable multimedia players (PMPs), vehicle-mounted terminals (e.g., vehicle-mounted navigation terminals), and fixed terminals such as digital TVs and desktop computers. The electronic device shown in FIG13 is merely an example and should not limit the functionality and scope of use of the embodiments of the present disclosure.

[0112] As shown in FIG13 , the electronic device 900 may include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. Various programs and data required for the operation of the electronic device 900 are also stored in the RAM 903. The processing device 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0113] Typically, the following devices can be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 can allow the electronic device 900 to communicate with other devices wirelessly or by wire to exchange data. Although FIG13 shows an electronic device 900 with various devices, it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or have alternatively.

[0114] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network via the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above-mentioned functions defined in the method of the embodiment of the present disclosure are performed.

[0115] It should be noted that the computer-readable medium mentioned above in the present disclosure may be a computer-readable signal medium or a computer-readable storage medium, or any combination of the two. A computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, a computer-readable storage medium may be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device, or component. In the present disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium may also be any computer-readable medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted using any suitable medium, including but not limited to wires, optical cables, RF (radio frequency), etc., or any suitable combination thereof.

[0116] The computer-readable medium may be included in the electronic device, or may exist independently without being incorporated into the electronic device.

[0117] The computer-readable medium carries one or more programs. When the one or more programs are executed by the electronic device, the electronic device executes the method shown in the above embodiment.

[0118] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer through any type of network, including a Local Area Network (LAN) or a Wide Area Network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet Service Provider).

[0119] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present disclosure. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the module, program segment, or a part of code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in a different order than that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of the boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0120] The units involved in the embodiments described in this disclosure may be implemented in software or hardware. In some cases, the name of a unit does not limit the unit itself. For example, the first acquisition unit may also be described as a "unit for acquiring at least two Internet Protocol addresses."

[0121] The functions described above herein may be performed, at least in part, by one or more hardware logic components. For example, and without limitation, exemplary types of hardware logic components that may be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chip (SOCs), complex programmable logic devices (CPLDs), and the like.

[0122] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0123] In a first aspect, according to one or more embodiments of the present disclosure, there is provided an image processing method, comprising:

[0124] An initial image is acquired and processed to obtain a segmented image, wherein the segmented image is used to indicate a target object in the initial image; a composite image is obtained by using a pre-trained image processing model and the segmented image, wherein the image processing model is used to generate a matching image background for the target object so that the target object in the output composite image does not visually physically interfere with the image background.

[0125] According to one or more embodiments of the present disclosure, the image processing model includes an encoder unit, a decoder unit and a diffusion unit; a composite image is obtained by using a pre-trained image processing model and the segmented image, including: extracting features from the initial image through the encoder unit to obtain a first feature map; adding noise to the first feature map through the diffusion unit to generate a noisy feature map; weightedly superimposing a target area of ​​the first feature map onto the noisy feature map through the diffusion unit to obtain superimposed features, and directionally denoising the superimposed features to obtain a denoised feature map, wherein the target area is an image area determined based on the segmented image; decoding the denoised feature map through the decoder unit to generate the composite image.

[0126] According to one or more embodiments of the present disclosure, the diffusion unit includes at least one denoising unit arranged in an orderly manner, and the target area of ​​the first feature map is weightedly superimposed on the noisy feature map through the diffusion unit to obtain a superimposed feature, and the superimposed feature is directionally denoised to obtain a denoised feature map, including: for each of the denoising units, the following steps are cyclically performed: based on the target weighting coefficient, the target area of ​​the first feature map and the noisy feature map are weightedly superimposed to obtain the target superimposed feature corresponding to the current denoising unit; using the current denoising unit to directionally denoise the current noisy feature map to obtain an intermediate denoising feature map corresponding to the current denoising unit; if the current denoising unit is not the last denoising unit, the noisy feature map is updated to the intermediate denoising feature map, and the current denoising unit is updated to the next denoising unit; if the current denoising unit is the last denoising unit, the denoising feature map is generated based on the intermediate denoising feature map.

[0127] According to one or more embodiments of the present disclosure, the method further includes: obtaining a segmentation confidence corresponding to the segmented image, the segmentation confidence representing the accuracy of segmentation of the target object; and obtaining the target weighting coefficient according to the segmentation confidence.

[0128] According to one or more embodiments of the present disclosure, a composite image is obtained by using a pre-trained image processing model and the segmented image, including: obtaining prompt word information, wherein the prompt word information is used to guide the image processing model to generate an image background with target image content; and inputting the initial image, the prompt word information, and the segmented image into the image processing model to generate the composite image.

[0129] According to one or more embodiments of the present disclosure, obtaining prompt word information includes: obtaining shooting parameters of the segmented image, wherein the shooting parameters represent shooting angle characteristics and / or lighting characteristics of the target object in the segmented image; obtaining corresponding description text based on the shooting parameters; and generating the prompt word information based on the description text.

[0130] According to one or more embodiments of the present disclosure, the method further includes: obtaining model scene information, wherein the model scene information is used to characterize the image style and / or image scene of the image background generated by the image processing model; obtaining a target image processing model based on the model scene information; and obtaining a composite image through the pre-trained image processing model and the segmented image, including: processing the segmented image through the target image processing model to obtain a composite image.

[0131] According to one or more embodiments of the present disclosure, the segmented image includes a first segmented image or a second segmented image, wherein the first segmented image is a transparent background image containing the target object; and the second segmented image is a mask image for locating the target object.

[0132] According to one or more embodiments of the present disclosure, after obtaining the composite image, at least one of the following items is further included: based on the segmented image, superimposing the image texture corresponding to the target object into the composite image; processing the composite image through a generative adversarial network model to obtain a super-resolution composite image, wherein the image resolution of the super-resolution composite image is greater than the image resolution of the composite image.

[0133] In a second aspect, according to one or more embodiments of the present disclosure, there is provided an image processing apparatus, comprising:

[0134] an acquisition module, configured to acquire an initial image and process the initial image to obtain a segmented image, wherein the segmented image is used to indicate a target object in the initial image;

[0135] A processing module is used to obtain a composite image using a pre-trained image processing model and the segmented image, wherein the image processing model is used to generate a matching image background for the target object so that the target object in the output composite image does not visually physically interfere with the image background.

[0136] According to one or more embodiments of the present disclosure, the image processing model includes an encoder unit, a decoder unit and a diffusion unit; the processing module is specifically used to: perform feature extraction on the initial image through the encoder unit to obtain a first feature map; add noise to the first feature map through the diffusion unit to generate a noisy feature map; weightedly superimpose the target area of ​​the first feature map onto the noisy feature map through the diffusion unit to obtain superimposed features, and perform directionally denoising on the superimposed features to obtain a denoised feature map, wherein the target area is an image area determined based on the segmented image; and decode the denoised feature map through the decoder unit to generate the composite image.

[0137] According to one or more embodiments of the present disclosure, the diffusion unit includes at least one denoising unit arranged in an orderly manner, and the processing module weightedly superimposes the target area of ​​the first feature map onto the noisy feature map through the diffusion unit to obtain a superimposed feature, and performs directionally denoising on the superimposed feature to obtain a denoised feature map, wherein when the target area is an image area determined based on the segmented image, the processing module is specifically used to: for each of the denoising units, cyclically execute the following steps: based on a target weighting coefficient, weightedly superimpose the target area of ​​the first feature map and the noisy feature map to obtain a target superimposed feature corresponding to the current denoising unit; use the current denoising unit to directionally denoise the current noisy feature map to obtain an intermediate denoising feature map corresponding to the current denoising unit; if the current denoising unit is not the last denoising unit, update the noisy feature map to the intermediate denoising feature map, and update the current denoising unit to the next denoising unit; if the current denoising unit is the last denoising unit, generate the denoising feature map based on the intermediate denoising feature map.

[0138] According to one or more embodiments of the present disclosure, the processing module is further used to: obtain a segmentation confidence corresponding to the segmented image, where the segmentation confidence represents the accuracy of segmentation of the target object; and obtain the target weighting coefficient based on the segmentation confidence.

[0139] According to one or more embodiments of the present disclosure, the processing module is specifically used to: obtain prompt word information, where the prompt word information is used to guide the image processing model to generate an image background with target image content; and input the initial image, the prompt word information and the segmented image into the image processing model to generate the composite image.

[0140] According to one or more embodiments of the present disclosure, when obtaining prompt word information, the processing module is specifically used to: obtain shooting parameters of the segmented image, wherein the shooting parameters represent shooting angle characteristics and / or lighting characteristics of the target object in the segmented image; obtain corresponding description text based on the shooting parameters; and generate the prompt word information based on the description text.

[0141] According to one or more embodiments of the present disclosure, the acquisition module is further used to: acquire model scene information, wherein the model scene information is used to characterize the image style and / or image scene of the image background generated by the image processing model; obtain a target image processing model based on the model scene information; and the processing module is specifically used to: process the segmented image through the target image processing model to obtain a composite image.

[0142] According to one or more embodiments of the present disclosure, the segmented image includes a first segmented image or a second segmented image, wherein the first segmented image is a transparent background image containing the target object; and the second segmented image is a mask image for locating the target object.

[0143] According to one or more embodiments of the present disclosure, after the composite image is obtained, the processing module is further used for at least one of the following: based on the segmented image, superimposing the image texture corresponding to the target object into the composite image; processing the composite image through a generative adversarial network model to obtain a super-resolution composite image, wherein the image resolution of the super-resolution composite image is greater than the image resolution of the composite image.

[0144] In a third aspect, according to one or more embodiments of the present disclosure, there is provided an electronic device, comprising: at least one processor and a memory;

[0145] The memory stores computer-executable instructions;

[0146] The at least one processor executes the computer-executable instructions stored in the memory, so that the at least one processor performs the image processing method described in the first aspect and various possible designs of the first aspect.

[0147] In a fourth aspect, according to one or more embodiments of the present disclosure, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the image processing method described in the first aspect and various possible designs of the first aspect is implemented.

[0148] In a fifth aspect, according to one or more embodiments of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the image processing method as described in the first aspect and various possible designs of the first aspect.

[0149] The above description is merely a preferred embodiment of the present disclosure and an illustration of the technical principles employed. Those skilled in the art should understand that the scope of disclosure involved in the present disclosure is not limited to the technical solutions formed by the specific combination of the above-mentioned technical features, but also includes other technical solutions formed by any combination of the above-mentioned technical features or their equivalents without departing from the above-mentioned disclosed concepts. For example, a technical solution formed by replacing the above-mentioned features with (but not limited to) technical features with similar functions disclosed in this disclosure.

[0150] In addition, although each operation is described in a specific order, this should not be understood as requiring these operations to be performed in the specific order shown or in a sequential order. Under certain circumstances, multitasking and parallel processing may be advantageous. Similarly, although some specific implementation details have been included in the above discussion, these should not be interpreted as limiting the scope of the present disclosure. Some features described in the context of a separate embodiment can also be implemented in a single embodiment in combination. On the contrary, the various features described in the context of a single embodiment can also be implemented in multiple embodiments individually or in any suitable sub-combination mode.

[0151] Although the subject matter has been described in language specific to structural features and / or methodological logical acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. An image processing method, comprising: Acquire an initial image, and process the initial image to obtain a segmented image, wherein the segmented image is used to indicate a target object in the initial image; A composite image is obtained by using a pre-trained image processing model and the segmented image, wherein the image processing model is used to generate a matching image background for the target object so that there is no visual physical interference between the target object and the image background in the output composite image.

2. The method according to claim 1, wherein: The image processing model includes an encoder unit, a decoder unit and a diffusion unit; A synthetic image is obtained by using a pre-trained image processing model and the segmented image, including: By means of the encoder unit, feature extraction is performed on the initial image to obtain a first feature map; Adding noise to the first feature map by the diffusion unit to generate a noisy feature map; Directionally denoising the noisy feature map by the diffusion unit, and in the directional denoising process, weightedly superimposing a target area of ​​the first feature map to obtain a denoised feature map, wherein the target area is an image area determined based on the segmented image; The de-noising feature map is decoded by the decoder unit to generate the synthetic image.

3. The method according to claim 2, wherein: The diffusion unit includes at least one de-noising unit arranged in an orderly manner, and the noisy feature map is subjected to directional de-noising by the diffusion unit, and in the directional de-noising process, the target area of ​​the first feature map is weightedly superimposed to obtain the de-noised feature map, including: For each of the de-noising units, the following steps are performed in sequence: Get the sequence value of the current de-noising unit; If the sequence value is less than the sequence threshold, the current de-noising unit is used to perform directionally de-noising on the current noisy feature map to obtain an intermediate de-noising feature map output by the current de-noising unit, and then the current de-noising unit is set as the next de-noising unit, the current noisy feature map is set as the intermediate de-noising feature map, and the step of obtaining the sequence value of the current de-noising unit is returned to execute; If the sequence value is greater than the sequence threshold and less than the last sequence corresponding to the last denoising unit, the target area of ​​the first feature map is weightedly superimposed on the current denoising feature map to obtain the superimposed feature map corresponding to the current denoising unit, and then the superimposed feature map is directionally denoised to obtain the intermediate denoising feature map output by the current denoising unit, and then the current denoising unit is set as the next denoising unit, the current denoising feature map is set as the intermediate denoising feature map, and the step of obtaining the sequence value of the current denoising unit is returned to be executed; If the sequence value is equal to the last sequence corresponding to the last denoising unit, the target area of ​​the first feature map is weightedly superimposed on the current denoised feature map to obtain the superimposed feature map corresponding to the current denoising unit, and then the superimposed feature map is directionally denoised to obtain the denoised feature map.

4. The method according to claim 3, further comprising: Acquire a segmentation confidence corresponding to the segmented image, wherein the segmentation confidence represents the accuracy of segmentation of the target object; Obtaining a target weighting coefficient according to the segmentation confidence; The weighted superposition of the target area of ​​the first feature map to the current noise-added feature map to obtain the superimposed feature map corresponding to the current de-noising unit includes: Based on the target weighting coefficient, the target area of ​​the first feature map is weighted and superimposed on the current noisy feature map to obtain the superimposed feature map corresponding to the current de-noising unit.

5. The method according to claim 1, wherein: A synthetic image is obtained by using a pre-trained image processing model and the segmented image, including: Acquire prompt word information, where the prompt word information is used to guide the image processing model to generate an image background having target image content; The initial image, the prompt word information and the segmented image are input into the image processing model to generate the synthesized image.

6. The method according to claim 5, wherein: The obtaining of prompt word information includes: Acquiring shooting parameters of the segmented image, where the shooting parameters represent shooting angle characteristics and / or illumination characteristics of a target object in the segmented image; According to the shooting parameters, a corresponding description text is obtained; The prompt word information is generated according to the description text.

7. The method according to claim 1, further comprising: Acquire model scene information, where the model scene information is used to characterize an image style and / or image scene of an image background generated by the image processing model; Obtaining a target image processing model according to the model scene information; A synthetic image is obtained by using a pre-trained image processing model and the segmented image, including: The segmented image is processed by the target image processing model to obtain a composite image.

8. The method according to claim 1, wherein: The segmented image includes a first segmented image or a second segmented image, wherein: The first segmented image is a transparent background image containing the target object; The second segmented image is a mask image used to locate the target object.

9. The method according to claim 1, wherein: After obtaining the composite image, the method further includes at least one of the following: Based on the segmented image, superimposing the image texture corresponding to the target object into the synthesized image; The synthetic image is processed by a generative adversarial network model to obtain a super-resolution synthetic image, wherein the image resolution of the super-resolution synthetic image is greater than the image resolution of the synthetic image.

10. An image processing device, comprising: An acquisition module is configured to acquire an initial image and process the initial image to obtain a segmented image, wherein the segmented image is used to indicate a target object in the initial image; The processing module is configured to obtain a composite image through a pre-trained image processing model and the segmented image, wherein the image processing model is used to generate a matching image background for the target object so that there is no visual physical interference between the target object and the image background in the output composite image.

11. An electronic device, comprising: Processor and memory; The memory stores computer-executable instructions; The processor executes the computer-executable instructions stored in the memory, so that the processor performs the image processing method according to any one of claims 1 to 9.

12. A computer-readable storage medium, wherein: The computer-readable storage medium stores computer-executable instructions, and when the processor executes the computer-executable instructions, the image processing method according to any one of claims 1 to 9 is implemented.

13. A computer program product comprising a computer program, wherein: When the computer program is executed by a processor, the image processing method according to any one of claims 1 to 9 is implemented.

Citation Information

Patent Citations

  • Image processing method and device, electronic equipment and storage medium

    CN112419328A

  • Image processing method, device and equipment and computer readable storage medium

    CN116704221A

  • Model training based on synthetic data

    US20230334834A1