Background separation for guided generative models

By combining unconditional and conditional masks and applying distance transformation and color mapping techniques, the problem of difficult separation of image foreground and background is solved, and a more natural and accurate image boundary processing is achieved.

CN119941776APending Publication Date: 2025-05-06ADOBE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410919261.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2023-11-02
Filing Date
2024-07-10
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The prior art is difficult to effectively separate the foreground and background of an image, resulting in unnatural boundaries of the generated image and requires manual editing to achieve smoothing.

Method used

By combining unconditional masks and conditional masks, combining distance transformation maps and color maps, a more accurate foreground mask is generated to achieve seamless separation of image foreground and background.

Benefits of technology

More accurate foreground area recognition and separation is achieved, eliminating the overlap and inconsistency of image boundaries, allowing characters to be seamlessly covered in any background.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119941776A_ABST
    Figure CN119941776A_ABST
Patent Text Reader

Abstract

Embodiments of the present disclosure relate to background separation for guiding generative models. Embodiments of the present disclosure include obtaining an input image and an approximation mask, the approximation mask approximately indicating a foreground region of the input image. Some embodiments generate an unconditional mask of a foreground region based on an input image. A conditional mask of the foreground region is generated based on the input image and the approximation mask. An output image is then generated based on the unconditional and conditional masks. In some cases, the output image includes a foreground region of the input image.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims the benefit of U.S. Provisional Application No. 63 / 495,194, filed in the U.S. Patent and Trademark Office on April 10, 2023, pursuant to 35 U.S.C. §119, the disclosure of which is incorporated herein by reference in its entirety. Technical Field

[0003] The following relates generally to image processing, and more specifically to image background separation using machine learning. Background Art

[0004] Image processing refers to the use of computers to edit digital images using algorithms or processing networks. Recently, machine learning models have been used in advanced image processing techniques. Among these machine learning models, diffusion models and other generative models such as generative adversarial networks (GANs) have been used for various tasks, including generating images with perceptual indicators, generating images with conditional settings, image inpainting, and image manipulation.

[0005] Image generation (a subfield of image processing) involves using machine learning models to synthesize images. Machine learning models can be used for a variety of image generation tasks, including image super-resolution, generating images with perceptual metrics, conditional generation (e.g., text-guided generation), image inpainting, and image manipulation. For example, a diffusion model is trained to take random noise as input and generate unseen images with features similar to the training data. Summary of the invention

[0006] The present disclosure describes a system and method for image processing. Embodiments of the present disclosure include an image processing device, which is configured to separate the foreground of an image (e.g., an area including one or more characters) from the background via masking. In some cases, an image generation model generates characters according to a specified font style. The font style is represented by a font mask (a binary image depicting a character) that guides character generation. Since characters may not follow the boundaries specified by the font style, the image processing device seamlessly covers the characters into any background by accurately identifying the foreground area. In some examples, the image processing device generates a combined foreground mask, which is based on a combination of a conditional mask and an unconditional mask. The conditional mask is based on a font mask (e.g., an approximate mask indicating the foreground area of ​​an input image) and an input image. In some examples, the combined foreground mask also incorporates a distance transform map and a color map.

[0007] Methods, apparatuses, and non-transitory computer-readable media for image processing are described. One or more embodiments of the method, apparatus, and non-transitory computer-readable media include: obtaining an input image and an approximate mask, the approximate mask approximately indicating a foreground region of the input image; based on the input image, generating an unconditional mask of the foreground region by an unconditional mask network; generating a conditional mask of the foreground region based on the input image and the approximate mask by a conditional mask network; and generating an output image based on the unconditional mask and the conditional mask, the output image including the foreground region of the input image.

[0008] Methods, apparatus, and non-transitory computer-readable media for image processing are described. One or more embodiments of the method, apparatus, and non-transitory computer-readable media include: initializing an unconditional mask network and a conditional mask network; receiving training data, the training data including an input image, an approximate mask indicating a foreground region of the input image, and a ground truth mask; using the training data, training the unconditional mask network to generate an unconditional mask of the foreground region based on the input image; and using the training data, training the conditional mask network to generate a conditional mask of the foreground region based on the input image and the approximate mask.

[0009] An apparatus and method for image processing are described. One or more embodiments of the apparatus and method include at least one processor; at least one memory, the at least one memory including instructions that can be executed by the at least one processor; a user interface, the user interface including parameters stored in at least one memory and configured to obtain an input image and an approximate mask, the approximate mask approximately indicating a foreground region of the input image; an unconditional mask network, the unconditional mask network including parameters stored in at least one memory and configured to generate an unconditional mask of the foreground region based on the input image; and a conditional mask network, the conditional mask network including parameters stored in at least one memory and configured to generate a conditional mask of the foreground region based on the input image and the approximate mask. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 An example of an image processing system according to aspects of the present disclosure is shown.

[0011] Figure 2 An example of an image generation application of the method according to aspects of the present disclosure is shown.

[0012] Figures 3 to 5 An example of a background separation effect according to aspects of the present disclosure is shown.

[0013] Figure 6 An example of a method for image processing according to the present disclosure is shown.

[0014] Figure 7 An example of an image processing apparatus according to aspects of the present disclosure is shown.

[0015] Figure 8 An example of a guided diffusion model according to aspects of the present disclosure is shown.

[0016] Fig. 9 An example of a U-shaped mesh according to aspects of the present disclosure is shown.

[0017] Fig.10 Examples of machine learning models according to aspects of the present disclosure are shown.

[0018] Fig.11 An example of a diffusion process according to aspects of the present disclosure is shown.

[0019] Fig.12 An example of a method for training a diffusion model according to aspects of the present disclosure is shown.

[0020] Fig.13 Examples of methods for training machine learning models according to aspects of the present disclosure are shown.

[0021] Fig.14 An example of a computing device according to aspects of the present disclosure is shown. DETAILED DESCRIPTION

[0022] The present disclosure describes a system and method for image processing. An embodiment of the present disclosure includes an image processing device, which is configured to separate the foreground of an image (e.g., an area including one or more characters) from the background via masking. In some cases, an image generation model generates characters according to a specified font style. The font style is represented by a font mask (a binary image depicting a character) that guides character generation. Since characters may not follow the boundaries specified by the font style, the image processing device seamlessly covers the characters into any background by accurately identifying the foreground area. In some examples, the image processing device generates a combined foreground mask, which is based on a combination of a conditional mask and an unconditional mask. The conditional mask is based on a font mask (e.g., an approximate mask indicating the foreground area of ​​an input image) and an input image. In some examples, the combined foreground mask also includes a distance transform map and a color map.

[0023] Recently, users have used software applications to modify properties related to text. For example, in a word editing application, a user can change properties such as font and text color. Generative models produce text effects when combined with technologies such as SDEdit and in-painting. Some methods generate characters in a style specified by a font style. The font style defines a font mask (a binary image depicting the character) that can guide character generation. However, these generated images do not necessarily follow the boundaries specified by the font mask. Naturally placing character sets adjacent to each other will lead to unsatisfactory results, such as overlapping boundaries and inconsistent edges. In order to seamlessly overlay characters into the target background, the characters (foreground area) need to be separated from the background.

[0024] In the field of image processing, background separation involves separating the foreground region of a generated image (as input) from the background region for further editing. Specifically, the background separation task involves disentangling foreground elements (such as objects or characters located in the foreground region) from the background.

[0025] Conventional models struggle to separate the foreground from the background of an image due to inadequate masking methods. Some methods result in poor quality generation with a "cut-off" effect. Content creators must manually edit the input image to "smooth out" the borders of the image to match other characters. The editing process is time-consuming and unfriendly to inexperienced editors.

[0026] Embodiments of the present disclosure include an improved image processing apparatus that generates a more accurate foreground mask. For example, the foreground mask generated by the improved image processing apparatus may be more suitable for use with characters having a defined shape (such as text characters). This enables a user to generate text characters with a synthetically generated texture that can be separated from the background (e.g., placing the text on another background).

[0027] Some embodiments of the present disclosure are configured to receive an input image and an approximation mask via a user interface, and generate an output image including a foreground region of the input image. The image processing device combines the unconditional mask and the conditional mask to obtain a combined and improved foreground for background separation. In some examples, the distance transform map and the color distance map are also incorporated into the combined foreground mask. Therefore, the combined foreground mask applies an accurate foreground probability mask representing the foreground region of the input image.

[0028] In some embodiments, the unconditional mask network generates an unconditional mask of the foreground region based only on the input image. The unconditional mask includes a larger area than the foreground region of the input image and applies a less accurate estimate near the boundary of the foreground region. Additionally, the conditional mask network generates a conditional mask of the foreground region based on the input image and the approximate mask. In this way, the approximate mask is refined to obtain a more accurate estimate of the boundary of the foreground region. In some examples, the approximate mask is a font mask representing a font style followed by characters in the input image. The approximate mask is an approximation of the foreground region of the input image.

[0029] In some examples, the image processing apparatus calculates a distance transform map by calculating a distance transform using the font mask (i.e., an additive estimate of a foreground probability mask via distance). Additionally, the image processing apparatus generates a color map by calculating a distance and a ratio of an average value of pixel values ​​inside the font mask to an average value of pixel values ​​outside the font mask (i.e., an additive estimate of a foreground probability mask via color).

[0030] Embodiments of the present disclosure utilize multiple probabilistic masks from various sources to generate accurate mask predictions for foreground regions via a mask combination component. An image processing apparatus based on the present disclosure uniquely combines unconditional masks, conditional masks, distance transform maps, and color distance maps to generate a combined foreground mask. An island removal operation is then applied to remove small objects located outside the region defined by the font mask. The combined and refined foreground mask accurately surrounds the boundaries of the foreground region of the input image.

[0031] In some examples, the mask or transparency can be generated after the style is applied to the glyph. Since the generated image may not conform to the exact boundaries of the original text font, it may be useful to determine the boundaries of the foreground after generation. This is useful to enable the user to apply text to another background.

[0032] To obtain a mask or transparency of foreground text including text effects, a combination of methods may be applied. For example, a subject selection method, an object selection method, or a color-based selection method may be used to identify and distinguish foreground and background pixels. In some cases, a single boundary selection method is used, but in some cases a combination of methods is applied. A distance transform may also be used to generate or refine a glyph boundary mask.

[0033] Thus, embodiments of the present disclosure provide improvements over conventional image processing and editing software by enabling the generation of image masks that more accurately distinguish image background from foreground objects.Some embodiments particularly enable accurate distinction between complex text characters and fonts and image background.

[0034] For example, the embodiment improves the font mask (binary image depicting an object or character) that guides character generation. The background-separated image more closely follows the boundaries specified by the font style. Character images processed in this way can be placed adjacent to each other and problems such as overlapping boundaries and inconsistent edges are eliminated. By separating the characters (foreground) from the background in a precise manner, content creators can easily overlay the characters seamlessly into the target background.

[0035] Embodiments of the present disclosure may be used in the context of image generation applications. For example, an image processing apparatus based on the present disclosure receives an input image generated by a guided generative model, separates a foreground region from a background, and generates an output image including the foreground region of the input image. Figure 2-Figure 5 Provides example applications in the context of image generation. Figure 1 and Figure 7-Figure 11 Provides details about the architecture of an example image processing system. Figure 2 and Figure 6 Provides details about the image processing process. Figure 12-13 Describe the example training process.

[0036] Background separation process

[0037] exist Figure 1-Figure 6 In the invention, a method, apparatus, and non-transitory computer-readable medium for image processing are described. One or more embodiments of the method, apparatus, and non-transitory computer-readable medium include: obtaining an input image and an approximate mask, the approximate mask approximately indicating a foreground region of the input image; generating an unconditional mask of the foreground region based on the input image through an unconditional mask network; generating a conditional mask of the foreground region based on the input image and the approximate mask through a conditional mask network; and generating an output image based on the unconditional mask and the conditional mask, the output image including the foreground region of the input image.

[0038] Some examples of methods, apparatus, and non-transitory computer-readable media further include: combining the unconditional mask and the conditional mask to obtain a combined mask, wherein the output image is generated based on the combined mask. In some examples, the input image is generated based on the approximate mask. Some examples of methods, apparatus, and non-transitory computer-readable media also include calculating a distance transform map based on the approximate mask, wherein the output image is generated based on the distance transform map.

[0039] Some examples of methods, apparatus, and non-transitory computer-readable media also include calculating a color distance map based on the approximation mask, wherein an output image is generated based on the color distance map. Some examples of methods, apparatus, and non-transitory computer-readable media also include performing island area removal on the input image, wherein the output image is generated based on the island area removal. In some examples, the foreground area includes text based on a font and modified using a text effect, and wherein the approximation mask is based on the text and the font without the text effect. In some examples, the unconditional mask includes a probabilistic mask and the conditional mask includes a refined probabilistic mask.

[0040] In some examples, an unconditional mask network is trained to generate an unconditional mask of a foreground region based on an input image, and a conditional mask network is trained to generate a conditional mask of a foreground region based on the input image and an approximate mask.

[0041] Figure 1 An example of an image processing system according to aspects of the present disclosure is shown. The example shown includes a user 100, a user device 105, an image processing apparatus 110, a cloud 115, and a database 120. The image processing apparatus 110 is a reference Figure 7 Examples of corresponding elements described, or including reference Figure 7 Aspects of the corresponding elements described.

[0042] exist Figure 1 In the example shown, the input image and the approximation mask are provided by the user 100 and transmitted to the image processing apparatus 110, for example, via the user device 105 and the cloud 115. The approximation mask approximately indicates the foreground area of ​​the input image. In some examples, the text effect model (see Figure 7 ) applies a text effect "butterfly" to the text "y" and generates an input image based on the approximate mask. With the text effect described by the text effect "butterfly", the input image includes wings and tentacles depicting the text "y". In some examples, the text effect model is an AI generative model such as a diffusion model.

[0043] The image processing device 110 generates an unconditional mask of the foreground area based on the input image through an unconditional mask network. The image processing device 110 generates a conditional mask of the foreground area based on the input image and the approximate mask through a conditional mask network. The image processing device 320 generates an output image including the foreground area. In the absence of other irrelevant objects in the input image, the foreground area more accurately surrounds the lowercase letter "y". The image processing device 110 separates the foreground (i.e., the area including one or more characters) from the background to obtain an output image.

[0044] In some examples, a text effect model (e.g., a pixel diffusion model) generates an input image. The image processing device 110 obtains the input image and generates an output image including a foreground area of ​​the input image based on an unconditional mask and a conditional mask via a background separation process. The image processing device 110 returns the output image to the user 100 via the cloud 115 and the user device 105. The background-separated image can then be combined with other similar character images to create visually appealing text. Reference Figure 2 The process of using the image processing device 110 is further described.

[0045] User device 105 may be a personal computer, laptop computer, mainframe computer, palmtop computer, personal assistant, mobile device, or any other suitable processing device. In some examples, user device 105 includes software incorporating an image processing application (e.g., an image editing application). In some examples, the image editing application on user device 105 may include the functionality of image processing apparatus 110.

[0046] The user interface can enable the user 100 to interact with the user device 105. In some embodiments, the user interface can include an audio device such as an external speaker system, an external display device such as a display screen, or an input device (e.g., a remote control device that interfaces with the user interface directly or with the aid of an I / O controller module). In some cases, the user interface can be a graphical user interface (GUI). In some examples, the user interface can be represented using code that is sent to the user device 105 and rendered locally by the browser.

[0047] The image processing device 110 includes a computer-implemented network including a user interface, an unconditional mask network, a conditional mask network, and a mask combination component. In some examples, the image processing device 110 includes an image generation model and a text effect model. The image processing device 110 may also include a processor unit, a memory unit, an I / O module, and a training component. The training component is used to train a machine learning model (or image processing network). Additionally, the image processing device 110 can communicate with a database 120 via a cloud 115. In some cases, the architecture of the image processing network is also referred to as a network, a machine learning model, or a network model. Reference Figure 7-Figure 11 More details about the architecture of the image processing device 110 are provided. Figure 2 and Figure 6 Further details regarding the operation of the image processing apparatus 110 are provided.

[0048] In some cases, the image processing device 110 is implemented on a server. The server provides one or more functions to the user through one or more network links in various networks. In some cases, the server includes a single microprocessor board, and the single microprocessor board includes a microprocessor responsible for controlling all aspects of the server. In some cases, the server uses a microprocessor and a protocol to exchange data with other devices / users on one or more networks via a hypertext transfer protocol (HTTP) and a simple mail transfer protocol (SMTP), but other protocols such as a file transfer protocol (FTP) and a simple network management protocol (SNMP) may also be used. In some cases, the server is configured to send and receive a hypertext markup language (HTML) formatted file (e.g., for displaying a web page). In various embodiments, the server includes a general-purpose computing device, a personal computer, a laptop computer, a mainframe computer, a supercomputer, or any other suitable processing device.

[0049] Cloud 115 is a computer network configured to provide on-demand availability of computer system resources such as data storage and computing power. In some examples, cloud 115 provides resources without active management by the user. The term cloud is sometimes used to describe a data center that is available to many users via the Internet. Some large cloud networks have functions distributed over multiple locations from a central server. A server is designated as an edge server if it has a direct or close connection to the user. In some cases, cloud 115 is limited to a single organization. In other examples, cloud 115 can be used for many organizations. In one example, cloud 115 includes a multi-layer communication network including multiple edge routers and core routers. In another example, cloud 115 is based on a local collection of switches in a single physical location.

[0050] Database 120 is an organized collection of data. For example, database 120 stores data in a specified format called a schema. Database 120 can be constructed as a single database, a distributed database, multiple distributed databases, or an emergency backup database. In some cases, a database controller can manage the storage and processing of data in database 120. In some cases, a user interacts with the database controller. In other cases, the database controller can operate automatically without user interaction.

[0051] Figure 2An example of a method 200 for image generation applications according to aspects of the present disclosure is shown. In some examples, these operations are performed by a system including a processor that executes a code set for a functional element of a control device. Additionally or alternatively, certain processes are performed using dedicated hardware. Typically, these operations are performed according to the methods and processes described in accordance with aspects of the present disclosure. In some cases, the operations described herein are composed of various sub-steps or performed in conjunction with other operations.

[0052] At operation 205, the user provides an input image and an approximation mask. In some cases, the operation of this step involves reference Figure 1 The user described or can be referenced by Figure 1 Describes the user execution.

[0053] As Figure 2 In the example shown, with the text effect described by the text effect "butterfly", the input image includes wings and antennae depicting the text "y". In some examples, the approximate mask is an approximation of the foreground mask, and the approximate foreground mask will be refined in the following operations. The approximate mask (i.e., the font mask) is used to guide the refinement process.

[0054] At operation 210, the system generates an unconditional mask and a conditional mask. In some cases, the operation of this step involves reference to Figure 1 and Figure 7 The image processing device described herein may be referred to as Figure 1 and Figure 7 The method is performed by the image processing device described in the embodiment of the present invention. In some examples, an unconditional mask prediction model (e.g., a first deep learning network) predicts a foreground mask based on an input image. Additionally, a conditional mask prediction model (e.g., a second deep learning network) takes the input image and the approximate mask as input. The conditional mask prediction model is configured to refine the approximate mask to obtain the predicted precise mask. One or more embodiments generate a precise mask by combining multiple sources of background priors (e.g., foreground mask probabilities) and combining background priors used to guide the generation model.

[0055] At operation 215, the system generates an output image including the foreground region of the input image based on the unconditional mask and the conditional mask. In some cases, the operation of this step involves referring to Figure 1 and Figure 7 The image processing device described may be referred to as Figure 1 and Figure 7The image processing device described in the present invention is used to perform. In the above example, the foreground area includes the character "y". In the absence of other irrelevant objects in the input image, the foreground area more accurately surrounds the lowercase letter "y". Therefore, after the background separation based on the present disclosure, the character "y" can be seamlessly incorporated into the target background.

[0056] At operation 220, the system displays the output image to the user. In some cases, the operation of this step involves referring to Figure 1 and Figure 7 The image processing device described may be referred to as Figure 1 and Figure 7 The image processing apparatus described is performed.

[0057] Figure 3 An example of a background separation effect according to aspects of the present disclosure is shown. The example shown includes an input image 300, an unconditional mask 305, a conditional mask 310, and a combined mask 315. In one embodiment, a mask combination component of an image processing device (see Figure 7 ) The unconditional mask 305 and the conditional mask 310 are combined to obtain a combined mask 315. An output image is generated based on the combined mask 315.

[0058] In one example, the input image 300 is an image of text, the text includes one or more characters based on a particular font, where additional shapes or textures (e.g., textures generated by a generative machine learning model) decorate the text. In this example, the approximate mask can be a font mask based on the original text, but without the additional shapes or textures. However, the approximate mask can represent any approximation of a foreground object. The input image 300 can be used alone to generate an unconditional mask 305, while the approximate mask (e.g., a font mask) is used to generate a conditional mask 310. In some examples, the conditional mask 310 is more likely to reflect the original shape (e.g., the original font shape), but may not incorporate some additional shapes or textures. In contrast, the unconditional mask 305 can capture additional shapes or textures, but may not correspond to the original shape.

[0059] In some examples, the text effect model (see Figure 7 ) applies the text effect "butterfly" to the text "y" and generates an input image 300. The input image 300 is generated based on the approximation mask. With the text effect described by the text effect hint "butterfly", the input image 300 includes wings and tentacles attached to the text "y". The approximation mask is used to approximately indicate the foreground area of ​​the input image 300. In some examples, the text effect model is an AI generated model such as a diffusion model. The input image 300 is a reference Figure 4 and Figure 5Examples of corresponding elements described, or including references Figure 4 and Figure 5 Describes aspects of the corresponding elements.

[0060] Here, the unconditional mask 305, the conditional mask 310, or the combined mask 315 may have the same resolution or size as the input image 300. The mask (e.g., the unconditional mask 305, the conditional mask 310, the combined mask 315) is used to indicate the location of the foreground region based on the value of the mask (the value is between 0 and 1). Figure 3 In the example shown, the foreground area more precisely shows the letter "y". That is, the foreground area is an area having the shape of the letter "y" as shown in the figure. The image processing device separates the foreground area (i.e., the area including one or more characters) of the input image 300 from the background area to obtain the output image.

[0061] In this example, the unconditional mask 305 includes additional content, such as butterfly wings and antennae extending from the letter "y". By incorporating the two masks, embodiments of the present disclosure provide a combined mask 315 and achieve a balanced result that reflects the desired shape while also capturing additional shapes and textures. For example, compared to the unconditional mask 305, the combined mask 315 includes less content, while compared to the conditional mask 310, the combined mask 315 includes more content.

[0062] In some embodiments, parameters are included to continuously adjust between the unconditional mask 305 and the conditional mask 310 such that the image processing device generates the combined mask 315 based on a weighted combination of the unconditional mask 305 and the conditional mask 310 .

[0063] Figure 4 An example of a background separation effect according to aspects of the present disclosure is shown. The example shown includes an input image 400, an image processing device 405, and a combined mask 410. The image processing device 405 is a reference Figure 1 and Figure 7 Examples of corresponding elements described, or including references Figure 1 and Figure 7 Describes aspects of the corresponding elements.

[0064] In some examples, the text effect model (see Figure 7) applies the text effect "butterfly" to the text "l" and generates an input image 400. The input image 400 is generated based on the approximate mask. The input image 400 includes wings and tentacles attached to the text "l", where the text effect is described by the text effect hint "butterfly". The approximate mask is used to approximately indicate the foreground area of ​​the input image 400. The image processing device 405 generates a combined mask 410. The combined mask 410 is generated based on a combination of the unconditional mask and the conditional mask.

[0065] Here, the unconditional mask, the conditional mask, or the combined mask 410 may have the same resolution or size as the input image 400. The mask (e.g., the unconditional mask, the conditional mask, the combined mask 410) is used to indicate the location of the foreground region based on the value of the mask (the value is between 0 and 1). Figure 4 In the example shown, the foreground area more precisely shows the letter "l". That is, the foreground area is an area having the shape of the letter "l" as shown in the figure. The image processing device 405 separates the foreground area (i.e., the area including one or more characters) of the input image 400 from the background area to obtain an output image.

[0066] The input image 400 is a reference Figure 3 and Figure 5 Examples of corresponding elements described, or including references Figure 3 and Figure 5 The combined mask 410 is a reference to the corresponding elements of the Figure 3 and Figure 5 Examples of corresponding elements described, or including references Figure 3 and Figure 5 Describes aspects of the corresponding elements.

[0067] Figure 5 An example of a background separation effect according to aspects of the present disclosure is shown. The example shown includes an input image 500, an image processing device 505, and a combined mask 510. The image processing device 505 is a reference Figure 1 and Figure 7 Examples of corresponding elements described, or including references Figure 1 and Figure 7 Describes aspects of the corresponding elements.

[0068] In some examples, the text effect model (see Figure 7) applies a text effect to the text "r" and generates an input image 500. The input image 500 is generated based on the approximate mask. Using the text effect described by the text effect prompt, the input image 500 depicts a lowercase "r". The approximate mask is used to approximately indicate the foreground area of ​​the input image 500. The image processing device 505 generates a combined mask 510. The combined mask 510 is generated based on a combination of an unconditional mask and a conditional mask.

[0069] Here, the unconditional mask, the conditional mask, or the combined mask 510 may have the same resolution or size as the input image 500. The mask (e.g., the unconditional mask, the conditional mask, the combined mask 510) is used to indicate the location of the foreground region based on the value of the mask (the value is between 0 and 1). Figure 5 In the example shown, the foreground area more precisely shows the letter "r". That is, the foreground area is an area having the shape of the letter "r" as shown in the figure. The image processing device 405 separates the foreground area (i.e., the area including one or more characters) of the input image 500 from the background area to obtain an output image.

[0070] The input image 500 is a reference Figure 3 and Figure 4 Examples of corresponding elements described, or including references Figure 3 and Figure 4 The combined mask 510 is a reference to the corresponding elements of the description. Figure 3 and Figure 4 Examples of corresponding elements described, or including references Figure 3 and Figure 4 Describes aspects of the corresponding elements.

[0071] Figure 6 An example of a method 600 for image processing according to aspects of the present disclosure is shown. In some examples, these operations are performed by a system including a processor that executes a code set for a functional element of a control device. Additionally or alternatively, certain processes are performed using dedicated hardware. Typically, these operations are performed according to the methods and processes described in various aspects of the present disclosure. In some cases, the operations described herein are composed of various sub-steps or are performed in conjunction with other operations.

[0072] At operation 605, the system obtains an input image and an approximate mask that approximately indicates a foreground region of the input image. In some cases, the operation of this step involves referring to Figure 7 and Fig.10 The user interface described may be referenced by Figure 7 and Fig.10 The user interface described is executed.

[0073] At operation 610, the system generates an unconditional mask of the foreground region based on the input image through an unconditional mask network. In some cases, the operation of this step involves reference Figure 7 and Fig.10 The unconditional mask network described can be obtained by referring to Figure 7 and Fig.10 The unconditional mask network described is implemented as follows.

[0074] According to some embodiments of the present disclosure, the machine learning model separates the foreground of the input image (ie, one or more characters of the input image) from the background. In some cases, the text effect model generates the input image.

[0075] In some examples, the unconditional mask network of the machine learning model calculates or generates an unconditional foreground probability mask based only on the input image. In some cases, the input image is also referred to as a generated image or a synthesized image. The generated foreground probability mask includes more desired or target areas and is not accurate near the boundaries.

[0076] At operation 615, the system generates a conditional mask of the foreground region based on the input image and the approximate mask through a conditional mask network. In some cases, the operation of this step involves reference Figure 7 and Fig.10 The conditional mask network described can be constructed by referring to Figure 7 and Fig.10 The conditional mask network described is used to perform

[0077] In some embodiments, a conditional mask network of the machine learning model computes or generates a conditional foreground probability mask based on the font mask and the input image. A second deep learning network takes as input an input image with a rough approximation of the characters and the foreground mask, and refines the approximate foreground mask. The font mask (i.e., as an approximate mask) is used to guide the refinement process. The unconditional mask network and the conditional mask network are separate deep learning networks.

[0078] At operation 620, the system generates an output image including the foreground region of the input image based on the unconditional mask and the conditional mask. In some cases, the operation of this step involves reference Figure 7 and Fig.10 The image generation model described can also be referenced by Figure 7 and Fig.10 The image generation model described is performed.

[0079] Network Architecture

[0080] exist Figure 7-Figure 11In the invention, an apparatus and method for image processing are described. One or more embodiments of the apparatus and method include at least one processor; at least one memory including instructions, the instructions being executable by the at least one processor; a user interface including parameters stored in at least one memory and configured to obtain an input image and an approximate mask, the approximate mask approximately indicating a foreground region of the input image; an unconditional mask network including parameters stored in at least one memory and configured to generate an unconditional mask of the foreground region based on the input image; and a conditional mask network including parameters stored in at least one memory and configured to generate a conditional mask of the foreground region based on the input image and the approximate mask.

[0081] Some examples of the apparatus and method further include generating an output image based on the unconditional mask and the conditional mask via a background separation process, wherein the output image includes a foreground region of the input image. Some examples of the apparatus and method further include a distance transform component configured to calculate a distance transform map based on the approximation mask, wherein the output image is generated based on the distance transform map.

[0082] Some examples of the apparatus and method also include a color distance component configured to calculate a color distance map based on an approximate mask, wherein an output image is generated based on the color distance map. Some examples of the apparatus and method also include an island area removal component configured to perform island area removal on an input image, wherein the output image is generated based on the island area removal. Some examples of the apparatus and method further include a mask combination component, which is configured to combine an unconditional mask and a conditional mask to obtain a combined mask, wherein the output image is generated based on the combined mask. Some examples of the apparatus and method also include a text effect model configured to generate an input image.

[0083] Figure 7 An example of an image processing device 700 according to aspects of the present disclosure is shown. The example shown includes the image processing device 700, a processor unit 705, an I / O module 710, a training component 715, a memory unit 720, and a machine learning model 725. The image processing device 700 is a reference Figure 1 Examples of corresponding elements described, or including references Figure 1 Describes aspects of the corresponding elements.

[0084] Machine learning model 725 is a reference Fig.10 Examples of corresponding elements described or include references Fig.10In one embodiment, the machine learning model 725 includes a user interface 730 , an unconditional mask network 735 , a conditional mask network 740 , a mask combination component 745 , an image generation model 750 , and a text effect model 755 .

[0085] Processor unit 705 is an intelligent hardware device (e.g., a general processing component, a digital signal processor (DSP), a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller, an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a programmable logic device, a discrete gate or transistor logic component, a discrete hardware component, or any combination thereof). In some cases, processor unit 705 is configured to operate a memory array using a memory controller. In other cases, the memory controller is integrated into the processor. In some cases, processor unit 705 is configured to execute computer-readable instructions stored in the memory to perform various functions. In some embodiments, processor unit 705 includes dedicated components for modem processing, baseband processing, digital signal processing, or transmission processing.

[0086] Examples of memory unit 720 include random access memory (RAM), read-only memory (ROM), or hard disk. Examples of memory unit 720 include solid-state memory and hard disk drive. In some examples, memory unit 720 is used to store computer-readable-computer executable software including instructions, which, when executed, cause the processor to perform various functions described herein. In some cases, memory unit 720 particularly includes a basic input / output system (BIOS), which controls basic hardware or software operations, such as interaction with peripheral components or devices. In some cases, a memory controller operates the memory unit. For example, a memory controller may include a row decoder, a column decoder, or both. In some cases, the memory cells within memory unit 720 store information in the form of logical states.

[0087] In some examples, at least one memory unit 720 includes instructions that can be executed by at least one processor unit 705. Memory unit 720 includes machine learning model 725 or stores parameters of machine learning model 725.

[0088] I / O module 710 (e.g., input / output interface) can include an I / O controller. An I / O controller can manage input and output signals for a device. An I / O controller can also manage peripheral devices that are not integrated into the device. In some cases, an I / O controller can represent a physical connection or port to an external peripheral device. In some cases, an I / O controller can utilize an operating system, such as Or other known operating systems. In other cases, the I / O controller may represent or interact with a modem, keyboard, mouse, touch screen, or similar device. In some cases, the I / O controller may be implemented as part of a processor. In some cases, a user may interact with the device via the I / O controller or via hardware components controlled by the I / O controller.

[0089] In some examples, the I / O module 710 includes a user interface. The user interface can enable a user to interact with the device. In some embodiments, the user interface can include an audio device such as an external speaker system, an external display device such as a display screen, or an input device (e.g., a remote control device that is directly or by means of an I / O controller module and docked with the user interface). In some cases, the user interface can be a graphical user interface (GUI). In some examples, the communication interface operates at the boundary between the communication entity and the channel, and can also record and process communication. A communication interface is provided herein to enable a processing system coupled to a transceiver (e.g., a transmitter and / or a receiver). In some examples, the transceiver is configured to transmit (or send) and receive signals for a communication device via an antenna.

[0090] According to some embodiments of the present disclosure, the image processing device 700 includes a computer-implemented artificial neural network (ANN) for mask prediction and image generation. ANN is a hardware or software component including multiple connection nodes (i.e., artificial neurons), which loosely correspond to neurons in the human brain. Each connection or edge transmits a signal from one node to another node (like a physical synapse in the brain). When a node receives a signal, it processes the signal and then transmits the processed signal to other connection nodes. In some cases, the signal between nodes includes real numbers and the output of each node is calculated by a function of the sum of its inputs. Each node and edge is associated with one or more node weights that determine how to process and transmit signals.

[0091] According to some embodiments, the image processing device 700 includes a convolutional neural network (CNN) for image processing (e.g., image encoding, image decoding). CNN is a type of neural network commonly used in computer vision or image classification systems. In some cases, CNN can enable digital images to be processed with minimal preprocessing. CNN can be characterized by using convolution (or cross-correlation) hidden layers. These layers apply convolution operations to the input before sending the results to the next layer. Each convolution node can process data of a finite domain (i.e., a receiving domain) of the input. During the forward propagation (forward pass) of the CNN, the filter at each layer can be convolved across the input volume, calculating the dot product between the filter and the input. During the training process, the filters can be modified so that they are activated when they detect specific features within the input.

[0092] According to some embodiments, the training component 715 initializes the unconditional mask network 735 and the conditional mask network 740. In some examples, the training component 715 receives training data, the training data including an input image, an approximate mask indicating a foreground region of the input image, and a ground truth mask. The training component 715 uses the training data to train the unconditional mask network 735 to generate an unconditional mask of the foreground region based on the input image. The training component 715 uses the training data to train the conditional mask network 740 to generate a conditional mask of the foreground region based on the input image and the approximate mask. In some cases, the training component 715 (shown in dotted lines) is implemented on a device other than the image processing device 700.

[0093] According to some embodiments, user interface 730 obtains an input image and an approximation mask that approximately indicates a foreground region of the input image. In some examples, the foreground region includes text based on a font and modified using a text effect, and wherein the approximation mask is based on the text and the font without the text effect.

[0094] According to some embodiments, the user interface 730 includes parameters stored in at least one memory and configured to obtain an input image and an approximate mask, the approximate mask approximately indicating a foreground area of ​​the input image. The user interface 730 is a reference Fig.10 Examples of corresponding elements described, or including references Fig.10 Describes aspects of the corresponding elements.

[0095] According to some embodiments, the unconditional mask network 735 generates an unconditional mask of the foreground region based on the input image. In some examples, the unconditional mask includes a probabilistic mask. The unconditional mask network 735 is a reference Fig.10 Examples of corresponding elements described, or including references Fig.10 Describes aspects of the corresponding elements.

[0096] According to some embodiments, the conditional mask network 740 generates a conditional mask of the foreground region based on the input image and the approximation mask. In some examples, the conditional mask includes a refined probability mask. The conditional mask network 740 is a reference Fig.10 Examples of corresponding elements described or include references Fig.10 Describes aspects of the corresponding elements.

[0097] According to some embodiments, the mask combination component 745 combines the unconditional mask and the conditional mask to obtain a combined mask, wherein an output image is generated based on the combined mask. Fig.10 Examples of corresponding elements described, or including references Fig.10 Describes aspects of the corresponding elements.

[0098] In some examples, the input image is generated based on the approximation mask. In some examples, the text effect model 755 is configured to generate the input image. According to some embodiments, the output image is generated based on the unconditional mask and the conditional mask via a background separation process, and the output image includes a foreground area of ​​the input image.

[0099] Figure 8 An example of a guided diffusion model 800 according to aspects of the present disclosure is shown. The example shown includes a guided diffusion model 800, an original image 805, a pixel space 810, a forward diffusion process 815, a noise image 820, a backward diffusion process 825, an output image 830, a text prompt 835, a text encoder 840, a guided feature 845, and a guided space 850. The guided diffusion model 800 is a reference Figure 7 Examples of corresponding elements described or include references Figure 7 Aspects of the corresponding element described (see text effect model 755).

[0100] Diffusion models are a class of generative neural networks that can be trained to generate new data with features similar to those found in the training data. Specifically, diffusion models can be used to generate new images. Diffusion models can be used for a variety of image generation tasks, including image super-resolution, generating images with perceptual metrics, conditional generation (e.g., text-guided generation), image inpainting, and image manipulation.

[0101] Types of diffusion models include denoising diffusion probabilistic models (DDPM) and denoising diffusion implicit models (DDIM). In DDPM, the generation process involves inverting a random Markov diffusion process. On the other hand, DDIM uses a deterministic process so that the same input produces the same output. Diffusion models can also be characterized by whether the noise is added to the image itself or to the image features generated by the encoder (i.e., latent diffusion).

[0102] The diffusion model works by iteratively adding noise to the data during the forward process, and then learning to recover the data by denoising the data during the reverse process. For example, during training, the guided diffusion model 800 can take as input an original image 805 in pixel space 810 and apply a forward diffusion process 815 to gradually add noise to the original image 805 to obtain noisy images 820 of various noise levels.

[0103] Next, the back diffusion process 825 (e.g., U-net ANN) gradually removes noise from the noisy image 820 at various noise levels to obtain an output image 830. In some cases, the output image 830 is created by each noise level in the various noise levels. The output image 830 can be compared with the original image 805 to train the back diffusion process 825.

[0104] The back diffusion process 825 may also be guided based on a textual cue 835 or another guiding cue such as an image, layout, segmentation map, etc. The textual cue 835 may be encoded using a text encoder 840 (e.g., a multimodal encoder) to obtain guiding features 845 in a guiding space 850. The guiding features 845 may be combined with the noise image 820 at one or more layers of the back diffusion process 825 to ensure that the output image 830 includes the content described by the textual cue 835. For example, the guiding features 845 may be combined with the noise features using a cross-attention block within the back diffusion process 825.

[0105] The original image 805 is a reference Fig.11 Examples of corresponding elements described or include references Fig.11 The forward diffusion process 815 is referenced Fig.11 Examples of corresponding elements described, or including references Fig.11 The reverse diffusion process 825 is referenced Fig.11 Examples of corresponding elements described or include references Fig.11 The output image 830 is a reference to the corresponding element of the description. Figure 3-Figure 5 Examples of corresponding elements described or include references Figure 3-Figure 5 Describes aspects of the corresponding elements.

[0106] Fig. 9An example of a U-net 900 according to aspects of the present disclosure is shown. The example shown includes a U-net 900, input features 905, an initial neural network layer 910, intermediate features 915, a downsampling layer 920, downsampled features 925, an upsampling process 930, upsampled features 935, skip connections 940, a final neural network layer 945, and output features 950.

[0107] In some examples, the diffusion model is based on a neural network architecture called U-net. U-net 900 takes input features 905 having an initial resolution and an initial number of channels and processes the input features 905 using an initial neural network layer 910 (e.g., a convolutional network layer) to produce intermediate features 915. The intermediate features 915 are then downsampled using a downsampling layer 920 so that the downsampled features 925 have a resolution smaller than the initial resolution and a number of channels larger than the initial number of channels.

[0108] This process is repeated multiple times and then the process is reversed. That is, the downsampled features 925 are upsampled using the upsampling process 930 to obtain upsampled features 935. The upsampled features 935 can be combined with the intermediate features 915 having the same resolution and number of channels via skip connections 940. These inputs are processed using the final neural network layer 945 to produce output features 950. In some cases, the output features 950 have the same resolution as the initial resolution and the same number of channels as the initial number of channels.

[0109] In some cases, U-net 900 employs additional input features to produce conditionally generated outputs. For example, the additional input features may include a vector representation of the input prompt. The additional input features may be combined with the intermediate features 915 within the neural network at one or more layers. For example, a cross-attention module may be used to combine the additional input features with the intermediate features 915.

[0110] Fig.10 An example of a machine learning model 1000 according to aspects of the present disclosure is shown. The example shown includes a machine learning model 1000, a user interface 1005, an unconditional mask network 1010, a conditional mask network 1015, a distance transformation component 1020, a color distance component 1025, a mask combination component 1030, an island region removal component 1035, and a background separation process 1040. The machine learning model 1000 is a reference Figure 7 Examples of corresponding elements described or include references Figure 7 Describes aspects of the corresponding elements.

[0111] According to some embodiments of the present disclosure, the user interface 1005 is configured to obtain an input image and an approximate mask, wherein the approximate mask approximately indicates a foreground region of the input image. Figure 7 Examples of corresponding elements described or include references Figure 7 In some examples, the input image is generated by a text effect model such as a diffusion model (see Figure 7 ).

[0112] The unconditional mask network 1010 is configured to generate an unconditional mask of the foreground region based on the input image. The conditional mask network 1015 is configured to generate a conditional mask of the foreground region based on the input image and the approximate mask. Figure 7 Examples of corresponding elements described or include Figure 7 Conditional mask network 1015 is referenced Figure 7 Examples of corresponding elements described or include references Figure 7 Describes aspects of the corresponding elements.

[0113] In one example, the input image is an image of text including one or more characters based on a specific font, wherein additional shapes or textures (e.g., textures generated by a generative machine learning model) decorate the text. In this example, the approximate mask can be a font mask based on the original text, but without additional shapes or textures. However, the approximate mask can represent any approximation of a foreground object. The input image can be used alone to generate an unconditional mask, while the approximate mask (e.g., a font mask) is used to generate a conditional mask. In some examples, the conditional mask is more likely to reflect the original shape (e.g., the original font shape), but may not incorporate some additional shapes or textures. In contrast, the unconditional mask can capture additional shapes or textures, but may not correspond to the original shape. By incorporating two masks, an embodiment of the present disclosure achieves a balanced result that reflects the desired shape while also capturing additional shapes and textures.

[0114] In some cases, additional factors such as a distance transform map or a color distance map can be used to combine the conditional mask and the unconditional mask. According to some embodiments, the distance transform component 1020 calculates the distance transform map based on the approximation mask, wherein the output image is generated based on the distance transform map. The distance transform component 1020 of the machine learning model 1000 is configured to calculate or generate another estimate of the foreground probability mask via the distance by calculating the distance transform using the font mask.

[0115] According to some embodiments, the color distance component 1025 calculates a color distance map based on the approximation mask, wherein the output image is generated based on the color distance map. The color distance component 1025 generates another estimate of the foreground probability mask by calculating the distance and the ratio of the average of the pixel values ​​inside the font mask and outside the font mask.

[0116] The mask combination component 1030 of the machine learning model 1000 combines the unconditional foreground probability mask, the conditional foreground probability mask, the foreground probability mask via distance transform, and the foreground probability mask via color into a combined estimate for the foreground mask. The combined mask more accurately surrounds the edges than the earlier foreground mask estimates. The mask combination component 1030 is a reference Figure 7 Examples of corresponding elements described, or including references Figure 7 Describes aspects of the corresponding elements.

[0117] In some examples, the mask combination component 1030 generates a combined mask as follows: Let: X=distancetransform map+conditional mask+color map, Y=unconditional mask, and the final mask is formulated as final mask=f(X)*g(Y)+C, where f and g are linear functions and C is a constant.

[0118] In some examples, the color map is the ratio of the distance to the foreground color to the distance to the background color. The foreground color is obtained by averaging all pixels located within the approximation mask, while the background color is obtained by averaging pixels located outside the approximation mask. The distance transform map is zero outside the approximation mask. Within the region defined by the approximation mask, pixels closer to the center (away from the border of the approximation mask) have higher values.

[0119] In some embodiments, the island removal component 1035 performs island removal on the input image, wherein the output image is generated based on the island removal. The island removal operation is applied to eliminate small objects located outside the area defined by the font mask via the island removal component 1035. The island removal is configured to find the connected component of the mask, and if the connected area overlaps with the approximate mask, the island removal component 1035 is configured to retain it. Otherwise, the island removal component 1035 is configured to discard it. The connected area, or object, in a binary image is a set of adjacent pixels. Determining which pixels are adjacent depends on how pixel connectivity is defined. For two-dimensional images, there are two types of connectivity. They are 4-connectivity (if the edges of the pixels touch, the pixels are connected) and 8-connectivity (if the edges or corners of the pixels touch, the pixels are connected).

[0120] In some embodiments, the ground truth masks are used to train the models, i.e., the unconditional mask network 1010 and the conditional mask network 1015 .

[0121] In some embodiments, the output image is generated based on the unconditional mask and the conditional mask via a background separation process, and the output image includes the foreground region of the input image.

[0122] Fig.11 An example of a diffusion process 1100 according to aspects of the present disclosure is shown. The example shown includes the diffusion process 1100, a forward diffusion process 1105, a backward diffusion process 1110, a noise image 1115, a first intermediate image 1120, a second intermediate image 1125, and an original image 1130.

[0123] As above reference Figure 8 As described above, the diffusion model may include a forward diffusion process 1105 for adding noise to an image (or a feature in a latent space) and a backward diffusion process 1110 for denoising the image (or feature) to obtain a denoised image. The forward diffusion process 1105 may be represented as q(x t ∣x t-1 ), and the reverse diffusion process 1110 can be expressed as p(x t-1 ∣x t ). In some cases, the forward diffusion process 1105 is used during training to generate images with successively larger noise, and the neural network is trained to perform the backward diffusion process 1110 (ie, to successively remove noise).

[0124] In the example forward pass for the latent diffusion model, the model uses a Markov chain to map the observed variable x0 (in pixel space or latent space) to intermediate variables x1,…,x T When the latent variable is passed through a neural network such as U-net, the Markov chain gradually adds Gaussian noise to the data to obtain an approximate posterior q(x 1∶T |x0), where x1,…,x T has the same dimensions as x0.

[0125] The neural network can be trained to perform the reverse process. During the back diffusion process 1110, the model is fed with noisy data x T (such as the noisy image 1115), and denoise the data to obtain p(x t-1 ∣x t ). At each step t-1, the back diffusion process 1110 uses x t(such as the first intermediate image 1120) and t as input. Here, t represents a step in the transformation sequence associated with different noise levels. The back diffusion process 1110 iteratively outputs x t-1 (such as the second intermediate image 1125), until x T is returned to x0, the original image 1130. The reverse process can be expressed as:

[0126] p θ (x t-1 |x t ):=N(x t-1 ;μ θ (x t , t), ∑ θ (x t , t)). (1)

[0127] The joint probability of a sample sequence in a Markov chain can be written as the product of the conditional probability and the marginal probability:

[0128]

[0129] Where p(x T )=N(x T 0, I) is a pure noise distribution because the reverse process takes the result of the forward process (a sample of pure noise) as input and Represents the Gaussian transformed sequence corresponding to a sequence with Gaussian noise added to samples.

[0130] At the interference time, the observation data x0 in the pixel space can be mapped into the latent space as input, and the generated data Mapping from the latent space back to the pixel space as the output. In some examples, x0 represents the original input image with low image quality, and the latent variables x1, ..., x T represents a noisy image, and Indicates the generated image with high image quality.

[0131] The forward diffusion process 1105 is referenced Figure 8 Examples of corresponding elements described, or including references Figure 8 The reverse diffusion process 1110 is a reference to the corresponding elements of the description. Figure 8 Examples of corresponding elements described, or including references Figure 8 The original image 1130 is a reference to the corresponding elements of the Figure 8 Examples of corresponding elements described, or including references Figure 8 Describes aspects of the corresponding elements.

[0132] Training and Evaluation

[0133] exist Figure 12-13 In the invention, a method, apparatus, and non-transitory computer-readable medium for image processing are described. One or more embodiments of the method, apparatus, and non-transitory computer-readable medium include: initializing an unconditional mask network and a conditional mask network; receiving training data, the training data including an input image, an approximate mask indicating a foreground region of the input image, and a ground truth mask; using the training data, training the unconditional mask network to generate an unconditional mask of the foreground region based on the input image; and using the training data, training the conditional mask network to generate a conditional mask of the foreground region based on the input image and the approximate mask.

[0134] Some examples of methods, apparatus, and non-transitory computer-readable media further include combining the unconditional mask and the conditional mask to obtain a combined mask, wherein the output image is generated based on the combined mask.

[0135] In some examples, the foreground region includes text based on a font and modified using a text effect, and wherein the approximate mask is based on the text and the font without the text effect. In some examples, the conditional mask includes a refined probabilistic mask.

[0136] Fig.12 An example of a method 1200 for training a diffusion model via forward and backward diffusion according to aspects of the present disclosure is shown. The method 1200 represents training as described above with reference to Fig.11 In some examples, these operations are performed by a system including a processor that executes a code set to control a device (such as Figure 7 Functional elements of the image processing device 700) described in .

[0137] Additionally or alternatively, some processes of method 1200 may be performed using dedicated hardware. Typically, these operations are performed according to the methods and processes described in various aspects of the present disclosure. In some cases, the operations described herein are composed of various sub-steps or are performed in conjunction with other operations.

[0138] At operation 1205, the user initializes the untrained model. Initialization may include defining the architecture of the model and establishing initial values ​​for model parameters. In some cases, initialization may include defining hyperparameters such as the number of layers, the resolution and channels of each layer block, the location of skip connections, etc.

[0139] At operation 1210, the system adds noise to the training image using a forward diffusion process of N stages. In some cases, the forward diffusion process is a fixed process in which Gaussian noise is continuously added to the image. In a latent diffusion model, Gaussian noise can be continuously added to features in the latent space.

[0140] At operation 1215, the system starts from stage N, and at each stage n, a back diffusion process is used to predict the image or image features at stage n-1. For example, the back diffusion process can predict the noise added by the forward diffusion process, and the predicted noise can be removed from the image to obtain the predicted image. In some cases, the original image is predicted at each stage of the training process.

[0141] At operation 1220, the system compares the image (or image features) predicted at stage n-1 with the actual image (or image features) (such as the image at stage n-1 or the original input image). For example, given observation data x, a diffusion model can be trained to transform the negative log-likelihood of the training data -logp θ Minimize the variational upper bound of (x).

[0142] At operation 1225, the system updates the parameters of the model based on the comparison. For example, the parameters of the U-net can be updated using gradient descent. The time-dependent parameters of the Gaussian transformation can also be learned.

[0143] Fig.13 An example of a method 1300 for training a machine learning model according to aspects of the present disclosure is shown. In some examples, the operations are performed by a system including a processor that executes a code set to control functional elements of the device. Additionally or alternatively, certain processes are performed using dedicated hardware. Typically, the operations are performed according to the methods and processes described in accordance with aspects of the present disclosure. In some cases, the operations described herein are composed of various sub-steps or are performed in conjunction with other operations.

[0144] Supervised learning is one of the three basic machine learning paradigms in addition to unsupervised learning and reinforcement learning. Supervised learning is a machine learning technique based on learning a function that maps input to output based on example input-output pairs. Supervised learning generates a function for predicting labeled data based on labeled training data including a training example set. In some cases, each example is a pair consisting of an input object (usually a vector) and an expected output value (i.e., a single value or output vector). The supervised learning algorithm analyzes the training data and generates an inference function that can be used to map new examples. In some cases, learning generates a function that correctly determines the class label of an unseen instance. In other words, the learning algorithm generalizes from the training data to unseen examples.

[0145] Thus, during the training process, the parameters and weights of the machine learning model are adjusted to increase the accuracy of the results (i.e., by trying to minimize a loss function that corresponds in some way to the difference between the current result and the target result). The weights of the edges increase or decrease the strength of the signal transmitted between the nodes. In some cases, the nodes have a threshold below which the signal is not transmitted. In some examples, the nodes are aggregated into layers. Different layers perform different transformations on their inputs. The initial layer is called the input layer, and the last layer is called the output layer. In some cases, the signal traverses certain layers multiple times.

[0146] At operation 1305, the system initializes the unconditional mask network and the conditional mask network. In some cases, the operation of this step involves reference Figure 7 The training components described or can be referenced by Figure 7 The training components described are executed.

[0147] At operation 1310, the system receives training data, the training data comprising an input image, an approximation mask indicating a foreground region of the input image, and a ground truth mask. In some cases, the operation of this step involves reference Figure 7 The training components described or can be referenced by Figure 7 The training components described are executed.

[0148] At operation 1315, the system trains an unconditional mask network using the training data to generate an unconditional mask of the foreground region based on the input image. In some cases, the operation of this step involves reference Figure 7 The training components described or can be referenced by Figure 7 In some embodiments, the unconditional mask network is excluded from receiving the approximate mask.

[0149] At operation 1320, the system trains a conditional mask network using the training data to generate a conditional mask of the foreground region based on the input image and the approximate mask. In some cases, the operation of this step involves reference Figure 7 The training components described or can be referenced by Figure 7 In some embodiments, the conditional mask network receives an approximate mask indicating a foreground region of an input image.

[0150] In some embodiments, the unconditional mask network and the conditional mask network are trained using training data for background separation tasks. In some examples, a pre-trained image generation model (e.g., a diffusion model) is used to generate the input image. Other types of generation models can also be used to generate the input image.

[0151] Fig.14An example of a computing device 1400 according to aspects of the present disclosure is shown. The example shown includes the computing device 1400, (multiple) processors 1405, a memory subsystem 1410, a communication interface 1415, an I / O interface 1420, (multiple) user interface components 1425, and a channel 1430. In one embodiment, the computing device 1400 includes (multiple) processors 1405, a memory subsystem 1410, a communication interface 1415, an I / O interface 1420, (multiple) user interface components 1425, and a channel 1430.

[0152] In some embodiments, computing device 1400 is Figure 1 An example of the image processing device 110 may include Figure 1 In some embodiments, the computing device 1400 includes one or more processors 1405, and the one or more processors 1405 can execute instructions stored in the memory subsystem 1410 to obtain an input image and an approximate mask, wherein the approximate mask approximately indicates a foreground region of the input image; generate an unconditional mask of the foreground region based on the input image through an unconditional mask network; generate a conditional mask of the foreground region based on the input image and the approximate mask through a conditional mask network; and generate an output image including the foreground region of the input image based on the unconditional mask and the conditional mask.

[0153] According to some embodiments, computing device 1400 includes one or more processors 1405. In some cases, the processor is an intelligent hardware device (e.g., a general processing component, a digital signal processor (DSP), a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller, an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a programmable logic device, a discrete gate or transistor logic component, a discrete hardware component, or a combination thereof). In some cases, the processor is configured to operate a memory array using a memory controller. In other cases, the memory controller is integrated into the processor. In some cases, the processor is configured to execute computer-readable instructions stored in the memory to perform various functions. In some embodiments, the processor includes dedicated components for modem processing, baseband processing, digital signal processing, or transmission processing.

[0154] According to some embodiments, the memory subsystem 1410 includes one or more memory devices. Examples of memory devices include random access memory (RAM), read-only memory (ROM), or hard disk. Examples of memory devices include solid-state memory and hard disk drive. In some examples, the memory is used to store computer-readable-computer executable software including instructions, which, when executed, cause the processor to perform various functions described herein. In some cases, the memory particularly includes a basic input / output system (BIOS), which controls basic hardware or software operations, such as interactions with peripheral components or devices. In some cases, a memory controller operates a memory cell. For example, a memory controller may include a row decoder, a column decoder, or both. In some cases, a memory cell within a memory stores information in the form of a logical state.

[0155] According to some embodiments, the communication interface 1415 operates at the boundary between the communication entity (such as the computing device 1400, one or more user devices, the cloud, and one or more databases) and the channel 1430 and can record and process the communication. In some cases, the communication interface 1415 is provided to enable a processing system coupled to a transceiver (e.g., a transmitter and / or a receiver). In some examples, the transceiver is configured to transmit (or send) and receive signals for the communication device via an antenna.

[0156] According to some embodiments, I / O interface 1420 is controlled by an I / O controller to manage input and output signals for computing device 1400. In some cases, I / O interface 1420 manages peripheral devices that are not integrated into computing device 1400. In some cases, I / O interface 1420 represents a physical connection or port to an external peripheral device. In some cases, an I / O controller uses an operating system, such as or other known operating systems. In some cases, an I / O controller represents or interacts with a modem, keyboard, mouse, touch screen, or similar device. In some cases, an I / O controller is implemented as a component of a processor. In some cases, a user interacts with the device via an I / O interface 1420 or via a hardware component controlled by an I / O controller.

[0157] According to some embodiments, user interface component(s) 1425 enable a user to interact with computing device 1400. In some cases, user interface component(s) 1425 include an audio device such as an external speaker system, an external display device such as a display screen, an input device (e.g., a remote control device that interfaces with a user interface directly or via an I / O controller), or a combination thereof. In some cases, user interface component 1425 includes a GUI.

[0158] The performance of the apparatus, system and method of the present disclosure has been evaluated and the results indicate that the embodiments of the present disclosure achieve improved performance over the prior art. Example experiments demonstrate that the image processing apparatus based on the present disclosure outperforms conventional systems.

[0159] The descriptions and drawings described herein represent example configurations and do not represent all implementations within the scope of the claims. For example, operations and steps may be rearranged, combined, or otherwise modified. In addition, structures and devices may be represented in the form of block diagrams to represent the relationships between components and avoid blurring the concepts described. Similar components or features may have the same name, but may have different reference numerals corresponding to different drawings.

[0160] Some modifications to the present disclosure will be apparent to those skilled in the art and the principles defined herein may be applied to other variations without departing from the scope of the present disclosure. Therefore, the present invention is not limited to the examples and designs described herein, but should be in accordance with the broadest scope consistent with the principles and novel features disclosed herein.

[0161] The described method can be implemented or executed by a device including a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic, a discrete hardware component or any combination thereof. The general-purpose processor can be a microprocessor, a conventional processor, a controller, a microcontroller or a state machine. The processor can also be implemented as a combination of computing devices (for example, a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in combination with a DSP core, or any other such configuration). Therefore, the functions described herein can be implemented in hardware or software and can be performed by a processor, firmware or any combination thereof. If implemented in software executed by a processor, the functions can be stored in the form of instructions or codes on a computer-readable medium.

[0162] Computer-readable media include non-transitory computer storage media and communication media with any medium that facilitates code or data transmission. Non-transitory storage media can be any available medium that can be accessed by a computer. For example, non-transitory computer-readable media can include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), compact disc (CD) or other optical disk storage device, magnetic disk storage device or any other non-transitory medium for carrying or storing data or code.

[0163] In addition, a connecting component may be appropriately referred to as a computer-readable medium. For example, if code or data is transmitted from a website, server or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology such as infrared, radio or microwave signals, the coaxial cable, fiber optic cable, twisted pair, DSL or wireless technology is included in the definition of medium. Combinations of media are also included within the scope of computer-readable media.

[0164] In the present disclosure and the appended claims, the word "or" indicates an inclusive list, so that, for example, a list of X, Y, or Z means X or Y or Z or XY or XZ or YZ or XYZ. The phrase "based on" is also not used to indicate a closed set of conditions. For example, a step described as "based on condition A" can be based on both condition A and condition B. In other words, the phrase "based on" should be interpreted to mean "based at least in part on". In addition, the word "a" or "an" indicates "at least one".

Claims

1. A method comprising: Acquire an input image and an approximate mask, wherein the approximate mask approximately indicates a foreground region of the input image; Generating an unconditional mask of the foreground region based on the input image through an unconditional mask network; Generating a conditional mask of the foreground region based on the input image and the approximate mask through a conditional mask network; as well as An output image including the foreground region of the input image is generated based on the unconditional mask and the conditional mask.

2. The method of claim 1 , wherein generating the output image comprises: The unconditional mask and the conditional mask are combined to obtain a combined mask, wherein the output image is generated based on the combined mask.

3. The method according to claim 1, wherein: The input image is generated based on the approximate mask.

4. The method of claim 1 , wherein generating the output image comprises: A distance transform map is calculated based on the approximation mask, wherein the output image is generated based on the distance transform map.

5. The method of claim 1 , wherein generating the output image comprises: A color distance map is calculated based on the approximation mask, wherein the output image is generated based on the color distance map.

6. The method of claim 1 , wherein generating the output image comprises: An island area removal is performed on the input image, wherein the output image is generated based on the island area removal.

7. The method according to claim 1, wherein: The foreground region includes text that is based on a font and modified with a text effect, and wherein the approximate mask is based on the text and the font without the text effect.

8. The method according to claim 1, wherein: The unconditional mask comprises a probabilistic mask and the conditional mask comprises a refined probabilistic mask.

9. The method according to claim 1, wherein: The unconditional mask network is trained to generate the unconditional mask of the foreground region based on the input image, and the conditional mask network is trained to generate the conditional mask of the foreground region based on the input image and the approximate mask.

10. A method comprising: Initialize the unconditional mask network and the conditional mask network; receiving training data, the training data comprising an input image, an approximation mask indicating a foreground region of the input image, and a ground truth mask; Using the training data to train the unconditional mask network to generate an unconditional mask of the foreground region based on the input image; as well as The conditional mask network is trained using the training data to generate a conditional mask of the foreground region based on the input image and the approximate mask.

11. The method according to claim 10, further comprising: The unconditional mask and the conditional mask are combined to obtain a combined mask, wherein an output image is generated based on the combined mask.

12. The method according to claim 10, wherein: The foreground region includes text that is based on a font and modified with a text effect, and wherein the approximate mask is based on the text and the font without the text effect.

13. The method of claim 10, wherein: The conditional mask includes a refined probability mask.

14. An apparatus comprising: at least one processor; at least one memory including instructions executable by the at least one processor; a user interface comprising parameters stored in the at least one memory and configured to acquire an input image and an approximate mask approximately indicating a foreground region of the input image; an unconditional mask network comprising parameters stored in the at least one memory and configured to generate an unconditional mask of the foreground region based on the input image; as well as A conditional mask network comprising parameters stored in the at least one memory and configured to generate a conditional mask of the foreground region based on the input image and the approximate mask.

15. The apparatus according to claim 14, further comprising: An image generation model is configured to generate an output image including the foreground region of the input image based on the unconditional mask and the conditional mask.

16. The apparatus according to claim 15, further comprising: A distance transform component is configured to compute a distance transform map based on the approximation mask, wherein the output image is generated based on the distance transform map.

17. The apparatus according to claim 15, further comprising: A color distance component is configured to calculate a color distance map based on the approximation mask, wherein the output image is generated based on the color distance map.

18. The apparatus according to claim 15, further comprising: An island area removal component is configured to perform island area removal on the input image, wherein the output image is generated based on the island area removal.

19. The apparatus according to claim 15, further comprising: A mask combination component is configured to combine the unconditional mask with the conditional mask to obtain a combined mask, wherein the output image is generated based on the combined mask.

20. The apparatus of claim 14, further comprising: A text effect model is configured to generate the input image.