Image augmentation based on image-to-image generation

The image augmentation method addresses the issue of monotonous background images in video communication by enhancing images with decorative effects and coherence through clutter removal, color modification, and element addition, improving user immersion.

WO2025213316A1PCT designated stage Publication Date: 2025-10-16MICROSOFT TECHNOLOGY LICENSING LLC +9
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/086498
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-04-08
Publication Date
2025-10-16

AI Technical Summary

Technical Problem

Existing video communication systems often utilize monotonous or mismatched background images that fail to enhance participant immersion, lacking decorative effects and coherence with user themes.

Method used

An image augmentation method that includes removing clutter objects, modifying object colors, and adding element images based on predefined themes using an augmentation engine and image generation models like StableDiffusion or DALL·E, ensuring the generated images have decorative effects and maintain structural consistency with the original image.

Benefits of technology

Enhances immersion in video communication by generating images with decorative effects corresponding to desired themes while maintaining structural coherence, thereby improving user engagement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024086498_16102025_PF_FP_ABST
    Figure CN2024086498_16102025_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure proposes a method, apparatus and computer program product for image augmentation. A first image and a selected theme may be received. The first image may be augmented based on the selected theme, the augmenting comprising: modifying the color of an object in the first image according to the selected theme, and / or adding element images corresponding to the selected theme in the first image. A second image may be generated based on the augmented first image and the selected theme, the second image being an image with decorative effects corresponding to the selected theme.
Need to check novelty before this filing date? Find Prior Art

Description

IMAGE AUGMENTATION BASED ON IMAGE-TO-IMAGE GENERATIONBACKGROUND

[0001] With the development of digital device, communication technology, video processing technology, etc., people may use terminal devices such as desktop computer, tablet computer, smart phone, etc., to conduct video communication with people located elsewhere for purposes such as chatting, work discussions, remote training, technical support, etc. Herein, video communication broadly refers to a communication method based on Internet technology that can transmit speeches and images of participants in real time. Video communication may include, e.g., video conference, video call, video streaming, etc.SUMMARY

[0002] This Summary is provided to introduce a selection of concepts that are further described below in the Detailed Description. It is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter.

[0003] Embodiments of the present disclosure propose a method, apparatus and computer program product for image augmentation. A first image and a selected theme may be received. The first image may be augmented based on the selected theme, the augmenting comprising: modifying the color of an object in the first image according to the selected theme, and / or adding element images corresponding to the selected theme in the first image. A second image may be generated based on the augmented first image and the selected theme, the second image being an image with decorative effects corresponding to the selected theme.

[0004] It should be noted that the above one or more aspects comprise the features hereinafter fully described and particularly pointed out in the claims. The following description and the drawings set forth in detail certain illustrative features of the one or more aspects. These features are only indicative of the various ways in which the principles of various aspects may be employed, and this disclosure is intended to include all such aspects and their equivalents.BRIEF DESCRIPTION OF THE DRAWINGS

[0005] The disclosed aspects will hereinafter be described in connection with the appended drawings that are provided to illustrate and not to limit the disclosed aspects.

[0006] FIG. 1 illustrates an exemplary process for image augmentation according to an embodiment of the present disclosure.

[0007] FIG. 2 illustrates an exemplary process for modifying the color of an object in an image according to a selected theme according to an embodiment of the present disclosure.

[0008] FIG. 3 illustrates an exemplary process for adding element images corresponding to a selected theme in an image according to an embodiment of the present disclosure.

[0009] FIG. 4 illustrates an exemplary process for detecting an area for adding element images in an image according to an embodiment of the present disclosure.

[0010] FIG. 5 illustrates an exemplary process for automatically building an element image set corresponding to a theme according to an embodiment of the present disclosure.

[0011] FIG. 6 illustrates an exemplary process for automatically generating an adding rule for an element according to an embodiment of the present disclosure.

[0012] FIG. 7A to FIG. 7C illustrate an example of image augmentation according to an embodiment of the present disclosure.

[0013] FIG. 8A to FIG. 8F illustrate another example of image augmentation according to an embodiment of the present disclosure.

[0014] FIG. 9 is a flowchart of an exemplary method for image augmentation according to an embodiment of the present disclosure.

[0015] FIG. 10 illustrates an exemplary apparatus for image augmentation according to an embodiment of the present disclosure.

[0016] FIG. 11 illustrates another exemplary apparatus for image augmentation according to an embodiment of the present disclosure.DETAILED DESCRIPTION

[0017] The present disclosure will now be discussed with reference to several example implementations. It is to be understood that these implementations are discussed only for enabling those skilled in the art to better understand and thus  implement the embodiments of the present disclosure, rather than suggesting any limitations on the scope of the present disclosure.

[0018] Video communication may be implemented through communicating speeches and images among electronic devices of users participating in the video communication. An image communicated during video communication may include a foreground image and a background image. The foreground image transmitted from a device may include a user image of a user of the device. The background image may be an actual background image captured in real time by a camera of the device, or a virtual background image set by the user. The actual background image or the virtual background image may be monotonous, or may not match the user's expected theme.

[0019] Embodiments of the present disclosure propose image augmentation based on image-to-image generation. A first image may be received. The first image may be an image depicting an indoor or outdoor environment. For example, the first image may be a background image in a video communication. A selected theme may be received. The selected theme indicates the style or effect to be achieved in the first image. The selected theme is one of a set of predefined themes. Different themes in the set of predefined themes correspond to different color sets and / or different element image sets. The first image may be augmented based on the selected theme, to obtain an augmented first image. The augmented first image and the selected theme may be provided to an image generation model, to generate a second image, the second image being an image with decorative effects corresponding to the selected theme.

[0020] The augmenting the first image based on the selected theme may comprise removing clutter objects from the first image and filling in the blank regions with inpainting techniques. The technical effect of removing clutter objects is to achieve a clean and tidy environment. Alternatively or additionally, the augmenting the first image based on the selected theme may comprise modifying colors of objects in the first image according to the selected theme. Alternatively or additionally, the augmenting the first image based on the selected theme may comprise adding element images corresponding to the selected theme in the first image. Herein, an element image refers to an image containing a specific element. The image generation model primarily relies on visual cues present in an input image to generate a meaningful and coherent output image. The technical effect of modifying object colors or adding element images is to provide explicit visual cues to the image generation model, so as to produce  decorative effects corresponding to the selected theme while ensuring that the generated second image is meaningful and coherent to the first image.

[0021] Through the image augmentation proposed by the embodiments of the present disclosure, an image that reflects decorative effects corresponding to a desirable theme and has a structure or layout that is substantially consistent with the initial image can be generated. When such an approach is applied in a video communication, the immersion of participants of the video communication can be enhanced.

[0022] Various embodiments of the present disclosure will hereinafter be described in detail in connection with the appended drawings.

[0023] FIG. 1 illustrates an exemplary process 100 for image augmentation according to an embodiment of the present disclosure.

[0024] A first image 102 may be received. The first image 102 may be an image depicting an indoor or outdoor environment. By way of example but not limitation, the first image 102 may be a background image in a video communication, e.g., an actual background image or a virtual background image. The actual background image may be obtained through removing a foreground image from an image captured by a camera of a device participating in the video communication, and then filling in the blank region with inpainting techniques.

[0025] A selected theme 104 may be received. The selected theme 104 indicates the style or effect to be achieved in the first image 104. The selected theme 104 is one of a set of predefined themes. The set of predefined themes may include, e.g., “celebration” , “garden” , “fancy” , “clean” , etc. The selected theme 104 may be recommended by an application receiving the first image 102, or may be chosen by a user from the set of predefined themes based on the desired style or effect applied to the first image 102. Different themes in the set of predefined themes correspond to different color sets and / or different element image sets. For example, if the selected theme is “celebration” , a color set corresponding to the selected theme may contain red, pink, yellow, etc., and an element image set corresponding to the selected theme may contain images of balloons, lanterns, flags, etc. Instead, if the selected theme is “garden” , a color set corresponding to the selected theme may contain green, red, brown, etc., and an element image set corresponding to the selected theme may contain images of flowers, grasses, vines, etc.

[0026] The first image 102 may be augmented based on the selected theme 104  through an augmentation engine 110, to obtain an augmented first image 112. The augmented first image 112 may be provided to an image generation model 170, to generate a second image 172. The second image 172 may be an image with decorative effects corresponding to the selected theme 104.

[0027] The augmentation engine 110 may comprise an object removal module 120. The object removal module 120 may remove clutter objects from the first image 102 and fill in the blank regions with inpainting techniques. A clutter object may be an object of a predetermined category, with a size below a predetermined threshold, and / or being located at a predetermined position. Objects of predetermined categories may include, e.g., spitball, bottle, paper cup, etc. Semantic segmentation may be performed on the first image 102 through a semantic segmentation model, to segment objects and surfaces in the first image 102. The semantic segmentation model is, e.g., a Semantic Segment Anything (SSA) model. Each object may be tagged with a corresponding category. The predetermined position may be floor, wall, etc. Clutter objects may be removed from the first image 102. The technical effect of this approach is to achieve a clean and tidy environment.

[0028] The augmentation engine 110 may comprise a color modification module 130. The color modification module 130 may modify colors of objects in the first image 102 based on the selected theme 104. An exemplary process for modifying the color of an object will be described later in conjunction with FIG. 2. The image generation model primarily relies on visual cues present in an input image to generate a meaningful and coherent output image. The technical effect of this approach is to provide explicit visual cues to the image generation model 170, so as to produce decorative effects corresponding to the selected theme 104 while ensuring that the generated second image 172 is meaningful and coherent to the first image 102.

[0029] The augmentation engine 110 may comprise an element addition module 140. The element addition module 140 may detect an area for adding element images in the first image 102, and add element images corresponding to the selected theme 104 in the area. An exemplary process for adding element images will be described later in conjunction with FIG. 3. The technical effect of this approach is to provide explicit visual cues to the image generation model 170, so as to produce decorative effects corresponding to the selected theme 104 while ensuring that the generated second image 172 is meaningful and coherent to the first image 102. The area for adding  element images may be determined based on a surface mask and a smoothness mask of the first image 102. In the case where the first image 102 is a background image in a video communication, when detecting the area for adding element images, a foreground mask 106 corresponding to the first image 102 may also be taken into account. An exemplary process for detecting the area for adding element images will be described later in conjunction with FIG. 4.

[0030] After the processing of the augmentation engine 110, the augmented first image 112 may be obtained. The augmented first image 112 may be provided to the image generation model 170. The image generation model 170 may be, e.g., StableDiffusion model, DALL·E model, Midjourney model, etc. The image generation model 170 may generate the second image 172 based on the augmented first image 112 and the selected theme 104. For example, the image generation model 170 may encode the augmented first image 112, to generate an image vector of the augmented first image 112, and then decode the image vector according to the selected theme 104, to generate the second image 172. The second image 172 may be an image with decorative effects corresponding to the selected theme 104. For example, the second image 172 may contain decorative elements corresponding to the selected theme 104. These decorative elements may not be included in the first image 102. Alternatively or additionally, some objects in the second image 172 may be decorated to be consistent with the selected theme 104. These objects may be objects present in the first image 102.

[0031] A prompt obtaining module 150 may randomly select a prompt 152 from a prompt set corresponding the selected theme 104. A prompt set corresponding a theme may contain a set of prompts predefined for the theme. Each prompt may include, e.g., a description of the theme, image quality requirements, etc. Taking the selected theme 104 being “clean” as an example, the prompt for the selected theme 104 may be “A neat room with minimalist decoration style. The environment is clean and tidy. Spotless floor. DLSR. HQ” . The image generation model 170 may decode the image vector of the augmented first image 112 according to the prompt 152, to generate the second image 172.

[0032] Preferably, when generating the second image 172 through the image generation model 170, in addition to the selected theme 104, a strength value may also be taken into account. The strength value may also be referred to as a strength factor, which can be used to balance deviation from original image content and level of  creativeness of the image generation model. A higher strength value results in more diverse and interesting images. Conversely, a lower strength value preserves more of the original image content. A strength value calculation module 160 may calculate a strength value 162, and the image generation model 170 may generate the second image 172 based on the augmented first image 112, the selected theme 104 and the strength value 162. For example, the image generation model 170 may encode the augmented first image 112, to generate an image vector of the augmented first image 112, and then decode the image vector of the augmented first image 112 according to both the prompt 152 and the strength value 162.

[0033] The strength value calculation module 160 may calculate complexity of the augmented first image 112 based on luma information of the augmented first image 112, and calculate the strength value 162 based on the complexity of the augmented first image 112.

[0034] In an implementation, the augmented first image 112 may be converted into a YCrCb color space. The Y channel image may be divided into k×k blocks with a stride of s. This will result in n patches in total. For each patch, its luma histogram on the Y channel may be calculated. This can be done by counting the number of pixels in each luminance value bin, as illustrated by the following equation:

[0035] Then, the entropy of each patch may be calculated using the luma histogram, as illustrated by the following equation:

[0036] Next, the entropy of all patches may be averaged to achieve the image complexity, as illustrated by the following equation:

[0037] Subsequently, the strength value 162 may be calculated based on the image complexity, as illustrated by the following equation: strength value=max (v1, min (v2, C (I) ) )          (4)

[0038] where v1 and v2 (v1<v2) defines a range of the strength value 162, which can be set depending on actual application requirements. As an example, v1 may be set as “0.3” ,  and v2 may be set as “0.6” . It should be appreciated that the above equations for calculating the strength value 162 are merely exemplary. The strength value 162 can also be calculated through other equations.

[0039] The technical effect of the above approach is to achieve dynamic selection of the strength value based on the image complexity, while keeping the strength value within a predefined range, that is, the range defined by v1 and v2. A higher strength value can be applied on a complex image to achieve creativeness, while a lower strength value is more suitable for a simple image to avoid structural change, as the simple image has less details and therefore it is difficult to constrain the image generation model.

[0040] Preferably, a content moderation model may be invoked to detect whether there is inappropriate content in the second image 172, such as watermark, people, unsafety content, etc. If no, the second image 172 may be returned to the user, while if yes, the image generation model 170 may regenerate the second image 172.

[0041] If the first image 102 is a background image in a video communication, e.g., an actual background image or a virtual background image, the second image 172 may be combined with foreground images, such as user images captured in real time by the user’s camera, to form video frames for the video communication.

[0042] In the process 100, the augmented first image 112 can be obtained through augmenting the first image 102 based on the selected theme 104, e.g., removing clutter objects from the first image 102, modifying colors of objects in the first image 102 according to the selected theme 104, adding element images corresponding to the selected theme 104 in the first image 102, etc. The second image 172 can be generated based on the augmented first image 112, the selected theme 104, and / or the strength value 162 through the image generation model 170. The second image 172 can contain decorative elements corresponding the selected theme 104. The technical effect of the above approach is to generate an image that reflects decorative effects corresponding to a desirable theme and has a structure or layout that is substantially consistent with the initial image. When such an approach is applied in a video communication, the immersion of participants of the video communication can be enhanced.

[0043] It should be appreciated that the process 100 in FIG. 1 is merely an example of the process for image augmentation. Depending on actual application requirements, the steps in the process for image augmentation may be replaced or modified in any manner, and the process may comprise more or fewer steps. For example, although in  FIG. 1, the object removal module 120, the color modification module 130 and the element addition module 140 are illustrated. However, depending on actual application requirements, the augmentation engine 110 may include more or fewer modules. In addition, the object removal module 120, the color modification module 130 and the element addition module 140 may operate independently or in combination with each other. Furthermore, the specific order or hierarchy of the steps in the process 100 is merely exemplary, and the process for image augmentation may be performed in an order different from the described order.

[0044] FIG. 2 illustrates an exemplary process 200 for modifying the color of an object in an image according to a selected theme according to an embodiment of the present disclosure. The process 200 may be executed by the color modification module 130 in FIG. 1.

[0045] At 202, an object may be identified from a first image. The first image may correspond to the first image 102 in FIG. 1. The object may be of a predetermined category, with a size below a predetermined threshold, and / or being located at a predetermined position. Objects of predetermined categories may include objects without inherent colors. For example, a flower pot, a cup, etc. may be objects without inherent colors, while flame, traffic light, etc. may be objects with inherent colors. The predetermined position may be floor, wall, etc.

[0046] At 204, a color may be randomly selected from a color set corresponding to a selected theme. The color set may also be referred to as a color palette. One or more color sets may be predefined for each theme, each color set containing a plurality of colors suitable for the theme. As an example, a color set corresponding to a theme “celebration” may contain red, pink, yellow, etc., while a color set corresponding to a theme “garden” may contain green, red, brown, etc.,

[0047] At 206, the color of the object may be changed to the randomly selected color.

[0048] Preferably, at 208, at least one of color grading, luma tuning and alpha blending may be performed on the object with the modified color. These operations can make the object more harmonious with the first image.

[0049] The transparency of the object with the modified color in the first image may be represented by an alpha value. The embodiments of the present disclosure propose to dynamically determine an alpha value for performing alpha  blending on the object with the modified color. The alpha value may be calculated as follows:

[0050] where represent a preset base alpha value, and h (iori, lcolor) represents an adjustment function for determining an adjustment alpha value based on image features of an original image Iori, i.e., the first image, and image features of the object with the modified color Icolor. As an example, the image features may be color, contrast, etc. The adjustment alpha value may be determined through calculating the color difference between the object and the first image and / or the contrast difference between the object and the first image.

[0051] The process for performing the alpha blending on the object with the modified color may be represented as follows:

[0052] where represents the image containing the object with the modified color.

[0053] The technical effect of this approach is to dynamically determine the optimal alpha value, which can ensure seamless integration of the object with the modified color into the first image.

[0054] It should be appreciated that the process 200 in FIG. 2 is merely an example of the process for modifying the color of an object in an image according to a selected theme. Depending on actual application requirements, the steps in the process for modifying the color of an object in an image may be replaced or modified in any manner, and the process may comprise more or fewer steps. For example, in the process 200, although the step 210 for performing color grading, luma tuning and alpha blending on the object is illustrated, in some embodiments, this step may be omitted. In addition, the specific order or hierarchy of the steps in the process 200 is merely exemplary, and the process for modifying the color of an object in an image may be performed in an order different from the described order.

[0055] FIG. 3 illustrates an exemplary process 300 for adding element images corresponding to a selected theme in an image according to an embodiment of the present disclosure. The process 300 may be executed by the element addition module 140 in FIG. 1. In the process 300, an area for adding element images in a first image may be detected, and then element images corresponding to a selected theme may be  added in the area.

[0056] At 302, an area for adding element images in a first image may be detected. The first image may correspond to the first image 102 in FIG. 1. The area for adding element images may be a blank or monotonous area containing no object or very few objects. An exemplary process for detecting an area for adding element images in an image will be described later in conjunction with FIG. 4.

[0057] Subsequently, element images may be iteratively added in the area, until the proportion of the added element images in the area exceeds a predetermined threshold. In each iteration, at 304, an element image may be randomly selected from an element image set corresponding to a selected theme. Element image sets corresponding to respective themes may be pre-built in an automatic manner. An element image set corresponding to a theme may contain a plurality of images of elements relevant to the theme. As an example, an element image set corresponding to a theme “celebration” may contain a plurality of images of balloons, lanterns, flags, etc.; while an element image set corresponding to a theme “garden” may contain a plurality of images of flowers, grasses, vines, etc. An exemplary process for automatically building an element image set corresponding to a theme will be described later in conjunction with FIG. 5.

[0058] At 306, the element image may be added in the area according to an adding rule for an element corresponding to the element image. Adding rules for respective elements may be pre-generated in an automatic manner. An adding rule for an element specifies values of a set of properties that should be complied with when adding an element image of the element in an image. The set of properties may include, e.g., position, rotation angle, scale, quality, number, area, specific category, etc. An exemplary process for automatically generating an adding rule for an element will be described later in conjunction with FIG. 6.

[0059] At 308, it may be determined whether the proportion of the added element images in the area exceeds a predetermined threshold. For example, it may be determined whether the total area of the added element images in the area exceeds the predetermined threshold. The predetermined threshold is, e.g., 10%.

[0060] If it is determined at 308 that the proportion of the added element images in the area does not exceed the predetermined threshold, the process 300 returns to the step 304, and continues the subsequent steps.

[0061] If it is determined at 308 that the proportion of the added element images in the area exceeds the predetermined threshold, the operation for adding element images may stop and the process 300 proceeds to a step 310. At 310, at least one of color grading, luma tuning, boundary blurring and alpha blending may be performed on the added element images. These operations can make the added element images more harmonious with the first image. The boundary blurring may be Gaussian blurring, which can smooth boundaries of the added element images.

[0062] The transparency of an added element image on the first image may be represented by an alpha value. The embodiments of the present disclosure propose to dynamically determine an alpha value for performing alpha blending on an element image. The process for dynamically determining this alpha value may be similar to that for dynamically determining the alpha value for performing alpha blending on the object with the modified color as described in conjunction with FIG. 2.

[0063] For example, an alpha value for performing the alpha blending on an element image may be calculated as follows:

[0064] where represent a preset base alpha value, and h (Iori, Iaddon) represents an adjustment function for determining an adjustment alpha value based on image features of an original image Iori, i.e., the first image, and image features of an add-on image Iaddon, i.e., an added element image. As an example, the image features may be color, contrast, etc. The adjustment alpha value may be determined through calculating the color difference between the element image and the first image and / or the contrast difference between the element image and the first image.

[0065] The process for performing the alpha blending on an element image may be represented as follows:

[0066] where represents the image with the added element image.

[0067] The technical effect of this approach is to dynamically determine the optimal alpha value, which can ensure seamless integration of the added element images with the first image.

[0068] It should be appreciated that the process 300 in FIG. 3 is merely an example of the process for adding element images corresponding to a selected theme in an image.  Depending on actual application requirements, the steps in the process for adding element images may be replaced or modified in any manner, and the process may comprise more or fewer steps. For example, in the process 300, although the step 310 for performing color grading, luma tuning, boundary blurring and alpha blending on the added element images is illustrated, in some embodiments, this step may be omitted. In addition, the specific order or hierarchy of the steps in the process 300 is merely exemplary, and the process for adding element images may be performed in an order different from the described order.

[0069] FIG. 4 illustrates an exemplary process 400 for detecting an area for adding element images in an image according to an embodiment of the present disclosure. The process 400 may correspond to the step 302 in FIG. 3.

[0070] At 410, semantic segmentation may be performed on the first image 402, to obtain a surface mask 412 Msurface of a first image 402. The first image 402 may correspond to the first image 102 in FIG. 1. The semantic segmentation may be performed through a semantic segmentation model. The semantic segmentation model is, e.g., the SSA model. The surface mask indicates regions of surfaces in the first image 102. The surfaces include, e.g., wall, ceiling, ground, etc.

[0071] At 420, pixel variances of image blocks in the first image 402 may be calculated, to obtain a smoothness mask 422 Msmooth of the first image 402. For each image block, pixel variance of the image block may be calculated. If the pixel variance is below a predetermined threshold, the image block may be classified as a smooth image block. The smoothness mask 422 indicates regions of smooth image blocks in the first image 402.

[0072] Subsequently, the area for adding element images may be determined based on the surface mask 412 Msurface and the smoothness mask 422 Msmooth. For example, at 430, an intersection operation may be performed on the surface mask 412 Msurfaceand the smoothness mask 422 Msmooth, to obtain an area mask 432 M, as illustrated by the following equation: M=Msurface∩Nsnooth         (9)

[0073] The technical effect of determining the area for adding element images based on the surface mask and the smoothness mask is to accurately detect a blank or monotonous area that contains no object or very few objects.

[0074] In the case where the first image 402 is a background image in a video  communication, when detecting the area for adding element images, a foreground mask 404 Mforeground corresponding to the first image 402 may also be taken into account. The foreground mask 404 Mforeground indicates the corresponding region of a foreground image in the first image 402. If the first image 402 is a virtual background image, the foreground mask 404 Mforeground may be obtained through mapping the user image to the first image 402. If the first image 402 is an actual background image, the foreground mask 404 Mforeground may be obtained when performing foreground and background segmentation on a video frame containing the first image 402. The area for adding element images may be determined based on the surface mask 412 Msurface, the smoothness mask 422 Msmooth and the foreground mask 404 Mf. For example, at 430, an intersection operation may be performed on the surface mask 412 Msurface, the smoothness mask 422 Msmooth, and the region outside the foreground mask 404 Mforeground in the first image 402, i.e., (1-Mforeground, to obtain an area mask 432 M, as illustrated by the following equation: M=Msurface∩Msmooth1-Mforeground)        (10)

[0075] The technical effect of determining the area for adding element images based on the surface mask, the smoothness mask and the foreground mask is to accurately detect a blank or monotonous area that contains no object or very few objects, while avoiding the foreground region that usually corresponds to the user image, so that the added element images are not obscured by the user image.

[0076] It should be appreciated that the process 400 in FIG. 4 is merely an example of the process for detecting an area for adding element images in an image. Depending on actual application requirements, the steps in the process for detecting an area for adding element images may be replaced or modified in any manner, and the process may comprise more or fewer steps. In addition, the specific order or hierarchy of the steps in the process 400 is merely exemplary, and the process for detecting an area for adding element images in an image may be performed in an order different from the described order.

[0077] FIG. 5 illustrates an exemplary process 500 for automatically building an element image set corresponding to a theme according to an embodiment of the present disclosure. The element image set may be used when adding element images in a blank or monotonous area in an image. In the process 500, an element list corresponding to a  theme may be obtained, and then element images of the element list may be generated, thereby building the element image set corresponding to the theme.

[0078] Multiple images relevant to a theme 502 may be searched through a search engine 510. The searched images may be denoted as an image 512-1, an image 512-2, …, an image 512-m, where m is the number of searched images.

[0079] For each image in the multiple images, an image description of the image may be generated through a large language model 520. Herein, a Large Language Model (LLM) refers to a deep learning model that can understand the meaning of natural language, generate natural language texts, or perform other natural language tasks. It should be appreciated that large language models include a multi-modal model that can perform processing tasks for natural language as well as other modalities. As an example, the large language model 520 may be ChatGPT. The large language model 520 may receive multi-modal inputs such as an image and a corresponding prompt, and output an image description of the image. The prompt might be “describe the image, list the objects, the objects’ scale, and the position of the objects” . The generated image description may be denoted as an image description 522-1, an image description 522-2, …, an image description 522-m.

[0080] Subsequently, for each image description, elements for decoration may be exacted from the image description through a large language model 530. The large language model 530 may be the same or different from the large language model 520. The extracted elements may be denoted as elements 532-1, elements 532-2, …, elements 532-m. These elements may form a first set of elements 532.

[0081] In addition, a second set of elements 542 relevant to the theme 502 and for decoration may be directly obtained through a large language model 540. The large language model 540 may be the same or different from the large language model 520. Taking the theme 502 being “garden” as an example, a prompt provides to the large language model 540 may be “list 10 typical elements usually used in interior garden decorated room” .

[0082] Next, an element list 552 corresponding to the theme 502 may be determined based on the first set of elements 532 and the second set of elements 542. For example, a filtering and ranking module 550 may filter out common elements appearing in both the first set of elements 532 and the second set of elements 542, and rank the common elements based on occurrence frequencies of the common elements in the multiple  image descriptions 522-1 to 522-m. A predetermined number of highest-ranked elements among the common elements may be selected, to form the element list 552 of the theme 502. Taking the theme 502 being “celebration” , the element list 552 of the theme 502 might be [ “balloon” , “lantern” , “flag” ] .

[0083] For each element in the element list 552, element images for the element may be generated through an image generation model 560, to obtain an element image set 562 corresponding to the theme 502. The image generation model 560 may be, e.g., StableDiffusion model, DALL·E model, Midjourney model, etc.

[0084] The technical effects of this approach are to automatically build an element image set corresponding to a theme without human intervention, while ensuring that the element image set contains element images relevant to the theme and suitable for decoration.

[0085] It should be appreciated that the process 500 in FIG. 5 is merely an example of the process for automatically building an element image set corresponding to a theme. Depending on actual application requirements, the steps in the process for automatically building the element image set may be replaced or modified in any manner, and the process may comprise more or fewer steps. In addition, the specific order or hierarchy of the steps in the process 500 is merely exemplary, and the process for automatically building the element image set may be performed in an order different from the described order.

[0086] FIG. 6 illustrates an exemplary process 600 for automatically generating an adding rule for an element according to an embodiment of the present disclosure. An adding rule set for an element may be used when adding an element image of the element in a blank or monotonous area in an image.

[0087] Multiple images relevant to an element 602 may be searched through a search engine 610. The searched images may be denoted as an image 612-1, an image 612-2, …, an image 612-k, where k is the number of searched images.

[0088] For each image in the multiple images, semantic segmentation may be performed on the image through a semantic segmentation model 620, to identify a set of properties of the element 602 in the image. The semantic segmentation model 620 is, e.g., the SSA model. The set of properties may include, e.g., position, rotation angle, scale, quality, number, area, specific category, etc. The number of properties in each set of properties may be denoted as f. Multiple sets of properties corresponding to the  multiple images may be obtained, which may be denoted as properties 622-1, properties 622-2, …, properties 622-k.

[0089] An adding rule 642 for the element 602 may be generated based on multiple sets of properties corresponding to the multiple images. For example, statistical analysis may be performed on the multiple sets of properties corresponding to the multiple images through a statistical analysis module 630, to obtain a set of average property values of the element 602. The set of average property values may be denoted as an average property value 632-1, an average property value 632-2, …, an average property value 632-f. For example, the average property value 632-1 may be an average position of the element 602 in the searched images, the average property value 632-2 may be an average rotation angle of the element 602 in the searched images, etc.

[0090] An adding rule 642 for the element 602 may be generated based on the set of average property values through a rule generation module 640. The rule generation module 640 may combine the element 602 with the set of average property values, to generate the adding rule 642. The set of average property values should be complied with when adding an element image of the element in an image.

[0091] The technical effects of this approach are to automatically generating an adding rule for an element without human intervention, while ensuring that the adding rule contains reasonable property values.

[0092] It should be appreciated that the process 600 in FIG. 6 is merely an example of the process for automatically generating an adding rule for an element. Depending on actual application requirements, the steps in the process for automatically generating an adding rule may be replaced or modified in any manner, and the process may comprise more or fewer steps. For example, if an element cannot be located through the semantic segmentation, an adding rule for the element may also be obtained through a large language model. Taking an element “indoor plant” as an example, a prompt provided to the large language model may be “How to use indoor plants to decorate a cozy and clean garden? For example, how many indoor plants are usually used? Where to put the indoor plants? Which kind of indoor plants are usually used? ”

[0093] FIG. 7A to FIG. 7C illustrate an example of image augmentation according to an embodiment of the present disclosure.

[0094] FIG. 7A illustrates an image 700a. The image 700a may correspond to the first image 102 in FIG. 1. The image 700a depicts an indoor environment, including  some surfaces and objects, such as a wall 702, a sofa 704, a floor lamp 706, a game controller 708, etc. The color of the floor lamp 706 may be white.

[0095] FIG. 7B illustrates an image 700b. The image 700b may be obtained through augmenting the image 700a based on a selected theme. The image 700b may correspond to the augmented first image 112 in FIG. 1. Assume that the selected theme is “celebration” . In the image 700b, the game controller 708 is identified as a clutter object and thus be removed. The color of the floor lamp 706 is changed to a color corresponding to the selected theme “celebration” , e.g., red. In addition, an element image 710 corresponding to the selected theme “celebration” is added on the wall 702.

[0096] FIG. 7C illustrates an image 700c. The image 700c may be generated based on the image 700b, the selected theme “celebration” and / or a strength value for the image 700b through an image generation model. The image 700c may correspond to the second image 172 in FIG. 1. In the image 700c, the sofa 704 is decorated to be consistent with the selected theme “celebration” . Colorful flags 712 corresponding to the selected theme “celebration” are generated and presented on the wall 702. In addition, a bundle of colorful balloons 714 corresponding to the selected theme “celebration” are generated and presented near the previous location of the floor lamp 706.

[0097] It can be seen that the image 700c reflects decorative effects corresponding to the selected theme “celebration” , while its structure and layout are substantially consistent with the image 700a.

[0098] FIG. 8A to FIG. 8F illustrate another example of image augmentation according to an embodiment of the present disclosure. This example may be applicable in the scenario of video communication.

[0099] FIG. 8A illustrates an image 800a, which may be an original video frame of a video communication. The image 800a may be captured by a camera of a device participating in the video communication. The image 800a include a background image and a foreground image. The background image depicts an indoor environment, including some surfaces and objects, such as a wall 802, a sofa 804, a floor lamp 806, a game controller 808, etc. The color of the floor lamp 806 may be white. The foreground image includes a user image 810.

[0100] FIG. 8B illustrates an image 800b, which may be obtained through removing the foreground image, i.e., the user image 810, from the image 800a. The foreground  image may be removed through techniques such as foreground and background segmentation, matting, etc. The image 800b has a blank region, as it lacks of the part obscured by the foreground image.

[0101] FIG. 8C illustrates an image 800c, which may be obtained through filling in the blank region in the image 800b with inpainting techniques. The inpainting techniques can fill in unknown parts of an image by utilizing the redundancy of the image itself or information from known parts of the image. The image 800c may correspond to the first image 102 in FIG. 1.

[0102] FIG. 8D illustrates an image 800d, which may be obtained through augmenting the image 800c based on a selected theme. The image 800d may correspond to the augmented first image 112 in FIG. 1. Assume that the selected theme is “celebration” . In the image 800d, the game controller 808 is identified as a clutter object and thus be removed. The color of the floor lamp 806 is changed to a color corresponding to the selected theme “celebration” , e.g., red. In addition, an element image 810 corresponding to the selected theme “celebration” is added on the wall 802.

[0103] FIG. 8E illustrates an image 800e. The image 800e may be generated based on the image 800d, the selected theme “celebration” and / or a strength value for the image 800d through an image generation model. The image 800e may correspond to the second image 172 in FIG. 1. In the image 800e, the sofa 804 is decorated to be consistent with the selected theme “celebration” . Colorful flags 812 corresponding to the selected theme “celebration” are generated and presented on the wall 802. In addition, a bundle of colorful balloons 814 corresponding to the selected theme “celebration” are generated and presented near the previous location of the floor lamp 806.

[0104] FIG. 8F illustrates an image 800f. The image 800f may be obtained through combining the image 800e and the foreground image, i.e., the user image 810.

[0105] It can be seen that the image 800e and the image 800f reflects decorative effects corresponding to the selected theme “celebration” , while their structures and layouts are substantially consistent with the image 800a or the image 800c. In addition, the generated colorful flags 812 are presented in a blank area and not obscured by the foreground image.

[0106] FIG. 9 is a flowchart of an exemplary method 900 for image augmentation according to an embodiment of the present disclosure.

[0107] At 910, a first image and a selected theme may be received.

[0108] At 920, the first image may be augmented based on the selected theme, the augmenting comprising: modifying the color of an object in the first image according to the selected theme, and / or adding element images corresponding to the selected theme in the first image.

[0109] At 930, a second image may be generated based on the augmented first image and the selected theme, the second image being an image with decorative effects corresponding to the selected theme.

[0110] In an implementation, the modifying the color of an object in the first image according to the selected theme may comprise: identifying an object from the first image, the object being of a predetermined category, with a size below a predetermined threshold, and / or being at a predetermined position; randomly selecting a color from a color set corresponding to the selected theme; and changing the color of the object to the randomly selected color.

[0111] In an implementation, the adding element images corresponding to the selected theme in the first image may comprise: detecting an area for adding element images in the first image; and adding element images corresponding to the selected theme in the area.

[0112] The detecting an area for adding element images in the first image may comprise: performing semantic segmentation on the first image, to obtain a surface mask of the first image; calculating pixel variances of image blocks in the first image, to obtain a smoothness mask of the first image; and determining the area for adding element images based on the surface mask and the smoothness mask.

[0113] The first image may be a background image in a video communication. The detecting an area for adding element images in the first image may further comprise: obtaining a foreground mask corresponding to the first image; and determining the area for adding element images based on the surface mask, the smoothness mask and the foreground mask.

[0114] The adding element images corresponding to the selected theme in the area may comprise: iteratively adding element images in the area, until the proportion of the added element images in the area exceeds a predetermined threshold, in each iteration: randomly selecting an element image from an element image set corresponding to the selected theme; and adding, in the area, the element image according to an adding rule  for an element corresponding to the element image.

[0115] The element image set corresponding to the selected theme may be automatically built through: searching multiple images relevant to the selected theme through a search engine; for each image in the multiple images, generating an image description of the image and exacting elements for decoration from the image description through a large language model, to obtain a first set of elements; obtaining a second set of elements relevant to the selected theme and for decoration through a large language model; filtering out common elements appearing in both the first set of elements and the second set of elements; ranking the common elements based on occurrence frequencies of the common elements in multiple image descriptions of the multiple element images; selecting highest-ranked elements among the common elements, to form an element list of the selected theme; and for each element in the element list, generating element images for the element through an image generation model, to obtain the element image set corresponding to the selected theme.

[0116] The adding rule for the element may be automatically generated through: searching multiple images relevant to the element through a search engine; for each image in the multiple images, performing semantic segmentation on the image to identify a set of properties of the element in the image; and generating the adding rule for the element based on multiple sets of properties corresponding to the multiple images.

[0117] The augmenting may further comprises: performing at least one of color grading, luma tuning, boundary blurring and alpha blending on the added element images.

[0118] In an implementation, the method 800 may further comprise: calculating complexity of the augmented first image based on luma information of the augmented first image; and calculating a strength value based on the complexity. The generating a second image may comprise: generating the second image based on the augmented first image, the selected theme and the strength value.

[0119] It should be appreciated that the method 900 may further comprise any other steps / processes for image augmentation according to the embodiments of the present disclosure as mentioned above.

[0120] FIG. 10 illustrates an exemplary apparatus 1000 for image augmentation according to an embodiment of the present disclosure.

[0121] The apparatus 1000 may comprise: an image and theme receiving module 1010, for receiving a first image and a selected theme; an image augmenting module 1020, for augmenting the first image based on the selected theme, the augmenting comprising: modifying the color of an object in the first image according to the selected theme, and / or adding element images corresponding to the selected theme in the first image; and an image generating module 1030, for generating a second image based on the augmented first image and the selected theme, the second image being an image with decorative effects corresponding to the selected theme. Furthermore, the apparatus 1000 may further comprise any other modules configured for image augmentation according to the embodiments of the present disclosure as mentioned above.

[0122] FIG. 11 illustrates another exemplary apparatus 1100 for image augmentation according to an embodiment of the present disclosure.

[0123] The apparatus 1100 may comprise: a processor 1110; and a memory 1120 storing computer-executable instructions. The computer-executable instructions, when executed, may cause the processor 1110 to: receive a first image and a selected theme; augment the first image based on the selected theme, the augmenting comprising: modifying the color of an object in the first image according to the selected theme, and / or adding element images corresponding to the selected theme in the first image; and generate a second image based on the augmented first image and the selected theme, the second image being an image with decorative effects corresponding to the selected theme.

[0124] In an implementation, the modifying the color of an object in the first image according to the selected theme may comprise: identifying an object from the first image, the object being of a predetermined category, with a size below a predetermined threshold, and / or being at a predetermined position; randomly selecting a color from a color set corresponding to the selected theme; and changing the color of the object to the randomly selected color.

[0125] In an implementation, the adding element images corresponding to the selected theme in the first image may comprise: detecting an area for adding element images in the first image; and adding element images corresponding to the selected theme in the area.

[0126] The detecting an area for adding element images in the first image may comprise: performing semantic segmentation on the first image, to obtain a surface  mask of the first image; calculating pixel variances of image blocks in the first image, to obtain a smoothness mask of the first image; and determining the area for adding element images based on the surface mask and the smoothness mask.

[0127] The first image may be a background image in a video communication. The detecting an area for adding element images in the first image may further comprise: obtaining a foreground mask corresponding to the first image; and determining the area for adding element images based on the surface mask, the smoothness mask and the foreground mask.

[0128] The adding element images corresponding to the selected theme in the area may comprise: iteratively adding element images in the area, until the proportion of the added element images in the area exceeds a predetermined threshold, in each iteration: randomly selecting an element image from an element image set corresponding to the selected theme; and adding, in the area, the element image according to an adding rule for an element corresponding to the element image.

[0129] The element image set corresponding to the selected theme may be automatically built through: searching multiple images relevant to the selected theme through a search engine; for each image in the multiple images, generating an image description of the image and exacting elements for decoration from the image description through a large language model, to obtain a first set of elements; obtaining a second set of elements relevant to the selected theme and for decoration through a large language model; filtering out common elements appearing in both the first set of elements and the second set of elements; ranking the common elements based on occurrence frequencies of the common elements in multiple image descriptions of the multiple element images; selecting highest-ranked elements among the common elements, to form an element list of the selected theme; and for each element in the element list, generating element images for the element through an image generation model, to obtain the element image set corresponding to the selected theme.

[0130] The adding rule for the element may be automatically generated through: searching multiple images relevant to the element through a search engine; for each image in the multiple images, performing semantic segmentation on the image to identify a set of properties of the element in the image; and generating the adding rule for the element based on multiple sets of properties corresponding to the multiple images.

[0131] In an implementation, the computer-executable instructions, when executed, may further cause the processor to: calculate complexity of the augmented first image based on luma information of the augmented first image; and calculate a strength value based on the complexity. The generating a second image may comprise: generating the second image based on the augmented first image, the selected theme and the strength value.

[0132] It should be appreciated that the processor 1110 may further perform any other steps / processes of the method for image augmentation according to the embodiments of the present disclosure as mentioned above.

[0133] The embodiments of the present disclosure propose a computer program product for image augmentation, comprising a computer program that is executed by a processor for: receiving a first image and a selected theme; augmenting the first image based on the selected theme, the augmenting comprising: modifying the color of an object in the first image according to the selected theme, and / or adding element images corresponding to the selected theme in the first image; and generating a second image based on the augmented first image and the selected theme, the second image being an image with decorative effects corresponding to the selected theme. Furthermore, the computer program may be further executed for implementing any other steps / processes of the method for image augmentation according to the embodiments of the present disclosure as mentioned above.

[0134] The embodiments of the present disclosure may be embodied in a computer-readable medium for image augmentation. The computer-readable medium may comprise instructions that, when executed, cause a processor to: receive a first image and a selected theme; augment the first image based on the selected theme, the augmenting comprising: modifying the color of an object in the first image according to the selected theme, and / or adding element images corresponding to the selected theme in the first image; and generate a second image based on the augmented first image and the selected theme, the second image being an image with decorative effects corresponding to the selected theme. Furthermore, the instructions, when executed, may further cause the processor to perform any other steps / processes of the method for image augmentation according to the embodiments of the present disclosure as mentioned above.

[0135] It should be appreciated that all the operations in the methods described  above are merely exemplary, and the present disclosure is not limited to any operations in the methods or sequence orders of these operations, and should cover all other equivalents under the same or similar concepts. In addition, the articles “a” and “an” as used in this specification and the appended claims should generally be construed to mean “one” or “one or more” unless specified otherwise or clear from the context to be directed to a singular form.

[0136] It should also be appreciated that all the modules in the apparatuses described above may be implemented in various approaches. These modules may be implemented as hardware, software, or a combination thereof. Moreover, any of these modules may be further functionally divided into sub-modules or combined together.

[0137] Processors have been described in connection with various apparatuses and methods. These processors may be implemented using electronic hardware, computer software, or any combination thereof. Whether such processors are implemented as hardware or software will depend upon the particular application and overall design constraints imposed on the system. By way of example, a processor, any portion of a processor, or any combination of processors presented in the present disclosure may be implemented with a microprocessor, microcontroller, digital signal processor (DSP) , a field-programmable gate array (FPGA) , a programmable logic device (PLD) , a state machine, gated logic, discrete hardware circuits, and other suitable processing components configured for performing the various functions described throughout the present disclosure. The functionality of a processor, any portion of a processor, or any combination of processors presented in the present disclosure may be implemented with software being executed by a microprocessor, microcontroller, DSP, or other suitable platform.

[0138] Software shall be construed broadly to mean instructions, instruction sets, code, code segments, program code, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, threads of execution, procedures, functions, etc. The software may reside on a computer-readable medium. A computer-readable medium may include, by way of example, memory such as a magnetic storage device (e.g., hard disk, floppy disk, magnetic strip) , an optical disk, a smart card, a flash memory device, random access memory (RAM) , read only memory (ROM) , programmable ROM (PROM) , erasable PROM (EPROM) , electrically erasable PROM (EEPROM) , a register, or a removable  disk. Although memory is shown separate from the processors in the various aspects presented throughout the present disclosure, the memory may be internal to the processors, e.g., cache or register.

[0139] The previous description is provided to enable any person skilled in the art to practice the various aspects described herein. Various modifications to these aspects will be readily apparent to those skilled in the art, and the generic principles defined herein may be applied to other aspects. Thus, the claims are not intended to be limited to the aspects shown herein. All structural and functional equivalents to the elements of the various aspects described throughout the present disclosure that are known or later come to be known to those of ordinary skilled in the art are expressly incorporated herein and intended to be encompassed by the claims.

Claims

1.A method for image augmentation, comprising:receiving a first image and a selected theme;augmenting the first image based on the selected theme, the augmenting comprising: modifying the color of an object in the first image according to the selected theme, and / or adding element images corresponding to the selected theme in the first image; andgenerating a second image based on the augmented first image and the selected theme, the second image being an image with decorative effects corresponding to the selected theme.2.The method of claim 1, wherein the modifying the color of an object in the first image according to the selected theme comprises:identifying an object from the first image, the object being of a predetermined category, with a size below a predetermined threshold, and / or being at a predetermined position;randomly selecting a color from a color set corresponding to the selected theme; andchanging the color of the object to the randomly selected color.3.The method of claim 1, wherein the adding element images corresponding to the selected theme in the first image comprises:detecting an area for adding element images in the first image; andadding element images corresponding to the selected theme in the area.4.The method of claim 3, wherein the detecting an area for adding element images in the first image comprises:performing semantic segmentation on the first image, to obtain a surface mask of the first image;calculating pixel variances of image blocks in the first image, to obtain a smoothness mask of the first image; anddetermining the area for adding element images based on the surface mask and the smoothness mask.5.The method of claim 4, wherein the first image is a background image in a video communication, and the detecting an area for adding element images in the first image further comprises:obtaining a foreground mask corresponding to the first image; anddetermining the area for adding element images based on the surface mask, the smoothness mask and the foreground mask.6.The method of claim 3, wherein the adding element images corresponding to the selected theme in the area comprises:iteratively adding element images in the area, until the proportion of the added element images in the area exceeds a predetermined threshold, in each iteration:randomly selecting an element image from an element image set corresponding to the selected theme; andadding, in the area, the element image according to an adding rule for an element corresponding to the element image.7.The method of claim 6, wherein the element image set corresponding to the selected theme is automatically built through:searching multiple images relevant to the selected theme through a search engine;for each image in the multiple images, generating an image description of the image and exacting elements for decoration from the image description through a large language model, to obtain a first set of elements;obtaining a second set of elements relevant to the selected theme and for decoration through a large language model;filtering out common elements appearing in both the first set of elements and the second set of elements;ranking the common elements based on occurrence frequencies of the common elements in multiple image descriptions of the multiple element images;selecting highest-ranked elements among the common elements, to form an element list of the selected theme; andfor each element in the element list, generating element images for the element through an image generation model, to obtain the element image set corresponding to the selected theme.8.The method of claim 6, wherein the adding rule for the element is automatically generated through:searching multiple images relevant to the element through a search engine;for each image in the multiple images, performing semantic segmentation on the image to identify a set of properties of the element in the image; andgenerating the adding rule for the element based on multiple sets of properties corresponding to the multiple images.9.The method of claim 3, wherein the augmenting further comprises:performing at least one of color grading, luma tuning, boundary blurring and alpha blending on the added element images.10.The method of claim 1, further comprising:calculating complexity of the augmented first image based on luma information of the augmented first image; andcalculating a strength value based on the complexity, andwherein the generating a second image comprises:generating the second image based on the augmented first image, the selected theme and the strength value.11.An apparatus for image augmentation, comprising:a processor; anda memory storing computer-executable instructions that, when executed, cause the processor to:receive a first image and a selected theme,augment the first image based on the selected theme, the augmenting comprising: modifying the color of an object in the first image according to the selected theme, and / or adding element images corresponding to the selected theme in the first image, andgenerate a second image based on the augmented first image and the selected theme, the second image being an image with decorative effects corresponding to the selected theme.12.The apparatus of claim 11, wherein the modifying the color of an object in the first image according to the selected theme comprises:identifying an object from the first image, the object being of a predetermined category, with a size below a predetermined threshold, and / or being at a predetermined position;randomly selecting a color from a color set corresponding to the selected theme; andchanging the color of the object to the randomly selected color.13.The apparatus of claim 11, wherein the adding element images corresponding to the selected theme in the first image comprises:detecting an area for adding element images in the first image; andadding element images corresponding to the selected theme in the area.14.The apparatus of claim 13, wherein the detecting an area for adding element images in the first image comprises:performing semantic segmentation on the first image, to obtain a surface mask of the first image;calculating pixel variances of image blocks in the first image, to obtain a smoothness mask of the first image; anddetermining the area for adding element images based on the surface mask and the smoothness mask.15.The apparatus of claim 14, wherein the first image is a background image in a video communication, and the detecting an area for adding element images in the first image further comprises:obtaining a foreground mask corresponding to the first image; anddetermining the area for adding element images based on the surface mask, the smoothness mask and the foreground mask.16.The apparatus of claim 13, wherein the adding element images corresponding to the selected theme in the area comprises:iteratively adding element images in the area, until the proportion of the added element images in the area exceeds a predetermined threshold, in each iteration:randomly selecting an element image from an element image set corresponding to the selected theme; andadding, in the area, the element image according to an adding rule for an element corresponding to the element image.17.The apparatus of claim 16, wherein the element image set corresponding to the selected theme is automatically built through:searching multiple images relevant to the selected theme through a search engine;for each image in the multiple images, generating an image description of the image and exacting elements for decoration from the image description through a large language model, to obtain a first set of elements;obtaining a second set of elements relevant to the selected theme and for decoration through a large language model;filtering out common elements appearing in both the first set of elements and the second set of elements;ranking the common elements based on occurrence frequencies of the common elements in multiple image descriptions of the multiple element images;selecting highest-ranked elements among the common elements, to form an element list of the selected theme; andfor each element in the element list, generating element images for the element through an image generation model, to obtain the element image set corresponding to the selected theme.18.The apparatus of claim 16, wherein the adding rule for the element is automatically generated through:searching multiple images relevant to the element through a search engine;for each image in the multiple images, performing semantic segmentation on the image to identify a set of properties of the element in the image; andgenerating the adding rule for the element based on multiple sets of properties corresponding to the multiple images.19.The apparatus of claim 11, wherein the computer-executable instructions, when executed, further cause the processor to:calculate complexity of the augmented first image based on luma information of the augmented first image; andcalculate a strength value based on the complexity, andwherein the generating a second image comprises:generating the second image based on the augmented first image, the selected theme and the strength value.20.A computer program product for image augmentation, comprising a computer program that is executed by a processor for:receiving a first image and a selected theme;augmenting the first image based on the selected theme, the augmenting comprising: modifying the color of an object in the first image according to the selected theme, and / or adding element images corresponding to the selected theme in the first image; andgenerating a second image based on the augmented first image and the selected theme, the second image being an image with decorative effects corresponding to the selected theme.

Citation Information

Patent Citations

  • Method and system for training image generation model using content information

    KR102624083B1

  • Electronic device and operating method thereof

    US11861769B2