AI facial decoration texture generation in social media platforms
By receiving basic images, image masks and user text prompts, and using artificial intelligence models to generate facial decorative textures, the problem of manual adjustment of AI generation effects in the existing technology is solved, and the automation of custom effects and skin color matching of ordinary users is achieved, which improves the user experience of social media platforms.
Patent Information
- Application Number
- CN202480005810.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2023-07-18
- Filing Date
- 2024-05-30
- Publication Date
- 2025-07-29
AI Technical Summary
In existing social media platforms, the image effects generated by AI need to be further manually adjusted, limiting the ability of ordinary users to customize effects, resulting in limited AI's usefulness in image editing.
Provides a computing system that generates facial decorative textures by receiving basic images, image masks and user text prompts using artificial intelligence models, especially using diffusion models, automatically adjusts the effects to ensure accurate positioning and adaptation to the needs of different skin tones.
It enables ordinary users to generate customized facial decorative textures instantly, ensuring that the effects are accurately positioned on the face, and automatically adjusting the skin tone matching, simplifying the effect creation process and improving user experience.
Smart Images

Figure CN120390937A_ABST
Abstract
Description
[0001] Cross - Reference to Related Applications
[0002] This application claims the benefit of priority of U.S. Provisional Application Serial No. 63 / 505,346, filed May 31, 2023, and U.S. Application Serial No. 18 / 354,546, filed Jul. 18, 2023, the entire contents of which are incorporated herein by reference for all purposes. BACKGROUND OF THE INVENTION
[0003] Many social media platforms provide tools for users to add effects to images and videos before posting them online. Some of these effects are applied to human faces, such as filters, stickers, and textures, which are designed to make it look like there are objects or materials present in the images and videos when in fact they are not, or otherwise alter or enhance real-world objects. These effects are typically provided in an effects library, and some social media platforms allow users to create new effects themselves. Creation is usually done manually, such as in image editing software, and is therefore often limited to advanced users. Artificial intelligence (AI) is becoming increasingly common as a tool for generating images without having to manually draw them from scratch. To date, attempts to use AI-generated images in social media to create new effects have required further manual adjustment to finalize the effects, which limits the usefulness of AI in this area and prevents laypersons from creating effects. SUMMARY OF THE INVENTION
[0004] A computing system is provided herein that provides a social media platform. In one example, the computing system includes one or more processors configured to execute instructions stored in an associated memory to receive a base image that includes a face, and to receive an image mask that defines regions for which inpainting occurs and regions for which inpainting does not occur. The regions for which inpainting does not occur include at least the eye regions. The one or more processors are configured to receive a user text prompt and use the base image, the image mask, and the user text prompt as inputs at an artificial intelligence (AI) model to generate a facial decoration texture.
[0005] This Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the Detailed Description. This Summary is not intended to identify key features or essential features of the claimed subject matter, nor is it intended to be used to limit the scope of the claimed subject matter. Additionally, the claimed subject matter is not limited to implementations that solve any or all disadvantages noted in any part of this disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0006] Figure 1A schematic diagram of a computing system including a server device providing a social media platform is shown.
[0007] Figure 2 An example artificial intelligence (AI) model used by the computing system of Figure 1 to generate a facial decoration texture is shown.
[0008] Figures 3A - 3B An example base image used by the computing system of Figure 1 is shown.
[0009] Figures 4A - 4D An example image mask used by the computing system of Figure 1 is shown.
[0010] Figure 5 An example graphical user interface (GUI) of the social media platform of Figure 1 displaying a prompt input screen is shown.
[0011] Figure 6 Another example prompt input screen of the GUI of Figure 5 is shown.
[0012] Figure 7 An example output selection screen of the GUI of Figure 5 is shown.
[0013] Figure 8A An example video editing screen of the GUI of Figure 5 is shown.
[0014] Figures 8B - 8C An example video edited in the video editing screen of Figure 8A is shown, where an "old woman" facial decoration texture is applied to a female's face.
[0015] Figures 9A - 9B An example video edited in the video editing screen of Figure 8A is shown, where the "old woman" facial decoration texture is applied to a female's face using blending for skin tone matching.
[0016] Figure 10 An example blending menu of the GUI of Figure 5 is shown.
[0017] Figures 11A - 11C Another example video edited in the video editing screen of Figure 8A is shown, where a makeup facial decoration texture is applied to a female's face.
[0018] Figure 12 An example effect menu listing trend effects in the GUI of Figure 5 is shown.
[0019] Figure 13 A flowchart of a method for a social media platform is shown.
[0020] Figure 14 Shows what can be implemented Figure 1 A schematic diagram of an example computing environment of a computing system. DETAILED DESCRIPTION
[0021] To solve the above problems, Figure 1 FIG. illustrates a computing system 100 that includes a server device 10 that provides a social media platform 12. The server device 10 includes one or more processors 14 that are configured to execute instructions stored in an associated memory 16 to implement various functions of the server device 10. The instructions can include, for example, a facial decoration texture generation module 18 and an application server program 20. It is to be understood that the server device 10 can include multiple different servers that work together to provide the social media platform 12, or can be a single server. The server device 10 can also include an effects data repository 22 for storing a facial decoration texture library for use by users of the social media platform 12 and a video data repository 24 for storing published video content for users to view.
[0022] One or more processors 14 can be configured to send instructions to a client device 26 to cause the client device 26 to display a graphical user interface (GUI) 28 of the social media platform 12. The server device 10 and the client device 24 can communicate with each other via a network 30 and one or more handlers 32 of the application server program 22. The client device 26 can be a smart phone, a tablet computer, a personal computer, etc., and includes one or more processors 34 configured to execute a client program 36 to display the GUI 28 on a display 30, a memory 32 for storing instructions, and one or more input devices 34 for receiving user input. The input device 34 can include, for example, a touch screen, a keyboard, a microphone, a camera, an accelerometer, etc. It is to be understood that the client program 36 can be a dedicated application for accessing the social media platform 12, or can alternatively be a general program (such as an Internet browser) for accessing content from various server devices including the social media platform 12 from the server device 10. It is to be further understood that in some implementations, the facial decoration texture generation module 18 can be executed locally by one or more processors 34 of the client device 26.
[0023] In short, the client device 24 can send a generation request 38 to the server device 10 seeking to generate a new effect. One or more processors 14 can be configured to receive the generation request 38, including a base image selection 40 and a user text prompt 42. Then, an artificial intelligence (AI) model 44 can receive as input a base image 46 including a face, an image mask 48 defining regions for which repair is to occur and regions for which repair is not to occur (see the example discussed below in Figures 4A - 4D ), and the user text prompt 42. It is to be understood that the AI model 44 can be configured to receive the base image 46 from a user of the client device 26, but more simply, receiving the base image 46 can include receiving a selection 40 of one of a plurality of base images 46 (see Figures 3A - 3B for two examples). Alternatively, the AI model 44 can be configured to retrieve a single stored base image 46 from the memory 16. Providing multiple base images 46 to the facial adornment texture generation module 18 will allow the user to test their creation on people with very different looks to ensure that the created effect is suitable for a wide range of users. In some implementations, the facial adornment texture module 18 can include multiple models, including other AI models 44A, but in other implementations, only one AI model 44 can be used.
[0024] The facial adornment texture generation module 18 can be configured to use the base image 46, the image mask 48, and the user text prompt 42 as inputs at the AI model 44 to generate a facial adornment texture 50. As will be discussed in more detail below, the regions for which repair is not to occur include at least the eye regions, which helps the AI model 44 center the generated facial adornment texture 50 at the correct location on the face in the base image and helps the facial adornment texture generation module 18 center the generated facial adornment texture 50 at the correct location on the face in a captured image or video using at least the eyes as an anchor. Finally, the server device 10 can be configured to store the facial adornment texture 50 in the effect data repository 22 and / or send the facial adornment texture 50 to the client device 26.
[0025] Figure 2 Shown by Figure 1An example of a computing system 100 that uses an AI model 44 to generate a facial decoration texture 50 is shown. The illustrated AI model 44 is merely an example, and any suitable generative AI model may be used. In the depicted example, the AI model 44 is a trained machine learning model, more specifically a diffusion model. Examples of known diffusion models may include Stable Diffusion, Real Vision, etc., and suitable diffusion models may include modified versions of these known models. In this example, an image encoder 52 is provided that has a pre-trained layer 52A (such as a pre-trained Contrastive Language-Image Pretraining (CLIP) Vision Transformer (ViT)), a fine-tuning layer 52B trained to extract visual features from input images (such as a base image 46 and an image mask 48), and a fully connected layer 52C configured to generate an embedding set 54 based at least on the visual features of the face in the base image 46 extracted by the fine-tuning layer 52B. The image mask 48 may be processed as an alpha channel, indicating which parts of the output should be opaque (included) and which parts should be transparent (excluded). For example, the base image 46 may first be masked based on the occluded and allowed regions of the image mask 48, and the resulting masked base image may then be used as the input for generation, as described below. In some implementations, the embedding set 54 may be associated with a user identifier 56 of a user of the client device 26.
[0026] The AI model 44 may be configured to receive a user text prompt 42 that describes what effect the user wishes the AI model 44 to create. The user text prompt 42 and the embedding set 54 may be provided as inputs to a text encoder 58, and the text encoder 58 may generate an input feature vector 60 based at least on the user text prompt 42 and the embedding set 54. The input feature vector 60 may be sent to a diffusion module 62 of the AI model 44, which is configured to generate a synthetic image as the facial decoration texture 50 based at least on the input feature vector 50.
[0027] Figures 4A - 4D An example of an image mask 48 used by Figure 1 the computing system 100 is shown. In Figure 4AIn this case, the facial decoration texture 50 to be generated is a mask, and the first image mask 48A includes a region for repair (white) and a region for no repair (black). Here, in addition to the eye regions 64 (specifically, two eye regions 64), the region for no repair also includes the mouth region 66. The general shape is like a mask for a Halloween costume, so the region for repair is limited to the mask shape and does not include the region 68 around the face. Using the first image mask 48A, a complete mask can be output as the facial decoration texture 50, with the mouth and eyes cut out so that the eyes and mouth of the person in the image are visible. In some implementations, the facial decoration texture 50 can be makeup, just as in the case when using Figures 4B - 4D the example image masks 48B to 48D shown. Figure 4B The second image mask 48B in includes the mouth region 66, which is included in the region for repair (white), while the region for no repair (black) also includes the region 70 around the mouth region 66. It is worth noting that this is the opposite of the first image mask 48, where the mouth region 66 is included in the region for no repair (black). This is because a mask usually shows the wearer's mouth, while makeup usually includes lipstick or lip gloss on the wearer's mouth. Therefore, the second image mask 48 can be lip-shaped and excludes the region 70 around the lips from the region for repair.
[0028] Figures 4C - 4D The regions 72 around the eye regions 64 are all included in the region for repair. For example, the third image mask 48C in FIG. 4 has a region for repair (white) that unevenly surrounds the eye region 64 to provide a palette for eyeshadow mainly above the wearer's eyes and additional under-eye makeup that is thinner around the lower side of the eye region 64. It can be understood that different shapes of the regions 72 around the eye regions 64 can be used, such as a larger region around the two eye regions 66. Meanwhile, Figure 4DThe fourth image mask 48D in [description] only provides the lash area 74 that radiates outward from the eye area 64 and surrounds the eye area 64 as the area for repair (white), while the rest of the image mask is black. The image mask 48 used by the AI model 44 can be appropriately selected based on the expected output of the facial decoration texture generation module 18 (i.e., whether the user requests to generate a mask or makeup). The social media platform 12 can provide these two features through separate channels, can determine which feature is requested based on the context of the user text prompt, or can provide only one of the features without the other. Additionally, facial decoration textures other than makeup and masks can be generated. For makeup, any suitable combination of the image masks 48B to 48D can be used as the image mask 48 to be input into the AI model 44. For example, a user who requests "dazzling pink lip gloss" can receive, in some instances, a generated facial decoration texture 50 that only covers the mouth area 66, and a user who requests "full glamour blue makeup" can receive a facial decoration texture 50 that covers the mouth area 66 and the area 72 surrounding the eye area 64.
[0029] Figure 5 A to Figure 12 illustrates various example screens and videos (or still images) related to the generation of the facial decoration texture 50 displayed by the GUI 28. In Figures 5 - 6 it, the GUI 28 is displaying a prompt input screen 76. Figure 5 The example shown can be customized for the desktop version of the client program 36, while Figure 6 the example shown can be customized for the mobile version. The desktop version can be targeted at more skilled users who can be given more options and control over effect creation, while the mobile version can be streamlined to produce effects for less experienced users. In Figure 5 it, the GUI 28 can display a prompt input box 78 configured to receive the user text prompt 42. Instructions 80 can explain how to use the facial decoration effect generation feature. The base image 46 selected by the user can be displayed for reference. The generation selector 82 can be selectable to send the input to the AI model 44 to start generation. In contrast, in Figure 6 it, one or more processors 14, 34 can be configured to present the GUI 28 to the user of the client device 26, in which case the client device 26 can be a mobile computing device. Here, the GUI 28 can ultimately be configured to display the facial decoration texture 50 on a person's face (e.g., see Figures 8B - 8C), but does not display the base image 46. In both versions, the GUI 28 may not display the image mask 48 to the user. By reducing the display of additional inputs that the user is not familiar with, the process can be streamlined and user confusion regarding the GUI 28 can be reduced. Additionally, the mobile version may include one or more suggestion prompts 84, which may be accompanied by an image or video of the corresponding facial decoration texture. Selection of one of the suggestion prompts 84 given by the user may cause the suggestion prompt 84 to be added to the prompt input box 78 for the user, and the user may freely modify or add the suggestion prompt 84 before finalizing the user text prompt 42.
[0030] Figure 7 shows Figure 5 An example output selection screen 86 of the GUI 28. Here, the output of the AI model 44 includes a plurality of facial decoration textures 50, four in the example shown here. The GUI 28 may include corresponding check selectors 88 for selecting any of the facial decoration textures 50 that the user wishes to retain. For example, each check selector 88 may be hidden until the user's cursor hovers over the corresponding facial decoration texture 50. The import selector 90 may be operable to download any selected facial decoration texture 50 to the client device 26, or all facial decoration textures 50 if none are selected. To the right of the facial decoration texture 50 is an option pane 92. A base image menu 92A may be included to allow the user to select which base image to display below the facial decoration texture 50. This may allow the user to test whether the generated facial decoration texture 50 is suitable for various facial types, particularly various skin tones. The option pane 92 may also include customizable options for the AI model 44, such as the generation steps 92B (which is the number of diffusion steps taken by the AI model 44) and the prompt strength 92C (which is the strength by which the AI model weights the user text prompt 42 during generation). Once the user is satisfied with the facial decoration texture 50 and downloads one or more to the client device 26, the user can use the facial decoration texture 50 in video and image editing.
[0031] Figure 8A shows Figure 5 An example video editing screen 94 of the GUI 28. A number of icons 96 arranged around the video editing screen 94 may be operable to perform various editing tasks to produce a final video. For simplicity, the icons 96 are only shown in Figure 8Ais shown. In this example, a female is shown in the video. Once the facial decoration texture 50 is created, one or more processors 14, 34 can also be configured to automatically apply the facial decoration texture 50 to the video captured by the camera of the client device 26, or it can be selectable from a menu, such as an effect menu screen 98 that can be opened via an effect selector 102 (see Figure 12 ). It is to be understood that the video can be a “viewfinder” live preview of a scene that the camera can capture before recording, a live shot that the camera is currently recording, or a previously recorded and stored shot. For example, Figures 8B - 8C shows a facial decoration texture 50 generated by applying a user text prompt “old woman” on the face 104 from a live video information stream. Generally, such a three-dimensional effect is applied as a texture on a mesh, where the mesh tracks the face as it moves in each frame. The face 104 in the live video information stream is detected using a face detection algorithm (e.g., finding the eyes and mouth), and a three-dimensional facial model (mesh) is generated from the detected face 104 using a three-dimensional reconstruction algorithm. The facial decoration texture 50 is applied to the three-dimensional facial model. Based on the changes in the position and orientation of the face 104 detected in each frame of the live video information stream, the position and orientation of the three-dimensional model and the facial decoration texture 50 applied to it are updated.
[0032] In some implementations, the facial decoration texture generation module 18 may even be able to adjust the mesh to create three-dimensional features, e.g., a tiger's muzzle protruding from the wearer's face, rather than a human nose with tiger stripes. This can be achieved, for example, by depth estimation performed by an algorithm and corresponding adjustment of the image mask 48. In this way, whether the mesh is original or modified, the user can try the facial decoration texture 50 in real time with a series of poses, gestures, and facial expressions, and start shooting with this effect, even though only the facial decoration texture 50 has been created before. Once the video is completed, the user can post the video content 106 on the social media platform 12 for other users to view on other client devices 108. Other users can view the video content 106 as well as other video content 110 stored in the video data repository 24 of the server device 10.
[0033] However, Figures 8B - 8CThe shown facial decoration texture 50 has a skin color significantly different from that of the applied female. If users feel that the facial decoration texture 50 does not take them into account, this may prevent users from utilizing the facial texture generation function. Therefore, one or more processors 14, 34 may also be configured to determine the skin color tone of the human face 104 at pixels on a pixel-by-pixel basis, and then compare the tones of the corresponding pixels of the facial decoration texture 50 to be superimposed on the human face 104. If the difference between the tone of the facial decoration texture 50 and the skin color tone is less than or equal to a threshold, then one or more processors 14, 34 may return the pixels of the facial decoration texture 50 as they are. However, if the difference is greater than the threshold, then one or more processors 14, 34 may multiply the tone of the facial decoration texture 50 and the skin color tone, and return the resulting value as the pixels of the facial decoration texture 50. Figures 9A - 9B Shown is the same female with a slightly different "old woman" facial decoration texture 50, which has been appropriately blended to match her skin color. Thus, even if the base image is a person with a skin color significantly different from that of the person to whom the generated facial decoration texture 50 is applied, the skin color can be blended, and a tone-matching effect can be provided for the user. Additionally, the blending process is based on the skin color tone, rather than another value such as brightness, so warmer tones may be blended more, while cooler tones may be blended less. Thus, for example, blue eyeshadow will be blended just right, without unnatural ink appearing on the human face, and at the same time, it will not seep into the skin of the human face. Additionally, as Figure 10 shown, one or more processors 14, 34 may also be configured to present multiple blending modes 112 to the user of the client device 26 for blending the facial decoration texture 50 with the human face 104. The blending modes 112 are shown in an example blending menu 114 here, and it is to be understood that they are merely examples and may include other suitable blending modes.
[0034] Figures 11A - 11C Shown in Figure 8A is another example video edited in the video editing screen 94 of Figure 11A where a makeup facial decoration texture is applied to the face of a female. Figure 8A Shown is a video of another human face 104 containing a female different from that in Figures 11B - 11C where no effect is applied. In Figures 11B - 11CAs can be seen, even if the woman in the video turns her head, the makeup effect is accurately placed around the eyes and mouth of the base image 46 because the AI model 44 is guided by using the image mask 48. Therefore, the generated facial decoration texture 50 can be accurately placed without manual intervention during or after generation.
[0035] In some instances, one or more processors 14, 34 may also be configured to store the facial decoration texture 50 and make the facial decoration texture 50 available to other users of the social media platform 12 via other client devices 108. The facial decoration texture 50 may be made available through different channels, such as the effect library of the effect data repository 22, for example, the effect library can be accessed via Figure 8A the effect menu screen 98 opened by the effect selector 102 as shown. In some cases, one or more processors 14, 34 may also be configured to present a plurality of trending facial decoration textures 50 including the facial decoration texture 50 to other users. Figure 12 is shown Figure 5 an example of an effect menu screen 98 listing trending effects in the GUI 28 of. As determined according to the algorithm, the trending facial decoration textures 50 may be those facial decoration textures 50 that other users of the social media platform 12 participate in (view, like, apply, post, edit, etc.). By allowing users to share and use user-created effects, the user community of the social media platform 12 can have an enhanced user experience, with increased options, including more options for users of various ethnicities and facial types, and thus user engagement can also be improved.
[0036] shows a flowchart of a method 1300 for a social media platform according to the present disclosure. The method 1300 may be performed by Implemented by the illustrated computing system 100. At 1302, method 1300 may include receiving a base image including a face. At 1304, method 1300 may include receiving an image mask that defines regions for which inpainting occurs and regions for which inpainting does not occur, and the region for which inpainting does not occur includes at least the eye region. At 1306, method 1300 may include receiving a user text prompt. At 1308, method 1300 may include using the base image, the image mask, and the user text prompt as inputs at an artificial intelligence (AI) model to generate a facial decoration texture. In this way, even users without graphic design skills can instantly generate custom facial decoration textures, and due to the use of an image mask that details regions for which inpainting occurs or does not occur, the generated texture will be easily and precisely applied. At 1310, method 1300 may include applying the facial decoration texture to a face in a real-time video information stream. Thus, once created, the user or other users can immediately use the facial decoration texture in a real-time video information stream with continuously updated frames, and the facial decoration texture can remain on the face with high accuracy.
[0037] In some implementations, the AI model may be a diffusion model. The diffusion model may be suitable for generating the desired effect. At 1312, method 1300 may include performing skin tone blending by performing the following sub-steps on a pixel-by-pixel basis: at 1314, determining the skin tone hue of the face at the pixel; at 1316, comparing the hue of the corresponding pixel of the facial decoration texture to be superimposed on the face; at 1318, if the difference between the hue of the facial decoration texture and the skin tone hue is less than or equal to a threshold, returning the pixel of the facial decoration texture as it is; and at 1320, if the difference is greater than the threshold, multiplying the hue of the facial decoration texture and the skin tone hue and returning the resulting value as the pixel of the facial decoration texture. In this way, even if the base image has pale skin and the face on which the facial decoration texture is to be applied has substantially darker skin, the hue of the facial decoration texture can be adjusted so that the skin tone blends naturally into the captured face image.
[0038] In some implementations, the facial decoration texture can be a mask, and the area where no repair occurs can also include the mouth area. Therefore, the generated facial decoration texture can accurately track the eyes and mouth of the human face, and correctly show these features through the facial decoration texture. In other implementations, the facial decoration texture can be makeup, the area where repair occurs can include the area around the mouth area and the eye area, and the area where no repair occurs can also include the area around the mouth area. Contrary to the mask implementation that should not cover the human mouth, makeup usually includes lip components, so it should cover the human mouth, but not blend outside the mouth. Therefore, the generated facial decoration texture can accurately track the eyes and mouth of the human face, and correctly show these features through the facial decoration texture, even when the mouth area is included rather than excluded in the mask implementation.
[0039] In some implementations, receiving the base image can include receiving a selection of one of a plurality of base images. For example, the desktop version of the client program can provide more customization features and options for the user, and can allow the user to select from various base images when requesting the generation of the facial decoration texture. Alternatively, at 1322, method 1300 can include presenting a graphical user interface (GUI) to the user of the mobile computing device, the GUI being configured to display the facial decoration texture on the human face without displaying the base image. Therefore, in a more streamlined method, the AI model can automatically receive the base image and the image mask from the storage device as inputs, without the user providing further input to select or provide these images. At 1324, method 1300 can include storing the facial decoration texture and making the facial decoration texture available to other users of the social media platform. In this way, users may be able to share their creations with other users, thus enhancing the user experience on the social media platform.
[0040] In some embodiments, the methods and processes described herein can be bound to the computing system of one or more computing devices. In particular, such methods and processes can be implemented as computer applications or services, application programming interfaces (APIs), libraries, and / or other computer program products.
[0041] A non-limiting embodiment of a computing system 1400 that can formulate one or more of the above methods and processes is schematically shown. The computing system 1400 is shown in a simplified form. The computing system 1400 can implement the above-described and in The computing device 10 illustrated in the figure. The computing system 1400 may take the form of one or more personal computers, server computers, tablet computers, home entertainment computers, network computing devices, gaming devices, mobile computing devices, mobile communication devices (such as smart phones) and / or other computing devices, as well as wearable computing devices (such as smart watches and head-mounted augmented reality devices).
[0042] The computing system 1400 includes a logical processor 1402, volatile memory 1404, and a non-volatile storage device 1406. The computing system 1400 may optionally include a display subsystem 1408, an input subsystem 1410, a communication subsystem 1412, and / or other components not shown.
[0043] The logical processor 1402 includes one or more physical devices configured to execute instructions. For example, the logical processor may be configured to execute instructions that are part of one or more applications, programs, routines, libraries, objects, components, data structures, or other logical constructs. Such instructions may be implemented to perform tasks, implement data types, transform the state of one or more components, achieve a technical effect, or otherwise achieve a desired result.
[0044] The logical processor may include one or more hardware processors (hardware) configured to execute software instructions. Additionally or alternatively, the logical processor may include one or more hardware logic circuits or firmware devices configured to execute hardware-implemented logic or firmware instructions. The processors of the logical processor 1402 may be single-core or multi-core, and the instructions executed thereon may be configured for sequential, parallel, and / or distributed processing. Optionally, the various components of the logical processor may be distributed among two or more separate devices, which may be remotely located and / or configured for coordinated processing. Various aspects of the logical processor may be virtualized and executed by remotely accessible networked computing devices configured in a cloud computing configuration. In such a case, it is to be understood that these virtualized aspects run on different physical logical processors of various different machines.
[0045] The non-volatile storage device 1406 includes one or more physical devices configured to store instructions executable by the logical processor to implement the methods and processes described herein. When such methods and processes are implemented, the state of the non-volatile storage device 1406 may be transformed, for example, to store different data.
[0046] The non-volatile storage device 1406 may include removable and / or built-in physical devices. The non-volatile storage device 1406 may include optical memories (e.g., CDs, DVDs, HD-DVDs, Blu-ray discs, etc.), semiconductor memories (e.g., ROMs, EPROMs, EEPROMs, flash memories, etc.), and / or magnetic memories (e.g., hard disk drives, floppy disk drives, tape drives, MRAMs, etc.) or other mass storage device technologies. The non-volatile storage device 1406 may include non-volatile, dynamic, static, read / write, read-only, sequential access, location-addressable, file-addressable, and / or content-addressable devices. It is to be understood that the non-volatile storage device 1406 is configured to store instructions even when the non-volatile storage device 1406 is powered off.
[0047] The volatile memory 1404 may include a physical device that includes a random access memory. The volatile memory 1404 is typically used by the logic processor 1402 to temporarily store information during the processing of software instructions. It is to be understood that when the volatile memory 1404 is powered off, the volatile memory 1404 generally does not continue to store instructions.
[0048] Aspects of the logic processor 1402, the volatile memory 1404, and the non-volatile storage device 1406 may be integrated together into one or more hardware logic components. For example, such hardware logic components may include field programmable gate arrays (FPGAs), program and application specific integrated circuits (PASIC / ASICs), program and application specific standard products (PSSP / ASSPs), systems on a chip (SOCs), and complex programmable logic devices (CPLDs).
[0049] The terms "module", "program", and "engine" may be used to describe an aspect of the computing system 1400, typically implemented in software by a processor, to perform a specific function using a portion of the volatile memory, which involves a transformational process that specifically configures the processor to perform the function. Thus, a module, program, or engine may be instantiated via the logic processor 1402 using a portion of the volatile memory 1404, which executes instructions saved by the non-volatile storage device 1406. It is to be understood that different modules, programs, and / or engines may be instantiated by the same application, service, code block, object, library, routine, API, function, etc. Similarly, the same module, program, and / or engine may be instantiated by different applications, services, code blocks, objects, routines, APIs, functions, etc. The terms "module", "program", and "engine" may encompass a single or a group of executable files, data files, libraries, drivers, scripts, database records, etc.
[0050] When included, the display subsystem 1408 can be used to present a visual representation of data stored by the non-volatile storage device 1406. The visual representation can take the form of a graphical user interface (GUI). Since the methods and processes described herein change the data stored by the non-volatile storage device and thus transform the state of the non-volatile storage device, correspondingly, the state of the display subsystem 1408 can be transformed to visually represent the changes in the underlying data. The display subsystem 1408 can include one or more display devices that actually utilize any type of technology. Such display devices can be combined with the logic processor 1402, the volatile memory 1404, and / or the non-volatile storage device 1406 in a shared housing, or such display devices can be peripheral display devices.
[0051] When included, the input subsystem 1410 can include or interface with one or more user input devices, such as a keyboard, mouse, touch screen, game controller, microphone, camera, accelerometer, gyroscope, and / or any other suitable sensors. When included, the communication subsystem 1412 can be configured to communicatively couple the various computing devices described herein to each other and to other devices. The communication subsystem 1412 can include wired and / or wireless communication devices that are compatible with one or more different communication protocols. As a non-limiting example, the communication subsystem can be configured to communicate via a wireless telephone network or a wired or wireless local area network or wide area network (such as HDMI over a wireless network connection). In some embodiments, the communication subsystem can allow the computing system 1400 to send messages to and / or receive messages from other devices via a network such as the Internet.
[0052] The following paragraphs provide additional description of the subject matter of the present disclosure. One aspect provides a computing system that provides a social media platform. The computing system includes one or more processors configured to execute instructions stored in an associated memory to: receive a base image including a face; receive an image mask that defines a region for inpainting and a region for non-inpainting, and the region for non-inpainting includes at least an eye region; receive a user text prompt; and at an artificial intelligence (AI) model, use the base image, the image mask, and the user text prompt as inputs to generate a facial decoration texture. In this aspect, additionally or alternatively, the AI model is a diffusion model. In this aspect, additionally or alternatively, the facial decoration texture is a mask, and the region for non-inpainting further includes a mouth region. In this aspect, additionally or alternatively, the facial decoration texture is makeup, the region for inpainting includes regions around the mouth region and the eye region, and the region for non-inpainting further includes regions around the mouth region. In this aspect, additionally or alternatively, receiving the base image includes receiving a selection of one of a plurality of base images. In this aspect, additionally or alternatively, the one or more processors are further configured to apply the facial decoration texture to a face in a real-time video information stream. In this aspect, additionally or alternatively, the one or more processors are further configured to: on a pixel-by-pixel basis, determine the skin color tone of the face at a pixel; compare the tone of the corresponding pixel of the facial decoration texture to be superimposed on the face; if the difference between the tone of the facial decoration texture and the skin color tone is less than or equal to a threshold, return the pixel of the facial decoration texture as it is; and if the difference is greater than the threshold, multiply the tone of the facial decoration texture and the skin color tone, and return the resulting value as the pixel of the facial decoration texture. In this aspect, additionally or alternatively, the one or more processors are further configured to present a plurality of blending modes for blending the facial decoration texture with the face to a user of a client device. In this aspect, additionally or alternatively, the one or more processors are further configured to present a graphical user interface (GUI) to a user of a mobile computing device, the GUI being configured to display the facial decoration texture on the face and not display the base image. In this aspect, additionally or alternatively, the one or more processors are further configured to store the facial decoration texture and make the facial decoration texture available to other users of the social media platform.
[0053] On the other hand, a method for a social media platform is provided. The method includes: receiving a base image including a human face; receiving an image mask that defines a region for repair to occur and a region for no repair to occur, and the region for no repair to occur includes at least the eye region; receiving a user text prompt; and at an artificial intelligence (AI) model, using the base image, the image mask, and the user text prompt as inputs to generate a facial decoration texture. In this aspect, additionally or alternatively, the AI model is a diffusion model. In this aspect, additionally or alternatively, the facial decoration texture is a mask, and the region for no repair to occur also includes the mouth region. In this aspect, additionally or alternatively, the facial decoration texture is makeup, the region for repair to occur includes the mouth region and the region around the eye region, and the region for no repair to occur also includes the region around the mouth region. In this aspect, additionally or alternatively, receiving the base image includes receiving a selection of one of a plurality of base images. In this aspect, additionally or alternatively, the method further includes applying the facial decoration texture to a human face in a real-time video information stream. In this aspect, additionally or alternatively, the method further includes, on a pixel-by-pixel basis: determining the skin color tone of the human face at a pixel; comparing the tone of the corresponding pixel of the facial decoration texture to be superimposed on the human face; if the difference between the tone of the facial decoration texture and the skin color tone is less than or equal to a threshold, returning the pixel of the facial decoration texture as it is; and if the difference is greater than the threshold, multiplying the tone of the facial decoration texture and the skin color tone, and returning the resulting value as the pixel of the facial decoration texture. In this aspect, additionally or alternatively, the method further includes storing the facial decoration texture and making the facial decoration texture available to other users of the social media platform. In this aspect, additionally or alternatively, a non-transitory computer-readable storage medium stores a computer program that, when executed by a processor, causes the processor to execute the method.
[0054] On the other hand, a server device for providing a social media platform is provided. The server device includes one or more processors configured to execute instructions stored in an associated memory to: receive a selection of a base image including a human face; receive a user text prompt; and at an artificial intelligence (AI) model, use the base image, an image mask, and the user text prompt as inputs to generate a facial decoration texture, the image mask defining a region for repair to occur and a region for no repair to occur, and the region for no repair to occur includes at least the eye region.
[0055] It is to be understood that the configurations and / or methods described herein are exemplary in nature, and these specific embodiments or examples should not be considered restrictive as many variations are possible. The specific routines or methods described herein may represent one or more strategies of any number of processing strategies. Accordingly, the various acts illustrated and / or described may be performed in the sequence illustrated and / or described, in other sequences, in parallel, or omitted. Similarly, the order of the above processes may be changed.
[0056] The subject matter of the present disclosure includes all novel and non-obvious combinations and subcombinations of the various processes, systems, and configurations and other features, functions, acts, and / or properties disclosed herein, as well as any and all equivalents thereof.
Claims
1. A computing system for providing a social media platform, the computing system comprising: One or more processors configured to execute instructions stored in an associated memory to: Receive a base image including a face; Receive an image mask that defines regions for inpainting and regions for not inpainting, the regions for not inpainting including at least the eye regions; Receive a user text prompt; And At an artificial intelligence (AI) model, use the base image, the image mask, and the user text prompt as inputs to generate a facial decoration texture.
2. The computing system according to claim 1, wherein the AI model is a diffusion model.
3. The computing system according to claim 1, wherein the facial decoration texture is a mask, and the regions for not inpainting further include the mouth region.
4. The computing system according to claim 1, wherein The facial decoration texture is makeup, The regions for inpainting include the mouth region and the regions around the eye regions, and The regions for not inpainting further include the regions around the mouth region.
5. The computing system according to claim 1, wherein receiving the base image includes receiving a selection of one base image from a plurality of base images.
6. The computing system according to claim 1, wherein the one or more processors are further configured to apply the facial decoration texture to a face in a real-time video stream.
7. The computing system according to claim 6, wherein the one or more processors are further configured to, on a per-pixel basis: Determine the skin color tone of the face at the pixel; Compare the tone of the corresponding pixel of the facial decoration texture to be superimposed on the face; If the difference between the tone of the facial decoration texture and the skin color tone is less than or equal to a threshold, return the pixel of the facial decoration texture as it is; and If the difference is greater than the threshold, multiply the tone of the facial decoration texture and the skin color tone, and return the resulting value as the pixel of the facial decoration texture.
8. The computing system according to claim 6, wherein the one or more processors are further configured to present a plurality of blending modes for blending the facial decoration texture with the face to a user of a client device.
9. The computing system according to claim 6, wherein the one or more processors are further configured to present a graphical user interface (GUI) to a user of a mobile computing device, the GUI being configured to display the facial decoration texture on the face and not display the base image.
10. The computing system according to claim 1, wherein the one or more processors are further configured to store the facial decoration texture and make the facial decoration texture available for other users of the social media platform.
11. A method for a social media platform, the method comprising: Receive a base image including a face; Receive an image mask that defines regions for which inpainting occurs and regions for which inpainting does not occur, where the regions for which inpainting does not occur include at least the eye regions; Receive a user text prompt; And At an artificial intelligence (AI) model, use the base image, the image mask, and the user text prompt as inputs to generate a facial decoration texture.
12. The method according to claim 11, wherein the AI model is a diffusion model.
13. The method according to claim 11, wherein the facial decoration texture is a mask, and the regions for which inpainting does not occur further include the mouth region.
14. The method according to claim 11, wherein The facial decoration texture is makeup, The regions for which inpainting occurs include the mouth region and the regions around the eye regions, and The regions for which inpainting does not occur further include the regions around the mouth region.
15. The method according to claim 11, wherein receiving the base image includes receiving a selection of one base image from a plurality of base images.
16. The method according to claim 11, further comprising applying the facial decoration texture to a face in a real-time video information stream.
17. The method according to claim 16, further comprising, on a pixel-by-pixel basis: Determine the skin color tone of the face at the pixel; Compare the tone of the corresponding pixel of the facial decoration texture to be superimposed on the face; If the difference between the tone of the facial decoration texture and the skin color tone is less than or equal to a threshold, return the pixel of the facial decoration texture as it is; and If the difference is greater than the threshold, multiply the tone of the facial decoration texture and the skin color tone, and return the resulting value as the pixel of the facial decoration texture.
18. The method according to claim 11, further comprising storing the facial decoration texture and making the facial decoration texture available for other users of the social media platform.
19. A non-transitory computer-readable storage medium storing a computer program that, when executed by a processor, causes the processor to perform the method according to claim 11.
20. A server device providing a social media platform, the server device comprising: One or more processors configured to execute instructions stored in an associated memory to: Receive a selection of a base image including a face; Receive a user text prompt; and At an artificial intelligence (AI) model, use the base image, an image mask, and the user text prompt as inputs to generate a facial decoration texture, the image mask defining regions for which inpainting occurs and regions for which inpainting does not occur, where the regions for which inpainting does not occur include at least the eye regions.