Whole-body styling of person

By using machine learning models to estimate and apply a stylized version of the real-world human body, the problem that existing AR systems are difficult to accurately identify and modify the human body without using depth sensors is solved, achieving an efficient and realistic AR experience.

CN120077409APending Publication Date: 2025-05-30SNAP INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380073701.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-10-17
Filing Date
2023-10-16
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing AR systems are difficult to accurately identify and modify the real-world human body depicted in images without using depth sensors, especially when the user is far away from the camera device, resulting in image quality degradation and partial facial or body error removal.

Method used

Replace the depiction of real-world humans by using machine learning models in real-time estimates or predicts stylized versions of real-world humans depicted in images and applies them to images.

Benefits of technology

It realizes a more effective and realistic AR experience without the need for depth sensors, and can apply AR graphics to different parts of the human body in real time when users move and posture changes, reducing processing complexity and resource requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120077409A_ABST
    Figure CN120077409A_ABST
Patent Text Reader

Abstract

Methods and systems for performing real-time stylized operations are disclosed. A system receives an image including a depiction of a whole body of a real-world person. The system applies a machine learning model to the image to generate a stylized version of the whole body of the real-world person corresponding to a given style, a machine learning model is trained using the training data to establish a relationship between a plurality of training images depicting the synthetically rendered whole body of the person and a corresponding truth valued stylized version of the whole body of the person of a given style. The system replaces the depiction of the whole body of the real-world person in the image with the generated stylized version of the whole body of the real-world person.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority Claim

[0002] This application claims the benefit of priority of U.S. Patent Application No. 17 / 967,230, filed Oct. 17, 2022, the entire content of which is incorporated herein by reference. Technical Field

[0003] The present disclosure generally relates to using a messaging application to provide an augmented reality (AR) experience. Background Art

[0004] Augmented reality (AR) is a modification of a virtual environment. For example, in virtual reality (VR), a user is fully immersed in a virtual world, while in AR, a user is immersed in a world where virtual objects are combined with the real world or virtual objects are superimposed on the real world. AR systems are designed to generate and present virtual objects that realistically interact with the real-world environment and with each other. Examples of AR applications can include single-player or multi-player video games, instant messaging systems, and the like. Brief Description of the Drawings

[0005] In the drawings, which are not necessarily drawn to scale, like reference numerals may describe similar components in different views. To easily identify the discussion of any particular element or action, one or more of the most significant digits in the reference numeral refers to the figure number in which the element is first introduced. Some non-limiting examples are shown in the drawings, in which:

[0006] Figure 1 is a graphical representation of a networked environment in which the present disclosure may be deployed, according to some examples.

[0007] Figure 2 is a graphical representation of a messaging client application, according to some examples.

[0008] Figure 3 is a graphical representation of a data structure maintained in a database, according to some examples.

[0009] Figure 4 is a graphical representation of a message, according to some examples.

[0010] Figures 5 to 7 is a block diagram showing components of an example body stylization system, according to some examples.

[0011] Figure 8 is a graphical representation of an output of a body stylization system, according to some examples.

[0012] Figure 9 is a flowchart showing an example operation of a body stylization system, according to some examples.

[0013] Figure 10 is a pictorial representation of a machine in the form of a computer system within which a set of instructions can be executed to cause the machine to perform any one or more of the methods discussed herein.

[0014] Figure 11 is a block diagram showing a software architecture within which an example can be implemented. DETAILED DESCRIPTION

[0015] The following description includes systems, methods, techniques, instruction sequences, and computer program products that embody illustrative examples of the present disclosure. In the following description, for purposes of explanation, numerous specific details are set forth to provide an understanding of the various examples. However, it will be apparent to one of ordinary skill in the art that the examples may be practiced without these specific details. Generally, well-known instruction instances, protocols, structures, and techniques have not been shown in detail.

[0016] Generally, VR and AR systems display an image representing a given user by capturing an image of the user and additionally obtaining a depth map of the real-world human body depicted in the image using a depth sensor. By processing the depth map and the image together, VR and AR systems can detect the user's positioning in the image and can appropriately modify the user or the background in the image. While such systems work well, the need for a depth sensor limits their scope of application. This is because adding a depth sensor to a user device for the purpose of modifying the image increases the total cost and complexity of the device, thereby making these devices less attractive.

[0017] Some systems do not require the use of depth sensors to modify images. For example, some systems allow users to replace the background in a video conference where the user's face is detected. Specifically, such systems can use specialized techniques optimized for identifying the user's face to identify the background in an image depicting the user's face. These systems can then replace only those pixels depicting the background such that the real-world background in the image is replaced with an alternative background. However, such systems typically cannot identify the user's entire body. Thus, if the user is at a distance greater than a threshold distance from the imaging device such that more than just the user's face is captured by the imaging device, using an alternative background to replace the background starts to fail. In such cases, the image quality is severely affected, and multiple parts of the user's face and body may be inadvertently removed by the system because the system incorrectly identifies such parts as belonging to the background of the image rather than the foreground. Additionally, when more than one user is depicted in an image or video feed, such systems cannot appropriately replace the background. Since such systems typically cannot distinguish the entire body of the user in the image from the background, these systems also cannot apply visual effects to certain parts of the user's body, such as transforming, blending, morphing, stylizing, or morphing (morphing) body parts (e.g., the face) into AR graphics.

[0018] Some AR systems allow AR graphics or AR elements to be added to an image or video to provide an engaging AR experience. Such systems can receive AR graphics from a designer and can scale and position the AR graphics within the image or video. To improve the placement and positioning of the AR graphics on the person depicted in the image or video, such systems detect the person depicted in the image or video and generate a skeleton (rig) representing the person's bones. The skeleton is then used to adjust the AR graphics based on changes in the movement of the skeleton. While such methods typically work well, generating the skeleton of the person in real time to adjust the placement of the AR graphics increases the processing complexity and power as well as memory requirements. This makes such systems inefficient or unable to run on small-scale mobile devices without sacrificing computational resources or processing speed. Additionally, the skeleton only represents the movement of the skeletal or bony structure of the person in the image or video and does not take into account any kind of external physical properties of the person, such as density, weight, skin characteristics, etc. Thus, any AR graphics in these systems can be adjusted in scale and position but cannot be deformed based on other physical properties of the person. Additionally, AR graphics designers typically need to create compatible skeletons for their AR graphics.

[0019] Some typical systems use grid - based reconstruction to modify portions of a face depicted in an image. These systems typically receive a grid that defines aspects of the face to be modified. The grid is then applied to a live image to modify the face depicted in the image. These systems are typically limited in their functionality because they can only realistically modify the face depicted in the image and do not perform correctly when the user's entire body is depicted in the image. Such systems often end up disproportionately adjusting parts of the body, which results in unrealistic modifications being applied to the image.

[0020] The techniques of the present disclosure improve the efficiency of using an electronic device by: using a machine - learning model to estimate or predict a stylized version of the entire body of a real - world person depicted in a received image or video in real time. The estimated or predicted stylized version of the person's entire body is then applied to a portion of the image to replace the depiction of the real - world person with the stylized version of the person's entire body. By using a trained machine - learning model to generate the stylized version of the person's entire body, the techniques of the present disclosure can apply one or more visual effects to an image or video associated with a real - world person depicted in the image or video in a more efficient and realistic manner. This can be done without the need to generate a skeleton or bone structure of the depicted object or perform any manual adjustments or image - processing techniques that may be time - consuming. Specifically, taking into account the movement and pose information of the person, the techniques of the present disclosure can deform, transform, change, stylize, and / or blend one or more body parts of the person's entire body depicted in the image or video into one or more AR elements.

[0021] This simplifies the process of adding AR graphics to an image or video, which significantly reduces the design constraints and costs when generating such AR graphics, and reduces the processing complexity, power, and storage requirements. This also improves the illusion that the AR graphics are part of the real - world environment depicted in an image or video depicting a person. This enables seamless and efficient addition of AR graphics to the underlying image or video in real time on a small mobile device. The techniques of the present disclosure can be applied specifically or primarily to a mobile device without the need for the mobile device to send the image / video to a server. In other examples, the techniques of the present disclosure are applied specifically or primarily to a remote server, or can be distributed between a mobile device and a server.

[0022] In addition, the techniques of the present disclosure allow AR graphic designers to generate a target style for their AR graphics without creating a compatible skeleton for the AR graphics, which saves time, effort, and reduces the complexity of creation. The techniques of the present disclosure use the target style to train one or more generative adversarial networks to generate training data to train a machine learning model to stylize the full body of a real-world person into the target style. Specifically, the techniques of the present disclosure can receive an image depicting the full body of a real-world person. The techniques of the present disclosure apply the machine learning model to the image to generate a stylized version of the full body of the real-world person corresponding to a given style. The training data can be used to train the machine learning model to establish a relationship between multiple training images (which depict the full body of a synthetically rendered person) and the corresponding ground truth stylized versions of the full body of a person of a given style. The techniques of the present disclosure replace the depiction of the full body of the real-world person in the image with the generated stylized version of the full body of the real-world person. Once trained, the machine learning model can operate in real time to generate or predict a stylized version of the full body of a person depicted in a received real-time image or video. The generated or predicted stylized version can be applied to the real-time image or video to generate a desired effect, where the real-world person depicted is deformed or modified to have dimensions, height, proportions, and / or appearance corresponding to the target style, which is used to train the machine learning model.

[0023] Accordingly, a realistic display can be provided that shows a user or person being stylized according to a selected or given style while the user or person moves around in a 3D video (including changes to body shape, body state, body style and visual attributes, body attributes, position, and rotation) in a manner that is intuitive for the user to interact with and select. This improves the overall user experience when using an electronic device. Additionally, by providing such an AR experience without using a depth sensor, the total amount of system resources required to complete a task is reduced.

[0024] Networked computing environment

[0025] Figure 1FIG. 0 is a block diagram illustrating an example messaging system 100 for exchanging data (e.g., messages and associated content) over a network. Messaging system 100 includes multiple instances of client devices 102, each of the multiple instances hosting several applications including a messaging client 104 and other external applications 109 (e.g., third-party applications). Each messaging client 104 is communicatively coupled via a network 112 (e.g., the Internet) to other instances of messaging clients 104, a messaging server system 108, and an external application server 110 (e.g., hosted on respective other client devices 102). Messaging client 104 can also communicate with locally hosted third-party applications (e.g., external application 109) using an application programming interface (API).

[0026] Client device 102 can operate as a stand-alone device or can be coupled (e.g., networked) to other machines. In a networked deployment, client device 102 can operate in the capacity of a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. Client device 102 can include, but is not limited to: server computers, client computers, personal computers (PCs), tablet computers, laptop computers, netbooks, set-top boxes (STBs), personal digital assistants (PDAs), entertainment media systems, cellular phones, smart phones, mobile devices, wearable devices (e.g., smart watches), smart home devices (e.g., smart appliances), other smart devices, web appliances, network routers, network switches, network bridges, or any machine capable of performing the disclosed operations. Additionally, although only a single client device 102 is shown, the term "client device" should also be considered to include a collection of machines that individually or jointly perform the disclosed operations.

[0027] In some examples, client device 102 can include AR glasses or an AR headset, where virtual content is displayed within the lenses of the glasses while the user views the real-world environment through the lenses. For example, an image can be presented on a transparent display that allows the user to view both the content presented on the display and real-world objects simultaneously.

[0028] The messaging client 104 is capable of communicating with other messaging clients 104 and the messaging server system 108 via the network 112 and exchanging data. The data exchanged between messaging clients 104 and between the messaging client 104 and the messaging server system 108 includes functionality (e.g., commands to activate functionality) and payload data (e.g., text, audio, video, or other multimedia data). In some examples, the client device 102 includes an eyewear device configured to generate augmented reality objects within the lenses of the eyewear device to provide the augmented reality experiences discussed herein.

[0029] The messaging server system 108 provides server-side functionality to a particular messaging client 104 via the network 112. While certain functions of the messaging system 100 are described herein as being performed by the messaging client 104 or by the messaging server system 108, the location of certain functions within the messaging client 104 or within the messaging server system 108 can be a design choice. For example, it may be technically preferable to initially deploy certain technologies and functionality within the messaging server system 108, but later migrate the technology and functionality to the messaging client 104 where the client device 102 has sufficient processing power.

[0030] The messaging server system 108 supports various services and operations provided to the messaging client 104. Such operations include transmitting data to the messaging client 104, receiving data from the messaging client 104, and processing data generated by the messaging client 104. As an example, the data can include message content, client device information, geographic location information, media enhancements and overlays, message content persistence conditions, social network information, and live event information. Data exchange within the messaging system 100 is initiated and controlled through functions available via the user interface (UI) of the messaging client 104.

[0031] Turning now specifically to the messaging server system 108, the API server 116 is coupled to the application server 114 and provides a programming interface to the application server 114. The application server 114 is communicatively coupled to the database server 120, which facilitates access to the database 126 that stores data associated with messages processed by the application server 114. Similarly, the web server 128 is coupled to the application server 114 and provides a web-based interface to the application server 114. To that end, the web server 128 processes incoming network requests via the Hypertext Transfer Protocol (HTTP) and several other related protocols.

[0032] The API server 116 receives and sends message data (e.g., commands and message payloads) between the client device 102 and the application server 114. Specifically, the API server 116 provides a set of interfaces (e.g., routines and protocols) that can be invoked or queried by the messaging client 104 to activate the functions of the application server 114. The API server 116 exposes various functions supported by the application server 114, including: account registration; login functionality; sending messages from a particular messaging client 104 to another messaging client 104 via the application server 114; sending media files (e.g., images or videos) from the messaging client 104 to the messaging server 118 and making them available for possible access by another messaging client 104; setting media data collections (e.g., stories); retrieving a friend list of the user of the client device 102; retrieving such collections; retrieving messages and content; adding and deleting entities (e.g., friends) in an entity graph (e.g., a social graph); locating friends within the social graph; and opening application events (e.g., related to the messaging client 104).

[0033] The application server 114 hosts a number of server applications and subsystems, including, for example, the messaging server 118, the image processing server 122, and the social network server 124. The messaging server 118 implements a number of message processing techniques and functions, particularly those related to the aggregation and other processing of the content (e.g., text and multimedia content) included in messages received from multiple instances of the messaging client 104. As will be described in further detail, text and media content from multiple sources can be aggregated into collections of content (e.g., referred to as stories or galleries). These collections are then made available to the messaging client 104. Given the hardware requirements for other processor- and memory-intensive processing of the data, such processing can also be performed on the server side by the messaging server 118.

[0034] The application server 114 also includes an image processing server 122 that is dedicated to performing various image processing operations, typically on images or videos within the payload of messages sent from or received at the messaging server 118.

[0035] The image processing server 122 is used to implement the enhancement system 208( Figure 2The scanning function (shown in). The scanning function includes activating and providing one or more AR experiences on the client device 102 when an image is captured by the client device 102. Specifically, the messaging client 104 on the client device 102 can be used to activate the imaging device. The imaging device displays one or more real-time images or videos along with one or more icons or identifiers of one or more AR experiences to the user. The user can select a given identifier among the identifiers to initiate the corresponding AR experience or perform a desired image modification (e.g., replacing the clothing worn by the user in the video or deforming, changing, blending, or transforming a part of the body of a person or user into an AR graphic such as an AR werewolf or an AR bat).

[0036] The social network server 124 supports various social networking functions and services and makes these functions and services available to the messaging server 118. To this end, the social network server 124 maintains and accesses the entity graph 308 within the database 126 (as Figure 3 shown). Examples of the functions and services supported by the social network server 124 include identifying other users with whom a particular user of the messaging system 100 has a relationship or whom the particular user is "following", and also include identifying the interests of a particular user and other entities.

[0037] Returning to the messaging client 104, the features and functions of external resources (e.g., external applications 109 or mini-programs) are available to the user via the interface of the messaging client 104. The messaging client 104 receives a user selection of an option for initiating or accessing the features of an external resource such as an external application 109 (e.g., a third-party resource). The external resource can be a third-party application (external application 109) installed on the client device 102 (e.g., a "native application"), or a scaled-down version of a third-party application (e.g., a "mini-program") hosted on the client device 102 or located at a remote end of the client device 102 (e.g., on a third-party server 110). The scaled-down version of the third-party application includes a subset of the features and functions of the third-party application (e.g., the full-scale, native version of a third-party stand-alone application) and is implemented using a markup language document. In one example, the scaled-down version of the third-party application (e.g., a "mini-program") is a web-based markup language version of the third-party application and is embedded in the messaging client 104. In addition to using a markup language document (e.g., a.*ml file), the mini-program can include a scripting language (e.g., a.*js file or a.json file) and a style sheet (e.g., a.*ss file).

[0038] In response to receiving a user selection of an option for a feature for launching or accessing an external resource (External Application 109), the messaging client 104 determines whether the selected external resource is a web-based external resource or a locally installed external application. In some cases, an external application 109 locally installed on the client device 102 can be launched independently of and separately from the messaging client 104, for example, by selecting an icon corresponding to the external application 109 on the home screen of the client device 102. A scaled-down version of such an external application can be launched or accessed via the messaging client 104, and in some examples, any part of the scaled-down external application cannot be (or only a limited part can be) accessed outside of the messaging client 104. The scaled-down external application can be launched by receiving a markup language document associated with the scaled-down external application from the external application server 110 by the messaging client 104 and processing such a document.

[0039] In response to determining that the external resource is a locally installed external application 109, the messaging client 104 instructs the client device 102 to launch the external application 109 by executing locally stored code corresponding to the external application 109. In response to determining that the external resource is a web-based resource, the messaging client 104 communicates with the external application server 110 to obtain a markup language document corresponding to the selected resource. The messaging client 104 then processes the obtained markup language document to present the web-based external resource within the user interface of the messaging client 104.

[0040] The messaging client 104 can notify a user of the client device 102 or other users related to such a user (e.g., "friends") of activities occurring in one or more external resources. For example, the messaging client 104 can provide notifications to participants in a conversation (e.g., a chat session) in the messaging client 104 regarding current or recent use of an external resource by one or more members of a user group. One or more users can be invited to join an active external resource or launch an external resource that was recently used but is currently inactive (within the group of friends). The external resource can provide the ability for respective participants in the conversation using the corresponding messaging client 104 to share items, conditions, statuses, or locations within the external resource with one or more members of the user group entering the chat session. The shared item can be an interactive chat card that members of the chat can interact with to, for example, launch a corresponding external resource, view specific information within the external resource, or take the members of the chat to a specific location or status within the external resource. Within a given external resource, a response message can be sent to the user on the messaging client 104. The external resource can selectively include different media items in the response based on the current context of the external resource.

[0041] The messaging client 104 can present a user with a list of available external resources (e.g., third-party or external applications 109 or mini-programs) to initiate or access a given external resource. The list can be presented in a context-sensitive menu. For example, icons representing different external applications in the external application 109 (or mini-program) can vary based on how the menu is launched (e.g., launched from a conversational interface or from a non-conversational interface) by the user.

[0042] The messaging client 104 can present one or more AR experiences to the user. As an example, the messaging client 104 can detect a person or user in an image or video captured by the client device 102. The messaging client 104 can apply a machine learning model (associated with a target style) to generate or estimate a stylized version of the person or object depicted in the image. The messaging client 104 can then generate a modified image depicting the stylized version of the person. Although the examples of the present disclosure are discussed with respect to modifying a person or user depicted in an image or video, similar techniques can also be applied to modify any other real-world object, such as an animal, furniture, a building, etc.

[0043] This provides the illusion that the stylized real-world object or person is actually included in the real-world environment depicted in the modified image or video, which improves the overall user experience.

[0044] System architecture

[0045] Figure 2 is a block diagram showing additional details regarding the messaging system 100 according to some examples. Specifically, the messaging system 100 is shown to include a messaging client 104 and an application server 114. The messaging system 100 includes several subsystems that are supported on the client side by the messaging client 104 and on the server side by the application server 114. These subsystems include, for example, an ephemeral timer system 202, a collection management system 204, an enhancement system 208, a map system 210, a game system 212, and an external resource system 220.

[0046] The ephemeral timer system 202 is responsible for implementing temporary or time-limited access to content by the messaging client 104 and the messaging server 118. The ephemeral timer system 202 includes several timers that selectively enable access to messages and associated content (e.g., for presentation and display) via the messaging client 104 based on the duration and display parameters associated with a message or a collection of messages (e.g., a story). Additional details regarding the operation of the ephemeral timer system 202 are provided below.

[0047] The collection management system 204 is responsible for managing collections or sets of media (e.g., collections of text, image, video, and audio data). Collections of content (e.g., messages, including images, videos, text, and audio) can be organized into an "event gallery" or "event story". Such collections can be available for a specified period of time (e.g., during the duration of an event related to the content). For example, content related to a concert can be made available as a "story" during the duration of the concert. The collection management system 204 can also be responsible for publishing an icon that provides a notification of the existence of a particular collection to the user interface of the messaging client 104.

[0048] In addition, the collection management system 204 also includes a curation interface 206 that allows a collection manager to manage and curate a particular content collection. For example, the curation interface 206 enables an event organizer to curate a collection of content related to a particular event (e.g., delete inappropriate content or redundant messages). Further, the collection management system 204 employs machine vision (or image recognition technology) and content rules to automatically curate content collections. In some examples, compensation can be paid to a user for including user-generated content in a collection. In such cases, the collection management system 204 operates to automatically pay such users for the use of their content.

[0049] The enhancement system 208 provides various functions that enable a user to enhance (e.g., annotate or otherwise modify or edit) media content associated with a message. For example, the enhancement system 208 provides functions related to generating and publishing a media overlay for a message processed by the messaging system 100. The enhancement system 208 operably provides a media overlay or enhancement (e.g., an image filter) to the messaging client 104 based on the geographical location of the client device 102. In another example, the enhancement system 208 operably provides a media overlay to the messaging client 104 based on other information such as the social network information of a user of the client device 102. The media overlay can include audio and visual content as well as visual effects. Examples of audio and visual content include pictures, text, logos, animations, and sound effects. Examples of visual effects include color overlays. The audio and visual content or visual effects can be applied to a media content item (e.g., a photo) at the client device 102. For example, the media overlay can include text, graphic elements, or images that can be superimposed on top of a photo taken by the client device 102. In another example, the media overlay includes a location identifier overlay (e.g., Venice Beach), a name of a live event, or a business name overlay (e.g., Beach Café). In another example, the enhancement system 208 uses the geographical location of the client device 102 to identify a media overlay that includes the business name at the geographical location of the client device 102. The media overlay can include other markers associated with the business. The media overlay can be stored in the database 126 and accessed via the database server 120.

[0050] In some examples, the enhancement system 208 provides a user-based publishing platform that enables a user to select a geographical location on a map and upload content associated with the selected geographical location. The user can also specify the circumstances in which a particular media overlay should be provided to other users. The enhancement system 208 generates a media overlay that includes the uploaded content and associates the uploaded content with the selected geographical location.

[0051] In other examples, the augmentation system 208 provides a merchant-based publishing platform that enables a merchant to select a specific media overlay associated with a geographical location via a bidding process. For example, the augmentation system 208 associates the media overlay of the highest bidding merchant with the corresponding geographical location for a predefined amount of time. The augmentation system 208 communicates with the image processing server 122 to obtain an augmented reality experience and presents an identifier of such an experience in one or more user interfaces (e.g., as an icon on a live image or video, or as a thumbnail or icon in an interface dedicated to presenting the augmented reality experience identifier). Once an AR experience is selected, one or more images, videos, or AR graphic elements are retrieved and presented as an overlay on the image or video captured by the client device 102. In some cases, the camera device is switched to a front view (e.g., the front camera device of the client device 102 is activated in response to the activation of a specific AR experience), and the image from the front camera device of the client device 102 rather than the rear camera device of the client device 102 starts to be displayed on the client device 102. One or more images, videos, or AR graphic elements are retrieved and presented as an overlay on the image captured and displayed by the front camera device of the client device 102.

[0052] In other examples, the augmentation system 208 is capable of communicating and exchanging data via the network 112 with another augmentation system 208 and a server on another client device 102. The exchanged data may include: a session identifier identifying a shared AR session; a transformation between the first client device 102 and the second client device 102 (e.g., multiple client devices 102 include a first device and a second device), the transformation being used to align the shared AR session to a common origin; a common coordinate system; and functions (e.g., commands for activating functions) and other payload data (e.g., text, audio, video, or other multimedia data).

[0053] The augmentation system 208 sends the transformation to the second client device 102 such that the second client device 102 can adjust the AR coordinate system based on the transformation. In this way, the first client device and the second client device 102 synchronize their coordinate systems and frames to display the content in the AR session. Specifically, the augmentation system 208 calculates the origin of the second client device 102 in the coordinate system of the first client device 102. Then, the augmentation system 208 can determine the offset in the coordinate system of the second client device 102 based on the position of the origin in the coordinate system of the second client device 102 from the perspective of the second client device 102. A transformation is generated using the offset such that the second client device 102 generates AR content according to the common coordinate system with the first client device 102.

[0054] The augmentation system 208 can communicate with the client device 102 to establish a separate or shared AR session. The augmentation system 208 can also be coupled to the messaging server 118 to establish an electronic group communication session (e.g., group chat, instant messaging) for the client device 102 in a shared AR session. The electronic group communication session can be associated with a session identifier provided by the client device 102 to obtain access to the electronic group communication session and the shared AR session. In one example, the client device 102 first obtains access to the electronic group communication session and then obtains a session identifier in the electronic group communication session that allows the client device 102 to access the shared AR session. In some examples, the client device 102 is able to access the shared AR session without the assistance of or communication with the augmentation system 208 in the application server 114.

[0055] The map system 210 provides various geographical location functions and supports the presentation of map-based media content and messages by the messaging client 104. For example, the map system 210 enables the display on a map of user icons or avatars (e.g., in the stored profile data 316) to indicate the current or past locations of a user's "friends", as well as media content (e.g., a collection of messages including photos and videos) generated by such friends within the context of the map. For example, a message posted by a user from a specific geographical location to the messaging system 100 can be displayed to the user's "friends" within the context of that specific location on the map in the map interface of the messaging client 104. A user can also share his or her location and status information with other users of the messaging system 100 (e.g., using an appropriate status avatar) via the messaging client 104, where the location and status information is similarly displayed to selected users within the context of the map interface of the messaging client 104.

[0056] The game system 212 provides various game functions within the context of the messaging client 104. The messaging client 104 provides a game interface that provides a list of available games (e.g., web-based games or web-based applications) that can be launched by the user within the context of the messaging client 104 and played with other users of the messaging system 100. The messaging system 100 also enables a specific user to invite such other users to participate in playing a specific game by sending an invitation from the messaging client 104 to the other users. The messaging client 104 also supports both voice messaging and text messaging (e.g., chat) within the context of playing a game, provides a leaderboard for the game, and also supports providing in-game rewards (e.g., game currency and items).

[0057] The external resource system 220 provides an interface for the messaging client 104 to communicate with the external application server 110 to initiate or access external resources. Each external resource (application) server 110 hosts an application based on a markup language (e.g., HTML5) or a scaled-down version of an external application (e.g., a game, utility, payment, or ride-sharing application external to the messaging client 104). The messaging client 104 can initiate the web-based resource (e.g., application) by accessing an HTML5 file from an external resource (application) server 110 associated with the web-based resource. In some examples, a software development kit (SDK) provided by the messaging server 118 is used to program the applications hosted by the external resource server 110 in JavaScript. The SDK includes APIs with functions that can be invoked or activated by web-based applications. In some examples, the messaging server 118 includes a JavaScript library that provides access to certain user data of the messaging client 104 to a given third-party resource. HTML5 is used as an example technology for programming games, but applications and resources programmed based on other technologies can be used.

[0058] To integrate the functionality of the SDK into a web-based resource, the SDK is downloaded from the messaging server 118 by the external resource (application) server 110 or otherwise received by the external resource (application) server 110. Once downloaded or received, the SDK is included as part of the application code of the web-based external resource. The code of the web-based resource can then call or activate certain functions of the SDK to integrate features of the messaging client 104 into the web-based resource.

[0059] The SDK stored on the messaging server 118 effectively provides a bridge between external resources (e.g., third-party or external applications 109 or applets and the messaging client 104). This provides users with a seamless experience of communicating with other users on the messaging client 104 while also retaining the look and feel of the messaging client 104. To bridge the communication between the external resources and the messaging client 104, in some examples, the SDK facilitates the communication between the external resource server 110 and the messaging client 104. In some examples, the WebView JavaScript Bridge running on the client device 102 establishes two one-way communication channels between the external resources and the messaging client 104. Messages are sent asynchronously between the external resources and the messaging client 104 via these communication channels. Each SDK function activation is sent as a message and a callback. Each SDK function is implemented by constructing a unique callback identifier and sending a message with that callback identifier.

[0060] By using the SDK, not all information from the messaging client 104 is shared with the external resource server 110. The SDK restricts which information is shared based on the needs of the external resources. In some examples, each external resource server 110 provides an HTML5 file corresponding to the web-based external resource to the messaging server 118. The messaging server 118 can add a visual representation of the web-based external resource (e.g., box art or other graphics) to the messaging client 104. Once the user selects the visual representation or indicates that the messaging client 104 accesses the features of the web-based external resource through the GUI of the messaging client 104, the messaging client 104 obtains the HTML5 file and instantiates the resources required to access the features of the web-based external resource.

[0061] The messaging client 104 presents a graphical user interface for an external resource (e.g., a login page or a splash screen). During, before, or after presenting the login page or splash screen, the messaging client 104 determines whether the launched external resource has previously been authorized to access the user data of the messaging client 104. In response to determining that the launched external resource has previously been authorized to access the user data of the messaging client 104, the messaging client 104 presents another graphical user interface of the external resource, which includes the functions and features of the external resource. In response to determining that the launched external resource has not previously been authorized to access the user data of the messaging client 104, after a threshold period (e.g., 3 seconds) of displaying the login page or splash screen of the external resource, the messaging client 104 slides out a menu for authorizing the external resource to access the user data (e.g., animating the menu to emerge from the bottom of the screen to the middle or other part of the screen). The menu identifies the type of user data that the external resource will be authorized to use. In response to receiving a user selection of the accept option, the messaging client 104 adds the external resource to the list of authorized external resources and enables the external resource to access the user data from the messaging client 104. In some examples, the messaging client 104 authorizes the external resource to access the user data according to the OAuth 2 framework.

[0062] The messaging client 104 controls the type of user data shared with an external resource based on the type of the authorized external resource. For example, access to a first type of user data (e.g., only 2D avatars of users with or without different avatar characteristics) is provided to an external resource that includes a full-scale external application (e.g., a third-party or external application 109). As another example, access to a second type of user data (e.g., payment information, 2D avatars of the user, 3D avatars of the user, and avatars with various avatar characteristics) is provided to an external resource that includes a scaled-down version of an external application (e.g., a web-based version of a third-party application). Avatar characteristics include different ways of customizing the appearance and feel of an avatar (e.g., different poses, facial features, clothing, etc.).

[0063] The body stylization system 224 receives an image that includes a depiction of a real-world object (e.g., a full body of a human, an animal, or other living or inanimate object). The body stylization system 224 applies a machine learning model to the image to generate a stylized version of the real-world object, such as a stylized version of a full body of a human. Training data can be used to train the machine learning model to establish a relationship between multiple training images (which depict full bodies of synthetically rendered humans) and the corresponding ground-truth stylized versions of full bodies of humans in a given style. The body stylization system 224 generates a modified image depicting the stylized real-world object.

[0064] In some examples, the body stylization system 224 is a component accessed by an AR / VR application implemented on the client device 102. The AR / VR application uses an RGB camera device to capture a monocular image of a user or a person. The AR / VR application applies various trained machine learning techniques to the captured user image to generate a stylized full body of the person, and applies one or more AR visual effects to the captured image to generate a modified image including the AR visual effects. In some implementations, the AR / VR application continuously captures images of the user and updates the stylized version of the person in real time or periodically to continuously or periodically update the one or more applied visual effects. This allows the user to move around in the real world and see the update of one or more visual effects in real time.

[0065] In training, the body stylization system 224 obtains training data, which includes: a plurality of training images depicting the full body of a synthetically rendered person; and the corresponding ground truth stylized version of the full body of a person in a given style. A machine learning technique (e.g., a deep neural network or other machine learning model) is trained based on the features of the plurality of training images in the training data. Specifically, the machine learning technique is applied to a first training data set, which includes a first training image among the plurality of training images depicting the full body of a synthetically rendered person, to generate an estimated stylized version of the full body of the person depicted in the first training image. The machine learning technique calculates the deviation between the estimated stylized version of the full body of the person depicted in the first training image and the ground truth stylized version of the full body of the person depicted in the first training image. One or more parameters of the machine learning technique (model) are updated based on the deviation between the estimated stylized version of the full body of the person depicted in the first training image and the ground truth stylized version of the full body of the person depicted in the first training image.

[0066] This process is repeated until a stopping criterion is reached. At this time, the trained machine learning model is output, stored, and used to generate a stylized version of the full body of a person depicted in an image or video in real time.

[0067] Data Architecture

[0068] Figure 3 is a schematic diagram showing a data structure 300 that can be stored in the database 126 of the messaging server system 108 according to certain examples. Although the content of the database 126 is shown as including several tables, it will be appreciated that the data can be stored in other types of data structures (e.g., as an object-oriented database).

[0069] The database 126 includes message data stored in the message table 302. For any particular message, the message data includes at least message sender data, message recipient (or receiver) data, and a payload. Refer to the followingFigure 4 Describe additional details about information that can be included in a message and within the message data stored in the message table 302.

[0070] The entity table 306 stores entity data and is linked (e.g., by reference) to the entity graph 308 and the profile data 316. Entities whose records are maintained within the entity table 306 can include individuals, corporate entities, organizations, objects, locations, events, etc. Any entity for which the messaging server system 108 stores data about it can be an identified entity, regardless of the entity type. Each entity is provided with a unique identifier and an entity type identifier (not shown).

[0071] The entity graph 308 stores information related to the relationships and associations between entities. By way of example only, such relationships can be social, professional (e.g., working in the same company or organization), interest-based, or activity-based.

[0072] The profile data 316 stores various types of profile data about a particular entity. Based on privacy settings specified by the particular entity, the profile data 316 can be selectively used and presented to other users of the messaging system 100. In the case where the entity is an individual, the profile data 316 includes, for example, a username, a phone number, an address, and settings (e.g., notification and privacy settings), as well as an avatar representation (or a set of such avatar representations) selected by the user. Then, a particular user can selectively include one or more of these avatar representations within the content of a message transmitted via the messaging system 100 and on a map interface displayed by the messaging client 104 to other users. The set of avatar representations can include “status avatars” that present a graphical representation of a status or activity that the user can choose to convey at a particular time.

[0073] In the case where the entity is a group, in addition to the group name, members, and various settings (e.g., notifications) for the relevant group, the profile data 316 for the group can similarly include one or more avatar representations associated with the group.

[0074] The database 126 also stores enhancement data, such as overlays or filters, in the enhancement table 310. The enhancement data is associated with videos (the data of which is stored in the video table 304) and images (the data of which is stored in the image table 312) and is applied to the videos and images.

[0075] The database 126 may also store data related to individual and shared AR sessions. This data may include data transmitted between the AR session client controller of the first client device 102 and another AR session client controller of the second client device 102, as well as data transmitted between the AR session client controller and the augmentation system 208. The data may include data for establishing a common coordinate system for a shared AR scene, transformations between devices, session identifiers, images depicting bodies, skeletal joint positions, wrist joint positions, feet, etc.

[0076] In one example, a filter is an overlay that is displayed as an overlay on an image or video during presentation to a receiving user. Filters can be of various types, including a filter selected by a user from a set of filters presented to the sending user by the messaging client 104 when the sending user is composing a message. Other types of filters include location-based filters (also known as geo-filters), which can be presented to the sending user based on a geographical location. For example, based on geographical location information determined by the global positioning system (GPS) unit of the client device 102, the messaging client 104 can present location-based filters specific to a nearby or specific location within the user interface.

[0077] Another type of filter is a data filter, which can be selectively presented to the sending user by the messaging client 104 based on other input or information collected by the client device 102 during the message creation process. Examples of data filters include the current temperature at a specific location, the current speed at which the sending user is traveling, the battery life of the client device 102, or the current time.

[0078] Other augmentation data that can be stored in the image table 312 includes AR content items (e.g., corresponding to an applied augmented reality experience). An AR content item or AR item can be a real-time special effect and sound that can be added to an image or video.

[0079] As described above, the enhanced data includes AR content items, overlays, image transformations, AR images, AR identifiers or markers, and similar items related to modifications that can be applied to image data (e.g., video or images). This includes real-time modifications that modify the image when it is captured using the device sensors of the client device 102 (e.g., one or more camera devices) and then display the modified image on the screen of the client device 102. This also includes modifications to stored content (e.g., video clips in a gallery that can be modified). For example, in the client device 102 that can access multiple AR content items, a user can use a single video clip with multiple AR content items to see how different AR content items will modify the stored clip. For example, by selecting different AR content items for the content, multiple AR content items applying different pseudo-random movement models can be applied to the same content. Similarly, real-time video capture can be used with the shown modifications to show how the video image currently captured by the sensors of the client device 102 will modify the captured data. Such data can be only displayed on the screen without being stored in the memory, or the content captured by the device sensors can be recorded and stored in the memory with or without modifications (or in both cases). In some systems, the preview feature can show simultaneously how different AR content items will look within different windows on the display. This can, for example, enable viewing multiple windows with different pseudo-random animations on the display at the same time.

[0080] Thus, using data of AR content items and various systems or other such transformation systems that use the data to modify content can involve the detection of real-world objects (e.g., faces, hands, bodies, cats, dogs, surfaces, objects, etc.) in video frames, the tracking of such objects as they leave the field of view, enter the field of view, and move around in the field of view, and the modification or transformation of such objects when tracking them. In various examples, different methods can be used to implement such transformations. Some examples can involve: generating a 3D mesh model of one or more objects; and using the transformation and animated textures of the model within the video to implement the transformation. In other examples, tracking points on the object can be utilized to place an image or texture (which can be 2D or 3D) at the tracking location. In yet another example, neural network analysis of video frames can be used to place an image, model, or texture in the content (e.g., a frame of an image or video). Thus, AR content items or elements refer both to the images, models, and textures used to create transformations in the content and to the additional modeling and analysis information required to implement such transformations using object detection, tracking, and placement.

[0081] Real-time video processing can be performed using any kind of video data (e.g., video streams, video files, etc.) stored in the memory of any kind of computerized system. For example, a user can load video files and save them in the device's memory, or can use the device's sensors to generate video streams. In addition, computer animation models can be used to process any object, such as a human face and parts of the human body, an animal, or a non-living thing (e.g., a chair, a car, or other objects).

[0082] In some examples, when a specific modification is selected along with the content to be transformed, the computing device identifies the elements to be transformed and then, if the elements exist in the frames of the video, detects and tracks them. The elements of the object are modified according to the modification request, thereby transforming the frames of the video stream. For different types of transformations, the transformation of the frames of the video stream can be performed by different methods. For example, for frame transformations mainly involving changing the form of the elements of an object, the feature points of each element of the object are calculated (e.g., using an Active Shape Model (ASM) or other known methods). Then, a grid based on the feature points is generated for each element in at least one element of the object. This grid is used for the subsequent stage of tracking the elements of the object in the video stream. During the tracking process, the grids mentioned for each element are aligned with the positions of the elements. Then, additional points are generated on the grid. A first set of first points is generated for each element based on the modification request, and a set of second points is generated for each element based on this set of first points and the modification request. Then, the frames of the video stream can be transformed by modifying the elements of the object based on this set of first points, this set of second points, and the grid. In such a method, the background of the modified object can also be changed or distorted by tracking and modifying the background.

[0083] In some examples, a transformation that uses the elements of an object to change some regions of the object can be performed by calculating the feature points of each element of the object and generating a grid based on the calculated feature points. Points are generated on the grid, and then various regions are generated based on the points. Then, the elements of the object are tracked by aligning the regions for each element with the positions of each element in at least one element, and the attributes of the regions can be modified based on the modification request, thereby transforming the frames of the video stream. Depending on the specific modification request, the attributes of the mentioned regions can be transformed in different ways. Such modifications can involve: changing the color of the region; removing at least some partial regions from the frames of the video stream; including one or more new objects in the region based on the modification request; and modifying or distorting the region or the elements of the object. In various examples, any combination of such modifications or other similar modifications can be used. For some models to be animated, some feature points can be selected as control points for the entire state space to be used to determine the options for the model animation.

[0084] In some examples of computer animation models for transforming image data using body / person detection, a specific body / person detection algorithm (e.g., 3D human pose estimation and mesh reconstruction processing) is used to detect the body / person in the image. Then, the ASM algorithm is applied to the body / person region of the image to detect body / person feature reference points.

[0085] Other methods and algorithms suitable for body / person detection can be used. For example, in some examples, landmarks are used to locate features, where a landmark represents a distinguishable point present in most of the images under consideration. For example, for body / person landmarks, the position of the left arm can be used. If the initial landmark is not identifiable, secondary landmarks can be used. Such a landmark identification process can be used for any such object. In some examples, a set of landmarks forms a shape. The coordinates of the points in the shape can be used to represent the shape as a vector. A similarity transformation that minimizes the average Euclidean distance between the shape points (which allows translation, scaling, and rotation) is used to align one shape with another. The mean shape is the average of the aligned training shapes.

[0086] In some examples, the search for landmarks starts from the mean shape aligned with the position and size of the body / person determined by the global body / person detector. Then, such a search repeats the following steps: a tentative shape is proposed by adjusting the positions of the shape points by template matching the image texture around each point, and then the tentative shape is made to conform to the global shape model until convergence occurs. In some systems, individual template matching is unreliable, and the shape model pools the results of weak template matches to form a stronger overall classifier. The entire search is repeated at each level in an image pyramid from coarse resolution to fine resolution.

[0087] The transformation system can capture an image or video stream on a client device (e.g., client device 102) and perform complex image manipulations locally on the client device 102 while maintaining an appropriate user experience, computation time, and power consumption. Complex image manipulations can include size and shape changes, emotion transitions (e.g., changing a face from a frown to a smile), state transitions (e.g., aging a subject, reducing apparent age, changing gender), style transitions, application of graphical elements, 3D human pose estimation, 3D body mesh reconstruction, and any other suitable image or video manipulations implemented by a convolutional neural network that has been configured to execute efficiently on the client device 102.

[0088] In some examples, a computer animation model for transforming image data can be used by a system in which a user can use a client device 102 having a neural network to capture an image or video stream of the user (e.g., a selfie), where the neural network operates as part of a messaging client 104 operating on the client device 102. A transformation system operating within the messaging client 104 determines the presence of a body / person within the image or video stream and provides a modification icon associated with the computer animation model for transforming the image data, or the computer animation model can be present in association with the interfaces described herein. The modification icon includes changes that can be the basis for modifying the body / person of the user within the image or video stream as part of a modification operation. Once the modification icon is selected, the transformation system initiates a process of transforming the user's image to reflect the selected modification icon (e.g., generating a smiling face for the user). Once an image or video stream is captured and a specified modification is selected, the modified image or video stream can be presented in a graphical user interface displayed on the client device 102. The transformation system can implement a complex convolutional neural network for a portion of the image or video stream to generate and apply the selected modification. That is, the user can capture an image or video stream and, once the modification icon is selected, the modified result can be presented to the user in real-time or near real-time. Additionally, in the case where a video stream is being captured and the selected modification icon remains toggled, the modification can be persistent. A neural network of machine learning can be used to implement such modifications.

[0089] A graphical user interface presenting the modifications performed by the transformation system can provide additional interaction options to the user. Such options can be based on the interface used to initiate the selection of a particular computer animation model and content capture (e.g., initiated from a content creator user interface). In various examples, after initially selecting the modification icon, the modification can be persistent. The user can turn the modification on or off and store it for later viewing or browsing to other areas of the imaging application by tapping or otherwise selecting the body / person being modified by the transformation system. In the case where multiple faces are modified by the transformation system, the user can globally toggle the modification on or off by tapping or selecting a single body / person modified and displayed within the graphical user interface. In some examples, individual body / persons within a group of multiple body / persons can be modified separately, or such modifications can be toggled individually by tapping or selecting a single body / person or a series of body / persons displayed within the graphical user interface.

[0090] The story table 314 stores data related to a collection of messages and associated image, video, or audio data, where the messages and associated image, video, or audio data are compiled into a collection (e.g., a story or a gallery). The creation of a particular collection can be initiated by a particular user (e.g., each user whose records are maintained in the entity table 306). A user can create a "personal story" in the form of a collection of content that has already been created and sent / broadcast by that user. To this end, the user interface of the messaging client 104 can include user-selectable icons to enable the sending user to add specific content to his or her personal story.

[0091] The collection can also constitute a "Live Story", which is a collection of content from multiple users created manually, automatically, or using a combination of manual and automatic techniques. For example, a "Live Story" can constitute a curated stream of user-submitted content from various locations and events. Users whose client devices have location services enabled and are at a common location event at a particular time can be presented, for example, via the user interface of the messaging client 104, with the option to contribute content to a particular Live Story. A Live Story can be identified to him or her by the messaging client 104 based on the user's location. The end result is a "Live Story" told from a community perspective.

[0092] Another type of content collection is called a "Location Story", which enables users whose client devices 102 are located within a particular geographical location (e.g., on a college or university campus) to contribute to a particular collection. In some examples, contributing to a Location Story may require secondary authentication to verify that the end user belongs to a particular organization or other entity (e.g., is a student on a university campus).

[0093] As mentioned above, the video table 304 stores video data, which, in one example, is associated with messages whose records are maintained within the message table 302. Similarly, the image table 312 stores image data associated with messages whose message data is stored in the entity table 306. The entity table 306 can associate various enhancements from the enhancement table 310 with the various images and videos stored in the image table 312 and the video table 304.

[0094] The trained machine learning techniques 307 store the parameters of one or more machine learning models that have been trained during the training of the body stylization system 224. For example, the trained machine learning techniques 307 store the trained parameters of one or more neural network machine learning techniques and / or generative adversarial networks (GANs).

[0095] Data communication architecture

[0096] Figure 4FIG. is a schematic diagram showing the structure of a message 400 according to some examples, which is generated by a messaging client 104 for transmission to another messaging client 104 or a messaging server 118. The content of a particular message 400 is used to populate a message table 302 stored in a database 126 accessible by the messaging server 118. Similarly, the content of the message 400 is stored in memory as "in-transit" or "in-flight" data of the client device 102 or the application server 114. The message 400 is shown to include the following example components:

[0097] · Message identifier 402: A unique identifier that identifies the message 400.

[0098] ● Message text payload 404: Text to be generated by a user via a user interface of the client device 102 and included in the message 400.

[0099] ● Message image payload 406: Image data captured by a camera device component of the client device 102 or retrieved from a memory component of the client device 102 and included in the message 400. The image data for a sent or received message 400 can be stored in an image table 312.

[0100] ● Message video payload 408: Video data captured by a camera device component or retrieved from a memory component of the client device 102 and included in the message 400. The video data for a sent or received message 400 can be stored in a video table 304.

[0101] ● Message audio payload 410: Audio data captured by a microphone or retrieved from a memory component of the client device 102 and included in the message 400.

[0102] ● Message enhancement data 412: Enhancement data (e.g., filters, stickers, or other annotations or enhancements) representing enhancements to be applied to the message image payload 406, the message video payload 408, or the message audio payload 410 of the message 400. The enhancement data for a sent or received message 400 can be stored in an enhancement table 310.

[0103] ● Message duration parameter 414: Indicates in seconds the content of the message (e.g., the message image payload 406, the message video payload 408, the message audio payload 410) will be transmitted via the message

[0104] The parameter value of the amount of time that the transceiver client 104 presents to or makes accessible to the user. ● Message geographic location parameter 416: Geographic location data associated with the content payload of the message (e.g., latitude coordinates and longitude coordinates). Multiple message geographic location parameter 416 values can be included in the payload, and each of these parameter values is associated with a content item included in the content (e.g., a specific image within the message image payload 406, or a specific video within the message video payload 408).

[0105] ● Message story identifier 418: An identifier value that identifies one or more content collections (e.g., the "stories" identified in the story table 314) associated with a specific content item in the message image payload 406 of the message 400. For example, the identifier value can be used to associate each of multiple images within the message image payload 406 with multiple content collections.

[0106] ● Message tag 420: Each message 400 can be tagged with multiple tags, and each of the multiple tags indicates the theme of the content included in the message payload. For example, in the case where a specific image included in the message image payload 406 depicts an animal (e.g., a lion), a tag value indicating the relevant animal can be included within the message tag 420. The tag value can be manually generated based on user input, or can be automatically generated using, for example, image recognition.

[0107] ● Message sender identifier 422: An identifier (e.g., a message transceiver system identifier, an email address, or a device identifier) that indicates the user of the client device 102 on which the message 400 is generated and from which the message 400 is sent.

[0108] · Message recipient identifier 424: An identifier (e.g., a message transceiver system identifier, an email address, or a device identifier) that indicates the user of the client device 102 to which the message 400 is addressed.

[0109] The content (e.g., value) of each component of the message 400 can be a pointer to a location in a table in which the content data value is stored. For example, the image value in the message image payload 406 can be a pointer to a location within the image table 312 (or the address of a location within the image table 312). Similarly, the value within the message video payload 408 can point to data stored within the video table 304, the value stored in the message enhancement data 412 can point to data stored in the enhancement table 310, the value stored in the message story identifier 418 can point to data stored in the story table 314, and the values stored in the message sender identifier 422 and the message recipient identifier 424 can point to user records stored in the entity table 306.

[0110] Body Styling System

[0111] Figure 5 is a block diagram showing an example body stylization system 224 according to some examples. The body stylization system 224 includes a set of components 510 that operate on a set of input data (e.g., training data 501), and the input data further includes one or more monocular images 502 depicting a person's full body. During the training phase, a set of input data (e.g., training data 501) is obtained from one or more databases ( Figure 3 ), and when using an AR / VR application (e.g., used by the messaging client 104), input data (e.g., one or more images 502) is obtained from the RGB camera device of the client device 102. The body stylization system 224 includes a training data generation module 512, a body stylization module 514, an AR effect module 519, an image modification module 518, an image display module 520, a 3D body tracking module 513, and a full body segmentation module 515. In some cases, the 3D body tracking module 513 performs 2D tracking rather than 3D tracking by extracting body key points in the image.

[0112] In some examples, the body stylization system 224 receives an image 502 depicting a person's full body in the real world. The body stylization system 224 applies a machine learning model to the image 502 to generate a stylized version of the person's full body in the real world. Training data can be used to train the machine learning model to establish a relationship between multiple training images (which depict a person's full body in synthetic rendering) and the corresponding ground truth stylized versions of the person's full body in a given style. The body stylization system 224 replaces the depiction of the person's full body in the real world in the image 502 with the generated stylized version of the person's full body in the real world. In some examples, the client device 102 and / or the messaging client 104 can implement different sets of machine learning models. Each of the machine learning models can have been trained to generate a corresponding stylized version of the person's full body based on a different respective target style. User input can be received to activate a given machine learning model among the machine learning models by selecting between various options, each of which is associated with a different style. In some examples, the image 502 is a frame of a video, and the generated stylized version is generated and applied to modify the frame and / or the video in real time.

[0113] Machine learning is a field of study that gives computers the ability to learn without being explicitly programmed. Machine learning explores the study and construction of algorithms (also referred to as tools in this article) that can learn from existing data and make predictions on new data. Such machine learning tools operate by building models based on example training data to make data-driven predictions or decisions, which are represented as outputs or evaluations. Although examples are given for several machine learning tools, the principles presented in this article can be applied to other machine learning tools.

[0114] In some examples, different machine learning tools can be used. For example, logistic regression (LR), naive Bayes, random forest (RF), neural network (NN), deep NN (DNN), matrix factorization, and support vector machine (SVM) tools can be used to classify or score videos.

[0115] Two common types of problems in machine learning are classification problems and regression problems. Classification problems (also known as categorization problems) aim to classify items into one of several categorical values (e.g., is the object an apple or an orange?). Regression algorithms aim to quantify some items (e.g., by providing values that are real numbers).

[0116] Machine learning algorithms use features to analyze data to generate evaluations. Each of the features is an individually measurable attribute of the observed phenomenon. The concept of a feature is related to the concept of an explanatory variable used in statistical techniques such as linear regression. Selecting features that provide useful information, are discriminative, and are independent is very important for the effective operation of MLP in pattern recognition, classification, and regression. Features can be of various types, such as numerical features, strings, and graphics.

[0117] In one example, by way of example only, features can be of different types and can include one or more of content, concepts, attributes, historical data, and / or user data. Machine learning algorithms utilize training data to find correlations between the identified features that affect the outcome or evaluation. In some examples, the training data includes labeled data, which is known data about one or more identifying features and one or more outcomes, such as detecting communication patterns, detecting the meaning of a message, generating a summary of a message, detecting action items in a message, detecting the urgency of a message, detecting the relationship between a user and a sender, calculating score attributes, calculating a message score, etc.

[0118] Using training data and identifying features, a machine learning program trains a machine learning tool. The machine learning tool evaluates the values of the features (when the features are relevant to the training data). The result of the training is a trained machine learning program. When performing an evaluation using the trained machine learning program, new data is provided as input to the trained machine learning program, and the trained machine learning program generates an evaluation as output.

[0119] A machine learning program supports two types of phases, namely, a training phase and a prediction phase. In the training phase, supervised learning, unsupervised learning, or reinforcement learning can be used. For example, the machine learning program (1) receives features (e.g., as structured data or labeled data in supervised learning) and / or (2) identifies features in the training data (e.g., unstructured data or unlabeled data for unsupervised learning). In the prediction phase, as an example of an evaluation, the machine learning program uses features for analyzing video frames to generate a stylized version of a person's full body or a prediction or result based on a target style.

[0120] In the training phase, feature engineering is used to identify features and can include identifying features that provide useful information, are discriminative, and are independent for the effective operation of the machine learning program in pattern recognition, classification, and regression. In some examples, the training data includes labeled data, which is known data of pre-identified features and one or more results. Each of the features can be a variable or an attribute, such as the various measurable properties of a process, an item, a system, or a phenomenon represented by a data set (e.g., the training data).

[0121] In the training phase, the machine learning program uses the training data to find the correlations between the features that affect the prediction result or evaluation. Using the training data and the identified features, the machine learning program is trained during the training phase of the machine learning program training. The machine learning program evaluates the values of the features (when the features are relevant to the training data). The result of the training is a trained machine learning program (e.g., a trained or learned model).

[0122] In addition, the training phase may involve machine learning in which the training data is structured (e.g., labeled during a preprocessing operation), and the trained machine learning program implements a relatively simple neural network that can, for example, perform classification and clustering operations. In other examples, the training phase may involve deep learning in which the training data is unstructured, and the trained machine learning program implements a DNN that can perform both feature extraction and classification / clustering operations.

[0123] A neural network generated during a training phase and implemented within a trained machine learning program can include a hierarchical (e.g., layered) organization of neurons. For example, neurons (or nodes) can be arranged hierarchically into a number of layers, including an input layer, an output layer, and multiple hidden layers. Each of the layers within the neural network can have one or many neurons, and each of these neurons operationally computes a small function (e.g., an activation function). For example, if the activation function generates a result that exceeds a particular threshold, the output can be transmitted from that neuron (e.g., the sending neuron) to a connected neuron (e.g., the receiving neuron) in a successive layer. The connections between neurons also have associated weights that define the influence of the input from the sending neuron to the receiving neuron. In some cases, these neurons implement one or more encoder or decoder networks.

[0124] In some examples, by way of example only, the neural network can also be one of several different types of neural networks, including a single-layer feedforward network, an artificial neural network (ANN), a GAN, a recurrent neural network (RNN), a symmetrically connected neural network and an unsupervised pre-trained network, a convolutional neural network (CNN), or a recursive neural network (RNN).

[0125] During a prediction phase, the trained machine learning program is used to perform an evaluation. In response to receiving video data, the video data is provided as input to the trained machine learning program, and the trained machine learning program generates an evaluation as output.

[0126] In some examples, the messaging client 104 can receive an input selecting a first AR experience associated with a first target style. In response, a first machine learning model within the machine learning model is accessed and used to generate a first stylized version of the full body of the person depicted in the image. The first stylized version can be used to modify the real-world object depicted in the input image or video to generate a modified image or video depicting a stylized version of the full body of the person.

[0127] In another example, the messaging client 104 receives an input selecting a second AR experience associated with a second target style. In response, a second machine learning model within the machine learning model is accessed and used to generate a second stylized version of the full body of the person depicted in the image. The second stylized version can be used to modify the real-world object depicted in the input image or video to generate a modified image or video depicting a stylized version of the full body of the person.

[0128] In some examples, the real-world object includes a person. In some examples, the body stylization system 224 receives an image depicting the full body of a real-world person and applies a machine learning model to the image to generate a stylized version of the full body of the real-world person corresponding to a given style. Training data can be used to train the machine learning model to establish a relationship between multiple training images (which depict the full body of a synthetically rendered person) and the corresponding ground truth stylized versions of the full body of a person of a given style. The body stylization system 224 replaces the depiction of the full body of the real-world person in the image with the generated stylized version of the full body of the real-world person.

[0129] In some examples, the full body of the real-world person includes a head, arms, a torso, and legs, and the stylized version of the full body of the real-world person includes stylized versions of the head, arms, torso, and legs. In some examples, the machine learning model includes a deep neural network.

[0130] In some examples, the body stylization system 224 receives an input for selecting a given style from a plurality of styles. The body stylization system 224 selects a machine learning model from a plurality of machine learning models, each of which is configured to generate a different stylized version of the full body of a person corresponding to a respective style among the plurality of styles. In some examples, the plurality of styles includes at least one of a zombie style, a bodybuilder style, a cartoon style, anime, Gollum, Neanderthal, and / or a Barbie style.

[0131] In some examples, the body stylization system 224 generates training data by performing a training operation. The body stylization system 224 accesses a first set of latent codes through a first full-body GAN and a second full-body GAN. The body stylization system 224 renders a first synthetic full body of a person corresponding to the first set of latent codes through the first full-body GAN and renders a second synthetic full body of a person corresponding to the first set of latent codes through the second full-body GAN. The body stylization system 224 calculates a direction loss based on the second synthetic full body of the person through a direction loss model associated with the given style. The body stylization system 224 updates one or more weights of the second GAN based on the direction loss and repeats the following operations until a stop criterion is reached: rendering the second synthetic full body of the person; calculating the direction loss; and updating one or more weights.

[0132] In some examples, the body stylization system 224 determines that a stopping criterion has been reached and, in response, repeats the operations on the second set of latent codes. The body stylization system 224 generates an image pair of training data by applying a new latent code to a first full-body GAN and a second full-body GAN to generate a first training image among a plurality of training images (which depict a full body of a synthetically rendered person) and a first ground truth stylized version of the full body of the person. In some examples, the body stylization system 224 trains a machine learning model based on the paired images of the training data.

[0133] In some examples, the body stylization system 224 accesses a first set of latent codes through a first face-based GAN and a second face-based GAN, and renders a first synthetic face of a person corresponding to the first set of latent codes through the first face-based GAN. The body stylization system 224 renders a second synthetic face of the person corresponding to the first set of latent codes through the second face-based GAN. The body stylization system 224 calculates a direction loss based on the second synthetic face of the person through a direction loss model associated with a given style, and updates one or more weights of the second GAN based on the direction loss. The body stylization system 224 repeats the following until the stopping criterion is reached again: rendering the second synthetic face of the person; calculating the direction loss; and updating one or more weights.

[0134] In some examples, the body stylization system 224 applies a new latent code to the first face-based GAN and the second face-based GAN, and generates a second pair of images of training data based on the outputs of the first face-based GAN and the second face-based GAN. In some examples, the body stylization system 224 superimposes a synthetic face of a person generated by the second face-based GAN on a second synthetic full body of the person in the paired image, where the second synthetic full body of the person is generated by the second full-body GAN. The body stylization system 224 smoothly fuses the second synthetic face of the person generated by the second face-based GAN with the second synthetic full body of the person generated by the second full-body GAN in the paired image. The body stylization system 224 stores the second synthetic face of the person generated by the second face-based GAN that has been smoothly fused with the second synthetic full body of the person generated by the second full-body GAN in the paired image as the second synthetic full body of the person in the paired image.

[0135] In some examples, the body stylization system 224 selects the part of the person's second synthetic full body that corresponds to the body part. The body stylization system 224 identifies a set of weights of the second GAN that corresponds to the body part, and calculates a directional loss based on the selected part of the person's second synthetic full body that corresponds to the body part through a directional loss model associated with a given style. The body stylization system 224 updates the set of weights corresponding to the body part based on the directional loss, without updating other weights associated with other body parts.

[0136] In some examples, the body stylization system 224 divides the person's second synthetic full body into separate parts corresponding to different body parts. The body stylization system 224 calculates the directional loss separately for each of the individual parts.

[0137] In some examples, the body stylization system 224 trains a machine learning model by performing training operations that include: accessing training data. The body stylization system 224 applies the machine learning model to a first training data set that includes a first training image among a plurality of training images depicting a synthetically rendered full body of a person to generate an estimated stylized version of the full body of the person depicted in the first training image. The body stylization system 224 calculates the deviation between the estimated stylized version of the full body of the person depicted in the first training image and the ground truth stylized version of the full body of the person depicted in the first training image. The body stylization system 224 updates one or more parameters of the machine learning model based on the deviation between the estimated stylized version of the full body of the person depicted in the first training image and the ground truth stylized version of the full body of the person depicted in the first training image.

[0138] In some examples, the body stylization system 224 generates full body key points for the full body of the person depicted in the first training image. The estimated stylized version of the full body of the person depicted in the first training image can be generated based on the full body key points. In some examples, an image is received as a frame of a video depicting a real-world person. In this case, the depiction of the full body of the real-world person in the video is replaced with a stylized version in real time.

[0139] The training data generation module 512 is configured to generate paired images that depict a synthetic (fake or computer-generated) version of a person's full body and a stylized version of the synthetic version of the person's full body. Specifically, the training data generation module 512 includes a first GAN that is configured to receive a latent code or vector as input and render a first image depicting the synthetic person's full body. The training data generation module 512 includes a second GAN that is configured to receive the same latent code or vector as input and render a second image depicting a stylized version of the synthetic person's full body according to a specific style. The first image and the second image together form a training data pair that is provided to the body stylization module 514 to train a machine learning model to estimate a stylized version of a person according to the received real-time image or video, where the stylized version depicts the person's full body according to a specific style.

[0140] For example, the body stylization module 514 receives the first image of the training data pair and estimates or generates an estimated image that includes an estimated stylized version of the person depicted in the first image. Then, the body stylization module 514 compares the estimated image with the second image to calculate a deviation. The stopping criterion is compared or analyzed with respect to the deviation. If the stopping criterion is not met, one or more parameters of the body stylization module 514 are updated, and the training data generation module 512 uses another latent code or vector to generate another pair of images, and the body stylization module 514 is retrained using the another pair of images. Once the stopping criterion is met, the body stylization module 514 is stored as a trained machine learning model. The trained body stylization module 514 is applied to a new image depicting the person's full body and generates a stylized version of the person's full body. The AR effect module 519 can modify the new image to replace the person's full body with the stylized version of the person's full body.

[0141] In some cases, to improve the quality of the output of the body stylization module 514 and improve convergence, a 2D body tracking or 3D body tracking module 513 can be used to identify 2D or 3D full body key points of the person's full body depicted in the training images and / or the images received from the client device 102. Then, the body stylization module 514 can generate an estimated stylized version of the person's full body based on the full body key points.

[0142] In some examples, the first GAN and the second GAN of the training data generation module 512 are trained separately and are trained before the body stylization module 514. Figure 6FIG. 600 shows a set of components of the training data generation module 512 according to some examples. Specifically, the components 600 include: a latent code 610, which can be generated by a latent code generator; a first full-body GAN 620; a second full-body GAN 622; and an orientation loss network 640, which is associated with a specific style 644.

[0143] The components 600 are configured to train the second full-body GAN 622 to generate a stylized version of an image 632 of a person's full body, which is generated by the first full-body GAN 620. Initially, the first full-body GAN 620 and the second full-body GAN 622 are configured with the same set of weights 624. The first full-body GAN 620 receives the latent code 610 from a latent code generator, which can be a random number generator. The first full-body GAN 620 generates a first image that depicts the normal non-stylized full body of a person 630. Since the first full-body GAN 620 and the second full-body GAN 622 are initialized with the same set of weights 624, in parallel, the second full-body GAN 622 also generates a second image that depicts the normal non-stylized full body of a person as its output image 632.

[0144] The images generated by the first full-body GAN 620 and the second full-body GAN 622 are provided to the orientation loss network 640. The orientation loss network 640 receives a style 644, which represents the parameters of a target style for a specific type of object (e.g., the full body of a human 642). The orientation loss network 640 is configured to generate a loss indicating how far the output image 632 is from looking like the target style based on the input text prompt. In some cases, the orientation loss network 640 implements an image language model. In some examples, the orientation loss network 640 calculates the loss based only on the text prompt or based on the text prompt and one or more example images of the target style.

[0145] Then, the loss is used to update one or more parameters of the second GAN 622, such as weights. At this stage, the parameters of the first full-body GAN 620 are not updated. Next, the second full-body GAN 622 is applied again to the same latent code 610 to generate a new image as the output image 632, which represents a stylized version of the person 630 depicted in the first image, and the first image is generated by the first full-body GAN 620. This image is again applied to the directional loss network 640 to recalculate the loss representing how far the output image 632 is from looking like the target style, and based on this loss, the parameters of the second full-body GAN 622 are updated again. This process continues until a stopping criterion based on the loss calculation of the directional loss network 640 reaches a threshold condition. In some examples, once the stopping criterion is met, the first full-body GAN 620 and the second full-body GAN 622 are used as part of the training data generation module 512 to generate new image pairs for training the body stylization module 514.

[0146] In some examples, the first full-body GAN 620 and the second full-body GAN 622 are trained in one or more iterations. At each iteration, the body stylization system 224 determines which layers of the second full-body GAN 622 are to be optimized. To this end, the body stylization system 224 samples a random latent vector / code from a normal distribution (e.g., of size 512). The latent code / vector is passed through a mapping fully connected network to generate a resulting vector (w) for modulating the layers of the first full-body GAN 620 and the second full-body GAN 622. The final result is two images output by the first full-body GAN 620 and the second full-body GAN 622. In some cases, the images are partitioned into body parts (e.g., head, torso, legs, etc.). Random selections and croppings are made of the same body parts from the two images output by the first full-body GAN 620 and the second full-body GAN 622. For example, as Figure 7 shown in a set of images 700, a random body part 710 is cropped from the image of the person 630 provided or generated by the first full-body GAN 620, and the same random body part 720 is cropped from the stylized version of the body 722 depicted in the image 632 provided or generated by the second full-body GAN 622. Then the directional loss network 640 calculates a loss based on the cropped portions of the images and the input text prompt pair (e.g., "human" and "zombie"), such as using a directional clip loss.

[0147] The loss is calculated by the orientation loss network 640 and backpropagated through the first full-body GAN 620 and the second full-body GAN 622 to update the w vector and obtain the w_prime vector. The absolute difference between the w vector and the w_prime vector can be calculated to select the k largest (most changed) elements among these vectors. The indices of these elements can correspond to the layer indices of the second full-body GAN 622. All layers in the second full-body GAN 622 except the layers corresponding to the k elements are frozen. In some examples, if the style has no pose change, the first six layers of the second full-body GAN 622 are frozen.

[0148] After selecting the layers of the second full-body GAN 622 to be trained, a random latent vector / code from a normal distribution (e.g., of size 512) is sampled again. This vector is passed through the mapping network to generate a result vector (w), which is then provided to the first full-body GAN 620 and the second full-body GAN 622 to generate corresponding images of a synthetic full body or entire body of a person. The full bodies depicted in the two images generated by the first full-body GAN 620 and the second full-body GAN 622 are divided into three or more parts (e.g., head, torso, and legs), and randomly selected body parts are cropped from the two images. A random viewpoint transformation can be applied to the cropped parts to simulate different viewpoints and avoid overfitting of the clip model. The cropped parts from the two images, along with their text prompts, are provided to the orientation loss network 640 to calculate the orientation loss. This can be done by: estimating the embeddings of each image and its text prompt (a total of four embeddings); calculating the displacement vectors (image2_embedding – image1_embedding and text prompt2_embedding – text prompt1_embedding); and finally obtaining the cosine similarity between these two result vectors. The loss is backpropagated through the first full-body GAN 620 and / or the second full-body GAN 622 to perform one step along the gradient direction and update the weights of the second full-body GAN 622.

[0149] In some examples, the training data generation module 512 includes a first face GAN and a second face GAN. Specifically, in addition to the first full-body GAN 620 and / or the second full-body GAN 622, the training data generation module 512 further includes first and second face GANs (not shown) that are only for faces and perform similar functions to the first full-body GAN 620 and / or the second full-body GAN 622. In this case, the training data generation module 512 includes: the first full-body GAN 620 and the second full-body GAN 622, which are configured to generate images depicting a synthetic full body of a person and a stylized version of the body; and the first face GAN and the second face GAN, which are configured to generate images depicting a synthetic face of a person and a stylized version of the face.

[0150] In this case, the face or face-only GAN is trained by performing a set of training operations. Specifically, the first face GAN is trained to receive a latent code / vector and generate a synthetic face-only image of a person. In parallel, the second face GAN is trained to receive the same latent code / vector and generate a stylized version of the face-only of a person. The output image of the second face GAN is provided to the orientation loss network 640 to calculate an orientation loss based on the target style. The orientation loss is fed back to the second face GAN to update one or more parameters of the face GAN. Once the stopping criterion is reached, the second face GAN is trained to generate a new image of a synthetic face stylized according to the target style, and the first face GAN generates a synthetic face without the target style.

[0151] In some examples, the training data generation module 512 includes four GANs (two full-body or whole-body GANs, such as the first full-body GAN 620 and the second full-body GAN 622, and two face-only GANs). Given a latent code / vector, a synthetic dataset of paired full-body images is generated using the first full-body GAN 620 and the second full-body GAN 622 (as described above). Given a pair of original full-body or whole-body images and stylized full-body or whole-body images (I1 and I2) generated by the first full-body GAN 620 and the second full-body GAN 622, the faces are cropped from the pair of images to provide cropped faces F1 and F2. These cropped faces are projected onto the latent space of the trained face-only GAN, which is configured to generate a stylized version of the synthetic face (e.g., using a StyleGAN encoder). After this projection is completed, a resulting set of latent vector pairs is provided. These two vectors are forward passed through the first face-only GAN and the second face-only GAN to produce high-resolution face images of F1 and F2, which can be referred to as F1_prime and F2_prime.

[0152] F1_prime and F2_prime are smoothly blended back onto the original full-body or whole-body image and the stylized full-body or whole-body image (I1 and I2). For this purpose, an optimization algorithm is used, which searches for the facial image that best matches the face in the full-body image. The optimization can start from a pair of latent vectors in the first iteration and search for a new latent vector that, given a pair of faces, does not result in visible seams or artifacts when fused with I1 and I2. In some cases, this can be done by formulating a loss for: 1) the faces in I1 and I2 and the background of the faces F1_prime, F2_prime, 2) the boundaries of F1, F2 (i.e., the locations where stitching occurs on I1 and I2). This optimization can occur for each pair in the paired full-body synthetic dataset. Finally, a new synthetic paired dataset of the original full-body human image and the stylized full-body human image is provided, where the faces have high quality and resolution (e.g., 1024x1024).

[0153] In an example, the AR effect module 519 selects one or more AR elements or graphics based on the segmentation mask associated with the stylized full-body of a person estimated by the training data generation module 512 and applies them to the object depicted in the image or video (e.g., the deformed object). These AR graphics combined with the real-world object depicted in the image or video are provided to the image modification module 518 to render an image or video depicting a stylized version of the person wearing the AR object (e.g., an AR wallet or earrings).

[0154] The image modification module 518 adjusts the image captured by the camera device based on the AR effects selected by the AR effect module 519. The image modification module 518 adjusts the way the AR elements placed on the stylized version of the person depicted in the image or video are presented in the image or video. The image display module 520 combines the adjustments made by the image modification module 518 into the received monocular image or video depicting the user's body. This image or video is provided by the image display module 520 to the client device 102 and can then be sent to another user or stored for later access and display.

[0155] In some examples, the image modification module 518 receives 2D or 3D body tracking information representing the 3D position of the user depicted in the image from the 3D body tracking module 513. The 3D body tracking module 513 generates the 2D or 3D body tracking information by processing the training data 501 using additional machine learning techniques. The image modification module 518 may also receive a full body segmentation from another machine learning technique, which represents which pixels in the image correspond to the user's full body. The full body segmentation may be received from the full body segmentation module 515. The full body segmentation module 515 generates the full body segmentation by processing the training data 501 using machine learning techniques.

[0156] Figure 8 is a diagrammatic representation of an output 800 of a body stylization system 224 according to some examples. Specifically, an input image 810 may be received by the body stylization system 224. The body stylization system 224 processes the input image 810 through a first trained machine learning model implemented by the body stylization module 514 to generate a first stylized version 812 of the body depicted in the input image 810 corresponding to the first style. In some cases, the input image 810 is processed through a second trained machine learning model implemented by the body stylization module 514 to generate a second stylized version 814 of the body depicted in the input image 810 corresponding to the second style.

[0157] In some examples, a second input image 820 may be received by the body stylization system 224. The body stylization system 224 processes the input image 820 through a first trained machine learning model implemented by the body stylization module 514 to generate a third stylized version 822 of the body depicted in the input image 820 corresponding to the first style. In some cases, the input image 820 is processed through a second trained machine learning model implemented by the body stylization module 514 to generate a fourth stylized version 824 of the body depicted in the input image 820 corresponding to the second style.

[0158] Figure 9 is a flowchart of a process 900 performed by a body stylization system 224 according to some examples. Although the flowchart may describe the operations as a sequential process, many of the operations may be performed in parallel or simultaneously. In addition, the order of the operations may be rearranged. The process terminates when its operations are completed. The process may correspond to a method, a program, etc. The steps of the method may be performed in whole or in part, may be performed in combination with some or all of the steps in other methods, and may be performed by any number of different systems or any part thereof (e.g., a processor in any system included in the system).

[0159] At operation 901, the body stylization system 224 (e.g., the client device 102 or the server) receives an image including a depiction of a full body of a real-world person, as discussed above.

[0160] At operation 902, the body stylization system 224 applies a machine learning model to the image to generate a stylized version of the full body of the real-world person corresponding to a given style, using training data to train the machine learning model to establish a relationship between multiple training images (which depict the full bodies of synthetically rendered people) and the corresponding ground-truth stylized versions of the full bodies of people of the given style (as discussed above).

[0161] At operation 903, the body stylization system 224 replaces the depiction of the full body of the real-world person in the image with the generated stylized version of the full body of the real-world person (as discussed above).

[0162] Machine architecture

[0163] Figure 10is an illustrative representation of a machine 1000 within which instructions 1008 (e.g., software, program, application, applet, app, or other executable code) can be executed to cause the machine 1000 to perform any one or more of the methods discussed herein. For example, the instructions 1008 can cause the machine 1000 to perform any one or more of the methods described herein. The instructions 1008 transform the general unprogrammed machine 1000 into a particular machine 1000 programmed to perform the described and illustrated functions in the described manner. The machine 1000 can operate as a stand-alone device or can be coupled (e.g., networked) to other machines. In a networked deployment, the machine 1000 can operate in a server-client network environment as a server machine or a client machine, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 1000 can include, but is not limited to: server computers, client computers, personal computers (PCs), tablet computers, laptop computers, netbooks, set-top boxes (STBs), personal digital assistants (PDAs), entertainment media systems, cellular telephones, smart phones, mobile devices, wearable devices (e.g., smart watches), smart home devices (e.g., smart appliances), other smart devices, web appliances, network routers, network switches, network bridges, or any machine capable of executing the instructions 1008 to perform the actions specified to be taken by the machine 1000, sequentially or otherwise. Further, while only a single machine 1000 is shown, the term "machine" shall also be taken to include a collection of machines that individually or jointly execute the instructions 1008 to perform any one or more of the methods discussed herein. For example, the machine 1000 can include a client device 102 or any one of a number of server devices forming part of a messaging server system 108. In some examples, the machine 1000 can also include both a client system and a server system, where certain operations of a particular method or algorithm are executed on the server side and certain operations of the particular method or algorithm are executed on the client side.

[0164] Machine 1000 may include a processor 1002, a memory 1004, and input / output (I / O) components 1038 that may be configured to communicate with each other via a bus 1040. In an example, the processor 1002 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, a processor 1006 and a processor 1010 that execute instructions 1008. The term "processor" is intended to include multi-core processors, which may include two or more independent processors (sometimes referred to as "cores") that may execute instructions simultaneously. Although Figure 10 multiple processors 1002 are shown, machine 1000 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.

[0165] Memory 1004 includes a main memory 1012, a static memory 1014, and a storage unit 1016, all of which may be accessed by the processor 1002 via the bus 1040. The main memory 1004, the static memory 1014, and the storage unit 1016 store instructions 1008 that embody any one or more of the methods or functions described herein. The instructions 1008 may also reside, completely or partially, within the main memory 1012, within the static memory 1014, within a machine-readable medium 1018 within the storage unit 1016, within at least one of the processors 1002 (e.g., within a cache memory of the processor), or within any suitable combination thereof during execution by the machine 1000.

[0166] The I / O components 1038 may include various components for receiving input, providing output, generating output, sending information, exchanging information, capturing measurement results, etc. The specific I / O components 1038 included in a particular machine will depend on the type of the machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, while a headless server machine will likely not include such a touch input device. It should be appreciated that the I / O components 1038 may include Figure 10Many other components not shown. In various examples, the I / O component 1038 may include a user output component 1024 and a user input component 1026. The user output component 1024 may include visual components (e.g., a display such as a plasma display panel (PDP), a light-emitting diode (LED) display, a liquid crystal display (LCD), a projector, or a cathode ray tube (CRT)), acoustic components (e.g., speakers), haptic components (e.g., a vibration motor, a resistance mechanism), other signal generators, etc. The user input component 1026 may include an alphanumeric input component (e.g., a keyboard, a touch screen configured to receive alphanumeric input, an optoelectronic keyboard, or other alphanumeric input components), a point-based input component (e.g., a mouse, a touchpad, a trackball, a joystick, a motion sensor, or other pointing instruments), a haptic input component (e.g., a physical button, a touch screen that provides the position and force of a touch or touch gesture, or other haptic input components), an audio input component (e.g., a microphone), etc.

[0167] In other examples, the I / O component 1038 may include a biometric component 1028, a motion component 1030, an environmental component 1032, or a positioning component 1034 and various other components. For example, the biometric component 1028 includes components for detecting expressions (e.g., hand expressions, facial expressions, voice expressions, body postures, or eye tracking), measuring biometric signals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), identifying a person (e.g., voice identification, retina identification, facial identification, fingerprint identification, or electroencephalogram-based identification), etc. The motion component 1030 includes an acceleration sensor component (e.g., an accelerometer), a gravity sensor component, a rotational sensor component (e.g., a gyroscope).

[0168] The environmental component 1032 includes, for example, one or more camera devices (with still image / photo and video capabilities), a lighting sensor component (e.g., a photometer), a temperature sensor component (e.g., one or more thermometers that detect the ambient temperature), a humidity sensor component, a pressure sensor component (e.g., a barometer), an acoustic sensor component (e.g., one or more microphones that detect background noise), a proximity sensor component (e.g., an infrared sensor that detects nearby objects), a gas sensor (e.g., a gas detection sensor that detects the concentration of a hazardous gas for safety or measures pollutants in the atmosphere), or other components that can provide an indication, measurement, or signal corresponding to the surrounding physical environment.

[0169] Regarding the imaging device, the client device 102 may have an imaging device system that includes, for example, a front imaging device on the front surface of the client device 102 and a rear imaging device on the rear surface of the client device 102. The front imaging device may be used, for example, to capture still images and videos of the user of the client device 102 (e.g., "selfies"), which can then be enhanced with the above-described enhancement data (e.g., filters). For example, the rear imaging device may be used to capture still images and videos in a more conventional imaging device mode, where these images are similarly enhanced with enhancement data. In addition to the front imaging device and the rear imaging device, the client device 102 may also include a 360° imaging device for capturing 360° photos and videos.

[0170] Furthermore, the imaging device system of the client device 102 may include a dual rear imaging device (e.g., a main imaging device and a depth sensing imaging device), or even a triple, quadruple, or quintuple rear imaging device configuration on the front and rear sides of the client device 102. For example, these multiple imaging device systems may include a wide-angle imaging device, an ultra-wide-angle imaging device, a telephoto imaging device, a macro imaging device, and a depth sensor.

[0171] The positioning component 1034 includes a position sensor component (e.g., a GPS receiver component), an altitude sensor component (e.g., an altimeter or barometer that detects air pressure, from which altitude can be obtained), an orientation sensor component (e.g., a magnetometer), and the like.

[0172] A variety of techniques can be used to implement communication. The I / O component 1038 also includes a communication component 1036 that is operable to couple the machine 1000 to the network 1020 or the device 1022 via a corresponding coupling or connection. For example, the communication component 1036 may include a network interface component or other suitable device that interfaces with the network 1020. In another example, the communication component 1036 may include a wired communication component, a wireless communication component, a cellular communication component, a near field communication (NFC) component, a Bluetooth component (e.g., Bluetooth low energy), a Wi-Fi component, and other communication components that provide communication via other modalities. The device 1022 may be another machine or any peripheral device among various peripheral devices (e.g., a peripheral device coupled via USB).

[0173] In addition, the communication component 1036 can detect an identifier or include components capable of operating to detect an identifier. For example, the communication component 1036 can include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional barcodes such as Universal Product Code (UPC) barcodes, multi-dimensional barcodes such as Quick Response (QR) codes, Aztec codes, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D barcodes, and other optical codes), or an acoustic detection component (e.g., a microphone for identifying a tagged audio signal). In addition, various information can be derived via the communication component 1036, such as a location obtained via Internet Protocol (IP) geolocation, a location obtained via Wi-Fi signal triangulation, a location obtained via detecting an NFC beacon signal that can indicate a specific location, etc.

[0174] Various memories (e.g., main memory 1012, static memory 1014, and the memory of the processor 1002) and the storage unit 1016 can store one or more sets of instructions and data structures (e.g., software) implemented or used by any one or more of the methods or functions described herein. When executed by the processor 1002, these instructions (e.g., instruction 1008) cause the various operations to implement the disclosed examples.

[0175] The instruction 1008 can be sent or received over the network 1020 using a transmission medium via a network interface device (e.g., the network interface component included in the communication component 1036) and using any one of several well-known transmission protocols (e.g., Hypertext Transfer Protocol (HTTP)). Similarly, the instruction 1008 can be sent or received using a transmission medium via a coupling (e.g., a peer-to-peer coupling) to the device 1022.

[0176] Software architecture

[0177] Figure 11FIG. 1100 is a block diagram showing a software architecture 1104 that may be installed on any one or more of the devices described herein. The software architecture 1104 is supported by hardware, such as a machine 1102 that includes a processor 1120, a memory 1126, and I / O components 1138. In this example, the software architecture 1104 may be conceptualized as a stack of layers, where each layer provides a specific function. The software architecture 1104 includes the following layers, such as an operating system 1112, libraries 1110, frameworks 1108, and applications 1106. In operation, the application 1106 activates API calls 1150 through the software stack and receives messages 1152 in response to the API calls 1150.

[0178] The operating system 1112 manages hardware resources and provides common services. The operating system 1112 includes, for example: a kernel 1114, services 1116, and drivers 1122. The kernel 1114 serves as an abstraction layer between the hardware and other software layers. For example, the kernel 1114 provides functions such as memory management, processor management (e.g., scheduling), component management, networking, and security settings. The services 1116 may provide other common services for other software layers. The drivers 1122 are responsible for controlling or interfacing with the underlying hardware. For example, the drivers 1122 may include a display driver, a camera device driver, a Bluetooth or Bluetooth Low Energy driver, a flash drive, a serial communication driver (e.g., a USB driver), a WI-FI driver, an audio driver, a power management driver, etc.

[0179] The libraries 1110 provide common low-level infrastructure used by the applications 1106. The libraries 1110 may include system libraries 1118 (e.g., C standard libraries) that provide functions such as memory allocation functions, string manipulation functions, mathematical functions, etc. In addition, the libraries 1110 may include API libraries 1124, such as media libraries (e.g., libraries for supporting the presentation and manipulation of various media formats, such as Moving Picture Experts Group-4 (MPEG4), Advanced Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Photographic Experts Group (JPEG or JPG), or Portable Network Graphics (PNG)), graphics libraries (e.g., the OpenGL framework for presenting graphical content in 2D and 3D on a display), database libraries (e.g., SQLite that provides various relational database functions), web libraries (e.g., WebKit that provides web browsing functions), etc. The libraries 1110 may also include various other libraries 1128 to provide many other APIs to the applications 1106.

[0180] The framework 1108 provides a common high-level infrastructure for use by the applications 1106. For example, the framework 1108 provides various graphical user interface functions, high-level resource management, and high-level location services. The framework 1108 may provide a wide range of other APIs that can be used by the applications 1106, some of which may be specific to a particular operating system or platform.

[0181] In an example, the applications 1106 may include a home application 1136, a contacts application 1130, a browser application 1132, a book reader application 1134, a location application 1142, a media application 1144, a messaging application 1146, a gaming application 1148, and various other applications such as an external application 1140. An application 1106 is a program that executes functions defined in a program. One or more of the applications 1106 may be created using various programming languages, such as object-oriented programming languages (e.g., Objective-C, Java, or C++) or procedural programming languages (e.g., C or assembly language), and may be structured in various ways. In a particular example, an external application 1140 (e.g., an application developed using an ANDROID or IOS SDK by an entity other than the vendor of a particular platform) may be mobile software running on a mobile operating system such as IOS, ANDROID, WINDOWS Phone, or another mobile operating system. In this example, the external application 1140 may activate an API call 1150 provided by the operating system 1112 to facilitate the functions described herein.

[0182] Glossary

[0183] A "carrier signal" refers to any non-tangible medium capable of storing, encoding, or carrying instructions for execution by a machine, and includes digital or analog communication signals or other non-tangible media to facilitate the transmission of such instructions. Instructions may be sent or received over a network via a network interface device using a transmission medium.

[0184] A "client device" refers to any machine that interfaces with a communication network to obtain resources from one or more server systems or other client devices. A client device may be, but is not limited to, a mobile phone, a desktop computer, a laptop computer, a portable digital assistant (PDA), a smartphone, a tablet computer, a superbook, a netbook, a laptop computer, a multiprocessor system, a microprocessor-based or programmable consumer electronics product, a gaming console, a set-top box, or any other communication device that a user may use to access a network.

[0185] "Communication network" refers to one or more portions of a network, which can be an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless WAN (WWAN), a metropolitan area network (MAN), the Internet, a portion of the Internet, a portion of the public switched telephone network (PSTN), a plain old telephone service (POTS) network, a cellular telephone network, a wireless network, a Wi-Fi network, other types of networks, or a combination of two or more such networks. For example, the network or a portion of the network can include a wireless network or a cellular network, and the coupling can be a code division multiple access (CDMA) connection, a global system for mobile communications (GSM) connection, or other types of cellular or wireless couplings. In this example, the coupling can implement any of a variety of data transfer technologies, such as single-carrier radio transmission technology (1xRTT), evolved data optimized (EVDO) technology, general packet radio service (GPRS) technology, GSM enhanced data rates for GSM evolution (EDGE) technology, the 3rd Generation Partnership Project (3GPP) including 3G, fourth-generation wireless (4G) networks, universal mobile telecommunications system (UMTS), high-speed packet access (HSPA), worldwide interoperability for microwave access (WiMAX), long-term evolution (LTE) standards, other data transfer technologies defined by various standards-setting organizations, other long-distance protocols, or other data transfer technologies.

[0186] "Component" refers to a device, physical entity, or logic having a boundary defined by a function or subroutine call, a branch point, an API, or other technology that provides partitioning or modularization for a particular processing or control function. Components can interface with other components via their interfaces to perform machine processes. A component can be an encapsulated functional hardware unit designed to be used with other components and is typically part of a program that performs a particular function among related functions.

[0187] Components can constitute software components (e.g., code implemented on a machine-readable medium) or hardware components. A "hardware component" is a tangible unit capable of performing certain operations and can be configured or arranged in some physical manner. In various examples, one or more computer systems (e.g., stand-alone computer systems, client computer systems, or server computer systems) or one or more hardware components of a computer system (e.g., a processor or a group of processors) can be configured by software (e.g., an application or a portion of an application) to operate as a hardware component that performs certain operations described herein.

[0188] The hardware components can also be implemented mechanically, electronically, or any suitable combination thereof. For example, the hardware components can include dedicated circuits or logic that are permanently configured to perform certain operations. The hardware components can be a dedicated processor, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC). The hardware components can also include programmable logic or circuits that are temporarily configured by software to perform certain operations. For example, the hardware components can include software executed by a general purpose processor or other programmable processor. Once configured by such software, the hardware components become a particular machine (or a particular component of a machine) that is uniquely customized to perform the configured functions and is no longer a general purpose processor. It will be appreciated that a decision can be made, for cost and time considerations, as to whether to implement the hardware components mechanically in dedicated and permanently configured circuits or in temporarily configured (e.g., software configured) circuits. Thus, the phrase "hardware component" (or "hardware-implemented component") should be understood to include a tangible entity, i.e., an entity that is physically constructed, permanently configured (e.g., hardwired) or temporarily configured (e.g., programmed) to operate in some manner or to perform certain operations described herein.

[0189] Consider an example where the hardware components are temporarily configured (e.g., programmed). It is not necessary to configure or instantiate each of the hardware components at any given time. For example, in the case where the hardware components include a general purpose processor that is configured by software to become a dedicated processor, the general purpose processor can be configured to different dedicated processors (e.g., including different hardware components) at different times. The software accordingly configures one or more particular processors to, for example, constitute a particular hardware component at one time and different hardware components at different times.

[0190] The hardware components can provide information to other hardware components and receive information from other hardware components. Thus, the described hardware components can be considered to be communicatively coupled. In the case where there are multiple hardware components present simultaneously, communication can be achieved through signal transmission (e.g., via appropriate circuits and buses) between or among two or more of the hardware components. In an example where multiple hardware components are configured or instantiated at different times, communication between such hardware components can be achieved, for example, by storing information in a memory structure to which the multiple hardware components have access and retrieving the information from the memory structure. For example, one hardware component can perform an operation and store the output of the operation in a memory device communicatively coupled thereto. Then, other hardware components can access the memory device at a subsequent time to retrieve the stored output and process it. The hardware components can also initiate communication with an input device or an output device and can operate on resources (e.g., a collection of information).

[0191] The various operations of the example methods described herein can be performed, at least in part, by one or more processors temporarily configured (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors may constitute processor-implemented components that operate to perform one or more of the operations or functions described herein. As used herein, a "processor-implemented component" refers to a hardware component implemented using one or more processors. Similarly, the methods described herein can be at least in part processor-implemented, where a particular one or more processors are examples of hardware. For example, at least some of the operations of the method can be performed by one or more processors 1002 or processor-implemented components. Additionally, one or more processors can operate to support the performance of relevant operations in a "cloud computing" environment or as "software as a service" (SaaS) operations. For example, at least some of the operations can be performed by a group of computers (as examples of machines including processors), where the operations can be accessed via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., APIs). The performance of certain operations can be distributed among processors, not residing only within a single machine but being deployed across a number of machines. In some examples, the processor or processor-implemented components can be located in a single geographical location (e.g., within a home environment, an office environment, or a server farm). In other examples, the processor or processor-implemented components can be distributed across a number of geographical locations.

[0192] "Computer-readable storage medium" refers to both machine storage media and transmission media. Thus, the term includes both storage devices / media and carrier / modulated data signals. The terms "machine-readable medium", "computer-readable medium", and "device-readable medium" mean the same thing and can be used interchangeably in this disclosure.

[0193] "Ephemeral message" refers to a message that is accessible for a limited duration. An ephemeral message can be text, an image, a video, etc. The access time of an ephemeral message can be set by the message sender. Alternatively, the access time can be a default setting or a setting specified by the recipient. Regardless of the setting technique, the message is transient.

[0194] "Machine storage medium" means a single or multiple storage devices and media that store executable instructions, routines, and data (e.g., centralized or distributed databases, and associated caches and servers). Accordingly, the term should be regarded as including, but not limited to, solid-state memory as well as optical and magnetic media, including memory internal or external to a processor. Specific examples of machine storage media, computer storage media, and device storage media include: non-volatile memory, including, for example, semiconductor memory devices such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), FPGA, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The terms "machine storage medium", "device storage medium", and "computer storage medium" mean the same thing and may be used interchangeably in this disclosure. The terms "machine storage medium", "computer storage medium", and "device storage medium" expressly exclude carrier waves, modulated data signals, and other such media, at least some of which are covered by the term "signal medium".

[0195] "Non-transitory computer-readable storage medium" means a tangible medium capable of storing, encoding, or carrying instructions executable by a machine.

[0196] "Signal medium" means any intangible medium capable of storing, encoding, or carrying instructions executable by a machine, and includes digital or analog communication signals or other intangible media that facilitate the communication of software or data. The term "signal medium" should be regarded as including any form of modulated data signal, carrier wave, etc. The term "modulated data signal" refers to a signal in which one or more of its characteristics are set or changed in such a way as to encode information in the signal. The terms "transmission medium" and "signal medium" mean the same thing and may be used interchangeably in this disclosure.

[0197] Without departing from the scope of this disclosure, changes and modifications may be made to the disclosed examples. These and other changes or modifications are intended to be included within the scope of this disclosure as expressed in the appended claims.

Claims

1. A method, comprising: receiving, by one or more processors, an image comprising a depiction of the full body of a real-world person; applying, by the one or more processors, a machine learning model to the image to generate a stylized version of the full body of the real-world person corresponding to a given style, the machine learning model being trained using training data to establish a relationship between a plurality of training images depicting the full body of a synthetically rendered person and corresponding ground-truth stylized versions of the full body of the person in the given style; and replacing, using the generated stylized version of the full body of the real-world person, the depiction of the full body of the real-world person in the image.

2. The method according to claim 1, wherein the full body of the real-world person includes a head, arms, a torso, and legs, and wherein the stylized version of the full body of the real-world person includes stylized versions of the head, arms, torso, and legs.

3. The method according to any one of claims 1 to 2, wherein the machine learning model includes a deep neural network.

4. The method according to any one of claims 1 to 3, wherein before applying the machine learning model to the image, the method comprises: receiving an input for selecting the given style from a plurality of styles; and selecting the machine learning model from a plurality of machine learning models, the plurality of machine learning models each being configured to generate different stylized versions of the full body of a person corresponding to respective styles in the plurality of styles.

5. The method according to claim 4, wherein the plurality of styles includes at least one of a zombie style, a bodybuilder style, a cartoon style, an anime style, a Gollum style, a Neanderthal style, or a Barbie style.

6. The method according to any one of claims 1 to 5, further comprising generating the training data by performing operations that comprise: accessing a first set of latent codes by a first full-body generative adversarial network (GAN) and a second full-body GAN; rendering, by the first full-body GAN, a first synthetic full body of a person corresponding to the first set of latent codes; rendering, by the second full-body GAN, a second synthetic full body of the person corresponding to the first set of latent codes; calculating, by a direction loss model associated with the given style, a direction loss based on the second synthetic full body of the person; updating, based on the direction loss, one or more weights of the second GAN; and repeating the following operations until a stop criterion is reached: rendering the second synthetic full body of the person; calculating the direction loss; and updating the one or more weights.

7. The method according to claim 6, further comprising: determining that the stop criterion has been reached; in response to determining that the stop criterion has been reached, repeating the operations for a second set of latent codes; and Generating an image pair in the training data by applying a new latent code to the first full-body GAN and the second full-body GAN to generate a first training image among a plurality of training images depicting the full body of a synthetically rendered person and a first ground-truth stylized version of the full body of the person.

8. The method according to any one of claims 1 to 7, further comprising: Training the machine learning model based on the image pair of the training data.

9. The method according to any one of claims 1 to 8, further comprising: Accessing the first set of latent codes by a first face-based GAN and a second face-based GAN; Rendering a first synthetic face of the person corresponding to the first set of latent codes by the first face-based GAN; Rendering a second synthetic face of the person corresponding to the first set of latent codes by the second face-based GAN; Calculating a direction loss based on the second synthetic face of the person by the direction loss model associated with the given style; Updating one or more weights of the second GAN based on the direction loss; and Repeating the following operations until the stopping criterion is reached again: rendering the second synthetic face of the person; calculating the direction loss; and updating the one or more weights.

10. The method according to claim 9, further comprising: Applying a new latent code to the first face-based GAN and the second face-based GAN; and Generating a second image pair of the training data based on the outputs of the first face-based GAN and the second face-based GAN.

11. The method according to any one of claims 1 to 10, further comprising: Overlaying the synthetic face of the person generated by the second face-based GAN with the second synthetic full body of the person generated by the second full-body GAN in the image pair; Smoothingly fusing the second synthetic face of the person generated by the second face-based GAN with the second synthetic full body of the person generated by the second full-body GAN in the image pair; and Storing the second synthetic face of the person generated by the second face-based GAN that has been smoothly fused with the second synthetic full body of the person generated by the second full-body GAN in the image pair as the second synthetic full body of the person in the image pair.

12. The method according to any one of claims 1 to 11, further comprising: Selecting a part of the second synthetic full body of the person corresponding to a body part; Identifying a set of weights of the second GAN corresponding to the body part; Calculating a direction loss based on the selection of the part of the second synthetic full body of the person corresponding to the body part by the direction loss model associated with the given style; and Updating the set of weights corresponding to the body part based on the direction loss without updating other weights associated with other body parts.

13. The method according to claim 12, further comprising: Divide the second synthetic full body of the person into separate parts corresponding to different body parts; And Calculate the orientation loss separately for each of the separate parts.

14. The method according to any one of claims 1 to 13, further comprising training the machine learning model by performing a training operation, the training operation Comprising: Access the training data; Apply the machine learning model to a first training data set including a first training image of a synthetically rendered full body of a person to generate an estimated stylized version of the full body of the person depicted in the first training image; Calculate the deviation between the estimated stylized version of the full body of the person depicted in the first training image and the ground truth stylized version of the full body of the person depicted in the first training image; And Update one or more parameters of the machine learning model based on the deviation between the estimated stylized version of the full body of the person depicted in the first training image and the ground truth stylized version of the full body of the person depicted in the first training image.

15. The method according to claim 14, further Comprising: Generate full body key points for the full body of the person depicted in the first training image, wherein the estimated stylized version of the full body of the person depicted in the first training image is generated based on the full body key points.

16. The method according to any one of claims 1 to 15, Wherein, The image is received as a frame of a video depicting a real-world person, and wherein the depiction of the full body of the real-world person in the video is replaced in real time with the stylized version.

17. A system, Comprising: A processor; And A memory component having instructions stored thereon that, when executed by the processor, cause the processor to perform operations, the operations including: Receive an image including a depiction of the full body of a real-world person; Apply a machine learning model to the image to generate a stylized version of the full body of the real-world person corresponding to a given style, the machine learning model being trained using training data to establish a relationship between a plurality of training images depicting a synthetically rendered full body of a person and the corresponding ground truth stylized versions of the full body of the person of the given style; and Replace the depiction of the full body of the real-world person in the image with the generated stylized version of the full body of the real-world person.

18. The system according to claim 17, Wherein, The full body of the real-world person includes a head, arms, a torso, and legs, and wherein the stylized version of the full body of the real-world person includes stylized versions of the head, arms, torso, and legs.

19. The system according to any one of claims 17 to 18, Wherein, The machine learning model includes a deep neural network.

20. A non-transitory computer-readable storage medium having instructions stored thereon that, when executed by a processor, cause the processor to perform operations, the operations Comprising: Receive an image including a depiction of the full body of a real-world person; Apply a machine learning model to the image to generate a stylized version of the full body of the real-world person corresponding to a given style, the machine learning model being trained using training data to establish a relationship between a plurality of training images depicting the full body of a synthetically rendered person and the corresponding ground truth stylized version of the full body of the person of the given style; And Replace the depiction of the full body of the real-world person in the image with the generated stylized version of the full body of the real-world person.