Augmented expression system
Patent Information
- Application Number
- KR1020267027751
- Authority / Receiving Office
- KR · KR
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2018-04-18
- Filing Date
- 2019-04-17
- Publication Date
- 2026-09-04
Smart Images

Figure P1020267027751_ABST
Abstract
Description
Technology Field
[0001] preference
[0002] The present application claims the benefit of priority of U.S. Provisional Patent Application No. 62 / 659,337 filed April 18, 2018, the benefit of which is claimed and the entirety thereof is incorporated herein by reference.
[0003] The embodiments of the present disclosure generally relate to mobile computing technology, and more specifically, are, not limited to, systems for creating and displaying augmented reality interfaces. Background Technology
[0004] Augmented reality (AR) is a live, direct, or indirect view of a physical real-world environment that has elements augmented by computer-generated sensory inputs. Prior art literature
[0005] International Patent Publication WO 2013 / 027893 US Patent Publication US 2013 / 235045 International Patent Publication WO 2014 / 194439 US Patent Publication US 2018 / 005420 US Patent Publication US 2012 / 130717 International Patent Publication WO 2008 / 096099 International Patent Publication WO 2016 / 090605 Brief explanation of the drawing To easily identify the discussion of any specific element or action, the top digit or numbers of the reference numbers refer to the drawing number where the element is first introduced. FIG. 1 is a block diagram illustrating an exemplary messaging system for exchanging data (e.g., messages and associated content) through a network according to some embodiments, wherein the messaging system includes an augmented representation system. FIG. 2 is a block diagram illustrating additional details regarding a messaging system according to exemplary embodiments. FIG. 3 is a block diagram illustrating various modules of an augmented representation system according to specific exemplary embodiments. FIG. 4 is a flowchart illustrating a method for presenting an augmented reality display according to specific exemplary embodiments. FIG. 5 is a flowchart illustrating a method for presenting an augmented reality display according to specific exemplary embodiments. FIG. 6 is a flowchart illustrating a method for presenting an augmented reality display according to certain exemplary embodiments. FIG. 7 is a flowchart illustrating a method for generating a 3D model based on face tracking inputs according to specific exemplary embodiments. FIG. 8 is an example of an interface for displaying augmented reality images according to specific exemplary embodiments. FIG. 9 is an example of an interface for displaying augmented reality images according to specific exemplary embodiments. FIG. 10 is an example of an interface for displaying augmented reality images according to specific exemplary embodiments. FIG. 11 is a block diagram illustrating a representative software architecture that can be used to implement various embodiments and can be used with various hardware architectures described in this specification. FIG. 12 is a block diagram illustrating components of a machine according to some exemplary embodiments capable of reading instructions from a machine-readable medium (e.g., a machine-readable storage medium) and performing any one or more of the methodologies discussed herein. Specific details for implementing the invention
[0006] The embodiments described herein relate to an augmented representation system that creates an interface specifically configured to present an augmented reality perspective and causes its display. The augmented representation system receives image and video data of a user and, based on the image and video data, tracks the user's facial landmarks in real time to create and present a three-dimensional (3D) bitmoji of the user. The system presents an augmented reality display comprising a depiction of the user including augmented reality 3D bitmoji heads and elements that are overlaid on locations in the presentation of the image data. For example, the user may display, stream, or record a video depicting himself and his surroundings. The system detects the presence of facial landmarks and, in response, "replaces" the user's head with a 3D bitmoji head that mimics and tracks the user's facial expression movements.
[0007] An augmented representation system receives image and video data from a camera (e.g., a front camera) associated with a client device, wherein the image and video data includes a set of facial landmarks of the user. In response to detecting the presence of a set of facial landmarks, the system generates a 3D model (e.g., a 3D BitMoji) based on the set of facial landmarks, overlays the 3D model on a position above the set of facial landmarks in a presentation of image data, and dynamically animates the 3D model based on the detected movements of the set of facial landmarks, for example, by face tracking.
[0008] In some embodiments, the system identifies a user or users based on a set of facial landmarks depicted by image data and generates a 3D model based on a user profile associated with the user. For example, the user may define the display characteristics of the 3D Bitmoji in the user profile by selecting a specific 3D Bitmoji from a set of 3D Bitmojis, or configure the 3D Bitmoji based on selections of individual components of the 3D Bitmoji, such as facial features, colors, hairstyles, eye color, as well as accessories (e.g., glasses, hats, sunglasses, monocles). In additional embodiments, the system generates the 3D Bitmoji based on attributes of the set of the user's facial landmarks. For example, attributes may include distances and slopes between various facial landmarks, the size and shape of the facial landmarks, as well as the user's complexion. The system selects individual components of the 3D Bitmoji based on the attributes of the facial landmarks.
[0009] The augmented expression system identifies a user's facial expression based on facial landmark tracking techniques applied to a set of facial landmarks depicted by image data. The augmented expression system receives inputs as the movements and positions of facial landmarks relative to one another through a face tracking module. For example, the face tracking module receives inputs including landmark identifiers (e.g., nose, mouth, left eye, right eye), input values indicating how much a facial landmark has moved or changed relative to a static position, and a direction of movement indicating the direction of the movement. Based on these inputs and combinations of inputs, the face tracking module of the augmented expression system determines the user's expression. For example, the system may track the simultaneous raising of the user's eyebrows while the user's mouth is open, and in response, determine that the movements and positions of the facial landmarks correspond to a "surprised" expression.
[0010] The augmented expression system includes a data repository containing expression definitions, and the expression definitions include mapping various face tracking inputs to expressions. The expressions include display commands for the 3D Bitmoji, going beyond simply causing the 3D Bitmoji to mimic the movements and positions of a set of the user's face landmarks. For example, the display commands may include 3D graphic elements (e.g., balloons, confetti, hearts) to be presented within the presentation of image data in response to identifying a specific expression or by exaggerating one or more face landmarks to emphasize the expression (e.g., dropping the jaw, popping the eyes out of the head, or changing the eyes into a stylized "X"). The user may additionally interact with the 3D graphic elements by providing additional inputs as movements of the face landmarks (e.g., winking, pursing, or blowing one eye) or as haptic inputs to a touchable device.
[0011] The augmented representation system facilitates the sharing and distribution of content including 3D BitMoji. For example, a user can stream or record a video containing a composite presentation that includes 3D BitMoji and video. In some embodiments, the user can select a previously recorded image or video and have the augmented representation system augment the previously recorded image or video with a display of 3D BitMoji that mimics the user's facial movements and expressions. In additional embodiments, the user can record or stream a composite presentation containing a display of 3D BitMoji in real time so that the 3D BitMoji is rendered on the image or video when image data is received by a client device.
[0012] The user may additionally share or distribute a synthetic presentation containing 3D Bitmoji by assigning 3D Bitmoji and associated image data to a message to be distributed to one or more recipients. In response to assigning 3D Bitmoji to a message, the system generates a flattened presentation of the synthetic presentation based on the image data and the 3D Bitmoji. The user may assign or transmit the flattened presentation to one or more recipients (e.g., as a message or by adding the flattened presentation to a story associated with the user's user profile).
[0013] In additional embodiments, the system transmits 3D Bitmoji and image data depicting a user to one or more recipients individually, and may enable the recipients to generate a presentation of the image data including a display of the 3D Bitmoji at a location within the image data based on a set of face landmarks. For example, the system may segment the 3D Bitmoji into a set of regions, wherein each region corresponds to a distinct face landmark among the set of face landmarks. When the recipient devices receive the image data including the set of face landmarks and the 3D Bitmoji, the system enables the recipient devices to generate a presentation of the image data by presenting the 3D Bitmoji at a location within the image data based on the locations of each of the face landmarks.
[0014] In some embodiments, a user may transmit 3D BitMoji commands that enable recipients of the 3D BitMoji commands to display the 3D BitMoji on content generated on their personal devices. For example, a user may configure a 3D BitMoji by selecting one or more BitMoji components and attributes (e.g., hair shape, color, accessories, etc.). The user assigns an identifier to the 3D BitMoji (e.g., "Angry Putin"), and the augmented representation system indexes and stores the 3D BitMoji commands and the identifier in a memory location associated with the user's user profile. Recipients of the 3D BitMoji commands may generate augmented reality presentations that include a display of the 3D BitMoji (e.g., "Angry Putin").
[0015] FIG. 1 is a block diagram illustrating an exemplary messaging system (100) for exchanging data (e.g., messages and associated content) over a network. The messaging system (100) includes a plurality of client devices (102), each of which hosts a plurality of applications including a messaging client application (104). Each messaging client application (104) is coupled to communicate with other instances of the messaging client application (104) and the messaging server system (108) over a network (106) (e.g., the Internet).
[0016] Accordingly, each messaging client application (104) can communicate with other messaging client applications (104) and a messaging server system (108) and exchange data through a network (106). The data exchanged between messaging client applications (104) and between a messaging client application (104) and a messaging server system (108) includes functions (e.g., commands to activate functions) as well as payload data (e.g., text, audio, video, or other multimedia data).
[0017] The messaging server system (108) provides server-side functionality to a specific messaging client application (104) via the network (106). Although specific functions of the messaging system (100) are described in this specification as being performed by the messaging client application (104) or by the messaging server system (108), it will be recognized that the location of specific functions within the messaging client application (104) or the messaging server system (108) is a design choice. For example, it may be technically desirable to initially place specific technologies and functions within the messaging server system (108), but later transfer these technologies and functions to the messaging client application (104) when the client device (102) has sufficient processing capacity.
[0018] The messaging server system (108) supports various services and operations provided to the messaging client application (104). Such operations include transmitting data to the messaging client application (104), receiving data from it, and processing data generated by it. In some embodiments, this data includes, for example, message content, client device information, geolocation information, media annotations and overlays, message content persistence conditions, social network information, and live event information. In other embodiments, other data is used. Data exchange within the messaging system (100) is initiated and controlled through functions available via the GUIs of the messaging client application (104).
[0019] Now, referring specifically to the messaging server system (108), an application program interface (API) server (110) is coupled to an application server (112) to provide a programmatic interface. The application server (112) is communicably coupled to a database server (118), which facilitates access to a database (120) where data associated with messages processed by the application server (112) is stored.
[0020] Specifically, regarding the application program interface (API) server (110), this server receives and transmits message data (e.g., commands and message payloads) between the client device (102) and the application server (112). Specifically, the application program interface (API) server (110) provides a set of interfaces (e.g., routines and protocols) that can be called or queried by a messaging client application (104) to activate the functionality of the application server (112). The application program interface (API) server (110) exposes various functions supported by the application server (112), including account registration, login functionality, transmission of messages via the application server (112) from a specific messaging client application (104) to another messaging client application (104), transmission of media files (e.g., images or videos) from the messaging client application (104) to the messaging server application (114), and setting of a collection of media data (e.g., stories) for possible access by another messaging client application (104), searching for a list of friends of the user of the client device (102), searching for such collections, searching for messages and content, adding and removing friends to a social graph, the location of friends within the social graph, and opening application events (e.g., related to the messaging client application (104)).
[0021] The application server (112) hosts a number of applications and subsystems, including a messaging server application (114), an image processing system (116), a social network system (122), and an augmented representation system (124). The messaging server application (114) implements a number of message processing techniques and functions, particularly those related to the aggregation and other processing of content (e.g., text and multimedia content) contained in messages received from multiple instances of the messaging client application (104). As described in more detail, text and media content from multiple sources may be aggregated into collections of content (e.g., called stories or galleries). These collections are then made available to the messaging client application (104) by the messaging server application (114). Other processor and memory-intensive data processing may also be performed on the server side by the messaging server application (114), taking into account the hardware requirements for such processing.
[0022] The application server (112) also includes an image processing system (116) dedicated to performing various image processing operations on an image or video received within the payload of a message in a messaging server application (114).
[0023] The social network system (122) supports various social networking functions and services and enables the messaging server application (114) to use these functions and services. To this end, the social network system (122) maintains and accesses an entity graph (304) within a database (120). Examples of functions and services supported by the social network system (122) include the identification of other users of the messaging system (100) with whom a specific user has a relationship or "follows," and also the identification of other entities and the interests of the specific user.
[0024] The application server (112) is communicably coupled to the database server (118), which facilitates access to the database (120) where data associated with messages processed by the messaging server application (114) is stored.
[0025] FIG. 2 is a block diagram illustrating further details regarding a messaging system (100) according to exemplary embodiments. Specifically, the messaging system (100) is illustrated as including a messaging client application (104) and an application server (112), which ultimately implements a number of subsystems, namely a short-term timer system (202), a collection management system (204), and an annotation system (206).
[0026] The short-term timer system (202) is responsible for enforcing temporary access to content, such as 3D Bitmoji, which is permitted by the messaging client application (104) and the messaging server application (114). To this end, the short-term timer system (202) includes a plurality of timers that selectively display messages and associated content through the messaging client application (104) and enable access thereto, based on durations and display parameters associated with messages, collections of messages, or graphic elements. Further details regarding the operation of the short-term timer system (202) are provided below.
[0027] The collection management system (204) is responsible for managing collections of media (e.g., collections of text, image, video and audio data, 3D BitMoji, and flattened presentations of 3D BitMoji). In some examples, collections of content (e.g., messages including images, videos, text and audio) may be organized into "event galleries" or "event stories." Such collections may be available for a specified period, such as the duration of the event to which the content is related. For example, content related to a music concert may be available as a "story" for the duration of that music concert. The collection management system (204) may also be responsible for posting an icon in the user interface of the messaging client application (104) that provides a notification of the existence of a specific collection.
[0028] The collection management system (204) further includes a curation interface (208) that allows a collection manager to manage and curate specific collections of content. For example, the curation interface (208) enables an event organizer to curate a collection of content related to a specific event (e.g., remove inappropriate content or duplicate messages). Additionally, the collection management system (204) automatically curates content collections using machine vision (or image recognition technology) and content rules. In certain embodiments, a reward may be paid to the user for including user-generated content in a collection. In such cases, the curation interface (208) operates to automatically pay such users for using their content.
[0029] The annotation system (206) provides various functions that enable the user to annotate media content associated with a message, or otherwise modify or edit it. For example, the annotation system (206) provides functions related to the creation and publication of media overlays for messages processed by the messaging system (100). The annotation system (206) effectively supplies media overlays to the messaging client application (104) based on the geolocation of the client device (102). In another example, the annotation system (206) effectively supplies media overlays to the messaging client application (104) based on other information, such as the social network information of the user of the client device (102). Media overlays may include audio and visual content and visual effects. Examples of audio and visual content include photographs, texts, logos, animations, and sound effects, as well as animated face models such as those generated by the augmented representation system (124). Examples of visual effects include color overlays. Audio and visual content or visual effects may be applied to media content items (e.g., photos) on the client device (102). For example, a media overlay may include text that can be overlaid on a photo that is taken and generated by the client device (102). In another example, the media overlay may include an identification of a location overlay (e.g., Venice beach), the name of a live event, or the name of a merchant overlay (e.g., Beach Coffee House). In another example, the annotation system (206) uses the geolocation of the client device (102) to identify a media overlay containing the name of a merchant at the geolocation of the client device (102). The media overlay may include other indications associated with the merchant.Media overlays are stored in a database (120) and can be accessed through a database server (118).
[0030] In one exemplary embodiment, the annotation system (206) provides a user-based publication platform that enables users to select a geolocation on a map and upload content associated with the selected geolocation. Users can also specify situations in which a particular media overlay should be provided to other users. The annotation system (206) generates a media overlay that includes the uploaded content and associates the uploaded content with the selected geolocation.
[0031] In another exemplary embodiment, the annotation system (206) provides a merchant-based publishing platform that enables merchants to select a specific media overlay associated with a geolocation through a bidding process. For example, the annotation system (206) associates the media overlay of the highest bidder with a corresponding geolocation for a predefined amount of time.
[0032] FIG. 3 is a block diagram illustrating components of an augmented representation system (124) configured to generate and display an augmented reality presentation based on image data and 3D bitmoji, according to some exemplary embodiments. The augmented representation system (124) is illustrated as comprising a feature tracking module (302), a modeling module (304), a presentation module (306), and a communication module (308), all of which are configured to communicate with each other (e.g., via a bus, shared memory, or switch). Any one or more of these modules may be implemented using one or more processors (310) (e.g., by configuring such one or more processors to perform the functions described for the module) and thus may include one or more of the processors (310).
[0033] Any one or more of the described modules may be implemented using hardware alone (e.g., one or more of the processors (310) of the machine) or a combination of hardware and software. For example, any described module of the augmented representation system (124) may physically include an array of one or more processors (e.g., one or more of the processors (310) of the machine or a subset thereof) configured to perform the operations described herein for that module. As another example, any module of the augmented representation system (124) may include software, hardware, or both that constitute an array of one or more processors (310) (e.g., one or more of the processors of the machine) to perform the operations described herein for that module. Thus, different modules of the augmented representation system (124) may include and be configured with different arrays of such processors (310) or a single array of such processors (310) at different times. Furthermore, any two or more modules of the augmented representation system (124) may be combined into a single module, and the functions described herein for the single module may be subdivided among multiple modules. Furthermore, according to various exemplary embodiments, the modules described herein as being implemented within a single machine, database, or device may be distributed across multiple machines, databases, or devices.
[0034] FIG. 4 is a flowchart illustrating a method (400) for presenting an augmented reality display according to certain exemplary embodiments. The operations of the method (400) may be performed by the modules described above with respect to FIG. 3. As illustrated in FIG. 4, the method (400) includes one or more operations (402, 404, 406, 408, and 410).
[0035] In operation 402, the presentation module (306) causes the display of a presentation of image data on a client device (e.g., client device (102A)), wherein the image data depicts a set of facial landmarks of the user. For example, the image data may be collected by a camera associated with the client device (102A), such as a front camera.
[0036] In operation 404, the modeling module (304) generates a 3D model (e.g., 3D Bitmoji) based on a set of face landmarks in response to the presentation module (306) displaying a presentation of image data. In some embodiments, the modeling module (304) generates a 3D model in response to detecting the presence of face landmarks within the image data. For example, the face tracking module (302) may utilize various face motion capture methodologies, including marker-based as well as markerless technologies. Markerless face motion capture methodologies identify and track face movements and expressions based on face landmarks, such as nostrils, lips and corners of the eyes, pupils, and wrinkles. In response to detecting the presence of such face landmarks within the image data, the face tracking module (302) causes the modeling module to generate a 3D model based on identifying the face landmarks.
[0037] In some embodiments, the face tracking module (302) identifies a user based on attributes of a set of identified face landmarks, such as distances and slopes between each face landmark, the size and shape of each face landmark, as well as device attributes of the client device (102A) (e.g., device identifier), and retrieves a user profile associated with the user, wherein the user profile includes display commands for 3D bitmoji.
[0038] In additional embodiments, the modeling module (304) generates a 3D bitmoji on the fly based on the attributes of face landmarks so that the 3D bitmoji resembles the user. For example, the modeling module (304) may access a bitmoji repository containing bitmoji elements, select a set of bitmoji elements based on the attributes of face landmarks, and generate a 3D bitmoji based on the selected bitmoji elements.
[0039] In some embodiments, the 3D bitmoji generated by the modeling module (304) comprises a set of bitmoji regions, wherein each region corresponds to a distinct face landmark among a set of face landmarks (e.g., a region corresponding to a face landmark representing the user's mouth, a region corresponding to a face landmark representing the user's left eye, etc.). In additional embodiments, the 3D bitmoji generated by the modeling module (304) comprises a blendshape representing a set of face landmarks as a series of vertex positions.
[0040] In operation 406, the presentation module (306) overlays a 3D bitmoji generated by the modeling module (304) on a location on the image data displayed on the client device (102A) based on the locations of a set of face landmarks. For example, the face tracking module (302) may identify a set of reference features based on a set of face landmarks and cause the presentation module (306) to overlay a presentation of the 3D bitmoji on the image data on the client device (102A).
[0041] In action 408, the feature tracking module (302) receives face tracking input (e.g., movement of a face landmark among a set of face landmarks). For example, the user may open or close his mouth, raise his eyebrows, or turn his head. In response to the feature tracking module (302) receiving the face tracking input, in action 410, the presentation module (302) dynamically animates an area of 3D bitmoji corresponding to the face landmark on the client device (102A). The 3D bitmoji displayed on the client device (102A) mimics the user's movements and facial expressions.
[0042] FIG. 5 is a flowchart illustrating a method (500) for presenting an augmented reality display according to certain exemplary embodiments. The operations of the method (500) may be performed by the modules described above with respect to FIG. 3. As illustrated in FIG. 5, the method (500) includes one or more operations (502, 504, 506, 508, and 510).
[0043] In operation 502, the face tracking module (302) identifies the user's expression (e.g., facial expression) based on inputs received as movements and positions of a set of face landmarks. For example, the face tracking module (302) receives inputs including a landmark identifier (e.g., nose, mouth, left eye, right eye), an input value indicating how much the face landmark has moved or changed relative to a static position, and a direction of movement indicating the direction of movement. In response to receiving the inputs, the face tracking module (302) accesses a repository containing a set of expression definitions, wherein the expression definitions include mapping various face tracking inputs to predefined expressions.
[0044] In some embodiments, the user of the client device (102A) explicitly provides expression definitions through one or more user inputs. For example, the user may form his face into a specific expression and capture an image depicting that expression on the client device (102A). The face tracking module (302) determines the relative positions of each face landmark among a set of face landmarks in the image and assigns the relative positions of the face landmarks to a specific expression, to which the user may assign expression identifiers as well as display commands. For example, the user may instruct the augmented expression system (124) to display a specific graphic element within a presentation of image data in response to detecting a specific expression based on the face tracking inputs.
[0045] In operation 504, the modeling module (304) selects a graphic element from a set of graphic elements based on an expression identified by the face tracking module (302). For example, in response to the face tracking module (302) identifying an expression, the modeling module (304) retrieves a corresponding expression definition that includes commands for displaying a specific graphic element, such as a 3D object, along with a presentation of image data from the client device (102A).
[0046] In operation 506, the presentation module (306) causes the display of a graphic element within the presentation of image data on the client device (102A). In some embodiments, the position and orientation of the graphic element are based on the user's facial expression. For example, each graphic element may include display commands that define how and where the graphic element should be displayed in the presentation of image data based on the positions, movements, and orientations of a set of facial landmarks tracked by the face tracking module (302).
[0047] In operation 508, the face tracking module (302) receives a second face tracking input (e.g., a movement distinct from the previous movement). In operation 510, the presentation module (306) applies changes to a graphic element based on attributes of the second face tracking input, such as input values.
[0048] For example, a graphic element displayed within a presentation of image data may include a 3D balloon displayed in close proximity to an area of a 3D bitmoji representing the user's lips. The user may purse their lips to indicate that they are blowing the balloon. In response to detecting the movement of the user's lips, the presentation module (306) causes the 3D balloon to inflate as if the user were inflating it.
[0049] FIG. 6 is a flowchart illustrating a method (600) for presenting an augmented reality display according to certain exemplary embodiments. The operations of the method (600) may be performed by the modules described above with respect to FIG. 3. As illustrated in FIG. 6, the method (600) includes one or more operations (602, 604, 606, and 608).
[0050] In operation 602, the communication module (308) generates a flattened presentation based on image data and 3D bitmoji. The flattened presentation includes a merging of the image data and 3D bitmoji. After merging the image data and 3D bitmoji into a single layer, in operation 604, the user of the client device (102A) assigns the flattened presentation to a message. For example, the message may be addressed to one or more recipients, including a second client device (102B).
[0051] In operation 606, the communication module (308) distributes a message containing a flattened presentation to recipients including a second client device (102B). In operation 608, the presentation module (306) causes a display of the flattened presentation on the second client device (102B), for example, as a short-term message.
[0052] FIG. 7 is a flowchart (700) illustrating a method for generating a 3D model based on face tracking inputs according to certain exemplary embodiments. As illustrated in FIG. 7, the flowchart (700) includes steps (700A, 700B, and 700C).
[0053] In step 700A, image data (702) is received from a client device (102A) as described in operation 402 of the method (400) in FIG. 4. For example, the image data (702) may be collected by a camera associated with the client device (102A), such as a front camera.
[0054] In step 700B, the face tracking module (302) identifies a set (704) of face landmarks within the image data (702). As illustrated in FIG. 7, the set (704) of face landmarks can be represented as a distribution of points representing the face landmarks of the user's face depicted in the image data (702). The face landmarks (704) may include, for example, the tip of the nose, the corners of the eyes, the corners of the eyebrows, the corners of the mouth, and the pupils. The distribution of points can be subdivided into collections of points so that each collection of points corresponds to an area of the 3D Bitmoji (706). For example, a collection (708) of points lining the lower part of the face depicted in the image data (702) may correspond to the lower part (710) of the 3D Bitmoji (706).
[0055] In step 700C, the modeling module (304) generates a 3D bitmoji (706) based on a set of face landmarks (704), as described in operation 404 of the method (400) in FIG. 4. As can be seen in step 700C of FIG. 7, the 3D bitmoji (706) is overlaid on the image data (702) at a location based on the set of face landmarks (704).
[0056] FIG. 8 is an example (800) of an interface for displaying an augmented reality image according to certain exemplary embodiments. As illustrated in FIG. 8, the example (800) includes a depiction of a 3D bitmoji (802) overlaid on image data (804) according to certain exemplary embodiments such as those described in the method (400) of FIG. 4.
[0057] FIG. 9 is an example (900) of an interface for displaying an augmented reality image according to certain exemplary embodiments. As illustrated in FIG. 9, the example (900) includes a depiction of a 3D bitmoji (902) overlaid on image data (904). The example (900) also includes an interface item (906) for receiving user input for recording content including composite presentations comprising image data (804) and the 3D bitmoji (902).
[0058] As illustrated in FIG. 9, the 3D bitmoji (902) includes an exaggerated depiction of an expression identified by the augmented expression system (124), as described in the method (500) of FIG. 5. In response to detecting a specific expression based on one or more face tracking inputs received from the client device (102A), the augmented expression system (124) causes the 3D bitmoji to present an exaggerated depiction of the corresponding expression.
[0059] FIG. 10 is an example (1000) of an interface for displaying an augmented reality image according to certain exemplary embodiments. As illustrated in FIG. 10, the example (10) includes a 3D bitmoji (1002) and a graphic element (1004) overlaid on image data (1006) according to certain exemplary embodiments such as those described in the method (500) of FIG. 5.
[0060] For example, in response to the face tracking module (302) receiving a face tracking input corresponding to a specific expression (e.g., the user pursing their lips), the presentation module (306) causes the display of a graphic element (1004) at a location within the image data (1006).
[0061] Software Architecture
[0062] FIG. 11 is a block diagram illustrating an exemplary software architecture (1106) that may be used with various hardware architectures described herein. FIG. 11 is a non-limiting example of a software architecture, and it will be recognized that many other architectures may be implemented to facilitate the functionality described herein. The software architecture (1106) may be executed on hardware such as the machine (1200) of FIG. 12, which includes processors (1204), memory (1214), and I / O components (1218), among others. A representative hardware layer (1152) is illustrated and may represent, for example, the machine (1100) of FIG. 11. The representative hardware layer (1152) includes a processing unit (1154) having associated executable instructions (1104). The executable instructions (1104) represent executable instructions of the software architecture (1106) that include implementations of methods, components, etc., described herein. The hardware layer (1152) also includes memory / storage (1156), which are memory and / or storage modules, and these also have executable instructions (1104). The hardware layer (1152) may also include other hardware (1158).
[0063] In the exemplary architecture of FIG. 11, the software architecture (1106) can be conceptualized as a stack of layers, each layer providing specific functionality. For example, the software architecture (1106) may include layers such as an operating system (1102), libraries (1120), applications (1116), and a presentation layer (1114). During operation, applications (1116) and / or other components within the layers may initiate application programming interface (API) API calls (1108) through the software stack and receive responses in response to API calls (1108). The layers illustrated are essentially representative, and not all software architectures have all layers. For example, some mobile or special-purpose operating systems may not provide frameworks / middleware (1118), while others may provide such layers. Other software architectures may include additional or different layers.
[0064] The operating system (1102) may manage hardware resources and provide common services. The operating system (1102) may include, for example, a kernel (1122), services (1124), and drivers (1126). The kernel (1122) may serve as an abstraction layer between the hardware and other software layers. For example, the kernel (1122) may be responsible for memory management, processor management (e.g., scheduling), component management, networking, security settings, etc. Services (1124) may provide other common services for other software layers. Drivers (1126) are responsible for controlling or interfacing with the underlying hardware. For example, depending on the hardware configuration, drivers (1126) include display drivers, camera drivers, Bluetooth® drivers, flash memory drivers, serial communication drivers (e.g., Universal Serial Bus (USB) drivers), Wi-Fi® drivers, audio drivers, power management drivers, etc.
[0065] Libraries (1120) provide a common infrastructure used by applications (1116) or other components or layers. Libraries (1120) provide functionality that enables other software components to perform tasks in a way that is easier than directly interfacing with the underlying operating system (1102) functionality (e.g., kernel (1122), services (1124) and / or drivers (1126)). Libraries (1120) may include system libraries (1144) (e.g., C standard library) that can provide functions such as memory allocation functions, string manipulation functions, mathematical functions, etc. In addition, the libraries (1120) may include API libraries (1146) such as media libraries (e.g., libraries for supporting the presentation and manipulation of various media formats such as MPREG4, H.264, MP3, AAC, AMR, JPG, PNG), graphics libraries (e.g., OpenGL frameworks that can be used to render 2D and 3D graphic content on a display), database libraries (e.g., SQLite that can provide various relational database functions), web libraries (e.g., WebKit that can provide web browsing functionality). The libraries (1120) may also include a wide variety of other libraries (1148) that provide many different APIs to applications (1116) and other software components / modules.
[0066] Frameworks / middleware (1118) (sometimes referred to as middleware) provide a high-level common infrastructure that can be used by applications (1116) and / or other software components / modules. For example, frameworks / middleware (1118) may provide various graphical user interface (GUI) functions, high-level resource management, high-level location services, etc. Frameworks / middleware (1118) may provide a wide spectrum of other APIs that can be used by applications (1116) and / or other software components / modules, some of which may be specific to a particular operating system (1102) or platform.
[0067] Applications (1116) include built-in applications (1138) and / or third-party applications (1140). Examples of representative built-in applications (1138) may include, but are not limited to, contact applications, browser applications, book reader applications, location applications, media applications, messaging applications, and / or game applications. Third-party applications (1140) may include applications developed using the ANDROID™ or IOS™ software development kit (SDK) by entities other than the vendor of a specific platform, and may be mobile software running on a mobile operating system such as IOS™, ANDROID™, WINDOWS® Phone, or other mobile operating systems. Third-party applications (1140) may initiate API calls (1108) provided by a mobile operating system (e.g., operating system (1102)) to facilitate the functionality described herein.
[0068] Applications (1116) may use built-in operating system functions (e.g., kernel (1122), services (1124) and / or drivers (1126)), libraries (1120), and frameworks / middleware (1118) to create user interfaces for interacting with users of the system. Alternatively or additionally, in some systems, interaction with the user may occur through a presentation layer such as a presentation layer (1114). In these systems, application / component "logic" may be separated from the modes of application / component that interact with the user.
[0069] FIG. 12 is a block diagram illustrating components of a machine (1200) according to some exemplary embodiments capable of reading instructions from a machine-readable medium (e.g., a machine-readable storage medium) and performing any one or more of the methodologies discussed herein. Specifically, FIG. 12 represents a schematic representation of a machine (1200) in an exemplary form of a computer system, in which instructions (1210) (e.g., software, program, application, applet, app, or other executable code) may be executed to cause the machine (1200) to perform any one or more of the methodologies discussed herein. Accordingly, instructions (1210) may be used to implement the modules or components described herein. Instructions (1210) convert a general unprogrammed machine (1200) into a specific machine (1200) programmed to perform the described and illustrated functions in the described manner. In alternative embodiments, the machine (1200) may operate as a standalone device or be coupled to other machines (e.g., networked). In a networked deployment, the machine (1200) may operate as a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment.The machine (1200) may include, but is not limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a personal digital assistant (PDA), an entertainment media system, a cellular phone, a smartphone, a mobile device, a wearable device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, a web appliance, a network router, a network switch, a network bridge, or any machine capable of executing commands (1210) that specify actions to be taken by the machine (1200) sequentially or otherwise. Furthermore, although only a single machine (1200) is exemplified, the term “machine” should also be considered to include a collection of machines that execute commands (1210) individually or jointly to perform any one or more of the methodologies discussed herein.
[0070] The machine (1200) may include processors (1204), memory / storage (1206), and I / O components (1218), which may be configured to communicate with each other, for example, via a bus (1202). The memory / storage (1206) may include memory (1214), such as main memory or other memory storage, and a storage unit (1216), both of which are accessible to the processors (1204), for example, via the bus (1202). The storage unit (1216) and the memory (1214) store instructions (1210) that implement any one or more of the methodologies or functions described herein. Instructions (1210) may also exist, wholly or partially, during the execution by the machine (1200), in memory (1214), in a storage unit (1216), in at least one of the processors (1204) (e.g., in the processor's cache memory), or any suitable combination thereof. Thus, memory (1214), storage unit (1216), and memory of the processors (1204) are examples of machine-readable media.
[0071] I / O components (1218) include a wide variety of components that perform functions such as receiving inputs, providing outputs, generating outputs, transmitting information, exchanging information, capturing measurements, etc. The specific I / O components (1218) included in a specific machine (1200) will depend on the type of machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, whereas a headless server machine may not include such a touch input device. It should be recognized that I / O components (1218) may include many other components not shown in FIG. 12. I / O components (1218) are grouped by functionality merely to simplify the discussion below, and this grouping is by no means limiting. In various exemplary embodiments, I / O components (1218) may include output components (1226) and input components (1228). The output components (1226) may include visual components (e.g., a display such as a plasma display panel (PDP), a light-emitting diode (LED) display, a liquid crystal display (LCD), a projector, or a cathode ray tube (CRT)), acoustic components (e.g., a speaker), haptic components (e.g., a vibration motor, a resistance mechanism), other signal generators, etc.Input components (1228) may include alphanumeric input components (e.g., a keyboard, a touch screen configured to receive alphanumeric input, a photo-optical keyboard, or other alphanumeric input components), point-based input components (e.g., a mouse, a touchpad, a trackball, a joystick, a motion sensor, or other pointing mechanism), haptic input components (e.g., a physical button, a touch screen providing the location and force of a touch or touch gesture, or other haptic input components), audio input components (e.g., a microphone), etc.
[0072] In additional exemplary embodiments, I / O components (1218) may include biometric components (1230), motion components (1234), environment components (1236), or position components (1238), among a wide range of other components. For example, biometric components (1230) may include components that detect expressions (e.g., hand expressions, facial expressions, voice expressions, body gestures, or eye tracking), measure biosignals (e.g., blood pressure, heart rate, body temperature, sweat, or brainwaves), and identify a person (e.g., voice identification, retinal identification, face identification, fingerprint identification, or brainwave-based identification). Motion components (1234) may include acceleration sensor components (e.g., accelerometers), gravity sensor components, rotation sensor components (e.g., gyroscopes), etc. Environmental components (1236) may include, for example, light sensor components (e.g., photometers), temperature sensor components (e.g., one or more thermometers for detecting ambient temperature), humidity sensor components, pressure sensor components (e.g., barometers), acoustic sensor components (e.g., one or more microphones for detecting background noise), proximity sensor components (e.g., infrared sensors for detecting nearby objects), gas sensors (e.g., gas detection sensors for detecting concentrations of hazardous gases for safety or measuring pollutants in the atmosphere), or other components capable of providing indications, measurements, or signals corresponding to the surrounding physical environment. Position components (1238) may include position sensor components (e.g., Global Position System (GPS) receiver components), altitude sensor components (e.g., altimeters or barometers for detecting atmospheric pressure from which altitude can be derived), direction sensor components (e.g., magnetometers), etc.
[0073] Communication can be implemented using a wide variety of technologies. I / O components (1218) may include communication components (1240) operable to connect the machine (1200) to a network (1232) or devices (1220) respectively through coupling (1222) and coupling (1224). For example, the communication component (1240) may include a network interface component, or other suitable device for interfacing with the network (1232). In additional examples, the communication components (1240) may include a wired communication component, a wireless communication component, a cellular communication component, a near-field communication (NFC) component, a Bluetooth® component (e.g., Bluetooth® Low Energy), a Wi-Fi® component, and other communication components that provide communication through other embodiments. The devices (1220) may be any of the other machine or a wide variety of peripheral devices (e.g., peripheral devices connected via a Universal Serial Bus (USB)).
[0074] Furthermore, the communication components (1240) may include components capable of detecting identifiers or operable to detect identifiers. For example, the communication components (1240) may include Radio Frequency Identification (RFID) tag reader components, NFC smart tag detection components, optical reader components (e.g., optical sensors for detecting 1-dimensional bar codes such as Universal Product Code (UPC) bar codes, multi-dimensional bar codes such as Quick Response (QR) codes, Aztec codes, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D bar codes, and other optical codes), or acoustic detection components (e.g., microphones for identifying tagged audio signals). In addition, various information, such as location via Internet Protocol (IP) geolocation, location via Wi-Fi® signal triangulation, and location via NFC beacon signal detection capable of indicating a specific location, can be derived through the communication components (1240).
[0075] Glossary
[0076] In this context, "carrier signal" refers to any intangible medium capable of storing, encoding, or carrying instructions for execution by a machine, and includes digital or analog communication signals or other intangible media to facilitate the communication of such instructions. Instructions may be transmitted or received over a network using a transmission medium through a network interface device and using any one of a number of well-known transmission protocols.
[0077] In this context, "client device" refers to any machine that interfaces with a communication network to obtain resources from one or more server systems or other client devices. A client device may be, but is not limited to, a mobile phone, desktop computer, laptop, portable information terminal (PDA), smartphone, tablet, ultrabook, netbook, laptop, multi-processor system, microprocessor-based or programmable consumer electronics, game console, set-top box, or any other communication device that a user can use to access a network.
[0078] In this context, “COMMUNICATIONS NETWORK” refers to one or more parts of a network that may be an ad-hoc network, intranet, extranet, virtual private network (VPN), local area network (LAN), wireless LAN (WLAN), wide area network (WAN), wireless WAN (WWAN), metropolitan area network (MAN), the Internet, part of the Internet, part of the Public Switched Telephone Network (PSTN), plain old telephone service (POTS) network, cellular telephone network, wireless network, Wi-Fi® network, other types of networks, or a combination of two or more such networks. For example, a network or part of a network may include a wireless or cellular network, and a combination may be a Code Division Multiple Access (CDMA) connection, a Global System for Mobile communications (GSM) connection, or other types of cellular or wireless combination.In this example, the combination can implement any various types of data transmission technology, such as 1xRTT (Single Carrier Radio Transmission Technology), EVDO (Evolution-Data Optimized) technology, GPRS (General Packet Radio Service) technology, EDGE (Enhanced Data rates for GSM Evolution) technology, 3GPP (third Generation Partnership Project) including 3G, 4G (fourth generation wireless) networks, UMTS (Universal Mobile Telecommunications System), HSPA (High Speed Packet Access), WiMAX (Worldwide Interoperability for Microwave Access), LTE (Long Term Evolution) standards, other things defined by various standards-setting organizations, other long-range protocols, or other data transmission technologies.
[0079] In this context, an "emphymeral message" refers to a message accessible for a limited duration. An emphymeral message can be text, images, videos, etc. The access time for an emphymeral message can be set by the message sender. Alternatively, the access time can be a default setting or a setting specified by the recipient. Regardless of the setting method, the message is temporary.
[0080] In this context, “Machine-Readable Medium” refers to a component, device, or other type of medium capable of storing instructions and data temporarily or permanently, and may include, but is not limited to, random access memory (RAM), read-only memory (ROM), buffer memory, flash memory, optical media, magnetic media, cache memory, other types of storage (e.g., erasable and programmable read-only memory (EEPROM)) and / or any suitable combination thereof. The term “Machine-Readable Medium” should be construed to include a single medium or multiple media capable of storing instructions (e.g., a centralized or distributed database, or an associated cache and server). The term “Machine-Readable Medium” should also be construed to include any medium or combination of multiple media capable of storing instructions (e.g., code) for execution by a machine, so that when the instructions are executed by one or more processors of the machine, the machine may perform any one or more of the methodologies described herein. Therefore, "machine-readable medium" refers not only to a single storage device or unit, but also to "cloud-based" storage systems or storage networks comprising multiple storage devices or units. The term "machine-readable medium" excludes the signal itself.
[0081] In this context, "component" refers to a device, physical entity, or logic having boundaries defined by function or subroutine calls, branch points, application program interfaces (APIs), or other technologies that provide the division or modularization of specific processing or control functions. Components can be combined with other components through their interfaces to execute machine processes. A component may be a packaged functional hardware unit designed to be used with other components and part of a program that performs a specific function among related functions. Components may constitute either software components (e.g., code implemented on machine-readable media) or hardware components. A "hardware component" is a tangible unit capable of performing specific actions and may be configured or arranged in a specific physical manner. In various exemplary embodiments, one or more computer systems (e.g., standalone computer systems, client computer systems, or server computer systems) or one or more hardware components of a computer system (e.g., processors or groups of processors) may be configured by software (e.g., applications or parts of applications) as hardware components that operate to perform specific operations as described herein. Hardware components may also be implemented mechanically, electronically, or any suitable combination thereof. For example, hardware components may include dedicated circuits or logic permanently configured to perform specific operations. Hardware components may be special-purpose processors such as Field-Programmable Gate Arrays (FPGAs) or Application Specific Integrated Circuits (ASICs).Hardware components may also include programmable logic or circuits that are temporarily configured by software to perform specific operations. For example, hardware components may include software executed by a general-purpose processor or another programmable processor. Once configured by such software, hardware components become specific machines (or specific components of a machine) uniquely customized to perform the configured functions and are no longer general-purpose processors. It will be recognized that the decision to implement a hardware component mechanically, in a dedicated, permanently configured circuit, or in a temporarily configured circuit (e.g., configured by software) may be driven by cost and time considerations. Therefore, the phrase "hardware component" (or "hardware-implemented component") should be understood to encompass type entities, that is, entities that are physically configured, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate in a specific manner or perform the specific operations described herein. When considering embodiments in which hardware components are temporarily configured (e.g., programmed), each hardware component does not need to be configured or instantiated at any single time instance. For example, if a hardware component includes a general-purpose processor configured by software to become a special-purpose processor, the general-purpose processor may be configured as different special-purpose processors (e.g., including different hardware components) at different times. Thus, the software configures a specific processor or processors to configure a specific hardware component at one time instance and a different hardware component at a different time instance. Hardware components may provide information to other hardware components and receive information from them.Accordingly, the described hardware components may be considered to be communicably coupled. In cases where multiple hardware components exist simultaneously, communication may be achieved through the transmission of signals between or between two or more of the hardware components (e.g., via appropriate circuits and buses). In embodiments where multiple hardware components are configured or instantiated at different times, communication between such hardware components may be achieved, for example, through the storage and retrieval of information within memory structures accessible to multiple hardware components. For example, one hardware component may perform an operation and store the output of that operation in a memory device communicably coupled thereto. Subsequently, additional hardware components may access the memory device to retrieve and process the stored output. Hardware components may also initiate communication with input or output devices and manipulate resources (e.g., collections of information). Various operations of the exemplary methods described herein may be performed at least partially by one or more processors configured temporarily (e.g., by software) or permanently to perform the relevant operations. Whether configured temporarily or permanently, such processors may comprise processor-implemented components that operate to perform one or more operations or functions described herein. As used herein, "processor-implemented component" refers to a hardware component implemented using one or more processors. Similarly, the methods described herein may be at least partially processor-implemented, and specific processors or processors are examples of hardware. For example, at least some of the operations of the method may be performed by one or more processors or processor-implemented components.Furthermore, one or more processors may also operate to support the execution of related operations in a "cloud computing" environment or as "software as a service" (SaaS). For example, at least some of the operations may be performed by a group of computers (as examples of machines containing processors), and these operations may be accessible via a network (e.g., the Internet) and through one or more appropriate interfaces (e.g., an Application Programming Interface (API)). The execution of specific operations may not only reside within a single machine but may also be distributed among processors deployed across multiple machines. In some exemplary embodiments, the processors or components implemented by the processors may be located in a single geographic location (e.g., a home environment, an office environment, or within a server farm). In other exemplary embodiments, the processors or components implemented by the processors may be distributed across multiple geographic locations.
[0082] In this context, "processor" refers to any circuit or virtual circuit (a physical circuit emulated by logic executed on an actual processor) that manipulates data values according to control signals (e.g., "instructions," "op codes," "machine code," etc.) and generates corresponding output signals applied to operate the machine. The processor may be, for example, a CPU (Central Processing Unit), a RISC (Reduced Instruction Set Computing) processor, a CISC (Complex Instruction Set Computing) processor, a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an RFIC (Radio-Frequency Integrated Circuit), or any combination thereof. The processor may also be a multi-core processor having two or more independent processors (sometimes referred to as "cores") capable of executing instructions simultaneously.
[0083] In this context, "timestamp" refers to a sequence of characters or encoded information that identifies when a particular event occurred, providing a date and time, for example, sometimes accurate to a fraction of a second.
[0084] In this context, "LIFT" is a measure of the performance of a targeted model in predicting or classifying cases, defined as having an improved response (with respect to the population as a whole) measured for a random selection targeting model.
[0085] In this context, "phoneme alignment" is a unit of speech that distinguishes one word from another. A phoneme can consist of a sequence of closure, plosive, and aspiration events; or, a diphthong can transition from a back vowel to a front vowel. Thus, a speech signal can be described not only by which phonemes it contains, but also by the positions of the phonemes. Therefore, phoneme alignment can be described as the "time-alignment" of phonemes within a waveform to determine the appropriate sequence and position of each phoneme within the speech signal.
[0086] In this context, "audio-to-visual conversion" refers to converting audible speech signals into visible speech, where the visible speech may include a mouth shape representing the audible speech signal.
[0087] In this context, a "Time Delayed Neural Network (TDNN)" is an artificial neural network architecture whose primary purpose is to operate on sequential data. An example would be converting continuous audio into a stream of classified phoneme labels for speech recognition.
[0088] In this context, "Bidirectional Long-Short Term Memory (BLSTM)" refers to a recurrent neural network (RNN) architecture that remembers values over arbitrary intervals. The stored values are not modified as learning progresses. RNNs allow for forward and backward connections between neurons. BLSTM is suitable for the classification, processing, and prediction of time series when given time lags of unknown magnitude and duration between events.
Claims
Claim 1 A method comprising: receiving an input at a client device for assigning a graphic element to a facial expression; causing a display of image data describing a user at the client device; creating a model based on the image data; receiving a tracking input corresponding to a movement described by the image data; detecting the facial expression based on the tracking input; retrieving the graphic element assigned to the facial expression; applying the graphic element to the model; and creating a presentation including the graphic element applied to the model. Claim 2 The method according to claim 1, wherein the step of generating the model based on the image data comprises: generating the model based on a set of face landmarks depicted in the image data. Claim 3 A method according to claim 1, wherein the step of receiving the tracking input comprises: receiving the tracking input based on the movement of a face landmark among a set of face landmarks depicted in the image data. Claim 4 The method of claim 1, wherein the step of detecting the expression based on the tracking input comprises: identifying the locations of a set of face landmarks relative to each other based on the movement described by the image data; and determining that the locations of the set of face landmarks correspond to the expression based on the mapping of face tracking inputs to a predefined set of expressions. Claim 5 The method of claim 1, wherein the step of receiving the input for assigning the graphic element to the expression comprises: receiving a user input that defines the expression based on the locations of a set of face landmarks depicted in an image captured in the client device; and assigning an expression identifier to the expression based on the user input. Claim 6 A method according to claim 1, further comprising: receiving a second tracking input corresponding to a second movement described by the image data; and applying a change to the graphic element based on the second tracking input. Claim 7 A method according to claim 1, further comprising the steps of: generating a flattened presentation based on the image data and the model, wherein the flattened presentation includes the graphic elements applied to the model; assigning the flattened presentation to a message; and transmitting the message to a second client device. Claim 8 A system comprising: a memory; and at least one hardware processor coupled to said memory and comprising instructions that cause said system to perform operations, wherein the operations include: receiving an input that assigns a graphic element to a facial expression from a client device; causing a display of image data describing a user from said client device; creating a model based on said image data; receiving a tracking input corresponding to a movement described by said image data; detecting said facial expression based on said tracking input; retrieving said graphic element assigned to said facial expression; applying said graphic element to said model; and generating a presentation including said graphic element applied to said model. Claim 9 In claim 8, generating the model based on the image data comprises: generating the model based on a set of face landmarks depicted in the image data. Claim 10 In claim 8, the system receiving the tracking input comprises: receiving the tracking input based on the movement of a face landmark among a set of face landmarks depicted in the image data. Claim 11 In claim 8, the system for detecting the expression based on the tracking input comprises: identifying the locations of a set of face landmarks relative to each other based on the movement described by the image data; and determining that the locations of the set of face landmarks correspond to the expression based on the mapping of face tracking inputs to a predefined set of expressions. Claim 12 In claim 8, receiving the input for assigning the graphic element to the expression comprises: receiving a user input defining the expression based on the locations of a set of face landmarks depicted in an image captured by the client device; and assigning an expression identifier to the expression based on the user input. Claim 13 A system according to claim 8, wherein the operations further include receiving a second tracking input corresponding to a second movement described by the image data; and applying a change to the graphic element based on the second tracking input. Claim 14 A system according to claim 8, wherein the operations further comprise generating a flattened presentation based on the image data and the model—the flattened presentation includes the graphic elements applied to the model—; assigning the flattened presentation to a message; and transmitting the message to a second client device. Claim 15 A non-transient machine-readable storage medium comprising instructions, wherein the instructions, when executed by one or more processors of a machine, cause the machine to perform operations including: receiving an input assigning a graphic element to a facial expression from a client device; causing a display of image data describing a user from the client device; creating a model based on the image data; receiving a tracking input corresponding to a movement described by the image data; detecting the facial expression based on the tracking input; retrieving the graphic element assigned to the facial expression; applying the graphic element to the model; and creating a presentation including the graphic element applied to the model. Claim 16 In paragraph 15, generating the model based on the image data comprises: generating the model based on a set of face landmarks depicted in the image data, a non-transient machine-readable storage medium. Claim 17 A non-transient machine-readable storage medium, wherein receiving the tracking input comprises: receiving the tracking input based on the movement of a face landmark among a set of face landmarks depicted in the image data. Claim 18 In claim 15, detecting the expression based on the tracking input comprises: identifying the locations of a set of face landmarks relative to each other based on the movement described by the image data; and determining that the locations of the set of face landmarks correspond to the expression based on the mapping of face tracking inputs to a predefined set of expressions, in a non-transient machine-readable storage medium. Claim 19 A non-transient machine-readable storage medium, wherein receiving the input for assigning the graphic element to the expression comprises: receiving user input defining the expression based on the locations of a set of face landmarks depicted in an image captured by the client device; and assigning an expression identifier to the expression based on the user input. Claim 20 A non-transient machine-readable storage medium according to claim 15, wherein the operations further comprise receiving a second tracking input corresponding to a second movement described by the image data; and applying a change to the graphic element based on the second tracking input.