Context-sensitive avatar subtitles
By introducing a context-sensitive avatar system in the messaging application, the avatar expression is automatically adjusted to match the text context input by the user, the problem of cumbersome processes when selecting a suitable avatar is solved, and the efficiency of message transmission and device resource utilization is improved.
Patent Information
- Application Number
- CN202080084728.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-12-09
- Filing Date
- 2020-12-09
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2040-12-09
AI Technical Summary
When users look for and select the right avatars on social media platforms to include in custom messages, the process is cumbersome and time-consuming, resulting in users losing interest in using advanced features of messaging applications.
By implementing a context-sensitive avatar system in messaging applications, the context of the subtitles is automatically determined based on the text string entered by the user, and the expression of the avatar is dynamically adjusted, thereby reducing the number of screens and interfaces that the user needs to browse.
Improves the efficiency of users sending messages on electronic devices, reduces the consumption of device resources, and allows users to create content-rich messages more seamlessly, thereby conveying information more accurately.
Smart Images

Figure CN114787813B_ABST
Abstract
Description
[0001] Priority Claim
[0002] This application claims priority to U.S. Patent Application Serial No. 16 / 707,635, filed December 9, 2019, the entire content of which is incorporated herein by reference. Technical Field
[0003] The present disclosure generally relates to providing captioned avatars using a messaging application. Background Art
[0004] Users are always looking for new ways to connect with their friends on social media platforms. One way users try to connect with friends is by sending custom messages with avatars. Many different types of avatars are available for users to choose to include in custom messages. Brief Description of the Drawings
[0005] In the drawings, which are not necessarily drawn to scale, the same reference numerals may describe the same components in different views. To easily identify the discussion of any particular element or action, one or more of the most significant digits in the reference numeral refer to the figure number in which the element was first introduced. Some embodiments are illustrated by way of example and not limitation in the figures of the drawings, in which:
[0006] Figure 1 is a block diagram showing an example messaging system for exchanging data (e.g., messages and associated content) over a network.
[0007] Figure 2 is a schematic diagram showing data that may be stored in a database of a messaging server system according to an example embodiment.
[0008] Figure 3 is a schematic diagram showing the structure of a message for communication generated by a messaging client application according to an example embodiment.
[0009] Figure 4 is a block diagram showing an example context-sensitive avatar system according to an example embodiment.
[0010] Figure 5 is a flowchart showing an example operation of a context-sensitive avatar system according to an example embodiment.
[0011] Figures 6 to 8 is an illustrative input and output of a context-sensitive avatar system according to an example embodiment.
[0012] Figure 9is a block diagram showing a representative software architecture that can be used in conjunction with various hardware architectures described herein according to an example embodiment.
[0013] Figure 10 is a block diagram showing components of a machine capable of reading instructions from a machine-readable medium (e.g., a machine-readable storage medium) and performing any one or more of the methods discussed herein. DETAILED DESCRIPTION
[0014] The following description includes systems, methods, techniques, instruction sequences, and computer program products that embody illustrative embodiments of the present disclosure. In the following description, for purposes of explanation, numerous specific details are set forth to provide an understanding of the various embodiments. However, it will be apparent to one skilled in the art that the embodiments may be practiced without these specific details. In general, well-known instruction instances, protocols, structures, and techniques are not necessarily shown in detail.
[0015] Typical user devices allow users to communicate with each other using graphics. To this end, users typically input search parameters to find the graphics that best represent the message the user is trying to convey. Specifically, many graphics are available for the user to choose from. Finding the correct graphic requires browsing multiple information pages and can be very tedious and time-consuming. Given the complexity and amount of time spent finding a graphic of interest to include in the message being sent, users are discouraged from including graphics in their messages. This results in users losing interest in using the advanced features of the messaging application, which wastes resources.
[0016] The disclosed embodiments improve the efficiency of using an electronic device by providing a messaging application that intelligently and automatically selects an avatar with a specific expression based on a message written by a user to be sent to a friend. This allows the user to spend less time composing a message that includes an avatar and enables the user to more seamlessly create a content-rich message, thereby more accurately conveying the message. Specifically, according to the disclosed embodiments, the messaging application receives a user's selection of the option to use a captioned avatar to generate a message. The messaging application presents the avatar and a caption entry area proximate to the avatar and populates the caption entry area with a text string provided by the user. As the user types one or more words in the text string, the messaging application determines the context of the caption and modifies the expression of the avatar based on the determined context.
[0017] In some cases, the subtitles are presented in a curved manner around the avatar. In some cases, the messaging application determines that the text string represents or identifies a second user. In response, the messaging application automatically retrieves a second avatar representing the second user. The messaging application then presents both the avatar of the user who composed the message and the retrieved avatar of the second user identified in the message. The avatar presentation has an expression corresponding to the context of the message. In some cases, the messaging application determines that the user is participating in a conversation with a second user. In response, the messaging application presents the avatar of each user participating in the conversation, the avatar having an expression corresponding to the context of the conversation or text in the avatar subtitle.
[0018] In this way, the disclosed embodiments improve the efficiency of using an electronic device by reducing the number of screens and interfaces that a user must navigate to select an appropriate avatar to convey a message to other users. This is accomplished by automatically determining the context of the message that the user is composing and modifying the expression of the avatar that is presented and will be included in the message sent to other users as the user types the words of the message. This also reduces the device resources (e.g., processor cycles, memory, and power usage) required to complete tasks using the device.
[0019] Figure 1 FIG. 7 is a block diagram illustrating an example messaging system 100 for exchanging data (e.g., messages and associated content) over a network 106. The messaging system 100 includes a plurality of client devices 102, each client device 102 hosting a number of applications including a messaging client application 104 and a third-party application 105. Each messaging client application 104 is communicatively coupled via a network 106 (e.g., the Internet) to other instances of the messaging client application 104, the third-party application 105, and the messaging server system 108.
[0020] Accordingly, each messaging client application 104 and third-party application 105 is able to communicate and exchange data via the network 106 with another messaging client application 104 and third-party application 105 and with the messaging server system 108. The data exchanged between the messaging client application 104 and the third-party application 105 and between the messaging client application 104 and the messaging server system 108 includes functionality (e.g., commands to activate functionality) as well as payload data (e.g., text, audio, video, or other multimedia data). Any disclosed communication between the messaging client application 104 and the third-party application 105 can be transmitted directly from the messaging client application 104 to the third-party application 105 and / or transmitted indirectly (e.g., via one or more servers) from the messaging client application 104 to the third-party application 105.
[0021] The third - party application 105 and the messaging client application 104 are applications that include a feature set allowing the client device 102 to access the context - sensitive avatar system 124. The third - party application 105 is an application separate and distinct from the messaging client application 104. The third - party application 105 is downloaded and installed by the client device 102 separately from the messaging client application 104. In some implementations, the third - party application 105 is downloaded and installed by the client device 102 either before or after downloading and installing the messaging client application 104. The third - party application 105 is provided by an entity or organization different from the entity or organization providing the messaging client application 104. The third - party application 105 is an application that can be accessed by the client device 102 using login credentials independent of the login credentials of the messaging client application 104. That is, the third - party application 105 can maintain a first user account while the messaging client application 104 can maintain a second user account. In an implementation, the third - party application 105 can be accessed by the client device 102 to perform various activities and interactions such as listening to music, watching videos, tracking workouts, viewing graphical elements (e.g., stickers), communicating with other users, etc.
[0022] As an example, the third - party application 105 can be a social networking application, a dating application, a ride - sharing or car - sharing application, a shopping application, a trading application, a gaming application, an imaging application, a music application, a video - browsing application, a workout - tracking application, a health - monitoring application, a graphical - element or sticker - browsing application, or any other suitable application.
[0023] The messaging client application 104 allows the user to access the camera device function. The camera device function of the messaging client application 104 activates the front - facing camera device of the client device 102 and presents on the display screen of the client device 102 the video or image captured or received by the front - facing camera device while the video or image is being captured. In an implementation, the front - facing camera device is integrated with the screen presenting the content captured by the front - facing camera device or is placed on the same side of the client device 102 as the screen presenting the content captured by the front - facing camera device.
[0024] After the user presses a suitable button of the messaging client application 104 to store the image or video captured by the front - facing camera device, the messaging client application 104 allows the user to view or edit the captured image. In some cases, one or more editing tools can be presented to modify or edit the stored image. Such editing tools can include a text tool that allows the user to add text to the image or video. Such text tools include large - text options, avatar - with - caption options, rainbow - text options, script - text options, etc.
[0025] In response to receiving a user's selection of a captioned avatar option, the messaging client application 104 retrieves the user's avatar. The avatar is presented on top of the captured image or video and includes a text entry area. In some cases, the avatar is initially presented with a neutral expression (e.g., not smiling or sad). In some embodiments, the avatar is presented with a default expression that can be selected by the user. The text entry area can be close to the avatar, such as above or below the avatar. The text entry area allows the user to enter a text string. As the user types words of the text string, the string wraps around the avatar in a circular manner to surround the avatar.
[0026] In some embodiments, the messaging client application 104 processes and analyzes one or more words of the text string entered by the user. The messaging client application 104 determines the context of the string based on one or more words of the text string. For example, if the string includes positive or happy words (e.g., glad, excited, ecstatic, impressed, etc.), then the messaging client application 104 determines the context as positive or happy. In response, the messaging client application 104 retrieves an avatar expression associated with or representing a positive or happy emotion. Once the context is determined to represent the retrieved avatar expression, the messaging client application 104 immediately modifies the avatar expression. In this way, as the user types words of the text string in the caption, the messaging client application 104 dynamically adjusts and modifies the expression of the avatar to represent the mood or context of the message in the text string. For example, if the user enters additional words after a positive or happy word and after the avatar expression has been modified, the avatar expression may change based on the context of the additional words. That is, the additional words can be determined to be associated with an excited context. In response, the messaging client application 104 modifies the expression of the avatar from an expression representing positive or happy to an expression representing excited.
[0027] In some embodiments, the messaging client application 104 may determine that one or more words in a text string represent another user of the messaging application. For example, one or more words may specifically identify the username of another user, or may include an attribute uniquely associated with another user. In response, the messaging client application 104 retrieves the second avatar of the other user and presents the second avatar together with the first avatar of the user who composed the message. In such a case, the expressions of the two avatars may be modified to represent the context of the message written by the user. In some cases, the expression corresponding to the first avatar of the user who composed the text string may be modified to represent a first context in the message associated with the user, and the expression corresponding to the second avatar of the other user mentioned in the text string may be modified to represent a second context. For example, a user may enter the string "I am happy, but John Smith is sad". In such a case, the expression of the user's avatar may be modified to represent the context of happiness, and the avatar of the other user John Smith may be modified to represent the context of sadness.
[0028] In some cases, the messaging client application 104 determines that the user has activated a camera device to capture an image or video from a conversation with one or more other users. In such a case, any image or video and the message consisting of the image or video may be automatically directed to one or more other users with whom the user is having a conversation. In such a case, when the user selects the captioned avatar option, the messaging client application 104 may automatically retrieve the avatars of each user participating in the conversation. The messaging client application 104 may modify the expressions of all the avatars being presented based on the context of the captions or text string entered by the user.
[0029] After the user finishes composing the caption using the avatar, the messaging client application 104 may receive an input from the user selecting a send option. In response, the messaging client application 104 sends the image or video captured by the user together with the avatar with captions that enhance the image or video to one or more specified recipients. The specified recipients may be manually entered by the user after composing the message. Alternatively, when the camera device for capturing an image or video including a captioned avatar is activated from a conversation with a conversation member, the specified recipients may be automatically populated to include all members of the conversation.
[0030] The messaging server system 108 provides server-side functionality to a specific messaging client application 104 via a network 106. Although certain functions of the messaging system 100 are described herein as being performed by the messaging client application 104 or by the messaging server system 108, it should be understood that the location of certain functions within the messaging client application 104 or the messaging server system 108 is a design choice. For example, it is technically preferred that certain technologies and functions can be initially deployed within the messaging server system 108, but later migrated to the messaging client application 104 on the client device 102 that has sufficient processing power.
[0031] The messaging server system 108 supports various services and operations provided to the messaging client application 104. Such operations include sending data to the messaging client application 104, receiving data from the messaging client application 104, and processing data generated by the messaging client application 104. The data can include, by way of example, message content, client device information, graphical elements, geographical location information, media annotations and overlays, virtual objects, message content persistence conditions, social network information, and life event information. Data exchange within the messaging system 100 is activated and controlled through functions available via a user interface (UI) (e.g., a graphical user interface) of the messaging client application 104.
[0032] Turning now specifically to the messaging server system 108, an API server 110 is coupled to an application server 112 and provides a programming interface to the application server 112. The application server 112 is communicatively coupled to a database server 118, and the database server 118 facilitates access to a database 120 in which data associated with messages processed by the application server 112 is stored.
[0033] Specifically handle the API server 110, which receives and sends message data (e.g., commands and message payloads) between the client device 102 and the application server 112. Specifically, the API server 110 provides a set of interfaces (e.g., routines and protocols) that can be called or queried by the messaging client application 104 and third-party applications 105 to activate the functions of the application server 112. The API server 110 exposes various functions supported by the application server 112, including: account registration; login functionality; sending messages from a specific messaging client application 104 to another messaging client application 104 or third-party application 105 via the application server 112; sending media files (e.g., graphical elements, images, or videos) from the messaging client application 104 to the messaging server application 114 and making them potentially accessible to another messaging client application 104 or third-party application 105; a list of graphical elements; setting up a collection of media data (e.g., stories); retrieving such a collection; retrieving a friend list of the user of the client device 102; retrieving messages and content; adding and deleting friends in the social graph; the location of friends in the social graph; accessing user conversation data; accessing avatar information stored on the messaging server system 108; and opening application events (e.g., related to the messaging client application 104).
[0034] The application server 112 hosts several applications and subsystems, including the messaging server application 114, the image processing system 116, the social network system 122, and the context-sensitive avatar system 124. The messaging server application 114 implements several message processing techniques and functions, particularly those related to the aggregation and other processing of content (e.g., text and multimedia content) included in messages received from multiple instances of the messaging client application 104. As will be described in further detail, text and media content from multiple sources can be aggregated into content collections (e.g., called stories or galleries). These collections are then made available to the messaging client application 104 by the messaging server application 114. Given the hardware requirements for such processing, other processor- and memory-intensive processing of data can also be performed by the messaging server application 114 on the server side.
[0035] The application server 112 also includes an image processing system 116, which is dedicated to performing various image processing operations typically on images or videos received within the payload of messages at the messaging server application 114. A part of the image processing system 116 can also be implemented by the context-sensitive avatar system 124.
[0036] The social networking system 122 supports various social networking functions and services and makes these functions and services available to the messaging server application 114. To this end, the social networking system 122 maintains and accesses an entity graph within the database 120. Examples of functions and services supported by the social networking system 122 include identifying other users of the messaging system 100 with whom a particular user has a relationship or "follows", and identifying other entities and interests of a particular user. Such other users may be referred to as friends of the user. The social networking system 122 can access location information associated with each of the user's friends to determine where they live or their current geographical location. The social networking system 122 can maintain a location profile for each of the user's friends indicating the geographical locations where the user's friends live.
[0037] The context-sensitive avatar system 124 dynamically modifies the expression of an avatar based on the context of the caption associated with the avatar. For example, the context-sensitive avatar system 124 receives a user's selection of a captioned avatar option. The context-sensitive avatar system 124 presents the user's avatar and a text entry area. As the user types words in the text entry area, the context-sensitive avatar system 124 determines the context of one or more words in the text entry area. In some cases, the context-sensitive avatar system 124 determines multiple contexts associated with one or more words in the text entry area. The context-sensitive avatar system 124 ranks the multiple contexts based on relevance and selects a given context associated with the highest ranking. The context-sensitive avatar system 124 retrieves the avatar expression associated with the given context and modifies the expression of the avatar to represent the retrieved avatar expression. In some cases, the context-sensitive avatar system 124 presents multiple captioned avatars and modifies the expression of all avatars based on the context of the captions.
[0038] The application server 112 is communicatively coupled to a database server 118 that facilitates access to a database 120 in which data associated with messages processed by the message server application 114 is stored. The database 120 can be a third-party database. For example, the application server 112 can be associated with a first entity, and the database 120 or a portion of the database 120 can be associated with and hosted by a second different entity. In some implementations, the database 120 stores user data collected by the first entity about each of the users of the services provided by the first entity. For example, user data includes username, phone number, password, address, friends, activity information, preferences, videos or content consumed by the user, etc.
[0039] Figure 2FIG. 200 is a schematic diagram showing data that may be stored in database 120 of messaging server system 108 according to some example embodiments. Although the contents of database 120 are shown as including a number of tables, it should be understood that the data may be stored in other types of data structures (e.g., as an object-oriented database).
[0040] Database 120 includes message data stored within message table 214. Entity table 202 stores entity data, including entity graph 204. Entities for which records are maintained within entity table 202 may include individuals, corporate entities, organizations, objects, locations, events, etc. Any entity for which the messaging server system 108 stores data about it may be an identified entity, regardless of type. Each entity is provided with a unique identifier, as well as an entity type identifier (not shown).
[0041] Entity graph 204 stores information about the relationships and associations between entities. By way of example only, such relationships may be social, professional (e.g., working in the same company or organization), interest-based, or activity-based.
[0042] Message table 214 may store a collection of conversations between a user and one or more friends or entities. Message table 214 may include various attributes of each conversation, such as a list of participants, the size of the conversation (e.g., the number of users and / or the number of messages), the chat color of the conversation, a unique identifier for the conversation, and any other conversation-related characteristics.
[0043] The database 120 also stores annotation data in the annotation table 212 in the form of examples of filters. The database 120 also stores the received annotation content in the annotation table 212. The filters for storing data in the annotation table 212 are associated with videos and / or images and are applied to videos (whose data is stored in the video table 210) and / or images (whose data is stored in the image table 208). In one example, the filter is an overlay that is displayed as an overlay on an image or video during presentation to the receiving user. Filters can be of various types, including filters selected by the user from a filter gallery presented to the sending user by the messaging client application 104 when the sending user is composing a message. Other types of filters include location-based filters (also known as geo-filters), which can be presented to the sending user based on the geographical location. For example, location-based filters specific to nearby or special locations can be presented within the UI by the messaging client application 104 based on geographical location information determined by the global positioning system (GPS) unit of the client device 102. Another type of filter is a data filter, which can be selectively presented to the sending user by the messaging client application 104 based on other inputs or information collected by the client device 102 during the message creation process. Examples of data filters include the current temperature at a specific location, the current speed at which the sending user is traveling, the battery life of the client device 102, or the current time.
[0044] Other annotation data that can be stored within the image table 208 is augmented reality data or LENSES. Augmented reality data can be real-time special effects and sounds that can be added to an image or video.
[0045] As described above, terms such as LENSES, overlays, image transforms, AR images, and the like refer to modifications that can be made to video or images. This includes real-time modifications, which modify an image as it is captured using a device sensor and then displayed on the device's screen in the modified state. This also includes modifications to stored content, such as video clips in a gallery that can be modified. For example, in a device with access to multiple LENSES, a user can use a single video clip with multiple LENSES to see how different LENSES would modify the stored clip. For example, by selecting different LENSES for the same content, multiple LENSES applying different pseudo-random motion models can be applied to the same content. Similarly, real-time video capture can be used with the illustrated modifications to show how the video image currently captured by the device's sensor would modify the captured data. Such data can be simply displayed on the screen without being stored in memory, or the content captured by the device sensor can be recorded and stored in memory with or without modification (or both). In some systems, a preview function can show how different LENSES would look in different windows on the display. For example, this can enable multiple windows with different pseudo-random animations to be viewed simultaneously on the display.
[0046] Data using LENSES and various systems or other such transform systems that use this data to modify content can thus involve: detection of objects (e.g., faces, hands, bodies, cats, dogs, surfaces, objects, etc.); tracking of such objects as they leave the field of view in a video frame, enter the field of view in a video frame, and move around in the field of view in a video frame; and modification or transformation of such objects as they are being tracked. In various embodiments, different methods can be used to implement such transformations. For example, some embodiments can involve generating a three-dimensional mesh model of one or more objects and using the transformation and animated textures of the model within the video to implement the transformation. In other embodiments, tracking of points on an object can be used to place an image or texture (which can be two-dimensional or three-dimensional) at the tracked location. In further embodiments, neural network analysis of video frames can be used to place an image, model, or texture in the content (e.g., an image or video frame). Thus, LENS data refers both to the images, models, and textures used to create transformations in content and to the additional modeling and analysis information required to implement such transformations using object detection, tracking, and placement.
[0047] Real-time video processing can be performed using any kind of video data (e.g., video streams, video files, etc.) stored in the memory of any kind of computerized system. For example, a user can load a video file and save it in the device's memory, or can use the device's sensors to generate a video stream. Additionally, computer animation models can be used to process any object, such as parts of a human face and body, animals, or inanimate objects such as chairs, cars, or other objects.
[0048] In some embodiments, when a particular modification is selected along with the content to be transformed, the elements to be transformed are identified by a computing device and then detected and tracked if they are present in the frames of the video. The elements of the object are modified according to the modification request, thereby transforming the frames of the video stream. For different kinds of transformations, the transformation of the frames of the video stream can be performed by different methods. For example, for a frame transformation that mainly refers to the changing form of the elements of an object, the feature points of each element of the object (e.g., using an Active Shape Model (ASM) or other known methods) are calculated. Then, a grid based on the feature points is generated for each of at least one element of the object. This grid is used for the subsequent stage of tracking the elements of the object in the video stream. During the tracking process, the mentioned grid for each element is aligned with the position of each element. Then, additional points are generated on the grid. A first set of first points is generated for each element based on the modification request, and a set of second points is generated for each element based on this set of first points and the modification request. Then, the frames of the video stream can be transformed by modifying the elements of the object based on this set of first points, this set of second points, and the grid. In such a method, the background of the modified object can also be changed or distorted by tracking and modifying the background.
[0049] In one or more embodiments, a transformation of changing some regions of an object using the elements of the object can be performed by calculating the feature points of each element of the object and generating a grid based on the calculated feature points. Points are generated on the grid, and then various regions are generated based on these points. Then, the elements of the object are tracked by aligning the regions of each element with the position of each of at least one element, and the attributes of the regions can be modified based on the modification request, thereby transforming the frames of the video stream. Depending on the specific modification request, the attributes of the mentioned regions can be transformed in different ways. Such modifications can involve: changing the color of the region; removing at least part of the region from the frames of the video stream; including one or more new objects in the region based on the modification request; and modifying or distorting the region or the elements of the object. In various embodiments, any combination of such modifications or other similar modifications can be used. For some models to be animated, some feature points can be selected as control points for determining the entire state space of the options for model animation.
[0050] In some embodiments of a computer animation model that uses face detection to transform image data, a specific face detection algorithm (e.g., Viola-Jones) is used to detect faces in an image. Then, an Active Shape Model (ASM) algorithm is applied to the face region of the image to detect facial feature reference points.
[0051] In other embodiments, other methods and algorithms suitable for face detection can be used. For example, in some embodiments, landmarks are used to locate features, which represent distinguishable points that exist in most images under consideration. For example, for facial landmarks, the location of the left eye pupil can be used. If the initial landmark is not recognizable (e.g., if a person has an eye patch), secondary landmarks can be used. Such a landmark recognition process can be used for any such object. In some embodiments, a set of landmarks forms a shape. The shape can be represented as a vector using the coordinates of the points in the shape. One shape is aligned with another using a similarity transformation (allowing translation, scaling, and rotation) that minimizes the average Euclidean distance between the shape points. The average shape is the average of the aligned training shapes.
[0052] In some embodiments, a search for landmarks begins from an average shape aligned with the position and size of the face determined by a full-face detector. Then, such a search repeats the steps of: suggesting a tentative shape by adjusting the positioning of the shape points through template matching of the image texture around each point; and conforming the tentative shape to a global shape model until convergence occurs. In some systems, individual template matches are unreliable, and the shape model pools the results of weak template matches to form a stronger overall classifier. The entire search is repeated at each level of an image pyramid from a coarse resolution to a fine resolution.
[0053] Embodiments of the transformation system can capture an image or video stream on a client device and perform complex image manipulations locally on the client device, such as client device 102, while maintaining an appropriate user experience, computation time, and power consumption. Complex image manipulations can include size and shape changes, emotion transfer (e.g., changing a face from a frown to a smile), state transfer (e.g., aging a subject, reducing apparent age, changing gender), style transfer, application of graphical elements, and any other suitable image or video manipulations implemented by a convolutional neural network that has been configured to execute efficiently on the client device.
[0054] In some example embodiments, a computer animation model for transforming image data can be used by a system in which a user can use a client device 102 having a neural network to capture an image or video stream of the user (e.g., a selfie), the neural network operating as part of a messaging client 104 operating on the client device 102. A transformation system operating within the messaging client 104 determines the presence of a face within the image or video stream and provides a modification icon associated with the computer animation model to transform the image data, or the computer animation model can be presented as being associated with an interface described herein. The modification icon includes a change that can be the basis for modifying the user's face within the image or video stream as part of a modification operation. Once the modification icon is selected, the transformation system initiates a process of transforming the user's image to reflect the selected modification icon (e.g., generating a smiling face on the user). In some embodiments, once the image or video stream is captured and a specified modification is selected, the modified image or video stream can be presented in a graphical user interface displayed on the mobile client device. The transformation system can implement a complex convolutional neural network on a portion of the image or video stream to generate and apply the selected modification. That is, once the modification icon is selected, the user can capture an image or video stream and the result of the modification can be presented in real time or near real time. Additionally, when a video stream is being captured, the modification can be persistent and the selected modification icon remains toggled. A machine-taught neural network can be used to implement such modifications.
[0055] In some embodiments, a graphical user interface that presents the modifications performed by the transformation system can supply additional interaction options to the user. Such options can be based on the interface used to initiate content capture and selection for a particular computer animation model (e.g., an initiation from a content creator user interface). In various embodiments, the modification can be persistent after an initial selection of the modification icon. The user can turn the modification on or off and store it for later viewing or browsing to other areas of the imaging application by tapping or otherwise selecting the face modified by the transformation system. In the case where multiple faces are modified by the transformation system, the user can globally turn the modification on or off by tapping or selecting an individual face modified and displayed within the graphical user interface. In some embodiments, each face within a group of multiple faces can be modified individually, or such modifications can be toggled individually by tapping or selecting each individual face or a series of individual faces displayed within the graphical user interface.
[0056] As described above, the video table 210 stores video data which, in one embodiment, is associated with a message whose record is maintained in the message table 214. Similarly, the image table 208 stores image data associated with a message whose message data is stored in the entity table 202. The entity table 202 can associate various annotations from the annotation table 212 with the various images and videos stored in the image table 208 and the video table 210.
[0057] The avatar expression list 207 stores a list of different avatar expressions associated with different contexts. For example, the avatar expression list 207 can store different avatar textures, each associated with a different context. The avatar textures can be retrieved and used to modify the avatars of one or more users to represent the context associated with the avatar texture.
[0058] The context list 209 stores a list of different contexts associated with different words or combinations of words. The avatar expression list 207 stores a list of rules used by the context-sensitive avatar system 124 to process text strings to obtain or determine the context of the text strings.
[0059] The story table 206 stores data related to a collection of messages and associated image, video, or audio data, which are compiled into a collection (e.g., a story or a gallery). The creation of a particular collection can be initiated by a particular user (e.g., each user for whom a record is maintained in the entity table 202). A user can create a "personal story" in the form of a collection of content that has been created and sent / broadcast by that user. To this end, the UI of the messaging client application 104 can include user-selectable icons to enable the sender user to add specific content to his or her personal story.
[0060] The collection can also constitute a "life story", which is a collection of content from multiple users that is created manually, automatically, or using a combination of manual and automatic techniques. For example, a "life story" can constitute a curated stream of user-submitted content from various locations and events. Users whose client devices have location services enabled and are at a common location event at a particular time can, for example, be presented with options via the UI of the messaging client application 104 to contribute content to a particular life story. The messaging client application 104 can identify a life story to a user based on the user's location. The end result is a "life story" told from a community perspective.
[0061] Another type of content collection is referred to as a "location story" that enables users of client device 102 located within a specific geographical location (e.g., on a college or university campus) to contribute to a specific collection. In some embodiments, contributing to a location story may require secondary authentication to verify that the end user belongs to a specific organization or other entity (e.g., is a student on a university campus).
[0062] Figure 3 FIG. 4 is a schematic diagram showing the structure of a message 300 according to some embodiments, which is generated by a messaging client application 104 for transmission to another messaging client application 104 or a messaging server application 114. The content of a particular message 300 is used to populate a message table 214 stored in a database 120, which can be accessed by the message server application 114. Similarly, the content of the message 300 is stored in memory as "in-transit" or "in-flight" data of the client device 102 or the application server 112. The message 300 is shown as including the following components:
[0063] · Message identifier 302: A unique identifier that identifies the message 300.
[0064] · Message text payload 304: The text to be generated by the user via the UI of the client device 102 and included in the message 300.
[0065] · Message image payload 306: Image data captured by a camera device component of the client device 102 or retrieved from the memory of the client device 102 and included in the message 300.
[0066] · Message video payload 308: Video data captured by a camera device component or retrieved from a memory component of the client device 102 and included in the message 300.
[0067] · Message audio payload 310: Audio data captured by a microphone or retrieved from a memory component of the client device 102 and included in the message 300.
[0068] · Message annotation 312: Annotation data (e.g., filters, stickers, or other enhancements) representing an annotation to be applied to the message image payload 306, the message video payload 308, or the message audio payload 310 of the message 300.
[0069] · Message duration parameter 314: A parameter value that indicates the amount of time in seconds that the content of the message (e.g., the message image payload 306, the message video payload 308, the message audio payload 310) will be presented to the user or made accessible to the user via the messaging client application 104.
[0070] · Message geographical location parameter 316: Geographical location data (e.g., latitude and longitude coordinates) associated with the content payload of the message. Multiple message geographical location parameter 316 values may be included in the payload, where each of these parameter values is associated with a content item included in the content (e.g., a specific image within the message image payload 306, or a specific video within the message video payload 308).
[0071] · Message story identifier 318: An identifier value that identifies one or more content collections (e.g., a "story") with which a specific content item within the message image payload 306 of message 300 is associated. For example, multiple images within the message image payload 306 may each be associated with multiple content collections using the identifier value.
[0072] · Message tag 320: Each message 300 may be tagged with multiple tags, each tag indicating a topic of the content included in the message payload. For example, in the case where a specific image included in the message image payload 306 depicts an animal (e.g., a lion), a tag value indicating the relevant animal may be included in the message tag 320. The tag values may be generated manually based on user input or may be generated automatically using, for example, image recognition.
[0073] · Message sender identifier 322: An identifier (e.g., a messaging system identifier, an email address, or a device identifier) of the user of the client device 102 on which the message 300 is generated and from which the message 300 is sent.
[0074] · Message recipient identifier 324: An identifier (e.g., a messaging system identifier, an email address, or a device identifier) of the user of the client device 102 for which the message 300 is addressed. In the case of a conversation between multiple users, the identifier may indicate each user involved in the conversation.
[0075] The content (e.g., values) of the various components of message 300 may be pointers to the locations of stored content data values in a table. For example, the image value in the message image payload 306 may be a pointer to a location (or address) within the image table 208. Similarly, the value within the message video payload 308 may point to data stored within the video table 210, the value stored in the message annotation 312 may point to data stored in the annotation table 212, the value stored in the message story identifier 318 may point to data stored in the story table 206, and the values stored in the message sender identifier 322 and the message recipient identifier 324 may point to user records stored in the entity table 202.
[0076] Figure 4 is a block diagram showing an example context - sensitive avatar system 124 according to an example embodiment. The context - sensitive avatar system 124 includes an avatar caption module 414, a context analysis module 416, and an avatar expression modification module 418.
[0077] The user activates the image capture component of the messaging client application 104. In response, the avatar caption module 414 presents an image or video captured by the front or rear camera device of the client device 102 on the display screen of the client device 102.
[0078] The avatar caption module 414 presents one or more editing tools to the user to modify the image or video presented to the user on the screen of the client device 102. The editing tools allow the user to select one or more graphical elements (e.g., avatars, text, emojis, images, videos, etc.) to add to or enhance the image or video presented to the user. For example, the user can add text to the image or video presented on the display screen at a location selected by the user. In some cases, the editing tools allow the user to select the option of a captioned avatar.
[0079] In response to receiving the user's selection of the captioned avatar option, the avatar caption module 414 determines whether the user is participating in a conversation with one or more other users. That is, the avatar caption module 414 determines whether an image or video has been captured in response to the user selecting the option to reply using the camera device from a conversation with another user. In response to determining that the option to reply using the camera device has been selected from the conversation to capture an image, the avatar caption module 414 presents the avatars of each user with whom the user is participating in the conversation. The avatars are presented with a caption entry area above or below the avatar. In response to determining that the user is not currently in contact with another member of the conversation (e.g., has not selected the option to reply using the camera device to capture an image), the avatar caption module 414 presents a single avatar with a caption entry area above or below the avatar. In both cases (single or multiple avatar presentations), as the user types words of a text string for the caption entry area, the text entered by the user into the caption entry area wraps around the avatar in a circular manner.
[0080] The context analysis module 416 processes one or more words in the text string input by the user. The context analysis module 416 analyzes one or more words as the user types the words, referring to the words and / or rules stored in the context list 209. The context analysis module 416 generates a context list of the words typed by the user and sorts the generated context list. The context analysis module 416 selects a given context associated with the highest ranking from the generated list.
[0081] In some cases, the context analysis module 416 determines which words in a string are associated with which presented avatar. The context analysis module 416 can perform semantic analysis and processing to determine which words are associated with a particular avatar based on the sentence structure. For example, if the text string is "I am happy while John Smith is sad", the context analysis module 416 determines that happy is associated with the user who is entering the expression, and sad is associated with the avatar of John Smith. In such a case, the context analysis module 416 retrieves and selects the context of each avatar associated with a particular word. Alternatively, the context analysis module 416 associates all avatars with the same top-ranked context from the context list.
[0082] The avatar expression modification module 418 receives one or more contexts selected by the context analysis module 416. The avatar expression modification module 418 retrieves one or more textures associated with the one or more selected contexts from the avatar expression list 207. For example, if the selected context is sad, the avatar expression modification module 418 retrieves the avatar texture associated with the sad context. For example, if the selected contexts are sad and happy, the avatar expression modification module 418 retrieves the avatar textures associated with the sad and happy contexts.
[0083] As the user types a text string in the caption entry area, the avatar expression modification module 418 uses the retrieved textures to modify the expressions of one or more avatars. The avatar expression modification module 418 can receive an indication from the context analysis module 416 identifying which avatar among the presented avatars is associated with which context. For example, when the context analysis module 416 determines that the first avatar is associated with the word corresponding to the first context and the second avatar is associated with the word corresponding to the second context, the context analysis module 416 provides this information to the avatar expression modification module 418. The avatar expression modification module 418 then uses the first texture associated with the first context to modify the expression of the first avatar, and uses the second texture associated with the second context to modify the expression of the second avatar.
[0084] The context analysis module 416 continuously and dynamically processes the words entered by the user in the caption entry area to continuously and dynamically rank and select the context that best fits the caption. When the context analysis module 416 selects a different context, the context analysis module 416 notifies the avatar expression modification module 418 to modify the expressions of the avatars based on the different or newly selected context. In this way, as the user types the words of the string in the caption entry area, the expressions of one or more avatars change to represent the context or the string.
[0085] Figure 5FIG. 0 is a flowchart showing an example operation of the context-sensitive avatar system 124 in the execution process 500 according to an example embodiment. The process 500 may be implemented by computer-readable instructions executed by one or more processors such that the operations of the process 500 may be performed in part or in whole by the functional components of the messaging server system 108 and / or the third-party application 105; thus, the process 500 is described below by way of example. However, in other embodiments, at least some of the operations of the process 500 may be deployed on various other hardware configurations. Thus, the process 500 is not intended to be limited to the messaging server system 108 and may be implemented in whole or in part by any other component. Some or all of the operations of the process 500 may be performed in parallel, out of order, or entirely omitted.
[0086] At operation 501, the context-sensitive avatar system 124 receives an input that selects an option for generating a message using a captioned avatar. For example, the context-sensitive avatar system 124 determines that the user has captured an image using the camera device of the mobile device and has entered a request to add text to the image. In response, the context-sensitive avatar system 124 presents a menu of text insertion options that includes the captioned avatar option. The context-sensitive avatar system 124 receives the user's selection of the captioned avatar option from the menu.
[0087] At operation 502, the context-sensitive avatar system 124 presents an avatar and a caption entry area adjacent to the avatar. For example, the context-sensitive avatar system 124 retrieves an avatar with a neutral expression associated with the user. The context-sensitive avatar system 124 presents a cursor above or below the avatar that enables the user to enter a text string that includes one or more words that wrap around the avatar in a circular manner.
[0088] At operation 503, the context-sensitive avatar system 124 populates the caption entry area with a text string that includes one or more words. For example, the context-sensitive avatar system 124 receives one or more words from the user who enters characters on the keypad.
[0089] At operation 504, the context-sensitive avatar system 124 determines the context based on one or more words in the text string. For example, the context-sensitive avatar system 124 processes the words as the user enters them to determine whether the words are associated with a positive or negative context or a happy or sad context.
[0090] At operation 505, the context-sensitive avatar system 124 modifies the expression of the avatar based on the determined context. For example, the context-sensitive avatar system 124 retrieves the texture of the avatar associated with the determined context and uses the retrieved texture to modify the expression of the avatar. That is, if the context is determined to be sad, the context-sensitive avatar system 124 modifies the expression of the avatar from neutral to sad.
[0091] Figures 6 to 8 is an illustrative input and output of the context-sensitive avatar system 124 according to an example embodiment. For example, as Figure 6 shown, after the user uses the camera device of the user device to capture an image, the user selects an option to add text to the image. In response, the context-sensitive avatar system 124 presents a text option menu 630 with multiple text entry types. For example, the first text entry type is the "@" type, which allows the user to add text with a reference to another user or topic. The second text entry type is a large text entry type that allows the user to input text and add the text to the image in a very large font or style. The third text entry type is the captioned avatar option 632, which allows the user to add an avatar and enter text in a caption entry area near the avatar. The fourth text entry type is the rainbow text entry type, which allows the user to input text with a color attribute. The fifth text entry type is the script type that allows the user to input text with a script attribute.
[0092] As shown in screen 601, the user selects the captioned avatar option 632. In response, the context-sensitive avatar system 124 retrieves the avatar 610 associated with the user and presents the avatar 610 together with the caption entry area 620. The context-sensitive avatar system 124 allows the user to type text, and as shown in screen 602, the string 622 entered by the user wraps around the avatar 612 in a circular manner. In addition, the context-sensitive avatar system 124 determines that the context in the string entered by the user is happy. For example, the context-sensitive avatar system 124 determines that the word "happy" is in the string 622 and that the word has a happy context. In response, the context-sensitive avatar system 124 modifies the expression of the avatar 610 to have a happy expression as shown in the avatar 612. That is, the avatar 610 in screen 601 may have a neutral expression, and the avatar 612 may have a happy expression to represent the context of the string 622 entered by the user.
[0093] Screen 603 shows another example where the string entered by the user is determined to have a sad context. For example, the string can include the word "worst" which is determined by the context-sensitive avatar system 124 to have a sad context. In response, the expression of avatar 610 is modified to have a sad expression as shown in screen 603. The user can select the send option to specify the recipient of the message, which includes an image captured by the user, enhanced with a captioned avatar having the modified expression.
[0094] As Figure 7 shown, the user types a string that includes a symbol 710 (e.g., "@") that references a topic or a user. Screen 701 includes the string with symbol 710 that references the username of another user. In response, the context-sensitive avatar system 124 retrieves a second avatar associated with the username of the other user being referenced. The context-sensitive avatar system 124 presents two avatars 720 (one representing the user who wrote the string and the other representing the second user). The avatars 720 are surrounded by the string with symbol 710 that wraps around the avatars 720 in a circular manner. The avatars 720 are modified to have an expression corresponding to the context of the string entered by the user. The user can select the send option to specify the recipient of the message, which includes an image captured by the user, enhanced with a captioned avatar 720 including symbol 710.
[0095] In some embodiments, the user selects the option to reply using the camera device from a conversation with another user. In response to the user selecting the option to reply using the camera device, the messaging client application 104 activates the camera device and allows the user to capture an image or video. After the user captures the image or video, the user can enhance the image or video with text by selecting the text entry option from the text tool menu. In response to the user selecting the captioned avatar option, the context-sensitive avatar system 124 determines that the user captured an image when the user selected the option to reply using the camera device and the user is participating in a conversation with one or more other users. In such a case, the context-sensitive avatar system 124 retrieves the avatars of each user with whom the user is participating in the conversation and presents all the avatars along with a caption entry area on screen 801. The user can enter a string in the caption entry area and the string wraps around the avatars in a circular manner. The expression of one or more avatars is modified based on the context of the string entered by the user. In some cases, the context-sensitive avatar system 124 modifies the expression of the avatars by analyzing one or more words in the conversation (e.g., words previously exchanged between the users) and / or by analyzing one or more words in the string entered by the user for the caption (e.g., words not previously exchanged between the users).
[0096] In some embodiments, the initial expression of the avatar presented when a user selects the captioned avatar option can be determined based on one or more words in the last message exchanged between users. The initial expression of the avatar is then modified based on one or more words of the string entered by the user in the caption entry area. For example, a user may be engaged in a conversation with another user, John. The last message the user sent to John or received from John may be "I am very happy today". The context-sensitive avatar system 124 can determine that the context of this message is happiness. In response to receiving the user's selection of the captioned avatar option (after selecting the option to reply using the camera device and capturing an image), the context-sensitive avatar system 124 can present avatars with a happy expression for the user and John. Then, the user can enter the string "Today is not a good day" into the caption entry area. The context-sensitive avatar system 124 can determine that this string is associated with a sad context, and in response, can modify the expressions of the user's and John's avatars from happy to sad. Then, the user can select the send option to send a message to all users involved in the conversation. The message includes the image captured by the user, the avatar with the modified expression, and the caption with the string entered by the user that surrounds the avatar in a circular manner. The recipients of the message are automatically selected based on the identities of the conversation members.
[0097] Figure 9 is a block diagram showing an example software architecture 906 that can be used in conjunction with the various hardware architectures described herein. Figure 9 is a non-limiting example of a software architecture, and it should be understood that many other architectures can be implemented to facilitate the functions described herein. The software architecture 906 can execute on hardware such as Figure 10 a machine 1000 that includes a processor 1004, a memory 1014, and input / output (I / O) components 1018, etc. A representative hardware layer 952 is shown and can represent, for example, Figure 10 the machine 1000. The representative hardware layer 952 includes a processing unit 954 with associated executable instructions 904. The executable instructions 904 represent the executable instructions of the software architecture 906, including the implementation of the methods, components, etc. described herein. The hardware layer 952 also includes a memory and / or storage module memory / storage device 956 that also has executable instructions 904. The hardware layer 952 may also include other hardware 958.
[0098] In Figure 9In the example architecture, the software architecture 906 can be conceptualized as a stack of layers, where each layer provides a specific function. For example, the software architecture 906 can include layers such as the operating system 902, libraries 920, framework / middleware 918, applications 916, and presentation layer 914. In operation, the applications 916 and / or other components within a layer can activate API calls 908 through the software stack and receive messages 912 in response to the API calls 908. The layers shown are representative in nature, and not all software architectures have all layers. For example, some mobile operating systems or specialized operating systems may not provide framework / middleware 918, while other operating systems may provide such a layer. Other software architectures can include additional layers or different layers.
[0099] The operating system 902 can manage hardware resources and provide common services. The operating system 902 can include, for example, a kernel 922, services 924, and drivers 926. The kernel 922 can act as an abstraction layer between the hardware and other software layers. For example, the kernel 922 can be responsible for memory management, processor management (e.g., scheduling), component management, networking, security settings, etc. The services 924 can provide other common services for other software layers. The drivers 926 are responsible for controlling or interfacing with the underlying hardware. For example, depending on the hardware configuration, the drivers 926 include a display driver, a camera device driver, a driver, a flash memory driver, a serial communication driver (e.g., a Universal Serial Bus (USB) driver), a driver, an audio driver, a power management driver, etc.
[0100] Library 920 provides a common infrastructure used by application 916 and / or other components and / or layers. Library 920 provides functions that allow other software components to perform tasks in an easier way than directly interfacing with the functions of the underlying operating system 902 (e.g., core 922, services 924, and / or drivers 926). Library 920 may include system libraries 944 (e.g., C standard library), and system libraries 944 may provide functions such as memory allocation functions, string manipulation functions, mathematical functions, etc. In addition, Library 920 may include API libraries 946, such as media libraries (e.g., libraries that support the rendering and manipulation of various media formats such as MPREG4, H.264, MP3, AAC, AMR, JPG, PNG), graphics libraries (e.g., OpenGL framework that can be used to render two-dimensional and three-dimensional graphical content on a display), database libraries (e.g., SQLite that can provide various relational database functions), web libraries (e.g., WebKit that can provide web browsing functions), etc. Library 920 may also include a variety of other libraries 948 to provide many other APIs to application 916 and other software components / modules.
[0101] Framework / middleware 918 (sometimes also referred to as middleware) provides a higher-level common infrastructure that can be used by application 916 and / or other software components / modules. For example, framework / middleware 918 may provide various graphical user interface functions, advanced resource management, advanced location services, etc. Framework / middleware 918 may provide a wide range of other APIs that can be utilized by application 916 and / or other software components / modules, some of which may be specific to a particular operating system 902 or platform.
[0102] Application 916 includes built-in applications 938 and / or third-party applications 940. Examples of representative built-in applications 938 may include, but are not limited to: contact applications, browser applications, e-book reader applications, location applications, media applications, messaging applications, and / or gaming applications. Third-party applications 940 may include applications developed using the ANDROID TM or IOS TM software development kit (SDK), and may be mobile software running on a mobile operating system such as IOS TM 、ANDROID TM 、 、Phone or other mobile operating systems. Third-party applications 940 may activate API calls 908 provided by the mobile operating system (e.g., operating system 902) to facilitate the functions described herein.
[0103] Application 916 can use built-in operating system functions (e.g., kernel 922, services 924, and / or drivers 926), libraries 920, and frameworks / middleware 918 to create a UI for interacting with the users of the system. Alternatively or additionally, in some systems, interaction with the users can occur through a presentation layer such as presentation layer 914. In these systems, the application / component "logic" can be separated from the aspects of the application / component that interact with the users.
[0104] Figure 10 is a block diagram showing the components of a machine 1000 capable of reading instructions from a machine-readable medium (e.g., a machine-readable storage medium) and performing any one or more of the methods discussed herein. Specifically, Figure 10 shows a schematic illustration of a machine 1000 in the example form of a computer system, in which instructions 1010 (e.g., software, program, application, applet, app, or other executable code) can be executed to cause the machine 1000 to perform any one or more of the methods discussed herein. Thus, the instructions 1010 can be used to implement the modules or components described herein. The instructions 1010 transform the general, non-programmed machine 1000 into a particular machine 1000 programmed to perform the described and shown functions in the described manner. In an alternative embodiment, the machine 1000 operates as a stand-alone device or can be coupled (e.g., networked) to other machines. In a networked deployment, the machine 1000 can operate as a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 1000 can include, but is not limited to: server computers, client computers, personal computers (PCs), tablet computers, laptop computers, netbooks, set-top boxes (STBs), personal digital assistants (PDAs), entertainment media systems, cellular telephones, smart phones, mobile devices, wearable devices (e.g., smart watches), smart home devices (e.g., smart appliances), other smart devices, web appliances, network routers, network switches, network bridges, or any machine capable of sequentially or otherwise executing the instructions 1010 specifying the actions to be taken by the machine 1000. Further, although only a single machine 1000 is shown, the term "machine" shall also be taken to include a collection of machines that individually or jointly execute the instructions 1010 to perform any one or more of the methods discussed herein.
[0105] Machine 1000 may include a processor 1004, a memory / storage 1006, and I / O components 1018, which may be configured to communicate with each other via, for example, a bus 1002. In an example embodiment, the processor 1004 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, a processor 1008 and a processor 1012 that may execute instructions 1010. The term "processor" is intended to include a multi-core processor 1004, which may include two or more independent processors (sometimes referred to as "cores") that may execute instructions simultaneously. Although Figure 10 a multi-processor 1004 is shown, machine 1000 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.
[0106] The memory / storage 1006 may include a memory 1014 such as a main memory or other memory storage device, and a storage unit 1016, both of which are accessible by the processor 1004 via, for example, the bus 1002. The storage unit 1016 and the memory 1014 store instructions 1010 that embody any one or more of the methods or functions described herein. The instructions 1010 may also reside, completely or partially, within the memory 1014, within the storage unit 1016, within at least one of the processors 1004 (e.g., within a cache memory of the processor), or within any suitable combination thereof during execution by the machine 1000. Accordingly, the memory 1014, the storage unit 1016, and the memory of the processor 1004 are examples of machine-readable media.
[0107] The I / O components 1018 may include various components for receiving input, providing output, generating output, sending information, exchanging information, capturing measurement results, and so on. The specific I / O components 1018 included in a particular machine 1000 will depend on the type of the machine. For example, a portable machine such as a mobile phone will likely include a touch input device or other such input mechanism, while a headless server machine will likely not include such a touch input device. It will be understood that the I / O components 1018 may include Figure 10Many other components not shown. The I / O components 1018 are grouped according to function only to simplify the following discussion, and this grouping is in no way restrictive. In various example embodiments, the I / O components 1018 may include output components 1026 and input components 1028. The output components 1026 may include visual components (e.g., displays such as plasma display panels (PDPs), light-emitting diode (LED) displays, liquid crystal displays (LCDs), projectors, or cathode ray tubes (CRTs)), auditory components (e.g., speakers), tactile components (e.g., vibration motors, resistance mechanisms), other signal generators, etc. The input components 1028 may include alphanumeric input components (e.g., keyboards, touchscreens configured to receive alphanumeric input, optical keyboards, or other alphanumeric input components), point-based input components (e.g., mice, touchpads, trackballs, joysticks, motion sensors, or other pointing instruments), tactile input components (e.g., physical buttons, touchscreens that provide the location and / or force of a touch or touch gesture, or other tactile input components), audio input components (e.g., microphones), etc.
[0108] In other example embodiments, the I / O components 1018 may include biometric components 1039, motion components 1034, environmental components 1036, or location components 1038 along with a wide variety of other components. For example, the biometric components 1039 may include components for detecting expressions (e.g., hand expressions, facial expressions, voice expressions, body postures, or eye tracking), measuring biometric signals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), identifying people (e.g., voice recognition, retina recognition, facial recognition, fingerprint recognition, or electroencephalogram-based recognition), etc. The motion components 1034 may include: acceleration sensor components (e.g., accelerometers), gravity sensor components, rotation sensor components (e.g., gyroscopes), etc. The environmental components 1036 may include, for example, lighting sensor components (e.g., photometers), temperature sensor components (e.g., one or more thermometers that detect the ambient temperature), humidity sensor components, pressure sensor components (e.g., barometers), auditory sensor components (e.g., one or more microphones that detect background noise), proximity sensor components (e.g., infrared sensors that detect nearby objects), gas sensors (e.g., gas detection sensors that detect the concentration of hazardous gases for safety or measure pollutants in the atmosphere), or other components that may provide an indication, measurement, or signal corresponding to the surrounding physical environment. The location components 1038 may include location sensor components (e.g., GPS receiver components), altitude sensor components (e.g., altimeters or barometers that detect the air pressure from which altitude can be obtained), orientation sensor components (e.g., magnetometers), etc.
[0109] A variety of techniques can be used to implement communication. The I / O component 1018 may include a communication component 1040 that is operable to couple the machine 1000 to the network 1037 or the device 1029 via the couplings 1024 and 1222, respectively. For example, the communication component 1040 may include a network interface component or other suitable device to interface with the network 1037. In another example, the communication component 1040 may include a wired communication component, a wireless communication component, a cellular communication component, a near field communication (NFC) component, components (e.g., low power), components, and other communication components that provide communication via other modalities. The device 1029 may be any one of other machines or various peripheral devices (e.g., a peripheral device coupled via USB).
[0110] In addition, the communication component 1040 may detect an identifier or include components operable to detect an identifier. For example, the communication component 1040 may include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional barcodes such as Universal Product Code (UPC) barcodes, multi-dimensional barcodes such as Quick Response (QR) codes, Aztec codes, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D barcodes, and other optical codes), or an auditory detection component (e.g., a microphone for identifying an audio signal of a tag). In addition, various information can be obtained via the communication component 1040, such as a location via Internet Protocol (IP) geolocation, a location via signal triangulation, a location via detecting an NFC beacon signal that may indicate a specific location, etc.
[0111] Glossary:
[0112] A "carrier signal" in this context refers to any non-tangible medium that is capable of storing, encoding, or carrying transient or non-transient instructions executed by a machine and includes digital or analog communication signals or other non-tangible media for facilitating the communication of such instructions. Instructions can be sent or received over a network via a network interface device using a transient or non-transient transmission medium and using any one of a plurality of well-known transmission protocols.
[0113] A "client device" in this context refers to any machine that interfaces with a communication network to obtain resources from one or more server systems or other client devices. A client device can be, but is not limited to, a mobile phone, desktop computer, laptop computer, PDA, smartphone, tablet computer, ultrabook, netbook, laptop, multi-processor system, microprocessor-based or programmable consumer electronics, game console, set-top box, or any other communication device that a user can use to access the network.
[0114] A "communication network" in this context refers to one or more portions of a network, which can be an ad hoc network, intranet, extranet, virtual private network (VPN), local area network (LAN), wireless LAN (WLAN), wide area network (WAN), wireless WAN (WWAN), metropolitan area network (MAN), the Internet, a portion of the Internet, a portion of the public switched telephone network (PSTN), a plain old telephone service (POTS) network, a cellular telephone network, a wireless network, a network, another type of network, or a combination of two or more such networks. For example, a network or a portion of a network can include a wireless network or a cellular network, and the coupling can be a code division multiple access (CDMA) connection, a global system for mobile communications (GSM) connection, or other types of cellular or wireless couplings. In this example, the coupling can implement any one of various types of data transfer technologies, such as single carrier radio transmission technology (1xRTT), evolution data optimized (EVDO) technology, general packet radio service (GPRS) technology, enhanced data rate GSM evolution (EDGE) technology, the 3rd Generation Partnership Project (3GPP) including 3G, 4th generation wireless (4G) networks, universal mobile telecommunications system (UMTS), high speed packet access (HSPA), worldwide interoperability for microwave access (WiMAX), long term evolution (LTE) standards, other standards defined by various standards setting organizations, other remote protocols, or other data transfer technologies.
[0115] A "transient message" in this context refers to a message that can be accessed within a limited duration of time. A transient message can be text, image, video, etc. The access time of a transient message can be set by the message sender. Alternatively, the access time can be a default setting or a setting specified by the recipient. Regardless of the setting technique, the message is temporary.
[0116] "Machine-readable medium" in this context refers to a component, device, or other tangible medium that can store instructions and data temporarily or permanently, and can include, but is not limited to, random access memory (RAM), read-only memory (ROM), buffer memory, flash memory, optical media, magnetic media, cache memory, other types of storage devices (e.g., erasable programmable read-only memory (EEPROM)), and / or any suitable combination thereof. The term "machine-readable medium" should be regarded as including a single medium or multiple media that can store instructions (e.g., a centralized or distributed database or associated cache and server). The term "machine-readable medium" will also be regarded as including any combination of media or multiple media that can store instructions executable by a machine (e.g., code) such that the instructions, when executed by one or more processors of the machine, cause the machine to perform any one or more of the methods described herein. Thus, "machine-readable medium" refers to a single storage device or apparatus, as well as "cloud"-based storage systems or storage networks that include multiple storage devices or apparatuses. The term "machine-readable medium" does not include the signal itself.
[0117] "Component" in this context refers to a device, physical entity, or logic having boundaries defined by a function or subroutine call, branch point, API, or other technical definition that provides partitioning or modularization of a particular processing or control function. Components can be combined with other components via their interfaces to perform machine processing. A component can be an encapsulated functional hardware unit designed to be used with other components and can be part of a program that typically performs a particular function among related functions. Components can constitute software components (e.g., code embodied on a machine-readable medium) or hardware components. A "hardware component" is a tangible unit capable of performing certain operations and can be configured or arranged in some physical manner. In various example embodiments, one or more computer systems (e.g., a stand-alone computer system, a client computer system, or a server computer system) or one or more hardware components of a computer system (e.g., a processor or group of processors) can be configured by software (e.g., an application or part of an application) to be a hardware component for performing some of the operations described herein.
[0118] Hardware components can be implemented mechanically, electronically, or in any suitable combination thereof. For example, a hardware component can include dedicated circuitry or logic that is permanently configured to perform certain operations. A hardware component can be a dedicated processor, such as a field programmable gate array (FPGA) or an ASIC. A hardware component can also include programmable logic or circuitry that is temporarily configured by software to perform certain operations. For example, a hardware component can include software executed by a general purpose processor or other programmable processor. Once configured by such software, the hardware component becomes a particular machine (or a particular component of a machine) uniquely customized to perform the configured functions, rather than a general purpose processor. It will be appreciated that the decision to implement a hardware component mechanically in dedicated and permanently configured circuitry or in temporarily configured (e.g., software-configured) circuitry can be made based on cost and time considerations. Accordingly, the phrase "hardware component" (or "hardware-implemented component") should be understood to include a tangible entity, i.e., an entity that is physically constructed, permanently configured (e.g., hardwired) or temporarily configured (e.g., programmed) to operate in some manner or to perform certain operations described herein. Considering implementations in which a hardware component is temporarily configured (e.g., programmed), there is no need to configure or instantiate each hardware component in the hardware component at any given time. For example, in a case where a hardware component includes a general purpose processor that is configured by software to be a dedicated processor, the general purpose processor can be configured as different dedicated processors (e.g., including different hardware components) at different times. The software accordingly configures one or more particular processors to, for example, constitute a particular hardware component at one time and a different hardware component at a different time.
[0119] Hardware components can provide information to, and receive information from, other hardware components. Accordingly, the described hardware components can be considered to be communicatively coupled. In cases where multiple hardware components are present simultaneously, communication can be achieved via signal transmission (e.g., via appropriate circuitry and buses) between or among two or more hardware components. In implementations in which multiple hardware components are configured or instantiated at different times, communication between such hardware components can be achieved, for example, by storing information in a memory structure accessed by the multiple hardware components and retrieving the information from the memory structure. For example, one hardware component can perform an operation and store the output of the operation in a memory device to which it is communicatively coupled. Then, another hardware component can access the memory device at a subsequent time to retrieve the stored output and process it.
[0120] The hardware component can also initiate communication with an input device or an output device and can operate on resources (e.g., collection of information). The various operations of the example methods described herein can be performed, at least in part, by one or more processors that are temporarily configured (e.g., by software) or permanently configured to perform the associated operations. Whether temporarily or permanently configured, such processors can constitute processor-implemented components that operate to perform one or more of the operations or functions described herein. As used herein, a "processor-implemented component" refers to a hardware component implemented using one or more processors. Similarly, the methods described herein can be at least in part processor-implemented, where a particular one or more processors are examples of hardware. For example, at least some of the operations of the method can be performed by one or more processors or processor-implemented components. Additionally, one or more processors can also operate to support the execution of associated operations in a "cloud computing" environment or operate as "software as a service" (SaaS). For example, at least some of the operations can be performed by a group of computers (as an example of machines including processors), where the operations can be accessed via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., APIs). The execution of certain operations can be distributed among the processors, not residing only within a single machine but deployed across multiple machines. In some example embodiments, the processor or processor-implemented components can be located in a single geographical location (e.g., within a home environment, an office environment, or a server farm). In other example embodiments, the processor or processor-implemented components can be distributed across several geographical locations.
[0121] A "processor" in this context refers to any circuit or virtual circuit (a physical circuit simulated by logic executed on an actual processor) that manipulates data values according to control signals (e.g., "commands", "opcodes", "machine codes", etc.) and produces corresponding output signals used to operate a machine. For example, a processor can be a CPU, a RISC processor, a CISC processor, a GPU, a DSP, an ASIC, an RFIC, or any combination thereof. A processor can also be a multi-core processor having two or more independent processors (sometimes referred to as "cores") that can execute instructions simultaneously.
[0122] A "timestamp" in this context refers to a sequence of characters or encoded information that identifies when an event occurred, e.g., so as to give a date and time of day, sometimes down to a fraction of a second.
[0123] Without departing from the scope of the present disclosure, changes and modifications can be made to the disclosed embodiments. These and other changes or modifications are intended to be included within the scope of the present disclosure and are expressed in the appended claims.
Claims
1. A method for avatar subtitles based on context, comprising: receiving, by one or more processors of a messaging application, an input that selects an option for generating a message using a captioned avatar; presenting, by the messaging application, a first avatar and a caption entry area near the first avatar, the first avatar being associated with a first user; populating, by the messaging application, the caption entry area with a text string including one or more words input by the first user; determining, by the messaging application, a context based on the one or more words in the text string, wherein determining the context includes: detecting a symbol in the text string that indicates a reference to a second user of the messaging application; and in response to detecting the symbol indicating a reference to the second user, adding a second avatar associated with the second user for display with the first avatar; and modifying, by the messaging application, an expression of the first avatar based on the determined context.
2. The method according to claim 1, wherein, the context is determined as each of the one or more words in the text string is input by the first user, and wherein the expression of the first avatar changes as the context is determined.
3. The method according to claim 2, further comprising: receiving an input from the first user including a first word of the one or more words; determining that the first word corresponds to a first context; and causing the first avatar to present a first expression corresponding to the first context.
4. The method according to claim 3, further comprising: after receiving an input from the first user including the first word, receiving an input from the first user including a second word of the one or more words; determining that the second word corresponds to a second context different from the first context; and changing the expression of the first avatar from the first expression to a second expression corresponding to the second context.
5. The method according to claim 1, wherein, the expression is at least one of happy, sad, excited, confused, or worried.
6. The method according to claim 1, further comprising: storing a database that associates different contexts with corresponding expressions; and retrieving, from the database, an expression associated with the determined context.
7. The method according to claim 1, further comprising: receiving a request to access a caption tool; and in response to the request, presenting a menu including a plurality of text entry types, wherein one of the plurality of text entry types includes the text entry type of the captioned avatar.
8. The method according to claim 1, further comprising sending the message to the second user, the message including the first avatar with the modified expression, the second avatar, and an image or video captured by the first user.
9. The method according to claim 1, wherein, The subtitle input area is presented above or below the first avatar and the second avatar, and wherein, as the first user inputs one or more words in the text string, the subtitle input area bends around the first avatar and the second avatar.
10. The method according to claim 1, further comprising: When receiving the input, determining that the first user is currently participating in a conversation with a third user; In response to determining that the first user is currently participating in a conversation with the third user, retrieving a third avatar associated with the third user, with whom the first user is currently participating in a conversation; Presenting the first avatar associated with the first user and the third avatar associated with the third user; and Presenting the subtitle input area near the first avatar and the third avatar.
11. The method according to claim 1, further comprising modifying the expressions of the first avatar and the second avatar based on the determined context.
12. The method according to claim 10, further comprising: Presenting a conversation interface including a plurality of messages exchanged between the first user and the third user; Receiving an input for selecting a camera device option from the conversation interface; Enabling the first user to capture an image using a camera device in response to receiving the input for selecting the camera device option; and Receiving an input for selecting an option to generate a message after the image is captured.
13. The method according to claim 1, further comprising: Determining the context of the conversation by processing one or more words in the conversation; and Modifying the expressions of the first avatar and the second avatar based on the one or more words in the conversation.
14. The method according to claim 10, further comprising: Determining a first context associated with the first user based on a first word among the one or more words in the text string; Determining a second context associated with the third user based on a second word among the one or more words in the text string; and Presenting the first avatar having a first expression corresponding to the first context together with the third avatar having a second expression corresponding to the second context.
15. The method according to claim 1, further comprising: Identifying a plurality of contexts based on the one or more words in the text string; Sorting the plurality of contexts based on relevance; and Selecting the context associated with the highest ranking among the plurality of contexts as the determined context.
16. A system for avatar subtitles based on context, comprising: A processor configured to perform operations including the following: Receiving, by a messaging application, an input that selects an option for generating a message using a subtitled avatar; Presenting, by the messaging application, a first avatar and a subtitle input area near the first avatar, the first avatar being associated with a first user; The messaging application fills the caption entry area with a text string including one or more words input by the first user; The messaging application determines context based on the one or more words in the text string, wherein determining the context includes: Detecting symbols in the text string that indicate a reference to a second user of the messaging application; and In response to detecting the symbols indicating a reference to the second user, adding a second avatar associated with the second user for display together with the first avatar; and The messaging application modifies the expression of the first avatar based on the determined context.
17. The system according to claim 16, wherein, The context is determined as each of the one or more words in the text string is input by the first user, and wherein the expression of the first avatar changes as the context is determined.
18. A non-transitory machine-readable storage medium including instructions that, when executed by one or more processors of a machine, cause the machine to perform operations including the following: The messaging application receives an input that selects an option to use captioned avatars to generate a message; The messaging application presents a first avatar and a caption entry area near the first avatar, the first avatar being associated with a first user; The messaging application fills the caption entry area with a text string including one or more words input by the first user; The messaging application determines context based on the one or more words in the text string, wherein determining the context includes: Detecting symbols in the text string that indicate a reference to a second user of the messaging application; and In response to detecting the symbols indicating a reference to the second user, adding a second avatar associated with the second user for display together with the first avatar; and The messaging application modifies the expression of the first avatar based on the determined context.
19. The non-transitory machine-readable storage medium according to claim 18, wherein, The context is determined as each of the one or more words in the text string is input by the first user, and wherein the expression of the first avatar changes as the context is determined.
Citation Information
Patent Citations
Avatars Reflecting User States
US20170357417A1