Chatting with micro sound clips

By linking audio clips with messages in messaging applications, the problem of inefficient image or avatar selection for users is solved, achieving faster selection and saving resources.

CN116349215BActive Publication Date: 2026-04-28SNAP INC
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SNAP INC
Filing Date
2021-09-21
Publication Date
2026-04-28

AI Technical Summary

Technical Problem

When users choose the appropriate image or avatar to convey a specific idea, existing technologies require them to manually browse multiple information pages, resulting in inefficiency and wasted resources.

Method used

By associating audio clips with messages in messaging applications, users can quickly select sounds or audio clips relevant to the message context and share them with others, reducing browsing steps and device resource consumption.

Benefits of technology

It improves the efficiency of user device usage, reduces the time spent browsing and selecting suitable images or avatars, and saves device resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116349215B_ABST
    Figure CN116349215B_ABST
Patent Text Reader

Abstract

Aspects of the disclosure relate to systems and methods for performing operations including receiving, by a messaging application implemented on a client device of a first user, an input comprising a message for a second user, obtaining, in response to receiving the input, contextual information associated with the message, identifying, based on the contextual information, a plurality of sounds representative of the message, receiving, from the first user, a selection of a given sound of the plurality of sounds, and transmitting, in response to receiving the selection of the given sound from the first user, the given sound to the second user.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority requirements

[0002] This application claims the benefit of priority to U.S. Patent Application No. 16 / 948,479, filed September 21, 2020, the entire contents of which are incorporated herein by reference. Technical Field

[0003] This disclosure generally relates to recording sound using messaging applications. Background Technology

[0004] The popularity of online interaction between users continues to grow. There are numerous ways for users to interact with other users online. Users can communicate with their friends using messaging apps, play multiplayer video games online with other users, or use various other applications to perform other actions. Attached Figure Description

[0005] In the accompanying drawings (which are not necessarily drawn to scale), similar reference numerals may describe similar parts in different views. To facilitate identification of any particular element or action being discussed, one or more of the highest-order digits in the reference numerals indicate the drawing number in which the element was first introduced. Some non-limiting examples are shown in the figures of the accompanying drawings, in which:

[0006] Figure 1 It is a graphical representation of a networked environment in which the present disclosure can be deployed, based on some examples.

[0007] Figure 2 It is a graphical representation of a messaging system with both client-side and server-side functionalities, based on some examples.

[0008] Figure 3 It is a graphical representation of the data structures maintained in the database based on some examples.

[0009] Figure 4 It is a graphical representation based on some example messages.

[0010] Figures 5 to 9 It is a graphical representation based on some examples of graphical user interfaces.

[0011] Figure 10 This is a flowchart illustrating example operations of a messaging application according to an example implementation.

[0012] Figure 11 It is a graphical representation of a machine in the form of a computer system, based on some examples, within which a set of instructions can be executed to cause the machine to perform any or more of the methods discussed herein.

[0013] Figure 12 It is a block diagram showing a software architecture in which examples can be implemented. Detailed Implementation

[0014] The following description includes systems, methods, techniques, instruction sequences, and computer program products embodying illustrative embodiments of this disclosure. In the following description, numerous specific details are set forth for illustrative purposes to provide an understanding of various embodiments. However, it will be apparent to those skilled in the art that embodiments can be practiced without these specific details. Generally, well-known examples of instructions, protocols, structures, and techniques need not be shown in detail.

[0015] Typically, users use messaging applications to exchange messages. Such applications allow users to choose from a predefined list of images and avatars to send messages to each other. Users increasingly use these images and avatars to communicate their thoughts. However, finding the right image or avatar to convey a specific idea can be tedious and time-consuming. Specifically, users have to manually search for a specific image or avatar to convey a given message using keywords. This requires navigating multiple pages of information until the desired image or avatar is found. Given the complexity and time commitment of finding the right image or avatar, users become discouraged from communicating using images or avatars, leading to wasted resources or underutilization.

[0016] The disclosed implementation improves the efficiency of using electronic devices by providing a system that allows users to associate sound or audio clips with messages shared with another user. When a receiving user accesses a message that has been associated with sound, the message is presented to the recipient and the sound is played back. To assist users in associating sound or audio clips with messages, a user interface is presented that intelligently and quickly displays the sound context-dependent on the message the user is typing or composing. Users can select one of the context-dependent sound or audio clips to preview the selected sound or audio clip or send it to other users.

[0017] Specifically, according to the disclosed implementation, a messaging application implemented on a first user's client device receives input including a message for a second user. In response to receiving the input, the messaging application obtains context information associated with the message and identifies multiple sounds representing the message based on the context information. The messaging application receives a selection of a given sound from the multiple sounds from the first user, and in response to receiving the selection of the given sound from the first user, sends the given sound to the second user.

[0018] In this way, the disclosed implementation improves the efficiency of using electronic devices by reducing the number of screens and interfaces that a user must navigate to find the sounds to share with other users in order to send messages to them. This is achieved by presenting a graphical user interface from which the user can quickly and efficiently select the sounds to be combined with the messages the user intends to send to other users. This reduces the device resources (e.g., processor cycles, memory, and power consumption) required to complete the task.

[0019] Networked computing environment

[0020] Figure 1 This is a block diagram illustrating an example messaging system 100 for exchanging data (e.g., messages and associated content) over a network. The messaging system 100 includes multiple instances of client devices 102, each hosting multiple applications including a messaging client 104 and other external applications 109 (e.g., third-party applications). Each messaging client 104 is communicatively coupled via a network 112 (e.g., the Internet) to (e.g., hosted on corresponding other client devices 102) other instances of the messaging client 104, a messaging server system 108, and an external app server 110. The messaging client 104 can also communicate with the locally hosted third-party applications 109 using an application programming interface (API).

[0021] The messaging client 104 can communicate and exchange data with other messaging clients 104 and messaging server system 108 via network 112. The data exchanged between messaging clients 104 and between messaging clients 104 and messaging server system 108 includes functions (e.g., commands for activating functions) and payload data (e.g., text, audio, video, or other multimedia data).

[0022] Message transceiver server system 108 provides server-side functionality to specific message transceiver clients 104 via network 112. While some functions of message transceiver system 100 are described herein as being performed by message transceiver client 104 or message transceiver server system 108, the location of certain functions within message transceiver client 104 or message transceiver server system 108 can be a design choice. For example, it may be technically preferred that certain technologies and functions are initially deployed within message transceiver server system 108, but later migrated to message transceiver client 104 with sufficient processing power on client device 102.

[0023] The messaging server system 108 supports various services and operations provided to the messaging client 104. Such operations include sending data to and receiving data from the messaging client 104, and processing data generated by the messaging client 104. As an example, this data may include message content, client device information, geolocation information, media enhancements and overlays, message content persistence conditions, social network information, and live event information. Data exchange within the messaging system 100 is activated and controlled through functions available via the user interface (UI) of the messaging client 104.

[0024] Specifically, turning to message transceiver server system 108, application programming interface (API) server 116 is coupled to application server 114 and provides a programming interface to application server 114. Application server 114 is communicatively coupled to database server 120, which facilitates access to database 126, which stores data associated with messages processed by application server 114. Similarly, web server 128 is coupled to application server 114 and provides a web-based interface to application server 114. For this purpose, web server 128 handles incoming network requests via Hypertext Transfer Protocol (HTTP) and several other related protocols.

[0025] Application Programming Interface (API) server 116 receives and sends message data (e.g., commands and message payloads) between client device 102 and application server 114. Specifically, API server 116 provides a set of interfaces (e.g., routines and protocols) that can be invoked or queried by messaging client 104 to activate the functionality of application server 114. Application Programming Interface (API) server 116 exposes various functions supported by application server 114, including: account registration; login functionality; sending messages from one messaging client 104 to another messaging client 104 via application server 114; sending media files (e.g., images or videos) from messaging client 104 to messaging server 118, and possible access for another messaging client 104; setting media data sets (e.g., stories); retrieving the friend list of the user of client device 102; retrieving such a set; retrieving messages and content; adding and deleting entities (e.g., friends) in an entity graph (e.g., a social graph); locating friends within the social graph; and opening application events (e.g., related to messaging client 104).

[0026] Application server 114 hosts multiple server applications and subsystems, including, for example, messaging server 118, image processing server 122, and social networking server 124. Messaging server 118 implements multiple messaging technologies and functions, particularly relating to the aggregation and other processing of content (e.g., text and multimedia content) included in messages received from multiple instances of messaging client 104. As will be described in further detail, text and media content from multiple sources can be aggregated into collections of content (e.g., referred to as stories or galleries). These collections are then made available to messaging client 104. Given the hardware requirements for additional processor- and memory-intensive processing of data, such processing can also be performed on the server side by messaging server 118.

[0027] Application server 114 also includes image processing server 122, which is dedicated to performing various image processing operations on images or videos typically within the payload of messages sent from or received at message transceiver server 118.

[0028] Social network server 124 supports various social networking functions and services and makes these functions and services available to messaging server 118. To this end, social network server 124 maintains and accesses entity graph 308 (such as...) within database 126. Figure 3 (As shown). Examples of the functions and services supported by the social networking server 124 include identifying other users in the messaging system 100 who are related to a particular user or who are being "followed" by that particular user, as well as identifying the interests and other entities of a particular user.

[0029] Returning to messaging client 104, the features and functionalities of the external resource (e.g., a third-party application 109 or an applet) are made available to the user via the interface of messaging client 104. Messaging client 104 receives user selections regarding options for launching or accessing features of the external resource (e.g., a third-party resource), such as external app 109. The external resource can be a third-party application (external app 109) installed on client device 102 (e.g., a “native app”), or a smaller version (e.g., an “app”) of a third-party application hosted on client device 102 or remotely on client device 102 (e.g., hosted on a third-party server 110). The smaller version of the third-party application comprises a subset of the features and functionalities of the third-party application (e.g., a full-scale native version of a third-party standalone application) and is implemented using markup language documentation. In one example, the smaller version of the third-party application (e.g., an “app”) is a web-based markup language version of the third-party application and is embedded in messaging client 104. In addition to using markup language documents (e.g., .*ml files), mini-programs can incorporate scripting languages ​​(e.g., .*js files or .json files) and style sheets (e.g., .*ss files).

[0030] In response to receiving a user selection of an option for launching or accessing an external resource (external app 109), messaging client 104 determines whether the selected external resource is a web-based external resource or a locally installed external application. In some cases, external application 109, locally installed on client device 102, can be launched independently of messaging client 104 and separately from messaging client 104, for example, by selecting an icon corresponding to external application 109 on the home screen of client device 102. A smaller version of such an external application can be launched or accessed via messaging client 104, and in some examples, no part of the smaller external application can be accessed (or only a limited part can be accessed) outside of messaging client 104. The smaller external application can be launched by receiving and processing a markup language document associated with the smaller external application from external app server 110 via messaging client 104.

[0031] In response to determining that the external resource is a locally installed external application 109, the messaging client 104 instructs the client device 102 to launch the external application 109 by executing locally stored code corresponding to the external application 109. In response to determining that the external resource is a web-based resource, the messaging client 104 communicates with the external app server 110 to obtain a markup language document corresponding to the selected resource. The messaging client 104 then processes the obtained markup language document to present the web-based external resource within the user interface of the messaging client 104.

[0032] The messaging client 104 can notify users of client device 102 or other users (e.g., "friends") associated with such users of one or more external resources of ongoing activity. For example, the messaging client 104 can provide participants in a conversation (e.g., a chat session) within the messaging client 104 with notifications related to the current or recent use of external resources by one or more members of a user group. One or more users can be invited to join a valid external resource or to activate a recently used but currently invalid external resource (within the group of friends). External resources can provide participants in the conversation, each using the corresponding messaging client 104, with the ability to share items, conditions, states, or locations of the external resource with one or more members of the user group who have entered the chat session. Shared items can be interactive chat cards that chat members can use to interact with, for example, activate the corresponding external resource, view specific information within the external resource, or take chat members to a specific location or state within the external resource. Within a given external resource, response messages can be sent to users on the messaging client 104. External resources can selectively include different media items in the response based on the current context of the external resource.

[0033] The messaging client 104 can present a list of available external resources (e.g., third-party or external applications 109 or mini-programs) to the user to launch or access a given external resource. This list can be presented as a context-sensitive menu. For example, the icons representing different external applications within external application 109 (or mini-program) can vary based on how the user launches the menu (e.g., from a conversational interface or a non-conversational interface).

[0034] System Architecture

[0035] Figure 2This is a block diagram illustrating further details of a messaging system 100 according to some examples. Specifically, the messaging system 100 is shown as including a messaging client 104 and an application server 114. The messaging system 100 includes multiple subsystems supported on the client side by the messaging client 104 and on the server side by the application server 114. These subsystems include, for example, a short-lived timer system 202, a collection management system 204, an enhancement system 208, a map system 210, a game system 212, and an external resource system 220.

[0036] The short-lived timer system 202 is responsible for implementing temporary or time-limited access to content by the message sending client 104 and the message sending server 118. The short-lived timer system 202 incorporates multiple timers that selectively implement access to (e.g., for rendering and displaying) messages and associated content via the message sending client 104 based on duration and display parameters associated with a message or set of messages (e.g., a story). Further details regarding the operation of the short-lived timer system 202 are provided below.

[0037] The collection management system 204 is responsible for managing collections or sets of media (e.g., collections of text, images, video, and audio data). Collections of content (e.g., messages, including images, videos, text, and audio) can be organized into "event galleries" or "event stories." Such collections can be made available for a specified time period (e.g., the duration of an event related to the content). For example, content related to a concert can be made available as a "story" for the duration of the concert. The collection management system 204 can also be responsible for publishing icons that provide notifications of the existence of specific collections to the user interface of the messaging client 104.

[0038] Furthermore, the collection management system 204 includes a curation interface 206, which enables the collection manager to manage and curate specific collections of content. For example, the curation interface 206 allows an event organizer to curate a collection of content related to a specific event (e.g., removing inappropriate content or redundant messages). Additionally, the collection management system 204 employs machine vision (or image recognition technology) and content rules to automatically curate content collections. In some examples, users may be compensated for including user-generated content in the collection. In such cases, the collection management system 204 operates to automatically pay such users for using their content.

[0039] Enhancement system 208 provides various functionalities that enable users to enhance (e.g., annotate or otherwise modify or edit) media content associated with a message. For example, enhancement system 208 provides functionalities related to generating and publishing media overlays for messages processed by messaging system 100. Enhancement system 208 can operable to provide media overlays or enhancements (e.g., image filters) to messaging client 104 based on the geolocation of client device 102. In another example, enhancement system 208 can operable to provide media overlays to messaging client 104 based on other information such as the social network information of the user of client device 102. Media overlays can include audio and visual content as well as visual effects. Examples of audio and visual content include pictures, text, logos, animations, and sound effects. Examples of visual effects include color overlays. Audio and visual content or visual effects can be applied to media content items (e.g., photos) at client device 102. For example, media overlays can include text or images that can be overlaid on a photograph taken by client device 102. In another example, media overlays include location identifier overlays (e.g., Venice Beach), live event names, or business name overlays (e.g., Beach Cafe). In yet another example, enhancement system 208 uses the geolocation of client device 102 to identify media overlays that include the business name at the location of client device 102. Media overlays may include additional tags associated with the business. Media overlays may be stored in database 126 and accessed through database server 120.

[0040] In some examples, enhancement system 208 provides a user-based publishing platform that allows users to select geographic locations on a map and upload content associated with those locations. Users can also specify scenarios where specific media overlays should be provided to other users. Enhancement system 208 generates a media overlay that includes the uploaded content and associates it with the selected geographic location.

[0041] In other examples, enhancement system 208 provides a merchant-based publishing platform that enables merchants to select specific media overlays associated with geographic locations via a bidding process. For example, enhancement system 208 associates the media overlay of the highest bidder with a corresponding geographic location for a predefined amount of time.

[0042] Map system 210 provides various geolocation functions and supports the presentation of map-based media content and messages by messaging client 104. For example, map system 210 enables the display (e.g., stored in profile data 316) of user icons or avatars on a map to indicate the current or past locations of the user's "friends" in the context of the map, as well as media content generated by these friends (e.g., a collection of messages including photos and videos). For example, a message posted by a user from a specific geolocation to messaging system 100 can be displayed to the specific user's "friends" in the context of that specific location on the map on the map interface of messaging client 104. A user can also share his or her location and status information with other users of messaging system 100 via messaging client 104 (e.g., using an appropriate status avatar), where the location and status information is similarly displayed to selected users in the context of the map interface of messaging client 104.

[0043] Game system 212 provides various game functions within the context of messaging client 104. Messaging client 104 provides a game interface that offers a list of available games (e.g., web-based games or web-based applications) that can be started by a user within the context of messaging client 104 and played with other users of messaging system 100. Messaging system 100 also enables specific users to invite other users to participate in a specific game by sending invitations from messaging client 104. Messaging client 104 also supports both voice and text messaging (e.g., chat) within the game context, provides leaderboards for the game, and also supports providing in-game rewards (e.g., game currency and items).

[0044] External resource system 220 provides messaging client 104 with an interface for communicating with external app server 110 to launch or access external resources. Each external resource (app) server 110 hosts, for example, an application based on a markup language (e.g., HTML5) or a smaller version of an external application (e.g., a game, utility, payment, or ride-sharing application external to messaging client 104). Messaging client 104 can launch a web-based resource by accessing an HTML5 file from the external resource (app) server 110 associated with the web-based resource (e.g., an application). In some examples, the application hosted by external resource server 110 is programmed in JavaScript using a software development kit (SDK) provided by messaging server 118. The SDK includes application programming interfaces (APIs) with functionality that can be invoked or activated by the web-based application. In some examples, messaging server 118 includes a JavaScript library that provides access to certain user data of messaging client 104 to a given third-party resource. HTML5 is used as an example technology for programming games, but applications and resources programmed based on other technologies can be used.

[0045] To integrate the SDK's functionality into a web-based resource, the external resource (app) server 110 downloads the SDK from the messaging server 118 or otherwise receives the SDK. Once downloaded or received, the SDK is included as part of the application code of the web-based external resource. The code of the web-based resource can then call or activate certain functions of the SDK to integrate the features of the messaging client 104 into the web-based resource.

[0046] The SDK stored on the messaging server 118 effectively bridges the gap between external resources (e.g., third-party or external applications 109 or mini-programs) and the messaging client 104. This provides users with a seamless experience communicating with other users on the messaging client 104 while preserving the look and feel of the messaging client 104. To bridge communication between external resources and the messaging client 104, in some examples, the SDK facilitates communication between the external resource server 110 and the messaging client 104. In some examples, a WebViewJavaScriptBridge running on the client device 102 establishes two unidirectional communication channels between the external resources and the messaging client 104. Messages are sent asynchronously between the external resources and the messaging client 104 via these communication channels. Each SDK function activation is sent as a message and a callback. Each SDK function is implemented by constructing a unique callback identifier and sending a message with that callback identifier.

[0047] By using the SDK, not all information from message client 104 is shared with external resource server 110. The SDK limits which information is shared based on the needs of the external resource. In some examples, each external resource server 110 provides the message client 118 with an HTML5 file corresponding to the web-based external resource. Message client 118 can add a visual representation (e.g., box art or other graphics) of the web-based external resource in message client 104. Once the user selects the visual representation or instructs message client 104 to access a feature of the web-based external resource through the message client 104's GUI, message client 104 obtains the HTML5 file and instantiates the resources required to access the feature of the web-based external resource.

[0048] The messaging client 104 presents a graphical user interface (GUI) for an external resource (e.g., a login page or title screen). During, before, or after presenting the login page or title screen, the messaging client 104 determines whether the initiated external resource has previously been authorized to access the messaging client 104's user data. In response to determining that the initiated external resource has previously been authorized to access the messaging client 104's user data, the messaging client 104 presents another GUI for the external resource, including its functionality and characteristics. In response to determining that the initiated external resource has not previously been authorized to access the messaging client 104's user data, after displaying the external resource's login page or title screen for a threshold time period (e.g., 3 seconds), the messaging client 104 slides up a menu for authorizing the external resource to access user data (e.g., animating the menu to appear from the bottom of the screen to the middle of the screen or other parts). This menu identifies the type of user data that the external resource will be authorized to use. In response to receiving a user's selection of the accept option, the messaging client 104 adds the external resource to the list of authorized external resources and allows that external resource to access user data from the messaging client 104. In some examples, the messaging client 104 authorizes external resources to access user data according to the OAuth 2 framework.

[0049] The messaging client 104 controls the type of user data shared with external resources based on the type of authorized external resource. For example, it grants access to a first type of user data (e.g., a two-dimensional avatar of a user, with or without different avatar characteristics) to external resources including full-scale external applications (e.g., a third party or external application 109). As another example, it grants access to a second type of user data (e.g., payment information, a user's two-dimensional avatar, a user's three-dimensional avatar, and avatars with various avatar characteristics) to external resources including smaller versions of external applications (e.g., a web-based version of a third-party application). Avatar characteristics include different ways of customizing the appearance and feel of an avatar, such as different poses, facial features, clothing, etc.

[0050] In some implementations, the messaging client 104 presents an interface that allows users to associate voice with messages exchanged with other users in a communication interface (e.g., a chat session of the messaging client 104). For example, such as Figure 5As shown, message transceiver client 104 presents a graphical user interface 500 to a first user of first client device 102. The graphical user interface 500 includes a communication interface 501 in which multiple messages are exchanged between the first user and one or more other users (e.g., a second user). The communication interface 501 includes a message composition section 510 (or text input area). Message transceiver client 104 receives user input by tapping the message composition section 510. In response, message transceiver client 104 presents a cursor and keyboard that allow the user to compose a message. The message entered in the message composition section 510 is transmitted or shared with one or more other users (e.g., John), where the user participates in a conversation within the communication interface 501 with said one or more other users.

[0051] As the user types message content in the message composition section 510, the message transceiver client 104 continuously or periodically analyzes keywords to determine the context of the message being entered. For example, the message transceiver client 104 searches a context database for one or more words in the message to retrieve or identify the message's context. Once the message's context is identified in the database, the message transceiver client 104 accesses a sound database. Each sound in the sound database is tagged or associated with one or more contexts. The message transceiver client 104 retrieves multiple sounds from the sound database that are associated with the context of the identified message.

[0052] In some implementations, the messaging client 104 presents a message enhancement interface including a first option 520 and a second option 530. The message enhancement interface allows the user to embed or include one or more enhancement elements with the message written by the user. In response to the messaging client 104 receiving a user selection for the first option 520, the messaging client 104 presents a sound selection area that includes a set of sounds that the user can associate with the written message. In response to the messaging client 104 receiving a user selection for the second option 530, the messaging client 104 presents a set of avatars or bitmojis that the user can associate with the written message.

[0053] In some implementations, in response to receiving a user selection of the first option 520, a sound pack selection option 540 is presented. The sound pack selection option 540 allows the user to select one or more sound packs, for which associated sounds are presented for selection. In one example, the user can select a celebrity sound pack. A celebrity sound pack includes a collection of micro-audio sound clips (e.g., audio files shorter than 3 seconds) recorded by various well-known celebrities or figures. The micro-audio sound clips may include catchphrases that the user typically associates with the corresponding celebrity or figure. In one example, the user can select a game audio sound pack. A game audio sound pack includes a collection of micro-audio sound clips of sounds typically played in various games familiar to the user who plays the game. In one example, the user can select a soundtrack sound pack. A soundtrack sound pack includes a collection of micro-audio sound clips of excerpts from famous songs. In one example, the user can select a humor sound pack. A humor sound pack includes a collection of funny micro-audio clips. In one example, the user can select an emotion sound pack. An emotion sound pack includes a collection of micro-audio clips corresponding to different emotions such as sadness, shame, embarrassment, affection, cuteness, etc.

[0054] Once the messaging client 104 receives a user selection for a sound packet, it filters the set of sounds associated with the message's context and presents graphical elements representing the filtered set of sounds to the user in the sound selection area. In another embodiment, once the messaging client 104 receives a user selection for a sound packet, the messaging client 102 searches only the sounds within the selected packet to identify one or more sounds within that packet that are associated with the context of the identified message. The identified one or more sounds are then presented along with corresponding graphical elements in the message enhancement interface.

[0055] Each sound identified and presented in the message enhancement interface includes a graphical element representing the sound (e.g., an icon, thumbnail, or image), a title for the sound, attribute information for the sound (e.g., artist or creator name), and various other relevant information (e.g., the context of the identified sound). The messaging client 104 receives a first type of user input associated with a given sound displayed in the message enhancement interface. This first type of user input may be a touch input including a tap on the graphical element representing the sound. In response to receiving the first type of user input, the messaging client 104 begins playing the audio associated with the tapped sound. The messaging client 104 presents a progress bar 522 associated with the sound being played. In one example, the progress bar 522 may be a circular element (e.g., a ring) surrounding the graphical element representing the sound. This circular element is filled with a circular progress bar that begins at a first point and ends at the same first point when the sound ends. In another embodiment, the progress bar 522 is presented adjacent to the graphical element representing the sound. The progress bar 522 travels from one point to another to indicate the playback progress of the sound. If the user taps on different graphic elements, a sound corresponding to the different graphic element is played, and a progress bar 522 is displayed near or around the different graphic element.

[0056] In one implementation, the messaging client 104 receives a second type of user input associated with a given sound displayed in the message enhancement interface. This second type of user input may be touch input including double-clicking a graphical element representing the sound (e.g., receiving two taps within a threshold time period). In response to receiving the second type of user input, the messaging client 104 associates the sound of the double-clicked graphical element with a message written in the message composition section 510. For example, the messaging client 104 sends the written message along with the sound associated with the double-clicked graphical element. In some implementations, the messaging client 104 sends the sound associated with the double-clicked graphical element instead of the message written in the message composition section 510. The message or sound is sent to one or more other users with whom the user is exchanging messages in the dialog interface 501.

[0057] like Figure 6As shown, the communication interface 600 addressed to the second user by the message sent by the first user is presented on the second client device 102 associated with the second user. The message written by the first user in the message composition section 510 is presented as message 610 in the communication interface 600. The communication interface 600 includes a message composition section 630, which allows the second user to reply to the first user with a message. In one example, in response to the first user sending a sound to the second user in a communication session, a sound icon 620 is presented according to the messages exchanged in the communication interface 600.

[0058] The sound icon 620 can be a generic icon that visually notifies a second user that a sound associated with a given message in the communication interface has been received. The messaging client 104 receives a user selection (e.g., a tap) on the sound icon 620. In response, the messaging client 104 retrieves the audio corresponding to the sound and begins playing the audio associated with the sound icon 620. Additionally, the messaging client 104 animates the sound icon 620 from a generic icon into a graphical element associated with the sound. For example, the messaging client 104 presents… Figure 8 The user interface 800 shown in the diagram has animated the sound icon 620 to be replaced by a graphical element 820. The graphical element 820 may be the same graphical element presented to the first user in the sound selection area, which can be double-clicked to associate the sound with a written message. The graphical element 820 may include a thumbnail or image representing the sound, the sound's title, attribute information, contextual information, and various other information associated with the sound. The graphical element 820 may include a play button 822. In response to a user selection of the play button 822, the messaging client 104 may replay the sound associated with the graphical element 820 once or in a loop. After the graphical element 820 has been presented for a threshold time period, the messaging client 104 may animate the graphical element 820 back to the sound icon 620. Alternatively, the messaging client 104 may maintain continuous display of the graphical element 820 after the sound icon 620 has been animated to the graphical element 820.

[0059] Return to reference Figure 7In some cases, the messaging client 104 may select from multiple different types of sound icons to display sound icon 720. Specifically, a first sound icon may correspond to the sound of a first sound pack (e.g., a game audio sound pack). A second sound icon may correspond to the voice of a celebrity. A third sound icon may correspond to the sound recorded by another user. A fourth sound icon may correspond to an emotional sound. In some implementations, the messaging client 104 determines the type of sound that the second user's second client device 102 has received. For example, the messaging client 104 determines the sound pack associated with the received sound. Based on the sound type, the messaging client 104 selects a sound icon to include in the second user's communication interface to indicate that the sound has been received and associated with message 610. That is, the messaging client 104 displays the sound icon 720 representing the type of sound received, rather than displaying a generic sound icon 620. For example, if the type of sound received is a sound recorded by a friend, the messaging client 104 may display the user's avatar or a generic avatar. If the received sound is part of an emoticon pack, the messaging client 104 displays a random or generic emoticon as the sound icon 720. If the received sound is a celebrity's voice, the messaging client 104 displays a famous sign representing the celebrity. In this way, without displaying the sound content, the second user is reminded that the sound is associated with the message 610. Furthermore, if the sound icon represents one type of sound rather than another, the second user may be more interested in selecting the sound icon 720.

[0060] In some implementations, the messaging client 104 receives user input, including touch and hold input associated with the graphical element 820. In response, the messaging client 104 presents an options window 910 associated with sound. For example, the options window 910 includes a first option 912 and a second option 914. The first option 912 is an option that allows the user to save a received sound to a sound library associated with a second user. The user can later access the sound library to view a list of previously saved sounds, modify sounds included in the sound library, and share one or more sounds in the sound library with one or more other users. The second option 914 is an option that allows the second user to edit or trim the sound for sharing with other users. Selecting the second option 914 presents a sound adjustment interface that allows the second user to adjust the start and end times of the sound to shorten or trim the sound. The second user can then respond to the message with another message or with a trimmed version of the sound.

[0061] Data Architecture

[0062] Figure 3This is a schematic diagram illustrating a data structure 300 that may be stored in a database 126 of a message transceiver server system 108, according to certain examples. Although the contents of the database 126 are shown as including multiple tables, it should be understood that data may be stored in other types of data structures (e.g., stored as an object-oriented database).

[0063] Database 126 includes message data stored in message table 302. For any given message, this message data includes at least message sender data, message receiver (or recipient) data, and a payload. See below for reference. Figure 4 Further details are provided regarding information that can be included in the message and in the message data stored in message table 302.

[0064] Entity table 306 stores entity data and (for example, links to entity diagram 308 and profile data 316). Entities whose records are stored in entity table 306 may include individuals, company entities, organizations, objects, locations, events, etc. Regardless of the entity type, any entity whose data is stored in message transceiver server system 108 can be an identifiable entity. Each entity is assigned a unique identifier and an entity type identifier (not shown).

[0065] Entity graph 308 stores information about the relationships and associations between entities. As an example only, such relationships can be social relationships based on interests or activities, or professional relationships (e.g., working in the same company or organization).

[0066] Profile data 316 stores various types of profile data about a specific entity. Based on privacy settings specified by the specific entity, profile data 316 can be selectively used and presented to other users of messaging system 100. In the case of an individual, profile data 316 includes, for example, a username, phone number, address, settings (e.g., notification and privacy settings), and an avatar representation (or a set of such avatar representations) selected by the user. A specific user can then selectively include one or more of these avatar representations in the content of messages transmitted via messaging system 100 and in a map interface displayed to other users by messaging client 104. The set of avatar representations may include “status avatars,” which present a graphical representation of a status or activity that the user can choose to transmit at a specific time.

[0067] In the case that the entity is a group, in addition to the group name, members and various settings for the relevant group (e.g., notifications), the group profile data 316 may similarly include one or more avatars associated with the group.

[0068] Database 126 also stores enhancement data, such as overlays or filters, in enhancement table 310. The enhancement data is associated with and applied to videos (video data is stored in video table 304) and images (image data is stored in image table 312).

[0069] In one example, a filter is an overlay displayed on top of an image or video during presentation to the receiving user. Filters can be of various types, including user-selected filters from a set of filters presented to the sending user by the messaging client 104 when the sending user is composing a message. Other types of filters include geolocation filters (also known as geographic filters), which can be presented to the sending user based on geolocation. For example, a geolocation filter specific to a nearby or particular location can be presented by the messaging client 104 within the user interface based on geolocation information determined by the Global Positioning System (GPS) unit of the client device 102.

[0070] Another type of filter is a data filter, which can be selectively presented to the sending user by the messaging client 104 based on other inputs or information collected by the client device 102 during the message creation process. Examples of data filters include the current temperature at a specific location, the current speed of the sending user, the battery life of the client device 102, or the current time.

[0071] Other augmented data that can be stored in image table 312 includes augmented reality content items (e.g., corresponding to an application lens or augmented reality experience). Augmented reality content items can be real-time effects and sounds that can be added to images or videos.

[0072] As described above, augmented data includes augmented reality content items, overlays, image transformations, AR images, and similar terms referring to modifications that can be applied to image data (e.g., video or images). This includes real-time modifications, which modify an image using the modifications as the device sensors (e.g., one or more cameras) of client device 102 capture the image and then display the image on the screen of client device 102. This also includes modifications to stored content, such as modifications to video clips in a gallery that can be modified. For example, in client device 102 with access to multiple augmented reality content items, a user can use a single video clip with multiple augmented reality content items to see how different augmented reality content items will modify the stored clip. For example, by selecting different augmented reality content items for the same content, multiple augmented reality content items with different pseudo-random motion models can be applied to the same content. Similarly, real-time video capture can be used with the illustrated modifications to show how the video image currently being captured by the sensors of client device 102 will modify the captured data. Such data can be displayed on the screen without being stored in memory, or the content captured by the device's sensors can be recorded and stored in memory with or without modification (or both). In some systems, the preview feature can simultaneously show different augmented reality content items displayed in different windows on the display. For example, this allows multiple windows with different pseudo-random animations to be viewed on the display at the same time.

[0073] Therefore, using augmented reality content items and various systems, or other such transformation systems that use this data to modify content, can involve: the detection of objects (e.g., faces, hands, bodies, cats, dogs, surfaces, objects, etc.) in video frames; tracking these objects as they leave, enter, and move around within the field of view; and modifying or transforming these objects while tracking them. In various examples, different methods can be used to implement such transformations. Some examples may involve generating a 3D mesh model of one or more objects, and using transformations and animated textures of the model within the video to implement the transformation. In other examples, points on the tracked objects can be used to place an image or texture (which can be two-dimensional or three-dimensional) at the tracked location. In still other examples, neural network analysis of video frames can be used to place images, models, or textures within content (e.g., images or video frames). Therefore, augmented reality content items refer both to the images, models, and textures used to create transformations within content, and to the additional modeling and analysis information required to implement such transformations through object detection, tracking, and placement.

[0074] Real-time video processing can be performed using any type of video data (e.g., video streams, video files, etc.) stored in the memory of any type of computerized system. For example, a user can load a video file and store it in the device's memory, or a video stream can be generated using the device's sensors. Furthermore, computer-animated models can be used to process any object, such as human faces and parts of the human body, animals, or inanimate objects (e.g., chairs, cars, or other objects).

[0075] In some examples, when a specific modification is selected along with the content to be transformed, the computing device identifies the element to be transformed, and then, if the element to be transformed exists in a frame of the video, detects and tracks the element to be transformed. The elements of the object are modified according to the modification request, thereby transforming the frames of the video stream. For different types of transformations, the transformation of frames of the video stream can be performed using different methods. For example, for frame transformations that primarily involve changing the form of elements of an object, feature points of each element of the object are calculated (e.g., using an Active Shape Model (ASM) or other known methods). Then, a feature point-based mesh is generated for each element of at least one element of the object. This mesh is used in a subsequent stage where the elements of the object in the video stream are tracked. In the tracking process, the aforementioned mesh for each element is aligned with the position of each element. Then, additional points are generated on the mesh. A first set of first points is generated for each element based on the modification request, and a second set of points is generated for each element based on this first set of points and the modification request. The frames of the video stream can then be transformed by modifying the elements of the object based on the first set of points, the second set of points, and the mesh. In this method, the background of the object being modified can also be changed or distorted by tracking and modifying the background.

[0076] In some examples, transformations that alter some regions of an object using its elements can be performed by calculating feature points for each element of the object and generating a mesh based on those calculated feature points. Points are generated on the mesh, and then various regions are generated based on these points. The elements of the object are then tracked by aligning the regions of each element with the positions of at least one of the elements, and the properties of the regions can be modified based on modification requests, thereby transforming frames of the video stream. Depending on the specific modification request, the properties of the mentioned regions can be transformed in different ways. Such modifications can involve: changing the color of the region; removing at least some portions of the region from the frames of the video stream; including one or more new objects in the region based on the modification request; and modifying or distorting the elements of the region or object. In various examples, any combination of such modifications or other similar modifications can be used. For certain models to be animated, some feature points can be selected as control points to determine the entire state space for options used in model animation.

[0077] In some examples of computer animation models that use face detection to transform image data, a specific face detection algorithm (e.g., Viola-Jones) is used to detect faces in the image. The Active Shape Model (ASM) algorithm is then applied to the facial regions of the image to detect facial feature reference points.

[0078] Other methods and algorithms suitable for face detection can be used. For example, in some examples, landmarks are used to locate features, which represent distinguishable points present in most of the images considered. For example, for facial landmarks, the location of the left pupil could be used. If the initial landmark is not recognizable (e.g., if the person is wearing an eye patch), secondary landmarks can be used. Such a landmark recognition process can be used for any such object. In some examples, the set of landmarks forms a shape. The shape can be represented as a vector using the coordinates of the points in the shape. One shape is aligned with another shape by a similarity transformation (allowing translation, scaling, and rotation) that minimizes the average Euclidean distance between the points of the shapes. The mean shape is the average of the aligned training shapes.

[0079] In some examples, a landmark search begins with an average shape aligned with the position and size of a face determined by a global face detector. This search then repeats the following steps: adjusting the position of shape points by template matching over the image texture around each point to suggest a provisional shape, and then conforming this provisional shape to a global shape model until convergence occurs. In some systems, individual template matching is unreliable, and the shape model pools the results of weak template matching to form a stronger overall classifier. The entire search is repeated at each level of the image pyramid, from coarse to fine resolution.

[0080] The transformation system can capture image or video streams on a client device (e.g., client device 102) and perform complex image manipulations locally on client device 102 while maintaining an appropriate user experience, computation time, and power consumption. Complex image manipulations can include size and shape changes, emotion shifts (e.g., changing a face from frowning to smiling), state shifts (e.g., aging an object, reducing its apparent age, or changing its gender), style shifts, application of graphical elements, and any other suitable image or video manipulations implemented by a convolutional neural network that has been configured to execute efficiently on client device 102.

[0081] In some examples, a computer animation model for transforming image data can be used by a system in which a user can capture an image or video stream (e.g., a selfie) using a client device 102 that operates as part of a messaging client 104 operating on client device 102. A transformation system operating within messaging client 104 determines the presence of a face in the image or video stream and provides a modification icon associated with the computer animation model to transform the image data, or the computer animation model can be presented as associated with the interface described herein. The modification icon includes changes that can be based on modifying the user's face within the image or video stream as part of a modification operation. Once a modification icon is selected, the transformation system initiates a process to transform the user's image to reflect the selected modification icon (e.g., generating a smiley face on the user). Once the image or video stream is captured and the specified modification is selected, the modified image or video stream can be presented in a graphical user interface displayed on client device 102. The transformation system can implement a complex convolutional neural network on a portion of the image or video stream to generate and apply the selected modification. That is, users can capture image or video streams, and once an edit icon is selected, the changes are presented to the user in real time or near real time. Furthermore, while a video stream is being captured, the changes can be persistent, and the selected edit icon continues to be toggled. Machine-trained neural networks can be used to implement such modifications.

[0082] A graphical user interface (GUI) presenting modifications performed by a transformation system can provide users with additional interactive options. Such options can be based on the interface used to initiate content capture and select specific computer-animated models (e.g., initiated from a content creator GUI). In various examples, modifications can be persistent after the modification icon is initially selected. Users can turn modifications on or off by tapping or otherwise selecting the face being modified by the transformation system and save it for later viewing or browsing to other areas of the imaging application. In cases where multiple faces are modified by the transformation system, users can globally turn modifications on or off by tapping or selecting a single face being modified and displayed within the GUI. In some examples, individual faces within a group of multiple faces can be modified separately, or such modifications can be toggled individually by tapping or selecting a single face or a series of individual faces displayed within the GUI.

[0083] Story table 314 stores data about collections of messages and associated image, video, or audio data, compiled into collections (e.g., stories or galleries). The creation of a specific collection can be initiated by a specific user (e.g., each user whose records are stored in entity table 306). A user can create a "personal story" in the form of a collection of content that has already been created and sent / broadcast by that user. For this purpose, the user interface of messaging client 104 may include user-selectable icons that allow the sending user to add specific content to his or her personal story.

[0084] The collection can also constitute a "live story" as a collection of content from multiple users, created manually, automatically, or using a combination of manual and automatic techniques. For example, a "live story" can constitute a curated stream of user-submitted content from various locations and events. Users whose client devices have location services enabled and who are at a co-located event at a specific time can be presented with the option to contribute content to a specific live story, for example, via the user interface of messaging client 104. The messaging client 104 can identify a live story to a user based on their location. The end result is a "live story" told from a collective perspective.

[0085] Another type of content collection is called a "location story," which allows users whose client devices 102 are located in a specific geographic location (e.g., on a college or university campus) to contribute to a specific collection. In some examples, contributing to a location story may require secondary authentication to verify that the end user belongs to a specific organization or other entity (e.g., a student on a university campus).

[0086] As mentioned above, video table 304 stores video data, which, in one example, is associated with a message whose record is stored in message table 302. Similarly, image table 312 stores image data associated with a message whose message data is stored in entity table 306. Entity table 306 allows various enhancements from enhancement table 310 to be associated with various images and videos stored in image table 312 and video table 304.

[0087] The external resource authorization table stores a list of all third-party resources (e.g., external applications, smaller versions of external applications such as web-based external applications and web-based game applications) that have been authorized to access user data of messaging client 104. The external resource authorization table also stores a timer for each authorized external resource, which is reset or refreshed each time the corresponding external resource is used. That is, the timer represents the frequency or relevance of each external resource's use. Whenever a user of messaging client 104 initiates or accesses a feature of an external resource, the timer for that external resource is reset or refreshed. In some cases, when the timer for a given external resource reaches a threshold (e.g., 90 days), the corresponding external resource is automatically revoked (e.g., the authorization for external resource access to user data is revoked until the user re-authorizes the external resource to access user data of messaging client 104). The external resource authorization table may also store the association between the user identifier provided by messaging client 104 and the corresponding user account information generated using a separate version of the external application. Message sending and receiving client 104 uses user identifiers to retrieve user account information generated using one version of an external application for incorporation or merging into another version of the external application.

[0088] Data communication architecture

[0089] Figure 4 This is a schematic diagram illustrating the structure of message 400 according to some examples, generated by message transceiver client 104 for transmission to another message transceiver client 104 or message transceiver server 118. The content of a particular message 400 is used to populate message table 302 in database 126 accessible by message transceiver server 118. Similarly, the content of message 400 is stored in memory as “in-transit” or “in-flight” data for client device 102 or application server 114. Message 400 is shown to include the following example components:

[0090] • Message Identifier 402: A unique identifier that identifies message 400.

[0091] • Message text payload 404: The text to be generated by the user via the user interface of the client device 102 and included in message 400.

[0092] • Message image payload 406: Image data captured by the camera component of the client device 102 or retrieved from the memory component of the client device 102 and included in the message 400. The image data for the sent or received message 400 can be stored in the image table 312.

[0093] • Message video payload 408: Video data captured by the camera device component or retrieved from the memory component of the client device 102 and included in the message 400. The video data for the sent or received message 400 can be stored in the video table 304.

[0094] • Message audio payload 410: Audio data captured by the microphone or retrieved from the memory component of the client device 102 and included in message 400.

[0095] • Message enhancement data 412: Enhancement data (e.g., filters, labels, or other annotations or enhancements) representing enhancements to be applied to the message image payload 406, message video payload 408, or message audio payload 410 of message 400. Enhancement data for the sent or received message 400 can be stored in enhancement table 310.

[0096] • Message duration parameter 414: A parameter value, in seconds, indicating the amount of time that the content of the message (e.g., message image payload 406, message video payload 408, message audio payload 410) will be presented to the user or made accessible to the user via the message sending and receiving client 104.

[0097] • Message geolocation parameter 416: Geolocation data (e.g., latitude and longitude coordinates) associated with the content payload of the message. The payload may include multiple message geolocation parameter 416 values, each of which is associated with a content item included in the content (e.g., a specific image within the message image payload 406, or a specific video within the message video payload 408).

[0098] • Message Story Identifier 418: An identifier value that identifies one or more sets of content (e.g., “Stories” identified in Story Table 314) associated with a specific content item in the message image payload 406 of message 400. For example, the identifier value can be used to associate multiple images within the message image payload 406 with multiple sets of content, respectively.

[0099] • Message Tag 420: Each message 400 can be labeled with multiple tags, each of which indicates the subject of the content included in the message payload. For example, in the case where a specific image in the message image payload 406 depicts an animal (e.g., a lion), a tag value can be included within the message tag 420 indicating the relevant animal. The tag value can be generated manually based on user input, or it can be automatically generated using, for example, image recognition.

[0100] • Message sender identifier 422: An identifier (e.g., message sending system identifier, email address, or device identifier) ​​indicating the user of the client device 102 on which message 400 is generated and from which message 400 is sent.

[0101] • Message receiver identifier 424: An identifier (e.g., message sending and receiving system identifier, email address, or device identifier) ​​indicating the user of the client device 102 to which message 400 is addressed.

[0102] The content (e.g., values) of each component of message 400 can be pointers to locations within tables where content data values ​​are stored. For example, image values ​​in message image payload 406 can be pointers (or addresses) to locations within image table 312. Similarly, values ​​in message video payload 408 can point to data stored in video table 304, values ​​stored in message enhancement data 412 can point to data stored in enhancement table 310, values ​​stored in message story identifier 418 can point to data stored in story table 314, and values ​​stored in message sender identifier 422 and message receiver identifier 424 can point to user records stored in entity table 306.

[0103] Figure 10 This is a flowchart illustrating example operations of a message transceiver client 104 performing process 1000 according to an exemplary embodiment. Process 1000 may be embodied in computer-readable instructions for execution by one or more processors, such that operations of process 1000 may be performed partially or entirely by functional components of client device 102; therefore, process 1000 is described below by way of example. However, in other embodiments, at least some of the operations of process 1000 may be deployed on various other hardware configurations, such as on application server 114. Operations in process 1000 may be executed in any order, in parallel, or may be completely skipped and omitted.

[0104] At operation 1001, client device 102 receives input including messages for a second user via a messaging application implemented on the first user's client device. For example, messaging client 104 receives text input in message composing section 510, which includes messages for another user on communication interface 501.

[0105] At operation 1002, client device 102 obtains context information associated with the message. For example, message sending and receiving client 104 continuously selects words from the text input and searches a context database to retrieve the context associated with the word as words are entered.

[0106] At operation 1003, client device 102 identifies multiple sounds representing a message based on contextual information. For example, message sending and receiving client 104 searches a sound database to identify a set of sounds associated with the context of the identified message.

[0107] At operation 1004, client device 102 receives a selection of a given sound from a plurality of sounds from a first user. For example, message transceiver client 104 receives a selection from a message enhancement region ( Figure 5 The first option 520 is the user input for selecting graphic elements in the associated sound selection area.

[0108] At operation 1004, in response to receiving a selection of a given sound from the first user, client device 102 sends the given sound to the second user. For example, messaging client 104 sends a message or sound, or both, to another user participating in a conversation with the user. The receiving messaging client 104 displays a sound icon 620 to indicate to the second user that the sound is associated with a received message.

[0109] Machine architecture

[0110] Figure 11This is a schematic representation of machine 1100, within which instructions 1108 (e.g., software, programs, applications, applets, or other executable code) can be executed to cause machine 1100 to perform any one or more of the methods discussed herein. For example, instructions 1108 can cause machine 1100 to perform any one or more of the methods described herein. Instructions 1108 transform a general, unprogrammed machine 1100 into a specific machine 1100 programmed to perform the described and illustrated functions in the described manner. Machine 1100 can operate as a standalone device or can be coupled (e.g., networked) to other machines. In a networked deployment, machine 1100 can operate as a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. Machine 1100 may include, but is not limited to: server computers, client computers, personal computers (PCs), tablet computers, laptop computers, netbooks, set-top boxes (STBs), personal digital assistants (PDAs), entertainment media systems, cellular phones, smartphones, mobile devices, wearable devices (e.g., smartwatches), smart home devices (e.g., smart appliances), other smart devices, web devices, network routers, network switches, network bridges, or any machine capable of sequentially or otherwise executing instructions 1108 specifying actions to be taken by machine 1100. Furthermore, although only a single machine 1100 is shown, the term "machine" should also be considered to include a collection of machines that individually or jointly execute instructions 1108 to perform any one or more of the methods discussed herein. For example, machine 1100 may include client device 102 or any of a plurality of server devices forming part of message transceiver server system 108. In some examples, machine 1100 may also include both client and server systems, wherein certain operations of a particular method or algorithm are performed on the server side and certain operations of said particular method or algorithm are performed on the client side.

[0111] Machine 1100 may include a processor 1102, a memory 1104, and an input / output (I / O) unit 1138 that can be configured to communicate with each other via a bus 1140. In the example, processor 1102 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, processor 1106 and processor 1110 that execute instruction 1108. The term "processor" is intended to include multi-core processors, which may include two or more independent processors (sometimes referred to as "cores") capable of executing instructions simultaneously. Although Figure 11 Multiprocessor 1102 is shown, but machine 1100 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.

[0112] Memory 1104 includes main memory 1112, static memory 1114, and memory cells 1116, all of which are accessible by processor 1102 via bus 1140. Main memory 1104, static memory 1114, and memory cells 1116 store instructions 1108 embodying any one or more of the methods or functions described herein. Instructions 1108 may also reside wholly or partially in main memory 1112, in static memory 1114, in machine-readable medium 1118 within memory cells 1116, within at least one processor in processor 1102 (e.g., within the processor's cache memory), or in any suitable combination thereof during execution by machine 1100.

[0113] I / O component 1138 may include various components for receiving input, providing output, generating output, sending information, exchanging information, capturing measurement results, etc. The specific I / O component 1138 included in a particular machine will depend on the type of machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, while a headless server machine is unlikely to include such a touch input device. It will be understood that I / O component 1138 may include... Figure 11Many other components are not shown. In various examples, I / O component 1138 may include user output component 1124 and user input component 1126. User output component 1124 may include visual components (e.g., displays such as plasma display panels (PDPs), light-emitting diode (LED) displays, liquid crystal displays (LCDs), projectors, or cathode ray tube (CRT) displays), acoustic components (e.g., speakers), haptic components (e.g., vibration motors, resistance mechanisms), other signal generators, etc. User input component 1126 may include alphanumeric input components (e.g., keyboards, touchscreens configured to receive alphanumeric input, photoelectric keyboards, or other alphanumeric input components), point-based input components (e.g., mice, touchpads, trackballs, joysticks, motion sensors, or other pointing instruments), haptic input components (e.g., physical buttons, touchscreens or other haptic input components that provide the position and force of a touch or touch gesture), audio input components (e.g., microphones), etc.

[0114] In other examples, I / O component 1138 may include biometric component 1128, motion component 1130, environmental component 1132, or position component 1134, as well as a wide range of other components. For example, biometric component 1128 includes components for detecting expressions (e.g., hand expressions, facial expressions, vocal expressions, body posture, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), and identifying people (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or EEG-based recognition). Motion component 1130 includes: accelerometer components (e.g., accelerometers), gravity sensor components, and rotation sensor components (e.g., gyroscopes).

[0115] The environmental component 1132 includes, for example, one or more camera devices (with still image / photograph and video capabilities), lighting sensor components (e.g., photometers), temperature sensor components (e.g., one or more thermometers that detect ambient temperature), humidity sensor components, pressure sensor components (e.g., barometers), acoustic sensor components (e.g., one or more microphones that detect background noise), proximity sensor components (e.g., infrared sensors that detect nearby objects), gas sensors (e.g., gas detection sensors that detect the concentration of hazardous gases for safety purposes or measure pollutants in the atmosphere), or other components that can provide indications, measurements, or signals corresponding to the surrounding physical environment.

[0116] Regarding the camera device, client device 102 may have a camera device system including, for example, a front-facing camera on the front surface of client device 102 and a rear-facing camera on the rear surface of client device 102. The front-facing camera may be used, for example, to capture still images and videos (e.g., "selfies") of the user of client device 102, which can then be enhanced using the aforementioned enhancement data (e.g., filters). For example, the rear-facing camera may be used to capture still images and videos in a more conventional camera device mode, which are similarly enhanced using the enhancement data. In addition to the front and rear-facing cameras, client device 102 may also include a 360° camera for capturing 360° photos and videos.

[0117] Furthermore, the camera system of the client device 102 may include dual rear cameras (e.g., a main camera and a depth-sensing camera), or even include triple, quadruple, or quintuple rear camera configurations on the front and rear sides of the client device 102. For example, these multi-camera systems may include wide-angle cameras, ultra-wide-angle cameras, telephoto cameras, macro cameras, and depth sensors.

[0118] The position component 1134 includes a positioning sensor component (e.g., a GPS receiver component), an altitude sensor component (e.g., an altimeter or barometer that detects air pressure and can determine altitude based on air pressure), an orientation sensor component (e.g., a magnetometer), etc.

[0119] Various technologies can be used to achieve communication. I / O component 1138 also includes communication component 1136, which is operable to couple machine 1100 to network 1120 or device 1122 via appropriate coupling or connection. For example, communication component 1136 may include a network interface component or another suitable device to interface with network 1120. In further examples, communication component 1136 may include wired communication components, wireless communication components, cellular communication components, near field communication (NFC) components, etc. Components (e.g.) (low power consumption) Components and other communication components that provide communication via other modes. Device 1122 can be any peripheral device from another machine or various peripheral devices (e.g., a peripheral device coupled via USB).

[0120] Furthermore, the communication component 1136 may detect identifiers or include components capable of operating to detect identifiers. For example, the communication component 1136 may include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional barcodes such as Universal Product Code (UPC) barcodes, multi-dimensional barcodes such as Quick Response (QR) codes, Aztec codes, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D barcodes, and other optical codes) or an acoustic detection component (e.g., a microphone for identifying audio signals of the tag). Additionally, various information can be obtained via the communication component 1136, such as location via Internet Protocol (IP) geolocation, etc. Location can be determined by signal triangulation or by detecting NFC beacon signals that indicate a specific location.

[0121] Various memories (e.g., main memory 1112, static memory 1114, and the memory of processor 1102) and storage units 1116 may store one or more sets of instructions and data structures (e.g., software) implemented or used by any one or more of the methods or functions described herein. When executed by processor 1102, these instructions (e.g., instruction 1108) enable various operations to implement the disclosed examples.

[0122] Instructions 1108 can be sent or received over network 1120 via a network interface device (e.g., the network interface component included in communication component 1136), using a transmission medium and employing any of several known transmission protocols (e.g., Hypertext Transfer Protocol (HTTP)). Similarly, instructions 1108 can be sent or received via a transmission medium through a coupling to device 1122 (e.g., peer-to-peer coupling).

[0123] Software Architecture

[0124] Figure 12This is a block diagram 1200 illustrating a software architecture 1204 that can be installed on any or more devices described herein. The software architecture 1204 is supported by hardware such as machine 1202, which includes a processor 1220, memory 1226, and I / O components 1238. In this example, the software architecture 1204 can be conceptualized as a stack of layers, where each layer provides specific functionality. The software architecture 1204 includes layers such as an operating system 1212, libraries 1210, frameworks 1208, and applications 1206. Operationally, application 1206 activates API calls 1250 via the software stack and receives messages 1252 in response to API calls 1250.

[0125] Operating system 1212 manages hardware resources and provides public services. Operating system 1212 includes, for example, a kernel 1214, services 1216, and drivers 1222. Kernel 1214 serves as an abstraction layer between hardware and other software layers. For example, kernel 1214 provides functions such as memory management, processor management (e.g., scheduling), component management, networking, and security settings. Services 1216 can provide other public services to other software layers. Driver 1222 is responsible for controlling the underlying hardware or interfacing with the underlying hardware. For example, driver 1222 may include a display driver, a camera driver, etc. or Low-power drives, flash drives, serial communication drives (e.g., USB drives), Drivers, audio drivers, power management drivers, etc.

[0126] Library 1210 provides common low-level infrastructure used by application 1206. Library 1210 may include system library 1218 (e.g., the C standard library) that provides functions such as memory allocation, string manipulation, and mathematical functions. Additionally, library 1210 may include API library 1224, such as media libraries (e.g., libraries for supporting the rendering and manipulation of various media formats, such as Moving Picture Experts Group-4 (MPEG4), Advanced Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codecs, Joint Picture Experts Group (JPEG or JPG) or Portable Web Graphics (PNG)), graphics libraries (e.g., OpenGL frameworks for rendering graphic content on a display in two-dimensional (2D) and three-dimensional (3D) formats), database libraries (e.g., SQLite for providing various relational database functions), web libraries (e.g., WebKit for providing web browsing functionality), etc. Library 1210 may also include various other libraries 1228 to provide many other APIs to application 1206.

[0127] Framework 1208 provides common high-level infrastructure for use by application 1206. For example, framework 1208 provides various graphical user interface (GUI) functions, high-level resource management, and advanced location services. Framework 1208 can provide a wide range of other APIs that can be used by application 1206, some of which may be specific to a particular operating system or platform.

[0128] In the example, application 1206 may include a home application 1236, a contacts application 1230, a browser application 1232, a book reader application 1234, a location application 1242, a media application 1244, a messaging application 1246, a game application 1248, and a wide variety of other applications such as external application 1240. Application 1206 is a program that performs the functions defined in the program. One or more applications of application 1206 can be created using various programming languages, such as object-oriented programming languages ​​(e.g., Objective-C, Java, or C++) or procedural programming languages ​​(e.g., C or assembly language). In a particular example, external application 1240 (e.g., used by an entity other than a vendor of a particular platform using Android) TM or iOS TM Applications developed using a Software Development Kit (SDK) can be used on platforms such as iOS. TM ANDROID TM , Mobile software running on the phone's mobile operating system or other mobile operating systems. In this example, external application 1240 can activate API calls 1250 provided by operating system 1212 to facilitate the functions described herein.

[0129] Glossary

[0130] "Carrier signal" refers to any intangible medium capable of storing, encoding, or carrying instructions to be executed by a machine, including digital or analog communication signals or other intangible media to facilitate the communication of such instructions. Instructions can be sent or received over a network using a transmission medium via a network interface device.

[0131] "Client device" refers to any machine that interfaces with a communication network to obtain resources from one or more server systems or other client devices. Client devices can be, but are not limited to, mobile phones, desktop computers, laptop computers, portable digital assistants (PDAs), smartphones, tablet computers, ultrabooks, netbooks, laptop computers, multiprocessor systems, microprocessor-based or programmable consumer electronics, game consoles, set-top boxes, or any other communication device that a user can use to access the network.

[0132] "Communication network" refers to one or more parts of a network, which can be an ad hoc network, intranet, extranet, virtual private network (VPN), local area network (LAN), wireless LAN (WLAN), wide area network (WAN), wireless WAN (WWAN), metropolitan area network (MAN), the Internet, a part of the Internet, a part of the Public Switched Telephone Network (PSTN), a POTS (Plain Old-Style Telephone Service) network, a cellular telephone network, a wireless network, etc. A network, other types of networks, or a combination of two or more such networks. For example, a network or part of a network may include a wireless network or a cellular network, and the coupling may be a Code Division Multiple Access (CDMA) connection, a Global System for Mobile Communications (GSM) connection, or other types of cellular or wireless coupling. In this example, the coupling can implement any data transmission technology of various types, such as Single Carrier Radio Transmission (1xRTT), Evolved Data Optimization (EVDO), General Packet Radio Service (GPRS), Enhanced Data Rate Evolution of GSM (EDGE), the 3rd Generation Partnership Project (3GPP) including 3G, fourth-generation wireless (4G) networks, Universal Mobile Telecommunications System (UMTS), High-Speed ​​Packet Access (HSPA), Global Microwave Access Interoperability (WiMAX), Long Term Evolution (LTE) standards, other data transmission technologies defined by various standards setting organizations, other long-distance protocols, or other data transmission technologies.

[0133] A "component" is a device, physical entity, or logic having boundaries defined by functional or subroutine calls, branch points, APIs, or other technologies that partition or modularize a particular processing or control function. A component can be combined with other components via its interface to perform machine processing. A component can be an encapsulated functional hardware unit designed for use with other components, and can be part of a program that typically performs a specific function within a related function.

[0134] Components can constitute software components (e.g., code embodied on a machine-readable medium) or hardware components. A “hardware component” is a tangible unit capable of performing certain operations and can be configured or arranged in some physical manner. In various example implementations, one or more computer systems (e.g., standalone computer systems, client computer systems, or server computer systems) or one or more hardware components (e.g., processors or processor groups) of a computer system can be configured by software (e.g., an application or application portion) to operate to perform certain operations described herein.

[0135] Hardware components can also be implemented mechanically, electronically, or in any suitable combination thereof. For example, a hardware component may include dedicated circuitry or logic permanently configured to perform certain operations. A hardware component may be a dedicated processor, such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). A hardware component may also include a programmable logic or circuit system temporarily configured by software to perform certain operations. For example, a hardware component may include software executed by a general-purpose processor or other programmable processor. Once configured by such software, the hardware component becomes a specific machine (or a specific part of a machine) uniquely tailored to perform the configured function, and no longer a general-purpose processor. It will be understood that the decision to implement a hardware component mechanically in a dedicated and permanently configured circuit or in a temporarily configured (e.g., software-configured) circuit may be made for cost and time considerations. Therefore, the phrase “hardware component” (or “hardware-implemented component”) should be understood to include tangible entities, i.e., entities physically constructed, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate or perform certain operations described herein.

[0136] Considering the implementation where hardware components are temporarily configured (e.g., programmed), it is not necessary to configure or instantiate each hardware component at any given time. For example, where the hardware components include a general-purpose processor that is configured by software to become a dedicated processor, the general-purpose processor can be configured at different times as its own different dedicated processors (e.g., including different hardware components). The software accordingly configures one or more specific processors to constitute a specific hardware component at one time and different hardware components at different times.

[0137] Hardware components can provide information to and receive information from other hardware components. Therefore, the described hardware components can be considered communicatively coupled. In the presence of multiple hardware components, communication can be achieved through signal transmission between or among two or more hardware components (e.g., via appropriate circuitry and buses). In embodiments where multiple hardware components are configured or instantiated at different times, such communication between hardware components can be achieved, for example, by storing information in a memory structure accessible to the multiple hardware components and retrieving information from that memory structure. For example, a hardware component can perform an operation and store the output of that operation in a memory device communicatively coupled to it. Other hardware components can then access the memory device at a subsequent time to retrieve and process the stored output. Hardware components can also initiate communication with input or output devices and can operate on resources (e.g., collections of information).

[0138] The various operations of the example methods described herein can be performed, at least in part, by one or more processors that are temporarily (e.g., by software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors can constitute components of a processor implementation that perform the operations or functions described herein. As used herein, "processor-implemented component" refers to a hardware component implemented using one or more processors. Similarly, the methods described herein can be implemented, at least in part, by processors, where a particular one or more processors are examples of hardware. For example, at least some operations of the methods can be performed by one or more processors 1102 or processor-implemented components. Furthermore, one or more processors can also operate to support the execution of relevant operations in a "cloud computing" environment or as a "Software as a Service" (SaaS) operation. For example, at least some operations can be performed by a group of computers (as an example of a machine including processors), where these operations are accessible via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., APIs). The execution of some operations can be distributed among processors, deployed across multiple machines rather than residing solely within a single machine. In some example implementations, the processor or processor-implemented components may be located in a single geographic location (e.g., within a home environment, office environment, or server cluster). In other example implementations, the processor or processor-implemented components may be distributed across multiple geographic locations.

[0139] "Computer-readable storage medium" refers to both machine-readable storage media and transmission media. Therefore, these terms include both storage devices / media and carrier / modulated data signals. The terms "machine-readable medium," "computer-readable medium," and "device-readable medium" refer to the same thing and may be used interchangeably in this disclosure.

[0140] A "brief message" is a message that is accessible for a limited period of time. Brief messages can be text, images, videos, etc. The access time for a brief message can be set by the message sender. Alternatively, the access time can be a default setting or a setting specified by the recipient. Regardless of the setting method, the message is temporary.

[0141] "Machine storage medium" refers to one or more storage devices and media (e.g., centralized or distributed databases, and associated caches and servers) that store executable instructions, routines, and data. Therefore, this term should be considered to include, but is not limited to, solid-state memory and optical and magnetic media, including memory internal or external to the processor. Specific examples of machine storage media, computer storage media, and / or device storage media include: non-volatile memory, including, for example, semiconductor memory devices such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), FPGAs, and flash memory devices; disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROMs and DVD-ROMs. The terms "machine storage medium," "device storage medium," and "computer storage medium" mean the same thing and may be used interchangeably in this disclosure. The terms "machine storage medium," "computer storage medium," and "device storage medium" expressly exclude carrier waves, modulated data signals, and other such media, at least some of which are covered by the term "signal medium."

[0142] "Non-transitory computer-readable storage medium" refers to a tangible medium capable of storing, encoding, or carrying instructions that can be executed by a machine.

[0143] "Signal medium" means any intangible medium capable of storing, encoding, or carrying machine-executable instructions and including digital or analog communication signals, or other intangible medium that facilitates the communication of software or data. The term "signal medium" should be considered to include any form of modulated data signal, carrier wave, etc. The term "modulated data signal" means a signal whose characteristics are set or altered in a manner that encodes information in the signal. The terms "transmission medium" and "signal medium" mean the same thing and may be used interchangeably in this disclosure.

[0144] Changes and modifications may be made to the disclosed embodiments without departing from the scope of this disclosure. Such and other changes or modifications are intended to be included within the scope of this disclosure as set forth in the appended claims.

Claims

1. A method comprising: Input, including messages for the second user, is received through a messaging application implemented on the first user's client device; In response to receiving the input, obtain context information associated with the message; Based on the context information, identify multiple sounds representing the message; Receive a selection of a given sound from the plurality of sounds from the first user; as well as In response to receiving a selection of the given sound from the first user, the message and the given sound are sent to the second user; as well as Without displaying the content of the given sound, the second user is alerted to the existence of a sound associated with the message, the alert including: In response to the first user sending the message and the given sound, before displaying the content of the given sound, a sound icon indicating that the given sound has been received is displayed to the second user at a display location on the second user's client device, wherein the sound icon is a generic icon that visually notifies the second user that the sound associated with the message has been received; and In response to receiving a user selection from the second user regarding the sound icon for displaying the content of the given sound, In place of the sound icon, a graphic element representing the given sound is displayed at the display location, the graphic element including information associated with the given sound; Retrieve the audio corresponding to the given sound; and Play the audio.

2. The method according to claim 1, further comprising: A display is generated that includes a sound selection area, the sound selection area comprising multiple graphic elements, each of which represents a corresponding sound from among the multiple sounds.

3. The method according to claim 1, further comprising: Generate a display of a communication interface, which includes multiple messages exchanged between the first user and the second user; A text input area for the communication interface is generated for display, wherein the input is received via the text input area; as well as As the text of the message is input into the text input area, context-dependent sounds are continuously identified and included as multiple sounds representing the message.

4. The method according to claim 3, further comprising: Within the communication interface, a message enhancement area is generated adjacent to the text input area. The message enhancement area includes a first option for selecting an avatar representing the message and a second option for selecting the sound representing the message. as well as In response to receiving input that selects the second option, a display of a sound selection area is generated within the message enhancement area, the sound selection area including a plurality of graphic elements, each of the plurality of graphic elements representing a corresponding sound among the plurality of sounds.

5. The method according to claim 1, further comprising: Multiple graphic elements are generated for display, each of the multiple graphic elements representing a corresponding sound among the multiple sounds; Receive input of a first type associated with a first graphic element among the plurality of graphic elements; as well as In response to receiving the first type of input, play the sound associated with the first graphical element.

6. The method according to claim 5, further comprising: Receive a second type of input associated with a first graphic element among the plurality of graphic elements; as well as In response to receiving the second type of input, the sound associated with the first graphic element is sent to the second user.

7. The method according to claim 5, further comprising: A progress bar is displayed near or around the first graphic element to indicate the playback of a sound associated with the first graphic element.

8. The method according to claim 5, further comprising: Select a sound packet from a plurality of sound packets, each of the plurality of sound packets comprising a different set of sounds; as well as Retrieve multiple graphic elements representing multiple sounds in the selected sound pack for inclusion in the display of the multiple graphic elements.

9. The method according to claim 8, wherein, The plurality of sound packs includes at least one of celebrity voices, game audio sounds, music soundtracks, humorous voices, or emotional voices.

10. The method according to claim 1, further comprising: In response to receiving a user selection of the sound icon, the given sound is played on the second user's audio output device.

11. The method according to claim 1, further comprising: The graphic element is persistently displayed after it has been replaced with the sound icon.

12. The method according to claim 1, wherein, The sound icon includes a play button, and the graphic element includes an avatar representing the creator of the given sound or a thumbnail of the given sound.

13. The method according to claim 1, further comprising: After a threshold time period, the display of the graphic element is animated into the sound icon.

14. The method according to claim 1, wherein, Based on the attributes of the given sound, a sound icon is selected from a plurality of sound icons for display, wherein a first sound icon among the plurality of sound icons represents a first type of sound, and a second sound icon among the plurality of sound icons represents a second type of sound.

15. The method according to claim 14, wherein, The attribute includes the creator of the given sound, wherein a first sound icon among the plurality of sound icons corresponds to a sound recorded by a celebrity, and wherein a second sound icon among the plurality of sound icons corresponds to a sound recorded by a friend of the first user.

16. The method of claim 10, further comprising: Receive input associated with the given sound to save the given sound in the sound library associated with the second user.

17. The method of claim 10, further comprising: This allows the second user to adjust or modify the given sound.

18. The method according to claim 1, wherein, The given sound includes micro-audio, which includes recordings of the voices of celebrities.

19. A system comprising: A processor configured to perform operations including: Input, including messages for the second user, is received through a messaging application implemented on the first user's client device; In response to receiving the input, obtain context information associated with the message; Based on the context information, identify multiple sounds representing the message; Receive selection of a given sound from the plurality of sounds from the first user; and In response to receiving a selection of the given sound from the first user, the message and the given sound are sent to the second user; and Without displaying the content of the given sound, the second user is alerted to the existence of a sound associated with the message, the alert including: In response to the first user sending the message and the given sound, before displaying the content of the given sound, a sound icon indicating that the given sound has been received is displayed to the second user at a display location on the second user's client device, wherein the sound icon is a generic icon that visually notifies the second user that the sound associated with the message has been received; and In response to receiving a user selection from the second user regarding the sound icon for displaying the content of the given sound, In place of the sound icon, a graphic element representing the given sound is displayed at the display location, the graphic element including information associated with the given sound; Retrieve the audio corresponding to the given sound; and Play the audio.

20. A non-transitory machine-readable storage medium, the non-transitory machine-readable storage medium comprising instructions that, when executed by one or more processors of a machine, cause the machine to perform operations, the operations including: Input, including messages for the second user, is received through a messaging application implemented on the first user's client device; In response to receiving the input, obtain context information associated with the message; Based on the context information, identify multiple sounds representing the message; Receive a selection of a given sound from the plurality of sounds from the first user; as well as In response to receiving a selection of the given sound from the first user, the message and the given sound are sent to the second user; as well as Without displaying the content of the given sound, the second user is alerted to the existence of a sound associated with the message, the alert including: In response to the first user sending the message and the given sound, before displaying the content of the given sound, a sound icon indicating that the given sound has been received is displayed to the second user at a display location on the second user's client device, wherein the sound icon is a generic icon that visually notifies the second user that the sound associated with the message has been received; and In response to receiving a user selection from the second user regarding the sound icon for displaying the content of the given sound, In place of the sound icon, a graphic element representing the given sound is displayed at the display location, the graphic element including information associated with the given sound; Retrieve the audio corresponding to the given sound; and Play the audio.

Citation Information

Patent Citations

  • Method and system of enhanced messaging

    US20050204309A1

  • Disambiguation of icons and other media in text-based applications

    US20080244446A1