Trigger gestures for selecting augmented reality content in messaging systems

KR103012780B1Active Publication Date: 2026-09-02SNAP INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
KR1020257011548
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-09-09
Filing Date
2023-09-08
Publication Date
2026-09-02
Estimated Expiration
2043-09-08

Smart Images

  • Figure 112025039834042-PCT00014_ABST
    Figure 112025039834042-PCT00014_ABST
Patent Text Reader

Abstract

The present technology detects a first gesture corresponding to an open trigger finger gesture. The present technology detects the position and location of a finger expression from the open trigger finger gesture. The present technology creates a first virtual object based at least partially on the position and location of the finger expression. The present technology detects a first collision event. The present technology detects a second gesture corresponding to a closed trigger finger gesture. The present technology selects the second virtual object. In response to the selection, the present technology renders the first virtual object as attached to the second virtual object. The present technology provides the first virtual object rendered as attached to the second virtual object within the first scene for display.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] Claim of priority

[0002] This patent application claims the benefit of priority to U.S. Application No. 17 / 941,612 filed September 9, 2022, the entirety of which is incorporated herein by reference. Background Technology

[0003] With the increasing use of digital images, the affordability of portable computing devices, the availability of increased capacity in digital storage media, and the increased bandwidth and accessibility of network connections, digital images have become a part of the daily lives of more and more people. Brief explanation of the drawing

[0004] In drawings that are not necessarily drawn to actual scale, similar numbers may describe similar components within different views. To facilitate the identification of any specific element or act, the top digit or numbers in a reference number indicate the drawing number where the element is first introduced. Some non-limiting examples are illustrated in the drawings of the attached drawings: FIG. 1 is a schematic representation of a networked environment in which the present disclosure may be arranged, according to some examples. Figure 2 is a schematic representation of a messaging system according to some examples that has functionality on both the client side and the server side. Figure 3 is a schematic representation of a data structure as maintained in a database, according to some examples. Figure 4 is a schematic representation of a message according to some examples. Figure 5 illustrates the process of creating and utilizing a user interface according to some examples. Figure 6 depicts a sequence diagram of an exemplary user interface process according to some examples. FIG. 7a illustrates an exemplary interface according to various embodiments. FIG. 7b illustrates an exemplary interface according to various embodiments. FIG. 8 illustrates exemplary interfaces according to various embodiments. FIG. 9 illustrates exemplary interfaces according to various embodiments. FIG. 10 illustrates exemplary interfaces according to various embodiments. FIG. 11 illustrates exemplary interfaces according to various embodiments. FIG. 12 illustrates exemplary interfaces according to various embodiments. FIG. 13 illustrates exemplary interfaces according to various embodiments. FIG. 14 illustrates exemplary interfaces according to various embodiments. FIG. 15 illustrates exemplary interfaces according to various embodiments. FIG. 16 illustrates exemplary interfaces according to various embodiments. FIG. 17 illustrates exemplary interfaces according to various embodiments. FIG. 18 illustrates exemplary interfaces according to various embodiments. FIG. 19 illustrates exemplary interfaces according to various embodiments. FIG. 20 illustrates exemplary interfaces according to various embodiments. FIG. 21 illustrates exemplary interfaces according to various embodiments. FIG. 22 illustrates exemplary interfaces according to various embodiments. FIG. 23 illustrates exemplary interfaces according to various embodiments. FIG. 24 illustrates exemplary interfaces according to various embodiments. FIG. 25 illustrates exemplary interfaces according to various embodiments. FIG. 26 is a flowchart illustrating a method according to specific exemplary embodiments. FIG. 27 is a schematic representation of a machine in the form of a computer system in which a set of instructions can be executed to enable the machine to perform any one or more of the methodologies discussed in this specification, according to some examples. FIG. 28 is a block diagram illustrating a software architecture in which examples can be implemented. Specific details for implementing the invention

[0005] Users with diverse interests from various locations can capture digital images of various subjects and make the captured images available to others through networks such as the Internet. To enhance users' experience with digital images and provide diverse features, enabling computing devices to perform image processing operations on various objects and / or features captured under a wide range of changing conditions (e.g., changes in image scales, noise, lighting, motion, or geometric distortion) can be a challenging and computationally intensive task.

[0006] Augmented reality technology aims to bridge the gap between virtual and real-world environments by providing enhanced real-world environments augmented with electronic information. As a result, the electronic information appears to be part of the real-world environment as perceived by the user. In one example, augmented reality technology further provides a user interface that allows interaction with electronic information overlaid on the enhanced real-world environment.

[0007] As mentioned above, due to the increasing use of digital images, the affordability of portable computing devices, the availability of increased capacity in digital storage media, and the increased bandwidth and accessibility of network connections, digital images have become an increasingly common part of people's daily lives. Users with diverse interests from various locations can capture digital images of various subjects and make them available to others through networks such as the Internet. To enhance users' experience with digital images and provide diverse features, enabling computing devices to perform image processing operations on various objects and / or features captured under a wide range of changing conditions (e.g., changes in image scales, noise, lighting, motion, or geometric distortion) can be a challenging and computationally intensive task.

[0008] Augmented reality technology aims to bridge the gap between virtual and real-world environments by providing enhanced real-world environments augmented with electronic information. As a result, the electronic information appears to be part of the real-world environment as perceived by the user. In one example, augmented reality technology further provides a user interface that allows interaction with electronic information overlaid on the enhanced real-world environment.

[0009] Augmented reality (AR) systems allow real and virtual environments to be combined to varying degrees to facilitate real-time interactions from users. Accordingly, as described herein, such AR systems may include various possible combinations of real and virtual environments, including augmented reality that primarily contains real elements and is closer to the real environment than a virtual environment (e.g., one without real elements). In this way, the real environment can be connected to the virtual environment by the AR system. A user immersed in the AR environment can explore this environment, and the AR system can track the user's viewpoint to provide visualizations based on how the user is positioned within the environment. Augmented reality (AR) experiences may be provided in a messaging client application (or messaging system) as described in the embodiments herein.

[0010] The embodiments of the technology described herein enable various operations involving AR content for capturing and modifying such content with a given electronic device, such as a mobile computing device.

[0011] Messaging systems are frequently and increasingly utilized by users of mobile computing devices in various settings to provide different types of functionality in a convenient manner. As described herein, the messaging system includes practical applications that provide improvements in capturing image data and rendering AR content (e.g., images, videos, etc.) based on captured image data by providing at least technical improvements in capturing image data using power and resource-constrained electronic devices. These improvements in capturing image data are made possible by techniques provided by the present technology, which reduce latency and increase efficiency in processing captured image data, thereby also reducing power consumption in capturing devices.

[0012] As further discussed herein, the infrastructure supports the creation and sharing of interactive media, referred herein as messages including 3D content or AR effects, across various components of the messaging system. In the exemplary embodiments described herein, messages may enter the system from a live camera or from storage (e.g., where messages including 3D content and / or AR effects are stored in memory or a database). The system supports motion sensor input and the loading of external effects and asset data.

[0013] As mentioned in this specification, the phrases “augmented reality experience,” “augmented reality content item,” and “augmented reality content generator” include or refer to various image processing operations corresponding to image modification, filters, AR content generators, media overlays, transformations, etc., as further described in this specification, and may additionally include the playback of audio or music content during the presentation of AR content or media content.

[0014] Networked computing environment

[0015] FIG. 1 is a block diagram illustrating an exemplary messaging system (100) for exchanging data (e.g., messages and associated content) over a network. The messaging system (100) includes multiple instances of client devices (102), each of which hosts multiple applications including a messaging client (104) and other applications (106). Each messaging client (104) is coupled to communicate with other instances of the messaging client (104) (e.g., hosted on their respective other client devices (102)), a messaging server system (108), and third-party servers (110) over a network (112) (e.g., the Internet). The messaging client (104) may also communicate with locally hosted applications (106) using Applications Program Interfaces (APIs).

[0016] A messaging client (104) can communicate with other messaging clients (104) and a messaging server system (108) and exchange data through a network (112). The data exchanged between messaging clients (104) and between a messaging client (104) and a messaging server system (108) includes functions (e.g., commands that invoke functions) as well as payload data (e.g., text, audio, video, or other multimedia data).

[0017] The messaging server system (108) provides server-side functionality to a specific messaging client (104) via the network (112). Although specific functions of the messaging system (100) are described herein as being performed by the messaging client (104) or by the messaging server system (108), the location of specific functions within the messaging client (104) or the messaging server system (108) may be a design choice. For example, it may be technically desirable to initially place specific technologies and functions within the messaging server system (108), but later migrate these technologies and functions to the messaging client (104) when the client device (102) has sufficient processing capacity.

[0018] The messaging server system (108) supports various services and operations provided to the messaging client (104). Such operations include transmitting data to the messaging client (104), receiving data from it, and processing data generated by it. This data may include, for example, message content, client device information, geolocation information, media augmentation and overlays, message content persistence conditions, social network information, and live event information. Data exchanges within the messaging system (100) are invoked and controlled through functions available via the user interface (UI) of the messaging client (104).

[0019] Now, referring specifically to the messaging server system (108), an application program interface (API) server (116) is coupled to application servers (114) to provide a programmatic interface. The application servers (114) are communicably coupled to a database server (120), which facilitates access to a database (126) that stores data associated with messages processed by the application servers (114). Similarly, a web server (128) is coupled to the application servers (114) and provides web-based interfaces to the application servers (114). To this end, the web server (128) processes incoming network requests via HTTP (Hypertext Transfer Protocol) and various other related protocols.

[0020] An API (Application Program Interface) server (116) receives and transmits message data (e.g., commands and message payloads) between a client device (102) and application servers (114). Specifically, the application program interface (API) server (116) provides a set of interfaces (e.g., routines and protocols) that can be called or queried by a messaging client (104) to invoke the functionality of the application servers (114). The application program interface (API) server (116) exposes various functions supported by the application servers (114), including account registration, login functionality, transmission of messages through the application servers (114) from a specific messaging client (104) to another messaging client (104), transmission of media files (e.g., images or videos) from the messaging client (104) to the messaging server (118), and settings of a collection of media data (e.g., stories) for possible access by another messaging client (104), searching of a list of friends of a user of the client device (102), searching of such collections, searching of messages and content, adding and deleting entities (e.g., friends) to an entity graph (e.g., social graph), locating friends within the social graph, and opening application events (e.g., related to the messaging client (104)).

[0021] Application servers (114) host a number of server applications and subsystems, including, for example, a messaging server (118), an image processing server (122), and a social network server (124). The messaging server (118) implements a number of message processing techniques and functions, particularly those related to the aggregation and other processing of content (e.g., text and multimedia content) contained in messages received from multiple instances of the messaging client (104). As described in more detail, text and media content from multiple sources may be aggregated into collections of content (e.g., called stories or galleries). These collections are then made available to the messaging client (104). Other processor and memory-intensive data processing may also be performed by the messaging server (118) on the server side, taking into account the hardware requirements for such processing.

[0022] The application servers (114) also include an image processing server (122) dedicated to performing various image processing operations with respect to images or videos in the payload of a message typically transmitted from or received from the messaging server (118).

[0023] The social network server (124) supports various social networking functions and services and makes these functions and services available to the messaging server (118). To this end, the social network server (124) maintains and accesses an entity graph (308) (shown in FIG. 3) within a database (126). Examples of functions and services supported by the social network server (124) include the identification of other users of the messaging system (100) with whom a specific user has a relationship or is "following," and also the identification of other entities and interests of the specific user.

[0024] Returning to the messaging client (104), features and functions of external resources (e.g., an application (106) or an applet) become available to the user through the interface of the messaging client (104). In this context, "external" refers to the fact that the application (106) or applet is outside the messaging client (104). External resources are often provided by a third party, but may also be provided by the creator or provider of the messaging client (104). The messaging client (104) receives a user selection of options to launch or access the features of these external resources. The external resources may be an application (106) installed on the client device (102) (e.g., a "native app"), or a small version of an application (e.g., an "applet") hosted on the client device (102) or located remotely from the client device (102) (e.g., on third-party servers (110)). A small version of an application contains a subset of the features and functions of the application (e.g., a full-scale native version of the application) and is implemented using a markup language document. In one example, a small version of the application (e.g., "Applet") is a web-based markup language version of the application and is embedded in a messaging client (104). In addition to using markup language documents (e.g., .*ml files), the applet can incorporate a scripting language (e.g., .*js files or .json files) and a style sheet (e.g., .*ss files).

[0025] In response to receiving a user selection of an option to launch or access features of an external resource, the messaging client (104) determines whether the selected external resource is a web-based external resource or a locally-installed application (106). In some cases, locally installed applications (106) on the client device (102) may be launched independently of and separately from the messaging client (104), for example, by selecting an icon corresponding to the application (106) on the home screen of the client device (102). Smaller versions of these applications may be launched or accessed through the messaging client (104), and in some examples, any part of the small application may not be accessible outside the messaging client (104), or only limited parts may be accessible. A small application can be launched by a messaging client (104) receiving, for example, a markup language document associated with the small application from a third-party server (110) and processing such document.

[0026] In response to determining that the external resource is a locally installed application (106), the messaging client (104) instructs the client device (102) to launch the external resource by executing locally-stored code corresponding to the external resource. In response to determining that the external resource is a web-based resource, the messaging client (104) communicates with (e.g.) third-party servers (110) to obtain a markup language document corresponding to the selected external resource. Subsequently, the messaging client (104) processes the obtained markup language document to present the web-based external resource within the user interface of the messaging client (104).

[0027] A messaging client (104) may notify a user of a client device (102) or other users associated with such user (e.g., “friends”) of activity occurring in one or more external resources. For example, the messaging client (104) may provide participants in a conversation (e.g., a chat session) in the messaging client (104) with notifications regarding current or recent use of an external resource by one or more members of a group of users. One or more users may be invited to participate in an active external resource or to launch an external resource that was recently used but is currently inactive (in the group of friends). The external resource may provide participants in the conversation, each using their own messaging clients (104), with the ability to share items, situations, states, or locations in the external resource with one or more members of a group of users in a chat session. The shared items may be interactive chat cards that allow members of the chat to interact, for example, to launch a corresponding external resource, view specific information within the external resource, or take a member of the chat to a specific location or state within the external resource. Within a given external resource, response messages can be sent to users on a messaging client (104). The external resource may optionally include different media items in the responses based on the current context of the external resource.

[0028] The messaging client (104) may present to the user a list of available external resources (e.g., applications (106) or applets) to launch or access a given external resource. This list may be presented in a context-sensitive menu. For example, icons representing different applications (106) (or applets) may vary based on how the menu is launched by the user (e.g., from a conversational interface or from a non-conversational interface).

[0029] System Architecture

[0030] FIG. 2 is a block diagram illustrating additional details regarding a messaging system (100) according to some examples. Specifically, the messaging system (100) is illustrated as including a messaging client (104) and application servers (114). The messaging system (100) implements a number of subsystems supported by the messaging client (104) on the client side and by the application servers (114) on the server side. These subsystems include, for example, a short-term timer system (202), a collection management system (204), an augmentation system (208), a map system (210), a game system (212), and an external resource system (214).

[0031] The short-term timer system (202) is responsible for enforcing temporary or time-limited access to content by the messaging client (104) and the messaging server (118). The short-term timer system (202) includes a plurality of timers that selectively enable access (e.g., for presentation and display) to messages and associated content through the messaging client (104) based on duration and display parameters associated with a message, or a collection of messages (e.g., a story). Further details regarding the operation of the short-term timer system (202) are provided below.

[0032] The collection management system (204) is responsible for managing sets or collections of media (e.g., collections of text, image, video, and audio data). Collections of content (e.g., messages including images, videos, text, and audio) may be organized into "event galleries" or "event stories." These collections may be available for a specific period, such as the duration of the event to which the content is related. For example, content related to a music concert may be available as a "story" for the duration of that music concert. The collection management system (204) may also be responsible for publishing an icon that provides notification of the existence of a specific collection in the user interface of the messaging client (104).

[0033] The collection management system (204) further includes a curation interface (206) that allows a collection manager to manage and curate collections of specific content. For example, the curation interface (206) enables an event organizer to curate collections of content related to a specific event (e.g., removing inappropriate content or duplicate messages). Additionally, the collection management system (204) automatically curates content collections using machine vision (or image recognition technology) and content rules. In certain examples, a reward may be paid to users for including user-generated content in a collection. In these cases, the collection management system (204) operates to automatically pay these users for using their content.

[0034] The augmentation system (208) provides various functions that enable a user to augment media content associated with a message (e.g., annotate or modify or edit in other ways). For example, the augmentation system (208) provides functions related to the creation and publication of media overlays for messages processed by the messaging system (100). The augmentation system (208) operatively supplies media overlays or augmentations (e.g., image filters) to the messaging client (104) based on the geolocation of the client device (102). In another example, the augmentation system (208) operatively supplies media overlays to the messaging client (104) based on other information, such as the social network information of the user of the client device (102). Media overlays may include audio and visual content and visual effects. Examples of audio and visual content include photos, text, logos, animations, and sound effects. Examples of visual effects include color overlays. Audio and visual content or visual effects may be applied to a media content item (e.g., a photo) on the client device (102). For example, the media overlay may include text or an image that can be overlaid on top of a photo taken by the client device (102). In another example, the media overlay may include a location identification overlay (e.g., Venice Beach), the name of a live event, or a seller name overlay (e.g., Beach Coffee House). In another example, the augmentation system (208) uses the geolocation of the client device (102) to identify a media overlay containing the seller name at the geolocation of the client device (102). The media overlay may include other indicia associated with the seller.Media overlays are stored in a database (126) and can be accessed through a database server (120).

[0035] In some examples, the augmentation system (208) provides a user-based publication platform that enables users to select a geolocation on a map and upload content associated with the selected geolocation. Users can also specify situations in which a particular media overlay should be provided to other users. The augmentation system (208) creates a media overlay that includes the uploaded content and associates the uploaded content with the selected geolocation.

[0036] In other examples, the augmentation system (208) provides a merchant-based publication platform that enables merchants to select a specific media overlay associated with a geolocation through a bidding process. For example, the augmentation system (208) associates the media overlay of the highest bidder with a corresponding geolocation for a predefined amount of time.

[0037] The map system (210) provides various geographic location functions and supports the presentation of map-based media content and messages by the messaging client (104). For example, the map system (210) enables the display of user icons or avatars (e.g., stored in profile data (316)) on the map to display the current or past locations of the user's "friends" as well as media content created by these friends (e.g., collections of messages including photos and videos) within the context of the map. For example, a message posted by a user to the messaging system (100) from a specific geographic location can be displayed to the specific user's "friends" on the map interface of the messaging client (104) within the context of the map at that specific location. Furthermore, the user can share their location and status information with other users of the messaging system (100) through the messaging client (104) (e.g., using an appropriate status avatar), and this location and status information is similarly displayed to selected users within the context of the map interface of the messaging client (104).

[0038] The game system (212) provides various gaming functions within the context of the messaging client (104). The messaging client (104) provides a game interface that provides a list of available games that can be launched by a user within the context of the messaging client (104) and played with other users of the messaging system (100). The messaging system (100) also enables a specific user to invite other users to participate in the play of a specific game by issuing invitations to other users from the messaging client (104). The messaging client (104) also supports both voice and text messaging (e.g., chats) within the context of gameplay, provides a leaderboard for games, and supports the provision of in-game rewards (e.g., coins and items).

[0039] The external resource system (214) provides an interface for the messaging client (104) to communicate with remote servers (e.g., third-party servers (110)) to launch or access external resources, namely applications or applets. Each third-party server (110) hosts, for example, a markup language (e.g., HTML5) based application or a small version of an application (e.g., a game, utility, payment, or ride-sharing application). The messaging client (104) can launch a web-based resource (e.g., an application) by accessing an HTML5 file from the third-party servers (110) associated with the web-based resource. In certain examples, applications hosted by the third-party servers (110) are programmed in JavaScript using a Software Development Kit (SDK) provided by the messaging server (118). The SDK includes Application Programming Interfaces (APIs) that have functions that can be called or invoked by web-based applications. In certain examples, the messaging server (118) includes a JavaScript library that provides access to a given external resource for specific user data of the messaging client (104). HTML5 is used as an exemplary technology for programming games, but applications and resources programmed based on other technologies may be used.

[0040] To integrate the functions of the SDK into a web-based resource, the SDK is downloaded from the messaging server (118) by the third-party server (110) or otherwise received by the third-party server (110). Once downloaded or received, the SDK is included as part of the application code of the web-based external resource. Subsequently, the code of the web-based resource may call or invoke specific functions of the SDK to integrate the features of the messaging client (104) into the web-based resource.

[0041] The SDK stored in the messaging server (118) effectively provides a bridge between external resources (e.g., applications (106) or applets) and the messaging client (104). This provides the user with an experience of communicating seamlessly with other users on the messaging client (104) while preserving the look and feel of the messaging client (104). To bridge communications between the external resources and the messaging client (104), in certain examples, the SDK facilitates communication between third-party servers (110) and the messaging client (104). In certain examples, a WebViewJavaScriptBridge running on a client device (102) establishes two unidirectional communication channels between the external resources and the messaging client (104). Messages are transmitted asynchronously between the external resources and the messaging client (104) through these communication channels. Each SDK function call is transmitted as a message and a callback. Each SDK function is implemented by constructing a unique callback identifier and sending a message containing that callback identifier.

[0042] By using the SDK, not all information from the messaging client (104) is shared with the third-party servers (110). The SDK limits which information is shared based on the needs of the external resources. In certain examples, each third-party server (110) provides the messaging server (118) with an HTML5 file corresponding to the web-based external resource. The messaging server (118) can add a visual representation of the web-based external resource (e.g., box art or other graphics) to the messaging client (104). Once the user selects the visual representation or instructs the messaging client (104) to access the features of the web-based external resource through the GUI of the messaging client (104), the messaging client (104) obtains the HTML5 file and instantiates the resources necessary to access the features of the web-based external resource.

[0043] The messaging client (104) presents a graphical user interface (e.g., a landing page or a title screen) for an external resource. While presenting the landing page or title screen, or before or after, the messaging client (104) determines whether the launched external resource was previously authorized to access the messaging client (104)'s user data. In response to the determination that the launched external resource was previously authorized to access the messaging client (104)'s user data, the messaging client (104) presents another graphical user interface of the external resource, including the functions and features of the external resource. In response to determining that the launched external resource has not previously been authorized to access the messaging client (104)'s user data, after a threshold period (e.g., 3 seconds) during which the external resource's landing page or title screen is displayed, the messaging client (104) slides up a menu to authorize the external resource to access user data (e.g., animates the menu as it surfaces from the bottom of the screen to the middle or another part of the screen). The menu identifies the type of user data to be authorized for the external resource to use. In response to receiving a user selection of an accept option, the messaging client (104) adds the external resource to a list of authorized external resources and allows the external resource to access user data from the messaging client (104). In some examples, the external resource is authorized by the messaging client (104) to access user data according to the OAuth 2 framework.

[0044] The messaging client (104) controls the type of user data shared with external resources based on the type of external resources authorized. For example, external resources including full-scale applications (e.g., application (106)) are provided with access to a first type of user data (e.g., only two-dimensional avatars of users with or without different avatar characteristics). As another example, external resources including small-scale versions of applications (e.g., web-based versions of applications) are provided with access to a second type of user data (e.g., payment information, two-dimensional avatars of users, three-dimensional avatars of users, and avatars with various avatar characteristics). Avatar characteristics include different ways of customizing the look and feel of the avatar, such as different poses, facial features, clothing, etc.

[0045] Data Architecture

[0046] FIG. 3 is a schematic diagram illustrating data structures (300) that may be stored in a database (126) of a messaging server system (108) according to specific examples. Although the contents of the database (126) are depicted as including multiple tables, it will be recognized that data may be stored in data structures of other types (e.g., as an object-oriented database).

[0047] The database (126) contains message data stored in the message table (302). For any specific message, this message data includes at least message sender data, message recipient (or receiver) data, and payload. Additional details regarding information that may be included in the message and may be included in the message data stored in the message table (302) are described below with reference to FIG. 4.

[0048] The entity table (306) stores entity data and is linked (e.g., for reference) to the entity graph (308) and profile data (316). Entities for which records are maintained within the entity table (306) may include individuals, legal entities, organizations, objects, places, events, etc. Regardless of the entity type, any entity for which the messaging server system (108) stores data may be a recognized entity. Each entity has an entity type identifier (not shown) as well as a unique identifier.

[0049] The entity graph (308) stores information regarding relationships and associations between entities. Such relationships may be based on interests or activities, for example, social, professional (for example, work in a general corporation or organization).

[0050] Profile data (316) stores multiple types of profile data for a specific entity. Profile data (316) may be optionally used and presented to other users of the messaging system (100) based on privacy settings specified by the specific entity. If the entity is an individual, profile data (316) includes, for example, a username, phone number, address, settings (e.g., notification and privacy settings), as well as a user-selected avatar representation (or a collection of such avatar representations). A specific user may then optionally include one or more of these avatar representations within the content of messages communicated through the messaging system (100) and on map interfaces displayed to other users by messaging clients (104). A collection of avatar representations may include "state avatars" that present a graphic representation of a state or activity that the user may select to communicate at a specific time.

[0051] If the entity is a group, the profile data (316) for the group may similarly include one or more avatar representations associated with the group, in addition to the group name, members, and various settings (e.g., notifications) for the relevant group.

[0052] The database (126) also stores augmentation data, such as overlays or filters, in the augmentation table (310). The augmentation data is associated with and applied to videos (data about which is stored in the video table (304)) and images (data about which is stored in the image table (312)).

[0053] In one example, the filters are overlays that are displayed overlaid on an image or video during presentation to the recipient user. The filters may be various types of filters, including user-selected filters from a set of filters presented to the sender user by the messaging client (104) when the sender user is composing a message. Other types of filters include geolocation filters (also known as geo-filters) that may be presented to the sender user based on geolocation. For example, geolocation filters specific to a neighborhood or a particular location may be presented within the user interface by the messaging client (104) based on geolocation information determined by the Global Positioning System (GPS) unit of the client device (102).

[0054] Another type of filter is a data filter that may be optionally presented to the sending user by the messaging client (104) based on other inputs or information collected by the client device (102) during the message generation process. Examples of data filters include the current temperature at a specific location, the current speed at which the sending user is moving, the battery life of the client device (102), or the current time.

[0055] Other augmented data that may be stored in the image table (312) includes augmented reality content items (e.g., corresponding to applying lenses or augmented reality experiences). Augmented reality content items may be real-time special effects and sounds that can be added to an image or video.

[0056] As described above, augmented data includes similar terms referring to augmented reality content items, overlays, image transformations, AR images, and modifications that can be applied to image data (e.g., videos or images). This includes real-time modifications that modify images, such as when an image is captured using device sensors (e.g., one or more cameras) of the client device (102) and then displayed on the screen of the client device (102) along with modifications. This also includes modifications to stored content, such as video clips within a gallery, that can be modified. For example, on a client device (102) accessing multiple augmented reality content items, the user may use a single video clip along with multiple augmented reality content items to determine how different augmented reality content items modify the stored clip. For example, multiple augmented reality content items applying different pseudorandom movement models may be applied to the same content by selecting different augmented reality content items for the content. Similarly, real-time video capture can be used with the illustrated modifications to show how video images currently being captured by the sensors of the client device (102) will modify the captured data. This data may simply be displayed on the screen and not stored in memory, or the content captured by the device sensors may be written and stored in memory with or without modifications (or both). In some systems, a preview feature may show how different augmented reality content items will appear simultaneously in different windows within the display. This may, for example, allow multiple windows with different pseudo-random animations to be displayed simultaneously on the display.

[0057] Therefore, data and various systems using augmented reality content items or other such transformation systems to modify content using this data may involve the detection of objects (e.g., faces, hands, bodies, cats, dogs, surfaces, objects, etc.), the tracking of these objects as they move out of, enter, and around video frames, and the modification or transformation of these objects when they are tracked. In various examples, different methods may be used to achieve these transformations. Some examples may involve generating a 3D mesh model of an object or objects, and using transformations of the model and animated textures within the video to achieve the transformation. In other examples, tracking points on an object may be used to place an image or texture (which may be 2D or 3D) at the tracked position. In yet other examples, neural network analysis of video frames may be used to place images, models, or textures on the content (e.g., images or video frames). Therefore, augmented reality content items refer to the images, models, and textures used to create transformations in the content, as well as all additional modeling and analysis information required to achieve these transformations through object detection, tracking, and placement.

[0058] Real-time video processing can be performed with any type of video data (e.g., video streams, video files, etc.) stored in the memory of any type of computerized system. For example, a user can load video files and store them in the device's memory, or generate video streams using the device's sensors. Additionally, any objects, such as human faces and parts of the human body, animals, or inanimate objects like chairs, cars, or other objects, can be processed using computer animation models.

[0059] In some examples, when a specific modification is selected along with the content to be transformed, the elements to be transformed are identified by a computing device and subsequently detected and tracked if they exist in the frames of the video. The elements of the object are modified according to the request for modification, thereby transforming the frames of the video stream. The transformation of the frames of the video stream may be performed by different methods for different types of transformations. For example, for frame transformations that primarily refer to changing the shapes of the elements of an object, characteristic points for each element of the object are calculated (e.g., using Active Shape Model (ASM) or other known methods). Then, a mesh based on the characteristic points is generated for each of at least one element of the object. This mesh is used in the next stage of tracking the elements of the object within the video stream. In the tracking process, the mesh mentioned for each element is aligned with the position of each element. Subsequently, additional points are generated on the mesh. A first set of first points is generated for each element based on the request for modification, and a second set of points is generated for each element based on the first set of points and the request for modification. Subsequently, the frames of the video stream can be transformed by modifying the elements of the object based on the mesh and the sets of first and second points. In this method, the background of the modified object can also be changed or distorted by tracking and modifying the background.

[0060] In some examples, transformations that modify parts of an object's regions using its elements can be performed by calculating characteristic points for each element of the object and generating a mesh based on the calculated characteristic points. Points are generated on the mesh, and then various regions based on the points are generated. Subsequently, the elements of the object are tracked by aligning the regions for each element with the positions for at least one element, and the characteristics of the regions can be modified based on requests for modification, thereby transforming the frames of the video stream. Depending on a specific request for modification, the characteristics of the mentioned regions can be transformed in different ways. Such modifications may involve changing the color of the regions; removing at least a portion of the regions from the frames of the video stream; including one or more new objects in the regions based on the modification request; and modifying or distorting the regions or elements of the objects. In various examples, any combination of these modifications or other similar modifications may be used. For specific models to be animated, some characteristic points may be selected as control points to be used to determine the entire state space of options for the model animation.

[0061] In some examples of computer animation models for transforming image data using face detection, faces are detected on the image using a specific face detection algorithm (e.g., Viola-Jones). Then, the Active Shape Model (ASM) algorithm is applied to the face region of the image to detect face feature reference points.

[0062] Other methods and algorithms suitable for face detection may be used. For example, in some examples, features are located using landmarks that represent distinguishable points present in most of the images under consideration. For face landmarks, for example, the location of the left pupil may be used. If the initial landmark is not identifiable (e.g., a person wearing an eye patch), secondary landmarks may be used. These landmark identification procedures can be applied to any such objects. In some examples, a set of landmarks forms a feature. Features can be represented as vectors using the coordinates of points within the feature. One feature is aligned with another using a similarity transformation (allowing translation, scaling, and rotation) that minimizes the average Euclidean distance between feature points. The average feature is the average of the aligned training features.

[0063] In some examples, a search for landmarks from the average shape aligned with the position and size of the face determined by the global face detector is initiated. This search then repeats the steps of proposing a provisional shape by adjusting the positions of the shape points through template matching of the image texture around each point, and then conforming the provisional shape to the global shape model until convergence occurs. In some systems, individual template matches are unreliable, and the shape model pools the results of weak template matches to form a stronger overall classifier. The overall search is repeated at each level of the image pyramid, from coarse to fine resolution.

[0064] The transformation system can capture an image or video stream on a client device (e.g., client device (102)) and perform complex image manipulations locally on the client device (102) while maintaining an appropriate user experience, computation time, and power consumption. Complex image manipulations may include size and shape changes, emotion transfers (e.g., changing a face from a frown to a smile), state transfers (e.g., aging a subject, reducing apparent age, changing gender), style transfers, application of graphic elements, and any other appropriate image or video manipulations implemented by a convolutional neural network configured to be executed efficiently on the client device (102).

[0065] In some examples, a computer animation model for transforming image data may be used by a system capable of capturing a user's image or video stream (e.g., a selfie) using a client device (102) having a neural network that operates as part of a messaging client (104) operating on the client device (102). A transformation system operating within the messaging client (104) determines the presence of a face in the image or video stream and provides modification icons associated with the computer animation model for transforming image data, or the computer animation model may exist as being associated with the interface described herein. The modification icons include changes that may be the basis for modifying the user's face in the image or video stream as part of a modification operation. When a modification icon is selected, the transformation system initiates a process of transforming the user's image to reflect the selected modification icon (e.g., generating a smiling face for the user). The modified image or video stream may be presented in a graphical user interface displayed on the client device (102) as soon as the image or video stream is captured and a specific modification is selected. The transformation system can generate and apply selected modifications by implementing complex convolutional neural networks on a portion of an image or video stream. That is, a user can capture an image or video stream and, if a modification icon is selected, be presented with the modified result in real-time or near real-time. Furthermore, modifications can be sustained as long as the video stream is being captured and the selected modification icon remains toggled. Machine-taught neural networks can be used to enable these modifications.

[0066] A graphical user interface presenting modifications performed by a transformation system may provide the user with additional interaction options. These options may be based on the interface used to initiate the capture and selection of content from a specific computer animation model (e.g., initiation from a content creator user interface). In various examples, modifications may be persistent after the initial selection of a modification icon. The user may toggle modifications on or off by tapping or otherwise selecting the face being modified by the transformation system, and may save it for later viewing or browsing to other areas of the imaging application. If multiple faces are modified by the transformation system, the user may toggle modifications on or off globally by tapping or selecting a single face that is modified and displayed within the graphical user interface. In some examples, among a group of multiple faces, individual faces may be modified individually, or these modifications may be toggled individually by tapping or selecting an individual face or a series of individual faces displayed within the graphical user interface.

[0067] The story table (314) stores data regarding collections of messages and associated image, video, or audio data that are compiled into collections (e.g., stories or galleries). The creation of a specific collection may be initiated by a specific user (e.g., each user for whom a record is maintained in the entity table (306)). The user may create a “personal story” in the form of a collection of content created and transmitted / broadcasted by that user. To this end, the user interface of the messaging client (104) may include a user-selectable icon to enable the sending user to add specific content to their personal story.

[0068] The collection may also constitute a “live story,” which is a collection of content from multiple users created manually, automatically, or using a combination of manual and automatic techniques. For example, a “live story” may constitute a curated stream of user-submitted content from various locations and events. Users who have location-enabled client devices and are at a common location event at a specific time may be presented with an option to contribute content to a specific live story, for example, through the user interface of a messaging client (104). A live story may be identified to the user by the messaging client (104) based on their location. The final result is a “live story” as described in the community.

[0069] An additional type of content collection is known as a “location story” that enables a user with a client device (102) located within a specific geographic location (e.g., a college or university campus) to contribute to a specific collection. In some examples, contribution to a location story may require a second degree of authentication to verify whether the end user belongs to a specific organization or other entity (e.g., a student on a university campus).

[0070] As mentioned above, the video table (304) stores video data associated with messages for which records are maintained in the message table (302), in one example. Similarly, the image table (312) stores image data associated with messages for which message data is stored in the entity table (306). The entity table (306) can associate various augmentations from the augmentation table (310) with various images and videos stored in the image table (312) and the video table (304).

[0071] Data communication architecture

[0072] FIG. 4 is a schematic diagram illustrating the structure of a message (400) according to some examples, generated by a messaging client (104) for communication with an additional messaging client (104) or a messaging server (118). The content of a particular message (400) is used to populate a message table (302) stored in a database (126) accessible by the messaging server (118). Similarly, the content of the message (400) is stored in memory as "in-transit" or "in-flight" data of the client device (102) or application servers (114). The message (400) is illustrated as comprising the following exemplary components:

[0073] ● Message identifier (402): A unique identifier that identifies the message (400).

[0074] ● Message text payload (404): Text generated by the user through the user interface of the client device (102) and included in the message (400).

[0075] ● Message image payload (406): Image data included in the message (400) that is captured by the camera component of the client device (102) or retrieved from the memory component of the client device (102). Image data for the transmitted or received message (400) may be stored in an image table (312).

[0076] ● Message video payload (408): Video data captured by the camera component or retrieved from the memory component of the client device (102) and included in the message (400). Video data for the transmitted or received message (400) may be stored in the video table (304).

[0077] ● Message audio payload (410): Audio data captured by a microphone or retrieved from a memory component of a client device (102) and included in a message (400).

[0078] ● Message augmentation data (412): Augmentation data (e.g., filters, stickers, or other annotations or enhancements) representing augmentations to be applied to the message image payload (406), message video payload (408), or message audio payload (410) of the message (400). Augmentation data for a transmitted or received message (400) may be stored in an augmentation table (310).

[0079] ● Message duration parameter (414): A parameter value indicating the amount of time in seconds during which the content of a message (e.g., message image payload (406), message video payload (408), message audio payload (410)) is presented to or made accessible to the user through the messaging client (104).

[0080] ● Message geolocation parameter (416): Geolocation data (e.g., latitude and longitude coordinates) associated with the content payload of the message. Multiple message geolocation parameter (416) values ​​may be included in the payload, and each of these parameter values ​​is associated with content items included in the content (e.g., a specific image in the message image payload (406), or a specific video in the message video payload (408)).

[0081] ● Message Story Identifier (418): Identifier values ​​that identify one or more content collections (e.g., “stories”) identified in the story table (314) to which a specific content item within the message image payload (406) of the message (400) is associated. For example, multiple images within the message image payload (406) may each be associated with multiple content collections using identifier values.

[0082] ● Message Tag (420): Each message (400) may be tagged with multiple tags, each of which indicates the subject of the content included in the message payload. For example, if a specific image included in the message image payload (406) depicts an animal (e.g., a lion), a tag value indicating the relevant animal may be included in the message tag (420). The tag values ​​may be manually generated based on user input or automatically generated using, for example, image recognition.

[0083] ● Message sender identifier (422): An identifier representing the user of the client device (102) to which the message (400) was created and the message (400) was sent (e.g., a messaging system identifier, an email address, or a device identifier).

[0084] ● Message recipient identifier (424): An identifier representing the user of the client device (102) to which the message (400) is addressed (e.g., a messaging system identifier, an email address, or a device identifier).

[0085] The contents (e.g., values) of the various components of the message (400) may be pointers to locations within tables where content data values ​​are stored. For example, an image value within the message image payload (406) may be a pointer (or its address) to a location within the image table (312). Similarly, values ​​within the message video payload (408) may point to data stored within the video table (304), values ​​within the message augmentations (412) may point to data stored within the augmentation table (310), values ​​within the message story identifier (418) may point to data stored within the story table (314), and values ​​within the message sender identifier (422) and message receiver identifier (424) may point to user records stored within the entity table (306).

[0086] FIG. 5 illustrates a sequence diagram of an exemplary user interface process according to some examples. During the process, a user interface engine (504) creates a user interface (510) comprising one or more virtual objects that constitute the interaction elements of the user interface. The virtual objects can be described as solids of 3D geometry having 3-tuple values ​​of X (horizontal), Y (vertical), and Z (depth). A render of the user interface is created, and the render data (512) is communicated to a graphics engine (506) and displayed to a user (516). The user interface engine (504) creates one or more virtual object colliders for one or more virtual objects (514). At least one camera (502) creates real-world video frame data (520) of the real world as seen by the user (518). The real-world video frame data (520) includes hand position video frame data of one or more of the user's hands within the render of the user interface by the graphics engine (506). Thus, the real-world video frame data (520) includes hand location video frame data and hand position video frame data of the user's hands when the user performs movements with their hands.

[0087] As mentioned in this specification, a collider (e.g., a virtual object collider) refers to a software configuration that may be attached to a specific area of ​​a virtual object to enable tracking the location of the collider and detecting when a collision occurs between the collider and another virtual object (e.g., when the collider intersects with another virtual object). In one example, when a collider is attached to a second virtual object, a collision event may be detected based on determining that the first collider of the first object has intersected with the collider of the second virtual object. As further discussed in this specification, in response to the detection of a collision event, the user interface engine (504) may transmit user interaction data containing such collision event to a specific application (e.g., application (508)) to enable the application to respond in a specific manner (e.g., perform a function or action, etc.).

[0088] The user interface engine (504) utilizes the hand position video frame data and hand position video frame data in the real-world video frame data (520) to extract landmarks (522) of the user's hands from the real-world video frame data (520) and generates landmark colliders (524) for one or more landmarks on one or more of the user's hands. The landmark colliders are used to determine user interactions between the user and virtual objects by detecting collisions (526) between the landmark colliders and the respective virtual object colliders of the virtual objects. The collisions are used by the user interface engine (504) to determine user interactions (528) with virtual objects by the user. The user interface engine (504) communicates user interaction data (530) of user interactions to the application (508) for use by the application (508).

[0089] In some examples, the application (508) utilizes various APIs and system libraries to receive and process real-world video frame data (520) and instructs the graphics engine (506) to perform specific action(s), thereby performing the functions of the user interface engine (504).

[0090] Although the above description relates to an application (508), it is understood that in some embodiments, a messaging client (104), an application (106), or an application (608) (discussed below) may perform the same operations as the application (508).

[0091] FIG. 6 depicts a sequence diagram of an exemplary user interface process according to some examples. At least one camera (604) generates real-world video frame data (610) of the real world as seen by the user (602). In one example, at least one camera (604) may be provided by a specific client device, such as a client device (102). The real-world video frame data (610) includes hand position video frame data of one or more of the user's hands. Thus, the real-world video frame data (610) includes hand position video frame data and hand position video frame data of the user's hands when the user performs movements with their hands. A gesture intent recognition engine (606) generates hand gesture data (614) containing hand gesture categorization information that indicates one or more hand gestures performed by the user by utilizing the hand position video frame data and hand position video frame data in the real-world video frame data (610) (612). The gesture intent recognition engine (606) communicates the hand gesture data (614) to an application (608) that utilizes the hand gesture data (614) as input from a user interface.

[0092] In some examples, the application (608) utilizes various APIs and system libraries to receive and process real-world video frame data (610) from at least one camera (604) and determine hand gesture data (614), thereby performing the functions of the gesture intent recognition engine (606).

[0093] Although the above description relates to an application (608), it is understood that in some embodiments, a messaging client (104), an application (106), or an application (508) may perform the same operations as the application (608).

[0094] Furthermore, it is understood that the user interface engine (504) and the graphics engine (506) discussed above in FIG. 5 can process hand gesture data (614) to perform similar operations discussed above in FIG. 5. For example, the user interface engine (504) can generate render data for a user interface based at least partially on the hand gesture data (614), and the graphics engine (506) can use the generated render data to render this user interface for display.

[0095] FIG. 7a illustrates an exemplary interface according to various embodiments. The exemplary interface may be provided for display on a client device (e.g., client device (102)) through the interface(s) of, for example, a messaging client (104), an application (106), and an application (608).

[0096] As described, the interface (702) includes real-world video frame data of the real world captured by the camera (604). The real-world video frame data includes one or more of the user's hands within the render of the interface (702) and additionally includes hand position video frame data and hand position video frame data of the user's hands when the user performs movements with their hands.

[0097] As previously discussed, the gesture intent recognition engine (606) generates hand gesture data including hand gesture categorization information that displays one or more hand gestures performed by a user by utilizing hand position video frame data and hand position video frame data in real-world video frame data. In one implementation, the gesture intent recognition engine (606) communicates the hand gesture data to an application that utilizes the hand gesture data as input from a user interface.

[0098] FIG. 7b illustrates an exemplary interface according to various embodiments. The exemplary interface may be provided for display on a client device (e.g., client device (102)) through the interface(s) of, for example, a messaging client (104), an application (106), an application (508), or an application (608). The examples described in the following discussion regarding FIG. 7b are a continuation of the above discussion regarding FIG. 7a.

[0099] In the example of FIG. 7b, the gesture intent recognition engine (606) analyzes real-world video frame data shown in the interface (704) to locate the hand position video frame data and hand position video frame data of the user's hand when the user performs movements with their hand. In this example, the gesture intent recognition engine (606) utilizes the hand position video frame data and hand position video frame data in the real-world video frame data to generate hand gesture data that includes hand gesture categorization information indicating a hand gesture performed by the user to indicate a gesture to start or stop recording by the camera (604). As illustrated in FIG. 7b, the start or stop recording gesture (706) corresponds to a gesture in which the user's hand is raised and open to show the palm. Thus, the start recording gesture corresponds to the first gesture of the aforementioned movements and positions, and the stop recording gesture corresponds to the second gesture of the same movements and positions.

[0100] In some examples, the user may perform a recording start / stop gesture by positioning the user's palm over a selectable graphic item shown in the interface (704). This selectable graphic item may be a graphic representation of a button, such as a recording start button or a recording stop button, or text information displayed as such (e.g., "Start", "Stop", etc.).

[0101] In one embodiment, the application (608) (or the messaging client (104), application (106), and application (508)) starts recording when it determines a recording start gesture from the hand gesture data, and stops recording when it determines a recording stop gesture from the hand gesture data. In this way, a "hands-free" approach to starting or stopping recording can be provided without the user having to return to the client device and perform these actions on the screen of the client device (e.g., via touch or tap gestures or input).

[0102] FIG. 8 illustrates exemplary interfaces according to various embodiments. The exemplary interfaces may be provided for display on a client device (e.g., client device (102)) through the interface(s) of, for example, a messaging client (104), an application (106), an application (508), or an application (608).

[0103] As illustrated, the interface (802) provides a graphic item (806), a graphic area (808), a graphic item (810), and a graphic item (812). In this example, the graphic area (808) includes a set of frames corresponding to real-world video frame data, and the set of frames is displayed as a timeline showing a chronological sequence of these frames. The graphic item (806) corresponds to a cut point where real-world video frame data can be trimmed between a start point corresponding to the graphic item (810) and an end point corresponding to the graphic item (812) (e.g., frames after the cut point are cut, discarded, or deleted).

[0104] As illustrated, the interface (804) shows that the graphic item (806) from the interface (802) has been moved to a later point in the frame sequence corresponding to the graphic item (814). The application (608) may perform a trimming operation to remove a second set of frames from the cut point corresponding to the graphic item (814). In one embodiment, this cut point is determined in an automated manner by the application (608) based on determining a point corresponding to the frame indicating the start of a record stop gesture within the sequence of frames, and then setting the cut point a threshold amount of frames before the frame indicating the start of the record stop gesture. In the implementation, it is understood that the threshold amount of frames may be a predetermined number of frames, such as five (5) frames, but the number may be set to any number of frames. In one embodiment, the user may further adjust or modify the cut point that moves the position of the graphic item (814) using touch inputs or gestures on the screen of the client device displaying the interface (804).

[0105] Although the above description relates to an application (608), it is understood that in some embodiments, a messaging client (104), an application (106), or an application (508) may perform the same operations as the application (608).

[0106] FIG. 9 illustrates exemplary interfaces according to various embodiments. An exemplary interface may be provided for display on a client device (e.g., client device (102)) through the interface(s) of, for example, a messaging client (104), an application (106), an application (508), or an application (608).

[0107] The examples in FIG. 9 illustrate embodiments for interacting with other virtual objects using a 2-D cursor.

[0108] As illustrated, the interface (902) displays real-world video frame data of the real world captured by the camera (502). The real-world video frame data includes one or more of the user's hands within the render of the interface (902) by the graphics engine (506), and additionally includes hand position video frame data and hand position video frame data of the user's hands when the user performs movements with their hands. Based on the location and position (910) of the representation of the user's fingers in the real-world video frame data, the user interface engine (504) creates a virtual object (912) corresponding to a cursor for interacting with other virtual objects rendered for display in the interface (902).

[0109] In this example, the user interface engine (504) enables the user to interact with a set of other virtual objects, including the virtual object (914), using the virtual object (912). The set of other virtual objects in this example is arranged in a grid-like configuration in which each virtual object is positioned with an equal spatial amount between adjacent virtual objects(s). In an implementation, the user interface engine (504) generates virtual object colliders for each of the virtual objects. In an implementation, the user interface engine (504) restricts the movements of the cursor corresponding to the virtual object (912) to a single control axis, such as the x-axis (e.g., horizontal). As a result, the virtual object (912) can select only one specific virtual object from a set of other virtual objects within the interface (902), such as virtual objects from a single "row" as arranged in the example of FIG. 9. As described, the virtual object (912) corresponding to the cursor selects the virtual object (914) based on detecting a collision event in which the first collider of the virtual object (912) intersects the second collider of the virtual object (914). In response to the collision event, the user interface engine (504) can animate the virtual object (914) in a specific manner and cause the graphics engine (506) to render the interface (902) accordingly. Furthermore, the user interface engine (504) transmits user interaction data containing data related to the collision event to a specific application (e.g., application (508)), wherein the application can perform a function or action in response to the collision event.

[0110] As illustrated in another example, the rendering of the interface (904) by the graphics engine (506) exemplifies that the position and location (930) of the user's finger has changed from the interface (902). Based on the changed position and location (910) of the representation of the user's finger in real-world video frame data, the user interface engine (504) moves the cursor from the position and location in the interface (902) to the position and location corresponding to the virtual object (932) in the interface (904). In the interface (904), the cursor corresponding to the virtual object (932) overlaps with the virtual object (934) to indicate a collision, and the user interface engine (504) can animate the virtual object (934) in response to detecting the collision event. Furthermore, the user interface engine (504) transmits user interaction data containing data related to the collision event to a specific application (e.g., application (508)), whereby the application can perform a function or action in response to the collision event.

[0111] Although the above description relates to an application (508), it is understood that in some embodiments, a messaging client (104), an application (106), or an application (608) may perform the same operations as the application (508).

[0112] FIG. 10 illustrates exemplary interfaces according to various embodiments. The exemplary interface may be provided for display on a client device (e.g., client device (102)) through the interface(s) of, for example, a messaging client (104), an application (106), an application (508), or an application (608). In the following discussion, the examples described in FIG. 10 are a continuation of the discussion in FIG. 9 above.

[0113] The examples in FIG. 10 illustrate embodiments for interacting with other virtual objects using a 3-D cursor.

[0114] As illustrated, the interface (1002) displays real-world video frame data of the real world captured by the camera (502). The real-world video frame data includes one or more of the user's hands within the render of the interface (1002) by the graphics engine (506), and additionally includes hand position video frame data and hand position video frame data of the user's hands when the user performs movements with their hands. Based on the position and position (1010) of the representation of the user's fingers in the real-world video frame data, the user interface engine (504) generates a virtual object (1012) corresponding to a cursor for interacting with other virtual objects rendered for display in the interface (1002).

[0115] In this example, the user interface engine (504) enables the user to use the virtual object (1012) to interact with a set of other virtual objects, including the virtual object (1014). The set of other virtual objects in this example is arranged in a grid-like configuration in which each virtual object is positioned with an equal spatial amount between adjacent virtual object(s). In an implementation, the user interface engine (504) generates virtual object colliders for each of the virtual objects. In an implementation, the user interface engine (504) restricts the movements of the cursor corresponding to the virtual object (1012) to two control axes (e.g., more than one axis), such as the x-axis (e.g., horizontal) and the y-axis (e.g., vertical). As a result, the virtual object (1012) can select a specific virtual object from a set of other virtual objects within the interface (1002), such as a virtual object from a top "row" as arranged in the example of FIG. 10. As described, the virtual object (1012) corresponding to the cursor selects the virtual object (1014) based on detecting a collision event in which the first collider of the virtual object (1012) intersects the second collider of the virtual object (1014). In response to the collision event, the user interface engine (504) can animate the virtual object (1014) in a specific manner and cause the graphics engine (506) to render the interface (1002) accordingly. Furthermore, the user interface engine (504) transmits user interaction data containing data related to the collision event to a specific application (e.g., application (508)), wherein the application can perform a function or action in response to the collision event.

[0116] As illustrated in another example, rendering of the interface (1004) by the graphics engine (506) exemplifies that the position and location (1030) of the user's finger has changed from the interface (902). Based on the changed position and location (910) of the representation of the user's finger in real-world video frame data, the user interface engine (504) moves the cursor from the position and location in the interface (1002) to the position and location corresponding to the virtual object (1032) in the interface (1004) in the lower row from the previously selected virtual object (1014). In the interface (1004), the cursor corresponding to the virtual object (1032) overlaps with the virtual object (1034) to indicate a collision, and the user interface engine (504) can animate the virtual object (1034) in response to detecting the collision event. Furthermore, the user interface engine (504) transmits user interaction data containing data related to a crash event to a specific application (e.g., application (508)), wherein the application may perform a function or action in response to the crash event. For example, the function or action may be a task related to navigation, games, creation, etc.

[0117] Although the above description relates to an application (508), it is understood that in some embodiments, a messaging client (104), an application (106), or an application (608) may perform the same operations as the application (508).

[0118] FIG. 11 illustrates exemplary interfaces according to various embodiments. An exemplary interface may be provided for display on a client device (e.g., client device (102)) through the interface(s) of, for example, a messaging client (104), an application (106), an application (508), or an application (608).

[0119] As described, the interface (1102) includes real-world video frame data of the real world captured by the camera (502). The real-world video frame data includes one or more of the user's hands within the render of the interface (902) by the graphics engine (506), and additionally includes hand position video frame data and hand position video frame data of the user's hands when the user performs movements with their hands. Based on the position and position (1110) of the representation of the user's finger (e.g., thumb) in the real-world video frame data and the second position and second position of the representation of the user's second finger (e.g., index finger), the user interface engine (504) detects a firearm-like or shooting hand sign.

[0120] Rendering of the interface (1104) by the graphics engine (506) exemplifies that the position and location (1130) of the user's first finger (e.g., thumb) has been changed from the interface (1102). Based on the changed position and location of the representation of the user's finger in real-world video frame data, the user interface engine (504) determines the direction that the user's second finger (e.g., index finger) is pointing and performs a ray casting technique to determine a vector or path for animating virtual objects (e.g., virtual bullets or lasers, etc.) that follow the user's second finger. As illustrated in this example, the graphics engine (506) renders virtual objects (1132) and virtual objects (1134) in the interface (1104) and animates these virtual objects to follow the path discussed above.

[0121] FIG. 12 illustrates exemplary interfaces according to various embodiments. The exemplary interface may be provided for display on a client device (e.g., client device (102)) through the interface(s) of, for example, a messaging client (104), an application (106), an application (508), or an application (608). The examples described in FIG. 12 below are a continuation of the discussion from FIG. 11.

[0122] As described, the interface (1202) includes real-world video frame data of the real world captured by the camera (502). The real-world video frame data includes one or more of the user's hands within the render of the interface (1202) by the graphics engine (506), and additionally includes hand position video frame data and hand position video frame data of the user's hands when the user performs movements with their hands. Based on the position and position (1210) of the representation of the user's second finger (e.g., index finger) in the real-world video frame data, the user interface engine (504) detects that the position and position of the user's second finger has changed from the interface (1102).

[0123] After such detection, the user interface engine (504) determines the direction in which the user's second finger (e.g., index finger) is pointing and performs a ray casting technique to determine a vector or path for animating virtual objects (e.g., virtual bullets or lasers, etc.) that follow from the user's first finger. As illustrated in this example, the graphics engine (506) renders the virtual object (1212) in the interface (1202) and animates the virtual object to follow the path discussed above.

[0124] Rendering of the interface (1204) by the graphics engine (506) exemplifies that the position and location (1230) of the user's second finger (e.g., index finger) has been changed from the interface (1202). Based on the changed position and location of the representation of the user's second finger in real-world video frame data, the user interface engine (504) determines the new direction that the user's second finger (e.g., index finger) is pointing and determines a vector or path for animating virtual objects following the user's second finger by performing a ray casting technique. As illustrated in this example, the graphics engine (506) renders a virtual object (1232) in the interface (1204) and animates this virtual object to follow the path discussed above.

[0125] FIG. 13 illustrates exemplary interfaces according to various embodiments. An exemplary interface may be provided for display on a client device (e.g., client device (102)) through the interface(s) of, for example, a messaging client (104) or an application (106).

[0126] As illustrated, the interface (1302) includes real-world video frame data of the real world captured by the camera (502). The real-world video frame data includes one or more of the user's hands within the render of the interface (1302) by the graphics engine (506), and additionally includes hand position video frame data and hand position video frame data of the user's hands when the user performs movements with their hands. Based on the position and position (1310) of the representation of the user's finger (e.g., thumb) in the real-world video frame data and the second position and second position of the representation of the user's second finger (e.g., index finger), the user interface engine (504) creates a virtual object (1312) corresponding to a cursor for interacting with other virtual objects rendered for display in the interface (1102). As further illustrated, the user interface engine (504) creates virtual objects (1314) and virtual objects (1316). A more detailed view of the user's hand is illustrated in the graphics area (1318). In one embodiment, the graphic area (1318) is hidden from the interface (1302) or is not displayed on the interface (1302).

[0127] As described, the virtual object (1312) corresponding to the cursor selects the virtual object (1316) based on detecting a collision event in which the first collider of the virtual object (1312) intersects the second collider of the virtual object (1316). In response to the collision event, the user interface engine (504) may attach the virtual object (1316) to the virtual object (1312) and cause the graphics engine (506) to render these two virtual objects accordingly (e.g., to make them appear connected to each other). In this example, the virtual object (1316) is a slider control that enables the selection of a specific color from a color palette represented by the virtual object (1314), and includes changing the selected color according to the position of the virtual object (1316) relative to the color palette. In particular, the selected color is displayed inside the virtual object (1316).

[0128] Rendering of the interface (1304) by the graphics engine (506) exemplifies that the position and location (1330) of the user's first finger (e.g., thumb) and the virtual object (1332) have changed from the interface (1302). In response to the changed position and location of the user's first finger (e.g., the thumb is now "closed" or adjacent to the second finger), the user interface engine (504) transmits user interaction data, including data related to the collision event, to a specific application (e.g., application (508)), wherein the application may perform a function or action in response to the collision event (e.g., activating a function such as a color selection corresponding to the position of the virtual object (1316) within the color palette provided by the virtual object (1314)). In this example, the virtual object (1336) is the same object as the virtual object (1316), and the virtual object (1332) is the same object as the virtual object (1312), each of which has a different position from the previous position in the interface (1302). A more detailed view of the user's hand is shown in the graphic area (1338).

[0129] In response to the changed position and position (1330) of the user's hand, the graphics engine (506) renders the virtual object (1336) and virtual object (1332) in the interface (1304) at a position slightly different from the previous position of the virtual object (1316) and virtual object (1312) in the interface (1302), so that the color selected from the color palette is different from the color shown on the virtual object (1316) in the interface (1302).

[0130] The above examples of FIG. 13 discuss a user's first finger (e.g., acting as a trigger when the thumb is "closed" or next to the second finger) as causing a selection using a virtual object (1336), but in some embodiments, another finger (e.g., index finger) may act as a "trigger" finger to cause a selection or initiate some functionality or action.

[0131] FIG. 14 illustrates exemplary interfaces according to various embodiments. The exemplary interface may be provided for display on a client device (e.g., client device (102)) through the interface(s) of, for example, a messaging client (104), an application (106), an application (508), or an application (608). The examples described in FIG. 14 below are a continuation of the discussion from FIG. 13.

[0132] Rendering of the interface (1402) by the graphics engine (506) exemplifies that the position and location (1410) of the user's hand and the virtual object (1412) have changed from the interface (1304). In response to the changed position and location (1410) of the user's hand and the virtual object (1412), the graphics engine (506) renders the virtual object (1416) and the virtual object (1412) in the interface (1402) at a position different from the previous position of the virtual object (1336) and the virtual object (1332) in the interface (1304), so that the color selected from the color palette is different from the color shown on the virtual object (1336) in the interface (1304). A more detailed view of the user's hand is shown in the graphics area (1418).

[0133] In response to the changed position and location of the user's hand, the user interface engine (504) transmits user interaction data to perform color selection. In this example, virtual object (1416) is the same object as virtual object (1336), virtual object (1412) is the same object as virtual object (1332), and each of these has a position different from the previous position in the interface (1304).

[0134] Rendering of the interface (1404) by the graphics engine (506) exemplifies that the position and location (1430) of the user's hand has changed from the interface (1402). In response to the changed position and location (1430) of the user's hand, the graphics engine (506) renders the virtual object (1436) and virtual object (1432) in the interface (1404) at a position different from the previous position of the virtual object (1416) and virtual object (1412) in the interface (1402), so that, as a result, the color selected from the color palette becomes different from the color shown on the virtual object (1416) in the interface (1402). A more detailed view of the user's hand is shown in the graphics area (1438).

[0135] In response to the changed position and location of the user's hand, the user interface engine (504) transmits user interaction data to perform color selection. In this example, virtual object (1436) is the same object as virtual object (1416), virtual object (1432) is the same object as virtual object (1412), and each of these has a position different from the previous position in the interface (1402).

[0136] FIG. 15 illustrates exemplary interfaces according to various embodiments. The exemplary interface may be provided for display on a client device (e.g., client device (102)) through the interface(s) of, for example, a messaging client (104), an application (106), an application (508), or an application (608). The examples described in FIG. 15 below are a continuation of the discussion from FIG. 14.

[0137] Rendering of the interface (1502) by the graphics engine (506) exemplifies that the position and location (1510) of the user's hand and the virtual object (1512) have changed from the interface (1404). In response to the changed position and location (1510) of the user's hand and the virtual object (1512), the graphics engine (506) renders the virtual object (1516) and the virtual object (1512) in the interface (1502) at a position different from the previous position of the virtual object (1436) and the virtual object (1432) in the interface (1404), so that the color selected from the color palette is different from the color shown on the virtual object (1436) in the interface (1404). A more detailed view of the user's hand is shown in the graphics area (1518).

[0138] In response to the changed position and location of the user's hand, the user interface engine (504) transmits user interaction data to perform color selection. In this example, virtual object (1516) is the same object as virtual object (1436), virtual object (1512) is the same object as virtual object (1432), and each of these has a position different from the previous position in the interface (1404).

[0139] A rendering of the interface (1504) by the graphics engine (506) illustrates that the position and location (1530) of the user's first finger (e.g., thumb) has been changed from the interface (1502). In response to the changed position and location (1430) of the user's first finger, the user interface engine (504) transmits user interaction data to perform a final color selection from a color palette to match the color shown on the virtual object (1516) in the interface (1502). A more detailed view of the user's hand is shown in the graphics area (1538).

[0140] FIG. 16 illustrates exemplary interfaces according to various embodiments. An exemplary interface may be provided for display on a client device (e.g., client device (102)) through the interface(s) of, for example, a messaging client (104), an application (106), an application (508), or an application (608).

[0141] As illustrated, the interface (1602) includes real-world video frame data of the real world captured by the camera (604). The real-world video frame data includes one or more of the user's hands within the render of the interface (1602) and additionally includes hand position video frame data and hand position video frame data of the user's hands when the user performs movements with their hands. Additionally, as illustrated, the interface (1602) includes a virtual object (1614) that can be manipulated using various gestures. The render of the interface (1602) illustrates the position and position (1610) of the user's first hand and the position and position (1612) of the user's second hand.

[0142] As previously discussed, the gesture intent recognition engine (606) generates hand gesture data including hand gesture categorization information that indicates one or more hand gestures performed by a user, utilizing hand position video frame data and hand position video frame data in real-world video frame data. In one implementation, the gesture intent recognition engine (606) communicates the hand gesture data to an application that utilizes the hand gesture data as input from a user interface. Since this hand gesture data discussed in the examples of FIGS. 16, 17, and 18 includes data for both hands of the user, the gesture intent recognition engine (606) is able to recognize multi-gesture movements involving both hands.

[0143] In the example of FIG. 16, hand gesture data indicates that a virtual object (1614) will be selected to interact with and respond to the user's hand movements. This virtual object is utilized by an application to perform a set of actions based at least partially on the position and location of the virtual object, including the size of the object.

[0144] As additionally illustrated, the interface (1604) includes the changed position and position (1630) of the user's first hand and the position and position (1632) of the user's second hand (e.g., each thumb of each hand is closed or closer to the index finger), and the gesture intent recognition engine (606) generates updated hand gesture data by utilizing the hand position video frame data and the hand position video frame data. In this other example of FIG. 16, the updated hand gesture data indicates that the virtual object (1614) is expanded to a larger size, which is illustrated as the virtual object (1634) in the render of the interface (1604).

[0145] FIG. 17 illustrates exemplary interfaces according to various embodiments. The exemplary interface may be provided for display on a client device (e.g., client device (102)) through the interface(s) of, for example, a messaging client (104), an application (106), an application (508), or an application (608). The examples described in FIG. 17 below are a continuation of the discussion from FIG. 16.

[0146] As illustrated, the interface (1702) includes the changed position and position of the user's first hand (1710) and the position and position of the user's second hand (1712), and the gesture intent recognition engine (606) generates updated hand gesture data by utilizing the hand position video frame data and the hand position video frame data. In this example of FIG. 17, the updated hand gesture data indicates that the virtual object (1634) in FIG. 16 is rotated in a specific direction, which is illustrated as the virtual object (1714) in the render of the interface (1702).

[0147] As additionally illustrated, the interface (1704) includes the changed position and position of the user's first hand (1730) and the position and position of the user's second hand (1732), and the gesture intent recognition engine (606) generates updated hand gesture data by utilizing the hand position video frame data and the hand position video frame data. In this other example of FIG. 17, the updated hand gesture data indicates that the virtual object (1714) is rotated in a different direction than in the interface (1702), which is illustrated as the virtual object (1734) in the render of the interface (1704).

[0148] FIG. 18 illustrates exemplary interfaces according to various embodiments. The exemplary interface may be provided for display on a client device (e.g., client device (102)) through the interface(s) of, for example, a messaging client (104), an application (106), an application (508), or an application (608). The examples described in FIG. 18 below are a continuation of the discussion from FIG. 17.

[0149] As illustrated, the interface (1802) includes the changed position and location (1810) of the user's first hand and the position and location (1812) of the user's second hand, and the gesture intent recognition engine (606) generates updated hand gesture data by utilizing the hand position video frame data and the hand position video frame data. In this example of FIG. 18, the updated hand gesture data indicates that the virtual object (1734) in FIG. 17 is expanded to full size, which is illustrated as the virtual object (1814) in the render of the interface (1802).

[0150] As additionally illustrated, the interface (1804) includes the changed position and location (1830) of the user's first hand and the position and location (1832) of the user's second hand, and the gesture intent recognition engine (606) generates updated hand gesture data by utilizing the hand position video frame data and the hand position video frame data. In this other example of FIG. 18, the updated hand gesture data indicates that the virtual object (1814) should be reduced from the full size of the interface (180202), which is illustrated as the virtual object (1834) in the render of the interface (1804).

[0151] FIG. 19 illustrates exemplary interfaces according to various embodiments. An exemplary interface may be provided for display on a client device (e.g., client device (102)) through the interface(s) of, for example, a messaging client (104), an application (106), an application (508), or an application (608).

[0152] As illustrated, the interface (1902) includes real-world video frame data of the real world captured by the camera (604). The real-world video frame data includes one or more of the user's hands within the render of the interface (1902) and additionally includes hand position video frame data and hand position video frame data of the user's hands when the user performs movements with their hands. Additionally, as illustrated, the interface (1902) includes a virtual object (1912) that can be manipulated using various gestures. The render of the interface (1902) exemplifies the position and position (1910) of the user's first hand.

[0153] As previously discussed, the gesture intent recognition engine (606) generates hand gesture data including hand gesture categorization information that displays one or more hand gestures performed by a user by utilizing hand position video frame data and hand position video frame data in real-world video frame data. In one implementation, the gesture intent recognition engine (606) communicates the hand gesture data to an application that utilizes the hand gesture data as input from a user interface.

[0154] In the example of FIG. 19, hand gesture data indicates that a part of the virtual object (1912) will be selected as a starting point for modifying the virtual object (912) based at least partially on the user's hand movements. This virtual object is utilized by an application to perform a set of actions based at least partially on the location and position of the virtual object, including other characteristics of the object.

[0155] As additionally illustrated, the interface (1904) includes the position and location (1930) of the user's first hand that has been changed, and the gesture intent recognition engine (606) generates updated hand gesture data by utilizing the hand position video frame data and the hand position video frame data. In this other example of FIG. 19, the updated hand gesture data indicates that a virtual object (1912) needs to be modified, which is illustrated as a virtual object (1932) in the render of the interface (1904), wherein the virtual object (1932) includes additional graphic data extended from a starting point in the interface (1902).

[0156] FIG. 20 illustrates exemplary interfaces according to various embodiments. The exemplary interface may be provided for display on a client device (e.g., client device (102)) through the interface(s) of, for example, a messaging client (104), an application (106), an application (508), or an application (608). The examples described in FIG. 20 below are a continuation of the discussion from FIG. 19.

[0157] As illustrated, the interface (2002) includes the position and location (2010) of the user's first hand that has been changed, and the gesture intent recognition engine (606) generates updated hand gesture data by utilizing the hand position video frame data and the hand position video frame data. In this example of FIG. 20, the updated hand gesture data indicates that a second part of the virtual object (1932) is selected as a starting point for modifying the virtual object (1932) based at least partially on the user's hand movements, which is illustrated as the virtual object (2012) in the render of the interface (2002).

[0158] As additionally illustrated, the interface (2004) includes the position and location (2030) of the user's first hand that has been changed, and the gesture intent recognition engine (606) generates updated hand gesture data by utilizing the hand position video frame data and the hand position video frame data. In this other example of FIG. 20, the updated hand gesture data indicates that a virtual object (2012) needs to be modified, which is illustrated as a virtual object (2032) in the render of the interface (2004), wherein the virtual object (2032) includes second additional graphic data extended from a second starting point in the interface (2002).

[0159] FIG. 21 illustrates exemplary interfaces according to various embodiments. The exemplary interface may be provided for display on a client device (e.g., client device (102)) through the interface(s) of, for example, a messaging client (104), an application (106), an application (508), or an application (608). The examples described in FIG. 21 below are a continuation of the discussion from FIG. 20.

[0160] As illustrated, the interface (2102) includes the position and location (2110) of the user's second hand that has been changed, and the gesture intent recognition engine (606) generates updated hand gesture data by utilizing the hand position video frame data and the hand position video frame data. In this example of FIG. 21, the updated hand gesture data indicates that a third part of the virtual object (2032) in FIG. 20 is selected as a third starting point for modifying the virtual object (2032) based at least partially on the user's hand movements, which is illustrated as the virtual object (2112) in the render of the interface (2102).

[0161] As additionally illustrated, the interface (2104) includes the position and location (2130) of the user's second hand that has been changed, and the gesture intent recognition engine (606) generates updated hand gesture data by utilizing the hand position video frame data and the hand position video frame data. In this other example of FIG. 21, the updated hand gesture data indicates that a virtual object (2112) needs to be modified, which is illustrated as a virtual object (2132) in the render of the interface (2104), wherein the virtual object (2132) includes third additional graphic data extended from a third starting point in the interface (2102).

[0162] FIG. 22 illustrates exemplary interfaces according to various embodiments. An exemplary interface may be provided for display on a client device (e.g., client device (102)) through the interface(s) of, for example, a messaging client (104), an application (106), an application (508), or an application (608).

[0163] As illustrated, the interface (2202) includes real-world video frame data of the real world captured by the camera (604). The real-world video frame data includes one or more of the user's hands within the render of the interface (2202) and additionally includes hand position video frame data and hand position video frame data of the user's hands when the user performs movements with their hands. Additionally, as illustrated, the interface (2202) includes a virtual object (2214) and a virtual object (2216) that can interact with various gestures and be moved using various gestures. The render of the interface (2202) illustrates the position and position (2210) of the user's first hand and the position and position (2212) of the user's second hand.

[0164] As previously discussed, the gesture intent recognition engine (606) generates hand gesture data including hand gesture categorization information that indicates one or more hand gestures performed by a user, utilizing hand position video frame data and hand position video frame data in real-world video frame data. In one implementation, the gesture intent recognition engine (606) communicates the hand gesture data to an application that utilizes the hand gesture data as input from a user interface. Since this hand gesture data discussed in the examples of FIGS. 22, 23, 24, and 25 includes data for both hands of the user, the gesture intent recognition engine (606) is able to recognize multi-gesture movements involving both hands or gesture movements involving one hand.

[0165] In the example of FIG. 22, hand gesture data indicates that additional virtual objects will be displayed in response to the user's hand movements. Each virtual object is selected based on the user's hand movements and then utilized by the application to perform a set of actions.

[0166] As additionally illustrated, the interface (2204) includes the changed position and position (2230) of the user's first hand and the position and position (2232) of the user's second hand (e.g., each hand moves outward toward the edges of the interface (2204)), and the gesture intent recognition engine (606) generates updated hand gesture data by utilizing the hand position video frame data and the hand position video frame data. In this other example of FIG. 22, the updated hand gesture data causes additional virtual objects to be rendered for display, which are illustrated in the render of the interface (2204) as virtual object (2234), virtual object (2235), virtual object (2236), virtual object (2238), and virtual object (2240).

[0167] FIG. 23 illustrates exemplary interfaces according to various embodiments. The exemplary interface may be provided for display on a client device (e.g., client device (102)) through the interface(s) of, for example, a messaging client (104), an application (106), an application (508), or an application (608). The examples described in FIG. 23 below are a continuation of the discussion from FIG. 22.

[0168] As illustrated, the interface (2302) includes the changed position and position (2310) of the user's first hand and the position and position (2312) of the user's second hand (e.g., an open hand gesture in which each thumb is spread apart from the other fingers in the same hand), and the gesture intent recognition engine (606) generates updated hand gesture data by utilizing the hand position video frame data and the hand position video frame data. In this example of FIG. 23, the updated hand gesture data indicates that virtual objects are set to remain fixed or stay at their current positions and positions in the interface (2302).

[0169] As additionally illustrated, the interface (2304) includes the position and location (2330) of the user's first hand that has been changed, and the gesture intent recognition engine (606) generates updated hand gesture data by utilizing the hand position video frame data and the hand position video frame data. In this other example of FIG. 23, the updated hand gesture data indicates that a virtual object (2216) is selected to cause the application to perform a set of corresponding actions.

[0170] FIG. 24 illustrates exemplary interfaces according to various embodiments. The exemplary interface may be provided for display on a client device (e.g., client device (102)) through the interface(s) of, for example, a messaging client (104), an application (106), an application (508), or an application (608). The examples described in FIG. 24 below are a continuation of the discussion from FIG. 23.

[0171] As illustrated, the interface (2402) includes the position and location (2410) of the user's first hand and the position and location (2412) of the user's second hand, which have been changed (e.g., from an open hand to a pinch gesture for two hands), and the gesture intent recognition engine (606) generates updated hand gesture data by utilizing the hand position video frame data and the hand position video frame data. In this example of FIG. 24, the updated hand gesture data indicates that virtual objects will be moved in response to the user's hand movements.

[0172] As additionally illustrated, the interface (2404) includes the changed position and position (2430) of the user's first hand and the position and position (2432) of the user's second hand (e.g., the second hand moved toward the upper edge of the interface (2404) and the first hand moved downward toward the lower edge of the interface (2404), and the gesture intent recognition engine (606) generates updated hand gesture data by utilizing the hand position video frame data and the hand position video frame data. In this other example of FIG. 22, the updated hand gesture data causes virtual objects to be positioned in a vertical arrangement in the render of the interface (2404).

[0173] FIG. 25 illustrates exemplary interfaces according to various embodiments. The exemplary interface may be provided for display on a client device (e.g., client device (102)) through the interface(s) of, for example, a messaging client (104), an application (106), an application (508), or an application (608). The examples described in FIG. 25 below are a continuation of the discussion from FIG. 24.

[0174] As illustrated, the interface (2502) includes the changed position and location (2510) of the user's first hand and the position and location (2512) of the user's second hand, and the gesture intent recognition engine (606) generates updated hand gesture data by utilizing the hand position video frame data and the hand position video frame data. In this example of FIG. 25, the updated hand gesture data indicates that virtual objects will be moved in response to the user's hand movements. In this example of FIG. 25, the updated hand gesture data causes the virtual objects to be positioned in an arched and horizontal arrangement in the render of the interface (2502).

[0175] As additionally illustrated, the interface (2504) includes the changed position and position (2530) of the user's first hand and the position and position (2532) of the user's second hand (e.g., the second hand has moved toward the center of the interface (2504) and the first hand has moved toward the center of the interface (2504)), and the gesture intent recognition engine (606) generates updated hand gesture data by utilizing the hand position video frame data and the hand position video frame data. In this other example of FIG. 25, the updated hand gesture data causes other virtual objects to be removed from the render of the interface (2504) and only the virtual object (2214) remains in the interface (2504).

[0176] FIG. 26 is a flowchart illustrating a method (2600) according to specific exemplary embodiments. The method (2600) may be embodied in computer-readable instructions for execution by one or more computer processors so that the operations of the method (2600) can be performed, in part or wholly, by a messaging client (104) in relation to the respective components described above in FIG. 5 and 6, or an application (e.g., application (106)) running on a given client device (e.g., client device (102)) communicating with the messaging server system (108) and its components; thus, the method (2600) is described below by way of example with reference to this. However, it will be understood that at least some of the operations of the method (2600) may be placed in various other hardware configurations, and that the method (2600) is not intended to be limited to the messaging client (104) or any of the components or systems mentioned above.

[0177] In the embodiment, the operations described in FIG. 26 correspond to the descriptions of at least FIG. 13, FIG. 14, and FIG. 15, as discussed above.

[0178] In operation 2602, the messaging client (104) detects a first gesture corresponding to an open trigger finger gesture from a set of frames.

[0179] In operation 2604, the messaging client (104) detects the position and location of the expression of the finger from the open trigger finger gesture.

[0180] In operation 2606, the messaging client (104) creates a first virtual object based at least partially on the position and location of the finger representation, and the first virtual object is extended from the finger representation.

[0181] In operation 2608, the messaging client (104) detects a first collision event corresponding to the first collider of the first virtual object crossing the second collider of the second virtual object.

[0182] In operation 2610, the messaging client (104) detects a second gesture corresponding to a closed trigger finger gesture from a second set of frames.

[0183] In operation 2612, the messaging client (104) selects a second virtual object in response to the first collision event and the detected second gesture.

[0184] In operation 2614, the messaging client (104) renders the first virtual object as attached to the second virtual object in response to the selection.

[0185] In operation 2616, the messaging client (104) provides for display the first virtual object rendered as being attached to the second virtual object in the first scene.

[0186] In one embodiment, an open trigger finger gesture includes a specific gesture comprising a thumb and an index finger indicating a gun or shooting hand signal, and a closed trigger finger gesture includes the thumb and index finger pointing in the same direction.

[0187] In one embodiment, detecting a first gesture corresponding to an open trigger finger gesture from a set of frames comprises: determining a set of distances between a first set of joints from the thumb and a second set of joints from the index finger, and determining that at least one distance from the set of distances is greater than a distance threshold to indicate an open trigger finger gesture.

[0188] In one embodiment, detecting a second gesture corresponding to a closed trigger finger gesture from a second set of frames comprises: determining a set of distances between a first set of joints from the thumb and a second set of joints from the index finger, and determining that at least one distance from the set of distances is smaller than a distance threshold to indicate a closed trigger finger gesture.

[0189] In one embodiment, a first scene includes a first representation of a real-world scene, a first virtual object, and a second virtual object, wherein the first scene includes real-world video frame data and virtual object data, and the virtual object data includes information used to render the first virtual object or the second virtual object.

[0190] In one embodiment, the messaging client (104) detects a second position and a second position of a finger expression from a closed trigger finger gesture, detects a change in the second position and a second position of the finger expression, moves a first virtual object to a different position and a different position based on the change, and moves a second virtual object to a second different position and a second different position based on the change, wherein moving the first virtual object and moving the second virtual object occur in cooperation with the two virtual objects.

[0191] In one embodiment, the messaging client (104) renders a set of moves based at least partially on moving a first virtual object and moving a second virtual object within a first scene, and provides the set of rendered moves for display.

[0192] In one embodiment, a messaging client (104) displays a selection of a specific option or specific item based on a third virtual object at least partially based on a second different location and a second different position of a second virtual object, wherein the third virtual object is overlaid by the second virtual object and is larger in size than the second virtual object.

[0193] In one embodiment, the messaging client (104) detects a third gesture corresponding to a second open trigger finger gesture from a third set of frames, renders a first virtual object as not attached to a second virtual object in response to the detection, and provides the first virtual object rendered as not attached to a second virtual object for display.

[0194] In one embodiment, the messaging client (104) displays a first selection of a first specific option or a first specific item based on a third virtual object in response to selecting a second virtual object and before rendering a set of moves, and the first specific option or the first specific item is different from the specific option or specific item.

[0195] Machine Architecture

[0196] FIG. 27 is a schematic representation of a machine (2700) on which instructions (2710) (e.g., software, program, application, applet, app, or other executable code) can be executed to cause the machine (2700) to perform any one or more of the methodologies discussed herein. For example, instructions (2710) can cause the machine (2700) to perform any one or more of the methods described herein. Instructions (2710) convert a general unprogrammed machine (2700) into a specific machine (2700) programmed to perform the described and illustrated functions in the described manner. The machine (2700) may operate as a standalone device or may be coupled to other machines (e.g., networked). In a networked deployment, the machine (2700) may operate as a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine (2700) may include, but is not limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a personal digital assistant (PDA), an entertainment media system, a cellular phone, a smartphone, a mobile device, a wearable device (e.g., a smartwatch), a smart home device (e.g., a smart appliance), other smart devices, a web appliance, a network router, a network switch, a network bridge, or any machine capable of executing commands (2710) that specify actions to be taken by the machine (2700) sequentially or otherwise.Furthermore, although only a single machine (2700) is exemplified, the term “machine” should also be considered to include a collection of machines that execute instructions (2710) individually or jointly to perform any one or more of the methodologies discussed herein. For example, the machine (2700) may include a client device (102) or any one of a plurality of server devices forming part of a messaging server system (108). In some examples, the machine (2700) may also include both client and server systems, where specific operations of a specific method or algorithm are performed on the server side and specific operations of a specific method or algorithm are performed on the client side.

[0197] The machine (2700) may include processors (2704), memory (2706), and I / O (input / output) components (2702) that can be configured to communicate with each other via a bus (2740). In one example, the processors (2704) (e.g., a CPU (Central Processing Unit), a RISC (Reduced Instruction Set Computing) processor, a CISC (Complex Instruction Set Computing) processor, a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an RFIC (Radio-Frequency Integrated Circuit), other processors, or any suitable combination thereof) may include, for example, a processor (2708) and a processor (2712) that execute instructions (2710). The term “processor” is intended to include multi-core processors that may include two or more independent processors (sometimes referred to as “cores”) capable of executing instructions simultaneously. FIG. 27 illustrates multiple processors (2704), but the machine (2700) may include a single processor having a single core, a single processor having multiple cores (e.g., a multi-core processor), multiple processors having a single core, multiple processors having multiple cores, or any combination thereof.

[0198] Memory (2706) includes main memory (2714), static memory (2716), and a storage unit (2718), both of which are accessible to processors (2704) via a bus (2740). The main memory (2706), static memory (2716), and storage unit (2718) store instructions (2710) that implement any one or more of the methodologies or functions described herein. The instructions (2710) may also exist, wholly or partially, during the execution by the machine (2700), in the main memory (2714), in the static memory (2716), in the machine-readable medium (2720) in the storage unit (2718), in at least one of the processors (2704) (e.g., in the processor's cache memory), or any suitable combination thereof.

[0199] The I / O components (2702) may include a wide variety of components for receiving inputs, providing outputs, generating outputs, transmitting information, exchanging information, capturing measurements, etc. The specific I / O components (2702) included in a particular machine will depend on the type of machine. For example, portable machines such as mobile phones may include touch input devices or other such input mechanisms, while headless server machines are unlikely to include such touch input devices. It will be acknowledged that the I / O components (2702) may include many other components not shown in FIG. 27. In various examples, the I / O components (2702) may include user output components (2726) and user input components (2728). User output components (2726) may include visual components (e.g., displays such as a plasma display panel (PDP), light-emitting diode (LED) display, liquid crystal display (LCD), projector, or cathode ray tube (CRT)), acoustic components (e.g., speakers), haptic components (e.g., vibration motors, resistance mechanisms), other signal generators, etc. User input components (2728) may include alphanumeric input components (e.g., keyboard, touch screen configured to receive alphanumeric input, photo-optical keyboard, or other alphanumeric input components), point-based input components (e.g., mouse, touchpad, trackball, joystick, motion sensor, or other pointing mechanism), tactile input components (e.g., physical buttons, touch screens providing the location and force of touch or touch gestures, or other tactile input components), audio input components (e.g., microphones), etc.

[0200] In additional examples, I / O components (2702) may include, among various other components, biometric components (2730), motion components (2732), environment components (2734), or position components (2736). For example, biometric components (2730) include components for detecting expressions (e.g., hand expressions, facial expressions, voice expressions, body gestures, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweat, or brain waves), identifying a person (e.g., voice identification, retinal identification, face identification, fingerprint identification, or EEG-based identification), and doing similar things. Motion components (2732) include acceleration sensor components (e.g., accelerometers), gravity sensor components, and rotation sensor components (e.g., gyroscopes).

[0201] Environmental components (2734) may include, for example, one or more cameras (having still image / photo and video capabilities), light sensor components (e.g., photometers), temperature sensor components (e.g., one or more thermometers for detecting ambient temperature), humidity sensor components, pressure sensor components (e.g., barometers), acoustic sensor components (e.g., one or more microphones for detecting background noise), proximity sensor components (e.g., infrared sensors for detecting nearby objects), gas sensors (e.g., gas detection sensors for detecting concentrations of hazardous gases for safety or measuring contaminants in the atmosphere), or other components capable of providing indications, measurements, or signals corresponding to the surrounding physical environment.

[0202] With respect to the cameras, the client device (102) may have a camera system including, for example, front cameras located on the front of the client device (102) and rear cameras located on the rear of the client device (102). The front cameras may be used, for example, to capture still images and videos (e.g., "selfies") of the user of the client device (102), which may then be augmented with the augmentation data (e.g., filters) described above. The rear cameras may be used, for example, to capture still images and videos in a more traditional camera mode, and these images are similarly augmented with augmentation data. In addition to the front and rear cameras, the client device (102) may also include a 360° camera for capturing 360° photos and videos.

[0203] Additionally, the camera system of the client device (102) may include dual rear cameras (e.g., a depth-sensing camera as well as a main camera) on the front and rear sides of the client device (102), or even triple, quadruple, or quintuple rear camera configurations. These multiple camera systems may include, for example, a wide camera, an ultra-wide camera, a telephoto camera, a macro camera, and a depth sensor.

[0204] Position components (2736) include position sensor components (e.g., GPS receiver components), altitude sensor components (e.g., altimeters or barometers that detect air pressure from which altitude can be derived), orientation sensor components (e.g., magnetometers), and similar ones.

[0205] Communication can be implemented using a wide variety of technologies. The I / O components (2702) further include communication components (2738) operable to connect the machine (2700) to a network (2722) or devices (2724) through their respective combinations or connections. For example, the communication components (2738) may include a network interface component, or other suitable device for interfacing with the network (2722). In additional examples, the communication components (2738) may include wired communication components, wireless communication components, cellular communication components, near-field communication (NFC) components, Bluetooth ® Components (e.g., Bluetooth) ® Low Energy), Wi-Fi ® It may include other communication components that provide communication through components and other aspects. The devices (2724) may be other machines or any various peripheral devices (e.g., peripheral devices connected via USB).

[0206] Furthermore, the communication components (2738) may include components capable of detecting identifiers or operable to detect identifiers. For example, the communication components (2738) may include Radio Frequency Identification (RFID) tag reader components, NFC smart tag detection components, optical reader components (e.g., optical sensors for detecting one-dimensional barcodes such as Universal Product Code (UPC) barcodes, multi-dimensional barcodes such as Quick Response (QR) codes, Aztec codes, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D barcodes, and other optical codes), or acoustic detection components (e.g., microphones for identifying tagged audio signals). In addition, various information, such as location via Internet Protocol (IP) geolocation, location via Wi-Fi® signal triangulation, and location via detection of NFC beacon signals that can indicate a specific location, can be derived through the communication components (2738).

[0207] Various memories (e.g., main memory (2714), static memory (2716), and memory of the processors (2704)) and storage unit (2718) may store one or more sets of instructions and data structures (e.g., software) that implement any one or more of the methodologies or functions described herein. These instructions (e.g., instructions (2710)) enable various operations to implement the disclosed examples when executed by the processors (2704).

[0208] Commands (2710) may be transmitted or received over a network (2722) using a transmission medium through a network interface device (e.g., a network interface component included in the communication components (2738)) and using any one of several well-known transmission protocols (e.g., HTTP (hypertext transfer protocol)). Similarly, commands (2710) may be transmitted or received using a transmission medium through a connection to the devices (2724) (e.g., a peer-to-peer connection).

[0209] Software Architecture

[0210] FIG. 28 is a block diagram (2800) illustrating a software architecture (2804) that may be installed in any one or more of the devices described herein. The software architecture (2804) is supported by hardware such as a machine (2802) that includes processors (2820), memory (2826), and I / O components (2838). In this example, the software architecture (2804) may be conceptualized as a stack of layers, each layer providing a specific function. The software architecture (2804) includes layers such as an operating system (2812), libraries (2810), frameworks (2808), and applications (2806). Operationally, applications (2806) invoke API calls (2850) through the software stack and receive messages (2852) in response to the API calls (2850).

[0211] The operating system (2812) manages hardware resources and provides common services. The operating system (2812) includes, for example, a kernel (2814), services (2816), and drivers (2822). The kernel (2814) acts as an abstraction layer between the hardware and other software layers. For example, the kernel (2814) provides memory management, processor management (e.g., scheduling), component management, networking, and security settings, among other functions. Services (2816) may provide other common services for other software layers. Drivers (2822) are responsible for controlling or interfacing with the underlying hardware. For example, drivers (2822) may include display drivers, camera drivers, BLUETOOTH® or BLUETOOTH® Low Energy drivers, flash memory drivers, serial communication drivers (e.g., USB drivers), WI-FI® drivers, audio drivers, power management drivers, etc.

[0212] Libraries (2810) provide common low-level infrastructure used by applications (2806). Libraries (2810) may include system libraries (2818) (e.g., C standard library) that provide functions such as memory allocation functions, string manipulation functions, mathematical functions, etc. Additionally, the libraries (2810) may include API libraries (2824) such as media libraries (e.g., libraries that support the presentation and manipulation of various media formats such as MPEG4 (Moving Picture Experts Group-4), Advanced Video Coding (H.264 or AVC), MP3 (Moving Picture Experts Group Layer-3), AAC (Advanced Audio Coding), AMR (Adaptive Multi-Rate) audio codecs, Joint Photographic Experts Group (JPEG or JPG), or PNG (Portable Network Graphics), graphics libraries (e.g., OpenGL frameworks used to render graphic content on a display in two dimensions (2D) and three dimensions (3D)), database libraries (e.g., SQLite providing various relational database functions), web libraries (e.g., WebKit providing web browsing functionality). The libraries (2810) may also include a wide variety of other libraries (2828) that provide many different APIs to applications (2806).

[0213] Frameworks (2808) provide common high-level infrastructure used by applications (2806). For example, frameworks (2808) provide various graphical user interface (GUI) functions, high-level resource management, and high-level location services. Frameworks (2808) may provide a wide range of other APIs that can be used by applications (2806), some of which may be specific to a particular operating system or platform.

[0214] In one example, applications (2806) may include a wide range of other applications such as a home application (2836), a contact application (2830), a browser application (2832), a book reader application (2834), a location application (2842), a media application (2844), a messaging application (2846), a game application (2848), and a third-party application (2840). Applications (2806) are programs that execute functions defined in programs. Various programming languages ​​may be used to create one or more of the applications (2806), which are structured in various ways, such as an object-oriented programming language (e.g., Objective-C, Java, or C++) or a procedural programming language (e.g., C or assembly language). In a specific example, a third-party application (2840) (e.g., an application developed by an entity other than a vendor of a specific platform using an ANDROID™ or IOS™ software development kit (SDK)) may be mobile software running on a mobile operating system such as IOS™, ANDROID™, WINDOWS® Phone, or other mobile operating systems. In this example, the third-party application (2840) may invoke API calls (2850) provided by the operating system (2812) to facilitate the functions described herein.

[0215] Glossary

[0216] "Carrier signal" refers to any intangible medium capable of storing, encoding, or carrying instructions for execution by a machine, and includes digital or analog communication signals or other intangible media to facilitate the communication of such instructions. Instructions may be transmitted or received over a network using a transmission medium through a network interface device.

[0217] "Client device" refers to any machine that interfaces with a communication network to obtain resources from one or more server systems or other client devices. A client device may be, but is not limited to, a mobile phone, a desktop computer, a laptop, PDAs (portable digital assistants), smartphones, tablets, ultrabooks, netbooks, laptops, multi-processor systems, microprocessor-based or programmable consumer electronics, game consoles, set-top boxes, or any other communication device that a user can use to access the network.

[0218] "Communication network" refers to one or more parts of a network that may be an ad-hoc network, intranet, extranet, VPN (virtual private network), LAN (local area network), wireless LAN (WLAN), WAN (wide area network), wireless WAN (WWAN), MAN (metropolitan area network), the Internet, part of the Internet, part of the PSTN (Public Switched Telephone Network), POTS (plain old telephone service) network, cellular telephone network, wireless network, Wi-Fi® network, other types of networks, or a combination of two or more of these networks. For example, a network or part of a network may include a wireless or cellular network, and a coupling may be a CDMA (Code Division Multiple Access) connection, a GSM (Global System for Mobile communications) connection, or other types of cellular or wireless coupling.In this example, the combination can implement any of the various types of data transmission technologies, such as 1xRTT (Single Carrier Radio Transmission Technology), EVDO (Evolution-Data Optimized) technology, GPRS (General Packet Radio Service) technology, EDGE (Enhanced Data rates for GSM Evolution) technology, 3GPP (third Generation Partnership Project) including 3G, 4th generation wireless (4G) networks, UMTS (Universal Mobile Telecommunications System), HSPA (High Speed ​​Packet Access), WiMAX (Worldwide Interoperability for Microwave Access), LTE (Long Term Evolution) standards, other things defined by various standard-setting organizations, other long-range protocols, or other data transmission technologies.

[0219] "Component" refers to a device, physical entity, or logic having boundaries defined by function or subroutine calls, branch points, APIs, or other techniques that provide the division or modularization of specific processing or control functions. Components may be combined with other components through their interfaces to complete machine processes. A component may be a packaged functional hardware unit designed to be used with other components and, typically, a part of a program that performs a specific function among the related functions. Components may constitute software components (e.g., code embodied on a machine-readable medium) or hardware components. "Hardware component" is a tangible unit capable of performing specific operations and may be configured or arranged in a specific physical manner. In various examples, one or more computer systems (e.g., standalone computer systems, client computer systems, or server computer systems) or one or more hardware components of a computer system (e.g., a processor or a group of processors) may be configured by software (e.g., an application or a part of an application) as hardware components that operate to perform specific operations as described herein. Hardware components may also be implemented mechanically, electronically, or any suitable combination thereof. For example, a hardware component may include dedicated circuits or logic that are permanently configured to perform specific operations. A hardware component may be a special-purpose processor, such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). A hardware component may also include programmable logic or circuits that are temporarily configured by software to perform specific operations.For example, a hardware component may include software executed by a general-purpose processor or another programmable processor. Once configured by such software, the hardware components become specific machines (or specific components of a machine) uniquely customized to perform the configured functions and are no longer general-purpose processors. It will be recognized that the decision to implement a hardware component mechanically, in a dedicated, permanently configured circuit, or in a temporarily configured circuit (e.g., configured by software) may be driven by cost and time considerations. Accordingly, the phrase "hardware component" (or "hardware-implemented component") should be understood to encompass tangible entities as long as they are physically configured, permanently configured (e.g., hardwired), or temporarily configured (e.g., programmed) to operate in a particular way or to perform the specific actions described herein. When considering examples where hardware components are temporarily configured (e.g., programmed), each hardware component does not need to be configured or instantiated at any single time instance. For example, in the case where a general-purpose processor is configured by software such that a hardware component becomes a special-purpose processor, the general-purpose processor may be configured as different special-purpose processors at different times (e.g., including different hardware components). Thus, the software configures a specific processor or processors to configure a specific hardware component at one time instance and a different hardware component at a different time instance. A hardware component may provide information to other hardware components and receive information from them. Therefore, the described hardware components may be considered as being coupled in a communicable manner.In cases where multiple hardware components exist simultaneously, communication may be achieved through signal transmission between two or more of the hardware components or between two or more (e.g., via appropriate circuits and buses). In examples where multiple hardware components are configured or instantiated at different times, communication between such hardware components may be achieved, for example, through the storage and retrieval of information within memory structures accessible to multiple hardware components. For example, one hardware component may perform an operation and store the output of that operation in a memory device to which it is communicably coupled. Subsequently, additional hardware components may access the memory device to retrieve and process the stored output. Hardware components may also initiate communication with input or output devices and operate on resources (e.g., collections of information). Various operations of the exemplary methods described herein may be performed at least partially by one or more processors configured temporarily (e.g., by software) or permanently to perform the relevant operations. Whether configured temporarily or permanently, such processors may constitute processor-implemented components that operate to perform one or more operations or functions described herein. As used herein, "processor-implemented component" refers to a hardware component implemented using one or more processors. Similarly, the methods described herein may be implemented at least partially by a processor, and specific processors or processors are examples of hardware.For example, at least some of the operations of the method may be performed by one or more processors or components implemented by processors. Furthermore, one or more processors may also operate to support the performance of the relevant operations in a "cloud computing" environment or as "software as a service" (SaaS). For example, at least some of the operations may be performed by a group of computers (as examples of machines containing processors), and these operations are accessible via a network (e.g., the Internet) and through one or more appropriate interfaces (e.g., APIs). The performance of specific operations may not only reside within a single machine but may also be distributed among processors deployed across multiple machines. In some examples, the processors or components implemented by processors may be located in a single geographic location (e.g., a home environment, an office environment, or within a server farm). In other examples, the processors or components implemented by processors may be distributed across multiple geographic locations.

[0220] "Computer-readable storage medium" refers to both machine storage media and transmission media. Accordingly, the terms include both storage devices / medias and carriers / modulated data signals. The terms "machine-readable medium," "computer-readable medium," and "device-readable medium" mean the same thing and may be used interchangeably in this disclosure.

[0221] An "ephemeral message" refers to a message accessible for a time-limited duration. Ephemeral messages can be text, images, videos, etc. The access time for an ephemeral message can be set by the message sender. Alternatively, the access time can be a default setting or a setting specified by the recipient. Regardless of the setting method, the message is transitory.

[0222] "Machine storage medium" refers to a single or multiple storage devices and media (e.g., centralized or distributed databases, and associated caches and servers) that store executable instructions, routines, and data. Accordingly, the term should be considered to include, but not be limited to, solid-state memories including memory inside or outside processors, and optical and magnetic media. Specific examples of machine storage medium, computer storage medium, and device storage medium include, by example, non-volatile memory including semiconductor memory devices, e.g., EPROM (erasable programmable read-only memory), EEPROM (electrically erasable programmable read-only memory), FPGAs, and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The terms "machine storage medium," "device storage medium," and "computer storage medium" mean the same thing and may be used interchangeably in this disclosure. The terms "machine storage medium," "computer storage medium," and "device storage medium" specifically exclude carrier waves, modulated data signals, and other such media, at least some of which are included under the term "signal medium."

[0223] "Non-transient computer-readable storage medium" refers to a type of medium capable of storing, encoding, or carrying instructions for execution by a machine.

[0224] "Signal medium" refers to any intangible medium capable of storing, encoding, or carrying instructions for execution by a machine, and includes digital or analog communication signals or other intangible media for facilitating the communication of software or data. The term "signal medium" should be considered to include any form of modulated data signal, carrier wave, etc. The term "modulated data signal" means a signal in which one or more of its characteristics are set or altered, such as in the matter of encoding information within the signal. The terms "transmission medium" and "signal medium" mean the same thing and may be used interchangeably in this disclosure.

Claims

Claim 1 A method comprising: using one or more hardware processors to detect a first gesture corresponding to an open trigger finger gesture from a set of frames; using one or more hardware processors to detect a position and location of a finger representation from the open trigger finger gesture; using one or more hardware processors to generate a first virtual object at least partially based on the position and location of the finger representation, wherein the first virtual object is extended from the finger representation; using one or more hardware processors to detect a first collision event corresponding to a first collider of the first virtual object intersecting a second collider of the second virtual object; using one or more hardware processors to detect a second gesture corresponding to a closed trigger finger gesture from a second set of frames; selecting the second virtual object in response to the first collision event and the detected second gesture; using one or more hardware processors to render the first virtual object as attached to the second virtual object in response to the selection. A method comprising the step of using one or more hardware processors to provide for display the rendered first virtual object attached to the second virtual object within the first scene. Claim 2 A method according to claim 1, wherein the open trigger finger gesture comprises a specific gesture including a thumb and an index finger indicating a firearm or a shooting hand sign, and the closed trigger finger gesture comprises the thumb and the index finger pointing in the same direction. Claim 3 In paragraph 2, the step of detecting a first gesture corresponding to the open trigger finger gesture from the set of frames comprises: determining a set of distances between a first set of joints from the thumb and a second set of joints from the index finger; and determining that at least one distance from the set of distances is greater than a distance threshold to indicate the open trigger finger gesture. Claim 4 In paragraph 2, the step of detecting a second gesture corresponding to the closed trigger finger gesture from a second set of frames comprises: determining a set of distances between a first set of joints from the thumb and a second set of joints from the index finger; and determining that at least one distance from the set of distances is smaller than a distance threshold to indicate the closed trigger finger gesture. Claim 5 A method according to claim 1, wherein the first scene includes real-world video frame data and virtual object data, and the virtual object data includes information used to render the first virtual object or the second virtual object. Claim 6 A method according to claim 1, further comprising: a step of detecting a second position and a second position of the expression of the finger from the closed trigger finger gesture; a step of detecting a change in the second position and a second position of the expression of the finger; a step of moving the first virtual object to a different position and a different position based on the change; and a step of moving the second virtual object to a second different position and a second different position based on the change, wherein moving the first virtual object and moving the second virtual object occur in cooperation with the two virtual objects. Claim 7 A method according to claim 6, further comprising: a step of rendering a set of moves based at least partially on moving the first virtual object and moving the second virtual object within the first scene; and a step of providing the set of rendered moves for display. Claim 8 In claim 7, the method further comprises the step of displaying a selection of a specific option or specific item based on a third virtual object, at least partially based on a second different location and a second different position of the second virtual object, wherein the third virtual object is overlaid by the second virtual object and is larger in size than the second virtual object. Claim 9 A method according to claim 8, further comprising: detecting a third gesture corresponding to a second open trigger finger gesture from a third set of frames; rendering the first virtual object as not attached to the second virtual object in response to detecting the third gesture; and providing the rendered first virtual object as not attached to the second virtual object for display. Claim 10 In claim 8, the method further comprises the step of displaying a first selection of a first specific option or a first specific item based on the third virtual object in response to selecting the second virtual object and before rendering the set of moves, wherein the first specific option or the first specific item is different from the specific option or the specific item. Claim 11 A system comprising: a processor; and a memory comprising instructions that, when executed by the processor, cause the processor to perform operations, wherein the operations include: detecting a first gesture corresponding to an open trigger finger gesture from a set of frames; detecting a position and location of a finger expression from the open trigger finger gesture; creating a first virtual object based at least partially on the position and location of the finger expression, wherein the first virtual object is extended from the finger expression; detecting a first collision event corresponding to a first collider of the first virtual object intersecting a second collider of a second virtual object; detecting a second gesture corresponding to a closed trigger finger gesture from a second set of frames; selecting the second virtual object in response to the first collision event and the detected second gesture; rendering the first virtual object as attached to the second virtual object in response to the selection; and providing the rendered first virtual object as attached to the second virtual object within a first scene for display. Claim 12 A system according to claim 11, wherein the open trigger finger gesture comprises a specific gesture including a thumb and an index finger indicating a firearm or a shooting hand sign, and the closed trigger finger gesture comprises the thumb and the index finger pointing in the same direction. Claim 13 A system according to claim 12, wherein detecting a first gesture corresponding to the open trigger finger gesture from a set of frames comprises: determining a set of distances between a first set of joints from the thumb and a second set of joints from the index finger; and determining that at least one distance from the set of distances is greater than a distance threshold to indicate the open trigger finger gesture. Claim 14 A system according to claim 12, wherein detecting a second gesture corresponding to the closed trigger finger gesture from a second set of frames comprises: determining a set of distances between a first set of joints from the thumb and a second set of joints from the index finger; and determining that at least one distance from the set of distances is smaller than a distance threshold to indicate the closed trigger finger gesture. Claim 15 A system according to claim 11, wherein the first scene includes real-world video frame data and virtual object data, and the virtual object data includes information used to render the first virtual object or the second virtual object. Claim 16 In claim 11, the operations further comprise: detecting a second position and a second position of the expression of the finger from the closed trigger finger gesture; detecting a change in the second position and a second position of the expression of the finger; moving the first virtual object to a different position and a different position based on the change; and moving the second virtual object to a second different position and a second different position based on the change, wherein moving the first virtual object and moving the second virtual object occur in cooperation with the two virtual objects. Claim 17 A system according to claim 16, wherein the operations further comprise: rendering a set of moves based at least partially on moving the first virtual object within the first scene and moving the second virtual object; and providing the set of rendered moves for display. Claim 18 In paragraph 17, the above operations further include: displaying a selection of a specific option or specific item based on a third virtual object, at least partially based on a second different location and a second different position of the second virtual object, wherein the third virtual object is overlaid by the second virtual object and is larger in size than the second virtual object, a system. Claim 19 A system according to claim 18, wherein the operations further comprise: detecting a third gesture corresponding to a second open trigger finger gesture from a third set of frames; rendering the first virtual object as not attached to the second virtual object in response to detecting the third gesture; and providing the rendered first virtual object as not attached to the second virtual object for display. Claim 20 A non-transient computer-readable medium comprising instructions that, when executed by a computing device, cause the computing device to perform operations, wherein the operations include: detecting a first gesture corresponding to an open trigger finger gesture from a set of frames; detecting a position and location of a finger representation from the open trigger finger gesture; creating a first virtual object based at least partially on the position and location of the finger representation, wherein the first virtual object is extended from the finger representation; detecting a first collision event corresponding to a first collider of the first virtual object intersecting a second collider of the second virtual object; detecting a second gesture corresponding to a closed trigger finger gesture from a second set of frames; selecting the second virtual object in response to the first collision event and the detected second gesture; and rendering the first virtual object as attached to the second virtual object in response to the selection. A non-transient computer-readable medium comprising providing the rendered first virtual object for display as being attached to the second virtual object within the first scene.

Citation Information

Patent Citations

  • Bare hand operation method and system in augmented reality

    CN113608619A

  • The analysis apparatus and method of user intention using video information in three dimensional space

    KR1020160133676A

  • Edge-Identifying Gesture-Driven User Interface Element Gating for Artificial Reality Systems

    KR1020220018562A

  • Micro hand gestures for controlling virtual and graphical elements

    US20220206588A1