Depth estimation using biological data

By generating location data by recognizing the user's skeletal or facial position or posture, and combining it with biometric data, the accuracy of distance measurement between the user and the device in augmented reality systems is solved, thus addressing the accuracy problem of distance measurement between the user and the device in existing technologies and achieving a better user experience.

CN118363456BActive Publication Date: 2025-12-30SNAP INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410393310.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-08-03
Filing Date
2021-03-30
Publication Date
2025-12-30
Estimated Expiration
2041-03-30

AI Technical Summary

Technical Problem

Existing augmented reality systems struggle to accurately determine the distance between users and devices, impacting the realistic interaction between virtual objects and the real world.

Method used

Location data is generated by recognizing the user's skeletal or facial position or posture, and combined with biometric data, the processor uses a processor to determine the distance between the user and the device.

Benefits of technology

It enables accurate estimation of the distance between the user and the device in augmented reality systems, thereby improving the interaction between virtual objects and the real world.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118363456B_ABST
    Figure CN118363456B_ABST
Patent Text Reader

Abstract

Implementations relate to a method comprising: generating, by a first device, positioning data based on an analysis of a data stream by identifying a position or pose of a skeleton or face of a user of a second device in a frame of a camera, wherein the data stream comprises an image of the user of the second device; receiving, by the first device, biometric data of the user from the second device, the biometric data based on output from a sensor or camera included in the second device; and determining, by the first device, a distance of the user from the first device using the positioning data and the biometric data of the user. Other implementations are described herein.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application is a divisional application of Chinese patent application No. 202180032389.6, filed on November 1, 2022, entitled "Depth Estimation Using Biological Data". The international filing date of the parent application is March 30, 2021, and the international application number is PCT / US2021 / 070341, with the earliest priority date being March 30, 2020.

[0002] Priority Statement

[0003] This application claims priority to U.S. Provisional Application Serial No. 63 / 002,128, filed March 30, 2020, and U.S. Patent Application Serial No. 16 / 983,751, filed August 3, 2020, both of which are incorporated herein by reference in their entirety. Technical Field

[0004] This invention relates to depth estimation using biological data. Background Technology

[0005] Augmented Reality (AR) is a modification of a virtual environment. For example, in Virtual Reality (VR), users are fully immersed in a virtual world, while in AR, users are immersed in a world where virtual objects are combined with or overlaid with the real world. AR systems aim to generate and present virtual objects that realistically interact with the real-world environment and with each other. Examples of AR applications can include single-player or multiplayer video games, instant messaging systems, and more. Summary of the Invention

[0006] Embodiments of the present invention provide a method comprising: generating location data by a first device based on analysis of a data stream by recognizing the position or pose of the bones or face of a user of a second device within a camera device frame, wherein the data stream includes an image of the user of the second device; receiving biometric data of the user from the second device by the first device, the biometric data being based on output from a sensor or camera device included in the second device; and determining the distance between the user and the first device by the first device using the location data and the biometric data of the user.

[0007] Embodiments of the present invention provide a device comprising: a processor; and a memory storing instructions thereon, the instructions, when executed by the processor, causing the device to perform operations including: capturing a data stream of images of a user including a second device using a camera device included in the device; generating location data based on analysis of the data stream by identifying the position or pose of the user's bones or face within the frame of the camera device; receiving biometric data of the user from the second device, the biometric data being based on output from a sensor or camera device included in the second device; and determining the distance between the user and the device using the location data and the user's biometric data.

[0008] Embodiments of the present invention provide a non-transitory computer-readable storage medium having instructions stored thereon, the instructions, when executed by a processor of a first device, causing the processor to perform operations including: using a camera device to capture a data stream of images of a user including a second device; generating location data based on analysis of the data stream by identifying the position or pose of the user's bones or face within the frame of the camera device; generating biometric data of the user from the second device, the biometric data being based on output from a sensor or camera device included in the second device; and using the location data and the user's biometric data to determine the distance between the user and the first device. Attached Figure Description

[0009] In the accompanying drawings (not necessarily drawn to scale), similar reference numerals may describe similar parts in different views. Similar reference numerals with different letter suffixes may indicate different instances of similar parts. For ease of identification of any particular element or action being discussed, one or more of the most significant digits in the reference numerals refer to the drawing number in which the element was first introduced. Some embodiments are shown by way of example, not limitation, in the figures, in which:

[0010] Figure 1 It is a graphical representation of a networked environment in which the present disclosure can be deployed, according to some example implementations.

[0011] Figure 2 It is a graphical representation of a message-transmitting client application based on some example implementations.

[0012] Figure 3 It is a graphical representation of a data structure maintained in a database, based on some example implementations.

[0013] Figure 4 It is a graphical representation of a message based on some example implementations.

[0014] Figure 5This is a flowchart of a process 500 for generating depth estimates based on biological data, according to one embodiment.

[0015] Figure 6 A process 600 for generating depth estimates based on biological data according to one embodiment is shown.

[0016] Figure 7 An example 700 is shown whereby a first user (user B) uses a first client device 102 to capture images or videos of a second user (user A) according to one embodiment.

[0017] Figure 8 An example 800 is shown using an image captured by a camera device included in a first client device 102, according to one embodiment.

[0018] Figure 9 It is a graphical representation of a machine in the form of a computer system according to some example embodiments, within which a set of instructions can be executed to cause the machine to perform any or more of the methods discussed herein.

[0019] Figure 10 This is a block diagram illustrating a software architecture in which the present disclosure can be implemented according to an example embodiment. Detailed Implementation

[0020] The following description includes systems, methods, techniques, instruction sequences, and computer program products embodying illustrative embodiments of the present disclosure. In this description, numerous specific details are set forth for purposes of explanation in order to provide an understanding of various embodiments of the subject matter of the invention. However, it will be apparent to those skilled in the art that embodiments of the subject matter of the invention can be practiced without these specific details. Generally, well-known examples of instructions, protocols, structures, and techniques are not necessarily shown in detail.

[0021] Furthermore, embodiments of this disclosure improve the functionality of augmented reality (AR) creation software and systems by using biometric data about the user to generate a depth estimate of the user. A first challenge in generating AR scenes is determining how far the tracked user is from the client device tracking the user.

[0022] More specifically, a first user may be viewing an AR scene that includes a second user, displayed on a first client device. In order to generate an AR scene including the second user on the first device, the AR system needs to determine the distance between the second user and the first client device in order to generate and render virtual objects in the AR scene that can realistically interact with the second user (e.g., real-world elements).

[0023] In one implementation, for example, when the first client device performs skeletal tracking that enables an augmented reality experience that enhances the body or clothing of a second user, the second client device sends known biometric data of the second user to generate a depth estimate (e.g., the distance of the second user to the first client device).

[0024] In one implementation, the messaging server system receives location data from a first client device and biometric data from a second client device. The messaging server system determines the distance between the user and the first client device based on the location data and / or biometric data. In another implementation, the second client device sends biometric data to the first client device, and the first client device determines the distance between the user and the first client device based on the location data and / or biometric data.

[0025] Figure 1 This is a block diagram illustrating an example messaging system 100 for exchanging data (e.g., messages and associated content) over a network. Messaging system 100 includes multiple instances of client devices 102, each instance hosting multiple applications including a messaging client application 104 and an AR session client controller 124. Each messaging client application 104 is communicatively coupled to other instances of the messaging client application 104 and a messaging server system 108 via a network 106 (e.g., the Internet). Each AR session client controller 124 is communicatively coupled to other instances of the AR session client controller 124 and an AR session server controller 126 in messaging server system 108 via the network 106.

[0026] The messaging client application 104 is able to communicate and exchange data with another messaging client application 104 and with the messaging server system 108 via the network 106. The data exchanged between messaging client applications 104 and between messaging client applications 104 and messaging server system 108 includes functions (e.g., commands that call functions) and payload data (e.g., text, audio, video, or other multimedia data).

[0027] Messaging server system 108 provides server-side functionality to a specific messaging client application 104 via network 106. While some functions of messaging system 100 are described herein as being performed by messaging client application 104 or by messaging server system 108, it is a design choice to locate certain functions within messaging client application 104 or messaging server system 108. For example, it may technically be preferred to initially deploy certain technologies and functions within messaging server system 108, but later migrate these technologies and functions to messaging client application 104, in which client device 102 has sufficient processing power.

[0028] The messaging server system 108 supports various services and operations provided to the messaging client application 104. Such operations include sending data to and receiving data from the messaging client application 104, and processing data generated by the messaging client application 104. For example, this data may include message content, client device information, geolocation information, media comments and overlays, message content persistence conditions, social network information, and live event information. Data exchange within the messaging system 100 is invoked and controlled via functions available through the user interface (UI) of the messaging client application 104.

[0029] AR session client controller 124 can communicate and exchange data with another AR session client controller 124 and AR session server controller 126 via network 106. The data exchanged between the multiple AR session client controllers 124 and between the AR session client controller 124 and AR session server controller 126 may include: location data, biometric data, depth estimates (e.g., the distance between a user and a client device tracking the user), requests to create a shared AR session, user identifiers in data streams (images or videos), session identifiers identifying the shared AR session, transformations between a first device and a second device (e.g., the multiple client devices 102 include a first device and a second device), a common coordinate system, functions (e.g., commands calling functions), and other payload data (e.g., text, audio, video, or other multimedia data).

[0030] Now, specifically to message server system 108, application programming interface (API) server 110 is coupled to application server 112 and provides a programming interface to application server 112. Application server 112 is communicatively coupled to database server 118, which facilitates access to database 120, which stores data associated with messages processed by application server 112.

[0031] Application Programming Interface (API) server 110 receives and sends message data (e.g., commands and message payloads) between client device 102 and application server 112. Specifically, API server 110 provides a set of interfaces (e.g., routines and protocols) that can be invoked or queried by messaging client application 104 to call functions of application server 112. Application Programming Interface (API) server 110 exposes various functions supported by application server 112, including account registration, login functionality, sending messages from one messaging client application 104 to another messaging client application 104 via application server 112, sending media files (e.g., images or videos) from messaging client application 104 to messaging server application 114, setting up collections of media data (e.g., stories) for possible access by another messaging client application 104, retrieving the friend list of the user of client device 102, retrieving such collections, retrieving messages and content, adding and deleting friends in the social graph, the location of friends in the social graph, and opening application events (e.g., involving messaging client application 104).

[0032] Application server 112 hosts multiple applications and subsystems, including messaging server application 114, image processing system 116, and social networking system 122. Messaging server application 114 implements numerous messaging techniques and functions, particularly involving the aggregation and other processing of content (e.g., text and multimedia content) received from multiple instances of messaging client application 104. As will be described in more detail, text and media content from multiple sources can be aggregated into collections of content (e.g., referred to as stories or libraries). Messaging server application 114 then makes these collections available to messaging client application 104. Given the hardware requirements for such processing, messaging server application 114 can also perform additional processor- and memory-intensive data processing on the server side.

[0033] Application server 112 also includes an image processing system 116 dedicated to performing various image processing operations, typically for images or videos received within the payload of a message at message delivery server application 114.

[0034] Social networking system 122 supports various social networking functions and services, and makes these functions and services available to messaging server application 114. To this end, social networking system 122 maintains and accesses entity graph 304 (such as...) within database 120. Figure 3(As shown). Examples of functions and services supported by the social networking system 122 include identifying other users of the messaging system 100 with whom a particular user has a relationship or who “follows” them, as well as identifying the interests of other entities and a particular user.

[0035] Application server 112 also includes an AR session server controller 126, which can communicate with AR session client controller 124 in client device 102 to establish individual or shared AR sessions. AR session server controller 126 can also be coupled to messaging server application 114 to establish electronic group communication sessions (e.g., group chat, instant messaging) for client devices within a shared AR session. The electronic group communication session can be associated with a session identifier provided by client device 102 to gain access to both the electronic group communication session and the shared AR session. In one embodiment, the client device first gains access to the electronic group communication session and then obtains a session identifier within the electronic group communication session that enables the client device to access the shared AR session. In some embodiments, client device 102 can access the shared AR session without the assistance of AR session server controller 126 in application server 112 or without communicating with AR session server controller 126 in application server 112.

[0036] Application server 112 is communicatively coupled to database server 118, which facilitates access to database 120, which stores data associated with messages processed by message server application 114.

[0037] Figure 2 This is a block diagram illustrating further details of a messaging system 100 according to an example embodiment. Specifically, the messaging system 100 is shown as including a messaging client application 104 and an application server 112, which in turn include several subsystems, namely a short-timer system 202, a collection management system 204, and an annotation system 206.

[0038] The short-lived timer system 202 is responsible for implementing short-lived access to content permitted by the messaging client application 104 and the messaging server application 114. To this end, the short-lived timer system 202 combines multiple timers that selectively display messages and associated content, and enable access to messages and associated content via the messaging client application 104, based on durations and display parameters associated with messages or sets of messages (e.g., stories). Further details regarding the operation of the short-lived timer system 202 are provided below.

[0039] The collection management system 204 is responsible for managing collections of media (e.g., collections of text, image, video, and audio data). In some examples, collections of content (e.g., messages, including images, videos, text, and audio) can be organized into “event libraries” or “event stories.” Such collections can be made available for a specified time period (e.g., the duration of the event the content relates to). For example, content related to a concert can be made available as a “story” during the duration of the concert. The collection management system 204 can also be responsible for publishing icons that notify the user interface of the messaging client application 104 of the existence of a specific collection.

[0040] The collection management system 204 also includes a curation interface 208 that allows collection managers to manage and curate specific collections of content. For example, the curation interface 208 enables event organizers to curate collections of content related to a specific event (e.g., removing inappropriate content or redundant messages). Additionally, the collection management system 204 employs machine vision (or image recognition technology) and content rules to automatically curate content collections. In some implementations, users may be paid compensation to include user-generated content in the collection. In such cases, the curation interface 208 operates to automatically pay such users for using their content.

[0041] Annotation system 206 provides various functions that enable users to annotate or otherwise modify or edit media content associated with messages. For example, annotation system 206 provides functions related to the generation and publication of media overlays for messages processed by messaging system 100. Annotation system 206 can operablely supply media overlays or supplements (e.g., image filters) to messaging client application 104 based on the geographic location of client device 102. In another example, annotation system 206 can operablely supply media overlays to messaging client application 104 based on other information (e.g., the social network information of the user of client device 102). Media overlays may include audio and visual content as well as visual effects. Examples of audio and visual content include pictures, text, logos, animations, and sound effects. Examples of visual effects include color overlays. Audio and visual content or visual effects may be applied to media content items (e.g., photos) at client device 102. For example, a media overlay may include text that can be overlaid on a photograph taken by client device 102. In another example, media overlays include location identifiers (e.g., Venice Beach) overlays, names of live events, or business names (e.g., Beach Cafe) overlays. In yet another example, annotation system 206 uses the geographic location of client device 102 to identify media overlays that include the name of a business located at the geographic location of client device 102. Media overlays may include additional tags associated with the business. Media overlays may be stored in database 120 and accessed via database server 118.

[0042] In one example implementation, annotation system 206 provides a user-based publishing platform that allows users to select geographic locations on a map and upload content associated with those locations. Users can also specify which media overlays should be provided to other users' environments. Annotation system 206 generates media overlays that include the uploaded content and associates the uploaded content with the selected geographic location.

[0043] In another example implementation, annotation system 206 provides a merchant-based publishing platform that enables merchants to select specific media coverage associated with geographic locations through a bidding process. For example, annotation system 206 associates the media coverage of the highest bidder with the corresponding geographic location for a predefined amount of time.

[0044] Figure 3 This is a schematic diagram illustrating a data structure 300 that can be stored in a database 120 of a message sending server system 108 according to some example embodiments. Although the contents of the database 120 are shown as including multiple tables, it should be understood that the data can be stored in other types of data structures (e.g., as an object-oriented database).

[0045] Database 120 includes message data stored in message table 314. Entity table 302 stores entity data, including entity diagram 304. Entities whose records are maintained in entity table 302 can include individuals, company entities, organizations, objects, locations, events, etc. Regardless of type, any entity whose data is stored in message server system 108 can be an identifiable entity. Each entity is assigned a unique identifier and an entity type identifier (not shown).

[0046] Entity Graph 304 also stores information about the relationships and associations between entities. As an example only, such relationships could be social or professional relationships based on interests or activities (e.g., working in a common company or organization).

[0047] Database 120 also stores annotation data in annotation table 312 in the form of filters. Filters stored in annotation table 312 are associated with and applied to videos (stored in video table 310) and / or images (stored in image table 308). In one example, a filter is an overlay displayed as an image or video during presentation to the recipient user. Filters can be of various types, including user-selected filters from a library of filters presented to the sending user by messaging client application 104 when the sending user is composing a message. Other types of filters include geolocation filters (also known as geographic filters), which can be presented to the sending user based on geographic location. For example, based on geographic location information determined by the GPS unit of client device 102, messaging client application 104 can present neighborhood-specific or location-specific geolocation filters within the user interface. Another type of file manager is a data file manager, which can be selectively presented to the sending user by messaging client application 104 based on other input or information collected by client device 102 during the message creation process. Examples of data filters include the current temperature at a specific location, the current speed of the user's movement, the battery life of the client device 102, or the current time.

[0048] Other annotation data that can be stored in image table 308 are augmented reality content items (e.g., corresponding to an applied lens or augmented reality experience). Augmented reality content items can be real-time special effects and sounds that can be added to images or videos.

[0049] As described above, augmented reality content items, overlays, image transformations, AR images, and similar terms refer to modifications that can be made to a video or image. This includes real-time modifications, i.e., modifying an image as it is captured using the device's sensors, and then using those modifications to display the image on the device's screen. This also includes modifications to stored content, such as video clips in a library that can be modified. For example, in a device with access to multiple augmented reality content items, a user can use a single video clip with multiple augmented reality content items to see how different augmented reality content items will modify the stored clip. For example, by selecting different augmented reality content items for the content, multiple augmented reality content items applying different pseudo-random motion models can be applied to the same content. Similarly, real-time video capture can be used in conjunction with the modifications shown to demonstrate how the video image currently captured by the device's sensors will modify the captured data. Such data can simply be displayed on the screen without being stored in memory, or the content captured by the device's sensors can be recorded and stored in memory with or without modification (or both). In some systems, a preview function can show how different augmented reality content items will be viewed simultaneously in different windows on the display. For example, this allows multiple windows with different pseudo-random animations to be viewed on the monitor simultaneously.

[0050] Using augmented reality content items or other such transformation systems to modify content data and various systems can therefore involve: detection of objects (e.g., faces, hands, bodies, cats, dogs, surfaces, objects, etc.); tracking these objects as they leave, enter, and move around the field of view in video frames; and modifying or transforming these objects while tracking them. In various implementations, different methods can be used to implement such transformations. For example, some implementations may involve generating a three-dimensional mesh model of one or more objects and using transformations of the model within the video and animated textures to implement the transformation. In other implementations, tracking points on objects can be used to place images or textures (which can be two-dimensional or three-dimensional) at the tracked locations. In yet another implementation, neural network analysis of video frames can be used to place images, models, or textures within content (e.g., images or video frames). Therefore, augmented reality content items refer both to images, models, and textures used to create transformations within content and to the additional modeling and analysis information required to implement such transformations through object detection, tracking, and placement.

[0051] Real-time video processing can be performed using any type of video data (e.g., video streams, video files, etc.) stored in the memory of any type of computerized system. For example, a user can load a video file and store it in the device's memory, or a video stream can be generated using the device's sensors. Furthermore, any object, such as a human face and body parts, animals, or inanimate objects like chairs, cars, or other objects, can be processed using computer-animated models.

[0052] In some implementations, when a specific modification is selected along with the content to be transformed, the elements to be transformed are identified by a computing device, and then their presence in video frames is detected and tracked. The elements of the object are modified according to the modification request, thereby transforming the frames of the video stream. Different methods can be used to transform the frames of the video stream for different types of transformations. For example, for transformations of frames that primarily involve changing the form of object elements, feature points of each element of the object are calculated (e.g., using an Active Shape Model (ASM) or other known methods). Then, a feature point-based mesh is generated for each of at least one element of the object. This mesh is used for the next stage of tracking the elements of the object in the video stream. During tracking, the aforementioned mesh for each element is aligned with the position of each element. Then, additional points are generated on the mesh. A first set of first points is generated for each element based on the modification request, and a second set of points is generated for each element based on the first set of points and the modification request. The frames of the video stream can then be transformed by modifying the elements of the object according to the first set of points, the second set of points, and the mesh. In such methods, the background of the modified object can also be changed or distorted by tracking and modifying the background.

[0053] In one or more embodiments, transformations that alter some regions of an object using its elements can be performed by calculating feature points for each element of the object and generating a mesh based on the calculated feature points. Points are generated on the mesh, and then various regions are generated based on these points. The elements of the object are then tracked by aligning the regions of each element with the positions of at least one element, and the properties of the regions can be modified according to modification requests, thereby transforming frames of the video stream. Transformations can be performed in different ways depending on the specific requirements for modifying the properties of the regions described above. Such modifications may involve: changing the color of the region; removing at least some portions of the region from the frames of the video stream; including one or more new objects in the regions based on modification requests; and modifying or distorting the elements of the regions or objects. In various embodiments, any combination of such modifications or other similar modifications can be used. For some models to be animated, some feature points can be selected as control points to determine the entire state space of the model's animation options.

[0054] In some implementations of computer animation models that use face detection to transform image data, faces are detected on the image using a specific face detection algorithm (e.g., Viola-Jones). The Active Shape Model (ASM) algorithm is then applied to the facial regions of the image to detect facial feature reference points.

[0055] In other implementations, other methods and algorithms suitable for face detection can be used. For example, in some implementations, landmarks are used to locate features that represent distinguishable points present in most of the images considered. For example, for a face landmark, the location of the left pupil could be used. Secondary landmarks can be used when the initial landmark is unrecognizable (e.g., if a person is wearing an eye patch). Such landmark recognition procedures can be used for any such object. In some implementations, a set of landmarks forms a shape. The shape can be represented as a vector using the coordinates of its midpoint. One shape is aligned with another shape by a similarity transformation (allowing translation, scaling, and rotation) that minimizes the average Euclidean distance between the points of the shapes. The average shape is the average of the aligned training shapes.

[0056] In some implementations, the search begins with a landmark search based on an average shape aligned with the position and size of the face determined by a global face detector. This search then repeats the following steps: proposing provisional shapes by adjusting the positions of shape points through template matching of the image texture around each point, and then conforming the provisional shapes to a global shape model until convergence occurs. In some systems, individual template matching is unreliable, and the shape model pools the results of weak template matchers to form a stronger overall classifier. The entire search is repeated at each level of the image pyramid, from coarse resolution to fine resolution.

[0057] The transformation system can be implemented by capturing image or video streams on a client device (e.g., client device 102) and performing complex image processing locally on client device 102 while maintaining a suitable user experience, computation time, and power consumption. Complex image processing may include size and shape changes, emotion transfer (e.g., changing a face from frowning to smiling), state transfer (e.g., aging an object, reducing apparent age, changing gender), style transfer, application of graphical elements, and any other suitable image or video processing implemented by a convolutional neural network that has been configured to execute efficiently on client device 102.

[0058] In some example implementations, a computer animation model for transforming image data can be used by a system in which a user can capture an image or video stream (e.g., a selfie) using a client device 102 having a neural network operating as part of a messaging client application 104 operating on client device 102. A transformation system operating within the messaging client application 104 determines the presence of a face within the image or video stream and provides a modification icon associated with the computer animation model to transform the image data, or the computer animation model can be presented as associated with the interface described herein. The modification icon includes changes that can be used to modify the user's face within the image or video stream as part of a modification operation. Once a modification icon is selected, the transformation system initiates a process to transform the user's image to reflect the selected modification icon (e.g., generating a smiley face on the user). In some implementations, the modified image or video stream can be presented in a graphical user interface displayed on a mobile client device as soon as the image or video stream is captured and the specified modification is selected. The transformation system can implement a complex convolutional neural network on a portion of the image or video stream to generate and apply the selected modification. In other words, once an edit icon has been selected, the user can capture an image or video stream and see the changes in real-time or near real-time. Furthermore, the changes can be persistent while a video stream is being captured and the selected edit icon continues to toggle. Machine learning neural networks can be used to achieve such modifications.

[0059] In some implementations, the graphical user interface (GUI) presenting the modifications performed by the transformation system can provide the user with additional interactive options. Such options may be based on the interface used to initiate content capture and selection for a specific computer animation model (e.g., initiated from a content creator user interface). In various implementations, modifications can be persistent after an initial selection of the modification icon. The user can turn modifications on or off by tapping or otherwise selecting a face modified by the transformation system and save the modification for later viewing or browsing to other areas of the imaging application. In cases where the transformation system modifies multiple faces, the user can globally turn modifications on or off by tapping or selecting a single face modified and displayed within the GUI. In some implementations, individual faces within a set of multiple faces can be modified individually, or such modifications can be toggled individually by tapping or selecting a single face or a series of individual faces displayed within the GUI.

[0060] As mentioned above, video table 310 stores video data, which in one embodiment is associated with messages whose records are maintained in message table 314. Similarly, image table 308 stores image data associated with messages whose message data is stored in entity table 302. Entity table 302 can associate various annotations from annotation table 312 with various images and videos stored in image table 308 and video table 310.

[0061] Story table 306 stores data about collections of messages and associated image, video, or audio data, compiled into collections (e.g., stories or libraries). The creation of a specific collection can be initiated by a specific user (e.g., each user whose records are maintained in entity table 302). A user can create "personal stories" in the form of collections of content that have already been created and sent / broadcast by that user. For this purpose, the user interface of messaging client application 104 may include user-selectable icons that allow the sending user to add specific content to his or her personal story.

[0062] Collections can also constitute "live stories," which are collections of content from multiple users created manually, automatically, or using a combination of manual and automatic technologies. For example, a "live story" can constitute a curated stream of user-submitted content from different locations and events. Users whose client devices have location services enabled and are at a common location event at a specific time can be presented with options to contribute content to a specific live story, for example, via the user interface of messaging client application 104. The messaging client application 104 can identify live stories to users based on their location. The end result is a "live story" told from a community perspective.

[0063] Another type of content collection is called a "location story," which allows users whose client devices 102 are located in a specific geographic location (e.g., at a university or on a university campus) to contribute to a specific collection. In some implementations, contributing to a location story may require a second level of authentication to verify that the end user belongs to a specific organization or other entity (e.g., is a student on a university campus).

[0064] Database 120 can also store data related to individual and shared AR sessions in AR session table 316. The data in AR session table 316 may include data communicated between AR session client controller 124 and another AR session client controller 124, as well as data communicated between AR session client controller 124 and AR session server controller 126. The data may include location data, biometric data, depth estimation, user identifiers in data streams (images or videos), data used to establish a common coordinate system for shared AR scenes, transformations between devices, session identifiers, etc.

[0065] Figure 4 This is a schematic diagram illustrating the structure of a message 400 according to some embodiments, which is generated by a messaging client application 104 for transmission to another messaging client application 104 or a messaging server application 114. The content of a particular message 400 is used to populate a message table 314 stored in a database 120, accessible to the messaging server application 114. Similarly, the content of the message 400 is stored in memory as “in transit” or “in flight” data of the client device 102 or application server 112. The message 400 is shown as including the following components:

[0066] ●Message Identifier 402: A unique identifier that identifies message 400.

[0067] ●Message text payload 404: The text to be generated by the user via the user interface of the client device 102 and included in message 400.

[0068] ●Message Image Payload 406: Image data captured by the camera component of the client device 102 or retrieved from the memory component of the client device 102 and included in message 400.

[0069] ● Message video payload 408: Video data captured by the camera device component or retrieved from the memory component of the client device 102 and included in message 400.

[0070] ●Message audio payload 410: Audio data captured by the microphone or retrieved from the memory component of the client device 102 and included in message 400.

[0071] ● Message annotation 412: Annotation data (e.g., filters, stickers, or other enhancements) representing annotations to be applied to message image payload 406, message video payload 408, or message audio payload 410 of message 400.

[0072] ● Message Duration Parameter 414: A parameter value, in seconds, indicating the amount of time that the content of the message (e.g., message image payload 406, message video payload 408, message audio payload 410) will be presented or made accessible to the user via the messaging client application 104.

[0073] ● Message geolocation parameter 416: Geographic location data (e.g., latitude and longitude coordinates) associated with the message's content payload. Multiple message geolocation parameter 416 values ​​may be included in the payload, each of which is associated with a content item included in the content (e.g., a specific image within the message image payload 406 or a specific video within the message video payload 408).

[0074] ● Message Story Identifier 418: An identifier value that identifies one or more sets of content (e.g., "story") associated with a specific content item in the message image payload 406 of message 400. For example, multiple images within the message image payload 406 may each be associated with multiple sets of content using their own identifier values.

[0075] ● Message Tag 420: Each message 400 can be labeled with multiple tags, each tag indicating the subject of the content included in the message payload. For example, in the case where a specific image included in the message image payload 406 depicts an animal (e.g., a lion), the tag value can be included within the message tag 420 indicating the relevant animal. The tag value can be manually generated based on user input, or it can be automatically generated using, for example, image recognition.

[0076] ●Message sender identifier 422: An identifier (e.g., a messaging system identifier, email address, or device identifier) ​​indicating the user of the client device 102 on which message 400 is generated and from which message 400 is sent.

[0077] ● Message Recipient Identifier 424: An identifier (e.g., a messaging system identifier, email address, or device identifier) ​​indicating the user of the client device 102 to which message 400 is addressed.

[0078] The content (e.g., values) of each component of message 400 can be pointers to locations in tables storing content data values. For example, image values ​​in message image payload 406 can be pointers to locations (or addresses) within image table 308. Similarly, values ​​in message video payload 408 can point to data stored in video table 310, values ​​in message annotation 412 can point to data stored in annotation table 312, values ​​in message story identifier 418 can point to data stored in story table 306, and values ​​in message sender identifier 422 and message receiver identifier 424 can point to user records stored in entity table 302.

[0079] Although the flowchart below describes operations as a sequential process, many operations can be executed in parallel or simultaneously. Furthermore, the order of operations can be rearranged. A process terminates when its operations are completed. A process can correspond to a method, a program, etc. The steps of a method can be executed in whole or in part, can be combined with some or all steps from other methods, and can be executed by any number of different systems (such as...). Figure 1 and / or Figure 9 The system described herein may be executed or any part of the system (such as a processor included in any system) may be executed.

[0080] Figure 5 This is a flowchart of a process 500 for generating depth estimates based on biological data, according to one embodiment. In one embodiment, process 500 may be executed by an AR session server controller 126 in a messaging server system 108 to determine the distance between a first client device 102 and a user (e.g., a second user) of a second client device 102. The first client device 102 and the second client device 102 may be included in a shared AR session. In one example, the shared AR session includes a real-time messaging session.

[0081] At operation 502, the AR session server controller 126 receives location data from a first client device 102 associated with a first user. In one embodiment, the first client device 102 generates the location data based on analysis of a data stream including images of a second user associated with a second client device 102. The data stream may be video of the second user. For example, such as... Figure 7 As shown, the first client device 102 can use a camera device included in the first client device 102 to capture data streams or video. Figure 8As shown, the field of view (or camera frame) of the camera device may include a second user and can be displayed on the display of the first client device 102. In this embodiment, the first client device 102 may use skeletal tracking to track the second user in the image. In one embodiment, to generate positioning data, the first client device 102 identifies the position or pose of the second user's skeleton within the camera frame. The first client device 102 may also generate positioning data by identifying the position or pose of the second user's face within the camera frame. At this stage, the proportions of the skeleton or face and the distance between the second user and the first user (or the first client device 102) may be ambiguous. For example, this distance may be ambiguous due to the lack of a binocular camera or depth sensor in the first client device 102.

[0082] At operation 504, the AR session server controller 126 receives biometric data of the second user from the second client device 102. The biometric data may be based on the output of sensors or cameras included in the second device. For example, the biometric data may be based on an image of the second user captured by a camera included in the second client device 102. For example, an image of the second user may be captured by a front-facing camera included in the second client device 102. The second client device 102 may also generate the biometric data of the second user based on the output of a depth sensor or binocular camera included in the second client device 102. In one embodiment, the biometric data of the second user is the dimensions of the second user's face.

[0083] At operation 506, the AR session server controller 126 uses location data and the second user's biometric data to determine the distance between the second user and the first client device 102. For example, using the location data, the AR session server controller 126 can determine the position of the second user's skeleton in the data stream and identify the second user's face in the data stream. The AR session server controller 126 can use biometric data, including the actual size of the second user's face, to determine the ratio between the size of the second user's face in the data stream and the actual size of the second user's face. The AR session server controller 126 can then use this ratio to determine the distance between the second user and the first client device 102 (e.g., depth estimation).

[0084] By determining a depth estimate of the second user in the data stream, the first client device 102 can, for example, generate an enhanced user experience including visual or audio effects to be applied to the image of the second user displayed on the first client device 102 in a shared AR session. For example, the visual or audio effects may include clothing to be worn by the second user, effects to be applied to the second user's face, voice effects to be applied to the second user, etc.

[0085] Figure 6 A process 600 for generating depth estimates based on biometric data according to one embodiment is illustrated. In one embodiment, process 600 may be executed by an AR session client controller 124 included in a first client device 102 to determine the distance between the first client device 102 and a user (e.g., a second user) of a second client device 102. The first client device 102 and the second client device 102 may be included in a shared AR session. In one example, the shared AR session includes a real-time messaging session.

[0086] At operation 602, the AR session client controller 124 of the first client device 102 uses a camera device included in the first client device 102 to capture a data stream containing an image of the user of the second client device 102. Figure 7 As shown, the first user (User B) uses the camera device in the first client device 102 to capture images or videos of the second user (User A) on the second client device 102. The data stream can be the video of the second user (User A). Figure 8 As shown, the field of view (or camera frame) of the camera device may include a second user (user A) and may be displayed on the display of the first client device 102.

[0087] In one implementation, at operation 604, the first client device 102 generates positioning data based on analysis of the data stream by identifying the position or pose of the user's skeleton or face within the camera frame. In this implementation, the first client device 102 may use skeletal tracking to track the user in the image. At this stage, the proportions of the skeleton or face and the distance between the user and the first client device 102 may be ambiguous. For example, this distance may be ambiguous due to the lack of a binocular camera or depth sensor in the first client device 102.

[0088] At operation 606, the AR session client controller 124 of the first client device 102 receives the user's biometric data from the second client device 102. The biometric data may be based on the output of sensors or camera devices included in the second client device 102. For example, the biometric data may be based on an image of the user captured by a camera device included in the second client device 102. For example, the image of the user may be captured by a front-facing camera device included in the second client device 102. The second client device 102 may also generate the user's biometric data based on the output of a depth sensor or binocular camera device included in the second client device 102. In one embodiment, the user's biometric data is the dimensions of the user's face.

[0089] At operation 608, the AR session client controller 124 of the first client device 102 uses location data and the user's biometric data to determine the distance between the user and the first device. For example, using the location data, the AR session server controller 126 can determine the position of the user's skeleton in the data stream and identify the user's face in the data stream. The AR session server controller 126 can use biometric data, including the actual size of the user's face, to determine the ratio between the size of the user's face in the data stream and the actual size of the user's face. The AR session server controller 126 can then use this ratio to determine the distance between the user and the first client device 102 (e.g., depth estimation).

[0090] By determining a depth estimate of the user in the data stream, the first client device 102 can, for example, generate an enhanced user experience including visual or audio effects to be applied to the image of the user displayed on the first client device 102 in a shared AR session. For example, the visual or audio effects may include clothing to be worn by the second user, effects to be applied to the second user's face, voice effects to be applied to the second user, etc.

[0091] Figure 7 An example 700 is shown whereby a first user (user B) uses a first client device 102 to capture images or videos of a second user (user A) according to one embodiment. Figure 7 As shown, the second user (user A) can be associated with the second client device 102. For example, the second user (user A) in Figure 7 Zhongzheng holds a second client device 102. The second client device 102 generates biometric data associated with a second user (user A) and sends the biometric data to the AR session server controller 126 in the messaging server system 108 and / or to the AR session client controller 124 in the first client device 102.

[0092] Figure 8An example 800 is shown using an image captured by a camera device included in a first client device 102, according to one embodiment. Figure 8 The image shows a second user (user A) within the field of view of the camera device of the first client device 102. In one embodiment, the field of view of the camera device frame or camera device, which includes the second user (user A) and is included in the first client device 102, is displayed on the display of the first client device 102.

[0093] Figure 9 This is a graphical representation of machine 900, within which instructions 908 (e.g., software, programs, applications, applets, or other executable code) can be executed to cause machine 900 to perform any or more of the methods discussed herein. For example, instructions 908 can cause machine 900 to perform any or more of the methods described herein. Instructions 908 transform a general, unprogrammed machine 900 into a specific machine 900 programmed to perform the described and illustrated functions in the described manner. Machine 900 can operate as a standalone device or can be coupled (e.g., networked) to other machines. In a networked deployment, machine 900 can operate as a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. Machine 900 may include, but is not limited to: server computers, client computers, personal computers (PCs), tablet computers, laptop computers, netbooks, set-top boxes (STBs), PDAs, entertainment media systems, cellular phones, smartphones, mobile devices, wearable devices (e.g., smartwatches), smart home devices (e.g., smart appliances), other smart devices, web devices, network routers, network switches, network bridges, or any machine capable of sequentially or otherwise executing instructions 908 specifying actions to be taken by machine 900. Furthermore, although only a single machine 900 is shown, the term "machine" should also be considered as a collection of machines that individually or jointly execute instructions 908 to perform any or more of the methods discussed herein.

[0094] Machine 900 may include processor 902, memory 904, and I / O components 938, which may be configured to communicate with each other via bus 940. In an example embodiment, processor 902 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an ASIC, a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, processor 906 and processor 910 that execute instruction 908. The term "processor" is intended to include multi-core processors, which may include two or more independent processors (sometimes referred to as "cores") capable of executing instructions simultaneously. Although Figure 9 Multiple processors 902 are shown, but machine 900 may include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.

[0095] Memory 904 includes main memory 912, static memory 914, and storage cells 916, all of which are accessible by processor 902 via bus 940. Main memory 904, static memory 914, and storage cells 916 store instructions 908 embodying any or more of the methods or functions described herein. Instructions 908 may also reside wholly or partially in main memory 912, in static memory 914, in machine-readable medium 918 within storage cells 916, within at least one processor in processor 902 (e.g., within the processor's cache memory), or in any suitable combination thereof during execution by machine 900.

[0096] I / O component 938 may include various components for receiving input, providing output, generating output, transmitting information, exchanging information, capturing measurement results, etc. The specific I / O component 938 included in a particular machine will depend on the type of machine. For example, portable machines such as mobile phones may include touch input devices or other such input mechanisms, while headless server machines may not include such touch input devices. It will be understood that I / O component 938 may include... Figure 9Many other components are not shown. In various example embodiments, I / O component 938 may include user output component 924 and user input component 926. User output component 924 may include visual components (e.g., displays, such as plasma display panels (PDPs), light-emitting diode (LED) displays, liquid crystal displays (LCDs), projectors, or cathode ray tubes (CRTs)), auditory components (e.g., speakers), tactile components (e.g., vibration motors, resistance mechanisms), other signal generators, etc. User input component 926 may include alphanumeric input components (e.g., keyboards, touchscreens configured to receive alphanumeric input, photoelectric keyboards, or other alphanumeric input components), pointing-based input components (e.g., mice, touchpads, trackballs, joysticks, motion sensors, or other pointing instruments), tactile input components (e.g., physical buttons, touchscreens or other tactile input components that provide the position and / or force of a touch or touch gesture), audio input components (e.g., microphones), etc.

[0097] In another example implementation, I / O component 938 may include biometric component 928, motion component 930, environmental component 932, or positioning component 934, as well as various other components. For example, biometric component 928 includes components for detecting expressions (e.g., hand gestures, facial expressions, voice expressions, body posture, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), and identifying a person (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or EEG-based recognition). Motion component 930 includes accelerometer components (e.g., accelerometers), gravity sensor components, and rotation sensor components (e.g., gyroscopes). Environmental component 932 includes, for example, one or more camera devices (with still image / photograph and video capabilities), lighting sensor components (e.g., photometers), temperature sensor components (e.g., one or more thermometers for detecting ambient temperature), humidity sensor components, pressure sensor components (e.g., barometers), acoustic sensor components (e.g., one or more microphones for detecting background noise), proximity sensor components (e.g., infrared sensors for detecting nearby objects), gas sensors (e.g., gas detection sensors for detecting the concentration of hazardous gases for safety reasons or for measuring pollutants in the atmosphere), or other components that can provide indications, measurements, or signals corresponding to the surrounding physical environment. Positioning component 934 includes position sensor components (e.g., GPS receiver components), altitude sensor components (e.g., altimeters or barometers for detecting air pressure from which altitude can be obtained), orientation sensor components (e.g., magnetometers), etc.

[0098] Various technologies can be used to achieve communication. I / O component 938 also includes communication component 936, which is operable to couple machine 900 to network 920 or device 922 via a suitable coupling or connection. For example, communication component 936 may include a network interface component or another suitable device interfaced with network 920. In further examples, communication component 936 may include wired communication components, wireless communication components, cellular communication components, near field communication (NFC) components, etc. Components (e.g.) (low power consumption) Components and other communication components that provide communication in other forms. Device 922 can be any peripheral device from another machine or various peripheral devices (e.g., a peripheral device coupled via USB).

[0099] Furthermore, the communication component 936 can detect identifiers or include components operable to detect identifiers. For example, the communication component 936 may include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional barcodes, such as Universal Product Code (UPC) barcodes; multi-dimensional barcodes, such as Quick Response (QR) codes, Aztec codes, data matrices, dataglyphs, MaxiCode, PDF417, Ultra Code, UCC RSS-2D barcodes, and other optical codes), or an auditory detection component (e.g., a microphone for identifying audio signals from the tag). Additionally, various information can be obtained via the communication component 936, such as location via Internet Protocol (IP) geolocation, etc. The location of signal triangulation, the location of NFC beacon signals that can be detected to indicate a specific location, etc.

[0100] Various memories (e.g., main memory 912, static memory 914, and / or the memory of processor 902) and / or storage units 916 may store one or more sets of instructions and data structures (e.g., software) that embody or are used by any one or more of the methods or functions described herein. These instructions (e.g., instruction 908), when executed by processor 902, cause various operations to implement the disclosed embodiments.

[0101] Instructions 908 can be sent or received on network 920 using a transmission medium, via a network interface device (e.g., a network interface component included in communication component 936), and using any of several known transmission protocols (e.g., Hypertext Transfer Protocol (HTTP)). Similarly, instructions 908 can be sent or received using a transmission medium via a coupling to device 922 (e.g., peer-to-peer coupling).

[0102] Figure 10 This is a block diagram 1000 illustrating a software architecture 1004 that can be installed on any one or more of the devices described herein. The software architecture 1004 is supported by hardware such as a machine 1002 including a processor 1020, memory 1026, and I / O components 1038. In this example, the software architecture 1004 can be conceptualized as a layer stack, where each layer provides specific functionality. The software architecture 1004 includes layers such as an operating system 1012, libraries 1010, frameworks 1008, and applications 1006. Operationally, application 1006 invokes API call 1050 through the software stack and receives message 1052 in response to API call 1050.

[0103] Operating system 1012 manages hardware resources and provides public services. Operating system 1012 includes, for example, kernel 1014, services 1016, and drivers 1022. Kernel 1014 acts as an abstraction layer between the hardware layer and other software layers. For example, kernel 1014 provides functions such as memory management, processor management (e.g., scheduling), component management, network and security settings. Services 1016 can provide other public services to other software layers. Drivers 1022 are responsible for controlling or interfacing with the underlying hardware. For example, drivers 1022 may include display drivers, camera drivers, etc. Driver or Low-power drivers, flash memory drivers, serial communication drivers (e.g., Universal Serial Bus (USB) drivers), Drivers, audio drivers, power management drivers, etc.

[0104] Library 1010 provides low-level public infrastructure used by application 1006. Library 1010 may include system library 1018 (e.g., the C standard library), which provides functions such as memory allocation functions, string manipulation functions, mathematical functions, etc. Additionally, library 1010 may include API library 1024, such as media libraries (e.g., libraries for supporting the rendering and manipulation of various media formats, such as Moving Picture Experts Group-4 (MPEG4), Advanced Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Picture Experts Group (JPEG or JPG) or Portable Web Graphics (PNG)), graphics libraries (e.g., OpenGL frameworks for rendering graphic content on a display in two-dimensional (2D) and three-dimensional (3D) formats), database libraries (e.g., SQLite for providing various relational database functionalities), web libraries (e.g., WebKit for providing web browsing functionality), etc. Library 1010 may also include a wide variety of other libraries 1028 to provide many other APIs to application 1006.

[0105] Framework 1008 provides advanced public infrastructure for use by Application 1006. For example, Framework 1008 provides various graphical user interface (GUI) functionalities, advanced resource management, and advanced location services. Framework 1008 can provide a wide range of other APIs that can be used by Application 1006, some of which may be specific to a particular operating system or platform.

[0106] In an example implementation, application 1006 may include home application 1036, contact application 1030, browser application 1032, book reader application 1034, location application 1042, media application 1044, messaging application 1046, game application 1048, and other applications in a broad category such as third-party application 1040. Application 1006 is a program that performs the functions defined in the program. One or more applications 1006 can be created in various ways using various programming languages, such as object-oriented programming languages ​​(e.g., Objective-C, Java, or C++) or procedural programming languages ​​(e.g., C or assembly language). In a particular example, third-party application 1040 (e.g., used by an entity other than the vendor of a particular platform using Android) TM or iOS TM Applications developed using a Software Development Kit (SDK) can be used on platforms such as iOS. TM ANDROID TM , Mobile software running on the mobile operating system or other mobile operating systems. In this example, third-party application 1040 can call API call 1050 provided by operating system 1012 to facilitate the functions described herein.

[0107] The present invention can also be implemented through the following embodiments.

[0108] Implementation Plan 1. A method comprising:

[0109] Location data is received from a first device associated with a first user, the first device generating the location data based on analysis of a data stream including images of a second user associated with a second device;

[0110] Receive biometric data of the second user from the second device, the biometric data being based on output from a sensor or camera device included in the second device; and

[0111] The location data and the second user's biometric data are used to determine the distance between the second user and the first device.

[0112] Implementation Scheme 2. The method according to Implementation Scheme 1, wherein the data stream is a video of the second user, wherein the video is captured by a camera device included in the first device.

[0113] Implementation Scheme 3. According to the method of Implementation Scheme 1, wherein the first device generating the positioning data based on the analysis of the data stream further includes:

[0114] The first device identifies the position or posture of the second user's skeleton within the camera device frame.

[0115] Implementation Scheme 4. The method according to Implementation Scheme 1, wherein the first device generating the positioning data based on the analysis of the data stream further includes:

[0116] The first device identifies the position or posture of the second user's face within the camera device frame.

[0117] Implementation Scheme 5. The method according to Implementation Scheme 1, wherein the biometric data of the second user is the size of the second user's face.

[0118] Implementation Scheme 6. The method according to Implementation Scheme 1, wherein the second user's biometric data is based on the output from a depth sensor or binocular camera device included in the second device.

[0119] Implementation Scheme 7. The method according to Implementation Scheme 1, wherein the first device and the second device are included in a shared augmented reality (AR) session.

[0120] Implementation Scheme 8. The method according to Implementation Scheme 7, wherein the shared AR session includes a real-time messaging session.

[0121] Implementation Scheme 9. A system comprising:

[0122] Processor; and

[0123] The memory thereon stores instructions that, when executed by the processor, cause the system to perform operations including:

[0124] Location data is received from a first device associated with a first user, the first device generating the location data based on analysis of a data stream including images of a second user associated with a second device;

[0125] Receive biometric data of the second user from the second device, the biometric data being based on output from a sensor or camera device included in the second device; and

[0126] The location data and the second user's biometric data are used to determine the distance between the second user and the first device.

[0127] Implementation Scheme 10. The system according to Implementation Scheme 9, wherein the data stream is a video of the second user, wherein the video is captured by a camera device included in the first device.

[0128] Implementation Scheme 11. The method according to Implementation Scheme 9, wherein the first device generating the positioning data based on the analysis of the data stream further includes:

[0129] The first device identifies the position or posture of the second user's skeleton within the camera device frame.

[0130] Implementation Scheme 12. The system according to Implementation Scheme 9, wherein the first device generating the positioning data based on the analysis of the data stream further includes:

[0131] The first device identifies the position or posture of the second user's face within the camera device frame.

[0132] Implementation Scheme 13. The system according to Implementation Scheme 9, wherein the biometric data of the second user is the size of the second user's face.

[0133] Implementation Scheme 14. The system according to Implementation Scheme 9, wherein the second user's biometric data is based on the output from a depth sensor or binocular camera device included in the second device.

[0134] Implementation Scheme 15. The system according to Implementation Scheme 9, wherein the first device and the second device are included in a shared augmented reality (AR) session.

[0135] Implementation Scheme 16. The system according to Implementation Scheme 15, wherein the shared AR session includes a real-time messaging session.

[0136] Implementation Scheme 17. A non-transitory computer-readable storage medium having instructions stored thereon, the instructions causing the processor to perform operations including the following when executed by a processor:

[0137] Location data is received from a first device associated with a first user, the first device generating the location data based on analysis of a data stream including images of a second user associated with a second device;

[0138] Receive biometric data of the second user from the second device, the biometric data being based on output from a sensor or camera device included in the second device; and

[0139] The location data and the second user's biometric data are used to determine the distance between the second user and the first device.

[0140] Implementation Scheme 18. A method comprising:

[0141] A data stream of images of a user, including a second device, is captured using a camera device included in the first device.

[0142] The first device generates location data based on the analysis of the data stream by recognizing the position or posture of the user's bones or face in the camera device frame;

[0143] The first device receives biometric data of the second user from the second device, the biometric data being based on output from a sensor or camera device included in the second device; and

[0144] The distance between the user and the first device is determined by the first device using the location data and the user's biometric data.

[0145] Implementation Scheme 19. The method according to Implementation Scheme 18, wherein the user's biometric data is the size of the user's face.

[0146] Implementation Scheme 20. The method according to Implementation Scheme 19, wherein the user's biometric data is based on the output from a depth sensor or binocular camera device included in the second device.

Claims

1. A method comprising: generating, by a first device, positioning data based on an analysis of a data stream, wherein the data stream comprises an image of a second user of a second device, by identifying a position or pose of a skeleton or face of the second user in a frame of a camera; receiving, by the first device, biometric data of the second user from the second device, the biometric data being based on a biometric feature of the second user from an output of a sensor or camera included in the second device; and determining, by the first device, a distance of the second user from the first device using the positioning data and the biometric data of the second user.

2. The method of claim 1, further comprising: capturing a video of the second user, wherein the data stream is the video of the second user.

3. The method of claim 1, wherein, the biometric data of the second user is a size of a face of the second user.

4. The method of claim 1, wherein, the biometric data of the second user is based on an output from a depth sensor or binocular camera included in the second device.

5. The method of claim 1, wherein, the first device and the second device are included in a shared augmented reality session.

6. The method of claim 5, wherein, the shared augmented reality session comprises a real-time messaging session.

7. A device comprising: a processor; and a memory having instructions stored thereon that, when executed by the processor, cause the device to perform operations comprising: capturing, using a camera included in the device, a data stream comprising an image of a second user of a second device; generating positioning data based on an analysis of the data stream by identifying a position or pose of a skeleton or face of the second user in a frame of a camera; receiving biometric data of the second user from the second device, the biometric data being based on a biometric feature of the second user from an output of a sensor or camera included in the second device; and determining a distance of the second user from the device using the positioning data and the biometric data of the second user. the data stream is a video of the second user. the biometric data of the second user is a size of a face of the second user.

8. The apparatus of claim 7, wherein, the biometric data of the second user is based on an output from a depth sensor or binocular camera included in the second device.

9. The apparatus of claim 7, wherein, the device and the second device are included in a shared augmented reality session.

10. The apparatus of claim 7, wherein, the shared augmented reality session comprises a real-time messaging session.

11. The apparatus of claim 7, wherein, 13. A non-transitory computer-readable storage medium having instructions stored thereon that, when executed by a processor of a first device, cause the processor to perform operations comprising:

12. The apparatus of claim 11, wherein, capturing, using a camera, a data stream comprising an image of a second user of a second device; generating positioning data based on an analysis of the data stream by identifying a position or pose of a skeleton or face of the second user in a frame of a camera; generating biometric data of the second user from the second device, the biometric data being based on an output of a sensor or camera included in the second device; and ​ ​ ​ determining a distance of the second user from the first device using the positioning data and the biometric data of the second user.

14. The non-transitory computer-readable storage medium of claim 13, wherein, the data stream is a video of the second user.

15. The non-transitory computer-readable storage medium of claim 13, wherein, the biometric data of the second user is a size of a face of the second user.

16. The non-transitory computer-readable storage medium of claim 13, wherein, the biometric data of the second user is based on output from a depth sensor or a binocular camera included in the second device.

17. The non-transitory computer-readable storage medium of claim 13, wherein, the first device and the second device are included in a shared augmented reality session.

18. The non-transitory computer-readable storage medium of claim 17, wherein, the shared augmented reality session includes a real-time messaging session.

Citation Information

Patent Citations

  • Determining size of virtual object

    CN109997175A

  • Biometric analysis of users to determine user locations

    CN110785766A