Scaling 3D volumes in augmented reality

Through pinch hand posture tracking technology, the XR system can scale virtual objects in real time, solving the problem of lack of effective user intention determination in the prior art and providing an intuitive user interaction method.

CN120359486APending Publication Date: 2025-07-22SNAP INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202380085630.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-12-13
Filing Date
2023-12-11
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

Existing XR systems have difficulty effectively determining the user's intentions to scale virtual objects without a keyboard or unable to track hands.

Method used

By pinching the hand posture or pinching posture, the user's hand posture is tracked by the camera device, the pinch position is determined, and the virtual object is scaled based on the pinch position and the center point of the virtual object.

Benefits of technology

It realizes real-time and continuous scaling of virtual objects through hand postures in the XR system, providing a user interaction method without keyboard or additional input devices, and enhancing the intuitiveness and interactivity of the user interface.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120359486A_ABST
    Figure CN120359486A_ABST
Patent Text Reader

Abstract

An augmented reality (XR) system provides a method for scaling a virtual object in an XR user interface of the XR system. The method includes providing an XR user interface of an XR system to a user, wherein the XR user interface includes a virtual object displayed to the user. The XR system determines a kneading position of a kneading hand gesture being made by the user, and scales the virtual object based on the kneading position and a virtual object center point of the virtual object. The XR system redisplays the scaled virtual object to the user in the XR user interface.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority Claim

[0002] This application claims the benefit of U.S. Patent Application Serial No. 18 / 065,201, filed on December 13, 2022, the entire content of which is incorporated herein by reference. Technical Field

[0003] This disclosure generally relates to user interfaces and, more particularly, to user interfaces for augmented reality or virtual reality. Background Art

[0004] A head-wearable device can be implemented with a transparent or semi-transparent display through which a user of the head-wearable device can view the surrounding environment. Such a head-wearable device enables a user to view the surrounding environment through the transparent or semi-transparent display and also enables the user to see objects (e.g., virtual objects such as renderings of 2D or 3D graphical models, images, videos, text, etc.) that are generated for display as part of and / or overlaid on the surrounding environment. This is generally referred to as "augmented reality" or "AR". The head-wearable device can also completely occlude the user's field of view and display a virtual environment through which the user can move or be moved. This is generally referred to as "virtual reality" or "VR". In a hybrid form, a view of the surrounding environment is captured using a camera device and then the view is displayed to the user together with augmentations on a display that occludes the user's eyes. As used herein, unless the context otherwise indicates, the term extended reality (XR) refers to augmented reality, virtual reality, and any hybrid of these technologies.

[0005] A user of a head-wearable device can access and use computer software applications to perform various tasks or engage in entertainment activities. To use a computer software application, the user interacts with a user interface provided by the head-wearable device. Brief Description of the Drawings

[0006] In the drawings (which are not necessarily drawn to scale), like reference numerals may describe similar components in different views. To easily identify the discussion of any particular element or action, one or more of the most significant digits in the reference numeral refers to the figure number in which the element was first introduced. Some non-limiting examples are shown in the figures of the drawings, in which:

[0007] Figure 1A is a perspective view of a head-mounted device according to some examples.

[0008] Figure 1B shows another view of the Figure 1A head-mounted device according to some examples.

[0009] Figure 2 is a diagrammatic representation of a machine in the form of a computer system within which a set of instructions can be executed to cause the machine to perform any one or more of the methods discussed herein.

[0010] Figure 3A shows a collaboration diagram of components of an XR system using hand tracking for user input, according to some examples.

[0011] Figure 3B shows a process flow diagram of a method using a pinch hand gesture or pinch pose as user input, according to some examples.

[0012] Figure 3C shows using a pinch hand gesture or pinch pose to rescale a virtual object, according to some examples.

[0013] Figure 3D is another illustration of using a pinch hand gesture or pinch pose to rescale a virtual object, according to some examples.

[0014] Figure 3E is another illustration of using a pinch hand gesture or pinch pose to rescale a virtual object, according to some examples.

[0015] Figure 4 shows a system of a head-mounted device, according to some examples.

[0016] Figure 5 is a diagrammatic representation of a networked environment in which the present disclosure may be deployed, according to some examples.

[0017] Figure 6 is a diagrammatic representation of a data structure maintained in a database, according to some examples.

[0018] Figure 7 is a diagrammatic representation of a messaging system having both client-side functionality and server-side functionality, according to some examples.

[0019] Figure 8 is a block diagram showing a software architecture, according to some examples. Detailed Description

[0020] Hand tracking is a way to provide user input from a user to an XR user interface provided by an XR system. The XR system uses a camera device and computer vision methods to track one or more hands of the user. The XR system determines the hand gesture or pose being made by the user based on video images captured by the camera device. In some XR systems, the XR user interface includes one or more virtual objects manipulated by the user, which is referred to as direct manipulation of virtual objects (DMVO). Manipulation of virtual objects can include operations for changing the appearance of the virtual objects, such as scaling the virtual object to be larger or smaller. Since the XR system may not have a keyboard, pointing device, or may not even be able to track both hands of the user, there is a need for a method for determining the user's intention during a scaling operation of a virtual object using one hand of the user.

[0021] Certain examples of the present disclosure provide methods for scaling virtual objects using a pinching hand gesture or pose. The user makes a pinching hand gesture or pose at a location near the virtual object. Depending on the location where the user pinches, the virtual object is scaled to increase in size or decrease in size.

[0022] In some examples, the XR system provides the XR user interface to the user, where the XR user interface includes a virtual object displayed to the user. The XR system determines the pinching location of the pinching hand gesture being made by the user and scales the virtual object based on the pinching location and the virtual object center point of the virtual object. The XR system redisplay the scaled virtual object to the user in the XR user interface.

[0023] In some examples, determining the pinching location further includes using one or more camera devices of the XR system to capture tracking video frame data. The XR system determines hand tracking data based on the tracking video frame data and determines the pinching location based on the hand tracking data.

[0024] In some examples, scaling the virtual object based on the pinching location and the virtual object center point of the virtual object further includes generating a virtual object vector based on the center point of the virtual object and the pinching location, and generating a pinching vector based on the center point of the virtual object and the virtual object pinching collider of the virtual object. The XR system scales the virtual object based on the virtual object vector and the pinching vector.

[0025] In some examples, the XR system detects that the user still maintains the pinching hand gesture, and in response to detecting that the pinching hand gesture is still maintained, continues to scale the virtual object based on the scaling vector.

[0026] In some examples, the XR system includes a head-worn device.

[0027] Other technical features may be readily apparent to those skilled in the art in light of the accompanying drawings, description, and claims.

[0028] Figure 1A is a perspective view of a head-mounted device 100 according to some examples. The head-mounted device 100 may be a client device of a computing system 502 of an XR system such as Figure 5 The head-mounted device 100 may include a frame 102 made of any suitable material such as plastic or metal including any suitable shape memory alloy. In one or more examples, the frame 102 includes a first optical element holder or left optical element holder 104 (e.g., a display or lens holder) and a second optical element holder or right optical element holder 106 connected by a bridge 112. A first optical element or left optical element 108 and a second optical element or right optical element 110 may be disposed within the left optical element holder 104 and the right optical element holder 106, respectively. The right optical element 110 and the left optical element 108 may be lenses, displays, display components, or a combination of the foregoing. Any suitable display component may be provided in the head-mounted device 100.

[0029] The frame 102 additionally includes a left arm piece or left temple piece 122 and a right arm piece or right temple piece 124. In some examples, the frame 102 may be formed from a single piece of material to have a unified or integral construction.

[0030] The head-mounted device 100 may include a computing device, such as a computer 120, which may be of any suitable type to be carried by the frame 102, and in one or more examples, may have a suitable size and shape to be partially disposed within one of the left temple piece 122 or the right temple piece 124. The computer 120 may include one or more processors having a memory, wireless communication circuitry, and a power source. As discussed below, the computer 120 includes low-power circuitry 426, high-speed circuitry 428, and a display processor. Various other examples may include these elements in different configurations or integrated in different ways. Additional details of aspects of the computer 120 may be implemented as shown for the machine 200 discussed herein.

[0031] The computer 120 additionally includes a battery 118 or other suitable portable power supply. In some examples, the battery 118 is disposed within the left temple piece 122 and is electrically coupled to the computer 120 disposed within the right temple piece 124. The head-mounted device 100 may include a connector or port (not shown) suitable for charging the battery 118, a wireless receiver, transmitter, or transceiver (not shown), or a combination of such devices.

[0032] The head-wearable device 100 includes a first camera device or a left camera device 114 and a second camera device or a right camera device 116. Although two camera devices are depicted, other examples contemplate the use of a single or additional (i.e., more than two) camera devices.

[0033] In some examples, in addition to the left camera device 114 and the right camera device 116, the head-wearable device 100 further includes any number of input sensors or other input / output devices. Such sensors or input / output devices may additionally include biometric sensors, positioning sensors, motion sensors, and the like.

[0034] In some examples, the left camera device 114 and the right camera device 116 provide tracking video frame data for the head-wearable device 100 to extract 3D information from real-world scenes.

[0035] The head-wearable device 100 may also include a touchpad 126 that is mounted to or integrated with one or both of the left temple piece 122 and the right temple piece 124. The touchpad 126 is typically arranged vertically and, in some examples, is approximately parallel to the user's temple. As used herein, being typically vertically aligned means that the touchpad is closer to vertical than horizontal, although it may be more vertical than that. Additional user input may be provided by one or more buttons 128, which in the example shown are disposed on the outer upper edges of the left optical element holder 104 and the right optical element holder 106. The one or more touchpads 126 and buttons 128 provide means by which the head-wearable device 100 can receive input from a user of the head-wearable device 100.

[0036] Figure 1B The head-wearable device 100 is shown from the perspective of a user wearing the head-wearable device 100. For clarity, many of the elements shown in Figure 1A are omitted. As Figure 1A described, Figure 1B the head-wearable device 100 shown includes a left optical element 140 and a right optical element 144, which are respectively fixed within a left optical element holder 132 and a right optical element holder 136.

[0037] The head-wearable device 100 includes: a right front optical assembly 130, the right front optical assembly 130 including a left near-eye display 150 and a right near-eye display 134; and a left front optical assembly 142, the left front optical assembly 142 including a left projector 146 and a right projector 152.

[0038] In some examples, the near-eye display is a waveguide. The waveguide includes reflective or diffractive structures (e.g., gratings and / or optical elements such as mirrors, lenses, or prisms). The light 138 emitted by the right projector 152 encounters the diffractive structure of the waveguide of the right near-eye display 134, which directs the light toward the user's right eye to provide an image on or in the right optical element 144 that is superimposed on the view of the real-world scene seen by the user. Similarly, the light 148 emitted by the left projector 146 encounters the diffractive structure of the waveguide of the left near-eye display 150, which directs the light toward the user's left eye to provide an image on or in the left optical element 140 that is superimposed on the view of the real-world scene seen by the user. The combination of the graphics processing unit, the image display driver, the right front optical assembly 130, the left front optical assembly 142, the left optical element 140, and the right optical element 144 provides the optical engine of the head-mounted device 100. The head-mounted device 100 uses the optical engine to generate a superimposition of the view of the real-world scene of the user, including displaying a user interface to the user of the head-mounted device 100.

[0039] However, it should be understood that other display technologies or configurations may be utilized within the optical engine to display images to the user within the user's field of view. For example, instead of projectors and waveguides, an LCD, LED, or other display panel or surface may be provided.

[0040] In use, the user of the head-mounted device 100 will see information, content, and various user interfaces on the near-eye display. As described in more detail herein, the user can then use the touchpad 126 and / or the buttons 128, touch inputs or voice inputs on an associated device (e.g., Figure 4 the mobile device 414 shown), and / or hand movements, positioning, and locations recognized by the head-mounted device 100 to interact with the head-mounted device 100.

[0041] In some examples, the optical engine of the XR system is incorporated into a lens that contacts the user's eye, such as a contact lens. The XR system uses the contact lens to generate images of the XR experience.

[0042] In some examples, the head-mounted device 100 includes an XR system. In some examples, the head-mounted device 100 is a component of an XR system that includes additional computing components. In some examples, the head-mounted device 100 is a component within an XR system that includes an additional user input system or device.

[0043] Machine architecture

[0044] Figure 2is an illustrative representation of a machine 200 within which instructions 202 (e.g., software, programs, applications, applets, apps, or other executable code) can be executed to cause the machine 200 to perform any one or more of the methods discussed herein. For example, the instructions 202 can cause the machine 200 to perform any one or more of the methods described herein. The instructions 202 transform a general, unprogrammed machine 200 into a particular machine 200 programmed to perform the described and illustrated functions in the described manner. The machine 200 can operate as a stand-alone device or can be coupled (e.g., networked) to other machines. In a networked deployment, the machine 200 can operate in the capacity of a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 200 can include, but is not limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a personal digital assistant (PDA), an entertainment media system, a cellular telephone, a smartphone, a mobile device, a wearable device (e.g., a smartwatch), a smart home device (e.g., a smart appliance), other smart devices, a web device, a network router, a network switch, a network bridge, or any machine capable of sequentially or otherwise executing the instructions 202 specifying actions to be taken by the machine 200. Further, although a single machine 200 is shown, the term "machine" shall also be taken to include a collection of machines that individually or jointly execute the instructions 202 to perform any one or more of the methods discussed herein. For example, the machine 200 can include a computing system 502 or any one of a plurality of server devices forming part of an interactive server system 510. In some examples, the machine 200 can also include both a client system and a server system, where certain operations of a particular method or algorithm are executed on the server side and certain operations of a particular method or algorithm are executed on the client side.

[0045] The machine 200 can include a processor 204, a memory 206, and input / output (I / O) components 208 that can be configured to communicate with each other via a bus 210. In an example, the processor 204 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) can include, for example, a processor 212 and a processor 214 that execute the instructions 202. The term "processor" is intended to include multi-core processors that can include two or more independent processors (sometimes referred to as "cores") that can execute instructions simultaneously. AlthoughFigure 2 Multiple processors 204 are shown, but machine 200 can include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.

[0046] Memory 206 includes main memory 216, static memory 240, and storage unit 218, all of which are accessible by processor 204 via bus 210. Main memory 206, static memory 240, and storage unit 218 store instructions 202 embodying any one or more of the methods or functions described herein. The instructions 202 may also reside, fully or partially, within main memory 216, within static memory 240, within machine-readable medium 220 within storage unit 218, within at least one of the processors within processor 204 (e.g., within a cache memory of the processor), or within any suitable combination thereof during execution by machine 200.

[0047] I / O components 208 can include various components for receiving input, providing output, generating output, sending information, exchanging information, capturing measurements, etc. The specific I / O components 208 included in a particular machine will depend on the type of the machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, while a headless server machine will likely not include such a touch input device. It should be understood that I / O components 208 can include Figure 2 many other components not shown. In various examples, I / O components 208 can include user output components 222 and user input components 224. User output components 222 can include visual components (e.g., a display such as a plasma display panel (PDP), a light-emitting diode (LED) display, a liquid crystal display (LCD), a projector, or a cathode ray tube (CRT)), acoustic components (e.g., speakers), haptic components (e.g., a vibration motor, a resistance mechanism), other signal generators, etc. User input components 224 can include alphanumeric input components (e.g., a keyboard, a touch screen configured to receive alphanumeric input, an optical keyboard, or other alphanumeric input components), point-based input components (e.g., a mouse, a touchpad, a trackball, a joystick, a motion sensor, or other pointing instruments), haptic input components (e.g., a physical button, a touch screen that provides the location and force of a touch or touch gesture, or other haptic input components), audio input components (e.g., a microphone), etc.

[0048] In additional examples, I / O component 208 may include a biometric component 226, a motion component 228, an environmental component 230, or a location component 232, as well as various other components. For example, biometric component 226 includes components for detecting expressions (e.g., gestures, facial expressions, vocal expressions, body postures, or eye tracking), measuring biometric signals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), identifying people (e.g., voice recognition, retina recognition, facial recognition, fingerprint recognition, or electroencephalogram-based recognition), and so on. Motion component 228 includes: an acceleration sensor component (e.g., an accelerometer), a gravity sensor component, a rotation sensor component (e.g., a gyroscope).

[0049] Environmental component 230 includes, for example, one or more camera devices (with still image / photo and video capabilities), a lighting sensor component (e.g., a photometer), a temperature sensor component (e.g., one or more thermometers for detecting ambient temperature), a humidity sensor component, a pressure sensor component (e.g., a barometer), an acoustic sensor component (e.g., one or more microphones for detecting background noise), a proximity sensor component (e.g., an infrared sensor for detecting nearby objects), a gas sensor (e.g., a gas detection sensor for detecting the concentration of hazardous gases for safety or for measuring pollutants in the atmosphere), a depth or distance sensor (e.g., a sensor for determining the distance to an object or for determining the depth of object features in a 3D coordinate system), or other components that can provide an indication, measurement, or signal corresponding to the surrounding physical environment.

[0050] Regarding the camera device, computing system 502 may have a camera device system that includes, for example, a front camera on the front surface of computing system 502 and a rear camera on the rear surface of computing system 502. The front camera may be used, for example, to capture still images and videos of the user of computing system 502 (e.g., "selfies"), which can then be enhanced with the enhanced data (e.g., filters) described above. The rear camera may be used, for example, to capture still images and videos in a more conventional camera mode, where these images are similarly enhanced with the enhanced data. In addition to the front camera and the rear camera, computing system 502 may also include a 360° camera device for capturing 360° photos and videos.

[0051] Furthermore, the camera device system of computing system 502 may include a dual rear camera (e.g., a main camera and a depth sensing camera), or even a triple, quadruple, or quintuple rear camera configuration on the front and rear sides of computing system 502. For example, these multiple camera device systems may include a wide-angle camera, an ultra-wide-angle camera, a telephoto camera, a macro camera, and a depth sensor.

[0052] The location component 232 includes a positioning sensor component (e.g., a GPS receiver component), an altitude sensor component (e.g., an altimeter or barometer that detects air pressure, from which altitude can be obtained), a direction sensor component (e.g., a magnetometer), etc.

[0053] A variety of techniques can be used to implement communication. The I / O component 208 also includes a communication component 234 that is operable to couple the machine 200 to the network 236 or the device 238 via corresponding couplings or connections. For example, the communication component 234 can include a network interface component or other suitable device that interfaces with the network 236. In another example, the communication component 234 can include a wired communication component, a wireless communication component, a cellular communication component, a near field communication (NFC) component, components (e.g., low power consumption), components, and other communication components for providing communication via other modalities. The device 238 can be another machine or any of a variety of peripheral devices (e.g., a peripheral device coupled via USB).

[0054] In addition, the communication component 234 can detect identifiers or include components operable to detect identifiers. For example, the communication component 234 can include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional barcodes such as universal product code (UPC) barcodes, multi-dimensional barcodes such as quick response (QR) codes, Aztec codes, data matrix, data glyph, MaxiCode, PDF417, UltraCode, UCC RSS-2D barcodes, and other optical codes), or an acoustic detection component (e.g., a microphone for identifying an audio signal of a tag). In addition, various information can be derived via the communication component 234, such as a location obtained via Internet protocol (IP) geolocation, a location obtained via signal triangulation, a location obtained via detecting an NFC beacon signal that can indicate a specific location, etc.

[0055] Various memories (e.g., main memory 216, static memory 240, and the memory of the processor 204) and the storage unit 218 can store one or more sets of instructions and data structures (e.g., software) embodied or used by any one or more of the methods or functions described herein. These instructions (e.g., instructions 202), when executed by the processor 204, cause various operations to implement the disclosed examples.

[0056] Instructions 202 can be sent or received via a network interface device (e.g., the network interface component included in communication component 234) using a transmission medium and any one of a number of well-known transmission protocols (e.g., Hypertext Transfer Protocol (HTTP)) over network 236. Similarly, instructions 202 can be sent or received using a transmission medium via a coupling (e.g., a peer-to-peer coupling) to device 238.

[0057] Figure 3A A collaboration diagram of components of an XR system using hand tracking for user input is shown, according to some examples. Figure 3B A process flow diagram of a method using a pinch hand gesture or a pinch pose as user input is shown, and Figure 3C 、 Figure 3D and Figure 3E using a pinch hand gesture or a pinch pose to rescale a virtual object is shown.

[0058] Although Figure 3B the method of rescaling virtual object 300 depicts a particular order of operations, the order can be changed without departing from the scope of the present disclosure. For example, some of the depicted operations can be performed in parallel, in a different order, or by different components of the XR system, which will not substantially affect the functionality of the method.

[0059] The method of rescaling virtual object 300 is used by an XR system (e.g., the head-mounted device 100 of Figure 1A ) to provide a continuous real-time input modality to a user of the XR system, where the user interacts with the XR user interface 346 using hand postures or hand gestures. The AR application can be a utility application such as an interactive game, a maintenance guide, an interactive map, an interactive tour guide, a tutorial, etc. The AR application can also be an entertainment application such as a video game, an interactive video, etc.

[0060] In operation 302, the XR system 332 generates an XR user interface 346 provided to the user 326. For example, the user interface engine 324 includes XR user interface control logic 372, which includes a dialogue script specifying a user interface dialogue implemented by the XR user interface 346, etc. The XR user interface control logic 372 also includes one or more actions taken by the XR system based on various dialogue events detected, such as user input. The user interface engine 324 also includes an XR user interface object model 370. The XR user interface object model 370 includes 3D coordinate data of one or more virtual objects (such as the virtual object 354), and 3D coordinate data of one or more pinch cuboid colliders associated with the virtual object 354 (such as the virtual object pinch collider 356). The virtual object pinch collider 356 is a virtual object with which the user 326 interacts to adjust parameters of the virtual object 354, such as but not limited to the scale of the virtual object 354. The XR user interface object model 370 also includes 3D graphic data of the virtual object 354 and the virtual object pinch collider 356. The optical engine 344 uses this 3D graphic data to generate the XR user interface 346 for display to the user 326.

[0061] In some examples, the shape of the virtual object pinch collider 356 may include one or more irregular or regular 3D geometric shapes, such as but not limited to regular or irregular polyhedrons, ellipsoids, conical shapes, etc. In some examples, the shape of the virtual object pinch collider 356 is based on one or more 3D geometric features of the virtual object 354.

[0062] The user interface engine 324 generates XR user interface graphic data 334 based on the XR user interface object model 370. The XR user interface graphic data 334 includes image video data of one or more virtual objects of the XR user interface 346. The user interface engine 324 transmits the XR user interface graphic data 334 to the image display driver 336 of the optical engine 344 of the XR system 332. The image display driver 336 receives the XR user interface graphic data 334 and generates a display control signal 338 based on the XR user interface graphic data 334. The image display driver 336 uses the display control signal 338 to control the operation of one or more optical components 320 of the optical engine 344. In response to the display control signal 338, one or more optical components 320 generate a visible image of the XR user interface 346 provided to the user 326.

[0063] In operation 304, the XR system 332 detects a pinching hand gesture or pose made by the user 326. For example, the XR system 332 uses one or more hand-tracking cameras 348 to capture tracking video frame data 350 of a hand gesture 340 or pose being made by the user 326 using one or more hands 352 in the user's hand. The hand-tracking camera 348 transmits the tracking video frame data 350 to the hand-tracking component 322 of the hand-tracking pipeline 342 of the XR system 332.

[0064] The hand-tracking component 322 receives the tracking video frame data 350 and generates hand-tracking data 330 based on the tracking video frame data 350. The hand-tracking data 330 includes bone model data in a 3D coordinate system of one or more bone models (such as bone model 358) of one or more hands 352 of the user based on landmark features extracted from the tracking video frame data 350, and hand gesture classification data of the hand gesture 340 or pose being made by one or more hands 352 of the user. The bone model includes bone model features, such as nodes 374, which correspond to visual landmarks of the respective parts of one or more hands 352 of the identified user 326. In some examples, the hand-tracking data 330 includes classification information of one or more landmarks associated with one or more hands 352 of the user 326, links (such as link 376) between the joints of the user's fingers, the physical positions of the landmarks, and landmark data (such as landmark identifiers). In some examples, the hand gesture classification data includes an indication of a pinching hand gesture or pose being made by one or more hands 352 in the user's hand.

[0065] For example, the hand-tracking component 322 identifies landmark features on the respective parts of one or more hands 352 of the user 326 captured in the tracking video frame data 350. In some examples, the hand-tracking component 322 uses computer vision methods to extract landmarks of one or more hands 352 of the user 326 from the tracking video frame data 350, and the computer vision methods include but are not limited to Harris corner detection, Shi-Tomasi corner detection, scale-invariant feature transform (SIFT), speeded-up robust features (SURF), features from accelerated segment test (FAST), oriented FAST and rotated BRIEF (ORB), etc.

[0066] In some examples, the hand tracking component 322 generates hand pose classification data and a series of bone models of the hand tracking data 330 based on landmarks extracted from the tracking video frame data 350 using artificial intelligence methods and an ML hand tracking model 328 previously generated using machine learning methods. In some examples, the ML hand tracking model 328 includes, but is not limited to, neural networks, learning vector quantization networks, logistic regression models, support vector machines, random decision forests, naive Bayes models, linear discriminant analysis models, and k-nearest neighbor models. In some examples, the machine learning methods for generating the ML hand tracking model 328 may include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, dimensionality reduction, self-learning, feature learning, sparse dictionary learning, and anomaly detection.

[0067] In some examples, the hand tracking component 322 generates hand pose classification data and a series of bone models of the hand tracking data 330 based on landmarks extracted from the tracking video frame data 350 using geometric methods.

[0068] In some examples, a pinching hand pose or a pinching gesture includes a user touching the tip of the index finger of one hand in their hand to the tip of the thumb of the same hand.

[0069] In some examples, a pinching hand pose or a pinching gesture is determined by detecting a prototype hand pose or a series of prototype gestures as the user moves the tip of the index finger of one hand in their hand closer to the tip of the thumb of the same hand.

[0070] The hand tracking component 322 of the XR system 332 transmits the hand tracking data 330 to the user interface engine 324. The user interface engine 324 detects that the user 326 is making a pinching hand pose or a pinching gesture based on the hand pose classification data of the hand tracking data 330.

[0071] In response to detecting that user 326 is making a pinching hand gesture or pinching pose, user interface engine 324 uses collision detector 368 to determine whether a pinching hand gesture or pinching pose is being made within virtual object pinching collider 356 of virtual object 354. The determination is made by collision detector 368 based on the skeletal model data of hand tracking data 330 and the 3D coordinate data of virtual object pinching collider 356 included in XR user interface object model 370. For example, user interface engine 324 determines the 3D coordinate data of the index finger tip node and thumb tip node of skeletal model 358. User interface engine 324 generates a pinching position collider 364 based on the 3D coordinate data of the index finger tip node and thumb tip node. Collision detector 368 determines whether the 3D geometry of virtual object pinching collider 356 collides with the 3D geometry of pinching position collider 364 by detecting the intersection between the 3D geometry of virtual object pinching collider 356 and the 3D geometry of pinching position collider 364. In the case of a point collider, if the point collider is located within the volume of another collider, a collision is determined to have occurred.

[0072] In some examples, the shape of pinching position collider 364 may include one or more irregular or regular 3D geometries, such as but not limited to regular or irregular polyhedrons, spheres, conical shapes, etc. In some examples, the shape of pinching position collider 364 is based on one or more characteristics of user's hand 352.

[0073] If collision detector 368 determines that pinching position collider 364 of the pinching hand gesture or pinching pose collides with virtual object pinching collider 356, user interface engine 324 also determines the 3D coordinates of pinching position 380 while user 326 is making the pinching hand gesture or pinching pose. In some examples, the 3D coordinates of pinching position 380 are determined based on the center point of pinching position collider 364. For example, the XR system determines the 3D coordinate data of pinching position 380 based on the 3D coordinate data of the center point of pinching position collider 364. In some examples, the 3D coordinates of pinching position 380 are determined based on the 3D coordinate data of the edge, vertex, or surface of pinching position collider 364. In some examples, the 3D coordinate data of pinching position 380 is determined based on the intersection of pinching position collider 364 and virtual object pinching collider 356.

[0074] In operation 306, the XR system determines pinching vector 362 based on the 3D coordinate data of virtual object center point 378 of virtual object 354 and the 3D coordinate data of pinching position 380. The resulting pinching vector 362 extends from virtual object center point 378 to pinching position 380.

[0075] In operation 308, the XR system determines a virtual object vector 360 based on the 3D coordinate data of the virtual object center point 378 and the 3D coordinate data of the virtual object collision body center point 382 of the virtual object pinch collision body 356. The resulting virtual object vector 360 extends from the virtual object center point 378 to the virtual object collision body center point 382.

[0076] In operation 310, the XR system 332 determines a scaling vector 366 based on the pinch vector 362 and the virtual object vector 360. For example, the XR system projects the pinch vector 362 onto the virtual object vector 360, as indicated by the projection 384 in Figure 3D .

[0077] In operation 312, the XR system 332 adjusts the scale of the virtual object 354 based on the virtual object vector 360 and the scaling vector 366. For example, the XR system 332 adjusts the scale of the object based on the ratio of the length of the scaling vector 366 divided by the length of the virtual object vector 360. For example, if the length of the scaling vector 366 is greater than the length of the virtual object vector 360, the scale of the virtual object 354 will be increased. If the length of the scaling vector 366 is less than the length of the virtual object vector 360, the scale of the virtual object 354 will be decreased.

[0078] In operation 314, the XR system 332 uses the optical engine 344 to redisplay the virtual object 354 based on the adjusted scale of the virtual object 354, as shown in Figure 3E .

[0079] In operation 316, the XR system 332 determines whether the user 326 is still maintaining a pinching hand gesture or a pinching pose. For example, the XR system 332 captures additional tracked video frame data 350 of the hand 352 of the user 326. The XR system generates additional hand tracking data 330 based on the additional tracked video frame data 350 as described herein. Included in the additional hand tracking data 330 is additional hand pose or hand gesture classification data. The user interface engine 324 detects whether the user is still maintaining a pinching pose or a pinching hand gesture based on the additional hand pose or hand gesture classification data of the hand tracking data 330.

[0080] In some examples, the release of the pinching hand gesture or the pinching pose includes the user moving the fingertip of the index finger of one of their hands away from the fingertip of the thumb of the same hand.

[0081] In some examples, the release of the pinching hand gesture or the pinching pose is determined by detecting a prototype hand gesture or a series of prototype poses when the user moves the fingertip of the index finger of one of their hands away from the fingertip of the thumb of the same hand.

[0082] In response to determining that the pinching hand gesture or pinching pose is still being held, the XR system transitions 386 to operation 304 and repeats the process of scaling the virtual object 354 as described herein. In response to determining that the pinching hand gesture or pinching pose is not being held, the XR system transitions 388 to operation 318 and ends.

[0083] In some examples, the XR system utilizes various APIs and system libraries to perform the functions of the hand tracking pipeline 342, the user interface engine 324, and the optical engine 344.

[0084] System with a head-mounted device

[0085] Figure 4 Illustrates a system 400 including a head-mounted device 100 with a selector input device according to some examples. Figure 4 Is a high-level functional block diagram of an example head-mounted device 100 communicatively coupled to a mobile device 414 and various server systems 404 (e.g., an interaction server system 510) via various networks 508.

[0086] The head-mounted device 100 includes one or more imaging devices, and each of the one or more imaging devices can be, for example, one or more imaging devices 408, a light emitter 410, and one or more wide-spectrum imaging devices 412.

[0087] The mobile device 414 connects to the head-mounted device 100 using both a low-power wireless connection 416 and a high-speed wireless connection 418. The mobile device 414 is also connected to the server system 404 and the network 406.

[0088] The head-mounted device 100 also includes two image displays in the image display 420 of the optical assembly. The two image displays 420 of the optical assembly include one image display associated with the left lateral side of the head-mounted device 100 and one image display associated with the right lateral side of the head-mounted device 100. The head-mounted device 100 also includes an image display driver 422 and a GPU 424. The image display 420 of the optical assembly, the image display driver 422, and the GPU 424 constitute the optical engine of the head-mounted device 100. The image display 420 of the optical assembly is used to present images and videos to the user of the head-mounted device 100, including images that may include a graphical user interface.

[0089] The image display driver 422 commands and controls the image display 420 of the optical component. The image display driver 422 can deliver the image data directly to the image display 420 of the optical component for presentation or can convert the image data into a signal or data format suitable for delivery to an image display device. For example, the image data can be video data formatted according to a compression format such as H.264 (MPEG-4 Part 10), HEVC, Theora, Dirac, RealVideo RV40, VP8, VP9, etc., and the still image data can be formatted according to a compression format such as Portable Network Graphics (PNG), Joint Photographic Experts Group (JPEG), Tagged Image File Format (TIFF), or Exchangeable Image File Format (EXIF), etc.

[0090] The head-wearable device 100 includes a frame and rods (or temples) extending from the lateral sides of the frame. The head-wearable device 100 also includes a user input device 430 (e.g., a touch sensor or a push button) including an input surface on the head-wearable device 100. The user input device 430 (e.g., a touch sensor or a push button) is used to receive an input selection from a user to manipulate a graphical user interface of the presented image.

[0091] Figure 4 The components shown for the head-wearable device 100 are located on one or more circuit boards such as a PCB or a flexible PCB in the rim or the temple. Alternatively or additionally, the depicted components can be located in the block, the frame, the hinge, or the nose bridge of the head-wearable device 100. The left camera device 408 and the right camera device 408 can include digital camera device elements such as complementary metal oxide semiconductor (CMOS) image sensors, charge-coupled devices, camera device lenses, or any other corresponding visible light or light-capturing elements that can be used to capture data including images of scenes with unknown objects.

[0092] The head-wearable device 100 includes a memory 402 that stores instructions for performing a subset or all of the functions described herein. The memory 402 can also include a storage device.

[0093] As Figure 4As shown, the high-speed circuit system 428 includes a high-speed processor 432, a memory 402, and a high-speed wireless circuit system 434. In some examples, the image display driver 422 is coupled to the high-speed circuit system 428 and is operated by the high-speed processor 432 to drive the left and right image displays of the image display 420 of the optical component. The high-speed processor 432 can be any processor capable of managing the high-speed communication and operation of any general computing system required for the head-worn device 100. The high-speed processor 432 includes the processing resources required to manage high-speed data transmission on the high-speed wireless connection 418 to a wireless local area network (WLAN) using the high-speed wireless circuit system 434. In some examples, the high-speed processor 432 executes an operating system of the head-worn device 100, such as the LINUX operating system or other such operating systems, and the operating system is stored in the memory 402 for execution. In addition to any other duties, the high-speed processor 432 that executes the software architecture for the head-worn device 100 is used to manage data transmission with the high-speed wireless circuit system 434. In certain examples, the high-speed wireless circuit system 434 is configured to implement the Institute of Electrical and Electronics Engineers (IEEE) 802.11 communication standard, which is also referred to herein as WiFi. In some examples, the high-speed wireless circuit system 434 can implement other high-speed communication standards.

[0094] The low-power wireless circuit system 436 and the high-speed wireless circuit system 434 of the head-worn device 100 can include a short-range transceiver (Bluetooth TM ) and a wireless wide area network, local area network, or wide area network transceiver (e.g., cellular or WiFi). The mobile device 414 - including transceivers that communicate via the low-power wireless connection 416 and the high-speed wireless connection 418 - can be implemented using the details of the architecture of the head-worn device 100 or using other elements of the network 406.

[0095] The memory 402 includes any storage device capable of storing various data and applications, and in addition includes camera device data generated by the left camera device 408 and the right camera device 408, the wide spectral band camera device 412, and the GPU 424, and images generated by the image display driver 422 for display on the image display of the image display 420 of the optical component. Although the memory 402 is shown as integrated with the high-speed circuitry 428, in some examples, the memory 402 may be a separate stand-alone component of the head-mounted device 100. In some such examples, electrical wiring may provide a connection from the GPU 424 or the low-power processor 438 to the memory 402 through a chip including the high-speed processor 432. In some examples, the high-speed processor 432 may manage the addressing of the memory 402 such that the low-power processor 438 will initiate the high-speed processor 432 whenever a read or write operation involving the memory 402 is required.

[0096] As Figure 4 shown, the low-power processor 438 or the high-speed processor 432 of the head-mounted device 100 may be coupled to a camera device (the camera device 408, the light emitter 410, or the wide spectral band camera device 412), the image display driver 422, the user input device 430 (e.g., a touch sensor or a push button), and the memory 402.

[0097] The head-mounted device 100 is connected to a host computer. For example, the head-mounted device 100 is paired with the mobile device 414 via the high-speed wireless connection 418 or connected to the server system 404 via the network 406. The server system 404 may be one or more computing devices that are part of a service or network computing system, e.g., including a processor, a memory, and a network communication interface to communicate with the mobile device 414 and the head-mounted device 100 via the network 406.

[0098] The mobile device 414 includes a processor and a network communication interface coupled to the processor. The network communication interface allows communication via the network 406, the low-power wireless connection 416, or the high-speed wireless connection 418. The mobile device 414 may also store at least part of the instructions for generating the stereophonic audio content in the memory of the mobile device 414 to implement the functions described herein.

[0099] The output components of the head-mounted device 100 include visual components, such as displays such as liquid crystal displays (LCDs), plasma display panels (PDPs), light-emitting diode (LED) displays, projectors, or waveguides. The image display of the optical component is driven by the image display driver 422. The output components of the head-mounted device 100 also include acoustic components (e.g., speakers), tactile components (e.g., vibration motors), other signal generators, and the like. The input components of the head-mounted device 100, the mobile device 414, and the server system 404, such as the user input device 430, may include alphanumeric input components (e.g., keyboards, touchscreens configured to receive alphanumeric input, optical keyboards, or other alphanumeric input components), point-based input components (e.g., mice, touchpads, trackballs, joysticks, motion sensors, or other pointing instruments), tactile input components (e.g., physical buttons, touchscreens that provide the location and force of a touch or touch gesture, or other tactile input components), audio input components (e.g., microphones), and the like.

[0100] The head-mounted device 100 may also include additional peripheral device elements. Such peripheral device elements may include biometric sensors, additional sensors, or display elements integrated with the head-mounted device 100. For example, the peripheral device elements may include any I / O components, which include output components, motion components, position components, or any other such elements described herein.

[0101] For example, biometric components include components for detecting expressions (e.g., gestures, facial expressions, vocal expressions, body postures, or eye tracking), measuring biometric signals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), identifying people (e.g., voice recognition, retina recognition, facial recognition, fingerprint recognition, or electroencephalogram-based recognition), and the like. Motion components include acceleration sensor components (e.g., accelerometers), gravity sensor components, rotational sensor components (e.g., gyroscopes), and the like. Position components include positioning sensor components for generating positioning coordinates (e.g., global positioning system (GPS) receiver components), Wi-Fi or Bluetooth TM transceivers, altitude sensor components (e.g., altimeters or barometers that detect air pressure, from which altitude can be obtained), direction sensor components (e.g., magnetometers), and the like. Such positioning system coordinates may also be received from the mobile device 414 via the low-power wireless circuitry 436 or the high-speed wireless circuitry 434 through the low-power wireless connection 416 and the high-speed wireless connection 418.

[0102] Networked computing environment

[0103] Figure 5is a block diagram showing an example interaction system 500 for facilitating interactions over a network (e.g., exchanging text messages, making text, audio, and video calls, or playing games). The interaction system 500 includes a plurality of XR systems 502, each of which hosts a plurality of applications, including interaction clients 504 and other applications 506. Each interaction client 504 is communicatively coupled via one or more communication networks including a network 508 (e.g., the Internet) to other instances of the interaction client 504 (e.g., hosted on corresponding other XR systems 502), an interaction server system 510, and a third-party server 512. The interaction client 504 can also communicate with the locally hosted applications 506 using an application programming interface (API).

[0104] Each computing system 502 can include one or more user devices, such as mobile devices 414, head-wearable devices 100, and computer client devices 514, which are communicatively connected to exchange data and messages.

[0105] The interaction client 504 interacts via the network 508 with other interaction clients 504 and with the interaction server system 510. The data exchanged between the interaction clients 504 (e.g., interaction 516) and between the interaction client 504 and the interaction server system 510 includes functionality (e.g., commands for activating functionality) and payload data (e.g., text, audio, video, or other multimedia data).

[0106] The interaction server system 510 provides server-side functionality to the interaction client 504 via the network 508. While certain functions of the interaction system 500 are described herein as being performed by the interaction client 504 or by the interaction server system 510, whether certain functions are located within the interaction client 504 or within the interaction server system 510 can be a design choice. For example, it may be technically preferable to initially deploy a particular technology and functionality within the interaction server system 510, but later migrate the technology and functionality to the interaction client 504 where the computing system 502 has sufficient processing power.

[0107] The interaction server system 510 supports various services and operations provided to the interaction client 504. Such operations include sending data to the interaction client 504, receiving data from the interaction client 504, and processing data generated by the interaction client 504. The data can include message content, client device information, geolocation information, media enhancements and overlays, message content persistence conditions, social network information, and live event information. Data exchange within the interaction system 500 is activated and controlled by functions available via the user interface (UI) of the interaction client 504.

[0108] Now turning specifically to the interaction server system 510, an application programming interface (API) server 518 is coupled to and provides a programming interface for an interaction server 520, making the functionality of the interaction server 520 accessible to interaction clients 504, other applications 506, and third-party servers 512. The interaction server 520 is communicatively coupled to a database server 522 to facilitate access to a database 524 that stores data associated with interactions processed by the interaction server 520. Similarly, a web server 526 is coupled to the interaction server 520 and provides a web-based interface to the interaction server 520. To this end, the web server 526 processes incoming network requests via the Hypertext Transfer Protocol (HTTP) and several other related protocols.

[0109] The application programming interface (API) server 518 receives and sends interaction data (e.g., commands and message payloads) between the interaction server 520 and the XR system 502 (and, for example, interaction clients 504 and other applications 506) and third-party servers 512. Specifically, the application programming interface (API) server 518 provides a set of interfaces (e.g., routines and protocols) that interaction clients 504 and other applications 506 can call or query to activate the functionality of the interaction server 520. The application programming interface (API) server 518 exposes various functions supported by the interaction server 520, including account registration; login functionality; sending interaction data from a particular interaction client 504 to another interaction client 504 via the interaction server 520; transferring media files (e.g., images or videos) from an interaction client 504 to the interaction server 520; setting a collection of media data (e.g., a story); retrieving a list of friends of a user of the computing system 502; retrieving messages and content; adding and deleting entities (e.g., friends) for an entity graph (e.g., a social graph); locating friends within a social graph; and opening application events (e.g., related to an interaction client 504).

[0110] The interaction server 520 hosts multiple systems and subsystems, which are described below with reference to Figure 7 as follows.

[0111] Linked applications

[0112] Returning to the interactive client 504, the features and functionality of external resources (e.g., linked application 506 or applet) are made available to the user via the interface of the interactive client 504. In this context, "external" refers to the fact that the application 506 or applet is external to the interactive client 504. Although external resources are typically provided by a third party, they can also be provided by the creator or provider of the interactive client 504. The interactive client 504 receives a user selection of an option for initiating or accessing the features of such an external resource. The external resource can be an application 506 installed on the computing system 502 (e.g., a "native app"), or a scaled-down version of an application hosted on or remote to the computing system 502 (e.g., on a third-party server 512) (e.g., an "applet"). The scaled-down version of the application includes a subset of the features and functionality of the application (e.g., the full-scale, native version of the application) and is implemented using a markup language document. In some examples, the scaled-down version of the application (e.g., an "applet") is a web-based markup language version of the application and is embedded within the interactive client 504. In addition to using a markup language document (e.g., a.*ml file), the applet can include a scripting language (e.g., a.*js file or a.json file) and a style sheet (e.g., a.*ss file).

[0113] In response to receiving a user selection of an option for initiating or accessing the features of an external resource, the interactive client 504 determines whether the selected external resource is a web-based external resource or a locally installed application 506. In some cases, an application 506 locally installed on the computing system 502 can be launched independently of and separately from the interactive client 504, e.g., by selecting an icon corresponding to the application 506 on the home screen of the computing system 502. A scaled-down version of such an application can be launched or accessed via the interactive client 504, and in some examples, no part or only a limited part of the scaled-down application can be accessed outside of the interactive client 504. The scaled-down application can be launched by receiving, e.g., a markup language document associated with the scaled-down application from a third-party server 512 via the interactive client 504 and processing such a document.

[0114] In response to determining that the external resource is a locally installed application 506, the interactive client 504 instructs the computing system 502 to launch the external resource by executing locally stored code corresponding to the external resource. In response to determining that the external resource is a web-based resource, the interactive client 504 communicates with, e.g., a third-party server 512 to obtain a markup language document corresponding to the selected external resource. The interactive client 504 then processes the obtained markup language document to render the web-based external resource within the user interface of the interactive client 504.

[0115] The interactive client 504 can notify a user of the computing system 502 or other users related to such a user (e.g., "friends") of an activity occurring in one or more external resources. For example, the interactive client 504 can provide a notification to participants in a conversation (e.g., a chat session) in the interactive client 504 regarding the current or recent use of an external resource by one or more members of a group of users. One or more users can be invited to join an active external resource or to initiate (within the group of friends) an external resource that was recently used but is currently inactive. The external resource can provide the ability to share items, conditions, statuses, or locations within the external resource with one or more members of a group of users in a chat session to the respective participants in the conversation who use the corresponding interactive client 504. The shared item can be an interactive chat card that chat members can use to interact, e.g., to initiate the corresponding external resource, view specific information within the external resource, or take the chat members to a specific location or status within the external resource. Within a given external resource, a response message can be sent to a user on the interactive client 504. The external resource can selectively include different media items in the response based on the current context of the external resource.

[0116] The interactive client 504 can present a list of available external resources (e.g., applications 506 or applets) to initiate or access a given external resource. The list can be presented in the form of a context-sensitive menu. For example, the icons representing different applications (or applets) of the application 506 (or applets) can vary based on how the menu is initiated by the user (e.g., from a conversation interface or from a non-conversation interface).

[0117] Data Architecture

[0118] Figure 6 is a schematic diagram showing a data structure 600 that can be stored in a database 604 of the interactive server system 510 according to certain examples. Although the contents of the database 604 are shown as including multiple tables, it will be appreciated that data can be stored in other types of data structures (e.g., as an object-oriented database).

[0119] The database 604 includes message data stored within a message table 606. For any particular message, the message data includes at least message sender data, message recipient (or receiver) data, and a payload. Further details regarding information that can be included in a message and that is included within the message data stored in the message table 606 are described below with reference to Figure 6 describe additional details regarding the information that can be included in a message and that is included within the message data stored in the message table 606.

[0120] The entity table 608 stores entity data and is linked (e.g., to a reference ground) to the entity graph 610 and the profile data 602. Entities whose records are maintained within the entity table 608 can include individuals, corporate entities, organizations, objects, locations, events, etc. Any entity for which the interaction server system 510 stores data about it can be an identified entity, regardless of the entity type. Each entity is set with a unique identifier and an entity type identifier (not shown).

[0121] The entity graph 610 stores information about the relationships and associations between entities. For example, such relationships can be social relationships based on interests or activities, professional relationships (e.g., working in the same company or organization). Some relationships between entities can be one-way, such as a personal user's subscription to digital content (e.g., a newspaper or other digital media channel or brand) of a business or publishing user. Other relationships can be two-way, such as the "friend" relationship between individual users of the interaction system 500.

[0122] Certain permissions and relationships can be attached to each relationship and also to each direction of the relationship. For example, a two-way relationship (e.g., the friend relationship between individual users) can include authorization for the public disclosure of digital content items between individual users, but certain restrictions or filters (e.g., based on content characteristics, location data, or time-of-day data) can be imposed on the public disclosure of these digital content items. Similarly, the subscription relationship between a personal user and a business user can impose different degrees of restrictions on the disclosure of digital content from the business user to the personal user, and can greatly limit or prevent the disclosure of digital content from the personal user to the business user. As an example of an entity, a particular user can record certain restrictions (e.g., in the form of privacy settings) in the record of that entity within the entity table 608. Such privacy settings can apply to all types of relationships within the context of the interaction system 500, or can be selectively applied to only certain types of relationships.

[0123] The profile data 602 stores various types of profile data about a particular entity. Based on the privacy settings specified by the particular entity, the profile data 602 can be selectively used and presented to other users of the interaction system 500. In the case where the entity is an individual, the profile data 602 includes, for example, a username, a phone number, an address, settings (e.g., notification and privacy settings), and a set of avatar representations (or a collection of such avatar representations) selected by the user. Then, a particular user can selectively include one or more of these avatar representations within the content of messages transmitted via the interaction system 500 and on the map interface displayed by the interaction client 504 to other users. The set of avatar representations can include "status avatars" that present graphical representations of the status or activities that the user can choose to convey at a particular time.

[0124] In the case where the entity is a group, in addition to the group name, members, and various settings for the associated group (e.g., notifications), the profile data 602 for the group can similarly include one or more avatar representations associated with the group.

[0125] The database 604 also stores enhancement data, such as overlays or filters, in an enhancements table 612. The enhancement data is associated with videos (the data of which is stored in a videos table 614) and images (the data of which is stored in an images table 616) and is applied to the videos and images.

[0126] In some examples, a filter is an overlay that is displayed as an overlay on an image or video during presentation to a message recipient. Filters can be of various types, including a filter selected by a user from a set of filters presented to the message sender by the interactive client 504 when the message sender is composing a message. Other types of filters include geolocation filters (also known as geo-filters), which can be presented to the message sender based on geolocation. For example, specific geolocation filters for a nearby or special location can be presented by the interactive client 504 within a user interface based on geolocation information determined by a global positioning system (GPS) unit of the computing system 502.

[0127] Another type of filter is a data filter, which can be selectively presented to the message sender by the interactive client 504 based on other input or information collected by the computing system 502 during the message creation process. Examples of data filters include the current temperature at a specific location, the current speed at which the message sender is traveling, the battery life of the computing system 502, or the current time.

[0128] Other enhancement data that can be stored within the images table 616 includes augmented reality content items (e.g., corresponding to an applied Lens or augmented reality experience). The augmented reality content items can be real-time special effects and sounds that can be added to an image or video.

[0129] As described above, augmented data includes AR, VR, and mixed reality (MR) content items, overlays, image transforms, images, and modifications that can be applied to image data (e.g., video or images). This includes real-time modifications, i.e., modifying an image when it is captured using a device sensor (e.g., one or more camera devices) of computing system 502 and then the modified image is displayed on the screen of computing system 502. This also includes modifications to stored content (e.g., video segments in a collection or group that can be modified). For example, in computing system 502 that accesses multiple augmented reality content items, a user can use a single video segment with multiple augmented reality content items to see how different augmented reality content items will modify the stored segment. Similarly, real-time video capture can use modifications to show how the video image currently captured by the sensors of computing system 502 will modify the captured data. Such data can be displayed only on the screen without being stored in memory, or the content captured by the device sensors can be recorded and stored in memory with or without (or both with and without) modifications. In some systems, a preview feature can simultaneously show how different augmented reality content items will look within different windows on the display. For example, this can enable multiple windows with different pseudo-random animations to be viewed simultaneously on the display.

[0130] Thus, systems that use data of augmented reality content items and various other such transformation systems that use this data to modify content can involve: detection of various objects (e.g., faces, hands, bodies, cats, dogs, surfaces, objects, etc.) in a video frame, tracking of such objects as they leave the field of view, enter the field of view, and move around in the field of view, and modification or transformation of such objects when tracking them. In various examples, different methods can be used to implement such transformations. Some examples can involve: generating a three-dimensional mesh model of one or more objects; and using the transformation and animated textures of the model within the video to implement the transformation. In some examples, tracking of points on an object can be used to place an image or texture (which can be two-dimensional or three-dimensional) at the tracking location. In yet another example, neural network analysis of the video frame can be used to place an image, model, or texture in the content (e.g., an image or video frame). Thus, augmented reality content items involve both the images, models, and textures for creating transformations in the content and the additional modeling and analysis information required to implement such transformations using object detection, tracking, and placement.

[0131] Real-time video processing can be performed using any kind of video data (e.g., video streams, video files, etc.) stored in the memory of any kind of computerized system. For example, a user can load video files and save them in the device's memory, or can use the device's sensors to generate video streams. Additionally, computer animation models can be used to process any object, such as a human face and various parts of a human body, an animal, or a non-biological object (e.g., a chair, a car, or other objects).

[0132] In some examples, when a specific modification is selected along with the content to be transformed, the element to be transformed is identified by a computing device and then, if the element to be transformed exists in a frame of the video, it is detected and tracked. The elements of the object are modified according to the modification request, thereby transforming the frame of the video stream. For different kinds of transformations, the frames of the video stream can be transformed by different methods. For example, for frame transformations that mainly involve changing the form of the elements of an object (e.g., using an Active Shape Model (ASM) or other known methods), characteristic points are calculated for each element of the object. Then, a grid based on the characteristic points is generated for each element of the object. This grid is used in subsequent stages to track the elements of the object in the video stream. During the tracking process, the grid for each element is aligned with the position of each element. Then, additional points are generated on the grid.

[0133] In some examples, the transformation of changing some regions of an object using the elements of the object can be performed by calculating the characteristic points for each element of the object and generating a grid based on the calculated characteristic points. Points are generated on the grid, and then various regions are generated based on these points. Then, the elements of the object are tracked by aligning the region of each element with the position of each element in at least one element, and the nature of the region can be modified based on the modification request, thereby transforming the frame of the video stream. Depending on the specific modification request, the nature of the mentioned region can be transformed in different ways. Such modifications can involve: changing the color of the region; removing some parts of the region from the frame of the video stream; including a new object in the region based on the modification request; and modifying or distorting the region or the elements of the object. In various examples, any combination of such modifications or other similar modifications can be used. For certain models to be animated, some characteristic points can be selected as control points for determining the entire state space of the options for model animation.

[0134] In some examples of computer animation models for transforming image data using face detection, a specific face detection algorithm (e.g., Viola-Jones) is used to detect faces in the image. Then, the Active Shape Model (ASM) algorithm is applied to the face region of the image to detect face feature reference points.

[0135] Other suitable methods and algorithms for face detection can be used. For example, in some examples, landmarks are used to locate virtual features, which represent distinguishable points present in most images under consideration. For example, for face landmarks, the location of the left eye pupil can be used. If the initial landmark is not recognizable (e.g., in the case where a person has an eye patch), secondary landmarks can be used. Such a landmark recognition process can be used for any such object. In some examples, a set of landmarks forms a shape. The coordinates of the points in the shape can be used to represent the shape as a vector. One shape is aligned with another shape using a similarity transformation (allowing translation, scaling, and rotation), which minimizes the average Euclidean distance between the shape points. The mean shape is the average of the aligned training shapes.

[0136] The transformation system can capture an image or video stream on a client device (e.g., computing system 502) and perform complex image manipulations locally on the computing system 502 while maintaining an appropriate user experience, computation time, and power consumption. Complex image manipulations can include size and shape changes, emotion transformation (e.g., changing a face from a frown to a smile), state transformation (e.g., aging the subject, reducing the apparent age, changing gender), style transformation, application of graphical elements, and any other suitable image or video manipulations implemented by a convolutional neural network that has been configured to execute efficiently on the computing system 502.

[0137] In some examples, a computer animation model for transforming image data can be used by the following system: In this system, a user can use a computing system 502 having a neural network operating as part of an interactive client 504 operating on the computing system 502 to capture an image or video stream of the user (e.g., a selfie). A transformation system operating within the interactive client 504 determines the presence of a face within the image or video stream and provides a modification icon associated with the computer animation model to transform the image data, or the computer animation model can be present in association with the interfaces described herein. The modification icon includes as part of the modification operation the changes that will be the basis for modifying the user's face within the image or video stream. Once the modification icon is selected, the transformation system initiates a process of transforming the user's image to reflect the selected modification icon (e.g., generating a smiling face on the user). Once the image or video stream is captured and the specified modification is selected, the modified image or video stream can be presented in a graphical user interface displayed on the computing system 502. The transformation system can implement a complex convolutional neural network on a portion of the image or video stream to generate and apply the selected modification. That is, the user can capture an image or video stream, and once the modification icon is selected, the modified result can be presented to the user in real time or near real time. Additionally, while a video stream is being captured, the modification can be persistent, and the selected modification icon remains toggled. Machine-taught neural networks can be used to implement such modifications.

[0138] A graphical user interface presenting the modifications performed by the transformation system can supply additional interaction options to the user. Such options can be based on the interface used to initiate the selection of a particular computer animation model and content capture (e.g., initiated from a content creator user interface). In various examples, after an initial selection of a modification icon, the modification can be persistent. The user can toggle the modification on or off by tapping or otherwise selecting the face modified by the transformation system and store it for later viewing or browsing to other areas of the imaging application. In the case where multiple faces are modified by the transformation system, the user can globally toggle the modification on or off by tapping or selecting an individual face modified and displayed within the graphical user interface. In some examples, an individual face within a group of multiple faces can be modified separately, or such modifications can be toggled individually by tapping or selecting an individual face or a series of individual faces displayed within the graphical user interface.

[0139] The story table 618 stores data regarding a collection of messages and associated image, video, or audio data, where the messages and associated image, video, or audio data are compiled into a collection (e.g., a story or a gallery). The creation of a particular collection can be initiated by a particular user (e.g., each user for whom a record is maintained in the entity table 608). A user can create a "personal story" in the form of a collection of content that has been created and sent / broadcast by that user. To this end, the user interface of the interaction client 504 can include user-selectable icons to enable a message sender to add particular content to his or her personal story.

[0140] The collection can also constitute a "Live Story" that is a collection of content from multiple users, which is created manually, automatically, or using a combination of manual and automatic techniques. For example, a "Live Story" can constitute a curated stream of user-submitted content from various locations and events. Options to contribute content to a particular Live Story can be presented, for example, via the user interface of the interaction client 504 to users whose client devices have location services enabled and are at a co-location event at a particular time. A Live Story can be identified to a user by the interaction client 504 based on the user's location. The end result is a "Live Story" told from a group perspective.

[0141] Another type of content collection is called a "Location Story" that enables users whose computing system 502 is located within a particular geolocation (e.g., on a college or university campus) to contribute to a particular collection. In some examples, contributing to a Location Story may require secondary authentication to verify that the end user belongs to a particular organization or other entity (e.g., is a student on a university campus).

[0142] As mentioned above, the video table 614 stores video data that, in some examples, is associated with messages whose records are maintained in the message table 606. Similarly, the image table 616 stores image data that is associated with messages whose message data is stored in the entity table 608. The entity table 608 can associate various enhancements from the enhancement table 612 with the various images and videos stored in the image table 616 and the video table 614.

[0143] The database 604 also includes social network information collected by the social network system 722.

[0144] System Architecture

[0145] Figure 7is a block diagram showing additional details regarding an interaction system 500 according to some examples. Specifically, the interaction system 500 is shown to include an interaction client 504 and an interaction server 520. The interaction system 500 includes multiple subsystems that are supported on the client side by the interaction client 504 and on the server side by the interaction server 520. Example subsystems are discussed below.

[0146] The image processing system 702 provides various functions that enable a user to capture and enhance (e.g., enhance or otherwise modify or edit) media content associated with a message.

[0147] The camera device system 704 includes control software (e.g., in a camera device application) that interacts with and controls the hardware camera device of the computing system 502 (e.g., directly or via the operating system) to modify and enhance real-time images captured and displayed via the interaction client 504.

[0148] The enhancement system 706 provides functions related to the generation and publication of enhancements (e.g., media overlays) for images captured in real time by the camera device of the computing system 502 or retrieved from the memory of the computing system 502. For example, the enhancement system 706 is operable to select, present, and display media overlays (e.g., image filters or image lenses) for the interaction client 504 for enhancing real-time images received via the camera device system 704 or stored images retrieved from the memory 402 of the computing system 502. These enhancements are selected by the enhancement system 706 based on some inputs and data such as:

[0149] · The geolocation of the computing system 502; and

[0150] · Social network information of the user of the computing system 502.

[0151] Enhancements can include audio and visual content as well as visual effects. Examples of audio and visual content include pictures, text, logos, animations, and sound effects. Examples of visual effects include color overlays. The audio and visual content or visual effects can be applied to media content items (e.g., photos or videos) at the computing system 502 for transmission in a message, or to video content such as a video content stream or feed sent from the interaction client 504. Thus, the image processing system 702 can interact with and support various subsystems of the communication system 708, such as the messaging system 710 and the video communication system 712.

[0152] The media overlay can include text or image data that can be overlaid on a photo taken by the computing system 502 or a video stream produced by the computing system 502. In some examples, the media overlay can be a location overlay (e.g., Venice Beach), a live event name, or a business name overlay (e.g., Beach Café). In additional examples, the image processing system 702 uses the geolocation of the computing system 502 to identify a media overlay that includes the business name at the geolocation of the computing system 502. The media overlay can include other markers associated with the business. The media overlay can be stored in the database 524 and accessed via the database server 522.

[0153] The image processing system 702 provides a user-based publishing platform that enables a user to select a geolocation on a map and upload content associated with the selected geolocation. The user can also specify the circumstances under which a particular media overlay should be provided to other users. The image processing system 702 generates a media overlay that includes the uploaded content and associates the uploaded content with the selected geolocation.

[0154] The augmentation creation system 714 supports an augmented reality developer platform and includes applications for content creators (e.g., artists and developers) to create and publish augmentations (e.g., augmented reality experiences) for the interactive client 504. The augmentation creation system 714 provides content creators with a library of built-in features and tools that includes, for example, custom shaders, tracking techniques, and templates.

[0155] In some examples, the augmentation creation system 714 provides a business-based publishing platform that enables a business to select a specific augmentation associated with a geolocation via an auction process. For example, the augmentation creation system 714 associates the media overlay of the highest bidding business with the corresponding geolocation for a predefined amount of time.

[0156] Communication system 708 is responsible for enabling and handling various forms of communication and interaction within interaction system 500, and includes messaging system 710, audio communication system 716, and video communication system 712. Messaging system 710 is responsible for implementing temporary or time-limited access to content by interaction client 504. Messaging system 710 includes multiple timers (e.g., within short-lived timer system 718), which selectively enable access (e.g., for presentation and display) to messages and associated content via interaction client 504 based on the duration and display parameters associated with a message or a set of messages (e.g., a story). Additional details regarding the operation of short-lived timer system 718 are provided below. Audio communication system 716 enables and supports audio communication (e.g., real-time audio chat) between multiple interaction clients 504. Similarly, video communication system 712 enables and supports video communication (e.g., real-time video chat) between multiple interaction clients 504.

[0157] User management system 720 is operationally responsible for managing user data and profiles, and includes social network system 722, which maintains social network information regarding the relationships between the users of interaction system 500.

[0158] Collection management system 724 is operationally responsible for managing collections or sets of media (e.g., collections of text, image, video, and audio data). Collections of content (e.g., messages, including images, videos, text, and audio) can be organized into "event galleries" or "event stories". Such collections can be made available for a specified period of time (e.g., the duration of the event to which the content pertains). For example, content related to a concert can be made available as a "story" for the duration of the concert. Collection management system 724 may also be responsible for publishing an icon that provides a notification of a particular collection to the user interface of interaction client 504. Collection management system 724 includes curation capabilities that enable a collection manager to manage and curate a particular collection of content. For example, a curation interface enables an event organizer to curate a collection of content related to a particular event (e.g., delete inappropriate content or redundant messages). Additionally, collection management system 724 employs machine vision (or image recognition technology) and content rules to automatically curate content collections. In some examples, compensation may be paid to users for including user-generated content in a collection. In such cases, collection management system 724 operates to automatically pay such users for the use of their content.

[0159] The map system 726 provides various geolocation functions and supports the presentation of map-based media content and messages by the interactive client 504. For example, the map system 726 enables the display on the map of user icons or avatars (e.g., stored in the profile data 602) to indicate the current or past locations of the user's "friends" within the context of the map and the media content (e.g., a collection of messages including photos and videos) generated by these friends. For example, on the map interface of the interactive client 504, messages posted by the user from a specific geolocation to the interactive system 500 can be displayed to the user's "friends" within the context of that specific location on the map. The user can also share his or her location and status information with other users of the interactive system 500 via the interactive client 504 (e.g., using an appropriate status avatar), where the location and status information are similarly displayed within the context of the map interface of the interactive client 504 to the selected users.

[0160] The game system 728 provides various game functions within the context of the interactive client 504. The interactive client 504 provides a game interface that presents a list of available games that can be launched by the user within the context of the interactive client 504 and played with other users of the interactive system 500. The interactive system 500 also enables a specific user to invite such other users to participate in playing a specific game by sending an invitation from the interactive client 504 to the other users. The interactive client 504 also supports audio, video, and text messaging (e.g., chat) within the context of playing a game, provides a leaderboard for the game, and also supports the provision of in-game rewards (e.g., game currency and items).

[0161] The external resource system 730 provides an interface for the interactive client 504 to communicate with remote servers (e.g., third-party servers 512) to launch or access external resources (i.e., applications or applets). Each third-party server 512 hosts an application or a scaled-down version of an application (e.g., a game application, a utility application, a payment application, or a ride-sharing application) based on, for example, a markup language (e.g., HTML5). The interactive client 504 can launch a web-based resource (e.g., an application) by accessing an HTML5 file from a third-party server 512 associated with the web-based resource. The application hosted by the third-party server 512 is programmed in JavaScript using a software development kit (SDK) provided by the interactive server 520. The SDK includes an application programming interface (API) with functions that can be called or activated by the web-based application. The interactive server 520 hosts a JavaScript library that provides access to the specific user data of the interactive client 504 for a given external resource. HTML5 is an example of a technology used to program games, but applications and resources programmed based on other technologies can be used.

[0162] To integrate the functionality of the SDK into a web-based resource, the SDK is downloaded by a third-party server 512 from the interaction server 520 or received by the third-party server 512 in some other way. Once downloaded or received, the SDK is included as part of the application code of a web-based external resource. The code of the web-based resource can then call or activate certain functions of the SDK to integrate the features of the interaction client 504 into the web-based resource.

[0163] The SDK stored on the interaction server system 510 effectively provides a bridge between an external resource (e.g., an application 506 or a mini-program) and the interaction client 504. This gives the user a seamless experience of communicating with other users on the interaction client 504 while still preserving the look and feel of the interaction client 504. To bridge the communication between the external resource and the interaction client 504, the SDK facilitates the communication between the third-party server 512 and the interaction client 504. The WebView JavaScript Bridge running on the computing system 502 establishes two one-way communication channels between the external resource and the interaction client 504. Messages are sent asynchronously between the external resource and the interaction client 504 via these communication channels. Each SDK function activation is sent as a message and a callback. Each SDK function is implemented by constructing a unique callback identifier and sending a message with that callback identifier.

[0164] By using the SDK, not all information from the interaction client 504 is shared with the third-party server 512. The SDK restricts which information is shared based on the needs of the external resource. Each third-party server 512 provides an HTML5 file corresponding to the web-based external resource to the interaction server 520. The interaction server 520 can add a visual representation (e.g., a box design or other graphics) of the web-based external resource in the interaction client 504. Once the user selects the visual representation or instructs the interaction client 504 to access the features of the web-based external resource via the GUI of the interaction client 504, the interaction client 504 obtains the HTML5 file and instantiates the resources for accessing the features of the web-based external resource.

[0165] The interactive client 504 presents a graphical user interface for an external resource (e.g., a landing page or a splash screen). During, before, or after presenting the landing page or splash screen, the interactive client 504 determines whether the launched external resource has been previously authorized to access the user data of the interactive client 504. In response to determining that the launched external resource has been previously authorized to access the user data of the interactive client 504, the interactive client 504 presents another graphical user interface of the external resource that includes the functions and features of the external resource. In response to determining that the launched external resource has not been previously authorized to access the user data of the interactive client 504, after a display threshold period (e.g., 3 seconds) of the landing page or splash screen of the external resource, the interactive client 504 slides up a menu for authorizing the external resource to access the user data (e.g., animating the menu to emerge from the bottom of the screen to the middle or other part of the screen). The menu identifies the types of user data for which the external resource will be authorized to use. In response to receiving a user selection of an accept option, the interactive client 504 adds the external resource to the list of authorized external resources and allows the external resource to access the user data from the interactive client 504. The external resource is authorized by the interactive client 504 to access the user data under the OAuth 2 framework.

[0166] The interactive client 504 controls the types of user data shared with the external resource based on the type of the authorized external resource. For example, access to a first type of user data (e.g., a two-dimensional avatar of a user with or without different avatar characteristics) is provided to an external resource that includes a full-scale application (e.g., application 506). As another example, access to a second type of user data (e.g., payment information, a two-dimensional avatar of the user, a three-dimensional avatar of the user, and avatars with various avatar characteristics) is provided to an external resource that includes a small-scale version of the application (e.g., a web-based version of the application). Avatar characteristics include different ways of customizing the appearance and feel of the avatar (e.g., different poses, facial features, clothing, etc.).

[0167] The advertising system 732 operationally enables third parties to purchase advertisements to be presented to end users via the interactive client 504 and also handles the delivery and presentation of these advertisements.

[0168] Software architecture

[0169] Figure 8FIG. 800 is a block diagram showing a software architecture 802 that may be installed on any one or more of the devices described herein. The software architecture 802 is supported by hardware such as a machine 804 that includes a processor 806, a memory 808, and I / O components 810. In this example, the software architecture 802 may be conceptualized as a stack of layers, where each layer provides a specific function. The software architecture 802 includes layers such as an operating system 812, libraries 814, frameworks 816, and applications 818. In operation, the application 818 activates API calls 820 through the software stack and receives messages 822 in response to the API calls 820.

[0170] The operating system 812 manages hardware resources and provides common services. The operating system 812 includes, for example: a kernel 824, services 826, and drivers 828. The kernel 824 serves as an abstraction layer between the hardware and other software layers. For example, the kernel 824 provides memory management, processor management (e.g., scheduling), component management, networking and security settings, and other functions. The services 826 may provide other common services to other software layers. The drivers 828 are responsible for controlling or interfacing with the underlying hardware. For example, the drivers 828 may include a display driver, a camera device driver, or a low-power driver, a flash driver, a serial communication driver (e.g., a USB driver), a driver, an audio driver, a power management driver, etc.

[0171] The libraries 814 provide common low-level infrastructure used by the applications 818. The libraries 814 may include system libraries 830 (e.g., the C standard library) that provide functions such as memory allocation functions, string manipulation functions, mathematical functions, etc. Additionally, the libraries 814 may include API libraries 832, such as media libraries (e.g., libraries for supporting the presentation and manipulation of various media formats such as Moving Picture Experts Group-4 (MPEG4), High Efficiency Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Photographic Experts Group (JPEG or JPG), or Portable Network Graphics (PNG)), graphics libraries (e.g., the OpenGL framework for 2D and 3D rendering in graphical content on a display), database libraries (e.g., SQLite that provides various relational database functions), web libraries (e.g., WebKit that provides web browsing functions), etc. The libraries 814 may also include various other libraries 834 to provide many other APIs to the applications 818.

[0172] The framework 816 provides a common high-level infrastructure used by the applications 818. For example, the framework 816 provides various graphical user interface (GUI) functions, high-level resource management, and high-level location services. The framework 816 can provide a wide range of other APIs that can be used by the applications 818, some of which may be specific to a particular operating system or platform.

[0173] In an example, the applications 818 can include a home application 836, a contacts application 838, a browser application 840, a book reader application 842, a location application 844, a media application 846, a messaging application 848, a gaming application 850, and various other applications such as third-party applications 852. The applications 818 are programs that execute functions defined in the program. One or more of the applications 818 can be created using various programming languages and structured in various ways, such as object-oriented programming languages (e.g., Objective-C, Java, or C++) or procedural programming languages (e.g., C language or assembly language). In a particular example, the third-party application 852 (e.g., an application developed using an ANDROID TM or IOS TM software development kit (SDK) by an entity other than the vendor of a particular platform) can be mobile software running on a mobile operating system such as IOS TM 、ANDROID TM 、 Phone, or other mobile operating systems. In this example, the third-party application 852 can activate API calls 820 provided by the operating system 812 to facilitate the functions described herein.

[0174] Conclusion

[0175] Changes and modifications can be made to the disclosed examples without departing from the scope of the present disclosure. These and other changes or modifications are intended to be included within the scope of the present disclosure as expressed in the appended claims.

[0176] Glossary

[0177] "Carrier signal" means any non-tangible medium capable of storing, encoding, or carrying instructions executed by a machine and includes digital or analog communication signals or other non-tangible media to facilitate the communication of such instructions. Instructions can be sent or received over a network using a transmission medium via a network interface device.

[0178] "Client device" refers to any machine that interfaces with a communication network to obtain resources from one or more server systems or other client devices. A client device can be, but is not limited to, a mobile phone, desktop computer, laptop computer, portable digital assistant (PDA), smartphone, tablet computer, ultrabook, netbook, laptop, multiprocessor system, microprocessor-based or programmable consumer electronics, game console, set-top box, or any other communication device that a user can use to access the network.

[0179] "Communication network" refers to one or more portions of a network, which can be an ad hoc network, intranet, extranet, virtual private network (VPN), local area network (LAN), wireless LAN (WLAN), wide area network (WAN), wireless WAN (WWAN), metropolitan area network (MAN), the Internet, a portion of the Internet, a portion of the public switched telephone network (PSTN), plain old telephone service (POTS) network, cellular telephone network, wireless network, a network, other types of networks, or a combination of two or more such networks. For example, a network or a portion of a network can include a wireless network or a cellular network, and the coupling can be a code division multiple access (CDMA) connection, global system for mobile communications (GSM) connection, or other type of cellular or wireless coupling. In this example, the coupling can implement any data transfer technology among various types of data transfer technologies, such as single carrier radio transmission technology (1xRTT), evolved data optimized (EVDO) technology, general packet radio service (GPRS) technology, enhanced data rate GSM evolution (EDGE) technology, 3rd Generation Partnership Project (3GPP) including 3G, 4th generation wireless (4G) network, universal mobile telecommunications system (UMTS), high speed packet access (HSPA), worldwide interoperability for microwave access (WiMAX), long term evolution (LTE) standard, other data transfer technologies defined by various standards setting organizations, other long distance protocols, or other data transfer technologies.

[0180] "Component" refers to a device, physical entity, or logic having the following boundaries: the boundaries are defined by functions or subroutine calls, branch points, APIs, or other techniques provided for partitioning or modularizing a particular processing or control function. Components can be combined with other components via their interfaces to perform machine processing. A component can be an encapsulated functional hardware unit designed to be used with other components and can be part of a program that generally performs a particular function among related functions. Components can constitute software components (e.g., code implemented on a machine-readable medium) or hardware components. A "hardware component" is a tangible unit capable of performing certain operations and can be configured or arranged in a physical manner. In various examples, one or more computer systems (e.g., a stand-alone computer system, a client computer system, or a server computer system) or one or more hardware components of a computer system (e.g., a processor or a group of processors) can be configured by software (e.g., an application or a portion of an application) to operate to perform certain operations as described herein as a hardware component. A hardware component can also be implemented mechanically, electronically, or in any suitable combination thereof. For example, a hardware component can include dedicated circuitry or logic permanently configured to perform certain operations. A hardware component can be a dedicated processor, such as a field-programmable gate array (FPGA) or an application-specific integrated circuit (ASIC). A hardware component can also include programmable logic or circuitry that is temporarily configured by software to perform certain operations. For example, a hardware component can include software executed by a general-purpose processor or other programmable processor. Once configured by such software, the hardware component becomes a particular machine (or a particular component of a machine) that is uniquely customized to perform the configured function and is no longer a general-purpose processor. It will be recognized that the decision of whether to implement a hardware component mechanically in dedicated and permanently configured circuitry or in temporarily configured (e.g., software-configured) circuitry can be made for cost and time considerations. Accordingly, the phrase "hardware component" (or "hardware-implemented component") should be understood to include a tangible entity, i.e., an entity that is physically constructed, permanently configured (e.g., hard-wired), or temporarily configured (e.g., programmed) to operate in a certain manner or to perform certain operations described herein. Considering an example where a hardware component is temporarily configured (e.g., programmed), it is not necessary to configure or instantiate each hardware component at any given time. For example, in the case where a hardware component includes a general-purpose processor that is configured by software to become a dedicated processor, the general-purpose processor can be configured separately as different dedicated processors (e.g., including different hardware components) at different times. The software accordingly configures one or more specific processors to, for example, constitute a particular hardware component at one moment and different hardware components at different moments. Hardware components can provide information to other hardware components and receive information from other hardware components.Accordingly, the described hardware components can be considered to be communicatively coupled. In cases where multiple hardware components are present simultaneously, communication can be achieved through signal transmission between or among two or more of the hardware components (e.g., via appropriate circuitry and buses). In examples where multiple hardware components are configured or instantiated at different times, communication between such hardware components can be achieved, for example, by storing information in a memory structure accessible to the multiple hardware components and retrieving the information from the memory structure. For example, one hardware component can perform an operation and store the output of the operation in a memory device to which it is communicatively coupled. Then, another hardware component can access the memory device at a subsequent time to retrieve the stored output and process it. Hardware components can also initiate communication with input devices or output devices and can operate on resources (e.g., collections of information). The various operations of the example methods described herein can be performed, at least in part, by one or more processors temporarily configured (e.g., via software) or permanently configured to perform the relevant operations. Whether temporarily or permanently configured, such processors can constitute processor-implemented components that operate to perform one or more of the operations or functions described herein. As used herein, a "processor-implemented component" refers to a hardware component implemented using one or more processors. Similarly, the methods described herein can be at least in part implemented by a processor, where a particular one or more processors are examples of hardware. For example, at least some of the operations of the method can be performed by one or more processors or processor-implemented components. Additionally, one or more processors can also operate to support the execution of relevant operations in a "cloud computing" environment or operate as "software as a service" (SaaS). For example, at least some of the operations can be performed by a group of computers (as an example of machines including processors), where the operations can be accessed via a network (e.g., the Internet) and via one or more appropriate interfaces (e.g., APIs). The execution of certain operations can be distributed among the processors, not residing only within a single machine but deployed across multiple machines. In some examples, the processor or processor-implemented components can be located in a single geographical location (e.g., within a home environment, an office environment, or a server farm). In other examples, the processor or processor-implemented components can be distributed across multiple geographical locations.

[0181] "Machine-readable storage medium" refers to both machine storage media and transmission media. Thus, these terms include both storage devices / media and carrier / modulated data signals. The terms "computer-readable medium", "machine-readable medium", and "device-readable medium" mean the same thing and can be used interchangeably in this disclosure.

[0182] "Machine storage medium" means a single or multiple storage devices and media (e.g., centralized or distributed databases, and associated caches and servers) that store executable instructions, routines, and data. Thus, the term should be regarded as including, but not limited to, solid state memories as well as optical and magnetic media, including memories internal or external to a processor. Specific examples of machine storage media, computer storage media, and device storage media include: non-volatile memories, including, for example, semiconductor memory devices such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), FPGA, and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The terms "machine storage medium", "device storage medium", and "computer storage medium" mean the same thing and may be used interchangeably in this disclosure. The terms "machine storage medium", "computer storage medium", and "device storage medium" expressly exclude carrier waves, modulated data signals, and other such media, at least some of which are covered by the term "signal medium".

[0183] "Non-transitory machine-readable storage medium" means a tangible medium capable of storing, encoding, or carrying instructions executable by a machine.

[0184] "Signal medium" means any intangible medium capable of storing, encoding, or carrying instructions executable by a machine, and includes digital or analog communication signals or other intangible media to facilitate the communication of software or data. The term "signal medium" should be regarded as including any form of modulated data signal, carrier wave, etc. The term "modulated data signal" means a signal whose one or more characteristics are set or changed in such a way as to encode information in the signal. The terms "transmission medium" and "signal medium" mean the same thing and may be used interchangeably in this disclosure.

[0185] Without departing from the scope of this disclosure, changes and modifications can be made to the disclosed examples. These and other changes or modifications are intended to be included within the scope of this disclosure as expressed in the appended claims.

Claims

1. A machine, comprising: One or more processors; And A memory that stores instructions which, when executed by the one or more processors, cause the machine to perform operations including: Providing a user with an XR user interface of an extended reality (XR) system, the XR user interface including virtual objects displayed to the user; Determining a pinch position of a pinch hand gesture being made by the user; Scaling the virtual object based on the pinch position and a virtual object center point of the virtual object; And Redisplaying the scaled virtual object to the user in the XR user interface.

2. The machine according to claim 1, wherein Determining the pinch position includes: Using one or more imaging devices of the XR system to capture tracking video frame data; Determining hand tracking data based on the tracking video frame data; and Determining the pinch position based on the hand tracking data.

3. The machine according to claim 2, wherein, Determining the pinch position based on the hand tracking data includes: Determining a pinch position collider based on bone model data of the hand tracking data; and Determining the pinch position based on the pinch position collider.

4. The machine according to claim 1, wherein, Scaling the virtual object based on the pinch position and the virtual object center point of the virtual object includes: Generating a virtual object vector based on the center point of the virtual object and the pinch position; Generating a pinch vector based on the center point of the virtual object and a virtual object pinch collider of the virtual object; and Scaling the virtual object based on the virtual object vector and the pinch vector.

5. The machine according to claim 4, wherein Scaling the virtual object includes: Determining a scaling vector based on the virtual object vector and the pinch vector; and Scaling the virtual object based on the scaling vector.

6. The machine according to claim 1, wherein The operations further include: After redisplaying the scaled object, detecting that the user is still holding the pinch hand gesture; and In response to detecting that the pinch hand gesture is still being held, continuing to scale the virtual object based on the scaling vector.

7. The machine according to claim 1, wherein, The XR system includes a head-mounted device.

8. A computer-implemented method, comprising: Providing, by one or more processors, an XR user interface of an XR system to a user, the XR user interface including virtual objects displayed to the user; Determining, by the one or more processors, a pinch position of a pinch hand gesture being made by the user; Scaling, by the one or more processors, the virtual object based on the pinch position and a virtual object center point of the virtual object; And Redisplaying, by the one or more processors, the scaled virtual object to the user in the XR user interface.

9. The computer-implemented method according to claim 8, wherein, Determining the pinch position includes: Capturing, by the one or more processors, tracking video frame data using one or more imaging devices of the XR system; Determining, by the one or more processors, hand tracking data based on the tracking video frame data; and Determining, by the one or more processors, the pinch position based on the hand tracking data.

10. The computer-implemented method according to claim 9, wherein, Determining the pinch position based on the hand tracking data includes: Determining a pinch position collider by the one or more processors based on the skeletal model data of the hand tracking data; and Determining the pinch position by the one or more processors based on the pinch position collider.

11. The computer-implemented method according to claim 8, wherein, Scaling the virtual object based on the pinch position and the virtual object center point of the virtual object includes:[[]] Generating a virtual object vector by the one or more processors based on the center point of the virtual object and the pinch position; Generating a pinch vector by the one or more processors based on the center point of the virtual object and the virtual object pinch collider of the virtual object; and Scaling the virtual object by the one or more processors based on the virtual object vector and the pinch vector.

12. The computer-implemented method according to claim 11, wherein, Scaling the virtual object includes:[[]] Determining a scaling vector by the one or more processors based on the virtual object vector and the pinch vector; and Scaling the virtual object by the one or more processors based on the scaling vector.

13. The computer-implemented method according to claim 8, further comprising:[[]] Detecting by the one or more processors that the user is still maintaining the pinching hand gesture after redisplaying the scaled virtual object; And Continuing to scale the virtual object by the one or more processors in response to detecting that the pinching hand gesture is still being maintained, based on the scaling vector.

14. The computer-implemented method according to claim 8, wherein, The XR system includes a head-wearable device.

15. A non-transitory machine-readable storage medium, the machine-readable storage medium comprising instructions that, when executed by a machine, cause the machine to perform operations including the following:[[]] Providing an XR user interface of an XR system to a user, the XR user interface including a virtual object displayed to the user; Determining a pinch position of a pinching hand gesture being made by the user; Scaling the virtual object based on the pinch position and the virtual object center point of the virtual object; And Redisplaying the scaled virtual object to the user in the XR user interface.

16. The non-transitory machine-readable storage medium according to claim 15, wherein, Determining the pinch position includes:[[]] Using one or more camera devices of the XR system to capture tracking video frame data; Determining hand tracking data based on the tracking video frame data; and Determining the pinch position based on the hand tracking data.

17. The non-transitory machine-readable storage medium according to claim 16, wherein, Determining the pinch position based on the hand tracking data includes:[[]] Determining a pinch position collider based on the skeletal model data of the hand tracking data; and Determining the pinch position based on the pinch position collider.

18. The non-transitory machine-readable storage medium according to claim 15, wherein, Scaling the virtual object based on the pinch position and the virtual object center point of the virtual object includes:[[]] Generating a virtual object vector based on the center point of the virtual object and the pinch position; Generating a pinch vector based on the center point of the virtual object and the virtual object pinch collider of the virtual object; and Scaling the virtual object based on the virtual object vector and the pinch vector.

19. The non-transitory machine-readable storage medium according to claim 18, wherein, Scaling the virtual object includes:[[]] Determining a scaling vector based on the virtual object vector and the pinch vector; and Scaling the virtual object based on the scaling vector.

20. The non-transitory machine-readable storage medium according to claim 15, wherein, The XR system includes a head-wearable device.