Steerable camera device for AR hand tracking

By implementing hand tracking input pipeline and optical axis manipulation technology in the head-mounted AR system, the problem of complex interaction between users in the AR system is solved, and a more efficient and convenient user input experience is achieved.

CN119948385APending Publication Date: 2025-05-06SNAP INC
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202380068002.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2022-09-22
Filing Date
2023-09-19
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When compared to other mobile devices, the indications and actions of user intentions are more complex, and the lack of physical input devices such as touch screens or keyboards leads to inconvenience in interaction.

Method used

By tracking the input pipeline by hand, the camera device manipulation component is used to adjust the optical axis of the camera device, and capture the movement of the user's hand with high resolution, achieving richer and more diverse user input.

Benefits of technology

It improves the user's interaction efficiency and convenience in the AR system, allows users to input through more natural and intuitive gestures, and reduces their dependence on physical input devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119948385A_ABST
    Figure CN119948385A_ABST
Patent Text Reader

Abstract

A system for hand tracking of an augmented reality (AR) system. The AR system uses a camera device of the AR system to capture tracking video frame data of a hand of a user of the AR system. The AR system generates a skeleton model based on the tracked video frame data, and determines a position of a hand of the user based on the skeleton model. The AR system focuses a steerable camera of the AR system on a user's hand.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] Priority claim

[0002] This application claims the benefit of U.S. patent application Ser. No. 17 / 950,825, filed on Sept. 22, 2022, which is hereby incorporated by reference in its entirety. Technical Field

[0003] The present disclosure relates generally to user interfaces, and more particularly to user interfaces used in augmented reality and virtual reality. Background Art

[0004] The head-mounted device can be implemented with a transparent or translucent display, through which the user of the head-mounted device can view the surrounding environment. Such a device enables the user to look through the transparent or translucent display to view the surrounding environment, and can also see objects generated for display to appear as part of the surrounding environment and / or superimposed on the surrounding environment (e.g., virtual objects such as renderings of 2D or 3D graphic models, images, videos, text, etc.). This is generally referred to as "augmented reality" or "AR". The head-mounted device can also completely block the user's field of view and display a virtual environment through which the user can move or be moved. This is generally referred to as "virtual reality" or "VR". In a hybrid form, a view of the surrounding environment is captured using a camera device, and then the view is displayed to the user together with the enhancement on a display that blocks the user's eyes. As used herein, unless the context indicates otherwise, the term AR refers to augmented reality, virtual reality, and any mixture of these technologies.

[0005] A user of a head mounted device can access and use computer software applications to perform various tasks or participate in entertainment activities. Performing tasks or participating in entertainment activities may require inputting various commands and text into the head mounted device. Therefore, it is desirable to have a mechanism for inputting commands and text. BRIEF DESCRIPTION OF THE DRAWINGS

[0006] To easily identify the discussion of any particular element or act, the most significant digit(s) in a reference number refers to the figure number in which the element is first introduced.

[0007] Figure 1 is a perspective view of a head mounted device according to some examples.

[0008] Figure 2 According to some examples Figure 1 Another view of the headset.

[0009] Figure 3is a diagrammatic representation of a machine in the form of a computing device according to some examples within which a set of instructions may be executed to cause the machine to perform any one or more of the methodologies discussed herein.

[0010] Figure 4 is a collaboration diagram of a hand tracking input pipeline for an AR system according to some examples.

[0011] Figure 5A is a diagram illustrating operation of a manipulable camera device through an AR system according to some examples.

[0012] Figure 5B is a block diagram of a steerable camera device according to some examples.

[0013] Figure 5C is a block diagram of another manipulable camera device according to some examples.

[0014] Figure 5D Aspects of the subject matter according to one implementation are shown.

[0015] Figure 6 is a process flow diagram of a method of manipulating a steerable camera device according to some examples.

[0016] Figure 7 is a block diagram illustrating a software architecture within which the present disclosure may be implemented, according to some examples.

[0017] Figure 8 is a block diagram showing a networked system including details of a head-mounted AR system according to some examples.

[0018] Fig. 9 is a block diagram illustrating an example messaging system for exchanging data (eg, messages and associated content) over a network according to some examples. DETAILED DESCRIPTION

[0019] Head-mounted AR systems, such as glasses, are limited when it comes to available user input modalities. Indicating user intent and invoking an action or application is more complex for users of head-mounted AR systems when compared to other mobile devices, such as mobile phones. When using a mobile phone, a user can go to the home screen and click on a specific icon to launch an application. However, due to the lack of physical input devices such as a touch screen or keyboard, such interactions are not easily performed on head-mounted AR systems. Typically, users can indicate their intent by pressing a limited number of hardware buttons or using a small touchpad. Therefore, it is desirable to have an input modality that allows a wider variety of inputs to be used by users to indicate their intent through user input.

[0020] In some examples, an input modality used by an AR system is used to recognize gestures made by a user that do not involve direct manipulation of a virtual object (DMVO). When the user is wearing the AR system, the gesture is made by the user moving and positioning parts of the user's body, and those parts of the user's body are detectable by the AR system. The detectable parts of the user's body may include parts of the user's upper body, arms, hands, and fingers. The components of the gesture may include the movement of the user's arms and hands, the position of the user's arms and hands in space, and the positioning of the user's upper body, arms, hands, and fingers. The gesture is useful in providing an AR experience for the user because it provides a way to provide user input to the AR system during the AR experience without causing the user to shift their attention away from the AR experience. As an example, in an AR experience that is an operating manual for a piece of machinery, the user can simultaneously view the piece of machinery in a real-world scene through the lens of the AR system, view the AR overlay on the real-world scene view of the machinery, and provide user input to the AR system.

[0021] The cost of low-level image transmission and processing for hand tracking is roughly proportional to the number of pixels in the captured camera image. Accurate inference of hand position, sign language gestures, and user intent depends on having a sufficient number of captured pixels in the camera image, i.e., the camera image should have a high enough resolution to discern fine details of the user's hand. Many image sensors used in cameras have uniform resolution across their Field Of View (FOV), and the user's hand occupies only a portion of this FOV. Therefore, for some image sensors, it is desirable to implement a narrow field of view that limits the physical space in which the user can issue hand input, or it is desirable for pixels captured by the image sensor that are not used to recognize gestures to be eliminated.

[0022] In some examples, a camera manipulation component of an AR system changes (referred to herein as "manipulating") the angle of an optical axis of a narrow FOV camera of a camera component of a hand tracking input pipeline to the position of a user's hand and captures that area at high resolution rather than capturing a larger area of ​​possible hand positions at high resolution. As used herein, an "AR FOV" is a FOV in which a camera's image sensor may be able to detect user input, and a "camera FOV" is a narrow FOV or sub-FOV of the AR FOV that corresponds to where the camera manipulation component manipulates the optical axis of the manipulable camera.

[0023] In some examples, an optical axis of a steerable camera is manipulated using one or more physical actuators that reposition the steerable camera, such as by positioning a camera assembly including a sensor and optical elements using pneumatic, hydraulic, or electromechanical actuators, etc.

[0024] In some examples, the optical axis of the steerable camera device is steered using one or more configurable optical elements comprised of a spatial light modulator (SLM) that spatially modulates the opacity of the one or more configurable optical elements.

[0025] In some examples, the optical axis of a controllable camera device is manipulated using one or more configurable optical elements comprised of an SLM that spatially modulates the phase of the one or more configurable optical elements, such as by modifying the refractive index of one or more portions of the SLM or modifying one or more physical dimensions of the SLM.

[0026] In some examples, one or more micro-electromechanical systems (MEMS) mirrors or the like are used to steer the optical axis of the steerable camera device.

[0027] The camera manipulation component determines the position of the user's hand based on the real-world scene frame data and manipulates the optical axis of the steerable camera to place the user's hand in the camera FOV of the steerable camera. The steerable camera captures hand tracking image data at high resolution within the camera FOV of the steerable camera.

[0028] In some examples, the camera manipulation component determines the position of the user's hands in the wider FOV by using a steerable narrow FOV camera to sweep within the ARFOV of the AR system until the camera manipulation component recognizes the user's hands in the AR FOV.

[0029] In some examples, the camera manipulation component uses a wide FOV camera that covers the AR FOV of the AR system to determine the position of the user's hand in the wider FOV. The camera manipulation component recognizes the user's hand and determines its position using the wide FOV camera, and then manipulates the narrow FOV camera to capture a video image from that position.

[0030] In some examples, once the camera manipulation component has located the user's hand and begins tracking the user's hand, the camera manipulation component predicts the future position of the hand for future frames and avoids reacquiring the position of the user's hand from scratch on each frame during continuous input.

[0031] Other technical features may be readily apparent to those skilled in the art from the accompanying drawings, descriptions and claims.

[0032] Figure 1 is a head mounted AR system according to some examples (e.g., Figure 1 100). The glasses 100 may include a frame 102, which is made of any suitable material, such as plastic or metal, including any suitable shape memory alloy. In one or more examples, the frame 102 includes a first optical element holder or a left optical element holder 104 (e.g., a display or lens holder) and a second optical element holder or a right optical element holder 106 connected by a bridge 112. A first optical element or a left optical element 108 and a second optical element or a right optical element 110 may be disposed in the left optical element holder 104 and the right optical element holder 106, respectively. The right optical element 110 and the left optical element 108 may be lenses, displays, display components, or a combination of the foregoing. Any suitable display component may be disposed in the glasses 100.

[0033] Frame 102 also includes a left arm or temple piece 122 and a right arm or temple piece 124. In some examples, frame 102 can be formed from a single piece of material to have a unitary or unitary construction.

[0034] The glasses 100 may include a computing device, such as a computer 120, which may be of any suitable type so as to be carried by the frame 102, and in one or more examples, the computing device may be of a suitable size and shape to be at least partially disposed in one of the temple pieces 122 or temple pieces 124. The computer 120 may include one or more processors as well as memory, wireless communication circuitry, and a power source. As discussed below, the computer 120 includes a low-power circuitry, a high-speed circuitry, and a display processor. Various other examples may include these elements in different configurations or integrated together in different ways. Additional details of various aspects of the computer 120 may be implemented as shown by the data processor 802 discussed below.

[0035] The computer 120 also includes a battery 118 or other suitable portable power source. In some examples, the battery 118 is disposed in the left temple piece 122 and is electrically coupled to the computer 120 disposed in the right temple piece 124. The glasses 100 may include a connector or port (not shown) suitable for charging the battery 118, a wireless receiver, transmitter or transceiver (not shown), or a combination of such devices.

[0036] The glasses 100 include a first camera or left camera 114 and a second camera or right camera 116. Although two cameras are depicted, other examples contemplate the use of a single or additional (i.e., more than two) cameras. In one or more examples, the glasses 100 include any number of input sensors or other input / output devices in addition to the left camera 114 and the right camera 116. Such sensors or input / output devices may also include biometric sensors, position sensors, motion sensors, etc.

[0037] In some examples, left camera 114 and right camera 116 provide video frame data for use by glasses 100 to extract 3D information from a real-world scene.

[0038] The glasses 100 may also include a touchpad 126 mounted to or integrated with one or both of the left temple piece 122 and the right temple piece 124. The touchpad 126 is typically arranged vertically, with the touchpad 126 being approximately parallel to the temple of the user in some examples. As used herein, typically vertical alignment means that the touchpad is more vertical than horizontal, although potentially more vertical than this. Additional user input may be provided by one or more buttons 128, which in the example shown are disposed on the outer upper edges of the left optical element holder 104 and the right optical element holder 106. One or more touchpads 126 and buttons 128 provide a means by which the glasses 100 can receive input from a user of the glasses 100.

[0039] Figure 2 The glasses 100 are shown from the user's perspective. Figure 1 Several elements shown in FIG. have been omitted. Figure 1 As described, Figure 2 The illustrated eyeglasses 100 include left and right optical elements 108, 110 secured within left and right optical element holders 104, 106, respectively.

[0040] The glasses 100 include a forward optical assembly 202 including a right projector 204 and a right near-eye display 206 , and a forward optical assembly 210 including a left projector 212 and a left near-eye display 216 .

[0041] In some examples, the near-eye display is a waveguide. The waveguide includes a reflective structure or a diffractive structure (e.g., a grating and / or an optical element such as a mirror, lens, or prism). Light 208 emitted by projector 204 encounters the diffractive structure of the waveguide of near-eye display 206, which directs the light toward the right eye of the user to provide an image on or in the right optical element 110 superimposed with a view of the real-world scene seen by the user. Similarly, light 214 emitted by projector 212 encounters the diffractive structure of the waveguide of near-eye display 216, which directs the light toward the left eye of the user to provide an image on or in the left optical element 108 superimposed with a view of the real-world scene seen by the user. The combination of the GPU, the forward optical assembly 202, the left optical element 108, and the right optical element 110 provides an optical engine of the glasses 100. The glasses 100 use the optical engine to generate an overlay with the user's view of the real-world scene, including displaying a user interface to the user of the glasses 100.

[0042] However, it will be appreciated that other display technologies or configurations may be utilized within the optical engine to display images to a user in the user's field of view. For example, instead of providing a projector 204 and a waveguide, an LCD, LED or other display panel or surface may be provided.

[0043] In use, the user of the glasses 100 will be presented with information, content, and various user interfaces on the near-eye display. As described in more detail herein, the user can then use the touch pad 126 and / or buttons 128, voice input, or on an associated device (e.g., Figure 8 The user can interact with the glasses 100 through touch input on a client device 826 shown in FIG. 1 and / or hand movements, positions, and locations recognized by the glasses 100.

[0044] Figure 3 is a diagrammatic representation of a machine 300 (e.g., a computing device) within which instructions 310 (e.g., software, a program, an application, an applet, an app, or other executable code) may be executed that cause the machine 300 to perform any one or more of the methodologies discussed herein. The machine 300 may be used as Figure 1The computer 120 of the glasses 100. For example, the instructions 310 can cause the machine 300 to perform any one or more of the methods described herein. The instructions 310 transform the general non-programmed machine 300 into a specific machine 300 that is programmed to perform the functions described and shown in the described manner. The machine 300 can operate as a standalone device, or can be coupled (e.g., networked) to other machines. In a networked deployment, the machine 300 can operate as a server machine or a client machine in a server-client network environment, or as a peer machine in a peer (or distributed) network environment. The machine 300 may include, but is not limited to: a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a PDA, an entertainment media system, a cellular phone, a smart phone, a mobile device, a head-mounted device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, a web (network) device, a network router, a network switch, a network bridge, or any machine capable of sequentially or otherwise executing instructions 310 specifying actions to be taken by the machine 300. Further, while a single machine 300 is illustrated, the term "machine" shall also be taken to include any collection of machines that individually or jointly execute instructions 310 to perform any one or more of the methodologies discussed herein.

[0045] Machine 300 may include processor 302, memory 304, and I / O components 306 that may be configured to communicate with each other via bus 344. In some examples, processor 302 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an ASIC, a radio frequency integrated circuit (RFIC), another processor, or any suitable combination thereof) may include, for example, processor 308 and processor 312 that execute instructions 310. The term "processor" is intended to include multi-core processors, which may include two or more independent processors (sometimes referred to as "cores") that may execute instructions simultaneously. Although Figure 3 Multiple processors 302 are shown, but machine 300 may include a single processor with a single core, a single processor with multiple cores (eg, a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.

[0046] The memory 304 includes a main memory 314, a static memory 316, and a storage unit 318, which are all accessible by the processor 302 via the bus 344. The main memory 304, the static memory 316, and the storage unit 318 store instructions 310 that implement any one or more of the methods or functions described herein. The instructions 310 may also reside, in whole or in part, within the main memory 314, within the static memory 316, within the machine-readable medium 320 within the storage unit 318, within one or more of the processors 302 (e.g., within a processor's cache memory), or within any suitable combination thereof during execution thereof by the machine 300.

[0047] The I / O components 306 may include various components that receive input, provide output, generate output, send information, exchange information, capture measurements, etc. The specific I / O components 306 included in a particular machine will depend on the type of machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, while a headless server machine will be less likely to include such a touch input device. It will be appreciated that the I / O components 306 may include Figure 3 306. In various examples, the I / O components 306 may include output components 328 and input components 332. The output components 328 may include visual components (e.g., displays such as plasma display panels (PDPs), light emitting diode (LED) displays, liquid crystal displays (LCDs), projectors, or cathode ray tubes (CRTs)), acoustic components (e.g., speakers), tactile components (e.g., vibration motors, resistance mechanisms), other signal generators, etc. The input components 332 may include alphanumeric input components (e.g., keyboards, touch screens configured to receive alphanumeric input, optical keyboards, or other alphanumeric input components), point-based input components (e.g., mice, touch pads, trackballs, joysticks, motion sensors, or other pointing instruments), tactile input components (e.g., physical buttons, touch screens that provide location and / or force of touch or touch gestures, or other tactile input components), audio input components (e.g., microphones), etc.

[0048] In another example, the I / O component 306 may include a biometric component 334, a motion component 336, an environmental component 338, or a positioning component 340, as well as various other components. For example, the biometric component 334 includes a component for recognizing expressions (e.g., hand expressions, facial expressions, voice expressions, body postures, or eye tracking), measuring biosignals (e.g., blood pressure, heart rate, body temperature, sweat, or brain waves), identifying people (e.g., voice recognition, retinal recognition, facial recognition, fingerprint recognition, or EEG-based recognition), etc. The motion component 336 may include an inertial measurement unit (IMU), an acceleration sensor component (e.g., an accelerometer), a gravity sensor component, a rotation sensor component (e.g., a gyroscope), etc. Environmental components 338 include, for example, lighting sensor components (e.g., photometers), temperature sensor components (e.g., one or more thermometers that detect ambient temperature), humidity sensor components, pressure sensor components (e.g., barometers), acoustic sensor components (e.g., one or more microphones that detect background noise), proximity sensor components (e.g., infrared sensors that detect nearby objects), gas sensors (e.g., gas detection sensors that detect concentrations of hazardous gases for safety or measure pollutants in the atmosphere), or other components that can provide indications, measurements, or signals associated with the surrounding physical environment. Positioning components 340 include position sensor components (e.g., GPS receiver components), altitude sensor components (e.g., an altimeter or barometer that detects air pressure from which altitude can be derived), orientation sensor components (e.g., magnetometers), etc.

[0049] Various technologies may be used to achieve communication. I / O component 306 also includes communication component 342, which is operable to couple machine 300 to network 322 or device 324 via coupling 330 and coupling 326, respectively. For example, communication component 342 may include a network interface component or other suitable device that interfaces with network 322. In other examples, communication component 342 may include a wired communication component, a wireless communication component, a cellular communication component, a near field communication (NFC) component, Components (e.g. Low energy consumption), Device 324 may be another machine or any of a variety of peripheral devices (eg, a peripheral device coupled via USB).

[0050] In addition, the communication component 342 can detect an identifier or include a component operable to detect an identifier. For example, the communication component 342 can include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting a one-dimensional barcode such as a universal product code (UPC) barcode, a multi-dimensional barcode such as a Quick Response (QR) code, an Aztec code, a data matrix, a data symbol (Dataglyph), a MaxiCode, PDF417, an UltraCode, a UCC RSS-2D barcode, and other optical codes) or an acoustic detection component (e.g., a microphone for an audio signal for identifying a tag). In addition, various information can be obtained via the communication component 342, such as a location obtained via an Internet Protocol (IP) geolocation, a location obtained via a NFC smart tag detection component, a location obtained via a NFC smart tag detection component, and a location obtained via a NFC smart tag detection component. Location derived from signal triangulation, location derived via detection of NFC beacon signals that can indicate a specific location, etc.

[0051] Various memories (e.g., memory 304, main memory 314, static memory 316, and / or memory of processor 302) and / or storage unit 318 may store one or more sets of instructions and data structures (e.g., software) embodying or used by any one or more of the methods or functions described herein. These instructions (e.g., instructions 310) when executed by processor 302 cause various operations to implement the disclosed examples.

[0052] Instructions 310 may be sent or received via a network interface device (e.g., a network interface component included in communication component 342) using a transmission medium and using any of a number of well-known transmission protocols (e.g., Hypertext Transfer Protocol (HTTP)) over network 322. Similarly, instructions 310 may be sent or received to device 324 via coupling 326 (e.g., a peer-to-peer coupling) using a transmission medium.

[0053] Figure 4 4 is a collaboration diagram of a hand tracking input pipeline 428 of an AR system such as glasses 100 according to some examples. When a user 432 interacts with an AR application (e.g., an AR DMVO application component 418 and an AR interactive application component 416) provided by the AR system, the hand tracking input pipeline 428 captures real-world scene video frame data 420 of a gesture 436 made by the user 432. The hand tracking input pipeline 428 recognizes gesture fragments, gestures, and signs captured in the real-world scene video frame data 420, and provides the gesture fragments, gestures, and signs as user input to the AR application.

[0054] The hand tracking input pipeline 428 includes a camera component 402 that includes one or more cameras that capture video frame data of a real-world scene environment from the perspective of a user 432, such as Figure 1 The camera 114 and the camera 116 are used to generate real-world scene video frame data 420 based on the captured video frame data. The real-world scene video frame data 420 includes tracking video frame data of detectable parts of the user's body when the user 432 makes a gesture, and the detectable parts of the user's body include parts of the user's upper body, arms, hands, and fingers. The tracking video frame data includes: video frame data of the movement of parts of the user's upper body, arms, and hands when the user 432 makes a gesture or moves his hands and fingers to interact with the real-world scene environment; video frame data of the position of the user's arms and hands in space when the user 432 makes a gesture or moves his hands and fingers to interact with the real-world scene environment; and video frame data of the user 432 maintaining the positioning of his upper body, arms, hands, and fingers when the user 432 makes a gesture or moves his hands and fingers to interact with the real-world scene environment. The camera component 402 transmits the real-world scene video frame data 420 to the skeleton model inference component 404.

[0055] The skeleton model inference component 404 identifies landmark features based on the real-world scene video frame data 420. The skeleton model inference component 404 generates skeleton model data 426 based on the identified landmark features. The landmark features include landmarks about the upper body, arms, and hands of the user in the real scene environment. The skeleton model data 426 includes data representing the skeleton model of the part of the user's body (e.g., his hands and arms). In some examples, the skeleton model data 426 also includes landmark data such as landmark identification, position in the real-world scene environment, segments between joints, and classification information of one or more landmarks associated with the upper body, arms, and hands of the user.

[0056] In some examples, the skeleton model inference component 404 uses an artificial intelligence method and a skeleton classifier model previously generated using a machine learning method to identify landmark features based on the real-world scene video frame data 420. In some examples, the skeleton classifier model includes but is not limited to a neural network, a learning vector quantization network, a logistic regression model, a support vector machine, a random decision forest, a naive Bayes model, a linear discriminant analysis model, and a K nearest neighbor model. In some examples, the machine learning method may include but is not limited to supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, dimensionality reduction, self-learning, feature learning, sparse dictionary learning, and anomaly detection.

[0057] In some examples, the camera component 402 transmits the real-world scene frame data 426 to the rough hand location inference component 412. The rough hand location inference component 412 generates coordinate transformation data 424 based on the real-world scene frame data 426. The rough hand location inference component 412 receives the real-world scene video frame data 420 of the real-world scene and extracts features of objects in the real-world scene including the upper body, arms and hands of the user from the real-world scene video frame data. The rough hand location inference component 412 generates the coordinate transformation data 424 based on the extracted features. The coordinate transformation data 424 includes a skeletal model of the hand of the user 432 when the user makes a gesture 436 when interacting with an AR application provided by the AR system. The skeletal model is continuously generated and the coordinate transformation of the skeletal model into the user coordinate system of the AR system is performed. The other components of the hand tracking input pipeline 428 can use the coordinate transformation data 424 to determine the position of the user's hand within the FOV of the camera component 402. The rough hand position inference component 412 transmits the coordinate transformation data 424 to the camera manipulation component 434 .

[0058] As reference Figure 6 More fully described, camera manipulation component 434 receives coordinate transformation data 424 and generates camera manipulation command data 438 based on coordinate transformation data 424. Camera manipulation command data 438 includes commands instructing camera component 402 to adjust the optical axis of a manipulable camera 440 of camera component 402 to align the optical axis with the position of the user's hand.

[0059] In some examples, the coarse hand positioning inference component 412 also transmits the coordinate transformation data 424 to the AR DMVO application component 418.

[0060] The gesture segment inference component 406 receives the skeletal model data 426 from the skeletal model inference component 404 and generates gesture segment data 422 based on the skeletal model data 426. The gesture is specified by the hand tracking input pipeline 428 according to the combination of gesture segments. The gesture segment is in turn composed of the combination and relationship of the landmarks included in the skeletal model data 426. When the hand tracking input pipeline 428 extracts the gesture segments from the skeletal model data 426 in a different layer than assembling the hand movements into gestures through the hand tracking input pipeline 428, the designer of the AR system can create new gestures constructed from existing gesture segments that make up known gestures without having to retrain the machine learning component of the hand tracking input pipeline 428.

[0061] In some examples, the posture segment inference component 406 compares one or more skeletal models included in the skeletal model data 426 with a previously generated posture segment model and generates one or more posture segment probabilities based on the comparison. The one or more posture segment probabilities indicate the probability that a specified posture segment can be identified from the skeletal model data 426. The posture segment inference component 406 generates the posture segment data 422 based on the one or more posture segment probabilities. In another example, the posture segment inference component 406 classifies the skeletal models in the skeletal model data 426 based on the posture segment models previously generated using the artificial intelligence method and the machine learning method to determine one or more posture segment probabilities. The posture segment inference component 406 transmits the posture segment data 422 to the posture inference component 408 and the posture text input recognition component 410.

[0062] The posture inference component 408 receives the posture fragment data 422 and determines the posture data 430 based on the posture fragment data 422. In some examples, the posture inference component 408 compares the posture fragments identified in the posture fragment data 422 with the posture identification data that identifies the specific posture. The posture identification consists of one or more posture fragments corresponding to the specific posture. The posture identification is defined using a grammar whose symbols correspond to the posture fragments. For example, a posture identifier for a posture is “LEFT_PALMAR_FINGERS EXTENDED_RIGHT PALMAR_FINGERS_EXTENDED”, where: “LEFT” is a symbol corresponding to a hand classifier indicating that the user's left hand has been identified; “PALMAR” is a symbol corresponding to a hand classifier indicating that the user's palm has been identified, and modifies “LEFT” to indicate that the user's left palm has been identified; “FINGERS” is a symbol corresponding to a hand classifier indicating that the user's fingers have been identified; and “EXTENDED” is a symbol corresponding to a hand classifier indicating that the user's fingers are extended and modifies “FINGERS”. In another example, a posture identifier is a single token, such as a number, that identifies a posture based on its constituent posture fragments. The posture identifier identifies the posture in the context of a physical description of the posture. The posture inference component 408 transmits the posture data 430 to the AR interactive application component 416.

[0063] The gesture text input recognition component 410 receives the gesture segment data 422 and generates the symbol data 414 based on the gesture segment data 422. In some examples, the gesture text input recognition component 410 compares the gesture segment identified in the gesture segment data 422 with the symbol data identifying specific characters, words, and commands. For example, the symbol data for the gesture is the character "V" as a gesture for finger spelling in American Sign Language (ASL). The individual gesture segments for gestures may be “LEFT” for the left hand, “PALMAR” for the palm of the left hand, “INDEXFINGER” for the index finger, “EXTENDED” for the modified “INDEXFINGER”, “MIDDLEFINGER” for the middle finger, “EXTENTED” for the modified “MIDDLEFINGER”, “RINGFINGER” for the ring finger, “CURLED” for the modified “RINGFINGER”, “LITTLEFINGER” for the little finger, “CURLED” for the modified “LITTLEFINGER”, “THUMB” for the thumb, and “CURLED” for the modified “THUMB”.

[0064] In some examples, the gesture text input recognition component 410 may recognize a whole word based on the gesture fragment indicated by the gesture fragment data 422. In other examples, the gesture text input recognition component 410 may also recognize a command based on the gesture fragment indicated by the gesture fragment data 422, such as a command corresponding to a specified keystroke set in an input system having a keyboard.

[0065] The gesture text input recognition component 410 transmits the symbol data 414 to the AR interactive application component 416 .

[0066] The AR application components (e.g., AR DMVO application component 418 and AR interactive application component 416) executed by the AR system are users of data (e.g., coordinate transformation data 424, skeletal model data 426, gesture data 430, and symbolic data 414) generated by the hand tracking input pipeline 428. The AR system executes the AR DMVO application component 418 to provide a user interface to a user of the AR system using direct manipulation of visual objects within a 2D or 3D user interface. The AR system executes the AR interactive application component 416 to provide a user interface, such as an AR experience, to a user of the AR system using gestures as an input modality.

[0067] In some examples, the camera component 402, the skeletal model inference component 404, and the coarse hand position inference component 412 communicate using an automatically synchronized shared memory buffer. In addition, the skeletal model inference component 404 and the coarse hand position inference component 412 publish the skeletal model data 426 and the coordinate transformation data 424, respectively, on a memory buffer that can be accessed by components and applications outside the hand tracking input pipeline 428 (e.g., the AR DMVO application component 418).

[0068] In many examples, gesture segment inference component 406, gesture inference component 408, and gesture text input recognition component 410 communicate gesture data 430 and symbol data 414, respectively, via an inter-process communication method.

[0069] In some examples, hand tracking input pipeline 428 continuously operates to generate and publish gesture data 430 , symbol data 414 , coordinate transformation data 424 based on real-world scene frame data 426 generated by one or more cameras of the AR system.

[0070] Figure 5A is a diagram showing a camera device that can be manipulated by an AR system operation, and Figure 5B and Figure 5C 5 is a block diagram of a steerable camera according to some examples. An AR system, such as glasses 100, changes (steers) an angle 518 of an optical axis 512 of a steerable camera (e.g., steerable camera 520 and steerable camera 526) to include one or more hands 510 of a user in a camera FOV 504 of the steerable camera when the user makes a gesture using the AR system. Figure 6 502 and its related description. The AR camera of the AR system captures video frame data of the real world scene 508 in the AR FOV 502 of the AR camera. The optical axis 506 of the AR camera is aligned with the optical axis of the user wearing the glasses 100. In some examples, the AR camera and the steerable camera are the same camera, and the camera manipulation component 434 of the hand tracking input pipeline 428 manipulates the steerable camera to alternately scan between one or more hands 510 of the user and the real world scene 508.

[0071] Figure 5Bis a diagram of a steerable camera 520 of the camera component 402 of the hand tracking input pipeline 428. The steerable camera 520 includes one or more actuators 516 connected to a camera 514 having an image sensor and a lens assembly. The camera 514 is movably attached to an inner surface of a housing 530 of the steerable camera 520 and is positioned so that the lens of the camera 514 is aligned with an aperture 532 of the housing 530. The pitch angle of the camera 514 is adjusted by sending a pitch adjustment command to the camera component 402. The camera component 402 receives the pitch adjustment command and generates an electrical signal that causes the pitch actuator of the one or more actuators 516 to move the camera and change the pitch angle of the camera 514 along the pitch optical axis angle 518 and, therefore, change the optical axis 512 of the steerable camera 520. In some examples, the yaw angle of camera 514 is adjusted by transmitting a yaw adjustment command to camera component 402. Camera component 402 receives the yaw adjustment command and generates an electrical signal so that a yaw actuator of one or more actuators 516 of manipulable camera 520 changes the yaw angle of camera 514 by a yaw optical axis angle (not shown), and thus changes the optical axis 512 of manipulable camera 520.

[0072] Figure 5C 5 is a diagram of a steerable camera 526 of the camera component 402 of the hand tracking input pipeline 428. The steerable camera 526 includes one or more actuators 524 that move a mirror 528 pivotally attached to an inner surface of a housing 534 of the steerable camera 526. The camera 522 having an image sensor and a lens assembly remains stationary while the pitch angle and / or yaw angle of the mirror are adjusted using the one or more actuators 524. The pitch angle of the steerable camera 526 is adjusted by sending a pitch adjustment command to the camera component 402. The camera component 402 receives the pitch adjustment command and generates an electrical signal that causes the pitch actuator of the steerable camera to change the pitch angle of the mirror, thereby changing (steering) the optical axis 512 of the steerable camera 526 by the pitch optical axis angle 518. In some examples, the yaw angle of the steerable camera 526 is adjusted by transmitting a yaw adjustment command to the camera assembly 402. The camera assembly 402 receives the yaw adjustment command and generates an electrical signal so that a yaw actuator in one or more actuators 524 changes the yaw angle of the mirror, thereby changing (steering) the optical axis 512 of the steerable camera 526 by a yaw optical axis angle (not shown).

[0073] Figure 5D5 is a diagram of a steerable camera 542 of a camera component 402 of a hand tracking input pipeline 428 according to some examples of the present disclosure. In some examples, the steerable camera 542 of the camera component 402 includes an optical assembly having one or more configurable SLMs 544 that spatially modulates the opacity of the optical assembly and / or the phase of one or more optical elements. When adjusting the spatial distribution of the opacity and / or phase of the configurable SLM 544 optical elements, the camera 522 having an image sensor and lens assembly remains stationary, and the steerable camera 542 remains stationary. The camera 540 and the configurable SLMs 544 are mounted in a housing 546, whereby the optical axis 536 of the camera 540 passes through an aperture 548 of the housing. The spatial distribution of the opacity and / or phase of the configurable SLM 544 optical elements is adjusted by sending a phase adjustment command to the camera component 402. Camera assembly 402 generates thermal or electrical signals that cause the spatial distribution of opacity and / or phase of configurable SLM 544 optical elements to be altered, thereby changing (steering) optical axis 536 of camera assembly 402 via pitch angle 538 and / or yaw angle (not shown).

[0074] Figure 6 is a process flow diagram of a steerable camera manipulation method 600 according to some examples. The AR system uses the steerable camera manipulation method 600 to steer the steerable camera 440 of the camera assembly 402 to align the optical axis of the steerable camera 440 with one or more hands of a user 432 of the AR system.

[0075] As reference Figure 4As previously described, the camera component 402 having the manipulable camera 440 generates real-world scene video frame data 420 based on the captured video frame data. The real-world scene video frame data 420 includes tracking video frame data of detectable portions of the user's body when the user 432 makes a gesture, and the detectable portions of the user's body include portions of the user's upper torso, arms, hands, and fingers. The camera component 402 transmits the real-world scene video frame data 420 to the skeletal model inference component 404. The skeletal model inference component 404 identifies landmark features based on the real-world scene video frame data 420. The skeletal model inference component 404 generates skeletal model data 426 based on the identified landmark features, and transmits the real-world scene frame data 426 to the coarse hand positioning inference component 412. The coarse hand positioning inference component 412 generates coordinate transformation data 424 based on the real-world scene frame data 426. For example, the coordinate transformation data 424 includes the coordinates of the skeletal model of one or more hands of the user 432 represented in a 3D spherical coordinate system with the user 432's viewpoint as the origin. That is, each joint of the skeletal model has the coordinates of the radius "r" of the joint from the point of origin, the inclination angle "θ" of the joint, and the azimuth angle "Φ" of the joint. The rough hand positioning inference component 412 transmits the coordinate transformation data 424 to the camera manipulation component 434.

[0076] In operation 602 , the camera manipulation component 434 receives the coordinate transformation data 424 from the skeleton model inference component 404 .

[0077] In operation 604, the camera manipulation component 434 determines the position of one or more hands of the user 432 within the camera FOV of the camera component 402 based on the coordinate transformation data 424. For example, the camera manipulation component 434 determines the center of mass of the skeletal model of the one or more hands of the user 432 and casts a ray extending from the viewpoint of the user 432 to the center of mass of the skeletal model. The ray has coordinates (r, θ, Φ), where r is the distance from the viewpoint of the user 432 to the center of mass of the skeletal model, θ is the tilt angle of the ray, and Φ is the azimuth angle of the ray.

[0078] In operation 606, the camera manipulation component 434 generates manipulation command data based on the position of the one or more hands of the user 432. For example, the pitch angle of the optical axis of the steerable camera 440 of the camera component 402 corresponds to the azimuth angle or Φ of the ray, and the yaw angle of the steerable camera 440 corresponds to the tilt angle or θ of the ray. The camera manipulation command data 438 includes a pitch adjustment command instructing the camera component 402 to set the pitch angle of the steerable camera 440 of the camera component 402 to the azimuth angle of the ray projected from the viewpoint of the user 432 to the center of mass of the skeletal model of the one or more hands of the user 432. In some examples, the camera manipulation command data 438 also includes a yaw adjustment command for the camera component to set the yaw angle of the manipulable camera 440 of the camera component 402 to the tilt angle of a ray projected from the viewpoint of the user 432 to the center of mass of the skeletal model of one or more hands of the user 432.

[0079] In operation 608, camera manipulation component 434 transmits camera manipulation command data 438 to camera component 402. Camera component 402 receives camera manipulation command data 438 and manipulates to align the optical axis of steerable camera 440 with the center of mass of one or more hands of user 432 based on the pitch adjustment command of camera manipulation command data 438. This focuses steerable camera 440 on one or more hands of user 432 and places the one or more hands of the user in the camera FOV of steerable camera 440. In some examples, camera component 402 receives camera manipulation command data 438 and manipulates to align the optical axis of steerable camera 440 with the center of mass of one or more hands of user 432 based on the yaw adjustment command of camera manipulation command data 438. This focuses the steerable camera 440 on the hand or hands of the user 432 and places the hand or hands of the user in the camera FOV of the steerable camera 440 .

[0080] In some examples, a non-steerable camera of camera component 402 having a camera FOV equal to the AR FOV captures real-world scene video frame data 420 for generating steering command data. In some examples, the non-steerable camera has a wider camera FOV than the steerable camera 440. In some examples, the non-steerable camera has a lower resolution than the steerable camera 440.

[0081] In some examples, to initially locate one or more hands of user 432, camera manipulation component 434 uses steerable camera 440 to scan one or more hands within AR FOV. Once camera manipulation component 434 finds one or more hands of user 432, camera manipulation component 434 predicts the next position of one or more hands of user 432 based on the current position of one or more hands of user 432 using a look-ahead process. Camera manipulation component 434 manipulates steerable camera to focus on the next position. In some examples, camera manipulation component 434 receives gesture segment data 422 from gesture segment inference component 406, and determines a possible next gesture segment and position based on gesture segment data 422 and language model. In some examples, language model is for American Sign Language (ASL), and language model is used to recognize gesture segments of sign language in ASL.

[0082] Camera manipulation component 434 determines a possible next gesture segment N based on previous gesture segments N-1, N-2, etc. and language gesture segment data 422. Camera manipulation component 434 generates a possible next position of one or more hands of user 432 based on possible next gesture segment N.

[0083] In an example, the camera manipulation component 434 determines a possible next gesture segment based on a language model that is a hidden Markov model that predicts what the possible next gesture segment N is based on one or more of the previous gesture segments N-1, N-2, etc.

[0084] In another example, the camera manipulation component 434 uses an AI method to determine the next gesture segment N based on a language model generated using a machine learning method. In some examples, the language model includes but is not limited to a neural network, a learning vector quantization network, a logistic regression model, a support vector machine, a random decision forest, a naive Bayes model, a linear discriminant analysis model, and a K nearest neighbor model. In some examples, the machine learning method may include but is not limited to supervised learning, unsupervised learning, semi-supervised learning, reinforcement learning, dimensionality reduction, self-learning, feature learning, sparse dictionary learning, and anomaly detection.

[0085] Figure 7700 is a block diagram illustrating a software architecture 704 that may be installed on any one or more of the devices described herein. The software architecture 704 is supported by hardware, such as a computing machine 702, including a processor 720, a memory 726, and an I / O component 738. In this example, the software architecture 704 may be conceptualized as a stack of layers in which each layer provides specific functionality. The software architecture 704 includes layers, such as an operating system 712, a library 708, a framework 710, and an application 706. In operation, the application 706 invokes an API call 750 through the software stack and receives a message 752 in response to the API call 750.

[0086] The operating system 712 manages hardware resources and provides public services. The operating system 712 includes, for example, a kernel 714, services 716, and drivers 722. The kernel 714 serves as an abstraction layer between the hardware and other software layers. For example, the kernel 714 provides functions such as memory management, processor management (e.g., scheduling), component management, networking, and security settings. Services 716 can provide other public services to other software layers. Drivers 722 are responsible for controlling or interfacing with the underlying hardware. For example, drivers 722 may include display drivers, camera drivers, or Low-power drivers, Flash drivers, Serial communications drivers (e.g., Universal Serial Bus (USB) drivers), Drivers, audio drivers, power management drivers, etc.

[0087] The library 708 provides a low-level common infrastructure used by the application 706. The library 708 may include a system library 718 (e.g., a C standard library) that provides functions such as memory allocation functions, string manipulation functions, mathematical functions, etc. In addition, the library 708 may include an API library 724, such as a media library (e.g., a library for supporting the presentation and manipulation of various media formats such as Moving Picture Experts Group 4 (MPEG4), Advanced Video Coding (H.264 or AVC), Moving Picture Experts Group Layer 3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Photographic Experts Group (JPEG or JPG), or Portable Network Graphics (PNG)), a graphics library (e.g., an OpenGL framework for presenting two-dimensional (2D) and three-dimensional (3D) graphics content on a display, GLMotif for implementing a user interface), an image feature extraction library (e.g., OpenIMAJ), a database library (e.g., SQLite providing various relational database functions), a web library (e.g., WebKit providing web browsing functions), etc. The library 708 may also include various other libraries 728 to provide many other APIs to the application 706 .

[0088] The framework 710 provides a high-level common infrastructure used by the applications 706. For example, the framework 710 provides various graphical user interface (GUI) functions, advanced resource management, and advanced positioning services. The framework 710 can provide a wide range of other APIs that can be used by the applications 706, some of which can be specific to a particular operating system or platform.

[0089] In some examples, applications 706 may include home applications 736, contact applications 730, browser applications 732, book reader applications 734, location applications 742, media applications 744, messaging applications 746, game applications 748, and a broad category of other applications such as third-party applications 740. Applications 706 are programs that perform functions defined in the program. Various programming languages ​​may be used to create one or more of the applications 706 constructed in various ways, such as object-oriented programming languages ​​(e.g., Objective-C, Java, or C++) or procedural programming languages ​​(e.g., C or assembly language). In a specific example, third-party applications 740 (e.g., those developed by entities other than the vendor of a particular platform using ANDROID TM or IOS TM Software Development Kit (SDK) can be used to develop applications on platforms such as IOS TM ANDROID TM , Mobile software running on the mobile operating system of the Phone or another mobile operating system. In this example, third-party applications 740 can activate API calls 750 provided by the operating system 712 to facilitate the functions described herein.

[0090] Figure 8 830 is a block diagram showing a networked system 800 including details of the glasses 100 according to some examples. The networked system 800 includes the glasses 100, a client device 826, and a server system 832. The client device 826 may be a smart phone, a tablet computer, a tablet phone, a laptop computer, an access point, or any other such device capable of connecting to the glasses 100 using a low-power wireless connection 836 and / or a high-speed wireless connection 834. The client device 826 is connected to the server system 832 via a network 830. The network 830 may include any combination of wired and wireless connections. The server system 832 may be one or more computing devices that are part of a service or network computing system. The client device 826 and any elements of the server system 832 and the network 830 may be connected using a wireless network, such as a wireless network or a wireless network. Figure 7 and Figure 3 The details of the software architecture 704 or machine 300 described in are implemented.

[0091] The glasses 100 include a data processor 802, a display 810, one or more cameras 808, and additional input / output elements 816. The input / output elements 816 may include a microphone, an audio speaker, a biometric sensor, an additional sensor, or an additional display element integrated with the data processor 802. Figure 7 and Figure 3 Examples of input / output elements 816 are further discussed. For example, input / output elements 816 may include any I / O component 306, including output component 328, motion component 336, etc. Figure 2 Examples of display 810 are discussed in . In the specific examples described herein, display 810 includes displays for the left and right eyes of a user.

[0092] The data processor 802 includes an image processor 806 (eg, a video processor), a GPU & display driver 838, a tracking module 840, an interface 812, low power circuitry 804, and high speed circuitry 820. The components of the data processor 802 are interconnected by a bus 842.

[0093] The interface 812 refers to any source of user commands provided to the data processor 802. In one or more examples, the interface 812 is a physical button that, when pressed, sends a user input signal from the interface 812 to the low-power processor 814. The low-power processor 814 can process pressing such a button and then immediately releasing it as a request to capture a single image, and vice versa. The low-power processor 814 can process pressing such a button for a first period of time as a request to capture video data when the button is pressed and stop video capture when the button is released, wherein the video captured when the button is pressed is stored as a single video file. Alternatively, pressing the button for a long period of time can capture a still image. In some examples, the interface 812 can be any mechanical switch or physical interface capable of accepting user input associated with requesting data from the camera 808. In other examples, the interface 812 can have a software component, or can be associated with a command received wirelessly from another source, such as from the client device 826.

[0094] The image processor 806 includes circuitry for receiving signals from the camera 808 and processing those signals from the camera 808 into a format suitable for storage in the memory 824 or for transmission to the client device 826. In one or more examples, the image processor 806 (e.g., a video processor) includes a microprocessor integrated circuit (IC) customized for processing sensor data from the camera 808, and volatile memory used by the microprocessor in operation.

[0095] The low power circuit system 804 includes a low power processor 814 and a low power wireless circuit system 818. These elements of the low power circuit system 804 can be implemented as separate elements or can be implemented on a single IC as part of a single system on a chip. The low power processor 814 includes logic for managing other elements of the glasses 100. As described above, for example, the low power processor 814 can accept user input signals from the interface 812. The low power processor 814 can also be configured to receive input signals or command communications from a client device 826 via a low power wireless connection 836. The low power wireless circuit system 818 includes circuit elements for implementing a low power wireless communication system. Bluetooth TM Smart, also known as Bluetooth TM Low power consumption is a standard implementation of a low power wireless communication system that may be used to implement the low power wireless circuitry 818. In other examples, other low power communication systems may be used.

[0096] High-speed circuit system 820 includes a high-speed processor 822, a memory 824, and a high-speed wireless circuit system 828. High-speed processor 822 can be any processor capable of managing high-speed communications and operations of any general-purpose computing system used by data processor 802. High-speed processor 822 includes processing resources used to manage high-speed data transmission over high-speed wireless connection 834 using high-speed wireless circuit system 828. In some examples, high-speed processor 822 executes an operating system such as a LINUX operating system or a program such as a UNIX operating system. Figure 7 The high-speed processor 822, which executes the software architecture of the data processor 802, manages data transmission with the high-speed wireless circuit system 828, in addition to any other duties. In some examples, the high-speed wireless circuit system 828 is configured to implement the Institute of Electrical and Electronics Engineers (IEEE) 802.11 communication standard, which is also referred to herein as Wi-Fi. In other examples, the high-speed wireless circuit system 828 can implement other high-speed communication standards.

[0097] The memory 824 includes any storage device capable of storing camera data generated by the camera 808 and the image processor 806. Although the memory 824 is shown as being integrated with the high-speed circuitry 820, in other examples, the memory 824 may be a separate, independent element of the data processor 802. In some such examples, electrical wiring may provide a connection from the image processor 806 or the low-power processor 814 to the memory 824 through a chip including the high-speed processor 822. In other examples, the high-speed processor 822 may manage addressing of the memory 824 so that the low-power processor 814 will initiate the high-speed processor 822 whenever a read or write operation involving the memory 824 is desired.

[0098] The tracking module 840 estimates the pose of the glasses 100. For example, the tracking module 840 uses image data and associated inertial data from the camera 808 and the positioning component 340, as well as GPS data, to track the position and determine the pose of the glasses 100 relative to a reference frame (e.g., a real-world scene). The tracking module 840 continuously collects and uses updated sensor data describing the movement of the glasses 100 to determine an updated three-dimensional pose of the glasses 100 that indicates changes in relative position and orientation relative to physical objects in the real-world scene. The tracking module 840 allows the glasses 100 to visually place virtual objects relative to physical objects within the user's field of view via the display 810.

[0099] The GPU & display driver 838 can use the pose of the glasses 100 to generate frames of virtual content or other content to be presented on the display 810 when the glasses 100 are operating in a traditional augmented reality mode. In this mode, the GPU & display driver 838 generates updated frames of virtual content based on the updated three-dimensional pose of the glasses 100, which reflects the changes in the user's position and orientation relative to physical objects in the user's real-world scene.

[0100] One or more functions or operations described herein may also be performed in an application resident on the glasses 100 or on the client device 826 or on a remote server. For example, one or more functions or operations described herein may be performed by one of the applications 706, such as the messaging application 746.

[0101] Fig. 9is a block diagram illustrating an example messaging system 900 for exchanging data (e.g., messages and associated content) over a network. The messaging system 900 includes multiple instances of a client device 826 that host several applications including a messaging client 902 and other applications 904. The messaging client 902 is communicatively coupled to other instances of the messaging client 902 (e.g., hosted on respective other client devices 826), a messaging server system 906, and a third-party server 908 via a network 830 (e.g., the Internet). The messaging client 902 may also communicate with the locally hosted application 904 using an application program interface (API).

[0102] The messaging clients 902 are able to communicate and exchange data with other messaging clients 902 and with a messaging server system 906 via the network 830. The data exchanged between the messaging clients 902 and between the messaging clients 902 and the messaging server system 906 include functions (e.g., commands for activating functions) and payload data (e.g., text, audio, video, or other multimedia data).

[0103] The messaging server system 906 provides server-side functionality to a particular messaging client 902 via the network 830. Although some functionality of the messaging system 900 is described herein as being performed by the messaging client 902 or by the messaging server system 906, it may be a design choice to locate some functionality within the messaging client 902 or within the messaging server system 906. For example, it may be technically preferable to initially deploy some technologies and functionality within the messaging server system 906, but then migrate the technologies and functionality to the messaging client 902 where the client device 826 has sufficient processing power.

[0104] The messaging server system 906 supports various services and operations provided to the messaging client 902. Such operations include sending data to the messaging client 902, receiving data from the messaging client 902, and processing data generated by the messaging client 902. As examples, the data may include message content, client device information, geographic location information, media enhancements and overlays, message content persistence conditions, social network information, and live event information. The data exchange within the messaging system 900 is activated and controlled by functions available via the user interface (UI) of the messaging client 902.

[0105] Turning now specifically to the messaging server system 906, an application program interface (API) server 910 is coupled to and provides a programming interface to an application server 914. The application server 914 is communicatively coupled to a database server 916, which facilitates access to a database 920 that stores data associated with messages processed by the application server 914. Similarly, a web server 924 is coupled to and provides a web-based interface to the application server 914. To this end, the web server 924 handles incoming network requests via the Hypertext Transfer Protocol (HTTP) and several other related protocols.

[0106] The application program interface (API) server 910 receives and sends message data (e.g., commands and message payloads) between the client device 826 and the application server 914. Specifically, the application program interface (API) server 910 provides a set of interfaces (e.g., routines and protocols) that can be called or queried by the messaging client 902 to activate the functions of the application server 914. The application program interface (API) server 910 exposes various functions supported by the application server 914, including: account registration; login functionality; sending messages from a particular messaging client 902 to another messaging client 902 via the application server 914, sending media files (e.g., images or videos) from the messaging client 902 to the messaging server 912, and possible access by another messaging client 902; setting up a collection of media data (e.g., a story); retrieving a friend list of a user of the client device 826; retrieving such a collection; retrieving messages and content; adding and removing entities (e.g., friends) to an entity graph (e.g., a social graph); locating friends in a social graph; and opening application events (e.g., related to the messaging client 902).

[0107] The application server 914 hosts several server applications and subsystems, including, for example, a messaging server 912, an image processing server 918, and a social network server 922. The messaging server 912 implements several message processing technologies and functions, particularly those related to the aggregation and other processing of content (e.g., text and multimedia content) included in messages received from multiple instances of the messaging client 902. As will be described in further detail, text and media content from multiple sources can be aggregated into collections of content (e.g., referred to as stories or galleries). These collections are then made available to the messaging client 902. In view of the hardware requirements for other processor and memory intensive processing of data, such processing can also be performed on the server side by the messaging server 912.

[0108] The application server 914 also includes an image processing server 918 that is dedicated to performing various image processing operations, typically with respect to images or videos within the payload of messages sent from or received at the messaging server 912.

[0109] The social network server 922 supports various social networking functions and services and makes these functions and services available to the messaging server 912. To this end, the social network server 922 maintains and accesses an entity graph within the database 920. Examples of functions and services supported by the social network server 922 include identifying other users in the messaging system 900 with whom a particular user has relationships or that the particular user is "following", and also includes identifying interests and other entities of a particular user.

[0110] The messaging client 902 may notify the user of the client device 826 or other users associated with such a user (e.g., "friends") of activities occurring in a shared or shareable session. For example, the messaging client 902 may provide notifications related to current or recent use of a game by one or more members of a user group to participants in a conversation (e.g., a chat session) in the messaging client 902. One or more users may be invited to join an active session or initiate a new session. In some examples, a shared session may provide a shared augmented reality experience in which multiple people may collaborate or participate.

[0111] "Carrier signal" refers to any intangible medium that can store, encode or carry instructions for execution by a machine and includes digital or analog communications signals or other intangible media to facilitate communication of such instructions. Instructions may be sent or received over a network using a transmission medium via a network interface device.

[0112] "Client Device" refers to any machine that interfaces with a communications network to obtain resources from one or more server systems or other client devices. A client device may be, but is not limited to, a mobile phone, a desktop computer, a laptop computer, a portable digital assistant (PDA), a smart phone, a tablet computer, an ultrabook, a netbook, a laptop computer, a multiprocessor system, a microprocessor-based or programmable consumer electronics product, a game console, a set-top box, or any other communications device that a user may use to access a network.

[0113] "Communications network" means one or more parts of a network, which may be an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless WAN (WWAN), a metropolitan area network (MAN), the Internet, a part of the Internet, a part of the public switched telephone network (PSTN), a plain old telephone service (POTS) network, a cellular telephone network, a wireless network, The coupling may be a network, another type of network, or a combination of two or more such networks. For example, the network or a portion of the network may include a wireless network or a cellular network, and the coupling may be a code division multiple access (CDMA) connection, a global system for mobile communications (GSM) connection, or other type of cellular or wireless coupling. In this example, the coupling may implement any of various types of data transmission technologies, such as single carrier radio transmission technology (1xRTT), evolution data optimized (EVDO) technology, general packet radio service (GPRS) technology, enhanced data rates for GSM evolution (EDGE) technology, the third generation partnership project (3GPP) including 3G, fourth generation wireless (4G) network, universal mobile telecommunications system (UMTS), high speed packet access (HSPA), world wide interoperability for microwave access (WiMAX), long term evolution (LTE) standards, other data transmission technologies defined by various standard setting organizations, other long distance protocols, or other data transmission technologies.

[0114] "Component" refers to a device, physical entity or logic with boundaries defined by function or subroutine calls, branch points, APIs or other technologies provided for partitioning or modularizing specific processing or control functions. Components can be combined with other components via their interfaces to perform machine processing. Components can be encapsulated functional hardware units designed to be used with other components and a part of a program that generally performs a specific function of related functions. Components can constitute software components (e.g., codes implemented on machine-readable media) or hardware components. "Hardware components" are tangible units that can perform certain operations and can be configured or arranged in a certain physical manner. In various examples, one or more computer systems (e.g., independent computer systems, client computer systems or server computer systems) or one or more hardware components (e.g., processors or processor groups) of a computer system can be configured by software (e.g., applications or application parts) to operate to perform hardware components of some operations as described herein. Hardware components can also be implemented mechanically, electronically or in any suitable combination thereof. For example, a hardware component may include a dedicated circuit system or logic that is permanently configured to perform some operations. The hardware component may be a dedicated processor, such as a field programmable gate array (FPGA) or an application specific integrated circuit (ASIC). The hardware component may also include a programmable logic or circuit system that is temporarily configured to perform some operations by software. For example, the hardware component may include software executed by a general-purpose processor or other programmable processor. Once configured by such software, the hardware component becomes a specific machine (or a specific component of a machine) that is customized to perform the configured function, and is no longer a general-purpose processor. It will be understood that the decision to mechanically implement the hardware component in a dedicated and permanently configured circuit system or in a temporarily configured (e.g., configured by software) circuit system can be driven due to cost and time considerations. Accordingly, the phrase "hardware component" (or "hardware-implemented component") is to be understood as containing a tangible entity, that is, an entity that is physically constructed, permanently configured (e.g., hardwired) or temporarily configured (e.g., programmed) to operate in a particular manner or perform some operations described herein. Considering an example where a hardware component is temporarily configured (e.g., programmed), the hardware component may not be configured or instantiated at any time. For example, where a hardware component includes a general purpose processor that is configured by software to be a special purpose processor, the general purpose processor may be configured as different special purpose processors (e.g., including different hardware components) at different times. The software configures one or more specific processors accordingly, such as to constitute a specific hardware component at one time and to constitute different hardware components at different times. Hardware components may provide information to other hardware components and receive information from other hardware components. Therefore, the described hardware components may be considered to be communicatively coupled.In the case of multiple hardware components being present at the same time, communication can be realized by signal transmission (e.g., by appropriate circuits and buses) between two or more hardware components or among two or more hardware components. In the example where multiple hardware components are configured or instantiated at different times, communication between such hardware components can be realized, for example, by storing information in a memory structure accessible to multiple hardware components and retrieving information in the memory structure. For example, a hardware component can perform an operation, and the output of the operation is stored in a memory device coupled to it in communication. Then, other hardware components can access the memory device at a subsequent time to retrieve the stored output and process it. The hardware component can also initiate communication with an input device or an output device, and can operate on resources (e.g., the collection of information). The various operations of the example methods described herein can be performed by a temporary configuration (e.g., by software) or a permanent configuration to perform one or more processors of the related operation. Whether it is a temporary configuration or a permanent configuration, such a processor can constitute a processor-implemented component that performs operations to perform one or more operations or functions described herein. As used herein, "processor-implemented components" refer to hardware components implemented using one or more processors. Similarly, the method described herein can be implemented in part by a processor, wherein one or more specific processors are examples of hardware. For example, at least some operations in the operation of the method can be performed by one or more processors or the parts implemented by the processor. In addition, one or more processors can also operate to support the execution of related operations in the "cloud computing" environment or operate as "software as a service" (SaaS). For example, at least some operations in the operation can be performed by a group of computers (as an example of a machine including a processor), wherein these operations can be accessed via a network (for example, the Internet) and via one or more appropriate interfaces (for example, API). The execution of some operations in the operation can be distributed between processors, not only resides in a single machine, but also deployed across multiple machines. In some examples, a processor or a part implemented by a processor can be located in a single geographical location (for example, in a home environment, an office environment or a server group). In other examples, a processor or a part implemented by a processor can be distributed across multiple geographical locations.

[0115] "Computer-readable media" refers to both machine storage media and transmission media. Therefore, these terms include both storage devices / media and carrier / modulated data signals. The terms "machine-readable medium", "computer-readable medium" and "device-readable medium" mean the same thing and can be used interchangeably in this disclosure.

[0116] "Machine storage media" refers to a single or multiple storage devices and / or media (e.g., centralized or distributed databases, and / or associated caches and servers) that store executable instructions, routines, and / or data. The term includes, but is not limited to, solid-state memory and optical and magnetic media, including memory internal or external to the processor. Specific examples of machine storage media, computer storage media, and / or device storage media include: non-volatile memory, including, for example, semiconductor memory devices such as erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), FPGA, and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The terms "machine storage media," "device storage media," and "computer storage media" mean the same thing and may be used interchangeably in this disclosure. The terms "machine storage media," "computer storage media," and "device storage media" expressly exclude carrier waves, modulated data signals, and other such media, at least some of which are encompassed by the term "signal media."

[0117] A "processor" refers to any circuit or virtual circuit (a physical circuit simulated by logic executed on an actual processor) that manipulates data values ​​according to control signals (e.g., "commands," "opcodes," "machine codes," etc.) and produces associated output signals that are applied to operate a machine. For example, a processor may be a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a radio frequency integrated circuit (RFIC), or any combination thereof. A processor may also be a multi-core processor having two or more independent processors (sometimes referred to as "cores") that can execute instructions simultaneously.

[0118] "Signal medium" refers to any intangible medium that is capable of storing, encoding or carrying instructions to be executed by a machine, and "signal medium" includes digital or analog communication signals or other intangible media to facilitate the communication of software or data. The term "signal medium" should be deemed to include any form of modulated data signal, carrier wave, etc. The term "modulated data signal" means a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal. The terms "transmission medium" and "signal medium" mean the same thing and may be used interchangeably in this disclosure.

[0119] Changes and modifications may be made to the disclosed examples without departing from the scope of the present disclosure. These and other changes or modifications are intended to be included within the scope of the present disclosure as expressed in the appended claims.

Claims

1. A computer-implemented method comprising: capturing, by one or more processors, tracking video frame data of a hand of a user of an augmented reality (AR) system using a first camera of the AR system; generating, by the one or more processors, a skeletal model based on the tracking video frame data; determining, by the one or more processors, a position of a hand of the user based on the skeletal model; generating, by the one or more processors, camera manipulation command data based on the position of the user's hand; as well as A steerable camera of the AR system is focused, by the one or more processors, on a hand of the user.

2. The method according to claim 1, wherein: The first camera comprises a non-steerable camera having a camera field of view (FOV), wherein the camera field of view (FOV) is equal to an AR FOV of the AR system.

3. The method according to claim 1, wherein: The first camera device is the steerable camera device, and the method further comprises: scanning, by the one or more processors, an AR FOV of the AR system using the steerable camera to initially locate a hand of the user; predicting, by the one or more processors, a next position of the user's hand based on the current position of the hand using a look-ahead process; and The steerable camera is focused, by the one or more processors, on the next position.

4. The method according to claim 3, wherein: The next position is also predicted based on a language model.

5. The method according to claim 4, wherein: The language model is for American Sign Language.

6. The method according to claim 1, wherein: The controllable camera device also includes: A camera device having a sensor and a lens assembly; and One or more actuators connected to the camera device.

7. The method according to claim 1, wherein: The AR system includes a head-mounted device.

8. A computing device, comprising: one or more processors; as well as a memory storing instructions that, when executed by one or more processors, cause the computing device to perform operations comprising: capturing, using a first camera device of the AR system, tracking video frame data of a hand of a user of the AR system; Generate a skeleton model based on the tracking video frame data; Determining a position of a hand of the user based on the skeletal model; generating camera manipulation command data based on the position of the user's hand; and Focusing the steerable camera of the AR system on the user's hand.

9. The computing device according to claim 8, wherein: The first camera comprises a non-steerable camera having a camera FOV equal to an AR FOV of the AR system.

10. The computing device according to claim 8, wherein: The first camera is the steerable camera, and Wherein, when the instructions are executed by the one or more processors, the computing device is further caused to perform operations, the operations comprising: Scanning the AR FOV of the AR system using the steerable camera to initially locate the user's hand; using a look-ahead process to predict a next position of the user's hand based on the current position of the hand; and The steerable camera is focused on the next position.

11. The computing device according to claim 10, wherein: The next position is also predicted based on a language model.

12. The computing device according to claim 11, wherein: The language model is for American Sign Language.

13. The computing device of claim 8, wherein: The controllable camera device comprises: A camera device having a sensor and a lens assembly; and One or more actuators connected to the camera device.

14. The computing device of claim 8, wherein: The AR system includes a head-mounted device.

15. A non-transitory computer-readable storage medium comprising instructions that, when executed by a computing device, cause the computing device to perform operations comprising: capturing, using a first camera device of the AR system, tracking video frame data of a hand of a user of the AR system; Generate a skeleton model based on the tracking video frame data; Determining a position of a hand of the user based on the skeletal model; generating camera device manipulation command data based on the position of the user's hand; as well as Focusing the steerable camera of the AR system on the user's hand.

16. The non-transitory computer readable storage medium of claim 15, wherein: The first camera comprises a non-steerable camera having a camera FOV equal to an AR FOV of the AR system.

17. The non-transitory computer-readable storage medium of claim 15, wherein: The first camera is the steerable camera, and Wherein, when the instruction is executed by the computing device, the computing device is further caused to perform an operation, the operation comprising: Scanning the AR FOV of the AR system using the steerable camera to initially locate the user's hand; using a look-ahead process to predict a next position of the user's hand based on the current position of the hand; and The steerable camera is focused on the next position.

18. The non-transitory computer readable storage medium of claim 17, wherein: The next position is also predicted based on a language model.

19. The non-transitory computer readable storage medium of claim 18, wherein: The language model is for American Sign Language.

20. The non-transitory computer readable storage medium of claim 15, wherein: The controllable camera device comprises: A camera device having a sensor and a lens assembly; and One or more actuators connected to the camera device.

21. The non-transitory computer readable storage medium of claim 15, wherein: The AR system includes a head-mounted device.

Citation Information

Cited By

  • Steerable camera for AR hand tracking

    US12710828B2