Sign language interpretation using cooperative agents

KR103024487B1Active Publication Date: 2026-09-29SNAP INC
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
KR1020257016095
Authority / Receiving Office
KR · KR
Patent Type
Patents
Current Assignee / Owner
Priority Date
2022-10-17
Filing Date
2023-10-16
Publication Date
2026-09-29
Estimated Expiration
2043-10-16

Smart Images

  • Figure 112025054772769-PCT00003_ABST
    Figure 112025054772769-PCT00003_ABST
Patent Text Reader

Abstract

A method for recognizing sign language using cooperative augmented reality devices is described. In one embodiment, the method comprises the steps of accessing a first image generated by a first augmented reality device and a second image generated by a second augmented reality device—the first image and the second image depict a hand gesture of a user of the first augmented reality device—synchronizing the first augmented reality device with the second augmented reality device, distributing one or more processes of a sign language recognition system between the first and second augmented reality devices in response to the synchronization, collecting results from one or more processes from the first and second augmented reality devices, and displaying text indicating a sign language translation of a hand gesture on a first display of the first augmented reality device in near real-time based on the results.
Need to check novelty before this filing date? Find Prior Art

Description

Technology Field

[0001] Claim of priority

[0002] This application claims the benefit of priority of U.S. Patent Application No. 17 / 967,209 filed October 17, 2022, the whole of which is incorporated herein by reference.

[0003] Technology field

[0004] The subject matter disclosed herein generally relates to augmented reality (AR) devices. Specifically, the present disclosure covers systems and methods for sign language interpretation using collaborative agents. Background Technology

[0005] Augmented reality (AR) devices allow users to observe a scene while simultaneously viewing relevant virtual content that can be aligned with items, images, objects, or environments within the device's field of view. Virtual reality (VR) devices provide a more immersive experience than AR devices. VR devices block the user's field of view with virtual content displayed based on the device's position and orientation. Brief explanation of the drawing

[0006] To facilitate the identification of a discussion of any specific element or action, the top digit or numbers in the reference number refer to the drawing number where the element is first introduced. FIG. 1 is a block diagram illustrating a network environment for sign language translation using collaborating augmented reality devices according to one exemplary embodiment. FIG. 2 is a block diagram illustrating a network environment for sign language translation using cooperative augmented reality devices according to one exemplary embodiment. FIG. 3 is a block diagram illustrating an augmented reality device according to one exemplary embodiment. FIG. 4 is a block diagram illustrating a cooperative hydration system according to one exemplary embodiment. FIG. 5 is an interaction diagram illustrating interactions between cooperative augmented reality devices according to one exemplary embodiment. FIG. 6 is an interaction diagram illustrating interactions between cooperative augmented reality devices according to one exemplary embodiment. FIG. 7 is an interaction diagram illustrating interactions between cooperative augmented reality devices according to one exemplary embodiment. FIG. 8 is an interaction diagram illustrating interactions between cooperative augmented reality devices according to one exemplary embodiment. FIG. 9 is a flowchart illustrating a method for sign language translation according to one exemplary embodiment. FIG. 10 is a flowchart illustrating a method for distributing sign language translation processes according to one exemplary embodiment. FIG. 11 illustrates a network environment in which a head-wearable device can be implemented according to one exemplary embodiment. FIG. 12 is a block diagram illustrating a software architecture in which the present disclosure can be implemented according to an exemplary embodiment. FIG. 13 is a schematic representation of a machine in the form of a computer system in which a set of instructions can be executed to enable the machine to perform any one or more of the methodologies discussed in this specification, according to one exemplary embodiment. Specific details for implementing the invention

[0007] The following description describes systems, methods, techniques, instruction sequences, and computing machine program products that illustrate exemplary embodiments of the subject matter. In the following description, for the purposes of explanation and to provide an understanding of various embodiments of the subject matter, numerous specific details are presented. However, it will be apparent to those skilled in the art that embodiments of the subject matter may be practiced without some or all of these specific details. Examples are merely representations of possible variations. Unless otherwise explicitly stated, structures (structural components, such as modules) are optional and may be combined or subdivided, and operations (e.g., in procedures, algorithms, or other functions) may be varied in sequence, combined, or subdivided.

[0008] The term "Augmented Reality (AR)" is used herein to refer to an interactive experience of a real-world environment in which physical objects existing in the real world are "augmented" or enhanced by computer-generated digital content (also referred to as virtual content or synthetic content). AR may also refer to a system that enables a combination of real and virtual worlds, real-time interaction, and 3D registration of virtual and real objects. A user of an AR system perceives virtual content that appears to be attached to or interacting with real-world physical objects.

[0009] The term "Virtual Reality" (VR) is used herein to refer to a simulated experience of a virtual world environment that is completely distinct from the real world environment. Computer-generated digital content is displayed in the virtual world environment. VR also refers to a system that enables a user of the VR system to be fully immersed in the virtual world environment and to interact with virtual objects presented in the virtual world environment.

[0010] The term "AR application" is used herein to refer to a computer-motion application that enables an AR experience. The term "VR application" is used herein to refer to a computer-motion application that enables a VR experience. The term "AR / VR application" refers to a computer-motion application that enables an AR experience or a combination of VR experiences.

[0011] The term “visual tracking system” is used herein to refer to a computer-operation application or system that enables the system to track visual features identified in images captured by one or more cameras of the visual tracking system. The visual tracking system constructs a model of the real-world environment based on the tracked visual features. Non-limiting examples of visual tracking systems include: a visual simultaneous localization and mapping system (VSLAM), and a visual inertial odometry (VIO) system. VSLAM may be used to construct targets from an environment or scene based on one or more cameras of the visual tracking system. A VIO system (also referred to as a visual-inertial tracking system) determines the latest pose (e.g., position and orientation) of the device based on data acquired from a number of sensors of the device (e.g., optical sensors, inertial sensors).

[0012] The term "Inertial Measurement Unit" (IMU) is used herein to refer to a device capable of reporting the inertial state of a moving object, including its acceleration, velocity, orientation, and position. The IMU enables the tracking of the object's motion by integrating the acceleration and angular velocity measured by the IMU. The IMU may also refer to a combination of accelerometers and gyroscopes capable of determining and quantifying linear acceleration and angular velocity, respectively. Values ​​obtained from the IMU gyroscopes can be processed to obtain the pitch, roll, and heading of the IMU and thus the pitch, roll, and heading of the object associated with the IMU. Signals from the IMU's accelerometers can also be processed to obtain the velocity and displacement of the IMU.

[0013] The term "three-degrees of freedom tracking system" (3DOF tracking system) is used herein to refer to a device that tracks rotational movement. For example, a 3DOF tracking system can track whether a user of a head-wearable device is looking to the left or right, rotating their head up or down, or pivoting to the left or right. However, a head-wearable device cannot use a 3DOF tracking system to determine whether the user has moved around the scene by moving in the physical world. As such, a 3DOF tracking system may not be accurate enough to be used for position signals. A 3DOF tracking system may be part of an AR / VR display device that includes IMU sensors. For example, a 3DOF tracking system uses sensor data from sensors such as accelerometers, gyroscopes, and magnetometers.

[0014] The term "six-degrees of freedom tracking system" (6DOF tracking system) is used herein to refer to a device that tracks rotational and translational motion. For example, a 6DOF tracking system can track whether a user rotates their head and moves forward or backward, sideways or vertically, and up or down. A 6DOF tracking system may include a Simultaneous Localization and Mapping (SLAM) system and / or a VIO system that relies on data acquired from a number of sensors (e.g., depth cameras, inertial sensors). A 6DOF tracking system analyzes data from the sensors to accurately determine the pose of the display device.

[0015] The present application describes a system for translating sign language using cooperative augmented reality devices. For example, a first user wears a head-wearable AR device (e.g., smart glasses) and performs hand gestures (e.g., sign language). A second user facing the first user may also wear another AR device or point the camera of their mobile communication device (e.g., a smartphone) toward the first user's hands. Cameras from the AR devices of the first and second users capture the first user's hand gestures from different perspectives or angles. The first AR device is first synchronized with the second AR device. For example, synchronization may be based on timestamps of images generated by the cameras of the first and second AR devices. In another example, synchronization may be based on the poses of the AR devices in the same aligned coordinate system. One or more processes of the sign language translation system are distributed between the first and second AR devices. For example, the first AR device may share images of hand gestures captured by the camera of the first AR device with the second AR device. The second AR device may have more computational resources (e.g., a smartphone may have a faster processor, larger memory, and a larger battery than smart glasses) and performs hand tracking, gesture recognition, and sign language translation. In one example, the distribution of processes in the sign language translation system may be based on a comparison of the specifications of the first AR device and the second AR device and other conditions (e.g., detection of hands being blocked from the cameras of the first AR device; the system will rely on images from the cameras of the second AR device when hands are blocked from the cameras of the first AR device).

[0016] In other exemplary embodiments, one AR device performs hand tracking and gesture recognition, and another AR device performs gesture recognition. In other examples, the results of the processes (sign language translation, hand skeletal detection, semantic gesture recognition) are shared between AR devices temporally (e.g., one every two frames) or spatially (e.g., left hand / right hand only). In another example, the first AR device displays context information, user historical inputs, a text corpus, and suggested words / phrases based on predefined terms. Upon detection of a triggering condition (operation of a specific gesture or user input), the first user may select suggested options to avoid finger spelling or accelerate interpretation.

[0017] In one exemplary embodiment, a method for recognizing sign language using cooperative AR devices is described. In one aspect, the method comprises the steps of accessing a first image generated by a first camera of a first augmented reality device and a second image generated by a second camera of a second augmented reality device—the first image and the second image depict a hand gesture of a user of the first augmented reality device—synchronizing the first augmented reality device with the second augmented reality device, distributing one or more processes of a sign language recognition system between the first augmented reality device and the second augmented reality device in response to the synchronization—one or more processes are performed on the corresponding augmented reality devices—collecting results from one or more processes from the first augmented reality device and from the second augmented reality device, and displaying text displaying a sign language translation of a hand gesture based on the results in near real-time on a first display of the first augmented reality device or a second display of the second augmented reality device.

[0018] As a result, one or more of the methodologies described herein facilitate the resolution of technical problems regarding resource management for sign language translation. The method described herein provides improvements to the operation of computer functions by distributing sign language processes among synchronized cooperative AR devices to provide reduced power consumption. As such, one or more of the methodologies described herein can eliminate the need for specific efforts or computing resources. Examples of such computing resources include processor cycles, network traffic, memory usage, data storage capacity, power consumption, network bandwidth, and cooling capacity.

[0019] FIG. 1 is a block diagram illustrating a network environment (100) for sign language translation using cooperative AR devices according to one exemplary embodiment. The network environment (100) includes AR device B (114), AR device A (110), and a server (116) coupled to communicate with each other via a network (104). AR device B (114), AR device A (110), and the server (116) may each be implemented in a computer system, wholly or partially, as described below in relation to FIG. 13. The server (116) may be part of a network-based system. For example, the network-based system may be or may include a cloud-based server system that provides additional information, such as reference frame alignment data of AR device B (114) and AR device A (110).

[0020] A user (106) wears an AR device A (110) (e.g., smart glasses) and operates the AR device A (110) through voice, hand gestures, and physical input (e.g., touching the surface of the AR device A (110) or pressing a button on the AR device A (110)). The user (106) performs hand gestures with his / her hands (108). The AR device A (110) captures images of the hands (108) using a camera (not shown) of the AR device A (110). The AR device A (110) may be a computing device having a display. The computing device may be detachably mounted on the user's (106) head. In one example, the display may be a screen that displays what is captured by the camera of the AR device B (114) / AR device A (110). In another example, the display of the device may be transparent, as in the lenses of wearable computing glasses, so that the user (106) can see content presented on the display while simultaneously seeing real-world objects visible through the display.

[0021] A user (112) operates an AR device B (114). The user (112) may be a human user (e.g., a human), a machine user (e.g., a computer configured by a software program to interact with the AR device B (114)), or any suitable combination thereof (e.g., a human assisted by a machine or a machine supervised by a human). In one example, the user (112) aims the camera of the AR device B (114) at the hands (108) of the user (106) in the real-world environment (102). The AR device B (114) may be a computing device having a display, such as a smartphone, a tablet computer, or a wearable computing device (e.g., a watch or glasses). The computing device may be handheld or detachably mounted on the head of the user (112). In one example, the display may be a screen that displays what is captured by the camera of the AR device B (114) / AR device A (110). In another example, the display of the device may be transparent, as in the lenses of wearable computing glasses, so that the user (112) can see content presented on the display while simultaneously seeing real-world objects visible through the display. In another example, the AR device B (114) includes a computing device (e.g., a mobile device capable of displaying text or playing audio).

[0022] In one example, AR device A (110) and AR device B (114) each generate images of the hands (108) from different viewpoints. As such, the images captured by AR device B (114) and AR device A (110) include overlapping regions.

[0023] AR device A (110) includes a tracking system (not shown). The tracking system tracks the pose (e.g., position and orientation) of AR device A (110) relative to the real-world environment (102) using optical sensors (e.g., image camera), inertial sensors (e.g., gyroscope, accelerometer), wireless sensors (Bluetooth, Wi-Fi), GPS sensors, and audio sensors to determine the location of AR device A (110) within the real-world environment (102). For example, the tracking system of AR device A (110) includes a 6DOF tracking system that tracks whether the user (106) has rotated their head, moved forward or backward, sideways or vertically, and up or down.

[0024] AR device B (114) also includes a tracking system (not shown). The tracking system uses optical sensors (e.g., image cameras), inertial sensors (e.g., gyroscopes, accelerometers), wireless sensors (Bluetooth, Wi-Fi), GPS sensors, and audio sensors to determine the location of AR device B (114) within the real-world environment (102) and tracks the pose (e.g., position and orientation) of AR device B (114) relative to the real-world environment (102). For example, the tracking system of AR device B (114) includes a 6DOF tracking system that tracks whether the user (112) (in the case of a head-wearable device) has rotated their head, moved forward or backward, sideways or vertically, and up or down.

[0025] In one exemplary embodiment, the server (116) coordinates the distribution of sign language translation between AR device B (114) and AR device A (110). For example, the server (116) identifies available computational resources and technical specifications of AR device B (114) and AR device A (110), and assigns one or more processes from sign language translation to AR device B (114) and AR device A (110) based on the available computational resources and technical specifications. In another example, the server (116) receives image data, image metadata (e.g., timestamps), and pose information from both AR device B (114) and AR device A (110), synchronizes the received information, performs sign language translation based on the synchronized information, and provides translation to both AR device B (114) and AR device A (110). In another example, the server (116) first provides the proposed translation to the AR device (A) (110) and waits for the user (106) to verify the proposed translation before sending the verified translation to the AR device B (114).

[0026] Any of the machines, databases, or devices illustrated in FIG. 1 may be implemented on a general-purpose computer modified (e.g., configured or programmed) by software to become a special-purpose computer for performing one or more of the functions described herein for that machine, database, or device. For example, a computer system capable of implementing any one or more of the methodologies described herein is discussed below in relation to FIG. 9. As used herein, "database" is a data storage resource and may store structured data as a text file, table, spreadsheet, relational database (e.g., object-relational database), triple store, hierarchical data store, or any suitable combination thereof. Additionally, any two or more of the machines, databases, or devices illustrated in FIG. 1 may be combined into a single machine, and the functions described herein for any single machine, database, or device may be subdivided among multiple machines, databases, or devices.

[0027] The network (104) may be any network that enables communication between machines (e.g., server (116)), databases, and devices (e.g., AR device B (114), AR device A (110)). Accordingly, the network (104) may be a wired network, a wireless network (e.g., a mobile or cellular network), or any suitable combination thereof. The network (104) may include one or more parts constituting a private network, a public network (e.g., the Internet), or any suitable combination thereof.

[0028] FIG. 2 is a network diagram illustrating a network environment (200) suitable for operating AR device B (114) and AR device A (110) according to some exemplary embodiments. The network environment (200) includes AR device B (114) and AR device A (110). AR device B (114) can communicate wirelessly (e.g., Bluetooth) with AR device A (110). In contrast to FIG. 1, AR device B (114) and AR device A (110) communicate directly with each other.

[0029] AR device A (110) and AR device B (114) can share data with each other. In one example, AR device A (110) receives image data and pose data from AR device B (114), and synchronizes AR device A (110) with AR device B (114) based on the data received from AR device B (114) and the image data and pose data of AR device A (110). In another example, AR device B (114) receives image data and pose data from AR device B (114), and synchronizes AR device B (114) with AR device A (110) based on the data received from AR device A (110) and the image data and pose data of AR device B (114).

[0030] FIG. 3 is a block diagram illustrating modules (e.g., components) of an AR device A (110) according to some exemplary embodiments. The AR device A (110) includes sensors (302), a display (304), a processor (308), a graphics processing unit (318), a display controller (320), and a storage device (306). Examples of the AR device A (110) include a wearable computing device. The AR device B (114) includes components similar to those of the AR device A (110). However, the AR device B (114) may include a wearable computing device, a tablet computer, or a smartphone.

[0031] The sensors (302) include optical sensors (314) and inertial sensors (316). The optical sensors (314) include stereo cameras. The inertial sensors (316) include a combination of gyroscopes, accelerometers, and magnetometers. Other examples of sensors (302) include proximity or location sensors (e.g., proximity communication, GPS, Bluetooth, Wifi), audio sensors (e.g., microphones), or any suitable combination thereof. It should be noted that the sensors (302) described herein are for illustrative purposes only and are therefore not limited to those described above. In one exemplary embodiment, the sensors (302) may include a structured optical sensor, a time-of-flight sensor, a passive stereo sensor, and a depth sensor such as an ultrasonic device, a time-of-flight sensor.

[0032] The display (304) includes a screen or monitor configured to display images generated by the processor (308). In one exemplary embodiment, the display (304) may be transparent or translucent so that the user (106) can view it through the display (304) (in the case of AR use). In another example, the display (304), such as an LCOS display, presents each frame of virtual content in a number of presentations.

[0033] The processor (308) includes an AR application (310), a 6DOF tracker (312), a depth system (324), and a cooperative sign language system (328). The AR application (310) uses computer vision to detect and identify hands (108) or items in a real-world environment (102). The AR application (310) searches for a virtual object (e.g., a 3D object model) based on the identified item (e.g., hands (108)) or physical environment. A display (304) displays the virtual object. The AR application (310) includes a local rendering engine that generates a visualization of the virtual object overlaid (e.g., superimposed, or otherwise displayed together with) an image of the item captured by an optical sensor (314). The visualization of the virtual object can be manipulated by adjusting the position of the item relative to the optical sensor (314) (e.g., its physical location, orientation, or both). Similarly, the visualization of a virtual object can be manipulated by adjusting the pose of the AR device A (110) relative to the item.

[0034] The 6DOF tracker (312) estimates the pose of AR device A (110). For example, the 6DOF tracker (312) tracks the location and pose of AR device A (110) relative to a reference frame (e.g., real-world environment (102)) using image data from an optical sensor (314) and an inertial sensor (316) and corresponding inertial data. In one example, the 6DOF tracker (312) determines the 3D pose of AR device A (110) using sensor data. The 3D pose is the determined orientation and position of AR device A (110) relative to the user's real-world environment (102). For example, AR device A (110) can identify the relative position and orientation of AR device A (110) from physical objects in the real-world environment (102) surrounding AR device A (110) using images of the user's real-world environment (102) as well as other sensor data. The 6DOF tracker (312) continuously collects and uses updated sensor data describing the movements of AR device A (110) to determine updated 3D poses of AR device A (110) that indicate changes in the relative position and orientation of AR device A (110) from physical objects in the real-world environment (102). The 6DOF tracker (312) provides the 3D poses of AR device A (110) to the depth system (324) and the cooperative sign language system (328).

[0035] The depth system (324) accesses images from the optical sensor (314) and sparse 3D points from the 6DOF tracker (312) to predict depths and generate a dense point cloud. In one exemplary embodiment, the optical sensor (314) includes a depth sensor or a stereo sensor.

[0036] The cooperative sign language system (328) coordinates the distribution of processes for a sign language translation application between AR device B (114) and AR device A (110) (and optionally a server (116)). For example, the cooperative sign language system (328) identifies available computational resources and technical specifications of AR device B (114) and AR device A (110), and assigns one or more processes from sign language translation between AR device B (114) and AR device A (110) based on the available computational resources and technical specifications. In another example, AR device A (110) receives image data, image metadata (e.g., timestamps), and pose information from AR device B (114), uses the received information to synchronize AR device A (110) with AR device B (114), performs sign language translation based on the synchronized information, displays the sign language translation on the display (304), and provides the sign language translation to AR device B (114). In another example, the cooperative sign language system (328) displays a proposed translation on a display (304), waits until the user (106) confirms the proposed translation, detects confirmation by detecting a predefined gesture that displays a sign language proposal confirmation gesture, and transmits the confirmed translation to AR device B (114).

[0037] The graphics processing unit (318) includes a render engine (not shown) configured to render frames of a 3D model of a virtual object based on virtual content provided by the AR application (310) and the pose of AR device B (114) (relative to AR device A (110)). That is, the graphics processing unit (318) generates frames of virtual content to be presented on the display (304) using the 3D pose of AR device B (114). For example, the graphics processing unit (318) renders frames of virtual content using the 3D pose so that the virtual content is presented in an orientation and position within the display (304) to appropriately augment the user's reality. For example, the graphics processing unit (318) may use 3D pose data to render frames of virtual content so that when presented on the display (304), the virtual content overlaps with physical objects in the user's real-world environment (102). The graphics processing unit (318) generates updated frames of virtual content based on the updated three-dimensional poses of the AR device B (114), which reflect changes in the user's position and orientation regarding physical objects in the user's real-world environment (102).

[0038] The graphics processing unit (318) transmits the rendered frame to the display controller (320). The display controller (320) is positioned as an intermediary between the graphics processing unit (318) and the display (304), receives image data (e.g., the rendered frame) from the graphics processing unit (318), and provides the rendered frame to the display (304).

[0039] The storage device (306) stores virtual object content (322) and gesture data (326) (e.g., a predefined gesture indicating a suggested word confirmation). The virtual object content (322) includes, for example, a database of visual references (e.g., images, QR codes) and corresponding virtual content (e.g., a three-dimensional model of virtual objects). The gesture data (326) displays predefined hand gestures or predefined user inputs to indicate confirmation of suggested words translated by a cooperative sign language system (328) or by another translation system (located on AR device B (114) or server (116)). The gesture data (326) may also define other triggering conditions that allow the user (106) to select suggested options to avoid finger spelling or accelerate sign language interpretation when detected.

[0040] Any one or more of the modules described herein may be implemented using hardware (e.g., a processor of a machine) or a combination of hardware and software. For example, any module described herein may be configured with a processor to perform the operations described herein for that module. Furthermore, any two or more of these modules may be combined into a single module, and the functions described herein for a single module may be subdivided among multiple modules. Additionally, according to various exemplary embodiments, modules described herein as being implemented within a single machine, database, or device may be distributed across multiple machines, databases, or devices.

[0041] FIG. 4 is a block diagram illustrating a cooperative sign language system (328) according to one exemplary embodiment. The cooperative sign language system (328) includes an AR device interface (402), a synchronization module (404), a hand tracking application (406), and a sign language recognition application (408).

[0042] The AR device interface (402) communicates with the AR device B (114). In one example, the AR device interface (402) accesses 6DoF pose data and camera data from the AR device B (114). In another example, the AR device interface (402) shares the 6DoF pose data of the 6DOF tracker (312) and the camera data of the optical sensor (314) with the AR device B (114). In another example, the AR device interface (402) communicates which processes the AR device B (114) will perform as part of a cooperative sign language translation process.

[0043] The synchronization module (404) accesses 6DoF pose data from the 6DOF tracker (312), camera data from the optical sensor (314), and 6DoF pose data and camera data (including timestamp metadata) from AR device B (114). In another example, the synchronization module (404) accesses the technical specifications and available computational resources of AR device A (110) and AR device B (114). The synchronization module (404) synchronizes AR device A (110) with AR device B (114) by matching the timestamps of images from AR device A (110) with images from AR device B (114). In another example, synchronization may be performed based on aligning / matching the 6DoF poses of the cameras of AR device A (110) and AR device B (114) in the same common coordinate system.

[0044] A hand tracking application (406) detects hands (108) using computer vision. In one example, the hand tracking application (406) identifies features (motion, skeleton, depth) of the hands and tracks the motion of the hands (108). The hand tracking application (406) may include a number of processes, such as, for example, a hand tracking process, a hand gesture detection process, and a sign language translation process.

[0045] A sign language recognition application (408) identifies gestures of hands (108) based on a hand tracking application (406) and identifies words corresponding to the gestures. In one example, the sign language recognition application (408) includes a sign language translation process. In another exemplary embodiment, the sign language recognition application (408) identifies a proposed word based on a combination of context information of AR device A (110) (e.g., user (106)'s profile, the geo-location of AR device A (110)), user (106)'s history inputs, a text corpus, and predefined terms. The sign language recognition application (408) provides text (representing the proposed word) to a display (304). The display (304) displays the text. In another example, the sign language recognition application (408) communicates the text to AR device B (114), where the text is displayed on the screen of AR device B (114).

[0046] FIG. 5 is an interaction diagram illustrating interactions between cooperative augmented reality devices (e.g., AR device A (110) and AR device B (114)) according to one exemplary embodiment. AR device B (114) captures images of hands (108) with its camera. AR device A (110) captures images of hands (108) with its camera. In 502, AR device B (114) transmits images, camera data, and 6DoF pose data to AR device A (110). In 504, AR device A (110) synchronizes AR device A (110) and AR device B (114) based on image data from both AR device A (110) and AR device B (114). In 506, AR device A (110) applies a hand tracking process based on synchronized images (e.g., a combination of image data from AR device A (110) and AR device B (114)). In 508, AR device A (110) applies a gesture recognition process to the synchronized images to generate a sign language translation. In 510, AR device A (110) displays the sign language translation on the display (304) and provides the sign language translation to AR device B (114).

[0047] FIG. 6 is an interaction diagram illustrating interactions between cooperative augmented reality devices (e.g., AR device A (110) and AR device B (114)) according to one exemplary embodiment. AR device B (114) captures images of hands (108) with its camera. AR device A (110) captures images of hands (108) with its camera. In 602, AR device B (114) transmits images, camera data, and 6DoF pose data to AR device A (110). In 604, AR device A (110) synchronizes AR device A (110) and AR device B (114) based on image data from both AR device A (110) and AR device B (114). In 606, AR device A (110) applies a hand tracking process based on synchronized images (e.g., a combination of image data from AR device A (110) and AR device B (114)). In 608, AR device A (110) applies a gesture recognition process to the synchronized images to generate a proposed sign language translation. In 610, AR device A (110) displays the proposed translation to the user (106) on the display (304). In 612, AR device A (110) detects confirmation of the proposed translation for the user (106). In 614, AR device A (110) shares the translation with AR device B (114).

[0048] FIG. 7 is an interaction diagram illustrating interactions between cooperative augmented reality devices (e.g., AR device A (110) and AR device B (114)) according to one exemplary embodiment. AR device B (114) captures images of hands (108) with its camera. AR device A (110) captures images of hands (108) with its camera. In 702, AR device B (114) transmits images, camera data, and 6DoF pose data to AR device A (110). In 704, AR device A (110) synchronizes AR device A (110) and AR device B (114) based on image data from both AR device A (110) and AR device B (114). In 706, AR device A (110) applies a hand tracking process based on synchronized images (e.g., a combination of image data from AR device A (110) and AR device B (114)). AR device A (110) transmits the results of the hand tracking process to AR device B (114). In 708, AR device B (114) applies a gesture recognition process to the results of the hand tracking processed by AR device A (110). In 710, AR device B (114) applies a sign language recognition process to the results of the gesture recognition process. In 712, AR device B (114) displays a sign language translation (e.g., text) on the display of AR device B (114). In 714, AR device B (114) provides the text to AR device A (110). In 716, AR device A (110) displays the text.

[0049] FIG. 8 is an interaction diagram illustrating interactions between cooperative augmented reality devices (e.g., AR device A (110) and AR device B (114)) according to one exemplary embodiment. AR device B (114) captures images of hands (108) with its camera. AR device A (110) captures images of hands (108) with its camera. In 802, AR device B (114) transmits images, camera data, and 6DoF pose data to AR device A (110). In 804, AR device A (110) synchronizes AR device A (110) and AR device B (114) based on image data from both AR device A (110) and AR device B (114). In 806, AR device A (110) provides an egocentric view image to AR device B (114). In 808, AR device A (110) applies a hand tracking process based on synchronized images (e.g., a combination of image data from AR device B (114) and self-centered view data from AR device A (110)). In 810, AR device B (114) applies a gesture recognition process and a sign language recognition process to the results of the hand tracking process. In 812, AR device B (114) displays a sign language translation (e.g., text) on the display of AR device B (114). In 814, AR device B (114) provides text to AR device A (110) for display.

[0050] FIG. 9 is a flowchart illustrating a method (900) for sign language translation according to one exemplary embodiment. Operations in the method (900) may be performed by a cooperative sign language system (328) using the components (e.g., modules, engines) described above in relation to FIG. 4. Accordingly, the method (900) is described as an example with reference to the cooperative sign language system (328). However, it should be understood that at least some of the operations of the method (900) may be performed by similar components placed on various other hardware configurations or located elsewhere.

[0051] In block (902), the synchronization module (404) synchronizes the signer's AR device (e.g., AR device A (110)) with the receiver's AR device (e.g., AR device B (114)). In block (904), for the synchronized images, the hand tracking application (406) performs hand tracking, and the sign language recognition application (408) performs sign language recognition. In block (906), the display (304) displays the proposed translation based on the sign language recognition on the signer's AR device. In block (908), the sign language recognition application (408) detects confirmation of the proposed translation on the signer's AR device. In block (910), the sign language recognition application (408) provides the proposed translation to the receiver's AR device.

[0052] It should be noted that other embodiments may use different sequencing, additional or fewer operations, and different nomenclature or terminology to achieve similar functions. In some embodiments, various operations may be performed in parallel with other operations in a synchronous or asynchronous manner. The operations described herein have been selected to exemplify some principles of the operations in a simplified form.

[0053] FIG. 10 is a flowchart illustrating a method for distributing sign language translation processes according to one exemplary embodiment. Operations in the method (1000) may be performed by a cooperative sign language system (328) using the components (e.g., modules, engines) described above in relation to FIG. 4. Accordingly, the method (1000) is described as an example with reference to the cooperative sign language system (328). However, it should be understood that at least some of the operations of the method (1000) may be performed by similar components placed on various other hardware configurations or located elsewhere.

[0054] In block (1002), the cooperative sign language system (328) synchronizes the signer's AR device with the receiver's AR device. In block (1004), the cooperative sign language system (328) distributes one or more processes from hand tracking, gesture recognition, and sign language recognition between the signer's AR device and the receiver's AR device. In block (1006), the cooperative sign language system (328) shares the results of the processes between the signer's AR device and the receiver's AR device. In block (1008), the display (304) displays a translation based on the shared results.

[0055] It should be noted that other embodiments may use different sequencing, additional or fewer operations, and different nomenclature or terminology to achieve similar functions. In some embodiments, various operations may be performed in parallel with other operations in a synchronous or asynchronous manner. The operations described herein have been selected to exemplify some principles of the operations in a simplified form.

[0056] System having a head-wearable device

[0057] FIG. 11 illustrates a network environment (1100) in which a head-wearable device (1102) can be implemented according to one exemplary embodiment. FIG. 11 is a high-level functional block diagram of an exemplary head-wearable device (1102) coupled to a mobile client device (1138) and a server system (1132) so as to be communicable through various networks (1140).

[0058] The head-wearable device (1102) includes a camera such as at least one of a visible light camera (1112), an infrared emitter (1114), and an infrared camera (1116). A client device (1138) can connect to the head-wearable device (1102) using both communication (1134) and communication (1136). The client device (1138) is connected to a server system (1132) and a network (1140). The network (1140) may include any combination of wired and wireless connections.

[0059] The head-wearable device (1102) further includes two image displays of the image display (1104) of the optical assembly. These two include one associated with the left lateral side of the head-wearable device (1102) and one associated with the right lateral side. The head-wearable device (1102) also includes an image display driver (1108), an image processor (1110), a low-power circuit (1126), and a high-speed circuit (1118). The image display (1104) of the optical assembly is intended to present images and videos including images that may include a graphical user interface for the user of the head-wearable device (1102).

[0060] The image display driver (1108) commands and controls the image display of the image display (1104) of the optical assembly. The image display driver (1108) may directly transmit image data to the image display of the image display (1104) of the optical assembly for presentation, or may convert the image data into a signal or data format suitable for transmission to the image display device. For example, the image data may be video data formatted according to compression formats such as H. 264 (MPEG-4 Part 10), HEVC, Theora, Dirac, RealVideo RV40, VP8, VP9, ​​or similar, and the still image data may be formatted according to compression formats such as PNG (Portable Network Group), JPEG (Joint Photographic Experts Group), TIFF (Tagged Image File Format), or Exif (exchangeable image file format), or similar.

[0061] As mentioned above, the head-wearable device (1102) includes a frame and stems (or temples) extending from the lateral sides of the frame. The head-wearable device (1102) further includes a user input device (1106) (e.g., a touch sensor or a push button) comprising an input surface on the head-wearable device (1102). The user input device (1106) (e.g., a touch sensor or a push button) is intended to receive an input selection from the user for manipulating a graphic user interface of a presented image.

[0062] The components illustrated in FIG. 11 for the head-wearable device (1102) are located on one or more circuit boards, e.g., a PCB or a flexible PCB, within the rims or temples. Alternatively or additionally, the illustrated components may be located on chunks, frames, hinges, or bridges of the head-wearable device (1102). The left and right sides may include digital camera elements, such as a CMOS (complementary metal-oxide-semiconductor) image sensor, a charge coupling device, a camera lens, or any other respective visible or light capture elements that can be used to capture data including images of scenes having unknown objects.

[0063] The head-wearable device (1102) includes a memory (1122) for storing instructions for performing a subset or all of the functions described herein. The memory (1122) may also include a storage device.

[0064] As illustrated in FIG. 11, the high-speed circuit section (1118) includes a high-speed processor (1120), memory (1122), and a high-speed wireless circuit section (1124). In this example, an image display driver (1108) is coupled to the high-speed circuit section (1118) to drive the left and right image displays of the image display (1104) of the optical assembly and is operated by the high-speed processor (1120). The high-speed processor (1120) may be any processor capable of managing high-speed communications and operations of any general-purpose computing system required for the head-wearable device (1102). The high-speed processor (1120) includes processing resources necessary to manage high-speed data transmissions over communication (1136) to a wireless local area network (WLAN) using the high-speed wireless circuit section (1124). In certain examples, the high-speed processor (1120) runs an operating system such as a LINUX operating system or other such operating system of the head-wearable device (1102), and the operating system is stored in memory (1122) for execution. In addition to any other tasks, the high-speed processor (1120) running the software architecture for the head-wearable device (1102) is used to manage data transmissions with the high-speed wireless circuit (1124). In certain examples, the high-speed wireless circuit (1124) is configured to implement the IEEE (Institute of Electrical and Electronic Engineers) 1102.11 communication standards, also referred to herein as Wi-Fi. In other examples, other high-speed communication standards may be implemented by the high-speed wireless circuit (1124).

[0065] The low-power wireless circuit section (1130) and the high-speed wireless circuit section (1124) of the head-wearable device (1102) are short-range transceivers (Bluetooth TMIt may include wireless wide-area, local, or wide-area network transceivers (e.g., cellular or WiFi). A client device (1138) including transceivers communicating via communication (1134) and communication (1136) may be implemented using details of the architecture of the head-wearable device (1102), just as other elements of the network (1140) may be.

[0066] The memory (1122) comprises any storage device capable of storing various data and applications, including, among other things, camera data generated by the left and right infrared cameras (1116) and the image processor (1110), as well as images generated to be displayed by the image display driver (1108) on the image displays of the image display (1104) of the optical assembly. Although the memory (1122) is depicted as being integrated with the high-speed circuit (1118), in other examples, the memory (1122) may be an independent, standalone element of the head-wearable device (1102). In such specific examples, electrical routing lines may provide a connection through a chip including a high-speed processor (1120) from the image processor (1110) or the low-power processor (1128) to the memory (1122). In other examples, the high-speed processor (1120) can manage the addressing of the memory (1122) so that the low-power processor (1128) boots the high-speed processor (1120) whenever a read or write operation involving the memory (1122) is required.

[0067] As illustrated in FIG. 11, a low-power processor (1128) or a high-speed processor (1120) of a head-wearable device (1102) may be coupled to a camera (visible light camera (1112); an infrared emitter (1114), or an infrared camera (1116)), an image display driver (1108), a user input device (1106) (e.g., a touch sensor or a push button), and a memory (1122).

[0068] The head-wearable device (1102) is connected to a host computer. For example, the head-wearable device (1102) is paired with a client device (1138) via communication (1136) or connected to a server system (1132) via a network (1140). The server system (1132) may be one or more computing devices as part of a service or network computing system, for example, including a processor, memory, and a network communication interface for communicating with the client device (1138) and the head-wearable device (1102) over the network (1140).

[0069] The client device (1138) includes a processor and a network communication interface coupled to the processor. The network communication interface allows communication, communication (1134), or communication (1136) over the network (1140). To implement the functionality described herein, the client device (1138) may additionally store at least some of the instructions for generating binaural audio content in the memory of the client device (1138).

[0070] The output components of the head-wearable device (1102) include visual components such as displays, for example, liquid crystal displays (LCDs), plasma display panels (PDPs), light emitting diode (LED) displays, projectors, or waveguides. The image displays of the optical assembly are driven by an image display driver (1108). The output components of the head-wearable device (1102) further include acoustic components (e.g., speakers), haptic components (e.g., vibration motors), other signal generators, etc. Input components of a head-wearable device (1102), a client device (1138), and a server system (1132), such as a user input device (1106), may include alphanumeric input components (e.g., a keyboard, a touch screen configured to receive alphanumeric input, a photo-optical keyboard, or other alphanumeric input components), point-based input components (e.g., a mouse, a touchpad, a trackball, a joystick, a motion sensor, or other pointing mechanism), haptic input components (e.g., a physical button, a touch screen providing location and force of touches or touch gestures, or other haptic input components), audio input components (e.g., a microphone), and similar ones.

[0071] The head-wearable device (1102) may optionally include additional peripheral device elements. These peripheral device elements may include biometric sensors integrated with the head-wearable device (1102), additional sensors, or display elements. For example, the peripheral device elements may include any I / O components including output components, motion components, position components, or any other such elements described herein.

[0072] For example, biometric components include components that detect expressions (e.g., hand expressions, facial expressions, voice expressions, body gestures, or eye tracking), measure biosignals (e.g., blood pressure, heart rate, body temperature, sweat, or brainwaves), identify a person (e.g., voice identification, retinal identification, facial identification, fingerprint identification, or EEG-based identification), and perform similar functions. Motion components include acceleration sensor components (e.g., accelerometers), gravity sensor components, rotation sensor components (e.g., gyroscopes), etc. Position components include location sensor components for generating location coordinates (e.g., Global Positioning System (GPS) receiver components), and WiFi or Bluetooth for generating positioning system coordinates. TM It includes transceivers, altitude sensor components (e.g., altimeters or barometers that detect atmospheric pressure from which altitude can be derived), orientation sensor components (e.g., magnetometers), and similar ones. These position determination system coordinates may also be received over communication (1136) from a client device (1138) via a low-power wireless circuit (1130) or a high-speed wireless circuit (1124).

[0073] Where phrases similar to "at least one of A, B, or C," "at least one of A, B, and C," "one or more of A, B, or C," or "one or more of A, B, and C" are used, such phrases are intended to be interpreted to mean that A may exist alone in the embodiment, B may exist alone in the embodiment, C may exist alone in the embodiment, or any combination of elements A, B, and C—e.g., A and B, A and C, B and C, or A, B, and C—may exist in a single embodiment.

[0074] Changes and modifications to the disclosed embodiments may be made without departing from the scope of the present disclosure. These and other changes or modifications are intended to be included within the scope of the present disclosure as expressed in the following claims.

[0075] FIG. 12 is a block diagram (1200) illustrating a software architecture (1204) that may be installed in any one or more of the devices described herein. The software architecture (1204) is supported by hardware such as a machine (1202) that includes processors (1220), memory (1226), and I / O components (1238). In this example, the software architecture (1204) may be conceptualized as a stack of layers, where each layer provides specific functionality. The software architecture (1204) includes layers such as an operating system (1212), libraries (1210), frameworks (1208), and applications (1206). Operationally, applications (1206) invoke API calls (1250) through the software stack and receive messages (1252) in response to the API calls (1250).

[0076] The operating system (1212) manages hardware resources and provides common services. The operating system (1212) includes, for example, a kernel (1214), services (1216), and drivers (1222). The kernel (1214) acts as an abstraction layer between the hardware and other software layers. For example, the kernel (1214) provides, among other functionalities, memory management, processor management (e.g., scheduling), component management, networking, and security settings. Services (1216) may provide other common services for other software layers. Drivers (1222) are responsible for controlling the underlying hardware or interfacing with it. For example, drivers (1222) may include display drivers, camera drivers, BLUETOOTH® or BLUETOOTH® Low Energy drivers, flash memory drivers, serial communication drivers (e.g., USB (Universal Serial Bus) drivers), WI-FI® drivers, audio drivers, power management drivers, etc.

[0077] Libraries (1210) provide low-level common infrastructure used by applications (1206). Libraries (1210) may include system libraries (1218) (e.g., C standard library) that provide memory allocation functions, string manipulation functions, mathematical functions, and similar functions. Additionally, the libraries (1210) may include media libraries (e.g., libraries that support the presentation and manipulation of various media formats such as MPEG4 (Moving Picture Experts Group-4), H.264 or AVC (Advanced Video Coding), MP3 (Moving Picture Experts Group Layer-3), AAC (Advanced Audio Coding), AMR (Adaptive Multi-Rate) audio codecs, JPEG or JPG (Joint Photographic Experts Group), or PNG (Portable Network Graphics), graphics libraries (e.g., OpenGL frameworks used to render two-dimensional (2D) and three-dimensional (3D) graphic content on a display), database libraries (e.g., SQLite, which provides various relational database functions), web libraries (e.g., WebKit, which provides web browsing functionality), and API libraries (1224) such as similar ones. The libraries (1210) may also include a wide variety of other libraries (1228) to provide many different APIs to applications (1206).

[0078] Frameworks (1208) provide high-level common infrastructure used by applications (1206). For example, frameworks (1208) provide various graphical user interface (GUI) functions, high-level resource management, and high-level location services. Frameworks (1208) may provide a wide spectrum of other APIs that can be used by applications (1206), some of which may be specific to a specific operating system or platform.

[0079] In an exemplary embodiment, applications (1206) may include a wide range of other applications such as a home application (1236), a contacts application (1230), a browser application (1232), a book reader application (1234), a location application (1242), a media application (1244), a messaging application (1246), a game application (1248), and a third-party application (1240). Applications (1206) are programs that execute functions defined in the programs. To create one or more of the applications (1206) structured in various ways, various programming languages ​​such as object-oriented programming languages ​​(e.g., Objective-C, Java, or C++) or procedural programming languages ​​(e.g., C or assembly language) may be used. In a specific example, third-party applications (1240) (e.g., ANDROID by an entity other than the vendor of a specific platform) TM or iOS TM Applications developed using a Software Development Kit (SDK) are iOS TM , ANDROID TMIt may be mobile software running on a mobile operating system, such as WINDOWS® Phone or other mobile operating systems. In this example, a third-party application (1240) may invoke API calls (1250) provided by the operating system (1212) to facilitate the functionality described herein.

[0080] FIG. 13 is a schematic representation of a machine (1300) on which instructions (1308) (e.g., software, programs, applications, applets, apps, or other executable code) can be executed to cause the machine (1300) to perform any one or more of the methodologies discussed herein. For example, instructions (1308) can cause the machine (1300) to perform any one or more of the methods described herein. Instructions (1308) convert a general unprogrammed machine (1300) into a specific machine (1300) programmed to perform the described and illustrated functions in the described manner. The machine (1300) may operate as a standalone device or may be coupled to other machines (e.g., networked). In a networked deployment, the machine (1300) may operate as a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine (1300) may include, but is not limited to, a server computer, a client computer, a personal computer (PC), a tablet computer, a laptop computer, a netbook, a set-top box (STB), a PDA, an entertainment media system, a cellular phone, a smartphone, a mobile device, a wearable device (e.g., a smart watch), a smart home device (e.g., a smart appliance), other smart devices, a web appliance, a network router, a network switch, a network bridge, or any machine capable of executing commands (1308) that specify actions to be taken by the machine (1300) sequentially or otherwise. Additionally, although only a single machine (1300) is exemplified, the term “machine” should also be considered to include a collection of machines that execute commands (1308) individually or jointly to perform any one or more of the methodologies discussed herein.

[0081] The machine (1300) may include processors (1302), memory (1304), and I / O components (1342) that can be configured to communicate with each other via a bus (1344). In an exemplary embodiment, the processors (1302) (e.g., a CPU (Central Processing Unit), a RISC (Reduced Instruction Set Computing) processor, a CISC (Complex Instruction Set Computing) processor, a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), an ASIC, a RFIC (Radio-Frequency Integrated Circuit), other processors, or any suitable combination thereof) may include, for example, a processor (1306) and a processor (1310) that execute instructions (1308). The term "processor" is intended to include multi-core processors that may include two or more independent processors (sometimes referred to as "cores") capable of executing instructions simultaneously. FIG. 13 illustrates multiple processors (1302), but the machine (1300) may include a single processor having a single core, a single processor having multiple cores (e.g., a multi-core processor), multiple processors having a single core, multiple processors having multiple cores, or any combination thereof.

[0082] Memory (1304) includes main memory (1312), static memory (1314), and storage unit (1316), both of which are accessible to processors (1302) via a bus (1344). The main memory (1304), static memory (1314), and storage unit (1316) store instructions (1308) that implement any one or more of the methodologies or functions described herein. The instructions (1308) may also exist, wholly or partially, during execution by the machine (1300), in the main memory (1312), in the static memory (1314), in the machine-readable medium (1318) in the storage unit (1316), in at least one of the processors (1302) (e.g., in the processor's cache memory), or any suitable combination thereof.

[0083] The I / O components (1342) may include a wide variety of components for receiving inputs, providing outputs, generating outputs, transmitting information, exchanging information, capturing measurements, etc. The specific I / O components (1342) included in a specific machine will depend on the type of machine. For example, portable machines such as mobile phones may include touch input devices or other such input mechanisms, whereas headless server machines may not include such touch input devices. It will be understood that the I / O components (1342) may include many other components not shown in FIG. 13. In various exemplary embodiments, the I / O components (1342) may include output components (1328) and input components (1330). The output components (1328) may include visual components (e.g., displays such as a plasma display panel (PDP), light emitting diode (LED) display, liquid crystal display (LCD), projector, or cathode ray tube (CRT)), acoustic components (e.g., speakers), haptic components (e.g., vibration motors, resistance mechanisms), other signal generators, etc.Input components (1330) may include alphanumeric input components (e.g., a keyboard, a touch screen configured to receive alphanumeric input, a photo-optical keyboard, or other alphanumeric input components), point-based input components (e.g., a mouse, a touchpad, a trackball, a joystick, a motion sensor, or other pointing mechanism), haptic input components (e.g., a physical button, a touch screen providing location and / or force of touches or touch gestures, or other haptic input components), audio input components (e.g., a microphone), and similar ones.

[0084] In additional exemplary embodiments, I / O components (1342) may include biometric components (1332), motion components (1334), environment components (1336), or position components (1338), among a wide range of other components. For example, biometric components (1332) include components that detect expressions (e.g., hand expressions, facial expressions, voice expressions, body gestures, or eye tracking), measure biosignals (e.g., blood pressure, heart rate, body temperature, sweat, or brainwaves), and identify a person (e.g., voice identification, retinal identification, facial identification, fingerprint identification, or brainwave-based identification). Motion components (1334) include acceleration sensor components (e.g., accelerometers), gravity sensor components, rotation sensor components (e.g., gyroscopes), etc. Environmental components (1336) include, for example, light sensor components (e.g., photometers), temperature sensor components (e.g., one or more thermometers for detecting ambient temperature), humidity sensor components, pressure sensor components (e.g., barometers), acoustic sensor components (e.g., one or more microphones for detecting background noise), proximity sensor components (e.g., infrared sensors for detecting nearby objects), gas sensors (e.g., gas detection sensors for detecting concentrations of hazardous gases for safety or measuring pollutants in the atmosphere), or other components capable of providing indications, measurements, or signals corresponding to the surrounding physical environment. Location components (1338) include location sensor components (e.g., GPS receiver components), altitude sensor components (e.g., altimeters or barometers for detecting atmospheric pressure from which altitude can be derived), orientation sensor components (e.g., magnetometers), and similar ones.

[0085] Communication can be implemented using a wide variety of technologies. The I / O components (1342) further include communication components (1340) operable to connect the machine (1300) to the network (1320) or devices (1322) respectively through coupling (1324) and coupling (1326). For example, the communication components (1340) may include a network interface component or other suitable device for interfacing with the network (1320). In additional examples, the communication components (1340) may include wired communication components, wireless communication components, cellular communication components, NFC (Near Field Communication) components, Bluetooth® components (e.g., Bluetooth® Low Energy), Wi-Fi® components, and other communication components that provide communication through other modalities. The devices (1322) may be any of the other machine or a wide variety of peripheral devices (e.g., peripheral devices coupled via USB).

[0086] Furthermore, the communication components (1340) may include components capable of detecting identifiers or detecting identifiers. For example, the communication components (1340) may include Radio Frequency Identification (RFID) tag reader components, NFC smart tag detection components, optical reader components (e.g., optical sensors for detecting multidimensional barcodes such as one-dimensional barcodes like Universal Product Code (UPC) barcodes, Quick Response (QR) codes, Aztec codes, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D barcodes, and other optical codes), or acoustic detection components (e.g., microphones for identifying tagged audio signals). In addition, various information such as locations via Internet Protocol (IP) geolocation, locations via Wi-Fi® signal triangulation, and locations via the detection of NFC beacon signals capable of indicating specific locations can be derived through the communication components (1340).

[0087] Various memories (e.g., memory (1304), main memory (1312), static memory (1314), and / or memory of processors (1302)) and / or storage unit (1316) may store one or more sets of instructions and data structures (e.g., software) that implement any one or more of the methodologies or functions described herein. These instructions (e.g., instructions (1308)) cause various operations to implement the disclosed embodiments when executed by processors (1302).

[0088] Commands (1308) may be transmitted or received through a network (1320) using a transmission medium, through a network interface device (e.g., a network interface component included in the communication components (1340)) and using any one of a number of well-known transmission protocols (e.g., Hypertext Transfer Protocol (HTTP)). Similarly, commands (1308) may be transmitted or received using a transmission medium through a connection (1326) (e.g., a peer-to-peer connection) to the devices (1322).

[0089] As used herein, the terms “machine storage medium,” “device storage medium,” and “computer storage medium” mean the same thing and may be used interchangeably in this disclosure. The terms refer to one or more storage devices and / or media (e.g., centralized or distributed databases, and / or associated caches and servers) that store executable instructions and / or data. Accordingly, the terms should be considered to include, but not be limited to, optical and magnetic media, including solid-state memories and memory inside or outside processors. Specific examples of machine storage medium, computer storage medium, and / or device storage medium include non-volatile memory, which, by example, include semiconductor memory devices, e.g., EPROM (erasable programmable read-only memory), EEPROM (electrically erasable programmable read-only memory), FPGA (field-programmable gate array), and flash memory devices; magnetic disks such as internal hard disks and removable disks; magneto-optical disks; and includes CD-ROM and DVD-ROM discs. The terms “machine storage medium,” “computer storage medium,” and “device storage medium” specifically exclude carriers, modulated data signals, and other such media, at least some of which are covered under the term “signal medium” discussed below.

[0090] The terms “transmission medium” and “signal medium” mean the same thing and may be used interchangeably in this disclosure. The terms “transmission medium” and “signal medium” should be considered to include any intangible medium capable of storing, encoding, or carrying instructions (1416) for execution by the machine (1400), and digital or analog communication signals or other intangible media to facilitate communication of such software. Accordingly, the terms “transmission medium” and “signal medium” should be considered to include any form of modulated data signal, carrier wave, etc. The term “modulated data signal” means a signal having one or more of the characteristics set or changed in a manner that encodes information in the signal.

[0091] The terms “machine-readable medium,” “computer-readable medium,” and “device-readable medium” mean the same thing and may be used interchangeably in this disclosure. The terms are defined to include both machine storage media and transmission media. Accordingly, the terms include both storage devices / medias and carriers / modulated data signals.

[0092] Although embodiments have been described with reference to specific exemplary embodiments, it will be apparent that various modifications and changes may be made to these embodiments without departing from the broader scope of the disclosure. Accordingly, the specification and drawings should be regarded as exemplary rather than restrictive. The accompanying drawings, which form part of this specification, illustrate specific embodiments in which the subject matter may be practiced as examples rather than limitations. The illustrated embodiments are described in sufficient detail to enable those skilled in the art to practice the teachings disclosed in this specification. Other embodiments may be used and derived therefrom so that structural and logical substitutions and changes may be made without departing from the scope of the disclosure. Accordingly, this detailed description should not be regarded in a restrictive sense, and the scope of various embodiments is defined only by the appended claims and the full scope of equivalents given to such claims.

[0093] These embodiments of the subject matter of the present invention may be referred to herein by the term “invention” individually and / or collectively merely for convenience, and without the intention of voluntarily limiting the scope of this application to any single invention or concept of invention where one or more are actually disclosed. Accordingly, while specific embodiments have been illustrated and described herein, it should be understood that any arrangement calculated to achieve the same purpose may replace the specific embodiments illustrated. The present disclosure is intended to cover any and all adaptations or variations of the various embodiments. Combinations of the above embodiments and other embodiments not specifically described herein will be apparent to those skilled in the art upon reviewing the above description.

[0094] A summary of the present disclosure is provided to enable the reader to quickly ascertain the essence of the technical disclosure. It is submitted with the understanding that it is not intended to be used to interpret or limit the scope or meaning of the claims. Additionally, in the foregoing detailed description, it may be seen that various features are grouped together in a single embodiment for the purpose of simplifying the present disclosure. This method of the present disclosure should not be interpreted as reflecting an intention that the claimed embodiments require more features than are explicitly mentioned in each claim. Rather, as reflected in the following claims, the subject matter of the present invention lies in fewer features than all the features of a single disclosed embodiment. Accordingly, the following claims are incorporated by reference to the detailed description, and each claim is independent as a separate embodiment.

[0095] Examples

[0096] Example 1 is a method comprising the steps of: accessing a first image generated by a first camera of a first augmented reality device and a second image generated by a second camera of a second augmented reality device—the first image and the second image depict a hand gesture of a user of the first augmented reality device—; synchronizing the first augmented reality device with the second augmented reality device; distributing one or more processes of a sign language recognition system between the first augmented reality device and the second augmented reality device in response to the synchronization—one or more processes are performed on the corresponding augmented reality devices—; collecting results from one or more processes from the first augmented reality device and from the second augmented reality device; and displaying text that displays a sign language translation of a hand gesture based on the results in near real-time on a first display of the first augmented reality device or on a second display of the second augmented reality device.

[0097] Example 2 includes the method of Example 1, and the step of synchronizing the first augmented reality device with the second augmented reality device further includes the step of mapping the first timestamp of the first image to the second timestamp of the second image.

[0098] Example 3 includes the method of Example 1, and the step of synchronizing the first augmented reality device with the second augmented reality device further includes the step of aligning the 6-degrees-of-freedom coordinate system of the first augmented reality device with the 6-degrees-of-freedom coordinate system of the second augmented reality device.

[0099] Example 4 includes the method of Example 1 and further includes the step of communicating a first image to a second augmented reality device—the first image provides an egocentric view of a user’s hand gesture of the first augmented reality device—and the second augmented reality device is configured to perform one or more processes of a sign language recognition system on a combination of the first image, a first hand skeleton based on the first image, a second image, and a second hand skeleton based on the second image, and one or more processes of the sign language recognition system include: a hand tracking process, a hand gesture detection process, and a sign language translation process.

[0100] Example 5 includes the method of Example 1 and further includes the step of identifying a proposed word based on a combination of context information of the first augmented reality device, history inputs of the user of the first augmented reality device, a text corpus, and predefined terms, and the text displays the proposed word.

[0101] Example 6 includes the method of Example 5 and further includes the steps of: displaying a proposed word on a first display of a first augmented reality device; receiving confirmation of the proposed word from a user of the first augmented reality device; and providing the proposed word to a second augmented reality device in response to receiving confirmation.

[0102] Example 7 includes the method of Example 6, and the step of receiving confirmation includes the step of detecting a predefined gesture from the user, and the predefined gesture indicates confirmation.

[0103] Example 8 includes the method of Example 1, and the step of distributing one or more processes of a sign language recognition system further includes the step of distributing one or more processes temporally or spatially between a first augmented reality device and a second augmented reality device, the temporal distribution includes processing alternating frames between the first augmented reality device and the second augmented reality device, and the spatial distribution includes processing a first gesture from the user's first hand only with the first augmented reality device, and a second gesture from the user's second hand only with the second augmented reality device.

[0104] Example 9 includes the method of Example 1, and the step of distributing one or more processes of a sign language recognition system further includes: detecting a blockage of a hand gesture in a first image by a first augmented reality device; and assigning a second augmented reality device to perform one or more processes based on a second image in response to detecting the blockage.

[0105] Example 10 includes the method of Example 1, the first augmented reality device includes a first head-wearable device, and the second augmented reality device includes a mobile hand-held device or a second head-wearable device.

[0106] Example 11 is a computing device comprising: a processor; and memory for storing instructions, wherein, when the instructions are executed by the processor, the device is configured to: access a first image generated by a first camera of a first augmented reality device and a second image generated by a second camera of a second augmented reality device—the first image and the second image depict a hand gesture of a user of the first augmented reality device—; synchronize the first augmented reality device with the second augmented reality device; in response to the synchronization, distribute one or more processes of a sign language recognition system between the first augmented reality device and the second augmented reality device—one or more processes are performed on the corresponding augmented reality devices—; collect results from one or more processes from the first augmented reality device and from the second augmented reality device; and display text displaying a sign language translation of a hand gesture based on the results in near real-time on a first display of the first augmented reality device or a second display of the second augmented reality device.

[0107] Example 12 includes the computing device of Example 11, and synchronizing the first augmented reality device with the second augmented reality device further includes mapping the first timestamp of the first image to the second timestamp of the second image.

[0108] Example 13 includes the computing device of Example 11, and synchronizing the first augmented reality device with the second augmented reality device further includes aligning the 6-degrees-of-freedom coordinate system of the first augmented reality device with the 6-degrees-of-freedom coordinate system of the second augmented reality device.

[0109] Example 14 includes the computing device of Example 11, and instructions further configure the device to communicate a first image to a second augmented reality device, wherein the first image provides an egocentric view of a user's hand gesture of the first augmented reality device, and the second augmented reality device is configured to perform one or more processes of a sign language recognition system on a combination of the first image, a first hand skeleton based on the first image, a second image, and a second hand skeleton based on the second image, and the one or more processes of the sign language recognition system include a hand tracking process, a hand gesture detection process, and a sign language translation process.

[0110] Example 15 includes the computing device of Example 11, and instructions further configure the device to identify a suggested word based on a combination of context information of the first augmented reality device, user history inputs of the first augmented reality device, a text corpus, and predefined terms, and text displays the suggested word.

[0111] Example 16 includes the computing device of Example 15, and the instructions further configure the device to: display a proposed word on a first display of a first augmented reality device; receive confirmation of the proposed word from a user of the first augmented reality device; and, in response to receiving confirmation, provide the proposed word to a second augmented reality device.

[0112] Example 17 includes the computing device of Example 16, and receiving acknowledgment includes: detecting a predefined gesture from a user, and the predefined gesture indicates acknowledgment.

[0113] Example 18 includes the computing device of Example 11, and distributing one or more processes of a sign language recognition system further includes distributing one or more processes temporally or spatially between a first augmented reality device and a second augmented reality device, wherein temporal distribution includes processing alternating frames between the first augmented reality device and the second augmented reality device; and spatial distribution includes processing a first gesture from the user's first hand only with the first augmented reality device, and a second gesture from the user's second hand only with the second augmented reality device.

[0114] Example 19 includes the computing device of Example 11 and further comprises distributing one or more processes of a sign language recognition system: detecting a blockage of a hand gesture in a first image by a first augmented reality device; and assigning a second augmented reality device to perform one or more processes based on a second image in response to detecting the blockage.

[0115] Example 20 is a non-transient computer-readable storage medium, and the computer-readable storage medium includes instructions that, when executed by a computer, cause the computer to: access a first image generated by a first camera of a first augmented reality device and a second image generated by a second camera of a second augmented reality device—the first image and the second image depict a hand gesture of a user of the first augmented reality device—; synchronize the first augmented reality device with the second augmented reality device; in response to the synchronization, distribute one or more processes of a sign language recognition system between the first augmented reality device and the second augmented reality device—one or more processes are performed on the corresponding augmented reality devices—; collect results from one or more processes from the first augmented reality device and from the second augmented reality device; and display text displaying a sign language translation of a hand gesture based on the results in near real-time on a first display of the first augmented reality device or on a second display of the second augmented reality device.

[0116] The described implementations of the subject may include one or more features alone or in combination, as exemplified below.

Claims

Claim 1 A method comprising: accessing a first image generated by a first camera of a first augmented reality device and a second image generated by a second camera of a second augmented reality device, wherein the first image and the second image depict a hand gesture of a user of the first augmented reality device; synchronizing the first augmented reality device with the second augmented reality device; distributing one or more processes of a sign language recognition system between the first augmented reality device and the second augmented reality device in response to the synchronization, wherein the one or more processes are performed on the corresponding augmented reality device; collecting results from the one or more processes from the first augmented reality device and from the second augmented reality device; and displaying text indicating a sign language translation of the hand gesture based on the results in near real-time on a first display of the first augmented reality device or on a second display of the second augmented reality device. Claim 2 A method according to claim 1, wherein the step of synchronizing the first augmented reality device with the second augmented reality device further comprises the step of mapping a first timestamp of the first image to a second timestamp of the second image. Claim 3 The method of claim 1, wherein the step of synchronizing the first augmented reality device with the second augmented reality device further comprises the step of registering the 6-degrees-of-freedom coordinate system of the first augmented reality device with the 6-degrees-of-freedom coordinate system of the second augmented reality device. Claim 4 The method of claim 1 further comprises the step of communicating the first image to the second augmented reality device, wherein the first image provides an egocentric view of the user's hand gesture of the first augmented reality device, and the second augmented reality device is configured to perform the one or more processes of the sign language recognition system on a combination of the first image, a first hand skeleton based on the first image, the second image, and a second hand skeleton based on the second image, and the one or more processes of the sign language recognition system include: a hand tracking process, a hand gesture detection process, and a sign language translation process. Claim 5 A method according to claim 1, further comprising the step of identifying a proposed word based on a combination of context information of the first augmented reality device, historical inputs of the user of the first augmented reality device, a text corpus, and predefined terms, wherein the text displays the proposed word. Claim 6 A method according to claim 5, further comprising: displaying the proposed word on the first display of the first augmented reality device; receiving confirmation of the proposed word from the user of the first augmented reality device; and providing the proposed word to the second augmented reality device in response to receiving the confirmation. Claim 7 In claim 6, the step of receiving the confirmation comprises: detecting a predefined gesture from the user, wherein the predefined gesture indicates the confirmation. Claim 8 A method according to claim 1, wherein the step of distributing one or more processes of the sign language recognition system further comprises the step of distributing the one or more processes temporally or spatially between the first augmented reality device and the second augmented reality device, wherein the temporal distribution comprises processing alternating frames between the first augmented reality device and the second augmented reality device, and the spatial distribution comprises processing a first gesture from the user's first hand solely with the first augmented reality device, and a second gesture from the user's second hand solely with the second augmented reality device. Claim 9 The method of claim 1, wherein the step of distributing one or more processes of the sign language recognition system further comprises: detecting an occlusion of the hand gesture in the first image by the first augmented reality device; and assigning the second augmented reality device to perform the one or more processes based on the second image in response to detecting the occlusion. Claim 10 A method according to claim 1, wherein the first augmented reality device comprises a first head-wearable device, and the second augmented reality device comprises a mobile hand-held device or a second head-wearable device. Claim 11 A computing device comprising: a processor; and a memory for storing instructions, wherein, when the instructions are executed by the processor, the device comprises: accessing a first image generated by a first camera of a first augmented reality device and a second image generated by a second camera of a second augmented reality device, wherein the first image and the second image depict a hand gesture of a user of the first augmented reality device; synchronizing the first augmented reality device with the second augmented reality device; distributing one or more processes of a sign language recognition system between the first augmented reality device and the second augmented reality device in response to the synchronization, wherein the one or more processes are performed on the corresponding augmented reality devices; collecting results from the one or more processes from the first augmented reality device and from the second augmented reality device; and configuring the device to display text displaying a sign language translation of the hand gesture based on the results in near real-time on a first display of the first augmented reality device or on a second display of the second augmented reality device. Claim 12 A computing device according to claim 11, wherein synchronizing the first augmented reality device with the second augmented reality device further comprises mapping a first timestamp of the first image to a second timestamp of the second image. Claim 13 A computing device according to claim 11, wherein synchronizing the first augmented reality device with the second augmented reality device further comprises: aligning the 6-degrees-of-freedom coordinate system of the first augmented reality device with the 6-degrees-of-freedom coordinate system of the second augmented reality device. Claim 14 In paragraph 11, the above commands further configure the device to communicate the first image to the second augmented reality device—the first image provides a self-centered view of the user's hand gesture to the first augmented reality device—and the second augmented reality device is configured to perform the one or more processes of the sign language recognition system on a combination of the first image, a first hand skeleton based on the first image, the second image, and a second hand skeleton based on the second image, and the one or more processes of the sign language recognition system include: a hand tracking process, a hand gesture detection process, and a sign language translation process, a computing device. Claim 15 In paragraph 11, the above commands further configure the device to identify a proposed word based on a combination of context information of the first augmented reality device, history inputs of the user of the first augmented reality device, a text corpus, and predefined terms, and the text displays the proposed word, a computing device. Claim 16 In paragraph 15, the above commands further configure the device to: display the proposed word on the first display of the first augmented reality device; receive confirmation of the proposed word from the user of the first augmented reality device; and, in response to receiving the confirmation, provide the proposed word to the second augmented reality device. Claim 17 In paragraph 16, receiving the above confirmation comprises: detecting a predefined gesture from the user—the predefined gesture indicates the confirmation. A computing device. Claim 18 A computing device according to claim 11, wherein distributing the one or more processes of the sign language recognition system further comprises distributing the one or more processes temporally or spatially between the first augmented reality device and the second augmented reality device, wherein temporal distribution comprises processing alternating frames between the first augmented reality device and the second augmented reality device, and spatial distribution comprises processing a first gesture from the user's first hand solely by the first augmented reality device, and a second gesture from the user's second hand solely by the second augmented reality device. Claim 19 A computing device according to claim 11, wherein distributing the one or more processes of the sign language recognition system further comprises: detecting a blockage of the hand gesture in the first image by the first augmented reality device; and assigning the second augmented reality device to perform the one or more processes based on the second image in response to detecting the blockage. Claim 20 A non-transient computer-readable storage medium comprising instructions that, when executed by a computer, cause the computer to: access a first image generated by a first camera of a first augmented reality device and a second image generated by a second camera of a second augmented reality device, wherein the first image and the second image depict a hand gesture of a user of the first augmented reality device; synchronize the first augmented reality device with the second augmented reality device; in response to the synchronization, distribute one or more processes of a sign language recognition system between the first augmented reality device and the second augmented reality device, wherein the one or more processes are performed on the corresponding augmented reality devices; collect results from the one or more processes from the first augmented reality device and from the second augmented reality device; and display text displaying a sign language translation of the hand gesture based on the results in near real-time on a first display of the first augmented reality device or on a second display of the second augmented reality device.

Citation Information

Patent Citations

  • sensory eyewear

    KR102257181B1

  • Hand shape matching VR sign language education system using gesture recognition technology

    KR102436239B1

  • Sign language communication with communication devices

    US20170277684A1

  • Eyewear including sign language to speech translation

    US20220188539A1