Sign language translation using collaborative agents
By allocating sign language processing among collaborative augmented reality devices, the problem of inefficient sign language translation resource management is solved, and the optimization of computing resources and real-time improvement of sign language translation is achieved.
Patent Information
- Application Number
- CN202380073035.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2022-10-17
- Filing Date
- 2023-10-16
- Publication Date
- 2025-05-30
AI Technical Summary
The prior art is difficult to effectively manage the resources of sign language translation, resulting in inefficient computer function operation.
By allocating sign language processing between synchronous collaborative augmented reality devices, power consumption reduction and computing resource optimization are provided, reducing the need for resources such as processor cycles, network traffic, and memory usage.
It has achieved improved resource management efficiency of sign language translation, reduced power consumption and resource consumption of computing devices, and improved the real-time and accuracy of translation processing.
Smart Images

Figure CN120077346A_ABST
Abstract
Description
[0001] Priority Claim
[0002] This application claims the benefit of priority of U.S. Patent Application No. 17 / 967,209, filed Oct. 17, 2022, which is hereby incorporated by reference in its entirety. Technical Field
[0003] The subject matter disclosed herein generally relates to augmented reality (AR) devices. Specifically, the present disclosure presents systems and methods for sign language translation using collaborative agents. Background Art
[0004] Augmented reality (AR) devices enable a user to observe a scene while seeing relevant virtual content that can be aligned with items, images, objects, or the environment within the device's field of view. Virtual reality (VR) devices provide a more immersive experience than AR devices. VR devices utilize virtual content displayed based on the positioning and orientation of the VR device to occlude the user's field of view. Brief Description of the Drawings
[0005] To facilitate an easy identification of the discussion of any particular element or act, one or more of the most significant digits in the reference numerals refer to the figure number in which that element is first introduced.
[0006] Figure 1 is a block diagram showing a network environment for sign language translation using collaborative augmented reality devices according to an example embodiment.
[0007] Figure 2 is a block diagram showing a network environment for sign language translation using collaborative augmented reality devices according to an example embodiment.
[0008] Figure 3 is a block diagram showing an augmented reality device according to an example embodiment.
[0009] Figure 4 is a block diagram showing a collaborative sign language system according to an example embodiment.
[0010] Figure 5 is an interaction diagram showing the interaction between collaborative augmented reality devices according to an example embodiment.
[0011] Figure 6 is an interaction diagram showing the interaction between collaborative augmented reality devices according to an example embodiment.
[0012] Figure 7 is an interaction diagram showing the interaction between collaborative augmented reality devices according to an example embodiment.
[0013] Figure 8It is an interaction diagram showing the interaction between collaborative augmented reality devices according to an example embodiment.
[0014] Figure 9 It is a flowchart showing a method for sign language translation according to an example embodiment.
[0015] Figure 10 It is a flowchart showing a method for allocating sign language translation processing according to an example embodiment.
[0016] Figure 11 It shows a network environment in which a head-mounted device can be implemented according to an example embodiment.
[0017] Figure 12 It is a block diagram showing a software architecture in which the present disclosure can be implemented according to an example embodiment.
[0018] Figure 13 It is a graphical representation of a machine in the form of a computer system, within which a set of instructions can be executed to cause the machine to perform any one or more of the methods discussed herein. Detailed Description
[0019] The following description describes systems, methods, techniques, instruction sequences, and computer program products that illustrate example embodiments of the subject matter. In the following description, for purposes of explanation, numerous specific details are set forth in order to provide an understanding of the various embodiments of the subject matter. However, it will be apparent to those skilled in the art that embodiments of the subject matter may be practiced without some or other of these specific details. The examples merely represent possible variations. Unless otherwise explicitly stated, structures (e.g., structural components such as modules) are optional and may be combined or subdivided, and operations (e.g., in a process, algorithm, or other function) may vary in order or be combined or subdivided.
[0020] The term "augmented reality" (AR) is used herein to refer to an interactive experience of a real-world environment in which physical objects residing in the real world are "augmented" or enhanced by computer-generated digital content (also referred to as virtual content or synthetic content). AR may also refer to a system that enables the combination of the real world and the virtual world, real-time interaction, and 3D registration of virtual objects and real objects. A user of an AR system perceives virtual content that appears to be connected to or interact with physical objects in the real world.
[0021] The term "virtual reality" (VR) is used herein to refer to a simulated experience of a virtual world environment that is completely different from the real-world environment. Digital content generated by a computer is displayed in the virtual world environment. VR also refers to a system that enables a user of the VR system to be fully immersed in the virtual world environment and interact with virtual objects presented in the virtual world environment.
[0022] The term "AR application" is used herein to refer to a computer-operated application that implements an AR experience. The term "VR application" is used herein to refer to a computer-operated application that implements a VR experience. The term "AR / VR application" refers to a computer-operated application that implements a combination of an AR experience or a VR experience.
[0023] The term "visual tracking system" is used herein to refer to a computer-operated application or system that enables the system to track visual features identified in an image captured by one or more camera devices of the visual tracking system. The visual tracking system constructs a model of the real-world environment based on the tracked visual features. Non-limiting examples of the visual tracking system include: a visual simultaneous localization and mapping system (VSLAM) and a visual inertial odometry (VIO) system. VSLAM can be used to construct an object from an environment or a scene based on one or more camera devices of the visual tracking system. The VIO system (also referred to as a visual inertial tracking system) determines the latest pose (e.g., localization and orientation) of a device based on data obtained from multiple sensors of the device (e.g., an optical sensor, an inertial sensor).
[0024] The term "inertial measurement unit" (IMU) is used herein to refer to a device that can report the inertial state of a moving body, which includes the acceleration, velocity, orientation, and localization of the moving body. The IMU tracks the movement of the body by integrating the acceleration and angular velocity measured by the IMU. The IMU can also refer to a combination of an accelerometer and a gyroscope, which can respectively determine and quantify linear acceleration and angular velocity. The values obtained from the IMU gyroscope can be processed to obtain the pitch, roll, and heading of the IMU, thereby obtaining the pitch, roll, and heading of the body associated with the IMU. The signals from the accelerometer of the IMU can also be processed to obtain the velocity and displacement of the IMU.
[0025] The term "three - degree - of - freedom tracking system" (3DOF tracking system) is used herein to refer to a device that tracks rotational movement. For example, a 3DOF tracking system can track whether a user of a head - wearable device looks left or right, rotates their head up or down, and turns left or right. However, a head - wearable device cannot use a 3DOF tracking system to determine whether the user moves around a scene by moving in the physical world. Thus, a 3DOF tracking system may not be accurate enough to be used for positioning signals. A 3DOF tracking system can be part of an AR / VR display device that includes IMU sensors. For example, a 3DOF tracking system uses sensor data from sensors such as accelerometers, gyroscopes, and magnetometers.
[0026] The term "six - degree - of - freedom tracking system" (6DOF tracking system) is used herein to refer to a device that tracks rotational and translational movement. For example, a 6DOF tracking system can track whether a user rotates their head and moves forward or backward, laterally or vertically, and up or down. A 6DOF tracking system can include a Simultaneous Localization and Mapping (SLAM) system and / or a Visual - Inertial Odometry (VIO) system that rely on data obtained from multiple sensors (e.g., depth cameras, inertial sensors). A 6DOF tracking system analyzes data from sensors to accurately determine the pose of a display device.
[0027] This application describes a system for translating sign language using collaborative augmented reality devices. For example, a first user wears a head-wearable AR device (e.g., smart glasses) and performs gestures (e.g., sign language). A second user facing the first user can also wear another AR device, or can point the camera device of his / her mobile communication device (e.g., smartphone) at the hands of the first user. The camera devices of the AR devices of the first user and the second user capture the gestures of the first user from different perspectives or angles. The first AR device first synchronizes with the second AR device. For example, the synchronization can be based on the timestamps of the images generated by the camera devices of the first AR device and the second AR device. In another example, the synchronization can be based on the poses of the AR devices in the same registration coordinate system. One or more processes of the sign language translation system are allocated among the first AR device and the second AR device. For example, the first AR device can share the images of the gestures captured by the camera device of the first AR device with the second AR device. The second AR device may have more computing resources (e.g., a smartphone may have a faster processor, a larger memory, and a larger battery than smart glasses), and performs hand tracking, pose recognition, and sign language translation. In one example, the allocation of the processes of the sign language translation system can be based on the comparison of the specifications of the first AR device and the second AR device and other conditions (e.g., detecting occlusion of the hand from the camera device of the first AR device; when the hand is occluded in the camera device of the first AR device, the system will rely on the image from the camera device of the second AR device).
[0028] In other example embodiments, one AR device performs hand tracking and pose recognition, and the other AR device performs pose recognition. In other examples, the results of the processes (sign language translation, hand bone detection, semantic pose recognition) are shared between the AR devices in time (e.g., every other frame) or in space (e.g., only the left / right hand). In another example, the first AR device displays suggested words / phrases based on context information, the user's historical input, a text corpus, predefined terms. When a trigger condition (a specific pose or an incentive input by the user) is detected, the first user can select the suggested option to avoid finger spelling or accelerate translation.
[0029] In an example embodiment, a method for using a collaborative AR device to recognize sign language is described. In one aspect, the method includes accessing a first image generated by a first camera device of a first augmented reality device and a second image generated by a second camera device of a second augmented reality device, the first image and the second image depicting a gesture of a user of the first augmented reality device, synchronizing the first augmented reality device with the second augmented reality device, and in response to the synchronization, allocating one or more processes of a sign language recognition system between the first augmented reality device and the second augmented reality device, wherein the one or more processes are executed on corresponding augmented reality devices, collecting results from the one or more processes from the first augmented reality device and from the second augmented reality device, and based on the results, displaying text of a sign language translation indicating the gesture in a first display of the first augmented reality device or in a second display of the second augmented reality device in near real time.
[0030] Accordingly, one or more of the methods described herein help to solve the technical problem of resource management for sign language translation. The methods currently described provide an improvement in the operation of the functions of a computer by distributing sign language processing among synchronized collaborative AR devices to provide reduced power consumption. Accordingly, one or more of the methods described herein can eliminate the need for certain workloads or computing resources. Examples of such computing resources include processor cycles, network traffic, memory usage, data storage capacity, power consumption, network bandwidth, and cooling capacity.
[0031] Figure 1 is a block diagram showing a network environment 100 for sign language translation using collaborative AR devices according to an example embodiment. The network environment 100 includes AR device B 114, AR device A 110, and server 116 communicatively coupled to each other via network 104. AR device B 114, AR device A 110, and server 116 may each be implemented in whole or in part in a computer system as described below with respect to Figure 13 The server 116 may be part of a network-based system. For example, the network-based system may be or include a cloud-based server system that provides additional information such as reference frame alignment data for AR device B 114 and AR device A 110.
[0032] User 106 wears AR device A 110 (e.g., smart glasses) and operates AR device A 110 via voice, gestures, physical input (e.g., touching the surface of AR device A 110 or pressing a button on AR device A 110). User 106 performs gestures with his / her hand 108. AR device A 110 captures an image of hand 108 using a camera device (not shown) of AR device A 110. AR device A 110 can be a computing device with a display. The computing device can be detachably mounted to the head of user 106. In one example, the display can be a screen that displays the content captured by the camera device of AR device B 114 / AR device A 110. In another example, the display of the device can be transparent, such as in the lenses of wearable computing glasses, which enables user 106 to view the content presented on the display while viewing real-world objects visible through the display.
[0033] User 112 operates AR device B 114. User 112 can be a human user (e.g., a human), a machine user (e.g., a computer configured by a software program to interact with AR device B 114), or any suitable combination thereof (e.g., a human assisted by a machine or a machine supervised by a human). In one example, user 112 aims the camera device of AR device B 114 at hand 108 of user 106 in the real-world environment 102. AR device B 114 can be a computing device with a display, such as a smartphone, a tablet computer, or a wearable computing device (e.g., a watch or glasses). The computing device can be handheld or can be detachably mounted to the head of user 112. In one example, the display can be a screen that displays the content captured by the camera device of AR device B114 / AR device A 110. In another example, the display of the device can be transparent, such as in the lenses of wearable computing glasses, which enables the user to view the content presented on the display while viewing real-world objects visible through the display. In another example, AR device B114 includes a computing device (e.g., a mobile device capable of displaying text or playing audio).
[0034] In one example, AR device A 110 and AR device B 114 each generate an image of hand 108 from a different viewpoint. Thus, the images captured by AR device B 114 and AR device A 110 include an overlapping region.
[0035] AR device A 110 includes a tracking system (not shown). The tracking system uses optical sensors (e.g., image capture devices), inertial sensors (e.g., gyroscopes, accelerometers), wireless sensors (Bluetooth, Wi-Fi), GPS sensors, and audio sensors to track the pose (e.g., position and orientation) of AR device A 110 relative to the real-world environment 102 to determine the position of AR device A 110 within the real-world environment 102. For example, the tracking system of AR device A 110 includes a 6DOF tracking system that tracks whether user 106 has rotated their head and moved forward or backward, laterally or vertically, and up or down.
[0036] AR device B 114 also includes a tracking system (not shown). The tracking system uses optical sensors (e.g., image capture devices), inertial sensors (e.g., gyroscopes, accelerometers), wireless sensors (Bluetooth, Wi-Fi), GPS sensors, and audio sensors to track the pose (e.g., position and orientation) of AR device B 114 relative to the real-world environment 102 to determine the position of AR device B 114 within the real-world environment 102. For example, the tracking system of AR device B 114 includes a 6DOF tracking system that tracks whether user 112 (in the case of a head-worn device) has rotated their head and moved forward or backward, laterally or vertically, and up or down.
[0037] In one example embodiment, server 116 coordinates the allocation of sign language translation between AR device B 114 and AR device A 110. For example, server 116 identifies the available computing resources and technical specifications of AR device B 114 and AR device A 110 and allocates one or more processes from the sign language translation to AR device B 114 and AR device A 110 based on the available computing resources and technical specifications. In another example, server 116 receives image data, image metadata (e.g., timestamps), and pose information from both AR device B 114 and AR device A 110, synchronizes the received information, performs sign language translation based on the synchronized information, and provides the translation to both AR device B 114 and AR device A 110. In another example, server 116 first provides a proposed translation to AR device A 110 and waits for user 106 to confirm the proposed translation before sending the confirmed translation to AR device B 114.
[0038] Figure 1Any of the machines, databases, or devices shown can be implemented in a general-purpose computer that has been modified by software (e.g., configured or programmed) to be a special-purpose computer to perform one or more of the functions described herein for that machine, database, or device. For example, a computer system that can implement any one or more of the methods described herein is discussed below with reference to Figure 9 As used herein, a "database" is a data storage resource and can store data structured as text files, tables, spreadsheets, relational databases (e.g., object-relational databases), triple stores, hierarchical data stores, or any suitable combination thereof. Additionally, Figure 1 Any two or more of the machines, databases, or devices shown in
[0039] Network 104 can be any network that enables communication between or among machines (e.g., server 116), databases, and devices (e.g., AR device B114, AR device A 110). Thus, network 104 can be a wired network, a wireless network (e.g., a mobile or cellular network), or any suitable combination thereof. Network 104 can include one or more portions that make up a private network, a public network (e.g., the Internet), or any suitable combination thereof.
[0040] Figure 2 is a network diagram showing a network environment 200 suitable for operating AR device B 114 and AR device A 110 according to some example embodiments. Network environment 200 includes AR device B 114 and AR device A 110. AR device B 114 can communicate wirelessly (e.g., via Bluetooth) with AR device A 110. In contrast to Figure 1 AR device B 114 and AR device A 110 communicate directly with each other.
[0041] AR device A 110 and AR device B 114 can share data with each other. In one example, AR device A 110 receives image data, pose data from AR device B 114, and synchronizes AR device A 110 with AR device B 114 based on the received data from AR device B 114 and the image data and pose data of AR device A 110. In another example, AR device B 114 receives image data, pose data from AR device B 114, and synchronizes AR device B 114 with AR device A 110 based on the received data from AR device A 110 and the image data and pose data of AR device B114.
[0042] Figure 3 is a block diagram showing modules (e.g., components) of an AR device A 110 according to some example embodiments. The AR device A 110 includes a sensor 302, a display 304, a processor 308, a graphics processing unit 318, a display controller 320, and a storage device 306. Examples of the AR device A 110 include wearable computing devices. The AR device B 114 includes components similar to those of the AR device A 110. However, the AR device B 114 may include a wearable computing device, a tablet computer, or a smart phone.
[0043] The sensor 302 includes an optical sensor 314 and an inertial sensor 316. The optical sensor 314 includes a stereo camera device. The inertial sensor 316 includes a combination of a gyroscope, an accelerometer, and a magnetometer. Other examples of the sensor 302 include a proximity sensor or a position sensor (e.g., near field communication, GPS, Bluetooth, Wifi), an audio sensor (e.g., a microphone), or any suitable combination thereof. Note that the sensor 302 described herein is for illustrative purposes and, thus, the sensor 302 is not limited to the sensors described above. In one example embodiment, the sensor 302 may include a depth sensor, such as a structured light sensor, a time-of-flight sensor, a passive stereo sensor, and ultrasonic devices, a time-of-flight sensor.
[0044] The display 304 includes a screen or a monitor configured to display images generated by the processor 308. In one example embodiment, the display 304 may be transparent or translucent such that the user 106 can view through the display 304 (in an AR use case). In another example, the display 304 (e.g., an LCOS display) presents each frame of the virtual content in multiple presentations.
[0045] The processor 308 includes an AR application 310, a 6DOF tracker 312, a depth system 324, and a collaborative sign language system 328. The AR application 310 uses computer vision to detect and recognize a hand 108 or an item in the real-world environment 102. The AR application 310 retrieves a virtual object (e.g., a 3D object model) based on the recognized item (e.g., the hand 108) or the physical environment. The display 304 displays the virtual object. The AR application 310 includes a local rendering engine that generates a visualization of the virtual object superimposed on an image of the item captured by the optical sensor 314 (e.g., superimposed thereon or otherwise displayed simultaneously therewith). The visualization of the virtual object can be manipulated by adjusting the positioning of the item relative to the optical sensor 314 (e.g., its physical position, orientation, or both). Similarly, the visualization of the virtual object can be manipulated by adjusting the pose of the AR device A 110 relative to the item.
[0046] The 6DOF tracker 312 estimates the pose of the AR device A 110. For example, the 6DOF tracker 312 uses the image data from the optical sensor 314 and the corresponding inertial data from the inertial sensor 316 to track the position and pose of the AR device A 110 relative to a reference frame (e.g., the real-world environment 102). In one example, the 6DOF tracker 312 uses the sensor data to determine the three-dimensional pose of the AR device A 110. The three-dimensional pose is the determined orientation and positioning of the AR device A 110 relative to the user's real-world environment 102. For example, the AR device A 110 may use the image of the user's real-world environment 102 and other sensor data to identify the relative positioning and orientation of the AR device A 110 and the physical objects in the real-world environment 102 around the AR device A 110. The 6DOF tracker 312 continuously collects and uses the updated sensor data describing the movement of the AR device A 110 to determine the updated three-dimensional pose of the AR device A 110, and the updated three-dimensional pose indicates the change in the relative positioning and orientation of the AR device A 110 and the physical objects in the real-world environment 102. The 6DOF tracker 312 provides the three-dimensional pose of the AR device A 110 to the depth system 324 and the collaborative sign language system 328.
[0047] The depth system 324 accesses the images from the optical sensor 314 and the sparse 3D points from the 6DOF tracker 312 to predict depth and generate a dense point cloud. In one example implementation, the optical sensor 314 includes a depth sensor or a stereo sensor.
[0048] The collaborative sign language system 328 coordinates the distribution of the processing of the sign language translation application between the AR device B 114 and the AR device A 110 (and optionally the server 116). For example, the collaborative sign language system 328 identifies the available computing resources and technical specifications of the AR device B 114 and the AR device A 110, and distributes one or more processes from the sign language translation between the AR device B 114 and the AR device A 110 based on the available computing resources and technical specifications. In another example, the AR device A 110 receives image data, image metadata (e.g., timestamp), and pose information from the AR device B 114, synchronizes the AR device A 110 with the AR device B 114 using the received information, performs sign language translation based on the synchronized information, displays the sign language translation in the display 304, and provides the sign language translation to the AR device B 114. In another example, the collaborative sign language system 328 displays a proposed translation in the display 304, waits for the user 106 to confirm the proposed translation, detects the confirmation by detecting a predefined pose indicating a sign language suggestion confirmation pose, and sends the confirmed translation to the AR device B 114.
[0049] The graphics processing unit 318 includes a rendering engine (not shown) configured to render frames of a 3D model of a virtual object based on virtual content provided by the AR application 310 and the pose of the AR device B 114 (relative to the AR device A 110). In other words, the graphics processing unit 318 uses the three-dimensional pose of the AR device B 114 to generate frames of virtual content to be presented on the display 304. For example, the graphics processing unit 318 uses the three-dimensional pose to render frames of virtual content such that the virtual content is presented in a certain orientation and position in the display 304 to appropriately enhance the user's sense of reality. As an example, the graphics processing unit 318 may use the three-dimensional pose data to render frames of virtual content such that when presented on the display 304, the virtual content overlaps with physical objects in the user's real-world environment 102. The graphics processing unit 318 generates updated frames of virtual content based on the updated three-dimensional pose of the AR device B 114, which reflects changes in the user's positioning and orientation relative to physical objects in the user's real-world environment 102.
[0050] The graphics processing unit 318 transmits the rendered frames to the display controller 320. The display controller 320 is positioned as an intermediary between the graphics processing unit 318 and the display 304, receives image data (e.g., the rendered frames) from the graphics processing unit 318, and provides the rendered frames to the display 304.
[0051] The storage device 306 stores virtual object content 322 and pose data 326 (e.g., predefined poses indicating confirmation of suggested words). The virtual object content 322 includes, for example, a database of visual references (e.g., images, QR codes) and corresponding virtual content (e.g., 3D models of virtual objects). The pose data 326 indicates predefined gestures or predefined user inputs to indicate confirmation of suggested words translated by the collaborative sign language system 328 or by another translation system (located on the AR device B 114 or the server 116). The pose data 326 may also define other trigger conditions that, when detected, enable the user 106 to select suggested options to avoid finger spelling or accelerate sign language translation.
[0052] Any one or more of the modules described herein may be implemented using hardware (e.g., a processor of a machine) or a combination of hardware and software. For example, any module described herein may configure a processor to perform the operations described herein for that module. Additionally, any two or more of these modules may be combined into a single module, and the functions described herein for a single module may be subdivided among multiple modules. Further, according to various example embodiments, modules described herein as being implemented within a single machine, database, or device may be distributed across multiple machines, databases, or devices.
[0053] Figure 4 It is a block diagram showing a collaborative sign language system 328 according to an example embodiment. The collaborative sign language system 328 includes an AR device interface 402, a synchronization module 404, a hand tracking application 406, and a sign language recognition application 408.
[0054] The AR device interface 402 communicates with the AR device B 114. In one example, the AR device interface 402 accesses 6DOF pose data and camera device data from the AR device B 114. In another example, the AR device interface 402 shares the 6DOF pose data of the 6DOF tracker 312 and the camera device data of the optical sensor 314 with the AR device B 114. In another example, the AR device interface 402 conveys what processes the AR device B 114 will perform as part of the collaborative sign language translation process.
[0055] The synchronization module 404 accesses the 6DOF pose data from the 6DOF tracker 312, the camera device data from the optical sensor 314, the 6DOF pose data and camera device data (including timestamp metadata) from the AR device B 114. In another example, the synchronization module 404 accesses the technical specifications and available computing resources of the AR device A 110 and the AR device B 114. The synchronization module 404 synchronizes the AR device A 110 with the AR device B 114 by matching the timestamps of the images from the AR device A 110 with the images from the AR device B 114. In another example, synchronization can be performed based on aligning / registering the 6DOF poses of the camera devices of the AR device A 110 and the AR device B 114 in the same common coordinate system.
[0056] The hand tracking application 406 uses computer vision to detect the hand 108. In one example, the hand tracking application 406 identifies the features (motion, bones, depth) of the hand and tracks the motion of the hand 108. The hand tracking application 406 can include multiple processes, such as, for example, a hand tracking process, a gesture detection process, and a sign language translation process.
[0057] The sign language recognition application 408 recognizes the gesture of the hand 108 based on the hand tracking application 406 and recognizes the word corresponding to the gesture. In one example, the sign language recognition application 408 includes a sign language translation process. In another exemplary implementation, the sign language recognition application 408 recognizes suggested words based on a combination of context information of the AR device A 110 (e.g., the profile of the user 106, the geographical location of the AR device A 110, the historical input of the user 106, the text corpus, and predefined terms). The sign language recognition application 408 provides the text (indicating the suggested words) to the display 304. The display 304 displays the text. In another example, the sign language recognition application 408 transmits the text to the AR device B 114, where the text is displayed on the screen of the AR device B 114.
[0058] Figure 5 is an interaction diagram showing the interaction between collaborative augmented reality devices (e.g., the AR device A 110 and the AR device B 114) according to an exemplary implementation. The AR device B 114 captures an image of the hand 108 with its camera device. The AR device A 110 captures an image of the hand 108 with its camera device. At 502, the AR device B 114 sends the image, camera device data, and 6DOF pose data to the AR device A 110. At 504, the AR device A 110 synchronizes the AR device A 110 and the AR device B 114 based on the image data from both the AR device A 110 and the AR device B 114. At 506, the AR device A 110 applies a hand tracking process based on the synchronized images (e.g., a combination of the image data from the AR device A 110 and the AR device B 114). At 508, the AR device A 110 applies a gesture recognition process to the synchronized images to generate a sign language translation. At 510, the AR device A 110 displays the sign language translation in the display 304 and provides the sign language translation to the AR device B 114.
[0059] Figure 6An interaction diagram showing the interaction between collaborative augmented reality devices (e.g., AR device A 110 and AR device B 114) according to an example embodiment. AR device B 114 captures an image of the hand 108 with its camera device. AR device A 110 captures an image of the hand 108 with its camera device. At 602, AR device B 114 sends the image, camera device data, and 6DOF pose data to AR device A 110. At 604, AR device A 110 synchronizes AR device A 110 and AR device B 114 based on the image data from both AR device A 110 and AR device B 114. At 606, AR device A 110 applies hand tracking processing based on the synchronized images (e.g., a combination of the image data from AR device A 110 and AR device B 114). At 608, AR device A 110 applies a gesture recognition process to the synchronized images to generate a proposed sign language translation. At 610, AR device A 110 displays the proposed translation to the user 106 in the display 304. At 612, AR device A 110 detects confirmation of the proposed translation for the user 106. At 614, AR device A 110 shares the translation with AR device B 114.
[0060] Figure 7 An interaction diagram showing the interaction between collaborative augmented reality devices (e.g., AR device A 110 and AR device B 114) according to an example embodiment. AR device B 114 captures an image of the hand 108 with its camera device. AR device A 110 captures an image of the hand 108 with its camera device. At 702, AR device B 114 sends the image, camera device data, and 6DOF pose data to AR device A 110. At 704, AR device A 110 synchronizes AR device A 110 and AR device B 114 based on the image data from both AR device A 110 and AR device B 114. At 706, AR device A 110 applies hand tracking processing based on the synchronized images (e.g., a combination of the image data from AR device A 110 and AR device B 114). AR device A 110 sends the result of the hand tracking processing to AR device B 114. At 708, AR device B 114 applies a gesture recognition process to the result of the hand tracking processed at AR device A 110. At 710, AR device B 114 applies a sign language recognition process to the result of the gesture recognition process. At 712, AR device B 114 displays the sign language translation (e.g., text) in the display of AR device B 114. At 714, AR device B 114 provides the text to AR device A 110. At 716, AR device A 110 displays the text.
[0061] Figure 8It is an interaction diagram showing the interaction between collaborative augmented reality devices (e.g., AR device A 110 and AR device B 114) according to an example embodiment. AR device B 114 captures an image of the hand 108 with its camera device. AR device A 110 captures an image of the hand 108 with its camera device. At 802, AR device B 114 sends the image, camera device data, and 6DOF pose data to AR device A 110. At 804, AR device A 110 synchronizes AR device A 110 and AR device B 114 based on the image data from both AR device A 110 and AR device B 114. At 806, AR device A 110 provides an egocentric view image to AR device B 114. At 808, AR device A 110 applies hand tracking processing based on the synchronized images (e.g., a combination of the image data from AR device B 114 and the egocentric view data from AR device A 110). At 810, AR device B 114 applies pose recognition processing and sign language recognition processing to the result of the hand tracking processing. At 812, AR device B 114 displays a sign language translation (e.g., text) on the display of AR device B 114. At 814, AR device B 114 provides the text to AR device A 110 for display.
[0062] Figure 9 It is a flowchart showing a method 900 for sign language translation according to an example embodiment. The operations in method 900 can be performed by the collaborative sign language system 328 using the components (e.g., modules, engines) described above with respect to Figure 4 However, it should be understood that at least some of the operations in method 900 can be deployed on various other hardware configurations, or be performed by similar components residing elsewhere.
[0063] In block 902, the synchronization module 404 synchronizes the signer's AR device (e.g., AR device A 110) with the receiver's AR device (e.g., AR device B 114). In block 904, the hand tracking application 406 performs hand tracking, and the sign language recognition application 408 performs sign language recognition on the synchronized images. In block 906, the display 304 displays a proposed translation based on sign language recognition on the signer's AR device. In block 908, the sign language recognition application 408 detects confirmation of the proposed translation on the signer's AR device. In block 910, the sign language recognition application 408 provides the proposed translation to the receiver's AR device.
[0064] Note that other embodiments may use different orderings, additional or fewer operations, and different nomenclatures or terms to accomplish similar functions. In some embodiments, the various operations may be performed in parallel with other operations in a synchronous or asynchronous manner. The operations described herein are chosen to illustrate some operational principles in a simplified form.
[0065] Figure 10 is a flowchart showing a method for allocating sign language translation processing according to an example embodiment. The operations in method 1000 may be performed by the collaborative sign language system 328 using the components (e.g., modules, engines) described above with respect to Figure 4 Thus, method 1000 is described by way of example with reference to the collaborative sign language system 328. However, it should be understood that at least some of the operations in method 1000 may be deployed on various other hardware configurations or performed by similar components residing elsewhere.
[0066] In block 1002, the collaborative sign language system 328 synchronizes the AR device of the signer with the AR device of the receiver. In block 1004, the collaborative sign language system 328 allocates one or more processes from hand tracking, gesture recognition, and sign language recognition between the AR device of the signer and the AR device of the receiver. In block 1006, the collaborative sign language system 328 shares the results of the processes between the AR device of the signer and the AR device of the receiver. In block 1008, the display 304 displays a translation based on the shared results.
[0067] Note that other embodiments may use different orderings, additional or fewer operations, and different nomenclatures or terms to accomplish similar functions. In some embodiments, the various operations may be performed in parallel with other operations in a synchronous or asynchronous manner. The operations described herein are chosen to illustrate some operational principles in a simplified form.
[0068] System with a head-mounted device
[0069] Figure 11 illustrates a network environment 1100 in which a head-mounted device 1102 may be implemented according to an example embodiment. Figure 11 is a high-level functional block diagram of an example head-mounted device 1102 that communicatively couples a mobile client device 1138 and a server system 1132 via various networks 1140.
[0070] The head-wearable device 1102 includes a camera device, such as at least one of a visible light camera device 1112, an infrared emitter 1114, and an infrared camera device 1116. The client device 1138 may be able to connect to the head-wearable device 1102 using both communications 1134 and 1136. The client device 1138 is connected to the server system 1132 and the network 1140. The network 1140 may include any combination of wired and wireless connections.
[0071] The head-wearable device 1102 also includes two image displays of an image display 1104 of an optical component. The two image displays include one image display associated with the left side of the head-wearable device 1102 and one image display associated with the right side of the head-wearable device 1102. The head-wearable device 1102 also includes an image display driver 1108, an image processor 1110, a low-power low-power circuitry 1126, and a high-speed circuitry 1118. The image display 1104 of the optical component is used to present images and videos to a user of the head-wearable device 1102, including images that may include a graphical user interface.
[0072] The image display driver 1108 commands and controls the image displays in the image display 1104 of the optical component. The image display driver 1108 may deliver image data directly to the image displays in the image display 1104 of the optical component for presentation, or may have to convert the image data into a signal or data format suitable for delivery to an image display device. For example, the image data may be video data formatted according to a compression format (e.g., H.264 (MPEG-4 Part 10), HEVC, Theora, Dirac, RealVideo RV40, VP8, VP9, etc.), and still image data may be formatted according to a compression format (e.g., Portable Network Graphics (PNG), Joint Photographic Experts Group (JPEG), Tagged Image File Format (TIFF), or Exchangeable Image File Format (Exif), etc.).
[0073] As described above, the head-wearable device 1102 includes a frame and a stem (or temple) extending from a side of the frame. The head-wearable device 1102 also includes a user input device 1106 (e.g., a touch sensor or a push button), including an input surface on the head-wearable device 1102. The user input device 1106 (e.g., a touch sensor or a push button) is used to receive input selections from a user to manipulate a graphical user interface of the presented images.
[0074] Figure 11The components for the head - wearable device 1102 shown are located on one or more circuit boards (e.g., PCB or flexible PCB) in the frame or temple. Alternatively or additionally, the depicted components may be located in the block, frame, hinge, or bridge of the head - wearable device 1102. The left and right may include digital camera device elements such as complementary metal - oxide - semiconductor (CMOS) image sensors, charge - coupled devices, camera device lenses, or any other corresponding visible - light or light - capturing elements that can be used to capture data, including images of scenes with unknown objects.
[0075] The head - wearable device 1102 includes a memory 1122 that stores instructions for performing a subset or all of the functions described herein. The memory 1122 may also include a storage device.
[0076] As Figure 11 shown, the high - speed circuitry 1118 includes a high - speed processor 1120, a memory 1122, and a high - speed wireless circuitry 1124. In this example, the image display driver 1108 is coupled to the high - speed circuitry 1118 and is operated by the high - speed processor 1120 to drive the left and right image displays in the image display 1104 of the optical assembly. The high - speed processor 1120 can be any processor capable of managing high - speed communication and the operation of any general - purpose computing system required for the head - wearable device 1102. The high - speed processor 1120 includes the processing resources required to manage high - speed data transfer over communication 1136 to a wireless local area network (WLAN) using the high - speed wireless circuitry 1124. In some examples, the high - speed processor 1120 executes an operating system (e.g., the LINUX operating system or other such operating system for the head - wearable device 1102), and the operating system is stored in the memory 1122 for execution. In addition to any other duties, the high - speed processor 1120 that executes the software architecture of the head - wearable device 1102 manages data transfer with the high - speed wireless circuitry 1124. In some examples, the high - speed wireless circuitry 1124 is configured to implement the Institute of Electrical and Electronics Engineers (IEEE) 802.11 communication standard (also referred to herein as Wi - Fi). In other examples, other high - speed communication standards may be implemented by the high - speed wireless circuitry 1124.
[0077] The low - power wireless circuitry 1130 and the high - speed wireless circuitry 1124 of the head - wearable device 1102 may include a short - range transceiver (Bluetooth TM) and a wireless wide area network, local area network, or wide area network transceiver (e.g., cellular or WiFi). The client device 1138, which includes a transceiver for communicating via communication 1134 and communication 1136, can be implemented using the details of the architecture of the head wearable device 1102, and so can other elements of the network 1140.
[0078] The memory 1122 includes any storage device capable of storing various data and applications, including camera device data generated by the left and right, infrared camera devices 1116 and the image processor 1110, and images for display generated by the image display driver 1108 on the image display of the optical component 1104, etc. Although the memory 1122 is shown as integrated with the high-speed circuitry 1118, in other examples, the memory 1122 can be a separate stand-alone element of the head wearable device 1102. In some such examples, wire-by-wire circuitry can provide a connection from the image processor 1110 or the low-power processor 1128 to the memory 1122 through a chip including the high-speed processor 1120. In other examples, the high-speed processor 1120 can manage the addressing of the memory 1122 such that the low-power processor 1128 will initiate the high-speed processor 1120 at any time when a read or write operation involving the memory 1122 is needed.
[0079] As Figure 11 shown, the low-power processor 1128 or the high-speed processor 1120 of the head wearable device 1102 can be coupled to a camera device (visible light camera device 1112; infrared emitter 1114 or infrared camera device 1116), the image display driver 1108, the user input device 1106 (e.g., touch sensor or push button), and the memory 1122.
[0080] The head wearable device 1102 is connected to a host computer. For example, the head wearable device 1102 is paired with the client device 1138 via communication 1136, or is connected to the server system 1132 via the network 1140. For example, the server system 1132 can be one or more computing devices that are part of a service or network computing system, which includes a processor, a memory, and a network communication interface to communicate with the client device 1138 and the head wearable device 1102 via the network 1140.
[0081] The client device 1138 includes a processor and a network communication interface coupled to the processor. The network communication interface allows communication via the network 1140, communication 1134, or communication 1136. The client device 1138 can also store at least a portion of the instructions for generating binaural audio content in the memory of the client device 1138 to implement the functions described herein.
[0082] The output components of the head-wearable device 1102 include visual components such as a display (e.g., a liquid crystal display (LCD), a plasma display panel (PDP), a light-emitting diode (LED) display, a projector, or a waveguide). The image display of the optical component is driven by an image display driver 1108. The output components of the head-wearable device 1102 also include acoustic components (e.g., speakers), haptic components (e.g., vibration motors), other signal generators, etc. The input components (e.g., user input device 1106) of the head-wearable device 1102, the client device 1138, and the server system 1132 may include alphanumeric input components (e.g., a keyboard, a touch screen configured to receive alphanumeric input, an optoelectronic keyboard, or other alphanumeric input components), pointing-based input components (e.g., a mouse, a touchpad, a trackball, a joystick, a motion sensor, or other pointing instruments), haptic input components (e.g., physical buttons, a touch screen that provides the position and force of a touch or touch gesture, or other haptic input components), audio input components (e.g., a microphone), etc.
[0083] The head-wearable device 1102 may optionally include additional peripheral device elements. Such peripheral device elements may include biometric sensors, additional sensors, or display elements integrated with the head-wearable device 1102. For example, the peripheral device elements may include any I / O components, including output components, motion components, positioning components, or any other such elements described herein.
[0084] For example, biometric components include components that detect expressions (e.g., hand expressions, facial expressions, voice expressions, body postures, or eye tracking), measure biological signals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), identify people (e.g., voice recognition, retina recognition, facial recognition, fingerprint recognition, or electroencephalogram-based recognition), etc. Motion components include acceleration sensor components (e.g., accelerometers), gravity sensor components, rotational sensor components (e.g., gyroscopes), etc. Positioning components include position sensor components that generate position coordinates (e.g., a global positioning system (GPS) receiver component), WiFi or Bluetooth TM transceivers, altitude sensor components (e.g., an altimeter or barometer that detects air pressure, from which altitude can be obtained), orientation sensor components (e.g., magnetometers), etc. Such positioning system coordinates may also be received from the client device 1138 via the low-power radio circuitry 1130 or the high-speed radio circuitry 1124 through communication 1136.
[0085] When using phrases such as "at least one of A, B, or C", "at least one of A, B, and C", "one or more of A, B, or C", or "one or more of A, B, and C", it is intended that the phrase be interpreted to mean that A can exist alone in an embodiment, B can exist alone in an embodiment, C can exist alone in an embodiment, or any combination of elements A, B, and C can exist in a single embodiment; for example, A and B, A and C, B and C, or A, B, and C.
[0086] Without departing from the scope of the present disclosure, changes and modifications may be made to the disclosed embodiments. These and other changes or modifications are intended to be included within the scope of the present disclosure as expressed in the appended claims.
[0087] Figure 12 is a block diagram 1200 showing a software architecture 1204 that may be installed on any one or more of the devices described herein. The software architecture 1204 is supported by hardware such as a machine 1202 that includes a processor 1220, a memory 1226, and I / O components 1238. In this example, the software architecture 1204 may be conceptualized as a stack of layers, where each layer provides a specific function. The software architecture 1204 includes layers such as an operating system 1212, libraries 1210, frameworks 1208, and applications 1206. In operation, the application 1206 makes API calls 1250 through the software stack and receives messages 1252 in response to the API calls 1250.
[0088] The operating system 1212 manages hardware resources and provides common services. The operating system 1212 includes, for example, a kernel 1214, services 1216, and drivers 1222. The kernel 1214 serves as an abstraction layer between the hardware and other software layers. For example, the kernel 1214 provides memory management, processor management (e.g., scheduling), component management, networking, and security settings, as well as other functions. The services 1216 may provide other common services to other software layers. The drivers 1222 are responsible for controlling or interfacing with the underlying hardware. For example, the drivers 1222 may include a display driver, a camera device driver, or a low-power driver, a flash driver, a serial communication driver (e.g., a universal serial bus (USB) driver), drivers, an audio driver, a power management driver, and the like.
[0089] The library 1210 provides low-level common infrastructure used by the application 1206. The library 1210 may include a system library 1218 (e.g., the C standard library), which provides functions such as memory allocation functions, string manipulation functions, mathematical functions, etc. Additionally, the library 1210 may include an API library 1224, such as a media library (e.g., a library for supporting the presentation and manipulation of various media formats, such as Moving Picture Experts Group-4 (MPEG4), High Efficiency Video Coding (H.264 or AVC), Moving Picture Experts Group Layer-3 (MP3), Advanced Audio Coding (AAC), Adaptive Multi-Rate (AMR) audio codec, Joint Photographic Experts Group (JPEG or JPG), or Portable Network Graphics (PNG)), a graphics library (e.g., the OpenGL framework for 2D and 3D rendering in graphical content on a display), a database library (e.g., SQLite that provides various relational database functions), a web library (e.g., WebKit that provides web browsing functions), etc. The library 1210 may also include various other libraries 1228 to provide many other APIs to the application 1206.
[0090] The framework 1208 provides high-level common infrastructure used by the application 1206. For example, the framework 1208 provides various Graphical User Interface (GUI) functions, advanced resource management, and advanced location services. The framework 1208 may provide a wide range of other APIs that can be used by the application 1206, some of which may be specific to a particular operating system or platform.
[0091] In an example embodiment, the application 1206 may include a home application 1236, a contacts application 1230, a browser application 1232, a book reader application 1234, a location application 1242, a media application 1244, a messaging application 1246, a gaming application 1248, and various other applications such as third-party applications 1240. The application 1206 is a program that executes functions defined in the program. One or more of the applications in the application 1206 can be created using various programming languages, such as object-oriented programming languages (e.g., Objective-C, Java, or C++) or procedural programming languages (e.g., the C language or assembly language). In a specific example, a third-party application 1240 (e.g., an application developed using an ANDROID TM or IOS TM software development kit (SDK)) can be an application on platforms such as IOS TM 、ANDROID TM 、 Mobile software running on the mobile operating system of a Phone or another mobile operating system. In this example, the third-party application 1240 can call the API call 1250 provided by the operating system 1212 to facilitate the functions described herein.
[0092] Figure 13 is a graphical representation of a machine 1300 within which instructions 1308 (e.g., software, program, application, applet, app, or other executable code) can be executed to cause the machine 1300 to perform any one or more of the methods discussed herein. For example, the instructions 1308 can cause the machine 1300 to perform any one or more of the methods described herein. The instructions 1308 transform the general unprogrammed machine 1300 into a particular machine 1300 programmed to perform the described and illustrated functions in the described manner. The machine 1300 can operate as a stand-alone device or can be coupled (e.g., networked) to other machines. In a networked deployment, the machine 1300 can operate as a server machine or a client machine in a server-client network environment, or as a peer machine in a peer-to-peer (or distributed) network environment. The machine 1300 can include, but is not limited to: server computers, client computers, personal computers (PCs), tablet computers, laptop computers, netbooks, set-top boxes (STBs), PDAs, entertainment media systems, cellular phones, smartphones, mobile devices, wearable devices (e.g., smartwatches), smart home devices (e.g., smart appliances), other smart devices, web devices, network routers, network switches, network bridges, or any machine capable of sequentially or otherwise executing the instructions 1308 specifying the actions to be taken by the machine 1300. Moreover, although only a single machine 1300 is shown, the term "machine" shall also be regarded as including a collection of machines that individually or jointly execute the instructions 1308 to perform any one or more of the methods discussed herein.
[0093] The machine 1300 can include a processor 1302, a memory 1304, and I / O components 1342 configured to communicate with each other via a bus 1344. In an example embodiment, the processor 1302 (e.g., a central processing unit (CPU), a reduced instruction set computing (RISC) processor, a complex instruction set computing (CISC) processor, a graphics processing unit (GPU), a digital signal processor (DSP), an ASIC, a radio frequency integrated circuit (RFIC), other processors, or any suitable combination thereof) can include, for example, a processor 1306 and a processor 1310 that execute the instructions 1308. The term "processor" is intended to include multi-core processors, which can include two or more independent processors (sometimes referred to as "cores") that can execute instructions simultaneously. AlthoughFigure 13 Multiple processors 1302 are shown, but machine 1300 can include a single processor with a single core, a single processor with multiple cores (e.g., a multi-core processor), multiple processors with a single core, multiple processors with multiple cores, or any combination thereof.
[0094] Memory 1304 includes main memory 1312, static memory 1314, and storage unit 1316, all of which are accessible by processor 1302 via bus 1344. Main memory 1304, static memory 1314, and storage unit 1316 store instructions 1308 embodying any one or more of the methods or functions described herein. The instructions 1308 may also reside, completely or partially, within main memory 1312, within static memory 1314, within machine-readable medium 1318 within storage unit 1316, within at least one of the processors 1302 (e.g., within a cache memory of the processor), or within any suitable combination thereof, during execution by machine 1300.
[0095] I / O components 1342 can include a variety of components that receive input, provide output, generate output, transmit information, exchange information, capture measurements, and so on. The specific I / O components 1342 included in a particular machine will depend on the type of machine. For example, a portable machine such as a mobile phone may include a touch input device or other such input mechanism, while a headless server machine is less likely to include such a touch input device. It should be understood that I / O components 1342 can include many other components not shown in Figure 13 In various example embodiments, I / O components 1342 can include output components 1328 and input components 1330. Output components 1328 can include visual components (e.g., a display such as a plasma display panel (PDP), a light-emitting diode (LED) display, a liquid crystal display (LCD), a projector, or a cathode ray tube (CRT)), acoustic components (e.g., speakers), haptic components (e.g., a vibration motor, a resistance mechanism), other signal generators, and so on. Input components 1330 can include alphanumeric input components (e.g., a keyboard, a touch screen configured to receive alphanumeric input, an optical keyboard, or other alphanumeric input components), pointing-based input components (e.g., a mouse, a touchpad, a trackball, a joystick, a motion sensor, or other pointing instrument), haptic input components (e.g., a physical button, a touch screen that provides the location and / or force of a touch or touch gesture, or other haptic input components), audio input components (e.g., a microphone), and so on.
[0096] In other example embodiments, the I / O component 1342 may include a biometric component 1332, a motion component 1334, an environmental component 1336, or a positioning component 1338, as well as a variety of other components. For example, the biometric component 1332 includes components for detecting expressions (e.g., hand expressions, facial expressions, voice expressions, body postures, or eye tracking), measuring biometric signals (e.g., blood pressure, heart rate, body temperature, sweating, or brain waves), identifying people (e.g., voice recognition, retina recognition, facial recognition, fingerprint recognition, or electroencephalogram-based recognition), and so on. The motion component 1334 includes an acceleration sensor component (e.g., an accelerometer), a gravity sensor component, a rotation sensor component (e.g., a gyroscope), and so on. The environmental component 1336 includes, for example, a lighting sensor component (e.g., a photometer), a temperature sensor component (e.g., one or more thermometers for detecting ambient temperature), a humidity sensor component, a pressure sensor component (e.g., a barometer), an auditory sensor component (e.g., one or more microphones for detecting background noise), a proximity sensor component (e.g., an infrared sensor for detecting nearby objects), a gas sensor (e.g., a gas detection sensor for detecting the concentration of hazardous gases to ensure safety or measuring pollutants in the atmosphere), or other components that can provide an indication, measurement, or signal corresponding to the surrounding physical environment. The positioning component 1338 includes a position sensor component (e.g., a GPS receiver component), an altitude sensor component (e.g., an altimeter or barometer that detects air pressure, from which altitude can be obtained), an orientation sensor component (e.g., a magnetometer), and so on.
[0097] A variety of techniques can be used to implement communication. The I / O component 1342 also includes a communication component 1340, which is operable to couple the machine 1300 to the network 1320 or the device 1322 via the couplings 1324 and 1326, respectively. For example, the communication component 1340 may include a network interface component or another suitable device to interface with the network 1320. In other examples, the communication component 1340 may include a wired communication component, a wireless communication component, a cellular communication component, a near field communication (NFC) component, components (e.g., low power), components, and other communication components for providing communication via other modalities. The device 1322 may be another machine or any of a variety of peripheral devices (e.g., a peripheral device coupled via USB).
[0098] In addition, the communication component 1340 can detect an identifier or include components operable to detect an identifier. For example, the communication component 1340 can include a radio frequency identification (RFID) tag reader component, an NFC smart tag detection component, an optical reader component (e.g., an optical sensor for detecting one-dimensional barcodes such as Universal Product Code (UPC) barcodes, multi-dimensional barcodes such as Quick Response (QR) codes, Aztec codes, Data Matrix, Dataglyph, MaxiCode, PDF417, Ultra Code, UCC RSS-2D barcodes, and other optical codes), or an acoustic detection component (e.g., a microphone for identifying an audio signal of a tag). Additionally, various information can be obtained via the communication component 1340, such as a location via Internet Protocol (IP) geolocation, a location via signal triangulation, a location via detecting an NFC beacon signal that can indicate a specific location, and so on.
[0099] Various memories (e.g., memory 1304, main memory 1312, static memory 1314, and / or the memory of the processor 1302) and / or the storage unit 1316 can store a set or more sets of instructions and data structures (e.g., software) that embody any one or more of the methods or functions described herein or are used by any one or more of the methods or functions described herein. These instructions (e.g., instruction 1308), when executed by the processor 1302, cause various operations to implement the disclosed embodiments.
[0100] The instructions 1308 can be sent or received via a network interface device (e.g., the network interface component included in the communication component 1340) using a transmission medium and using any one of a plurality of well-known transmission protocols (e.g., Hypertext Transfer Protocol (HTTP)) over the network 1320. Similarly, the instructions 1308 can be sent or received to the device 1322 via the coupling 1326 (e.g., a peer-to-peer coupling) using a transmission medium.
[0101] As used herein, the terms "machine storage medium", "device storage medium", and "computer storage medium" mean the same thing and may be used interchangeably in this disclosure. These terms refer to a single or multiple storage devices and / or media that store executable instructions and / or data (e.g., a centralized or distributed database, and / or associated caches and servers). Thus, these terms should be regarded as including, but not limited to, solid state memories as well as optical and magnetic media, including memories internal or external to a processor. Specific examples of machine storage medium, computer storage medium, and / or device storage medium include: non-volatile memories, including, for example, semiconductor memory devices such as erasable programmable read only memory (EPROM), electrically erasable programmable read only memory (EEPROM), field programmable gate array (FPGA), and flash memory devices; magnetic disks, such as internal hard disks and removable disks; magneto-optical disks; and CD-ROM and DVD-ROM disks. The terms "machine storage medium", "computer storage medium", and "device storage medium" expressly exclude carrier waves, modulated data signals, and other such media, at least some of which are covered by the term "signal medium" discussed below.
[0102] The terms "transmission medium" and "signal medium" mean the same thing and may be used interchangeably in this disclosure. The terms "transmission medium" and "signal medium" should be understood to include any non-transitory medium that is capable of storing, encoding, or carrying instructions 1416 for execution by machine 1400, and includes digital or analog communication signals or other non-transitory media that facilitate the communication of such software. Thus, the terms "transmission medium" and "signal medium" should be regarded as including any form of modulated data signal, carrier wave, etc. The term "modulated data signal" means a signal that has one or more of its characteristics set or changed in such a manner as to encode information in the signal.
[0103] The terms "machine-readable medium", "computer-readable medium", and "device-readable medium" mean the same thing and may be used interchangeably in this disclosure. These terms are defined to include both machine storage medium and transmission medium. Thus, these terms include both storage devices / media and carrier waves / modulated data signals.
[0104] Although embodiments have been described with reference to specific example embodiments, it will be apparent that various modifications and changes can be made to these embodiments without departing from the broader scope of the disclosure. Accordingly, the specification and drawings are to be regarded in an illustrative rather than a restrictive sense. The accompanying drawings, which form a part of this invention, illustrate, by way of example and not of limitation, specific embodiments in which the subject matter can be practiced. The illustrated embodiments are described in sufficient detail to enable those skilled in the art to practice the teachings disclosed herein. Other embodiments can be utilized and logical substitutions and changes can be made therefrom without departing from the scope of the disclosure. Accordingly, this detailed description should not be construed in a limiting sense, and the scope of the various embodiments is defined only by the appended claims and the full scope of equivalents to which such claims are entitled.
[0105] Such embodiments of the subject matter of the present invention can be referred to herein individually and / or collectively by the term "invention" merely for convenience and are not intended to voluntarily limit the scope of this application to any single invention or inventive concept if more than one invention or inventive concept is in fact disclosed. Accordingly, although specific embodiments have been shown and described herein, it should be understood that any arrangement calculated to achieve the same purpose can be substituted for the specific embodiments shown. The disclosure is intended to cover any and all changes or variations of various embodiments. Combinations of the above-described embodiments and other embodiments not specifically described herein will be apparent to those skilled in the art upon reviewing the above description.
[0106] A summary of the disclosure is provided to enable the reader to quickly ascertain the nature of the technical disclosure. It is submitted with the understanding that it will not be used to interpret or limit the scope or meaning of the claims. Additionally, in the foregoing detailed description, it can be seen that for the purposes of simplifying the disclosure, various features are combined in a single embodiment. This method of the disclosure should not be interpreted as reflecting an intention that the claimed embodiments require more features than are expressly recited in each claim. Rather, as reflected in the appended claims, the subject matter of the present invention lies in less than all of the features of a single disclosed embodiment. Accordingly, the claims are hereby incorporated into the detailed description, where each claim stands on its own as a separate embodiment.
[0107] Example
[0108] Example 1 is a method that includes: accessing a first image generated by a first imaging device of a first augmented reality device and a second image generated by a second imaging device of a second augmented reality device, where the first image and the second image depict a gesture of a user of the first augmented reality device; synchronizing the first augmented reality device with the second augmented reality device; in response to the synchronization, allocating one or more processes of a sign language recognition system between the first augmented reality device and the second augmented reality device, where the one or more processes are executed on corresponding augmented reality devices; collecting results from the one or more processes from the first augmented reality device and from the second augmented reality device; and based on the results, displaying text indicating a sign language translation of the gesture in a first display of the first augmented reality device or in a second display of the second augmented reality device in near real-time.
[0109] Example 2 includes the method of Example 1, where synchronizing the first augmented reality device with the second augmented reality device further includes: mapping a first timestamp of the first image to a second timestamp of the second image.
[0110] Example 3 includes the method of Example 1, where synchronizing the first augmented reality device with the second augmented reality device further includes: registering a six-degree-of-freedom coordinate system of the first augmented reality device with a six-degree-of-freedom coordinate system of the second augmented reality device.
[0111] Example 4 includes the method of Example 1, and further includes: transmitting the first image to the second augmented reality device, where the first image provides an egocentric view of a gesture of a user of the first augmented reality device, and where the second augmented reality device is configured to perform one or more processes of a sign language recognition system on a combination of the first image, a first hand skeleton based on the first image, the second image, and a second hand skeleton based on the second image, where the one or more processes of the sign language recognition system include: a hand tracking process, a gesture detection process, and a sign language translation process.
[0112] Example 5 includes the method of Example 1, and further includes: identifying suggested words based on a combination of context information of the first augmented reality device, historical inputs of a user of the first augmented reality device, a text corpus, and predefined terms, where the text indicates the suggested words.
[0113] Example 6 includes the method of Example 5, and further includes: displaying the suggested words in a first display of the first augmented reality device; receiving confirmation of the suggested words from a user of the first augmented reality device; and in response to receiving the confirmation, providing the suggested words to the second augmented reality device.
[0114] Example 7 includes the method of Example 6, where receiving the confirmation includes: detecting a predefined gesture from the user, where the predefined gesture indicates confirmation.
[0115] Example 8 includes the method of Example 1, wherein allocating one or more processes of the sign language recognition system further includes: allocating one or more processes between a first augmented reality device and a second augmented reality device either temporally or spatially, and wherein the temporal allocation includes processing alternate frames between the first augmented reality device and the second augmented reality device, and wherein the spatial allocation includes processing a first gesture of a first hand from a user only with the first augmented reality device and processing a second gesture of a second hand from the user only with the second augmented reality device.
[0116] Example 9 includes the method of Example 1, wherein allocating one or more processes of the sign language recognition system further includes: detecting an occlusion of a gesture in a first image by the first augmented reality device; and in response to detecting the occlusion, assigning the second augmented reality device to perform one or more processes based on a second image.
[0117] Example 10 includes the method of Example 1, wherein the first augmented reality device includes a first head-wearable device, and wherein the second augmented reality device includes a mobile hand-held device or a second head-wearable device.
[0118] Example 11 is a computing device, including: a processor; and a memory storing instructions that, when executed by the processor, configure the device to: access a first image generated by a first imaging device of a first augmented reality device and a second image generated by a second imaging device of a second augmented reality device, the first image and the second image depicting a gesture of a user of the first augmented reality device; synchronize the first augmented reality device with the second augmented reality device; in response to the synchronization, allocate one or more processes of the sign language recognition system between the first augmented reality device and the second augmented reality device, wherein the one or more processes are executed on corresponding augmented reality devices; collect results from the one or more processes from the first augmented reality device and from the second augmented reality device; and based on the results, display text of a sign language translation indicating the gesture in a first display of the first augmented reality device or in a second display of the second augmented reality device in near real time.
[0119] Example 12 includes the computing device of Example 11, wherein synchronizing the first augmented reality device with the second augmented reality device further includes: mapping a first timestamp of the first image to a second timestamp of the second image.
[0120] Example 13 includes the computing device of Example 11, wherein synchronizing the first augmented reality device with the second augmented reality device further includes: registering a six-degree-of-freedom coordinate system of the first augmented reality device with a six-degree-of-freedom coordinate system of the second augmented reality device.
[0121] Example 14 includes the computing device of Example 11, wherein the instructions further configure the device to: transmit a first image to a second augmented reality device, the first image providing an egocentric view of the gestures of a user of the first augmented reality device, and wherein the second augmented reality device is configured to perform one or more processes of a sign language recognition system on a combination of the first image, a first hand skeleton based on the first image, a second image, and a second hand skeleton based on the second image, wherein the one or more processes of the sign language recognition system include: a hand tracking process, a gesture detection process, and a sign language translation process.
[0122] Example 15 includes the computing device of Example 11, wherein the instructions further configure the device to: identify suggested words based on a combination of context information of the first augmented reality device, historical inputs of a user of the first augmented reality device, a text corpus, and predefined terms, wherein the text indicates the suggested words.
[0123] Example 16 includes the computing device of Example 15, wherein the instructions further configure the device to: display the suggested words in a first display of the first augmented reality device; receive confirmation of the suggested words from a user of the first augmented reality device; and in response to receiving the confirmation, provide the suggested words to the second augmented reality device.
[0124] Example 17 includes the computing device of Example 16, wherein receiving the confirmation includes: detecting a predefined gesture from the user, the predefined gesture indicating confirmation.
[0125] Example 18 includes the computing device of Example 11, wherein allocating one or more processes of the sign language recognition system further includes: allocating one or more processes between the first augmented reality device and the second augmented reality device either temporally or spatially, and wherein the temporal allocation includes processing alternate frames between the first augmented reality device and the second augmented reality device, and wherein the spatial allocation includes processing a first gesture of a first hand of the user only with the first augmented reality device and processing a second gesture of a second hand of the user only with the second augmented reality device.
[0126] Example 19 includes the computing device of Example 11, wherein allocating one or more processes of the sign language recognition system further includes: detecting an occlusion of a gesture in the first image by the first augmented reality device; and in response to detecting the occlusion, assigning the second augmented reality device to perform one or more processes based on the second image.
[0127] Example 20 is a non-transitory computer-readable storage medium that includes instructions that, when executed by a computer, cause the computer to: access a first image generated by a first imaging device of a first augmented reality device and a second image generated by a second imaging device of a second augmented reality device, the first image and the second image depicting a gesture of a user of the first augmented reality device; synchronize the first augmented reality device with the second augmented reality device; in response to the synchronization, allocate one or more processes of a sign language recognition system between the first augmented reality device and the second augmented reality device, wherein the one or more processes are executed on corresponding augmented reality devices; collect results from the one or more processes from the first augmented reality device and from the second augmented reality device; and based on the results, display text of a sign language translation indicating the gesture in a first display of the first augmented reality device or in a second display of the second augmented reality device in near real time.
[0128] As shown below by way of example, implementations of the described subject matter may include one or more features, individually or in combination.
Claims
1. A method, comprising: accessing a first image generated by a first imaging device of a first augmented reality device and a second image generated by a second imaging device of a second augmented reality device, the first image and the second image depicting a gesture of a user of the first augmented reality device; synchronizing the first augmented reality device with the second augmented reality device; in response to the synchronization, allocating one or more processes of a sign language recognition system between the first augmented reality device and the second augmented reality device, wherein the one or more processes are executed on corresponding augmented reality devices; collecting results from the one or more processes from the first augmented reality device and from the second augmented reality device; and based on the results, displaying text indicating a sign language translation of the gesture in a first display of the first augmented reality device or in a second display of the second augmented reality device in near real time.
2. The method according to claim 1, wherein, synchronizing the first augmented reality device with the second augmented reality device further comprises: mapping a first timestamp of the first image to a second timestamp of the second image.
3. The method according to claim 1, wherein, synchronizing the first augmented reality device with the second augmented reality device further comprises: registering a six-degree-of-freedom coordinate system of the first augmented reality device with a six-degree-of-freedom coordinate system of the second augmented reality device.
4. The method according to claim 1, further comprising: transmitting the first image to the second augmented reality device, the first image providing an egocentric view of the gesture of the user of the first augmented reality device, and wherein the second augmented reality device is configured to perform the one or more processes of the sign language recognition system on a combination of the first image, a first hand skeleton based on the first image, the second image, and a second hand skeleton based on the second image, wherein the one or more processes of the sign language recognition system include: a hand tracking process, a gesture detection process, and a sign language translation process.
5. The method according to claim 1, further comprising: identifying suggested words based on a combination of context information of the first augmented reality device, historical inputs of the user of the first augmented reality device, a text corpus, and predefined terms, wherein the text indicates the suggested words.
6. The method according to claim 5, further comprising: displaying the suggested words in the first display of the first augmented reality device; receiving confirmation of the suggested words from the user of the first augmented reality device; and in response to receiving the confirmation, providing the suggested words to the second augmented reality device.
7. The method according to claim 6, wherein, receiving the confirmation includes: detecting a predefined gesture from the user, the predefined gesture indicating the confirmation.
8. The method according to claim 1, wherein, allocating the one or more processes of the sign language recognition system further includes: Allocate the one or more processes between the first augmented reality device and the second augmented reality device either temporally or spatially, and wherein the temporal allocation includes processing alternating frames between the first augmented reality device and the second augmented reality device, wherein the spatial allocation includes processing a first gesture of a first hand from the user only with the first augmented reality device and processing a second gesture of a second hand from the user only with the second augmented reality device.
9. The method according to claim 1, wherein, allocating the one or more processes of the sign language recognition system further includes: detecting, by the first augmented reality device, an occlusion of the gesture in the first image; and in response to detecting the occlusion, assigning the second augmented reality device to perform the one or more processes based on the second image.
10. The method according to claim 1, wherein, the first augmented reality device includes a first head-wearable device, and wherein the second augmented reality device includes a mobile hand-held device or a second head-wearable device.
11. A computing device, comprising: a processor; and a memory storing instructions that, when executed by the processor, configure the device to: access a first image generated by a first imaging device of a first augmented reality device and a second image generated by a second imaging device of a second augmented reality device, the first image and the second image depicting a gesture of a user of the first augmented reality device; synchronize the first augmented reality device with the second augmented reality device; in response to synchronization, allocate one or more processes of a sign language recognition system between the first augmented reality device and the second augmented reality device, wherein the one or more processes are executed on corresponding augmented reality devices; collect results from the one or more processes from the first augmented reality device and from the second augmented reality device; and based on the results, display text indicating a sign language translation of the gesture in a first display of the first augmented reality device or in a second display of the second augmented reality device in near real-time.
12. The computing device according to claim 11, wherein, synchronizing the first augmented reality device with the second augmented reality device further includes: mapping a first timestamp of the first image to a second timestamp of the second image.
13. The computing device according to claim 11, wherein, synchronizing the first augmented reality device with the second augmented reality device further includes: registering a six-degree-of-freedom coordinate system of the first augmented reality device with a six-degree-of-freedom coordinate system of the second augmented reality device.
14. The computing device according to claim 11, wherein, the instructions further configure the device to: transmit the first image to the second augmented reality device, the first image providing an egocentric view of the gesture of the user of the first augmented reality device, and Wherein, the second augmented reality device is configured to perform the one or more processes of the sign language recognition system on a combination of the first image, the first hand skeleton based on the first image, the second image, and the second hand skeleton based on the second image. Wherein, the one or more processes of the sign language recognition system include: a hand tracking process, a gesture detection process, and a sign language translation process.
15. The computing device according to claim 11, Wherein, the instructions further configure the device to: Identify suggested words based on a combination of the context information of the first augmented reality device, the historical input of the user of the first augmented reality device, a text corpus, and predefined terms, wherein the text indicates the suggested words.
16. The computing device according to claim 15, Wherein, the instructions further configure the device to: Display the suggested words on the first display of the first augmented reality device; Receive confirmation of the suggested words from the user of the first augmented reality device; And In response to receiving the confirmation, provide the suggested words to the second augmented reality device.
17. The computing device according to claim 16, Wherein, Receiving the confirmation includes: Detecting a predefined gesture from the user, the predefined gesture indicating the confirmation.
18. The computing device according to claim 11, Wherein, Allocating the one or more processes of the sign language recognition system further includes: Allocating the one or more processes temporally or spatially between the first augmented reality device and the second augmented reality device, and wherein temporal allocation includes processing alternate frames between the first augmented reality device and the second augmented reality device, wherein spatial allocation includes processing only the first gesture of the first hand of the user with the first augmented reality device and processing only the second gesture of the second hand of the user with the second augmented reality device.
19. The computing device according to claim 11, Wherein, Allocating the one or more processes of the sign language recognition system further includes: Detecting an occlusion of the gesture in the first image by the first augmented reality device; and In response to detecting the occlusion, assigning the second augmented reality device to perform the one or more processes based on the second image.
20. A non-transitory computer-readable storage medium, the computer-readable storage medium includes instructions that, when executed by a computer, cause the computer to: Access a first image generated by a first imaging device of a first augmented reality device and a second image generated by a second imaging device of a second augmented reality device, the first image and the second image depicting gestures of a user of the first augmented reality device; Synchronize the first augmented reality device with the second augmented reality device; In response to synchronization, allocate one or more processes of a sign language recognition system between the first augmented reality device and the second augmented reality device, wherein the one or more processes are executed on corresponding augmented reality devices. Collect results from the one or more processes from the first augmented reality device and from the second augmented reality device; and Based on the results, display text indicating a sign language translation of the gesture in near real-time in a first display of the first augmented reality device or in a second display of the second augmented reality device.