Physical Keyboard Tracking
By employing gradient- and variance-based techniques to detect T, X, and L shapes between keyboard keys, the method addresses the challenge of accurate keyboard tracking in augmented and virtual reality, achieving sub-millimeter precision.
Patent Information
- Application Number
- JP2023533741
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2020-12-01
- Filing Date
- 2021-11-29
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-11-29
AI Technical Summary
High-precision keyboard tracking in augmented and virtual reality environments is challenging due to low-resolution, noisy, and distorted images captured by head-mounted displays, which are often occluded by the user's hand and have suboptimal lighting, making accurate feature detection difficult.
The method involves detecting prominent features like T, X, and L shapes formed by the spaces between keyboard keys using gradient- and variance-based techniques, correcting distorted images, and comparing them to pre-mapped keyboard models to determine the pose of the physical keyboard accurately.
Enables sub-millimeter accurate tracking of physical keyboards by rendering a virtual model that matches the physical keyboard's pose, allowing precise interaction in augmented and virtual reality environments.
Smart Images

Figure 0007708858000009 
Figure 0007708858000010 
Figure 0007708858000011
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to rendering augmented reality (AR) or virtual reality (VR) content on a user device. The present disclosure generally relates to an input controller for use in an AR / VR environment.
Background Art
[0002] Virtual reality is a computer-generated simulation of an environment (e.g., a 3D environment) in which a user can interact in an ostensibly real or physical way. A virtual reality system, which can be a single device or a group of devices, can generate this simulation for display to a user, e.g., on a virtual reality headset or some other display device. The simulation can include images, sounds, tactile feedback, and / or other sensations to mimic a real or imaginary environment. As virtual reality becomes increasingly prominent, the range of useful applications of virtual reality is rapidly expanding. The most common applications of virtual reality involve games or other interactive content, but just behind that are other applications such as viewing visual media items (e.g., photos, videos) for entertainment or training purposes. The possibility of using virtual reality to simulate real conversations and other user interactions is also being explored.
[0003] Augmented reality provides a view of the real world or physical world with computer-generated sensory input (e.g., visual, auditory) added. In other words, computer-generated virtual effects can enhance or supplement the view of the real world. For example, a camera on a virtual reality headset or head-mounted display (HMD) can capture a real-world scene (as an image or video) and display a composite with computer-generated virtual objects of the captured scene. The virtual objects can be, for example, two-dimensional and / or three-dimensional objects and can be static or animated.
SUMMARY OF THE INVENTION
[0004] According to a first aspect of the present disclosure, a computing device captures an image from a camera perspective, the image depicting a physical keyboard, compares pixels of the image with a predetermined shape template to detect one or more shape features of the physical keyboard depicted in the image, the predetermined shape template representing visual characteristics of spaces between keyboard keys, accesses predetermined shape features of a keyboard model associated with the physical keyboard, and determines a pose of the physical keyboard based on a comparison between (1) the detected one or more shape features of the physical keyboard and (2) a projection of the predetermined shape features of the keyboard model towards the camera perspective.
[0005] In some embodiments, the predetermined shape template is a T-shaped template having a body portion and a roof portion, the body portion representing visual characteristics of a space between two keyboard keys within a first keyboard row.
[0006] In some embodiments, the roof portion of the T-shaped template represents visual characteristics of a space between (1) a portion of the first keyboard row corresponding to two keyboard keys within the first keyboard row and (2) one or more keyboard keys within a second keyboard adjacent to the first keyboard row.
[0007]
[0008] In some embodiments, the predetermined shape template is a T-shaped template having a body portion and a roof portion, the body portion representing the visual characteristics of the space between a first keyboard key in a first keyboard row and a second keyboard key in a second keyboard row adjacent to the first keyboard row, the first keyboard key being the rightmost key in the first keyboard row and the second keyboard key being the rightmost key in the second keyboard row, and the roof portion of the T-shaped template representing the visual characteristics of the space to the right of the first keyboard key and the second keyboard key, or the first keyboard key being the leftmost key in the first keyboard row and the second keyboard key being the leftmost key in the second keyboard row, and the roof portion of the T-shaped template representing the visual characteristics of the space to the left of the first keyboard key and the second keyboard key.
[0009] In some embodiments, comparing the pixels of an image to a predetermined shape template comprises dividing the image into a plurality of pixel blocks, each pixel block including a plurality of pixels, determining the visual characteristics of each pixel block based on the pixel intensities associated with the plurality of pixels within the pixel block, comparing the visual characteristics of the pixel block to a predetermined shape template representing the visual characteristics of the space between keyboard keys, and detecting one or more shape features of a physical keyboard depicted in the image based on a comparison between the visual characteristics of the pixel blocks of the image and the predetermined shape template representing the visual characteristics of the space between keyboard keys.
[0010] In some embodiments, determining the visual characteristics of a pixel block based on the pixel intensities associated with the plurality of pixels within the pixel block comprises comparing the pixel intensities associated with at least one portion of the pixel block to another portion of the pixel block to determine the gradients and variances of the pixel intensities associated with the various portions of the pixel block.
[0011] In some embodiments, the method further includes rendering a representation of a keyboard model to match the pose of a physical keyboard.
[0012] In some embodiments, the method further includes identifying a keyboard model from a database including a plurality of keyboard models by comparing one or more detected shape features of the physical keyboard with predetermined shape features of a plurality of keyboard models in the database prior to accessing the predetermined shape features of the keyboard model associated with the physical keyboard.
[0013] In some embodiments, the predetermined shape template is an L-shaped template or an X-shaped template.
[0014] According to a second aspect of the present disclosure, there is provided one or more non-transitory computer-readable storage media embodying software that, when executed, is operable to perform capturing an image from a camera perspective, wherein the image depicts a physical keyboard, detecting one or more shape features of the physical keyboard depicted in the image by comparing pixels of the image with a predetermined shape template, wherein the predetermined shape template represents visual characteristics of spaces between keyboard keys, accessing predetermined shape features of a keyboard model associated with the physical keyboard, and determining a pose of the physical keyboard based on a comparison between (1) the one or more detected shape features of the physical keyboard and (2) a projection of the predetermined shape features of the keyboard model towards the camera perspective.
[0015] In some embodiments, the predetermined shape template is a T-shaped template having a body portion and a roof portion, and the body portion represents visual characteristics of a space between two keyboard keys within a first keyboard column.
[0016] In some embodiments, the roof portion of the T-shaped template represents the visual characteristics of the space between (1) a portion of the first keyboard row corresponding to two keyboard keys within the first keyboard row and (2) one or more keyboard keys within a second keyboard adjacent to the first keyboard row.
[0017] In some embodiments, comparing the pixels of an image to a predetermined shape template comprises dividing the image into a plurality of pixel blocks, each pixel block including a plurality of pixels, determining the visual characteristics of each pixel block based on the pixel intensities associated with the plurality of pixels within the pixel block, comparing the visual characteristics of the pixel block to a predetermined shape template representing the visual characteristics of the space between keyboard keys, and detecting one or more shape features of a physical keyboard depicted within the image based on the comparison between the visual characteristics of the pixel blocks of the image and the predetermined shape template representing the visual characteristics of the space between keyboard keys.
[0018] In some embodiments, determining the visual characteristics of a pixel block based on the pixel intensities associated with the plurality of pixels within the pixel block includes comparing the pixel intensities associated with at least one portion of the pixel block to another portion of the pixel block to determine the gradients and variances of the pixel intensities associated with the various portions of the pixel block.
[0019] According to a third aspect of the present disclosure, a system is provided that includes one or more processors and one or more computer-readable non-transitory storage media that communicate with the one or more processors. The one or more computer-readable non-transitory storage media contain instructions that, when executed by the one or more processors, cause the system to: capture an image from a camera perspective, where the image depicts a physical keyboard; detect one or more shape features of the physical keyboard depicted in the image by comparing pixels of the image with a predetermined shape template, where the predetermined shape template represents visual characteristics of spaces between keyboard keys; access predetermined shape features of a keyboard model associated with the physical keyboard; and determine a pose of the physical keyboard based on a comparison between (1) the detected one or more shape features of the physical keyboard and (2) a projection of the predetermined shape features of the keyboard model towards the camera perspective.
[0020] In some embodiments, the predetermined shape template is a T-shaped template having a body portion and a roof portion, and the body portion represents visual characteristics of a space between two keyboard keys within a first keyboard row.
[0021] In some embodiments, the roof portion of the T-shaped template represents visual characteristics of a space between (1) a portion of the first keyboard row corresponding to two keyboard keys within the first keyboard row and (2) one or more keyboard keys within a second keyboard adjacent to the first keyboard row.
[0022] In some embodiments, comparing the pixels of an image to a predetermined shape template comprises dividing the image into a plurality of pixel blocks, each pixel block including a plurality of pixels, determining the visual characteristics of each pixel block based on the pixel intensities associated with the plurality of pixels within the pixel block, comparing the visual characteristics of the pixel block to a predetermined shape template representing the visual characteristics of the space between keyboard keys, and detecting one or more shape features of a physical keyboard depicted in the image based on the comparison between the visual characteristics of the pixel blocks of the image and the predetermined shape template representing the visual characteristics of the space between keyboard keys.
[0023] In some embodiments, determining the visual characteristics of a pixel block based on the pixel intensities associated with the plurality of pixels within the pixel block comprises comparing the pixel intensities associated with at least one portion of the pixel block to another portion of the pixel block to determine the gradients and variances of the pixel intensities associated with the various portions of the pixel block.
Brief Description of the Drawings
[0024]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5A
Figure 5B
Figure 6A
Figure 6B
Figure 7
Figure 8A
Figure 8B
Figure 9
Figure 10
Figure 11
DETAILED DESCRIPTION OF THE INVENTION
[0025] The embodiments disclosed herein are described, but not limited to, determining the exact position and orientation (pose) of a physical keyboard and rendering an image of a virtual model corresponding to the physical keyboard to exactly match the pose of the physical keyboard. High-precision keyboard tracking is difficult. Machine learning techniques are too slow to provide real-time applications of keyboard tracking and are too inaccurate due to the low quality of the cameras used for tracking. Therefore, computer vision techniques are required. However, various factors also make computer vision techniques difficult. For example, accurate feature detection based on images captured by an outward-facing camera of a head-mounted display (HMD) is difficult because the captured images may have low resolution and high noise, and the features of the keyboard have low contrast. Furthermore, the features of the keyboard are often blocked by the user's hand, and the lighting conditions may not be optimal for accurate feature detection. Some devices use a fish-eye lens for the outward-facing camera to maximize tracking coverage, so the captured images may also be distorted or warped.
[0026] To address these issues, the invention disclosed herein looks for certain T, X, and L features formed by the spaces between keys on a physical keyboard, which are prominent features that can be detected even in low-resolution, distorted, or partially occluded keyboard images. The method can start by capturing an image of the keyboard. If the image is distorted (e.g., due to being captured by a fisheye lens), the image may be corrected by generating a straight-line image of the keyboard. Next, gradient- and variance-based techniques are used to detect the T, X, and L features of the keyboard. The gradient- and variance-based techniques utilize the distribution and differences in pixel intensities within the boundary regions of the keys (the uniform regions surrounding the symbols on each key) and within the spaces between adjacent keys. For any given keyboard, based on the differences in pixel intensities, unique patterns in the shape of a T, an L, and / or an X (or cross-shaped, “+” shaped) can be detected. These unique patterns can be compared to pre-mapped models of various keyboards in a database. When a virtual model that matches the physical keyboard is found, the virtual model can be rendered to exactly match the pose of the physical keyboard by mapping it to the detected features of the physical keyboard. This technique enables the physical keyboard to be tracked with sub-millimeter accuracy. This technique may also be generalized for tracking any real-world object that often has T, X, and L features.
[0027] Embodiments of the present invention may include or be implemented in relation to an artificial reality system. Artificial reality is reality in a form that has been adjusted in some manner prior to presentation to a user, which may include, for example, virtual reality (VR), augmented reality (AR), mixed reality (MR), hybrid reality, or some combination and / or derivative thereof. Artificial reality content may include fully generated content or generated content combined with captured content (e.g., a photograph of the real world). Artificial reality content may include video, audio, tactile feedback, or some combination thereof, any of which may be presented in a single channel or multiple channels (such as stereoscopic video that provides a three-dimensional effect to an observer). Further, in some embodiments, artificial reality may be associated with, for example, applications, products, accessories, services, or some combination thereof that are used to create content in artificial reality and / or that are used in artificial reality (such as to perform an activity in artificial reality). An artificial reality system that provides artificial reality content may be implemented on various platforms, including a head-mounted display (HMD) connected to a host computer system, a stand-alone HMD, a mobile device or computing system, or any other hardware platform capable of providing artificial reality content to one or more observers.
[0028] The embodiments disclosed in this specification are merely examples and the scope of the present disclosure is not limited thereto. A particular embodiment may include all, some, or none of the components, elements, features, functions, operations, or steps of the embodiments disclosed above. Embodiments according to the present invention are disclosed particularly in the appended claims directed to methods, storage media, systems, and computer program products, and any feature described in one claim category, for example, in a method, may also be claimed in another claim category, for example, in a system. The dependencies or references in the appended claims are only selected for formal reasons. However, any subject matter as a result of an intentional forward reference (in particular, multiple dependencies) to any preceding claim is equally claimable, so that any combination of claims and their features is disclosed and claimable regardless of the dependencies selected in the appended claims. The claimable subject matter includes not only combinations of features described in the appended claims but also any other combination of features in the claims, and each feature described in the claims may be combined with any other feature or combination of features in the claims. Further, any of the embodiments and features described or shown in this specification may be claimed in a separate claim and / or in any combination with any of the embodiments or features described or shown in this specification or with any of the features of the appended claims.
[0029] Embodiments are disclosed herein related to systems and methods for transplanting physical objects in a physical space into a virtual space of virtual reality. In one aspect, transplanting a physical object in a physical space into a virtual space includes activating the physical object in the physical space to obtain a virtual model of the physical object and rendering an image of the virtual model in the virtual space.
[0030] In some embodiments, the physical object is activated during a pass-through mode in which the HMD presents or renders a view of the physical space to the user of the head-mounted display. For example, a virtual model of the physical object may be identified or selected during the pass-through mode and rendered within the virtual space. The physical object within the physical space may be tracked, and the position and orientation of the virtual model within the virtual space may be adjusted according to the position and orientation of the physical object within the physical space. In one aspect, an indication of the user's interaction with the physical object within the physical space may be presented on the virtual model within the virtual space as feedback to the user.
[0031] In one aspect, the physical object can be a general-purpose input device (such as a keyboard or a mouse) that can be manufactured or produced by a company different from the company that manufactures or produces the head-mounted display and / or a dedicated hand-held input device (such as a pointing device). By rendering a virtual model of the input device within the user's field of view as a reference or guidance (such as acting as a proxy for the input device), the user can easily reach for the virtual model and thus the input device and provide input to the virtual reality through the input device.
[0032] In one aspect, with respect to a virtual model in a virtual space (e.g., using a virtual model in a virtual space for spatial guidance), spatial feedback regarding a user's interaction with an input device in the physical space can be provided visually to the user. In one approach, the input device in the physical space is detected with respect to the user of the input device. A virtual model of the detected input device in the virtual space in terms of position and orientation (pose) may be presented to the user by a display device. The position and orientation of the virtual model in the virtual space can correspond (e.g., track or mirror) to the position and orientation of the input device in the physical space with respect to the user. With respect to the virtual model in the virtual space (and e.g., a virtual representation of the user's hand), spatial feedback regarding the user's interaction with the input device in the physical space can be provided visually to the user via the virtual space. Thus, via the spatial feedback to the virtual model, the user can easily find the input device, reach out, and provide input via the input device in the physical space while enjoying a virtual reality experience (e.g., while looking at the virtual space instead of the physical space).
[0033] The systems and methods disclosed herein may relate to transplanting physical objects into virtual reality, but the general principles disclosed herein may be applicable to augmented reality or mixed reality.
[0034] FIG. 1 is a block diagram of an exemplary artificial reality system environment 100 in which a console 110 operates. In some embodiments, the artificial reality system environment 100 includes a head-mounted display (HMD) 150 worn by a user and a console 110 that provides artificial reality content to the HMD 150. In one aspect, the HMD 150 can detect its position, orientation, and / or the direction of the user's line of sight wearing the HMD 150, and can provide the detected position and line of sight direction to the console 110. The console 110 can identify a view within the artificial reality space corresponding to the detected position, orientation, and / or line of sight direction, and can generate an image depicting the identified view. The console 110 can provide the image to the HMD 150 for rendering. In some embodiments, the artificial reality system environment 100 includes an input device 120 communicatively coupled to the console 110 or the HMD 150 via a wired cable, a wireless link (e.g., Bluetooth, Wi-Fi, etc.), or both. The input device 120 can be a dedicated hardware with motion sensors (e.g., a pointing device or a controller), a general-purpose keyboard, a mouse, etc. Via the input device 120, the user can provide inputs associated with the presented artificial reality. In some embodiments, the artificial reality system environment 100 includes more, fewer, or different components than those shown in FIG. 1. In some embodiments, the functionality of one or more components of the artificial reality system environment 100 can be distributed among the components in a manner different from that described herein. For example, some of the functions of the console 110 may be performed by the HMD 150. For example, some of the functions of the HMD 150 may be performed by the console 110. In some embodiments, the console 110 is integrated as part of the HMD 150.
[0035] In some embodiments, the HMD 150 can include or correspond to electronic components that can be worn by a user and present or provide an artificial reality experience to the user. The HMD 150 can render one or more images, videos, audio, or some combination thereof to provide an artificial reality experience to the user. In some embodiments, the audio is presented via an external device (e.g., speakers and / or headphones) that receives audio information from the HMD 150, the console 110, or both and presents audio data based on the audio information. In some embodiments, the HMD 150 includes a sensor 155, a communication interface 165, an image renderer 170, an electronic display 175, and / or an object transfer controller 180. These components can work together to detect the position and orientation of the HMD 150 and / or the direction of the user's line of sight while wearing the HMD 150, and can render an image of a view within the artificial reality corresponding to the detected position and orientation of the HMD 150 and / or the direction of the user's line of sight. In other embodiments, the HMD 150 includes more, fewer, or different components than those shown in FIG. 1. In some embodiments, the object transfer controller 180 may be activated or deactivated according to the control from the user of the HMD 150.
[0036] In some embodiments, sensor 155 includes an electronic component, or a combination of an electronic component and a software component, that detects the position, orientation, and / or the direction of the user's line of sight of HMD 150. Examples of sensor 155 can include one or more imaging sensors, one or more accelerometers, one or more gyroscopes, one or more magnetometers, a global positioning system, or another suitable type of sensor that detects motion and / or position. For example, one or more accelerometers can measure translational motion (e.g., forward / backward, up / down, left / right), and one or more gyroscopes can measure rotational motion (e.g., pitch, yaw, roll). In some embodiments, the imaging sensor can capture an image for detecting physical objects, user gestures, hand shapes, user interactions, etc. In some embodiments, sensor 155 detects translational and rotational motion and determines the orientation and position of HMD 150. In one aspect, sensor 155 detects translational and rotational motion relative to the previous orientation and position of HMD 150, and can identify the new orientation and / or position of HMD 150 by accumulating or integrating the detected translational and / or rotational motion. As an example, assuming HMD 150 is oriented in a direction 25 degrees from a reference direction, in response to detecting that HMD 150 has rotated 20 degrees, sensor 155 can identify that HMD 150 is now facing or oriented in a direction 45 degrees from the reference direction. As another example, assuming HMD 150 is located 2 feet away from a reference point in a first direction, in response to detecting that HMD 150 has moved 3 feet in a second direction, sensor 155 can identify that HMD 150 is now located at the vector sum of 2 feet in the first direction and 3 feet in the second direction from the reference point. In one aspect, the direction of the user's line of sight can be determined or estimated according to the position and orientation of HMD 150.
[0037] In some embodiments, sensor 155 can include an electronic component, or a combination of an electronic component and a software component, that generates sensor measurements of the physical space. Examples of sensor 155 for generating sensor measurements can include one or more imaging sensors, thermal sensors, and the like. In one example, the imaging sensor can capture an image corresponding to the field of view of a user within the physical space (or the view from the position of HMD 150 following the orientation of HMD 150). Image processing can be performed on the captured image to detect a physical object or a portion of the user within the physical space.
[0038] In some embodiments, communication interface 165 includes an electronic component, or a combination of an electronic component and a software component, that communicates with console 110. Communication interface 165 can communicate with communication interface 115 of console 110 via a communication link. The communication link can be a wireless link, a wired link, or both. Examples of wireless links can include cellular communication links, short-range wireless communication links, Wi-Fi, Bluetooth, or any communication wireless communication link. Examples of wired links can include Universal Serial Bus (USB), Ethernet, Firewire, High-Definition Multimedia Interface (HDMI), or any wired communication link. In embodiments where console 110 and HMD 150 are implemented on a single system, communication interface 165 can communicate with console 110 via at least a bus connection or conductive traces. Via the communication link, communication interface 165 can send data indicating the position of the specified HMD 150 and the orientation of HMD 150, and / or the direction of the user's line of sight to console 110. Moreover, via the communication link, communication interface 165 can receive data indicating the image to be rendered from console 110.
[0039] In some embodiments, the image renderer 170 includes, for example, an electronic component or a combination of an electronic component and a software component that generates one or more images for display according to changes in the view of the virtual reality space. In some embodiments, the image renderer 170 is implemented as a processor (or a graphics processing unit (GPU)). The image renderer 170 can receive data depicting the image to be rendered through the communication interface 165 and render the image through the electronic display 175. In some embodiments, the data from the console 110 may be compressed or encoded, and the image renderer 170 can decompress or decode the data to generate and render an image. The image renderer 170 can receive a compressed image from the console 110 and decompress the compressed image, thereby reducing the communication bandwidth between the console 110 and the HMD 150. In one aspect, the process of detecting the position, orientation of the HMD 150, and / or the line-of-sight direction of the user wearing the HMD 150 by the HMD 150 and generating a high-resolution image (e.g., 1920×1080 pixels) corresponding to the detected position, orientation, and / or line-of-sight direction by the console 110 and transmitting it to the HMD 150 may be computationally intensive and may not be executed within the frame time (e.g., less than 11 ms). The image renderer 170 can generate one or more images through a shading process and a reprojection process when an image from the console 110 is not received within the frame time. For example, the shading process and the reprojection process can be adaptively executed according to changes in the view of the virtual reality space.
[0040] In some embodiments, the electronic display 175 is an electronic component that displays an image. The electronic display 175 can be, for example, a liquid crystal display or an organic light emitting diode display. The electronic display 175 can be a transparent display that enables a user to see through it. In some embodiments, when the HMD 150 is worn by the user, the electronic display 175 is disposed proximate to the user's eyes (e.g., less than 3 inches). In one aspect, the electronic display 175 radiates or projects light towards the user's eyes in accordance with an image generated by the image renderer 170.
[0041] In some embodiments, the object transplantation controller 180 includes an electronic component, or a combination of an electronic component and a software component, that activates a physical object and generates a virtual model of the physical object. In one approach, the object transplantation controller 180 detects a physical object in the physical space during the pass-through mode, the sensor 155 can capture an image of the user's view (or field of view) of the physical space therein, and the electronic display 175 can present the captured image to the user. The object transplantation controller 180 can generate a virtual model of the physical object and present the virtual model in the virtual space, and the electronic display 175 can display the user's field of view of the virtual space therein. A detailed description of the activation of the physical object and the rendering of the virtual model of the physical object is provided below.
[0042] In some embodiments, the console 110 is an electronic component, or a combination of an electronic component and a software component, that provides content to be rendered via the HMD 150. In one aspect, the console 110 includes a communication interface 115 and a content provider 130. These components can operate together to identify an augmented reality view corresponding to the position of the HMD 150, the orientation of the HMD 150, and / or the direction of the user's line of sight of the HMD 150, and can generate an augmented reality image corresponding to the identified view. In other embodiments, the console 110 includes more, fewer, or different components than those shown in FIG. 1. In some embodiments, the console 110 performs some or all of the functions of the HMD 150. In some embodiments, the console 110 is integrated as part of the HMD 150 as a single device.
[0043] In some embodiments, the communication interface 115 is an electronic component, or a combination of an electronic component and a software component, that communicates with the HMD 150. The communication interface 115 can be the counterpart component to a communication interface 165 that communicates via a communication link (e.g., a USB cable). Via the communication link, the communication interface 115 can receive data from the HMD 150 indicating the position of the identified HMD 150, the orientation of the HMD 150, and / or the direction of the line of sight of the identified user. Moreover, via the communication link, the communication interface 115 can transmit data depicting the image to be rendered to the HMD 150.
[0044] Content provider 130 is a component that generates content to be rendered according to the position of HMD 150, the orientation of HMD 150, and / or the line-of-sight direction of the user of HMD 150. In one aspect, content provider 130 determines a view of the virtual reality according to the position of HMD 150, the orientation of HMD 150, and / or the line-of-sight direction of the user of HMD 150. For example, content provider 130 maps the position of HMD 150 in the physical space to a position in the virtual space, and determines a view of the virtual space along the line-of-sight direction from the mapped position in the virtual space. Content provider 130 can generate image data depicting an image of the determined view of the virtual space, and transmit the image data to HMD 150 via communication interface 115. In some embodiments, content provider 130 generates metadata including motion vector information, depth information, edge information, object information, etc. associated with the image, and transmits the metadata together with the image data to HMD 150 via communication interface 115. Content provider 130 can compress and / or encode the data for depicting the image, and transmit the compressed and / or encoded data to HMD 150. In some embodiments, content provider 130 generates an image periodically (e.g., every 11 ms) and provides it to HMD 150.
[0045] FIG. 2 is a diagram of the HMD 150 according to an exemplary embodiment. In some embodiments, the HMD 150 includes a front rigid body 205 and a band 210. The front rigid body 205 includes an electronic display 175, sensors 155A, 155B, 155C, and an image renderer 170 (not shown in FIG. 2). The sensor 155A can be an accelerometer, gyroscope, magnetometer, or another suitable type of sensor that detects movement and / or position. The sensors 155B, 155C can be imaging sensors that capture images for detecting physical objects, user gestures, hand shapes, user interactions, etc. In some embodiments, the sensors 155B and 155C can be imaging sensors with fisheye lenses. The HMD 150 may include additional components (e.g., GPS, wireless sensors, microphones, thermal sensors, etc.). In other embodiments, the HMD 150 has a configuration different from that shown in FIG. 2. For example, the image renderer 170, and / or the sensors 155A, 155B, 155C may be arranged at positions different from those shown in FIG. 2.
[0046] Figure 3 is a diagram of the object transplantation controller 180 of FIG. 1 according to an exemplary implementation of the present disclosure. In some embodiments, the object transplantation controller 180 includes an object detector 310, a VR model generator 320, a VR model renderer 330, and a feedback controller 340. These components can operate together to detect physical objects and present virtual models of the physical objects. The virtual model may be identified, activated, or generated, and presented, such that a user of the HMD 150 can identify the position of the physical object while wearing the HMD 150. In some embodiments, these components may be implemented as hardware, software, or a combination of hardware and software. In some embodiments, these components may be implemented as application specific integrated circuits (ASICs) or field programmable gate arrays (FPGAs). In some embodiments, these components may be implemented as a processor and a non-transitory computer-readable medium storing instructions that cause the processor to execute the various processes disclosed herein when executed by the processor. In some embodiments, the object transplantation controller 180 includes more, fewer, or different components than those shown in FIG. 3. In some embodiments, the functions of some of the components may be performed by, or in cooperation with, the content provider 130 or a remote server. For example, some of the functions of the object detector 310, the VR model generator 320, or both may be performed by the content provider 130 or a remote server. In some embodiments, the object transplantation controller 180 includes more, fewer, or different components than those shown in FIG. 3.
[0047] In some embodiments, the object detector 310 is or includes components that detect physical objects in the physical space according to the captured images. In one application, the object detector 310 detects an input device (e.g., a keyboard or a mouse) in the physical space by performing image processing on the captured images. In one approach, the object detector 310 detects the contours, outlines, and / or layouts of the keys or buttons of the physical objects in the captured images (e.g., formed by the keys themselves or the spaces between the keys), or combinations thereof, and identifies the type of the physical object according to the detected contours, outlines, and / or layouts of the keys or buttons. For example, the object detector 310 determines whether the physical object in the user's perspective of the physical space is a keyboard or a mouse according to the detected contours, outlines, and / or layouts of the keys or buttons of the physical object. The object detector 310 can also identify the position of the physical object according to a thermal sensor that detects a heat map of the physical object. In one example, the object detector 310 can detect that the physical object has a specific number of keys according to the outline of the physical object in the captured image and determine that the physical object is a keyboard.
[0048] In one aspect, the object detector 310 detects a physical object and presents a view of the physical space or a portion of the view to the user of the HMD 150 via the electronic display 175. For example, an image (e.g., of a physical object and / or a part of the user) captured by an imaging sensor of the HMD 150 (e.g., sensors 155B, 155C) can be presented to the user via the electronic display 175 (e.g., regardless of whether there is a fusion with a virtual model and / or an image of a virtual space). Accordingly, a user wearing the HMD 150 can detect a physical object in the physical space and / or identify its position via the HMD 150, for example, using image processing on the image acquired by the imaging sensor.
[0049] In some embodiments, the VR model generator 320 is or includes a component that generates, obtains, or identifies a virtual model of the detected physical object. In one approach, the VR model generator 320 stores a plurality of candidate models for different manufacturing companies, brands, and / or product models. The VR model generator 320 can compare the detected contour, shape, and / or layout of a key or button (e.g., formed by the key itself or the space between keys) to the contour, shape, and / or layout of the keys or buttons of the plurality of candidate models to identify or specify a candidate model that matches or has the closest contour, shape, and / or layout of the detected key or button of the physical object. The VR model generator 320 can detect or receive product identification information of the physical object and identify or specify a candidate model corresponding to the detected product identification information. The virtual model generator 320 can generate, determine, obtain, or select the specified candidate model as the virtual model of the physical object.
[0050] In some embodiments, the VR model renderer 330 is or includes a component that renders an image of a virtual model of a physical object. In one approach, the VR model renderer 330 tracks physical objects in the captured image and determines the position and orientation of the physical object relative to the user or the HMD 150. In one aspect, since the user can move around while wearing the HMD 150, the position and orientation of the physical object relative to the user or the HMD 150 may change. The VR model renderer 330 can determine the six degrees of freedom of the virtual model (e.g., forward / backward (surge), up / down (heave), left / right (sway) translation, left / right tilt (roll), forward / backward tilt (pitch), left / right rotation (yaw)) such that the position and orientation of the virtual model in the virtual space relative to the user or the HMD 150 that is seen or displayed correspond to the position and orientation of the physical object relative to the user or the HMD 150 in the captured image. In certain embodiments, the VR model renderer 330 performs image processing on the captured image to track specific features of the physical object (e.g., the four corners and / or sides of a keyboard, the shape formed by the spaces between keys, and / or a pattern formed by a group of such shapes), and determines the position and orientation (pose) of the virtual model to match, correspond to, track, or conform to the features of the physical object in the captured image. The VR model renderer 330 can present the virtual model according to the position and orientation of the virtual model via the electronic display 175. In one aspect, the VR model renderer 330 tracks the physical object while the user is wearing the HMD 150, updates the position and orientation of the virtual model, and the electronic display 175 presents the user's viewpoint / field of view of the virtual space therein. With the virtual model presented in the virtual space and acting as a spatial guide or reference, the user can easily find and reach for the physical object even though the user cannot actually see the physical object because the user is wearing the HMD.
[0051] In some embodiments, the feedback controller 340 is or includes a component that generates spatial feedback for a user interaction with a physical object. In one aspect, the feedback controller 340 detects and tracks the user's hand in the captured image within the HMD 150 and visually provides spatial feedback regarding the user's movement and / or interaction with the physical object via the electronic display 175. The spatial feedback may be provided with respect to a virtual model. In one example, the feedback controller 340 determines whether the user's hand is within a predetermined distance from (or in proximity to) the physical object. When the user's hand is within a predetermined distance from a physical object (e.g., a keyboard), the feedback controller 340 can generate or render a virtual model of the user's hand and present the virtual model of the user's hand within the virtual space via the electronic display 175. When the user's hand is not within a predetermined distance from a physical object (e.g., a keyboard), the feedback controller 340 cannot present or render a virtual model of the user's hand via the electronic display 175. In some embodiments, the feedback controller 340 can determine or generate a region (e.g., a rectangular region or other region) surrounding the virtual model within the virtual space and present the region via the electronic display 175. When the user's hand is within the region, a passing image of a portion of the virtual model of the hand or a portion of the hand within the region can be presented as spatial feedback (e.g., regardless of whether it is fused with other images).
[0052] In one example, the feedback controller 340 determines that a portion of the physical object is interacting with the user and indicates that the corresponding portion of the virtual model is interacting with the user via the electronic display 175. For example, the feedback controller 340 determines that a key or button on the keyboard has been pressed by performing image processing on the captured image or by receiving an electrical signal corresponding to user input via the keyboard. The feedback controller 340 can highlight the corresponding key or corresponding button of the virtual model to indicate which key on the keyboard has been pressed. Accordingly, although the user cannot actually see the physical object because the HMD is worn, the user can confirm whether the input provided via the physical object is correct.
[0053] In practice, it can be difficult to provide accurate and real-time spatial feedback of user interactions with physical objects. Typically, images of physical objects captured by an outward-facing imaging sensor of an HMD are often not directly usable for providing accurate spatial feedback to a user because the images are occluded (e.g., by the user's hand) or have characteristics of low resolution, high noise, or low contrast. For example, FIG. 4 shows images 401, 402, 403 captured by an imaging sensor of an HMD, which are occluded by the user's hand or have characteristics of low resolution, high noise, or low contrast. One solution for using such images to provide spatial feedback to a user is to determine the exact pose of the physical object depicted in the image and render a virtual model in a virtual space that accurately tracks the pose of the physical object. However, it is difficult to determine the pose of any physical object accurately enough to provide spatial feedback when the image is partially occluded or has characteristics of low resolution, high noise, or low contrast. The invention of the present disclosure provides a method for accurately determining the pose of a physical object, such as a keyboard, by detecting prominent features that can be detected even in low-resolution, high-noise, or partially occluded images. Although many embodiments of the present disclosure describe determining the pose of a physical keyboard, the present disclosure is applicable to any physical object having prominent features similar to those described herein.
[0054] In one embodiment, the prominent features of a keyboard may be identified based on visual characteristics formed by the space between keys or buttons on the keyboard or by a specific shape present at the corner of the keyboard. For example, FIG. 5A shows an exemplary image of a physical keyboard having T-shaped features 501, 502, and 503, L-shaped feature 510, and X-shaped feature 520. Typically, the most common shape feature of a keyboard is T-shaped. For example, the keyboard shown in FIG. 5A has approximately 120 T-shapes and 3 X-shapes. The T-shape is composed of an upper part called the roof part and a lower part called the body part. FIG. 5B shows an exemplary T-shaped template 590 used to detect the T-shape of the keyboard and an enlarged view of a portion of the keyboard corresponding to one of the detected T-shapes 580. FIG. 5B also provides an indication of where the roof part 592 and body part 591 of the T-shape are located. The T-shape may be formed by the visual characteristics of the space between two or more keys on the keyboard. The space between two adjacent keys in the same column forms the body part of the T-shape. For example, the T-shape may be formed by two keys in one column and a third key in an adjacent column (e.g., the T-shape 580 shown in FIG. 5B). For example, the T-shape may be formed by two keys in one column and another two keys in an adjacent column (e.g., the X-shape 520 shown in FIG. 5A). As described below, the X-shape may be formed by combining two T-shapes facing in opposite directions. When each of the two keys is separate but the rightmost or leftmost key in an adjacent column, the T-shape may be formed by the two keys without other keys, and the roof part of the T-shape is formed by the space on the right or left side of the two keys (e.g., the T-shape 503 shown in FIG. 5A). When both of the two keys are in the topmost or bottommost row, the T-shape may be formed by the two keys without other keys, and the roof part of the T-shape is formed by the space above or below the two keys. Each detected T-shape may be associated with an orientation.For example, in FIG. 5A, the T-shaped 501 is oriented such that its body faces downward with respect to its roof, the T-shaped 502 is oriented such that its body faces upward with respect to its roof, and the T-shaped 503 is oriented such that its body faces leftward with respect to its roof. Each shape feature may be associated with a positive or negative sign depending on the appearance of the shape feature. A positive sign may be assigned to a shape feature if the space between the keys forming the shape feature is brighter compared to the surrounding keys, and a negative sign may be assigned to a shape feature if the space between the keys forming the shape feature is darker compared to the surrounding keys. Alternatively, a sign opposite to the sign described above may be assigned to the shape feature. For example, in FIG. 5A, since the keys surrounding the T-shaped 503 appear darker than the space between the keys, the T-shaped 503 may be assigned a positive sign. In one embodiment, two T-shaped shapes having opposite orientations may be joined to form an X-shaped shape such as the X-shaped 520 shown in FIG. 5A.
[0055] In one embodiment, the T-shaped features of the keyboard are detected by evaluating the pixel intensity of the image depicting the keyboard. In one approach, a gradient and variance-based approach is used to identify T-shaped shapes within the image depicting the keyboard. FIGS. 6A-6B illustrate one approach for detecting T-shaped features based on the gradient and variance approach. Typically, in an image depicting a keyboard, one of the most prominent features of the keyboard corresponds to the difference in pixel intensity between the keyboard keys and the space between the keys. For example, in FIG. 4, image 401 shows keyboard keys contrasting with the dark spaces between the keys, and images 402 and 403 show keyboard keys contrasting with the bright spaces between the keys. The gradient and variance-based approach utilizes such differences in pixel intensity to detect the shape features of the keyboard. This approach can be divided into two steps. The first step includes evaluating the gradient of the pixel intensity for the roof portion of the T-shaped shape, and the second step includes evaluating the gradient of the pixel intensity for the body portion of the T-shaped shape.
[0056] FIG. 6A shows, according to one embodiment, the step of evaluating the gradient of the pixel intensity of the roof portion. This step includes identifying four groups of pixels within the pixel block, namely an upper portion 610 that may correspond to the lower ends of one or more specific keys, two lower portions 630 that may each correspond to the upper part of a specific key, and a central portion 620 between the upper portion 610 and the lower portions 630. In one embodiment, identifying the four groups of pixels includes identifying groups of pixels having pixel intensities that are substantially uniform or have a dispersion less than a specific amount. For example, the dispersion of the pixel intensities of the pixels in each of the upper portion 610, the central portion 620, and the lower portion 630 may be a dispersion less than a predetermined minimum amount. In addition to calculating the dispersion across the four regions, the gradient of the pixel intensity between the groups of pixels is determined. The upper gradient is determined based on the difference between the average pixel intensity value of the upper portion 610 and the average pixel intensity value of the central portion 620. ∇ top =μ center -μ top
[0057] The lower gradient is determined based on the difference between the average pixel intensity value of the central portion 610 and the average pixel intensity value of the lower portion 630. ∇ bottom =μ bottom -μ center
[0058] The combined gradient is ∇ horz =∇ top -∇ bottom is
[0059] The deviation is TIFF0007708858000001.tif16170.
[0060] The horizontal response referred to in FIG. 6A combines the dispersion of different regions corresponding to the T-shaped roof portion and the gradient between these regions so as not to respond to changing luminance, contrast, and image noise. TIFF0007708858000002.tif16170
[0061] If both the upper slope and the lower slope are greater than a specific minimum slope value, the horizontal response is calculated based on that. Alternatively, if either the upper slope or the lower slope is determined to be less than the minimum slope value, a zero value is assigned to the horizontal response. TIFF0007708858000003.tif22170
[0062] If the sign of the combined slope is incorrect, a zero value is assigned to the vertical response. TIFF0007708858000004.tif15170
[0063] FIG. 6B shows, according to one embodiment, the step of evaluating the slope of the pixel intensity of the body portion. This step includes identifying three groups of pixels within the pixel block, namely a left portion 660 that may correspond to the right end of a specific key on the keyboard, a right portion 680 that may correspond to the left end of a specific key on the keyboard, and a central portion 670 that is between the left portion 660 and the right portion 680. In one embodiment, identifying the three groups of pixels includes identifying groups of pixels having pixel intensities that are substantially uniform or have a dispersion less than a specific amount. For example, the dispersion of the pixel intensities of the pixels in each of the left portion 660, the central portion 670, and the right portion 680 may be less than a predetermined minimum amount of dispersion. In addition to calculating the dispersion across the three regions, the slope of the pixel intensity between the groups of pixels is determined. The left slope is determined based on the difference between the average pixel intensity value of the left portion 660 and the average pixel intensity value of the central portion 670. ∇ left =μ middle -μ left
[0064] The right slope is determined based on the difference between the average pixel intensity values of the central portion 670 and the right portion 680. ∇ right =μ right -μ middle
[0065] The combined gradient is ∇ vert = ∇ left - ∇ right as follows.
[0066] The deviation is TIFF0007708858000005.tif16170.
[0067] The vertical response referred to in FIG. 6B combines the dispersions of different regions corresponding to the T-shaped body portion so as not to respond to changing luminance, contrast, and image noise. TIFF0007708858000006.tif16170
[0068] If both the left gradient and the right gradient are greater than a specific minimum gradient value, the vertical response is calculated based on that. Alternatively, if either the left gradient or the right gradient is determined to be less than the minimum gradient value, a zero value is assigned to the vertical response. TIFF0007708858000007.tif23170
[0069] If the sign of the combined gradient is incorrect, a zero value is assigned to the vertical response. TIFF0007708858000008.tif15170
[0070] In one embodiment, the T-shape is identified by combining a horizontal response and a vertical response (e.g., by either summation or multiplication) and determining whether the combined response value is greater than a predetermined minimum response value. If the combined response value is less than the minimum response value, the pixel block is determined not to contain a T-shape. In one embodiment where multiple adjacent combined responses are detected, the combined response having the strongest response value, or the strongest absolute response value, may be selected for that group. In some embodiments, referring to FIG. 6A, the upper and lower portions may be configured such that the size and shape of the upper portion 610 substantially match the size and shape of the lower portion 630. This enables the horizontal response to be uniform regardless of whether the horizontal response is calculated from top to bottom or bottom to top, and further enables the horizontal response to be combined with the vertical response corresponding to a T-shape oriented downward or the vertical response corresponding to a T-shape oriented upward. For example, the upper portion 610 may be divided into two parts corresponding to the two lower portions 630. Alternatively, the two lower portions 630 may be combined into a single part instead of being separated by the body portion of the T-shape template. In some embodiments, referring to FIG. 6B, the height of the central portion 670 may be adjusted to substantially match the heights of the left portion 660 and the right portion 680. This enables the vertical response to be uniform regardless of whether the vertical response is calculated from right to left or left to right, and further enables the vertical response to be combined with the horizontal response corresponding to a T-shape oriented downward or upward.
[0071] In one embodiment, the VR model generator 320 generates, obtains, or identifies a virtual model for a keyboard based on shape features and / or corner features. FIG. 5A shows examples of shape features (e.g., 501, 502, 503, 510, and 520) and corner features (e.g., 550). In one embodiment, the virtual model of the physical keyboard is generated based on features detected on the physical keyboard, such as corner features and shape features including the orientation and sign of the shape features. For example, FIG. 7 shows an exemplary keyboard 710 having some of the detected shape features based on which the virtual model 720 can be generated. In one embodiment, the virtual model can be a three-dimensional virtual object, and each detected feature may be assigned a three-dimensional position on the virtual model. In some embodiments, the VR model generator 320 can generate a simplified version of the virtual model as a two-dimensional virtual object and assign a two-dimensional position to each of the detected features. The simplified virtual model can enable the matching process described below to be performed more quickly and efficiently. In one embodiment, the virtual model may be generated to include information regarding the topology of the keyboard. The topology of the keyboard includes information regarding, for example, the number of columns of the keyboard, the identification of the column to which the shape feature belongs, the size of the keys, the distance between the keyboard columns, and the like.
[0072] Figures 8A-8B show exemplary keyboards demonstrating a method for determining the pose of a physical keyboard based on the characteristics of the keyboard. The method can start by capturing an image from a viewpoint corresponding to the imaging sensor of the HMD. An exemplary image is shown in Figure 8A. In some embodiments where the image is captured by an imaging sensor with a fish-eye lens, the image may appear to be distorted or warped. In such embodiments, the image may be processed to remove the distortion. For example, Figure 8A shows an image 810 that has been processed to remove distortion from a fish-eye lens. In some embodiments, if the keyboard depicted in the image does not appear to be rectangular (e.g., appears trapezoidal), the keyboard may be corrected to be rectangular. For example, Figure 8B shows a keyboard image 840 that has been corrected to be rectangular. In some embodiments, a stereo camera may be used to capture two images of the keyboard, i.e., one image from each of the two lenses of the stereo camera. In such embodiments, the shape features may be detected separately in the two images, and it becomes possible to calculate the pose of the keyboard based on the shape features detected in the two images.
[0073] In one embodiment, edge-based detection may be used to determine whether an arbitrary keyboard is depicted in an image (e.g., at an initial detection stage). Edge-based detection differs from shape-based detection in that it looks for the outer edges of the keyboard (e.g., the four sides of the keyboard) rather than the shape features of the keyboard (e.g., the T-shaped feature). Edge-based detection is usually not as accurate in determining the pose of the keyboard, but it is still possible to determine the approximate pose of the keyboard. Moreover, assuming that edge-based detection has a lower computational cost than shape-based detection, it may be more suitable for the initial detection stage where the goal is to determine whether an arbitrary keyboard is depicted in the image. In some embodiments, shape-based detection may be used instead of edge-based detection at the initial keyboard detection stage. For example, in a situation where a part of the keyboard is blocked (e.g., by the user's hand) or the features of the keyboard are not detected, the number of feature correspondence relationships detected using shape-based detection is much larger than the number of edge-based detections (e.g., usually, a keyboard has more than 100 shape features for four edge features), so shape-based detection may be more suitable for keyboard detection. In other words, if a part of the keyboard edge is blocked, since a typical keyboard has only four edges (sides), edge-based detection may not be able to detect the keyboard, but shape-based detection may be able to detect the keyboard even if a part of many shape features depends on other unblocked shape features and is blocked. In some embodiments, a deep learning approach may be applied to the image to detect the approximate positions of the four corners of the keyboard in order to determine whether an arbitrary keyboard is depicted in the image (e.g., at the initial detection stage).
[0074] In one embodiment, after identifying a keyboard in an image (e.g., based on edge-based detection) and determining a rough pose of the keyboard, shape-based detection may be used to determine the exact pose of the keyboard. In one embodiment, an image of the keyboard may be divided into blocks of pixels, and then the shape features of the keyboard may be detected by comparing each of the pixel blocks with a shape pattern template (e.g., the T-shaped template, X-shaped template, or L-shaped template shown in FIGS. 6A and 6B). For example, FIG. 8B shows a keyboard 840 having shape features detected on the keyboard. In one embodiment, the detected shape features are clustered into groups based on the keyboard columns. The cluster of features provides a unique characteristic pattern that can be compared with the characteristic pattern of the virtual models stored in the database to find a virtual model that matches the keyboard. When a matching virtual model is found, the virtual model in the virtual space is projected towards the viewpoint of the camera that captured the image, and the exact pose of the keyboard can be determined by adjusting the pose of the virtual model until the feature correspondence between the keyboard and the virtual model matches (e.g., by minimizing the projection error). In one embodiment, if a virtual model of the physical keyboard does not exist in or cannot be found in the database, the virtual model of the keyboard may be generated based on the detected features of the physical keyboard and stored in the database.
[0075] In one embodiment, after an accurate pose for the keyboard has been determined, the keyboard may be continuously tracked to determine whether the keyboard is still at the expected position. Otherwise, the pose of the keyboard may be determined again by performing edge-based detection or shape-based detection. In one embodiment, in order to minimize the computational cost associated with the tracking process, since edge-based detection has a lower computational cost than shape-based detection, edge-based detection may be used first to determine whether the keyboard is at the expected position. If edge-based detection fails to detect the keyboard, shape-based detection may be used instead. In one embodiment, the tracking process may include an active stage in which the object detector 310 actively tracks the keyboard and an idle stage in which the keyboard is not being tracked. The object detector 310 can only perform the active stage intermittently in order to minimize computational cost and improve the battery life of the HMD. In one embodiment, the object detector 310 can adjust the length of time for which the active stage and the idle stage are performed based on various conditions. For example, the object detector 310 can first perform the active stage for a period longer than the idle stage, but if the keyboard remains stationary for a significant period of time, the object detector 310 can adjust the execution so that the idle stage is performed longer than the active stage. In some embodiments, the object detector 310 can manually perform the active stage in situations where there is an indication that the keyboard may have been moved by the user (e.g., when the user's hand is detected to be near the left and right sides of the keyboard, or when the user presses a specific key).
[0076] FIG. 9 shows an exemplary method 900 for determining the pose of a keyboard. The method can start, in step 901, by capturing an image from a camera viewpoint, where the image depicts a physical keyboard. In step 902, the method can continue by detecting one or more shape features of the physical keyboard depicted in the image by comparing the pixels of the image to a predetermined shape pattern, where the predetermined shape pattern represents the visual characteristics of the spaces between the keyboard keys. In step 903, the method can continue by accessing the predetermined shape features of a keyboard model associated with the physical keyboard. In step 904, the method can continue by determining the pose of the physical keyboard based on a comparison between (1) the detected one or more shape features of the physical keyboard and (2) the projection of the predetermined shape features of the keyboard model towards the camera viewpoint. Particular embodiments can repeat one or more steps of the method of FIG. 9, as needed. Although the present disclosure describes and illustrates the particular steps of the method of FIG. 9 as occurring in a particular order, the present disclosure contemplates that any suitable steps of the method of FIG. 9 can occur in any suitable order. Moreover, although the present disclosure describes and illustrates an exemplary method for determining the pose of a keyboard, the present disclosure contemplates any suitable method for determining the pose of a keyboard that, as needed, can include all, some, or none of the steps of the method of FIG. 9 and any suitable additional steps. Further, although the present disclosure describes and illustrates particular components, devices, or systems for performing the particular steps of the method of FIG. 9, the present disclosure contemplates that any suitable combination of any suitable components, devices, or systems can perform any suitable steps of the method of FIG. 9.
[0077] FIG. 10 shows an exemplary network environment 1000 associated with a social networking system. The network environment 1000 includes client systems 1030, a social networking system 1060, and third-party systems 1070 interconnected with each other by a network 1010. FIG. 10 shows a particular configuration of the client systems 1030, the social networking system 1060, the third-party systems 1070, and the network 1010, but the present disclosure contemplates any suitable configuration of the client systems 1030, the social networking system 1060, the third-party systems 1070, and the network 1010. By way of example and not limitation, two or more of the client systems 1030, the social networking system 1060, and the third-party systems 1070 may be directly connected to each other, bypassing the network 1010. As another example, two or more of the client systems 1030, the social networking system 1060, and the third-party systems 1070 may be located wholly or partially in the same physical or logical location as each other. For instance, an AR / VR headset 1030 may be connected to a local computer or a mobile computing device 1070 via short-range wireless communication (e.g., Bluetooth). Moreover, FIG. 10 shows a particular number of client systems 1030, the social networking system 1060, the third-party systems 1070, and the network 1010, but the present disclosure contemplates any suitable number of client systems 1030, the social networking system 1060, the third-party systems 1070, and the network 1010. By way of example and not limitation, the network environment 1000 may include multiple client systems 1030, the social networking system 1060, the third-party systems 1070, and the network 1010.
[0078] The present disclosure contemplates any suitable network 1010. By way of example and not limitation, one or more portions of network 1010 may include a short-range wireless network (e.g., Bluetooth, Zigbee, etc.), an ad hoc network, an intranet, an extranet, a virtual private network (VPN), a local area network (LAN), a wireless LAN (WLAN), a wide area network (WAN), a wireless WAN (WWAN), a metropolitan area network (MAN), a portion of the Internet, a portion of the public switched telephone network (PSTN), a cellular phone network, or a combination of two or more of these. Network 1010 may include one or more networks 1010.
[0079] Link 1050 can connect the client system 1030, the social networking system 1060, and the third-party system 1070 to the communication network 1010 or to each other. The present disclosure contemplates any suitable link 1050. In certain embodiments, one or more links 1050 are one or more wired links (such as, for example, a digital subscriber line (DSL) or a data over cable service interface specification (DOCSIS)), wireless links (such as, for example, Wi-Fi, Worldwide Interoperability for Microwave Access (WiMAX), Bluetooth), or optical links (such as, for example, Synchronous Optical Network (SONET) or Synchronous Digital Hierarchy (SDH)). In certain embodiments, one or more links 1050 each include an ad hoc network, an intranet, an extranet, a VPN, a LAN, a WLAN, a WAN, a WWAN, a MAN, a portion of the Internet, a portion of the PSTN, a cellular technology-based network, a satellite communication technology-based network, another link 1050, or a combination of two or more such links 1050. The links 1050 need not necessarily be the same throughout the network environment 1000. One or more first links 1050 may differ from one or more second links 1050 in one or more respects.
[0080] In certain embodiments, client system 1030 can be an electronic device that includes hardware, software, or embedded logic components, or a combination of two or more such components, and is capable of performing the appropriate functions implemented or supported by client system 1030. By way of example and not limitation, client system 1030 can include a computer system such as a VR / AR headset, a desktop computer, a notebook or laptop computer, a netbook, a tablet computer, an e-book reader, a GPS device, a camera, a personal digital assistant (PDA), a handheld electronic device, a cellular phone, a smartphone, an augmented / virtual reality device, other suitable electronic devices, or any suitable combination thereof. The present disclosure contemplates any suitable client system 1030. Client system 1030 can enable a network user of client system 1030 to access network 1010. Client system 1030 can enable its user to communicate with other users in other client systems 1030.
[0081] In certain embodiments, the social networking system 1060 can be a network-addressable computing system that hosts an online social network. The social networking system 1060 can generate, store, receive, and transmit social networking data such as, for example, user profile data, concept profile data, social graph information, or other suitable data related to the online social network. The social networking system 1060 may be accessed by other components of the network environment 1000, either directly or via the network 1010. By way of non-limiting example, the client system 1030 can access the social networking system 1060, either directly or via the network 1010, using a web browser, or a native application associated with the social networking system 1060 (e.g., a mobile social networking application, a messaging application, another suitable application, or any combination thereof). In certain embodiments, the social networking system 1060 may include one or more servers 1062. Each server 1062 can be a single server or a distributed server spanning multiple computers or multiple data centers. The servers 1062 can be various types of servers such as, for example, but not limited to, web servers, news servers, mail servers, message servers, advertising servers, file servers, application servers, exchange servers, database servers, proxy servers, or other servers suitable for performing the functions or processes described herein, or any combination thereof. In certain embodiments, each server 1062 may include hardware, software, or embedded logic for performing the appropriate functions implemented or supported by the server 1062, or a combination of two or more such components. In certain embodiments, the social networking system 1060 may include one or more data stores 1064.Data store 1064 may be used to store various types of information. In certain embodiments, the information stored in data store 1064 may be organized according to a proprietary data structure. In certain embodiments, each data store 1064 may be a relational database, columnar database, correlation database, or other suitable database. The present disclosure describes or illustrates particular types of databases, but the present disclosure contemplates any suitable type of database. Certain embodiments may provide an interface that enables client system 1030, social networking system 1060, or third party system 1070 to manage, retrieve, modify, add, or delete information stored in data store 1064.
[0082] In certain embodiments, social networking system 1060 may store one or more social graphs in one or more data stores 1064. In certain embodiments, a social graph may include a plurality of nodes (each corresponding to a particular user) or a plurality of concept nodes (each corresponding to a particular concept), and a plurality of edges connecting the nodes. Social networking system 1060 may provide users of the online social network with the ability to communicate and interact with other users. In certain embodiments, a user may join the online social network via social networking system 1060 and then add connections (e.g., relationships) to some of the other users of social networking system 1060 that the user wishes to be connected to. As used herein, the term “friend” may refer to any other user of social networking system 1060 with whom a user has connected, associated, or formed a relationship via social networking system 1060.
[0083] In certain embodiments, the social networking system 1060 can provide a user with the ability to take actions on various types of items or objects supported by the social networking system 1060. By way of non-limiting example, items and objects can include groups or social networks to which a user of the social networking system 1060 can belong, events or calendar entries that a user might be interested in, computer-based applications that a user can use, transactions that enable a user to purchase or sell items via a service, interactions with advertisements that a user can perform, or other suitable items or objects. The user can interact with anything that can be represented within the social networking system 1060 or by an external system of a third-party system 1070 separate from the social networking system 1060 and coupled to the social networking system 1060 via the network 1010.
[0084] In certain embodiments, the social networking system 1060 may be able to link various entities. By way of non-limiting example, the social networking system 1060 can enable a user to interact with each other and receive content from a third-party system 1070 or other entities, or enable a user to interact with these entities via an application programming interface (API) or other communication channels.
[0085] In certain embodiments, third-party system 1070 may include a local computing device communicatively coupled to client system 1030. For example, if client system 1030 is an AR / VR headset, third-party system 1070 may be a local laptop configured to perform the necessary graphic rendering and provide the rendering results to AR / VR headset 1030 for subsequent processing and / or display. In certain embodiments, third-party system 1070 may be able to execute software associated with client system 1030 (e.g., a rendering engine). Third-party system 1070 may generate a sample data set having sparse pixel information of video frames and transmit the sparse data to client system 1030. Client system 1030 may then be able to generate a frame restored from the sample data set.
[0086] In certain embodiments, third-party system 1070 may also include one or more types of servers, one or more data stores, one or more interfaces including but not limited to APIs, one or more web services, one or more content sources, one or more networks, or any other suitable components that the server can communicate with. Third-party system 1070 may be operated by an entity different from the entity operating social networking system 1060. However, in certain embodiments, social networking system 1060 and third-party system 1070 may operate in cooperation with each other to provide social networking services to users of social networking system 1060 or third-party system 1070. In this sense, social networking system 1060 can provide a platform or backbone that other systems such as third-party system 1070 can use to provide social networking services and functions to users over the Internet.
[0087] In certain embodiments, third-party system 1070 may include a third-party content object provider (e.g., including the sparse sample data sets described herein). The third-party content object provider may include one or more sources of content objects that can be communicated to client system 1030. By way of example and not limitation, the content object may include information about things or activities of interest to the user, such as, for example, movie showtimes, movie reviews, restaurant reviews, restaurant menus, product information and reviews, or other suitable information. By way of another example and not limitation, the content object may include incentive content objects such as coupons, discount tickets, gift vouchers, or other suitable incentive objects.
[0088] In certain embodiments, social networking system 1060 also includes user-generated content objects that can improve a user's interaction with social networking system 1060. User-generated content may include anything that a user can add, upload, send, or "post" to social networking system 1060. By way of example and not limitation, a user communicates a post from client system 1030 to social networking system 1060. The post may include data such as a status update or other text data, location information, photographs, videos, links, music, or other similar data or media. The content may also be added to social networking system 1060 by a third party via a "communication channel" such as a news feed or stream.
[0089] In certain embodiments, the social networking system 1060 may include various servers, subsystems, programs, modules, logs, and data stores. In certain embodiments, the social networking system 1060 may include one or more of the following: a web server, an action logger, an API request server, a relevance and ranking engine, a content object classifier, a notification controller, an action log, a third party content object publication log, an inference module, an authorization / privacy server, a search module, an advertising targeting module, a user interface module, a user profile store, a connections store, a third party content store, or a location store. The social networking system 1060 may also include suitable components such as a network interface, a security mechanism, a load balancer, a failover server, a management and network operations console, other suitable components, or any suitable combination thereof. In certain embodiments, the social networking system 1060 may include one or more user profile stores for storing user profiles. A user profile may include, for example, historical information, demographic information, behavioral information, social information, or other types of descriptive information such as employment history, educational history, hobbies or preferences, interests, affinities, or location. Interest information may include interests related to one or more categories. Categories may be general or specific. By way of example and not limitation, if a user indicates "like" for an article related to a brand of shoes, the category may be the brand or a general category such as "shoes" or "clothing". A connections store may be used to store connections information about users. Connections information may indicate users who have similar or common employment histories, group memberships, hobbies, educational histories, or are related in any way or share common attributes. Connections information may also include user-defined connections between different users (both internal and external) and content.The web server may be used to link the social networking system 1060 via the network 1010 to one or more client systems 1030 or one or more third - party systems 1070. The web server may include a mail server or other messaging function for receiving and routing messages between the social networking system 1060 and one or more client systems 1030. The API request server can enable third - party systems 1070 to access information from the social networking system 1060 by calling one or more APIs. The action logger may be used to receive communications from the web server regarding a user's actions on or away from the social networking system 1060. Along with the action log, a third - party content object log regarding user exposure to third - party content objects may be maintained. The notification controller can provide information regarding content objects to the client system 1030. The information may be pushed to the client system 1030 as a notification, or the information may be pulled from the client system 1030 in response to a request received from the client system 1030. The authorization server may be used to enforce one or more privacy settings of users of the social networking system 1060. The user's privacy settings determine how specific information associated with the user can be shared. The authorization server can enable the user to opt - in or opt - out of having their actions recorded by the social networking system 1060 or shared with other systems (e.g., third - party systems 1070), for example, by setting appropriate privacy settings.The third-party content object store may be used to store content objects received from a third party such as third-party system 1070. The location store may be used to store location information received from client system 1030 associated with a user. The advertising price setting module may combine social information, the current time, location information, or other suitable information to provide relevant advertisements to the user in the form of notifications.
[0090] FIG. 11 shows an exemplary computer system 1100. In certain embodiments, one or more computer systems 1100 perform one or more steps of one or more of the methods described or illustrated herein. In certain embodiments, one or more computer systems 1100 provide the functions described or illustrated herein. In certain embodiments, software running on one or more computer systems 1100 performs one or more steps of one or more of the methods described or illustrated herein, or provides the functions described or illustrated herein. Certain embodiments include one or more portions of one or more computer systems 1100. As used herein, reference to a computer system may, where appropriate, include a computing device, and vice versa. Moreover, reference to a computer system may, where appropriate, include one or more computer systems.
[0091] The present disclosure contemplates any suitable number of computer systems 1100. The present disclosure contemplates computer systems 1100 in any suitable physical form. By way of example and not limitation, computer system 1100 can be an embedded computer system, a system-on-chip (SOC), a single-board computer system (SBC) (such as, for example, a computer-on-module (COM) or a system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive kiosk, a mainframe, a mesh of computer systems, a cellular phone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more of these. Optionally, computer system 1100 can include one or more computer systems 1100, can be single or distributed, can span multiple locations, can span multiple machines, can span multiple data centers, or can exist within a cloud that optionally includes one or more cloud components within one or more networks. Optionally, one or more computer systems 1100 can perform one or more steps of one or more of the methods described or illustrated herein without substantial spatial or temporal limitation. By way of example and not limitation, one or more computer systems 1100 can perform one or more steps of one or more of the methods described or illustrated herein in real time or in batch mode. One or more computer systems 1100 can perform one or more steps of one or more of the methods described or illustrated herein at different times or at different locations, as required.
[0092] In certain embodiments, computer system 1100 includes a processor 1102, a memory 1104, a storage 1106, an input / output (I / O) interface 1108, a communication interface 1110, and a bus 1112. Although this disclosure describes and shows a particular computer system having a particular number of particular components in a particular arrangement, the disclosure contemplates any suitable computer system having any suitable number of any suitable components in any suitable arrangement.
[0093] In certain embodiments, processor 1102 includes hardware for executing instructions, such as instructions to create a computer program. By way of non-limiting example, to execute instructions, processor 1102 can retrieve (or fetch) instructions from internal registers, internal caches, memory 1104, or storage 1106, decode and execute them, and then write one or more results to internal registers, internal caches, memory 1104, or storage 1106. In certain embodiments, processor 1102 may include one or more internal caches for data, instructions, or addresses. The present disclosure contemplates processor 1102 including any suitable number of any suitable internal caches, as appropriate. By way of non-limiting example, processor 1102 may include one or more instruction caches, one or more data caches, and one or more translation lookaside buffers (TLBs). Instructions in the instruction cache can be copies of instructions in memory 1104 or storage 1106, and the instruction cache can speed up the retrieval of those instructions by processor 1102. Data in the data cache can be a copy of data in memory 1104 or storage 1106 that the instructions being executed by processor 1102 operate on, results of previous instructions executed by processor 1102 for access by subsequent instructions being executed by processor 1102, or for writing to memory 1104 or storage 1106, or other suitable data. The data cache can speed up read or write operations by processor 1102. The TLB can speed up virtual address translation for processor 1102. In certain embodiments, processor 1102 may include one or more internal registers for data, instructions, or addresses. The present disclosure contemplates processor 1102 including any suitable number of any suitable internal registers, as appropriate.Optionally, processor 1102 may include one or more arithmetic logic units (ALUs), be a multi-core processor, or include one or more processors 1102. This disclosure describes and shows particular processors, but this disclosure contemplates any suitable processor.
[0094] In certain embodiments, memory 1104 includes main memory for storing instructions for execution by processor 1102 or data on which processor 1102 operates. By way of example and not limitation, computer system 1100 can load instructions into memory 1104 from storage 1106 or another source (such as another computer system 1100). Processor 1102 can then load instructions from memory 1104 into internal registers or an internal cache. To execute the instructions, processor 1102 can fetch the instructions from the internal registers or internal cache and decode them. During or after execution of the instructions, processor 1102 can write one or more results (which can be intermediate or final results) into internal registers or an internal cache. Processor 1102 can then write one or more of those results into memory 1104. In certain embodiments, processor 1102 executes only instructions within one or more internal registers or an internal cache or in memory 1104 (as contrasted with storage 1106 or other locations) and operates only on data within one or more internal registers or an internal cache or in memory 1104 (as contrasted with storage 1106 or other locations). One or more memory buses (which may each include an address bus and a data bus) can couple processor 1102 to memory 1104. Bus 1112 can include one or more memory buses, as described below. In certain embodiments, one or more memory management units (MMUs) exist between processor 1102 and memory 1104 to facilitate access to memory 1104 requested by processor 1102. In certain embodiments, memory 1104 includes random access memory (RAM). This RAM can be volatile memory, if appropriate. If appropriate, this RAM can be dynamic RAM (DRAM) or static RAM (SRAM). Moreover, if appropriate, this RAM can be single-port or multi-port RAM. The present disclosure contemplates any suitable RAM.Memory 1104 may, if necessary, include one or more memories 1104. Although the present disclosure describes and shows specific memories, the present disclosure contemplates any suitable memory.
[0095] In certain embodiments, storage 1106 includes mass storage for data or instructions. By way of example and not limitation, storage 1106 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or universal serial bus (USB) drive, or a combination of two or more of these. Storage 1106 may, if necessary, include removable media or non-removable (or fixed) media. Storage 1106 may, if necessary, be internal or external to computer system 1100. In certain embodiments, storage 1106 is non-volatile solid state memory. In certain embodiments, storage 1106 includes read only memory (ROM). Where appropriate, this ROM may be mask-programmed ROM, programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), electrically rewritable ROM (EAROM), or flash memory, or a combination of two or more of these. The present disclosure contemplates mass storage 1106 in any suitable physical form. Storage 1106 may, if necessary, include one or more storage control units to facilitate communication between processor 1102 and storage 1106. If necessary, storage 1106 may include one or more storages 1106. Although the present disclosure describes and shows specific storages, the present disclosure contemplates any suitable storage.
[0096] In certain embodiments, I / O interface 1108 includes one or more interfaces for communication between computer system 1100 and one or more I / O devices, including hardware, software, or both. Computer system 1100 may optionally include one or more of these I / O devices. One or more of these I / O devices can enable communication between a person and computer system 1100. By way of example and not limitation, I / O devices can include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, steel camera, stylus, tablet, touch screen, trackball, video camera, another suitable I / O device, or a combination of two or more of these. The I / O devices can include one or more sensors. The present disclosure contemplates any suitable I / O devices and any suitable I / O interface 1108 therefor. Optionally, I / O interface 1108 can include one or more devices or software drivers that enable processor 1102 to drive one or more of these I / O devices. I / O interface 1108 can optionally include one or more I / O interfaces 1108. The present disclosure describes and shows particular I / O interfaces, but the present disclosure contemplates any suitable I / O interface.
[0097] In certain embodiments, communication interface 1110 includes one or more interfaces for communication (such as packet-based communication) between computer system 1100 and one or more other computer systems 1100 or one or more networks, including hardware, software, or both. By way of non-limiting example, communication interface 1110 may include a network interface controller (NIC) or network adapter for communicating with Ethernet or other wire-based networks, or a wireless NIC (WNIC) or wireless adapter for communicating with wireless networks such as WI-FI networks. The present disclosure contemplates any suitable network and any suitable communication interface 1110 therefor. By way of non-limiting example, computer system 1100 may communicate with an ad hoc network, a personal area network (PAN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), or one or more portions of the Internet, or a combination of two or more of these. One or more portions of one or more of these networks may be wired or wireless. By way of example, computer system 1100 may communicate with a wireless personal area network (WPAN) (such as BLUETOOTH WPAN), a WI-FI network, a WI-MAX network, a cellular phone network (such as a global system for mobile communications (GSM) network for mobile communication), or other suitable wireless network, or a combination of two or more of these. Computer system 1100 may optionally include any suitable communication interface 1110 for any of these networks. Communication interface 1110 may optionally include one or more communication interfaces 1110. The present disclosure describes and shows specific communication interfaces, but the present disclosure contemplates any suitable communication interface.
[0098] In certain embodiments, bus 1112 includes hardware, software, or both that couple components of computer system 1100 to each other. By way of example and not limitation, bus 1112 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Extended Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), a HYPERTRANSPORT (HT) interconnect, an Industry Standard Architecture (ISA) bus, an INFINIBAND interconnect, a Low Pin Count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or another suitable bus, or a combination of two or more of these. Bus 1112 may optionally include one or more buses 1112. Although the present disclosure describes and shows specific buses, the present disclosure contemplates any suitable bus or interconnect.
[0099] As used herein, one or more computer-readable non-transitory storage media may include, where appropriate, one or more semiconductor-based or other integrated circuits (ICs) (such as a Field Programmable Gate Array (FPGA) or an Application Specific Integrated Circuit (ASIC)), a hard disk drive (HDD), a hybrid hard drive (HHD), an optical disk, an optical disk drive (ODD), a magneto-optical disk, a magneto-optical drive, a floppy disk, a floppy disk drive (FDD), magnetic tape, a solid state drive (SSD), a RAM drive, a Secure Digital card or drive, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these. A computer-readable non-transitory storage media may, where appropriate, be volatile, non-volatile, or a combination of volatile and non-volatile.
[0100] As used herein, unless explicitly stated otherwise or indicated otherwise by the context, "or" is inclusive and not exclusive. Thus, as used herein, "A or B" means "A, B, or both", unless explicitly stated otherwise or indicated otherwise by the context. Further, unless explicitly stated otherwise or indicated otherwise by the context, "and" is both conjunctive and disjunctive. Thus, as used herein, "A and B" means "A and B, jointly or severally", unless explicitly stated otherwise or indicated otherwise by the context.
[0101] The scope of the present disclosure encompasses all changes, substitutions, variations, modifications, and alterations to the exemplary embodiments described or illustrated herein that would be understood by those of ordinary skill in the art. The scope of the present disclosure is not limited to the exemplary embodiments described or illustrated herein. Moreover, the present disclosure describes and illustrates each of the embodiments herein as including particular components, elements, features, functions, operations, or steps, but any of these embodiments may include any combination or substitution of any of the components, elements, features, functions, operations, or steps described or illustrated anywhere herein that would be understood by those of ordinary skill in the art. Further, references in the appended claims to an apparatus or system or a component of an apparatus or system that is adapted, arranged, capable, configured, enabled, operable, or operative to perform a particular function include that apparatus, system, or component, whether or not the particular function is activated, turned on, or unlocked, so long as that apparatus, system, or component is so adapted, arranged, capable, configured, enabled, operable, or operative. Additionally, the present disclosure describes or illustrates particular embodiments as providing particular advantages, but a particular embodiment may provide none, some, or all of these advantages.
Claims
1. capturing an image from a camera viewpoint, wherein the image depicts a physical keyboard; detecting one or more shape features of the physical keyboard depicted in the image by comparing pixels of the image with a predetermined shape template, wherein the predetermined shape template represents visual characteristics of spaces between keyboard keys; accessing predetermined shape features of a keyboard model associated with the physical keyboard; and determining a pose of the physical keyboard based on a comparison between (1) the detected one or more shape features of the physical keyboard and (2) a projection of the predetermined shape features of the keyboard model towards the camera viewpoint. A method comprising the above steps.
2. The method of claim 1, wherein the predetermined shape template is a T-shaped template having a body portion and a roof portion, and the body portion represents visual characteristics of a space between two keyboard keys within a first keyboard row.
3. The method of claim 2, wherein the roof portion of the T-shaped template represents visual characteristics of a space between (1) a portion of the first keyboard row corresponding to the two keyboard keys within the first keyboard row and (2) one or more keyboard keys within a second keyboard adjacent to the first keyboard row.
4. The method of claim 2, wherein the first keyboard row is either: the topmost row of the keyboard, and the roof portion of the T-shaped template represents visual characteristics of a space above the two keyboard keys within the first keyboard row; or the bottommost row of the keyboard, and the roof portion of the T-shaped template represents visual characteristics of a space below the two keyboard keys within the first keyboard row.
5. The method of claim 1, wherein the predetermined shape template is a T-shaped template having a body portion and a roof portion, and the body portion represents visual characteristics of a space between a first keyboard key within a first keyboard row and a second keyboard key within a second keyboard row adjacent to the first keyboard row. The first keyboard key is the rightmost key within the first keyboard column, the second keyboard key is the rightmost key within the second keyboard column, and the roof portion of the T-shaped template represents the visual characteristics of the space to the right of the first keyboard key and the second keyboard key, or The first keyboard key is the leftmost key within the first keyboard column, the second keyboard key is the leftmost key within the second keyboard column, and the roof portion of the T-shaped template represents the visual characteristics of the space to the left of the first keyboard key and the second keyboard key, The method according to claim 1.
6. Comparing the pixels of the image with the predetermined shape template is Dividing the image into a plurality of pixel blocks, each pixel block including a plurality of pixels, and dividing, For each of the pixel blocks, Determining the visual characteristics of the pixel block based on the pixel intensities associated with the plurality of pixels within the pixel block, Comparing the visual characteristics of the pixel block with the predetermined shape template representing the visual characteristics of the space between the keyboard keys, Detecting the one or more shape features of the physical keyboard depicted in the image based on the comparison between the visual characteristics of the pixel block of the image and the predetermined shape template representing the visual characteristics of the space between the keyboard keys Including, preferably, determining the visual characteristics of the pixel block based on the pixel intensities associated with the plurality of pixels within the pixel block is Comparing the pixel intensities associated with at least one portion of the pixel block with another portion of the pixel block to determine the gradients and variances of the pixel intensities associated with the various portions of the pixel block The method according to any one of claims 1 to 5, including.
7. Rendering the representation of the keyboard model to match the pose of the physical keyboard The method according to any one of claims 1 to 6, further including.
8. Before accessing the predetermined shape features of the keyboard model associated with the physical keyboard, identifying the keyboard model from the database including the plurality of keyboard models by comparing the detected one or more shape features of the physical keyboard with the predetermined shape features of the plurality of keyboard models in the database The method according to any one of claims 1 to 7, further comprising.
9. The method according to claim 1, wherein the predetermined shape template is an L-shaped template or an X-shaped template.
10. capturing an image from a camera viewpoint, wherein the image depicts a physical keyboard, detecting one or more shape features of the physical keyboard depicted in the image by comparing the pixels of the image with a predetermined shape template, wherein the predetermined shape template represents visual characteristics of spaces between keyboard keys, accessing predetermined shape features of a keyboard model associated with the physical keyboard, determining the pose of the physical keyboard based on a comparison between (1) the detected one or more shape features of the physical keyboard and (2) a projection of the predetermined shape features of the keyboard model towards the camera viewpoint One or more computer-readable non-transitory storage media embodying software that is operable when executed to perform the above.
11. The one or more computer-readable non-transitory storage media according to claim 10, wherein the predetermined shape template is a T-shaped template having a body portion and a roof portion, the body portion representing visual characteristics of a space between two keyboard keys in a first keyboard row, and preferably, the roof portion of the T-shaped template representing visual characteristics of a space between (1) a portion of the first keyboard row corresponding to the two keyboard keys in the first keyboard row and (2) one or more keyboard keys in a second keyboard adjacent to the first keyboard row.
12. comparing the pixels of the image with the predetermined shape template is dividing the image into a plurality of pixel blocks, each pixel block including a plurality of pixels, For each of the pixel blocks, determining visual characteristics of the pixel block based on pixel intensities associated with the plurality of pixels within the pixel block; comparing the visual characteristics of the pixel block with the predetermined shape template representing the visual characteristics of the space between the keyboard keys; detecting the one or more shape features of the physical keyboard depicted in the image based on the comparison between the visual characteristics of the pixel block of the image and the predetermined shape template representing the visual characteristics of the space between the keyboard keys including, preferably, determining the visual characteristics of the pixel block based on the pixel intensities associated with the plurality of pixels within the pixel block, comparing the pixel intensities associated with at least one portion of the pixel block with another portion of the pixel block to determine gradients and variances of the pixel intensities associated with the various portions of the pixel block One or more computer-readable non-transitory storage media according to claim 10 or claim 11, including.
13. A system comprising one or more processors and one or more computer-readable non-transitory storage media in communication with the one or more processors, wherein the one or more computer-readable non-transitory storage media include instructions that, when executed by the one or more processors, capturing an image from a camera perspective, the image depicting a physical keyboard; detecting one or more shape features of the physical keyboard depicted in the image by comparing pixels of the image with a predetermined shape template, the predetermined shape template representing visual characteristics of a space between keyboard keys; accessing predetermined shape features of a keyboard model associated with the physical keyboard; determining a pose of the physical keyboard based on a comparison between (1) the detected one or more shape features of the physical keyboard and (2) a projection of the predetermined shape features of the keyboard model towards the camera perspective causing the system to perform. A system.
14. The predetermined shape template is a T-shaped template having a body portion and a roof portion, the body portion representing the visual characteristics of the space between two keyboard keys within a first keyboard row, and preferably, the roof portion of the T-shaped template representing the visual characteristics of the space between (1) a portion of the first keyboard row corresponding to the two keyboard keys within the first keyboard row and (2) one or more keyboard keys within a second keyboard adjacent to the first keyboard row. The system according to claim 13.
15. comparing the pixels of the image with the predetermined shape template includes dividing the image into a plurality of pixel blocks, each pixel block including a plurality of pixels, for each of the pixel blocks, determining the visual characteristics of the pixel block based on the pixel intensities associated with the plurality of pixels within the pixel block, comparing the visual characteristics of the pixel block with the predetermined shape template representing the visual characteristics of the space between the keyboard keys, detecting the one or more shape features of the physical keyboard depicted in the image based on the comparison between the visual characteristics of the pixel blocks of the image and the predetermined shape template representing the visual characteristics of the space between the keyboard keys and preferably, determining the visual characteristics of the pixel block based on the pixel intensities associated with the plurality of pixels within the pixel block includes comparing the pixel intensities associated with at least one portion of the pixel block with another portion of the pixel block to determine the gradients and variances of the pixel intensities associated with the various portions of the pixel block The system according to claim 13 or claim 14.
Citation Information
Patent Citations
Keyboard for virtual reality, augmented reality, and mixed reality display systems
JP2020521217A