Physical keyboard tracking
By detecting T, X, and L features on the physical keyboard and using gradient and variance methods to correct image distortion, the problem of inaccurate keyboard tracking in virtual reality systems is solved, achieving high-precision physical keyboard tracking.
Patent Information
- Application Number
- CN202180090297.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2020-12-01
- Filing Date
- 2021-11-29
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2041-11-29
AI Technical Summary
Existing technologies struggle to accurately track the position and orientation of physical keyboards in virtual reality systems, especially under low-resolution, distorted, or partially occluded image conditions, leading to inaccurate keyboard feature detection.
By detecting predefined T, X, and L features between keys on a physical keyboard, gradient and variance methods are used to correct image distortion, and the virtual model is rendered to match the posture of the physical keyboard by matching it with keyboard models in the database.
It achieves high-precision keyboard tracking, capable of tracking physical keyboards with sub-millimeter accuracy, and is suitable for virtual reality, augmented reality, and mixed reality systems.
Smart Images

Figure CN116710968B_ABST
Abstract
Description
Technical Field
[0001] This disclosure generally relates to rendering augmented reality (AR) or virtual reality (VR) content on a user device. This disclosure generally relates to input controllers used in AR / VR environments. Background Technology
[0002] Virtual reality (VR) is a computer-generated simulation of an environment (e.g., a 3D environment) that a user can interact with in a seemingly real or physically realistic way. A VR system (which can be a single device or a group of devices) can generate this simulation for display to a user, for example, on a VR headset or some other display device. The simulation can include images, sound, haptic feedback, and / or other senses to mimic a real or fictional environment. As VR becomes increasingly important, its useful applications are rapidly expanding. The most common applications of VR involve games or other interactive content, while others include viewing visual media projects (e.g., photos, videos) for entertainment or training purposes. The feasibility of using VR to simulate real-world conversations and other user interactions is also being explored.
[0003] Augmented reality (AR) provides additional, computer-generated sensory input (e.g., visual, auditory) to the real or physical world. In other words, computer-generated virtual effects can enhance or supplement the real-world field of vision. For example, cameras on a virtual reality headset or head-mounted display (HMD) can capture real-world scenes (as images or videos) and display a composite of the captured scene and computer-generated virtual objects. These virtual objects can be, for example, two-dimensional and / or three-dimensional objects, and can be static or lifelike. Summary of the Invention
[0004] According to a first aspect of this disclosure, a method is provided, the method comprising: by a computing device: acquiring an image from a camera viewpoint, the image depicting a physical keyboard; detecting one or more shape features of the physical keyboard depicted in the image by comparing pixels of the image with a predetermined shape template, the predetermined shape template representing visual features of the space between keyboard keys; accessing predetermined shape features of a keyboard model associated with the physical keyboard; and determining the pose of the physical keyboard based on a comparison between: (1) the detected one or more shape features of the physical keyboard; and (2) the projection of the predetermined shape features of the keyboard model toward the camera viewpoint.
[0005] In some embodiments, the predetermined shape template is a T-shaped template having a trunk portion and a top cover portion, the trunk portion representing the visual characteristics of the space between two keyboard keys in the first keyboard row.
[0006] In some embodiments, the top cover portion of the T-shaped template represents a visual feature of the space between: (1) the portion of the first keyboard row corresponding to two keyboard keys in the first keyboard row; and (2) one or more keyboard keys in a second keyboard row adjacent to the first keyboard row.
[0007] In some embodiments, the first keyboard row is: the top row of the keyboard, where the top cover portion of the T-shaped template represents the visual characteristics of the space above two keyboard keys in the first keyboard row; or the bottom row of the keyboard, where the top cover portion of the T-shaped template represents the visual characteristics of the space below two keyboard keys in the first keyboard row.
[0008] In some embodiments, the predetermined shape template is a T-shaped template having a trunk portion and a top cover portion. The trunk portion represents the visual characteristics of the space between a first key in a first keyboard row and a second key in a second keyboard row, the second keyboard row being adjacent to the first keyboard row. Specifically: the first key is the rightmost key in the first keyboard row, and the second key is also the rightmost key in the second keyboard row; the top cover portion of the T-shaped template represents the visual characteristics of the space to the right of the first and second key; or the first key is the leftmost key in the first keyboard row, and the second key is also the leftmost key in the second keyboard row; the top cover portion of the T-shaped template represents the visual characteristics of the space to the left of the first and second key.
[0009] In some embodiments, comparing the pixels of the image with a predetermined shape template includes: dividing the image into a plurality of pixel blocks, each pixel block including a plurality of pixels; for each of the plurality of pixel blocks: determining a visual feature of the pixel block based on the pixel intensity associated with the plurality of pixels in the pixel block; comparing the visual feature of the pixel block with a predetermined shape template representing visual features of the space between keyboard keys; and detecting one or more shape features of a physical keyboard depicted in the image based on the comparison of the visual feature of the pixel block of the image with the predetermined shape template representing visual features of the space between keyboard keys.
[0010] In some embodiments, determining the visual characteristics of a pixel block based on pixel intensities associated with a plurality of pixels in the pixel block includes comparing pixel intensities associated with at least a portion of the pixel block with pixel intensities associated with another portion of the pixel block to determine the gradient and variance of the pixel intensities associated with each portion of the pixel block.
[0011] In some embodiments, the method further includes: rendering a representation of the keyboard model to match the pose of the physical keyboard.
[0012] In some embodiments, the method further includes: before accessing predetermined shape features of a keyboard model associated with a physical keyboard: identifying the keyboard model from a database comprising a plurality of keyboard models by comparing one or more shape features of the detected physical keyboard with predetermined shape features of the plurality of keyboard models in the database.
[0013] In some embodiments, the predetermined shape template is an L-shaped template or an X-shaped template.
[0014] According to a second aspect of this disclosure, one or more computer-readable non-transitory storage media are provided, the one or more computer-readable non-transitory storage media comprising software that, when executed, is operable to: acquire an image from a camera viewpoint depicting a physical keyboard; detect one or more shape features of the physical keyboard depicted in the image by comparing pixels of the image with a predetermined shape template, the predetermined shape template representing visual features of the space between keyboard keys; access predetermined shape features of a keyboard model associated with the physical keyboard; and determine the orientation of the physical keyboard based on comparisons between: (1) the detected one or more shape features of the physical keyboard; and (2) the projection of the predetermined shape features of the keyboard model toward the camera viewpoint.
[0015] In some embodiments, the predetermined shape template is a T-shaped template having a trunk portion and a top cover portion, the trunk portion representing the visual characteristics of the space between two keyboard keys in the first keyboard row.
[0016] In some embodiments, the top cover portion of the T-shaped template represents a visual feature of the space between: (1) the portion of the first keyboard row corresponding to two keyboard keys in the first keyboard row; and (2) one or more keyboard keys in a second keyboard row adjacent to the first keyboard row.
[0017] In some embodiments, comparing the pixels of the image with a predetermined shape template includes: dividing the image into a plurality of pixel blocks, each pixel block including a plurality of pixels; for each of the plurality of pixel blocks: determining a visual feature of the pixel block based on the pixel intensity associated with the plurality of pixels in the pixel block; comparing the visual feature of the pixel block with a predetermined shape template representing visual features of the space between keyboard keys; and detecting one or more shape features of a physical keyboard depicted in the image based on the comparison of the visual features of the pixel blocks of the image with the predetermined shape template representing visual features of the space between keyboard keys.
[0018] In some embodiments, determining the visual characteristics of a pixel block based on pixel intensities associated with a plurality of pixels in the pixel block includes comparing pixel intensities associated with at least a portion of the pixel block with pixel intensities associated with another portion of the pixel block to determine the gradient and variance of the pixel intensities associated with each portion of the pixel block.
[0019] According to a third aspect of the invention, a system is provided, comprising: one or more processors; and one or more computer-readable non-transitory storage media communicating with the one or more processors, the one or more computer-readable non-transitory storage media including instructions that, when executed by the one or more processors, cause the system to perform the following operations: acquiring an image from a camera viewpoint, the image depicting a physical keyboard; detecting one or more shape features of the physical keyboard depicted in the image by comparing pixels of the image with a predetermined shape template, the predetermined shape template representing visual features of the space between keyboard keys; accessing predetermined shape features of a keyboard model associated with the physical keyboard; and determining the orientation of the physical keyboard based on a comparison between: (1) the detected one or more shape features of the physical keyboard; and (2) the projection of the predetermined shape features of the keyboard model toward the camera viewpoint.
[0020] In some embodiments, the predetermined shape template is a T-shaped template having a trunk portion and a top cover portion, the trunk portion representing the visual characteristics of the space between two keyboard keys in the first keyboard row.
[0021] In some embodiments, the top cover portion of the T-shaped template represents a visual feature of the space between: (1) the portion of the first keyboard row corresponding to two keyboard keys in the first keyboard row; and (2) one or more keyboard keys in a second keyboard row adjacent to the first keyboard row.
[0022] In some embodiments, comparing the pixels of the image with a predetermined shape template includes: dividing the image into a plurality of pixel blocks, each pixel block including a plurality of pixels; for each of the plurality of pixel blocks: determining a visual feature of the pixel block based on the pixel intensity associated with the plurality of pixels in the pixel block; comparing the visual feature of the pixel block with a predetermined shape template representing visual features of the space between keyboard keys; and detecting one or more shape features of a physical keyboard depicted in the image based on the comparison of the visual feature of the pixel block of the image with the predetermined shape template representing visual features of the space between keyboard keys.
[0023] In some embodiments, determining the visual characteristics of a pixel block based on pixel intensities associated with a plurality of pixels in the pixel block includes comparing pixel intensities associated with at least a portion of the pixel block with pixel intensities associated with another portion of the pixel block to determine the gradient and variance of the pixel intensities associated with each portion of the pixel block. Attached Figure Description
[0024] Figure 1 This is a block diagram of an example artificial reality system environment in which the console runs.
[0025] Figure 2 This is a schematic diagram of an HMD based on an example embodiment.
[0026] Figure 3 This is a schematic diagram of the object porting controller.
[0027] Figure 4 The image shown is an example image acquired by the HMD's imaging sensor.
[0028] Figure 5A and Figure 5B An example image of a physical keyboard is shown.
[0029] Figure 6A and Figure 6B A method for detecting the shape features of a keyboard is shown.
[0030] Figure 7 An example keyboard and a virtual model of that keyboard are shown.
[0031] Figure 8A and Figure 8B An example keyboard is shown, which demonstrates a method for determining the orientation of a physical keyboard based on its characteristics.
[0032] Figure 9 An example method for determining the keyboard's orientation is shown.
[0033] Figure 10 An example network environment associated with a social networking system is shown.
[0034] Figure 11 An example computer system is shown. Detailed Implementation
[0035] The embodiments disclosed herein describe, but are not limited to, determining the precise position and orientation (pose) of a physical keyboard, and rendering an image of a virtual model corresponding to the physical keyboard to accurately match the pose of the physical keyboard. High-precision keyboard tracking is challenging. Machine learning methods are too slow for real-time applications providing keyboard tracking and are too inaccurate due to the low quality of the cameras used for tracking. Therefore, a computer vision approach is needed. However, various factors also make computer vision approaches difficult. For example, accurate feature detection based on images captured by an outward-facing camera of a head-mounted display (HMD) is difficult because the captured images may have low resolution and high noise, and the keyboard features may have low contrast. Furthermore, the user's hands often obscure keyboard features, and lighting conditions may not be optimal for accurate feature detection. Since some devices use fisheye lenses for outward-facing cameras to maximize tracking coverage, the captured images may also become distorted or warped.
[0036] To address these issues, the invention disclosed herein seeks predefined T-features, X-features, and L-features formed by the spaces between keys on a physical keyboard. These features are readily detectable even in low-resolution, distorted, or partially occluded keyboard images. The method begins by acquiring an image of the keyboard. If the image is distorted (e.g., due to acquisition via a fisheye lens), it can be corrected by generating a rectilinear image of the keyboard. The gradient and variance-based method then detects the T-features, X-features, and L-features of the keyboard. This gradient and variance-based method utilizes the distribution and differences in pixel intensity: pixel intensity in the boundary regions of each key (a uniform region surrounding the symbol on each key) and pixel intensity in the space between adjacent keys. For any given keyboard, unique T-shaped, L-shaped, and / or X-shaped (or cross-shaped, "+") patterns can be detected based on the differences in pixel intensity. These unique patterns can be compared with pre-mapped models of various keyboards in a database. Once a virtual model matching the physical keyboard is found, it can be rendered to accurately match the keyboard's pose by mapping the virtual model to the detected features of the physical keyboard. This technique allows for tracking of a physical keyboard with sub-millimeter precision. This technique can also be extended to tracking any real-world object that typically has T-features, X-features, and L-features.
[0037] Various embodiments of the present invention may include an artificial reality system or a combination thereof. Artificial reality is a form of reality that has been adjusted in some way before being presented to a user. This artificial reality may include, for example, virtual reality (VR), augmented reality (AR), mixed reality (MR), hybrid reality, or some combination and / or derivative thereof. Artificial reality content may include fully generated content or generated content combined with captured content (e.g., real-world photographs). Artificial reality content may include video, audio, haptic feedback, or some combination thereof, any of which may be presented in a single channel or multiple channels (e.g., stereoscopic video that provides a three-dimensional effect to the viewer). Furthermore, in some embodiments, artificial reality may be associated with applications, products, accessories, services, or some combination thereof, which are used, for example, to create content in artificial reality and / or for use in artificial reality (e.g., to perform actions in artificial reality). Artificial reality systems that deliver artificial reality content can be implemented on a variety of platforms, including head-mounted displays (HMDs) connected to a host computer system, standalone HMDs, mobile devices or computing systems, or any other hardware platform capable of delivering artificial reality content to one or more viewers.
[0038] The embodiments disclosed herein are merely examples, and the scope of this disclosure is not limited to these embodiments. Specific embodiments may include all or some of the components, elements, features, functions, operations, or steps of the embodiments disclosed above, or may exclude the components, elements, features, functions, operations, or steps of the embodiments disclosed above. In particular, embodiments according to the invention are disclosed in the appended claims, which are directed to a method, a storage medium, a system, and a computer program product, wherein any feature mentioned in one claim class (e.g., method) may also be claimed in another claim class (e.g., system). Dependencies or references in the appended claims are chosen solely for formal reasons. However, any subject matter arising from intentional references to any prior claim (particularly multiple dependencies) may also be claimed, such that any combination of the claims and their features is disclosed and can be claimed regardless of the dependencies chosen in the appended claims. Claimable subject matter includes not only combinations of features set forth in the appended claims but also any other combination of features in these claims, wherein each feature mentioned in the claims may be combined with any other feature or combination of features in these claims. Furthermore, any of the embodiments and features described or depicted herein may be claimed in a separate claim, and / or in any combination with any embodiment or feature described or depicted herein or in any combination with any feature of the appended claims.
[0039] This document discloses the following embodiments: These embodiments relate to systems and methods for transferring physical objects from physical space to virtual space in virtual reality. In one aspect, transferring a physical object from physical space to virtual space includes: activating the physical object in physical space to obtain a virtual model of the physical object, and rendering an image of the virtual model in the virtual space.
[0040] In some embodiments, a physical object is activated during pass-through mode, in which the HMD presents or renders the field of view of the physical space to the user of the head-mounted display. For example, a virtual model of a physical object can be identified or selected during pass-through mode, and that virtual model can be rendered in virtual space. Physical objects in the physical space can be tracked, and the position and orientation of the virtual model in the virtual space can be adjusted based on the position and orientation of the physical object in the physical space. In one aspect, instructions for user interaction with physical objects in the physical space can be presented on the virtual model in the virtual space as feedback to the user.
[0041] In one aspect, the physical object is a general-purpose input device (e.g., a keyboard or mouse), which may be manufactured or produced by a different company than those that manufacture or produce head-mounted displays and / or specialized handheld input devices (e.g., pointing devices). By rendering a virtual model of the input device within the user's field of view as a reference or guide for the user (e.g., as a proxy for the input device), the user can easily interact with the virtual model and thus with the input device, providing input to the virtual reality through it.
[0042] In one aspect, spatial feedback regarding the user's interaction with an input device in the physical space can be visually provided to the user relative to a virtual model in a virtual space (e.g., using the virtual model in the virtual space for spatial guidance). In one method, the user's input device in the physical space is detected relative to the input device. A virtual model of the detected input device at a position and orientation (posture) in the virtual space can be presented to the user via a display device. The position and orientation of the virtual model in the virtual space can correspond to (e.g., track or reflect) the position and orientation of the input device relative to the user in the physical space. Relative to this virtual model in the virtual space (and, for example, a virtual representation of the user's hand), spatial feedback regarding the user's interaction with the input device in the physical space can be visually provided to the user through the virtual space. Therefore, through spatial feedback relative to the virtual model, the user can easily locate and touch the input device and provide input through the input device in the physical space while enjoying a virtual reality experience (e.g., while viewing a virtual space rather than a physical space).
[0043] Although the systems and methods disclosed herein may involve porting physical objects to virtual reality, the general principles disclosed herein can be applied to augmented reality or mixed reality.
[0044] Figure 1This is a block diagram of an example artificial reality system environment 100, in which a console 110 operates. In some embodiments, the artificial reality system environment 100 includes an HMD 150 worn by a user and a console 110 that provides artificial reality content to the HMD 150. In one aspect, the HMD 150 can detect its position, orientation, and / or the gaze direction of the user wearing the HMD 150, and can provide the detected position and gaze direction to the console 110. The console 110 can determine a field of view within the artificial reality space corresponding to the detected position, orientation, and / or gaze direction, and generate an image depicting the determined field of view. The console 110 can provide the image to the HMD 150 for rendering. In some embodiments, the artificial reality system environment 100 includes an input device 120 communicatively coupled to the console 110 or the HMD 150 via a wired cable, a wireless link (e.g., Bluetooth, WiFi, etc.), or both. The input device 120 can be dedicated hardware (e.g., a pointing device or controller) with motion sensors, a universal keyboard, a mouse, etc. Users can provide input associated with the presented artificial reality through input device 120. In some embodiments, the artificial reality system environment 100 includes... Figure 1 The components shown may include more or fewer components, or may include components with... Figure 1 The components shown are different components. In some embodiments, the functionality of one or more components of the artificial reality system environment 100 may be distributed among multiple components in a manner different from that described herein. For example, some functions of the console 110 may be performed by the HMD 150. For example, some functions of the HMD 150 may be performed by the console 110. In some embodiments, the console 110 is integrated as part of the HMD 150.
[0045] In some embodiments, HMD 150 includes or corresponds to electronic components that can be worn by a user and can present or provide an artificial reality experience to the user. HMD 150 can render one or more images, videos, audio, or some combination thereof to provide an artificial reality experience to the user. In some embodiments, audio is presented via an external device (e.g., a speaker and / or headphones) that receives audio information from HMD 150, console 110, or both, and presents audio based on that audio information. In some embodiments, HMD 150 includes a sensor 155, a communication interface 165, an image renderer 170, an electronic display 175, and / or an object transfer controller 180. These components can work together to detect the position and orientation of HMD 150, and / or the gaze direction of the user wearing HMD 150, and can render an image of the field of view within the artificial reality corresponding to the detected position and orientation of HMD 150, and / or the user's gaze direction. In other embodiments, HMD 150 includes a larger... Figure 1 The components shown may include more or fewer components, or may include components with... Figure 1 The components shown are different components. In some embodiments, the object migration controller 180 can be activated or deactivated according to control from the user of the HMD 150.
[0046] In some embodiments, sensor 155 includes electronic components, or a combination of electronic and software components, for detecting the position, orientation, and / or user gaze direction of HMD 150. Examples of sensor 155 may include one or more imaging sensors, one or more accelerometers, one or more gyroscopes, one or more magnetometers, a global positioning system, or another suitable type of sensor for detecting motion and / or position. For example, one or more accelerometers may measure translational motion (e.g., forward / backward, up / down, left / right), and one or more gyroscopes may measure rotational motion (e.g., pitch, yaw, roll). In some embodiments, the imaging sensor may acquire images for detecting physical objects, user gestures, hand shapes, user interactions, etc. In some embodiments, sensor 155 detects translational and rotational motion and determines the orientation and position of HMD 150. In one aspect, sensor 155 can detect translational and rotational movements relative to the previous orientation and position of HMD 150, and determine the new orientation and / or position of HMD 150 by accumulating or integrating the detected translational and / or rotational movements. For example, assuming HMD 150 is facing a direction at 25 degrees with respect to a reference direction, sensor 155 can determine that HMD 150 is now facing or oriented at 45 degrees with respect to the reference direction in response to detecting that HMD 150 has rotated 20 degrees. As another example, assuming HMD 150 is located two feet from the reference point in a first direction, sensor 155 can determine that HMD 150 is now located at the vector product of two feet from the reference point in the first direction and three feet from the reference point in the second direction in response to detecting that HMD 150 has moved three feet in a second direction. In one aspect, the user's gaze direction can be determined or estimated based on the position and orientation of HMD 150.
[0047] In some embodiments, sensor 155 may include electronic components, or a combination of electronic and software components, for generating sensor measurement results of the physical space. Examples of sensor 155 for generating sensor measurement results may include one or more imaging sensors, thermal sensors, etc. In one example, the imaging sensor may acquire an image that corresponds to the user's field of view in the physical space (or the field of view seen from the position of HMD 150 according to the orientation of HMD 150). Image processing may be performed on the acquired image to detect physical objects or a portion of the user in the physical space.
[0048] In some embodiments, the communication interface 165 includes electronic components, or a combination of electronic and software components, that communicate with the console 110. The communication interface 165 can communicate with the communication interface 115 of the console 110 via a communication link. This communication link can be a wireless link, a wired link, or both. Examples of a wireless link may include a cellular communication link, a near-field communication link, Wi-Fi, Bluetooth, or any wireless communication link for communication purposes. Examples of a wired link may include a universal serial bus (USB), Ethernet, FireWire, a high-definition multimedia interface (HDMI), or any wired communication link. In embodiments where the console 110 and HMD 150 are implemented on a single system, the communication interface 165 can communicate with the console 110 via at least one bus connection or conductive trace. Through the communication link, the communication interface 165 can send data to the console 110 indicating the determined location and orientation of the HMD 150, and / or the user's gaze direction. Furthermore, through this communication link, the communication interface 165 can receive data from the console 110 indicating the image to be rendered.
[0049] In some embodiments, the image renderer 170 includes electronic components, or a combination of electronic and software components, that generate one or more images for display (e.g., based on changes in field of view in an artificial reality space). In some embodiments, the image renderer 170 is implemented as a processor (or graphics processing unit (GPU)). The image renderer 170 can receive data describing an image to be rendered via communication interface 165 and render the image via electronic display 175. In some embodiments, the data from console 110 may be compressed or encoded, and the image renderer 170 can decompress or decode the data to generate and render the image. The image renderer 170 can receive compressed images from console 110 and decompress the compressed images to reduce the communication bandwidth between console 110 and HMD 150. In one respect, the following process may be computationally expensive and cannot be performed within a single frame (e.g., less than 11 ms): HMD 150 detects the position of HMD 150, the orientation of HMD 150, and / or the gaze direction of the user wearing HMD 150, and console 110 generates a high-resolution image (e.g., 1920 × 1080 pixels) corresponding to the detected position, orientation, and / or gaze direction and sends the high-resolution image to HMD 150. When no image is received from console 110 within that single frame, image renderer 170 can generate one or more images through a shading process and a reprojection process. For example, the shading process and reprojection process can be performed adaptively according to changes in the field of view in artificial reality space.
[0050] In some embodiments, the electronic display 175 is an electronic component that displays images. The electronic display 175 may be, for example, a liquid crystal display or an organic light-emitting diode display. The electronic display 175 may be a transparent display that allows the user to see through it. In some embodiments, when the HMD 150 is worn by a user, the electronic display 175 is located near the user's eyes (e.g., less than 3 inches). In one aspect, the electronic display 175 emits or projects light toward the user's eyes based on an image generated by the image renderer 170.
[0051] In some embodiments, the object migration controller 180 includes electronic components, or a combination of electronic and software components, for activating a physical object and generating a virtual model of the physical object. In one method, the object migration controller 180 detects a physical object in physical space during a pass-through mode, in which a sensor 155 can acquire an image of a user's physical space field of view (or field of view), and an electronic display 175 can present the acquired image to the user. The object migration controller 180 can generate a virtual model of the physical object and present the virtual model in a virtual space, wherein the electronic display 175 can display the user's virtual space field of view. Specific implementations regarding activating a physical object and rendering a virtual model of the physical object are provided below.
[0052] In some embodiments, console 110 is an electronic component, or a combination of electronic and software components, that provides content to be rendered by HMD 150. In one aspect, console 110 includes a communication interface 115 and a content provider 130. These components can work collaboratively to determine an artificial reality field of view corresponding to the location of HMD 150, the orientation of HMD 150, and / or the user's gaze direction of HMD 150, and can generate an artificial reality image corresponding to the determined field of view. In other embodiments, console 110 includes a... Figure 1 The components shown may include more or fewer components, or may include components with... Figure 1 The components shown are different components. In some embodiments, console 110 performs some or all of the functions of HMD 150. In some embodiments, console 110 is integrated as part of HMD 150 as a single device.
[0053] In some embodiments, communication interface 115 is an electronic component, or a combination of electronic and software components, that communicates with HMD 150. Communication interface 115 may be a corresponding component of communication interface 165 for communication via a communication link (e.g., a USB cable). Through the communication link, communication interface 115 can receive data from HMD 150 indicating the determined location of HMD 150, the orientation of HMD 150, and / or the determined gaze direction of the user. Furthermore, through this communication link, communication interface 115 can send data describing an image to be rendered to HMD 150.
[0054] Content provider 130 generates components of the content to be rendered based on the position of HMD 150, the orientation of HMD 150, and / or the user's gaze direction of HMD 150. In one aspect, content provider 130 determines the field of view of artificial reality (HWD) based on the position of HMD 150, the orientation of HMD 150, and / or the user's gaze direction of HMD 150. For example, content provider 130 maps the position of HMD 150 in physical space to its position in virtual space, and determines the field of view of the virtual space from the mapped position along the gaze direction. Content provider 130 may generate image data describing the determined field of view of the virtual space and send the image data to HMD 150 via communication interface 115. In some embodiments, content provider 130 generates metadata associated with the image (including motion vector information, depth information, edge information, object information, etc.) and sends the metadata, along with the image data, to HWD 150 via communication interface 115. Content provider 130 may compress and / or encode the data describing the image, and may send the compressed and / or encoded data to HMD 150. In some embodiments, content provider 130 periodically (e.g., every 11 ms) generates an image and provides the image to HMD 150.
[0055] Figure 2 This is a schematic diagram of an HMD 150 according to an example embodiment. In some embodiments, the HMD 150 includes a front rigid body 205 and a strap 210. The front rigid body 205 includes an electronic display 175. Figure 2 (Not shown in the image), sensors 155A, 155B, and 155C, and an image renderer 170. Sensor 155A may be an accelerometer, gyroscope, magnetometer, or another suitable type of sensor for detecting motion and / or position. Sensors 155B and 155C may be imaging sensors that acquire images for detecting physical objects, user gestures, hand shapes, user interactions, etc. In some embodiments, sensors 155B and 155C may be imaging sensors with fisheye lenses. HMD 150 may include additional components (e.g., GPS, wireless sensors, microphones, thermal sensors, etc.). In other embodiments, HMD 150 has... Figure 2 The configurations shown are different. For example, the image renderer 170, and / or sensors 155A, 155B, and 155C can be set to different configurations. Figure 2 The positions shown are different.
[0056] Figure 3 It is based on the exemplary embodiments of this disclosure. Figure 1A schematic diagram of the object transfer controller 180. In some embodiments, the object transfer controller 180 includes an object detector 310, a VR model generator 320, a VR model renderer 330, and a feedback controller 340. These components can work together to detect physical objects and render virtual models of those physical objects. The virtual model can be identified, activated, or generated, and can be rendered so that a user of the HMD 150 can locate the physical object while wearing the HMD 150. In some embodiments, these components can be implemented as hardware, software, or a combination of hardware and software. In some embodiments, these components are implemented as application-specific integrated circuits (ASICs) or field-programmable gate arrays (FPGAs). In some embodiments, these components are implemented as processors and non-transitory computer-readable media storing instructions that, when executed by the processor, cause the processor to perform the various processes disclosed herein. In some embodiments, the object transfer controller 180 includes a... Figure 3 The components shown may include more or fewer components, or may include components with... Figure 3 The components shown are different from those in the diagram. In some embodiments, the functionality of some components may be performed by the content provider 130 or a remote server, or in combination with the content provider or the remote server. For example, some functions of the object detector 310, the VR model generator 320, or both may be performed by the content provider 130 or a remote server. In some embodiments, the object migration controller 180 includes components that are more complex than those in the diagram. Figure 3 The components shown may include more or fewer components, or may include components with... Figure 3 The components shown are different components.
[0057] In some embodiments, object detector 310 is or includes a component that detects physical objects in physical space based on acquired images. In one application, object detector 310 detects input devices (e.g., keyboards or mice) in physical space by performing image processing on the acquired images. In one method, object detector 310 detects the shape, outline, and / or (e.g., the layout of keys or buttons formed by the keys themselves or the space between keys) or combinations thereof of a physical object in the acquired image, and determines the type of the physical object based on the detected shape, outline, and / or key or button layout. For example, object detector 310 determines whether the physical object is a keyboard or a mouse from the user's physical space perspective based on the detected shape, outline, and / or key or button layout of the physical object. Object detector 310 may also locate physical objects based on a thermal sensor that detects a heat map of the physical object. In one example, object detector 310 may detect that the physical object has a specific number of keys based on the outline of the physical object in the acquired image and determine that the physical object is a keyboard.
[0058] In one aspect, object detector 310 detects physical objects and presents a view of the physical space or a portion thereof to a user of HMD 150 via electronic display 175. For example, the electronic display 175 may present to the user images acquired by the imaging sensors of HMD 150 (e.g., sensors 155B and 155C) (e.g., images of the physical object and / or a portion of the user's image, either mixed with or not mixed with images of the virtual model and / or virtual space). Therefore, a user wearing HMD 150 can detect and / or locate physical objects in the physical space through HMD 150, for example, by image processing of one or more images acquired by the imaging sensors.
[0059] In some embodiments, the VR model generator 320 is or includes a component that generates, acquires, or identifies a virtual model of a detected physical object. In one method, the VR model generator 320 stores multiple candidate models of different manufacturing companies, brands, and / or product models. The VR model generator 320 can compare the detected shape, outline, and / or (e.g., the layout of buttons or keys formed by the buttons themselves or the space between buttons) with the shape, outline, and / or button or key layout of the multiple candidate models, and identify or determine a candidate model having a shape, outline, and / or button or key layout that matches or is closest to the shape, outline, and / or key layout of the detected physical object. The VR model generator 320 can detect or receive a product identifier of the physical object and identify or determine a candidate model corresponding to the detected product identifier. The VR model generator 320 can generate, determine, acquire, or select the determined candidate model as a virtual model of the physical object.
[0060] In some embodiments, the VR model renderer 330 is or includes a component that renders an image of a virtual model of a physical object. In one method, the VR model renderer 330 tracks a physical object in an acquired image and determines the position and orientation of that physical object relative to the user or HMD 150. In one aspect, the position and orientation of the physical object relative to the user or HMD 150 may change because the user may move around while wearing the HMD 150. The VR model renderer 330 can determine six degrees of freedom of the virtual model (e.g., forward / backward (swing), up / down (heavy), left / right (lateral), left / right (tilt), forward / backward (pitch), left / right (rotate)) such that the position and orientation of the seen or displayed virtual model relative to the user or HMD 150 in virtual space can correspond to the position and orientation of the physical object in the acquired image relative to the user or HMD 150. In a particular embodiment, the VR model renderer 330 can perform image processing on the acquired images to track certain features of the physical object (e.g., the four corners and / or edges of a keyboard, the shape formed by the space between the keys, and / or a pattern formed by a set of such shapes), and determine the position and orientation (pose) of the virtual model to match, correspond to, track, or adapt to the features of the physical object in the acquired images. The VR model renderer 330 can present the virtual model via the electronic display 175 based on the position and orientation of the virtual model. In one aspect, when the user is wearing the HMD 150, the VR model renderer 330 tracks the physical object and updates the position and orientation of the virtual model, wherein the electronic display 175 presents the user's virtual spatial viewpoint / field of view. Because the virtual model is presented in virtual space and acts as a spatial guide or reference, the user can easily locate and touch the physical object when they cannot actually see it due to wearing the HMD.
[0061] In some embodiments, the feedback controller 340 is or includes a component that generates spatial feedback on user interaction with a physical object. In one aspect, the feedback controller 340 detects and tracks the user's hand in the acquired image of the HMD 150 and visually provides spatial feedback regarding the user's movement and / or interaction with the physical object via the electronic display 175. The spatial feedback may be provided relative to a virtual model. In one example, the feedback controller 340 determines whether the user's hand is within (or close to) a predetermined distance from a physical object. If the user's hand is within the predetermined distance from the physical object (e.g., a keyboard), the feedback controller 340 may generate or render a virtual model of the user's hand and present that virtual model of the user's hand in virtual space via the electronic display 175. If the user's hand is not within the predetermined distance from the physical object (e.g., a keyboard), the feedback controller 340 may not present or render the virtual model of the user's hand via the electronic display 175. In some embodiments, the feedback controller 340 determines or generates an area (e.g., a rectangular area or other area) surrounding the virtual model in virtual space and may present that area via the electronic display 175. When a user's hand is within the area, a portion of the virtual model of the hand or a transparent image of the part of the hand within the area (e.g., mixed with other images or not mixed with other images) can be presented as spatial feedback.
[0062] In one example, the feedback controller 340 determines that a portion of the physical object is interacting with the user and indicates the corresponding portion of the virtual model is interacting with the user via an electronic display 175. For example, the feedback controller 340 determines that a key or button on the keyboard has been pressed by performing image processing on a captured image or by receiving an electrical signal corresponding to user input via the keyboard. The feedback controller 340 can highlight the corresponding key or button on the virtual model to indicate which key on the keyboard has been pressed. Therefore, the user can confirm the accuracy of input provided through the physical object even when they cannot actually see the physical object because they are wearing an HMD.
[0063] In practice, providing accurate, real-time spatial feedback on user interactions with physical objects can be challenging. Typically, images of physical objects acquired by the outward-facing imaging sensors of an HMD cannot be directly used to provide accurate spatial feedback to the user because these images are often occluded (e.g., blocked by the user's hand) or have low resolution, high noise, or low contrast characteristics. For example, Figure 4Images 401, 402, and 403, acquired by the imaging sensor of an HMD, are shown. These images are either occluded by a user's hand or have low resolution, high noise, or low contrast characteristics. One solution for providing spatial feedback to the user using these images is to determine the accurate pose of the physical object depicted in these images and render a virtual model in virtual space that accurately tracks the pose of that physical object. However, if the image is partially occluded or has low resolution, high noise, or low contrast characteristics, it is difficult to accurately determine the pose of any physical object to provide spatial feedback. The invention disclosed herein provides a method for accurately determining the pose of a physical object (e.g., a keyboard) by detecting prominent features that are detectable even in low-resolution, high-noise, or partially occluded images. Although many embodiments of this disclosure describe determining the pose of a physical keyboard, this disclosure corresponds to any physical object having prominent features similar to those described herein.
[0064] In one embodiment, the distinctive features of the keyboard can be identified based on visual characteristics formed by certain shapes existing between the keys or buttons on the keyboard, or by the corners of the keyboard. For example, Figure 5A An example image of a physical keyboard is shown, which has T-shaped features 501, 502, and 503, L-shaped feature 510, and X-shaped feature 520. Typically, the most common shape feature of a keyboard is T-shaped. For example, Figure 5A The keyboard shown has approximately 120 T-shaped sections and 3 X-shaped sections. The T-shaped section consists of a top portion (called the top cover) and a bottom portion (called the trunk). Figure 5B An example T-shaped template 590 for detecting the T-shape of the keyboard is shown, as well as a close-up view of the keyboard portion corresponding to one of the detected T-shapes 580. Figure 5B Indicators are also provided regarding the location of the T-shaped top portion 592 and the main portion 591. The T-shape can be formed by the visual feature of the space between two or more keys on the keyboard, with the space between two adjacent keys in the same row forming the main portion of the T-shape. For example, the T-shape can be formed by two keys in one row and a third key in an adjacent row (e.g., Figure 5B The T-shape shown is 580. A T-shape can be formed by two keys in one row and two additional keys in the adjacent row (e.g., Figure 5A The lower half of the X-shape 520 shown). As described below, an X-shape can be formed by combining two T-shapes in opposite orientations. If the individual keys in the two buttons are different but the rightmost or leftmost keys in adjacent rows, the T-shape can also be formed by these two keys without any other keys, with the top portion of the T-shape formed by the space to the right or left of these two keys (e.g., Figure 5A The T-shape 503 shown is also possible if both buttons are in the top or bottom row, and the T-shape can be formed by these two buttons without any other buttons. The top portion of the T-shape is formed by the space above or below the two buttons. Each detected T-shape can be associated with an orientation. For example, in... Figure 5A In this design, T-shape 501 is oriented such that its main stem faces downwards relative to its top cover, T-shape 502 is oriented such that its main stem faces upwards relative to its top cover, and T-shape 503 is oriented such that its main stem faces left relative to its top cover. Depending on the appearance of each shape feature, it can be associated with a positive or negative sign: if the space between the keys forming the shape feature is brighter than the surrounding keys, a positive sign can be assigned to that shape feature; conversely, if the space between the keys forming the shape feature is darker than the surrounding keys, a negative sign can be assigned to that shape feature. Alternatively, a sign opposite to the above symbols can be assigned to each shape feature. For example, in... Figure 5A In this design, because the buttons around T-shaped 503 appear darker than the space between the buttons, a positive sign can be assigned to T-shaped 503. In one embodiment, two T-shapeds with opposite orientations can be combined to form an X-shape, for example... Figure 5A The X-shape 520 is shown.
[0065] In one embodiment, the T-shaped feature of the keyboard is detected by evaluating the pixel intensity of an image depicting the keyboard. In another approach, a gradient- and variance-based method is used to identify the T-shape in an image depicting the keyboard. Figure 6A and Figure 6B This paper demonstrates a method for detecting T-shaped features based on gradient and variance approaches. Typically, in an image depicting a keyboard, one of the most prominent features corresponds to the difference in pixel intensity between the keys and the space between them. For example, in... Figure 4 In the images, image 401 shows keyboard keys contrasting with the darker spaces between them, while images 402 and 403 show keyboard keys contrasting with the brighter spaces between them. A gradient- and variance-based method utilizes this difference in pixel intensity to detect the keyboard's shape features. This method can be divided into two steps: the first step involves evaluating the gradient of pixel intensity in the top portion of the T-shape, and the second step involves evaluating the gradient of pixel intensity in the main body of the T-shape.
[0066] Figure 6AThe steps for evaluating the gradient of pixel intensity in a top cover portion according to one embodiment are illustrated. This step involves identifying four pixel groups in a pixel block: a top portion 610, which may correspond to the bottom edge portion of one or more specific buttons; two bottom portions 630, each of which may correspond to the top portion of a specific button; and a middle portion 620, located between the top portion 610 and the bottom portions 630. In one embodiment, identifying the four pixel groups involves identifying pixel groups with pixel intensities that are substantially uniform or have a variance less than a specific amount. For example, the variance of the pixel intensity of pixels in each of the top portion 610, the middle portion 620, and the bottom portion 630 may be less than a predetermined minimum variance. In addition to calculating the variance over the four regions, a gradient of pixel intensity between pixel groups is determined. The top gradient is determined based on the difference between the average pixel intensity value of the top portion 610 and the average pixel intensity value of the middle portion 620.
[0067]
[0068] The bottom gradient is determined based on the difference between the average pixel intensity value of the middle portion (620) and the average pixel intensity value of the bottom portion (630).
[0069]
[0070] The combined gradient is:
[0071]
[0072] The variance is:
[0073]
[0074] Figure 6A The horizontal response mentioned above combines the variance of different regions corresponding to the T-shaped top portion and the gradient between these regions to make it insensitive to variations in brightness, contrast, and image noise.
[0075]
[0076] If both the top and bottom gradients are greater than a specific minimum gradient value, the level response is calculated based on those gradients. Alternatively, if either the top or bottom gradient is determined to be less than the minimum gradient value, the level response is assigned a value of zero.
[0077]
[0078] If the sign of the combined gradient is incorrect, the level response is assigned a value of zero:
[0079]
[0080] Figure 6B The steps for evaluating pixel intensity gradients of a backbone portion according to one embodiment are illustrated. This step involves identifying three pixel groups within a pixel block: a left portion 660, which may correspond to the right edge of a specific key on a keyboard; a right portion 680, which may correspond to the left edge of a specific key on a keyboard; and a middle portion 670, located between the left portion 660 and the right portion 680. In one embodiment, identifying the three pixel groups involves identifying pixel groups with pixel intensities that are substantially uniform or have a variance less than a specific amount. For example, the variance of the pixel intensity of pixels in each of the left portion 660, the middle portion 670, and the right portion 680 may be less than a predetermined minimum variance. In addition to calculating the variances over the three regions, a gradient of pixel intensity between the pixel groups is determined. The left gradient is determined based on the difference between the average pixel intensity value of the left portion 660 and the average pixel intensity value of the middle portion 670.
[0081]
[0082] The right gradient is determined based on the difference between the average pixel intensity value of the middle portion (670) and the average pixel intensity value of the right portion (680).
[0083]
[0084] The combined gradient is:
[0085]
[0086] The variance is:
[0087]
[0088] Figure 6B The vertical response mentioned above combines the variance of different regions corresponding to the main body of the T-shape with the gradients between these regions to make it insensitive to variations in brightness, contrast, and image noise.
[0089]
[0090] If both the left and right gradients are greater than a specific minimum gradient value, the vertical response is calculated based on those left and right gradients. Alternatively, if either the left or right gradient is determined to be less than the minimum gradient value, the vertical response is assigned a value of zero.
[0091]
[0092] If the sign of the combined gradient is incorrect, the vertical response is assigned a value of zero:
[0093]
[0094] In one embodiment, a T-shape is identified by combining the horizontal and vertical responses (e.g., by summing or multiplying) and determining whether the combined response value is greater than a predetermined minimum response value. If the combined response value is less than the predetermined minimum response value, the pixel block is determined not to contain a T-shape. In embodiments where multiple adjacent combined responses are detected, the combined response with the strongest response value or the strongest absolute response value can be selected for the group. In some embodiments, reference is made to... Figure 6A The top and bottom portions can be configured such that the size and shape of the top portion 610 substantially match the size and shape of the bottom portion 630. This ensures that the horizontal response calculated either from a top-to-bottom direction or from a bottom-to-top direction is consistent, thereby further allowing the horizontal response to be combined with a vertical response corresponding to a downward-facing T-shape, or with a vertical response corresponding to an upward-facing T-shape. For example, the top portion 610 can be divided into two parts corresponding to the two bottom portions 630. Alternatively, the two bottom portions 630 can be combined into a single part instead of being separated by the main body of the T-shaped template. In some embodiments, reference... Figure 6B The height of the middle section 670 can be adjusted to roughly match the heights of the left section 660 and the right section 680. This ensures that the vertical response is consistent whether calculated from right to left or from left to right, thereby further enabling the vertical response to be combined with the horizontal response corresponding to a downward or upward T-shape.
[0095] In one embodiment, the VR model generator 320 generates, acquires, or identifies a virtual model of a keyboard based on shape features and / or corner features. Figure 5A Examples of shape features (e.g., 501, 502, 503, 510, and 520) and corner features (e.g., 550) are shown. In one embodiment, the virtual model of the physical keyboard is generated based on features detected on the physical keyboard, such as corner features and shape features, including the orientation and symbol of the shape features. For example, Figure 7An example keyboard 710 with some of the detected shape features is shown, upon which a virtual model 720 can be generated. In one embodiment, the virtual model can be a three-dimensional virtual object, and a three-dimensional position on the virtual model can be assigned to each detected feature. In some embodiments, the VR model generator 320 can generate a simplified version of the virtual model as a two-dimensional virtual object and assign a two-dimensional position to each of the detected features. The simplified virtual model can enable the matching process described below to be performed faster and more efficiently. In one embodiment, the virtual model can be generated to include information about the topology of the keyboard. The topology of the keyboard includes information about, for example, the number of rows of the keyboard, the identifier of the row to which the shape features belong, the size of the keys, the distance between the keyboard rows, etc.
[0096] Figure 8A and Figure 8B An example keyboard is shown, demonstrating a method for determining the orientation of a physical keyboard based on its features. The method can begin by acquiring images from a viewpoint corresponding to the imaging sensor of the HMD. Figure 8A An example image is shown. In some embodiments where the image is acquired by an imaging sensor with a fisheye lens, the image may appear distorted or warped. In such embodiments, the image can be processed to eliminate the distortion. For example, Figure 8A Image 810 is shown, which has been processed to eliminate distortion caused by a fisheye lens. In some embodiments, if the keyboard depicted in the image does not appear rectangular (e.g., it appears trapezoidal), the keyboard can be corrected to appear rectangular. For example, Figure 8B An image 840 of a keyboard that has been corrected to a rectangle is shown. In some embodiments, a stereo camera can be used to acquire two images of the keyboard, one image acquired through each of the two lenses of the stereo camera. In such an embodiment, shape features can be detected in both images separately, thereby enabling the keyboard's pose to be calculated based on the shape features detected in both images.
[0097] In one embodiment, edge-based detection can be used to determine whether any keyboard is depicted in the image (e.g., in the initial detection phase). Edge-based detection differs from shape-based detection because it looks for the outer edges of the keyboard (e.g., the four sides of the keyboard) rather than the keyboard's shape features (e.g., T-shaped features). While edge-based detection is generally less accurate in determining the keyboard's pose, it can still determine a coarse pose. Furthermore, edge-based detection may be more suitable for the initial detection phase, where the goal is to determine whether any keyboard is depicted in the image, given its lower computational cost compared to shape-based detection. In some embodiments, shape-based detection can be used instead of edge-based detection for the initial keyboard detection phase. For example, when the keyboard is partially occluded (e.g., by a user's hand) or keyboard features are otherwise not detected, shape-based detection may be more suitable for detecting the keyboard because the number of features detected using shape-based detection is much greater than the number detected using edge-based detection (e.g., a keyboard typically has over 100 shape features compared to 4 edge features). In other words, if some edges of the keyboard are occluded, edge-based detection might fail to detect the keyboard because a typical keyboard only has four edges. However, shape-based detection can detect the keyboard by relying on the remaining unoccluded shape features, even if some of the shape features are occluded. In some embodiments, to determine whether any keyboard is depicted in the image (e.g., in the initial detection phase), deep learning methods can be applied to the image to detect the approximate locations of the four corners of the keyboard.
[0098] In one embodiment, after recognizing a keyboard in an image and determining its coarse pose (e.g., based on edge-based detection), shape-based detection can be used to determine the keyboard's precise pose. In one embodiment, the image of the keyboard can be divided into multiple pixel blocks, and then each of these pixel blocks can be compared with a shape pattern template (e.g., ...). Figure 6A and Figure 6B The shape features of the keyboard are detected by comparing it with the T-shaped, X-shaped, or L-shaped templates shown. For example, Figure 8BA keyboard 840 is shown, having shape features detected on the keyboard. In one embodiment, the detected shape features can be aggregated into multiple groups based on multiple keyboard rows. Aggregation of these features provides unique feature patterns, which can be compared with feature patterns of virtual models stored in a database to find a virtual model that matches the keyboard. Once a matching virtual model is found, the accurate orientation of the keyboard can be determined by projecting the virtual model in virtual space toward the viewpoint of the camera that captured the image, and adjusting the orientation of the virtual model (e.g., by minimizing projection errors) until the feature correspondences of the keyboard and the virtual model match each other. In one embodiment, if a virtual model of the physical keyboard is not present or cannot be found in the database, a virtual model of the keyboard can be generated based on the detected features of the physical keyboard, and the virtual model of the keyboard can be stored in the database.
[0099] In one embodiment, after determining the accurate orientation of the keyboard, the keyboard can be continuously tracked to determine if it is still in the expected position. If the keyboard is not in the expected position, its orientation can be determined again by performing edge-based detection or shape-based detection. In one embodiment, to minimize the computational cost associated with the tracking process, edge-based detection can be used first to determine if the keyboard is in the expected position because it is less computationally expensive than shape-based detection. If edge-based detection fails to detect the keyboard, shape-based detection can be used instead. In one embodiment, the tracking process may include an active phase and an idle phase, in which the object detector 310 actively tracks the keyboard, and in the idle phase, it does not track the keyboard. The object detector 310 may only implement the active phase intermittently to minimize computational cost and improve the battery life of the HMD. In one embodiment, the object detector 310 may adjust the duration of the active and idle phases based on various conditions. For example, object detector 310 may initially implement the active phase for a longer period than the idle phase, but if the keyboard remains stationary for a considerable period, object detector 310 may adjust its implementation so that the idle phase is implemented for a longer period than the active phase. In some embodiments, object detector 310 may manually implement the active phase when there is an indication that the keyboard may be moved by the user (e.g., when the user's hand is detected near the left and right sides of the keyboard, or when the user presses one or more specific keys).
[0100] Figure 9An example method 900 for determining keyboard pose is shown. The method may begin at step 901 by acquiring an image from a camera viewpoint depicting a physical keyboard. At step 902, the method may continue by detecting one or more shape features of the physical keyboard depicted in the image by comparing pixels of the image with a predetermined shape pattern representing visual features of the space between keyboard keys. At step 903, the method may continue by accessing predetermined shape features of a keyboard model associated with the physical keyboard. At step 904, the method may continue by determining the pose of the physical keyboard based on a comparison between: (1) one or more detected shape features of the physical keyboard; and (2) the projection of the predetermined shape features of the keyboard model toward the camera viewpoint. Certain embodiments may be repeated where appropriate. Figure 9 One or more steps in the method. Although this disclosure will Figure 9 The specific steps of the method are described and shown as being performed in a particular order, but this disclosure contemplates performing them in any suitable order. Figure 9 The method includes any suitable steps. Furthermore, although this disclosure describes and illustrates example methods for determining the keyboard's orientation, this disclosure contemplates any suitable methods for determining the keyboard's orientation, including any suitable steps, where appropriate, such steps may include... Figure 9 The method may include all steps, some steps, or may exclude them. Figure 9 Any steps of the method. Furthermore, although this disclosure describes and illustrates the implementation of... Figure 9 The method may refer to a specific component, device, or system of a particular step, but this disclosure is contemplated for the execution of... Figure 9 Any suitable step of the method, any suitable component, device or system, or any suitable combination thereof.
[0101] Figure 10 An example network environment 1000 associated with a social networking system is shown. Network environment 1000 includes a client system 1030, a social networking system 1060, and a third-party system 1070, which are connected to each other via network 1010. Although Figure 10A specific arrangement of client system 1030, social networking system 1060, third-party system 1070, and network 1010 is shown, but this disclosure contemplates any suitable arrangement of client system 1030, social networking system 1060, third-party system 1070, and network 1010. By way of example and not limitation, two or more of client system 1030, social networking system 1060, and third-party system 1070 may bypass network 1010 and connect directly to each other. As another example, two or more of client system 1030, social networking system 1060, and third-party system 1070 may be physically or logically located in the same place as each other, wholly or partially. For example, an AR / VR headset (corresponding to client system 1030) may connect to a local computer or mobile computing device (corresponding to third-party system 1070) via short-range wireless communication (e.g., Bluetooth). Furthermore, although... Figure 10 A specific number of client systems 1030, social networking systems 1060, third-party systems 1070, and networks 1010 are shown, but this disclosure contemplates any suitable number of client systems 1030, social networking systems 1060, third-party systems 1070, and networks 1010. As an example and not a limitation, network environment 1000 may include multiple client systems 1030, multiple social networking systems 1060, multiple third-party systems 1070, and multiple networks 1010.
[0102] This disclosure considers any suitable network 1010. By way of example and not limitation, one or more portions of network 1010 may include short-range wireless networks (e.g., Bluetooth, Zigbee, etc.), ad hoc networks, intranets, extranets, virtual private networks (VPNs), local area networks (LANs), wireless LANs (WLANs), wide area networks (WANs), wireless wide area networks (WWANs), metropolitan area networks (MANs), a portion of the Internet, a portion of the Public Switched Telephone Network (PSTN), a cellular telephone network, or a combination of two or more of these networks. Network 1010 may include one or more networks 1010.
[0103] Link 1050 enables client system 1030, social networking system 1060, and third-party system 1070 to connect to network 1010 or to each other. This disclosure contemplates any suitable link 1050. In certain embodiments, one or more links 1050 may include one or more wired links (e.g., Digital Subscriber Line (DSL) or Data Over Cable Service Interface Specification (DOCSIS)), one or more wireless links (e.g., Wi-Fi, Worldwide Interoperability for Microwave Access (WiMAX), Bluetooth), or one or more optical links (e.g., Synchronous Optical Network (SONET) or Synchronous Digital Hierarchy (SDH)). In a particular embodiment, one or more links 1050 each include an ad hoc network, an intranet, an extranet, a VPN, a LAN, a WLAN, a WAN, a WWAN, a MAN, a portion of the Internet, a portion of the PSTN, a cellular-based network, a satellite-based network, another link 1050, or a combination of two or more such links 1050. In the entire network environment 1000, the multiple links 1050 need not all be identical. In one or more aspects, one or more first links may differ from one or more second links.
[0104] In a particular embodiment, client system 1030 may be an electronic device that includes hardware, software, or embedded logic components, or a combination of two or more such components, and is capable of performing suitable functions implemented or supported by client system 1030. By way of example and not limitation, client system 1030 may include a computer system, such as a VR / AR headset, desktop computer, laptop or notebook computer, netbook, tablet computer, e-book reader, GPS device, camera, personal digital assistant (PDA), handheld electronic device, cellular phone, smartphone, augmented reality / virtual reality device, other suitable electronic device, or any suitable combination thereof. This disclosure contemplates any suitable client system 1030. Client system 1030 enables network users at client system 1030 to access network 1010. Client system 1030 enables its users to communicate with other users at other client systems 1030.
[0105] In a particular embodiment, the social networking system 1060 may be a network-addressable computing system capable of hosting online social networks. The social networking system 1060 may generate, store, receive, and transmit social networking data, such as user profile data, concept profile data, social graph information, or other suitable data related to the online social network. The social networking system 1060 may be accessed directly by other components in the network environment 1000 or via network 1010. By way of example and not limitation, the client system 1030 may use a web browser or a native application associated with the social networking system 1060 (e.g., a mobile social networking application, a messaging application, another suitable application, or any combination thereof) to access the social networking system 1060 directly or via network 1010. In a particular embodiment, the social networking system 1060 may include one or more servers 1062. Each server 1062 may be a single server or a distributed server spanning multiple computers or multiple data centers. Server 1062 can be of various types, such as, but not limited to, web server, news server, mail server, message server, advertising server, file server, application server, exchange server, database server, proxy server, another server suitable for performing the functions or processes described herein, or any combination thereof. In a particular embodiment, each server 1062 may include hardware, software, or embedded logic components, or combinations of two or more such components, for performing suitable functions implemented or supported by server 1062. In a particular embodiment, social networking system 1060 may include one or more data storage devices 1064. Data storage devices 1064 can be used to store various types of information. In a particular embodiment, the information stored in data storage devices 1064 may be organized according to a particular data structure. In a particular embodiment, each data storage device 1064 may be a relational database, columnar database, correlation database, or other suitable database. Although this disclosure describes or illustrates specific types of databases, this disclosure contemplates any suitable type of database. A particular embodiment may provide an interface that enables client system 1030, social network system 1060, or third-party system 1070 to manage, retrieve, modify, add, or delete information stored in data storage 1064.
[0106] In a particular embodiment, the social network system 1060 may store one or more social graphs in one or more data storage devices 1064. In a particular embodiment, the social graph may include multiple nodes—which may include multiple user nodes (each user node corresponds to a specific user) or multiple concept nodes (each concept node corresponds to a specific concept)—and multiple edges connecting these nodes. The social network system 1060 may provide users of the online social network with the ability to communicate and interact with other users. In a particular embodiment, a user can join an online social network through the social network system 1060 and subsequently add connections (e.g., relationships) to some other users in the social network system 1060 that they wish to connect with. As used herein, the term "friend" may refer to any other user with whom a user has already formed a connection, association, or relationship through the social network system 1060.
[0107] In a particular embodiment, the social networking system 1060 may provide users with the ability to take action on various types of items or objects supported by the social networking system 1060. By way of example, and not limitation, these items and objects may include groups or social networks to which the user of the social networking system 1060 may belong, events or calendar entries that the user may be interested in, computer-based applications that the user may use, transactions that allow the user to buy or sell items through services, interactions with advertisements that the user may perform, or other suitable items or objects. Users may interact with anything that can be represented in the social networking system 1060 or through an external system 1070, separate from and coupled to the social networking system 1060 via network 1010.
[0108] In a particular embodiment, the social networking system 1060 may be able to connect various entities. By way of example and not limitation, the social networking system 1060 may enable users to interact with each other and receive content from third-party systems 1070 or other entities, or may enable users to interact with these entities through application programming interfaces (APIs) or other communication channels.
[0109] In a particular embodiment, the third-party system 1070 may include a local computing device communicatively coupled to the client system 1030. For example, if the client system 1030 is an AR / VR headset, the third-party system 1070 may be a local laptop configured to perform necessary graphics rendering and provide the rendered results to the AR / VR headset (corresponding to the client system 1030) for subsequent processing and / or display. In a particular embodiment, the third-party system 1070 may execute software associated with the client system 1030 (e.g., a rendering engine). The third-party system 1070 may generate a sample dataset with sparse pixel information of video frames and send this sparse data to the client system 1030. The client system 1030 may then generate frames reconstructed from the sample dataset.
[0110] In certain embodiments, the third-party system 1070 may also include one or more types of servers, one or more data storage devices, one or more interfaces (including, but not limited to, APIs), one or more web services, one or more content sources, one or more networks, or any other suitable component (e.g., with which the server can communicate). The third-party system 1070 may be operated by an entity different from the entity operating the social networking system 1060. However, in certain embodiments, the social networking system 1060 and the third-party system 1070 may operate collaboratively to provide social networking services to users of the social networking system 1060 or to the third-party system 1070. In this sense, the social networking system 1060 may provide a platform or backbone that other systems (e.g., the third-party system 1070) can use to provide social networking services and functionality to users on the Internet.
[0111] In a particular embodiment, third-party system 1070 may include a third-party content object provider (e.g., including the sparse sample dataset described herein). The third-party content object provider may include one or more content object sources that can be transmitted to client system 1030. As an example, and not a limitation, content objects may include information about things or activities of interest to the user, such as movie showtimes, movie reviews, restaurant reviews, restaurant menus, product information and reviews, or other suitable information. As another example, and not a limitation, content objects may include incentivized content objects, such as coupons, discount tickets, gift certificates, or other suitable incentives.
[0112] In a particular embodiment, the social networking system 1060 also includes user-generated content objects that can enhance user interaction with the social networking system 1060. User-generated content can include any content that a user can add, upload, send, or "post" to the social networking system 1060. As an example, and not a limitation, a user transmits a post from client system 1030 to social networking system 1060. A post can include data such as status updates or other text data, location information, photos, videos, links, music, or other similar data or media. Content can also be added to the social networking system 1060 by a third party via a "communication channel" (e.g., a news feed or stream).
[0113] In certain embodiments, the social networking system 1060 may include various servers, subsystems, programs, modules, logs, and data storage. In certain embodiments, the social networking system 1060 may include one or more of the following: a web server, an action logger, an API request server, a relevance and ranking engine, a content object classifier, a notification controller, action logs, third-party content object exposure logs, an inference module, an authorization / privacy server, a search module, an ad targeting module, a user interface module, a user profile storage, a contact storage, a third-party content storage, or a location storage. The social networking system 1060 may also include suitable components, such as network interfaces, security mechanisms, load balancers, failover servers, management and network operations consoles, other suitable components, or any suitable combination thereof. In certain embodiments, the social networking system 1060 may include one or more user profile storage devices for storing user profiles. User profiles may include, for example, biometric information, demographic information, behavioral information, social information, or other types of descriptive information (e.g., work experience, educational history, hobbies or preferences, interests, kinship, or location). Interest information may include interests associated with one or more categories. Categories may be general or specific. As an example, and not a limitation, if a user “likes” items related to a shoe brand, the category could be that brand, or the generic categories “shoes” or “clothing.” A contact store can be used to store contact information about users. This contact information can indicate users who have similar or shared work experience, group memberships, hobbies, educational history, or who are associated with or share common attributes in any way. Contact information can also include user-defined connections (both internal and external) between different users and content. A web server can be used to link the social networking system 1060 to one or more client systems 1030 or one or more third-party systems 1070 via network 1010. The web server can include a mail server or other messaging functionality for receiving and routing messages between the social networking system 1060 and one or more client systems 1030. An API request server can allow third-party systems 1070 to access information from the social networking system 1060 by calling one or more APIs. An action logger can be used to receive information from the web server regarding user actions of opening or closing the social networking system 1060. Combined with the action log, a log of third-party content objects that a user exposes to can be maintained. The notification controller can provide information about content objects to the client system 1030. This information can be pushed to the client system 1030 as a notification, or it can be retrieved from the client system 1030 in response to a request received from the client system 1030.An authorization server can be used to enforce one or more privacy settings for users of the social networking system 1060. A user's privacy settings determine how specific information associated with that user can be shared. The authorization server can allow users, for example, by setting appropriate privacy settings, to choose whether or not their actions are recorded by the social networking system 1060 or shared with other systems (e.g., third-party system 1070). A third-party content object storage can be used to store content objects received from third parties (e.g., third-party system 1070). A location storage can be used to store location information received from a client system 1030 associated with the user. An advertising targeting module can combine social information, current time, location information, or other suitable information to provide relevant advertisements to the user in the form of notifications.
[0114] Figure 11 An example computer system 1100 is illustrated. In a particular embodiment, one or more computer systems 1100 perform one or more steps of one or more methods described or illustrated herein. In a particular embodiment, one or more computer systems 1100 provide the functionality described or illustrated herein. In a particular embodiment, software running on one or more computer systems 1100 performs one or more steps of one or more methods described or illustrated herein, or provides the functionality described or illustrated herein. The particular embodiments include one or more portions of one or more computer systems 1100. Throughout this document, references to computer systems may include, where appropriate, computing devices, and vice versa. Furthermore, references to computer systems may encompass one or more computer systems where appropriate.
[0115] This disclosure contemplates any suitable number of computer systems 1100. This disclosure contemplates computer systems 1100 employing any suitable physical form. By way of example and not limitation, computer system 1100 may be an embedded computer system, a system-on-chip (SOC), a single-board computer system (SBC) (e.g., a computer-on-module (COM) or system-on-module (SOM)), a desktop computer system, a laptop or notebook computer system, an interactive self-service machine, a mainframe, a mesh architecture of computer systems, a mobile phone, a personal digital assistant (PDA), a server, a tablet computer system, an augmented / virtual reality device, or a combination of two or more of these computer systems. Where appropriate, computer system 1100 may include one or more computer systems 1100; may be single or distributed; span multiple locations; span multiple machines; span multiple data centers; or reside in a cloud (which may include one or more cloud components in one or more networks). Where appropriate, one or more computer systems 1100 can perform one or more steps of the methods described or illustrated herein without significant space or time constraints. By way of example and not limitation, one or more computer systems 1100 can perform one or more steps of the methods described or illustrated herein in real time or in batch mode. Where appropriate, one or more computer systems 1100 can perform one or more steps of the methods described or illustrated herein at different times or in different locations.
[0116] In a particular embodiment, computer system 1100 includes a processor 1102, memory 1104, storage 1106, input / output (I / O) interface 1108, communication interface 1110, and bus 1112. Although this disclosure describes and illustrates a particular computer system having a particular number of components in a particular arrangement, this disclosure contemplates any suitable computer system having any suitable number of components in any suitable arrangement.
[0117] In a particular embodiment, processor 1102 includes hardware for executing instructions (e.g., those that constitute a computer program). By way of example, and not limitation, to execute instructions, processor 1102 may retrieve (or read) multiple instructions from internal registers, internal cache, memory 1104, or memory 1106; decode and execute these instructions; and subsequently write one or more results to internal registers, internal cache, memory 1104, or memory 1106. In a particular embodiment, processor 1102 may include one or more internal caches for data, instructions, or addresses. Where appropriate, this disclosure contemplates processor 1102 including any suitable number of suitable internal caches. By way of example, and not limitation, processor 1102 may include one or more instruction caches, one or more data caches, and one or more page table caches (translation lookaside buffers, TLBs). Instructions in the instruction cache may be copies of instructions in memory 1104 or memory 1106, and the instruction cache may accelerate the retrieval of these instructions by processor 1102. The data in the data cache may be a copy of the data in memory 1104 or memory 1106 for operation by instructions executed at processor 1102; it may be the result of a previous instruction executed at processor 1102 for access by subsequent instructions executed at processor 1102 or for writing to memory 1104 or memory 1106; or it may be other suitable data. The data cache can accelerate read or write operations of processor 1102. The TLB can accelerate virtual address translation of processor 1102. In a particular embodiment, processor 1102 may include one or more internal registers for data, instructions, or addresses. Where appropriate, this disclosure contemplates processor 1102 including any suitable number of suitable internal registers. Where appropriate, processor 1102 may include one or more arithmetic logic units (ALUs); may be a multi-core processor; or may include one or more processors 1102. Although this disclosure describes and illustrates specific processors, this disclosure contemplates any suitable processor.
[0118] In a particular example, memory 1104 includes main memory used to store instructions for execution by processor 1102 or data for operation by processor 1102. By way of example and not limitation, computer system 1100 may load instructions from memory 1106 or another source (e.g., another computer system 1100) into memory 1104. Processor 1102 may then load these instructions from memory 1104 into internal registers or internal cache. To execute these instructions, processor 1102 may retrieve and decode these instructions from internal registers or internal cache. During or after the execution of these instructions, processor 1102 may write one or more results (which may be intermediate or final results) to internal registers or internal cache. Processor 1102 may then write one or more of these results into memory 1104. In a particular embodiment, processor 1102 executes only instructions in one or more internal registers, internal cache, or memory 1104 (not memory 1106 or elsewhere), and operates only on data in one or more internal registers, internal cache, or memory 1104 (not memory 1106 or elsewhere). One or more memory buses (each of which may include an address bus and a data bus) couple processor 1102 to memory 1104. As described below, bus 1112 may include one or more memory buses. In a particular embodiment, one or more memory management units (MMUs) are located between processor 1102 and memory 1104 and facilitate access to memory 1104 requested by processor 1102. In a particular embodiment, memory 1104 includes random access memory (RAM). Where appropriate, the RAM is volatile memory. Where appropriate, the RAM may be dynamic RAM (DRAM) or static RAM (SRAM). Furthermore, where appropriate, the RAM may be single-port or multi-port RAM. This disclosure considers any suitable RAM. Where appropriate, memory 1104 includes one or more memory modules. Although this disclosure describes and illustrates specific memory, it considers any suitable memory.
[0119] In a particular embodiment, memory 1106 includes a large-capacity storage for data or instructions. By way of example and not limitation, memory 1106 may include a hard disk drive (HDD), a floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or a Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 1106 may include removable or non-removable (or fixed) media. Where appropriate, memory 1106 may be located internally or externally to computer system 1100. In a particular embodiment, memory 1106 is a non-volatile solid-state memory. In a particular embodiment, memory 1106 includes read-only memory (ROM). Where appropriate, the ROM may be a mask-programmed ROM, a programmable ROM (PROM), an erasable PROM (EPROM), an electrically erasable PROM (EEPROM), an electrically alterable ROM (EAROM), or flash memory, or a combination of two or more of these. This disclosure contemplates a large-capacity memory 1106 in any suitable physical form. Where appropriate, memory 1106 may include one or more memory control units that facilitate communication between processor 1102 and memory 1106. Where appropriate, memory 1106 may include one or more memory units. Although this disclosure describes and illustrates specific memories, it contemplates any suitable memory.
[0120] In a particular embodiment, I / O interface 1108 includes hardware, software, or both that provide one or more interfaces for communication between computer system 1100 and one or more I / O devices. Where appropriate, computer system 1100 may include one or more of these I / O devices. These one or more I / O devices enable communication between a person and computer system 1100. By way of example and not limitation, I / O devices may include a keyboard, keypad, microphone, monitor, mouse, printer, scanner, speaker, still camera, stylus, tablet computer, touchscreen, trackball, camera, another suitable I / O device, or a combination of two or more of these. I / O devices may include one or more sensors. This disclosure contemplates any suitable I / O device and any suitable I / O interface 1108 for such I / O devices. Where appropriate, I / O interface 1108 may include one or more device or software drivers that enable processor 1102 to drive one or more of these I / O devices. Where appropriate, I / O interface 1108 may include one or more I / O interfaces 1108. Although this disclosure describes and illustrates specific I / O interfaces, this disclosure considers any suitable I / O interface.
[0121] In a particular embodiment, communication interface 1110 includes hardware, software, or both that provides one or more interfaces for communication (e.g., packet-based communication) between computer system 1100 and one or more other computer systems 1100 or one or more networks. By way of example and not limitation, communication interface 1110 may include a network interface controller (NIC) or network adapter for communicating with Ethernet or other wired networks, or may include a wireless NIC (WNIC) or wireless adapter for communicating with wireless networks (e.g., Wi-Fi networks). This disclosure contemplates any suitable network and any suitable communication interface 1110 for that network. By way of example and not limitation, computer system 1100 may communicate with one or more portions of ad hoc networks, personal area networks (PANs), local area networks (LANs), wide area networks (WANs), metropolitan area networks (MANs), or the Internet, or a combination of two or more of these. One or more portions of one or more of these networks may be wired or wireless. As an example, computer system 1100 may communicate with networks such as wireless PAN (WPAN) (e.g., Bluetooth WPAN), Wi-Fi networks, Wi-Fi networks, cellular telephone networks (e.g., Global System for Mobile Communication (GSM) networks), or other suitable wireless networks, or combinations of two or more of these networks. Where appropriate, computer system 1100 may include any suitable communication interface 1110 for any of these networks. Where appropriate, communication interface 1110 may include one or more communication interfaces 1110. Although specific communication interfaces are described and illustrated in this disclosure, any suitable communication interface is contemplated in this disclosure.
[0122] In a particular embodiment, bus 1112 includes hardware, software, or both that couples multiple components of computer system 1100 to each other. By way of example and not limitation, bus 1112 may include: an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a front-side bus (FSB), a HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth interconnect, a low-pin-count (LPC) bus, a memory bus, a Micro Channel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCIe) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association local (VLB) bus, or another suitable bus, or a combination of two or more of these. Where appropriate, bus 1112 may include one or more buses 1112. Although this disclosure describes and illustrates a particular bus, this disclosure considers any suitable bus or interconnect.
[0123] In this document, where appropriate, one or more computer-readable non-transitory storage media may include: one or more semiconductor-based integrated circuits (ICs) or other integrated circuits (e.g., field-programmable gate arrays (FPGAs) or application-specific ICs (ASICs)), hard disk drives (HDDs), hybrid hard drives (HHDs), optical disc drives (ODDs), magneto-optical disk drives (ODDs), floppy disk drives (FDDs), magnetic tape, solid-state drives (SSDs), RAM drives, secure digital cards or drives, any other suitable computer-readable non-transitory storage media, or any suitable combination of two or more of these. Where appropriate, computer-readable non-transitory storage media may be volatile, non-volatile, or a combination of volatile and non-volatile.
[0124] In this document, unless otherwise expressly stated or the context otherwise requires, "or" is open-ended rather than exclusive. Therefore, in this document, unless otherwise expressly stated or the context otherwise requires, "A or B" means "A, B, or both." Furthermore, unless otherwise expressly stated or the context otherwise requires, "and" is both common and separate. Therefore, in this document, unless otherwise expressly stated or the context otherwise requires, "A and B" means "A and B, commonly or separately."
[0125] The scope of this disclosure includes all changes, substitutions, variations, transformations, and modifications to the exemplary embodiments described or shown herein, which will be understood by those skilled in the art. The scope of this disclosure is not limited to the exemplary embodiments described or shown herein. Furthermore, although this disclosure describes and illustrates various embodiments herein as including specific components, elements, features, functions, operations, or steps, any embodiment in these embodiments may include any combination or arrangement of any components, elements, features, functions, operations, or steps described or shown anywhere herein as will be understood by those skilled in the art. Moreover, references in the appended claims to an apparatus or system, or a component in an apparatus or system, that is suitable for, arranged to, capable of, configured to, enable, operable, or operable to perform a particular function include that apparatus, system, component, whether or not the apparatus, system, component, or the particular function is activated, turned on, or unlocked, provided that the apparatus, system, or component is so suitable for, arranged to, capable of, configured to, enable, operable, or operable. Furthermore, although this disclosure describes or illustrates specific embodiments to provide particular advantages, specific embodiments may not provide these advantages, or may provide some or all of these advantages.
Claims
1. A method for determining keyboard posture, comprising: From computing devices: Images are captured from a camera viewpoint, and the images depict the physical keyboard; Identify the physical keyboard in the captured image; One or more shape features of the physical keyboard depicted in the image are detected by comparing the pixels of the image with a predetermined shape template, the predetermined shape template representing the visual features of the space between the keyboard keys, wherein each of the one or more shape features of the physical keyboard corresponds to a T-feature, an X-feature, or an L-feature formed by the space between the keyboard keys on the physical keyboard. A keyboard model associated with the physical keyboard is determined from the database by comparing the detected one or more shape features of the physical keyboard with predetermined shape features of the plurality of keyboard models in a database that includes a plurality of keyboard models. Access a predetermined shape feature of the keyboard model associated with the physical keyboard; and The orientation of the physical keyboard is determined based on a comparison between: one or more shape features of the detected physical keyboard; and the projection of the predetermined shape features of the keyboard model toward the camera viewpoint.
2. The method according to claim 1, wherein, The predetermined shape template is a T-shaped template with a main body and a top cover, wherein the main body represents the visual feature of the space between two keyboard keys in the first keyboard row.
3. The method according to claim 2, wherein, The top cover portion of the T-shaped template represents the visual characteristics of the space between the following: the portion of the first keyboard row corresponding to the two keyboard keys in the first keyboard row; and one or more keyboard keys in the second keyboard row adjacent to the first keyboard row.
4. The method according to claim 2, wherein, The first keyboard row is: The top row of the keyboard, the top cover portion of the T-shaped template represents the visual characteristics of the space above the two keyboard keys in the first keyboard row; or The bottom row of the keyboard, and the top cover portion of the T-shaped template, represent the visual characteristics of the space below the two keyboard keys in the first keyboard row.
5. The method according to claim 1, wherein, The predetermined shape template is a T-shaped template with a main body and a top cover. The main body represents the visual characteristics of the space between the first key in the first keyboard row and the second key in the second keyboard row. The second keyboard row is adjacent to the first keyboard row, wherein: The first keyboard key is the rightmost key in the first keyboard row, and the second keyboard key is the rightmost key in the second keyboard row. The top cover portion of the T-shaped template represents the visual characteristics of the space to the right of the first and second keyboard keys; or The first keyboard key is the leftmost key in the first keyboard row, and the second keyboard key is the leftmost key in the second keyboard row. The top cover portion of the T-shaped template represents the visual characteristics of the space to the left of the first keyboard key and the second keyboard key.
6. The method according to any one of claims 1 to 5, wherein, Comparing the pixels of the image with the predetermined shape template includes: The image is divided into multiple pixel blocks, and each pixel block includes multiple pixels; For each of the plurality of pixel blocks: The visual features of the pixel block are determined based on the pixel intensities associated with the plurality of pixels in the pixel block; and The visual features of the pixel block are compared with a predetermined shape template representing the visual features of the space between the keyboard keys; and The one or more shape features of the physical keyboard depicted in the image are detected by comparing the visual features of the pixel blocks of the image with the predetermined shape template representing the visual features of the space between the keyboard keys.
7. The method according to claim 6, wherein, Determining the visual features of the pixel block based on the pixel intensities associated with the plurality of pixels in the pixel block includes: The pixel intensity associated with at least a portion of the pixel block and the pixel intensity associated with another portion of the pixel block are compared to determine the gradient and variance of the pixel intensity associated with each portion of the pixel block.
8. The method according to any one of claims 1 to 5, further comprising: Render a representation of the keyboard model to match the posture of the physical keyboard.
9. The method according to claim 1, wherein, The predetermined shape template is an L-shaped template or an X-shaped template.
10. One or more computer-readable non-transitory storage media, said one or more computer-readable non-transitory storage media comprising software, said software being operable to: Images are captured from a camera viewpoint, and the images depict the physical keyboard; Identify the physical keyboard in the captured image; One or more shape features of the physical keyboard depicted in the image are detected by comparing the pixels of the image with a predetermined shape template, wherein the predetermined shape template represents the visual features of the space between the keyboard keys. Each of the one or more shape features of the physical keyboard corresponds to a T-feature, X-feature, or L-feature formed by the space between the keyboard keys on the physical keyboard; A keyboard model associated with the physical keyboard is determined from the database by comparing the detected one or more shape features of the physical keyboard with predetermined shape features of the plurality of keyboard models in a database that includes a plurality of keyboard models. Access a predetermined shape feature of the keyboard model associated with the physical keyboard; as well as The orientation of the physical keyboard is determined based on a comparison between the following: one or more shape features of the physical keyboard detected; The projection of the predetermined shape features of the keyboard model toward the camera viewpoint.
11. One or more computer-readable non-transitory storage media according to claim 10, wherein, The predetermined shape template is a T-shaped template with a main body and a top cover, wherein the main body represents the visual feature of the space between two keyboard keys in the first keyboard row.
12. One or more computer-readable non-transitory storage media according to claim 11, wherein, The top cover portion of the T-shaped template represents the visual characteristics of the space between the following: the portion of the first keyboard row corresponding to the two keyboard keys in the first keyboard row; and one or more keyboard keys in the second keyboard row adjacent to the first keyboard row.
13. One or more computer-readable non-transitory storage media according to any one of claims 10 to 12, wherein, Comparing the pixels of the image with the predetermined shape template includes: The image is divided into multiple pixel blocks, and each pixel block includes multiple pixels; For each of the plurality of pixel blocks: The visual features of the pixel block are determined based on the pixel intensities associated with the plurality of pixels in the pixel block; and The visual features of the pixel block are compared with a predetermined shape template representing the visual features of the space between the keyboard keys; and The one or more shape features of the physical keyboard depicted in the image are detected by comparing the visual features of the pixel blocks of the image with the predetermined shape template representing the visual features of the space between the keyboard keys.
14. One or more computer-readable non-transitory storage media according to claim 13, wherein, Determining the visual features of the pixel block based on the pixel intensities associated with the plurality of pixels in the pixel block includes: The pixel intensity associated with at least a portion of the pixel block and the pixel intensity associated with another portion of the pixel block are compared to determine the gradient and variance of the pixel intensity associated with each portion of the pixel block.
15. A system for determining keyboard posture, comprising: One or more processors; and one or more computer-readable non-transitory storage media communicating with the one or more processors, the one or more computer-readable non-transitory storage media including instructions that, when executed by the one or more processors, cause the system to perform: Images are captured from a camera viewpoint, and the images depict the physical keyboard; Identify the physical keyboard in the captured image; One or more shape features of the physical keyboard depicted in the image are detected by comparing the pixels of the image with a predetermined shape template, the predetermined shape template representing the visual features of the space between the keyboard keys, wherein each of the one or more shape features of the physical keyboard corresponds to a T-feature, an X-feature, or an L-feature formed by the space between the keyboard keys on the physical keyboard. A keyboard model associated with the physical keyboard is determined from the database by comparing the detected one or more shape features of the physical keyboard with predetermined shape features of the plurality of keyboard models in a database that includes a plurality of keyboard models. Access a predetermined shape feature of the keyboard model associated with the physical keyboard; and The orientation of the physical keyboard is determined based on a comparison between: one or more shape features of the detected physical keyboard; and the projection of the predetermined shape features of the keyboard model toward the camera viewpoint.
16. The system according to claim 15, wherein, The predetermined shape template is a T-shaped template with a main body and a top cover, wherein the main body represents the visual feature of the space between two keyboard keys in the first keyboard row.
17. The system according to claim 16, wherein, The top cover portion of the T-shaped template represents the visual characteristics of the space between the following: the portion of the first keyboard row corresponding to the two keyboard keys in the first keyboard row; and one or more keyboard keys in the second keyboard row adjacent to the first keyboard row.
18. The system according to any one of claims 15 to 17, wherein, Comparing the pixels of the image with the predetermined shape template includes: The image is divided into multiple pixel blocks, and each pixel block includes multiple pixels; For each of the plurality of pixel blocks: The visual features of the pixel block are determined based on the pixel intensities associated with the plurality of pixels in the pixel block; and The visual features of the pixel block are compared with a predetermined shape template representing the visual features of the space between the keyboard keys; and The one or more shape features of the physical keyboard depicted in the image are detected by comparing the visual features of the pixel blocks of the image with the predetermined shape template representing the visual features of the space between the keyboard keys.
19. The system according to claim 18, wherein, Determining the visual features of the pixel block based on the pixel intensities associated with the plurality of pixels in the pixel block includes: The pixel intensity associated with at least a portion of the pixel block and the pixel intensity associated with another portion of the pixel block are compared to determine the gradient and variance of the pixel intensity associated with each portion of the pixel block.
Citation Information
Patent Citations
Display method and device based on physical keyboard, terminal equipment and storage medium
CN110442245A
Keyboards for virtual, augmented, and mixed reality display systems
US20180350150A1